Underwater passive acoustic sea surface wind speed inversion method based on convolutional neural network

CN122817848APending Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611301879.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该阈值的设定缺乏客观依据,阈值设置过高则有效样本被大量剔除,阈值设置过低则船舶辐射噪声、海洋生物活动声、降雨击溅声等干扰成分难以被有效排除,而此类干扰源的时频特征与风成噪声之间存在一定程度的重叠,传统阈值筛选方式难以准确区分

Benefits of technology

[0028](1)本发明从水下声学信号中提取50Hz至20kHz范围内的1/3倍频程声压级,并按采集时间顺序排列为二维时频声压级矩阵,将二维时频声压级矩阵作为卷积神经网络的输入。相比于现有技术仅利用单个或少数几个离散频点构建经验模型的方案,二维时频声压级矩阵完整保留了风成噪声主导频段内相邻频带之间的频率关联性和相邻时间帧之间的时间连续性。卷积神经网络通过多层卷积核的局部感受野扫描,能够自动学习不同频带之间的能量分布模式和不同时间帧之间的演变规律,从大量数据中自主学习到完整的时频拓扑结构,解决了传统模型输入特征割裂、表征能力不足的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817848A_ABST
    Figure CN122817848A_ABST
Patent Text Reader

Abstract

The application discloses a kind of underwater passive acoustic sea surface wind speed inversion methods based on convolutional neural network, belong to sea surface wind speed inversion technical field.The application is by pre-processing to underwater acoustic signal, extracts the sound pressure level of each time frame in wind noise dominant frequency band, obtains two-dimensional time-frequency sound pressure level matrix, after tensorization and global normalization, input to convolutional neural network and carry out multilayer convolution operation, output three-dimensional feature map, the three-dimensional feature map is handled after global average pooling layer and outputs global feature vector;Global feature vector is mapped into normalized wind speed prediction value by fully connected regression layer, and the sea surface wind speed inversion value is obtained after inverse normalization processing.Convolutional neural network and the training process of fully connected regression layer use Huber loss function as optimization target.The application extracts time-frequency structure features by convolutional neural network, reduces interference by global average pooling, enhances training robustness by Huber loss, and improves the precision of sea surface wind speed inversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sea surface wind speed inversion technology, specifically to an underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks. Background Technology

[0002] Sea surface wind speed is one of the core parameters driving momentum, heat, and mass exchange at the air-sea interface. Accurately acquiring large-scale, long-term sea surface wind speed data is a crucial prerequisite for improving the accuracy of ocean numerical forecasting and understanding ocean dynamic processes. Traditional methods for measuring sea surface wind speed are mainly divided into two categories: contact and non-contact. Contact measurement methods include buoy-based anemometers and shipborne meteorological equipment. These methods can provide point wind speed data with high temporal resolution and relatively reliable measurement accuracy. However, due to the special nature of the marine environment, the deployment and maintenance costs of contact equipment are high, the spatial coverage is extremely limited, and the equipment is easily damaged under extreme sea conditions such as typhoons and cyclones, leading to the loss of critical data. Non-contact measurement methods are represented by satellite scatterometers, synthetic aperture radar, and spaceborne GNSS-R (Global Navigation Satellite System-Reflectometry) remote sensing. While these technologies can achieve global-scale sea surface wind field observation, their temporal resolution is limited by the satellite orbital period, making it difficult to meet the continuous monitoring needs of rapidly evolving sea conditions.

[0003] Compared to traditional methods, underwater passive acoustic inversion technology has advantages such as strong observation concealment, strong continuous working capability, and low operation and maintenance costs, and has become one of the research hotspots in the field of marine acoustic observation. The physical basis of underwater passive acoustic inversion technology lies in the fact that the sea surface wind field drives wave breaking and induces bubble generation. The bubbles radiate energy in the form of sound waves through oscillation and resonance processes, forming characteristic wind-generated noise in the marine environmental noise. The spectral level of this noise shows a significant statistical correlation with the sea surface wind speed.

[0004] Regarding underwater passive acoustic wind speed inversion, existing technologies have proposed various empirical models, including those using linear functions, logarithmic functions, and low- or high-order polynomials to establish the functional relationship between noise spectral levels at specific frequencies and sea surface wind speed. However, these existing empirical models all suffer from two technical limitations in practical applications: First, the model parameters are highly fixed to limited data segments from specific experimental sea areas, resulting in insufficient extrapolation and scene transfer capabilities. Second, the model inputs are based on sound pressure level values ​​at single or multiple discrete frequency points, neglecting the structural and statistical correlations between adjacent frequency bands and adjacent time frames within the dominant wind noise frequency band. This fragmented input strategy restricts the model's representational capabilities.

[0005] Furthermore, existing solutions rely on preset spectral morphology thresholds to screen input data for quality, in order to remove samples affected by non-wind-generated noise. The setting of this threshold lacks objective basis; if the threshold is set too high, a large number of valid samples are removed; if the threshold is set too low, interference components such as ship-radiated noise, marine life activity sounds, and rain splashing sounds are difficult to effectively eliminate. Moreover, the time-frequency characteristics of these interference sources overlap to some extent with wind-generated noise, making it difficult for traditional threshold screening methods to accurately distinguish them.

[0006] Therefore, how to make full use of the time-frequency structure information in the dominant frequency band of wind noise to improve the accuracy of sea surface wind speed inversion is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides an underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks. By constructing a two-dimensional time-frequency sound pressure level matrix covering the dominant frequency band of wind-generated noise as the model input, the convolutional neural network automatically extracts the time-frequency structure features between adjacent frequency bands and adjacent time frames, and introduces a global average pooling layer and Huber loss function to achieve robust training, thereby improving the accuracy of sea surface wind speed inversion.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] This invention proposes an underwater passive acoustic wind speed inversion method based on convolutional neural networks, comprising the following steps:

[0010] S1. Preprocess the collected underwater acoustic signals, extract the sound pressure level of each time frame in the wind noise-dominant frequency band from the preprocessed acoustic signals, and arrange the extracted sound pressure levels into a two-dimensional time-frequency sound pressure level matrix according to the acquisition time order.

[0011] S2. Tensor quantization is performed on the two-dimensional time-frequency sound pressure level matrix to obtain a three-dimensional tensor. After global normalization of the three-dimensional tensor, a normalized tensor is obtained. The normalized tensor is then input into the trained convolutional neural network.

[0012] S3. The convolutional neural network performs multiple convolution operations on the input normalized tensor in sequence to output a three-dimensional feature map; the three-dimensional feature map is processed by a global average pooling layer to output a global feature vector.

[0013] S4. The global feature vector is mapped to a standardized wind speed prediction value through a trained fully connected regression layer. The standardized wind speed prediction value is then de-standardized to obtain the sea surface wind speed inversion value.

[0014] Furthermore, S1 specifically includes:

[0015] S101. Divide the acquired acoustic signal into non-overlapping frames according to the preset frame length, and treat each frame as a time frame.

[0016] S102. Apply a Hanning window to the acoustic signal of each time frame and then perform a fast Fourier transform to obtain the power spectral density of each time frame.

[0017] S103. For the power spectral density of each time frame, extract the sound pressure level at the center frequency of its 1 / 3 octave band in the range of 50 Hz to 20 kHz to form a 1 / 3 octave band sound pressure level matrix.

[0018] S104. Arrange the 1 / 3 octave band sound pressure level matrix of each time frame in the order of acquisition time to construct the two-dimensional time-frequency sound pressure level matrix.

[0019] Furthermore, during the training of the convolutional neural network and the fully connected regression layer, the samples in the training dataset are stratified and sampled according to the magnitude of the actual wind speed value to obtain the training set and the test set, so that the number of training samples and test samples in each wind speed interval is equal. During the training process, the Huber loss function is used as the optimization objective to optimize the parameters of the convolutional neural network and the fully connected regression layer end-to-end.

[0020] Furthermore, in S2, the three-dimensional tensor is subjected to global z-score normalization.

[0021] Furthermore, in S3, the convolutional neural network includes multiple convolutional blocks connected in series, adjacent convolutional blocks are connected through a max pooling layer, and the last convolutional block is connected to the global average pooling layer.

[0022] Each convolutional block consists of a convolutional layer, a batch normalization layer, and a non-linear activation layer connected in sequence.

[0023] The normalized tensor is sequentially passed through concatenated convolutional blocks to obtain the three-dimensional feature map. The global average pooling layer calculates the mean of the three-dimensional feature map along the spatial dimension and outputs the global feature vector.

[0024] Furthermore, the number of convolutional blocks is three, and the number of output channels of the three convolutional blocks is 32, 64 and 128 respectively. The kernel size of the convolutional layer in each convolutional block is 3×3.

[0025] Furthermore, in S4, the fully connected regression layer includes a first fully connected layer, a Dropout layer, and a second fully connected layer connected in sequence;

[0026] The global feature vector is sequentially mapped to a standardized wind speed prediction value through a first fully connected layer, a Dropout layer, and a second fully connected layer.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] (1) This invention extracts the 1 / 3 octave band sound pressure level from underwater acoustic signals in the range of 50Hz to 20kHz and arranges them into a two-dimensional time-frequency sound pressure level matrix according to the acquisition time sequence. The two-dimensional time-frequency sound pressure level matrix is ​​used as the input of a convolutional neural network. Compared with the existing technology that only uses a single or a few discrete frequency points to build an empirical model, the two-dimensional time-frequency sound pressure level matrix completely preserves the frequency correlation between adjacent frequency bands in the wind noise-dominant frequency band and the temporal continuity between adjacent time frames. The convolutional neural network can automatically learn the energy distribution patterns between different frequency bands and the evolution law between different time frames through the local receptive field scanning of multiple convolutional kernels. It can autonomously learn the complete time-frequency topology from a large amount of data, solving the problems of fragmented input features and insufficient representation ability of traditional models.

[0029] (2) This invention introduces a global average pooling layer after the last convolutional block of the convolutional neural network to calculate the mean of the three-dimensional feature map along the spatial dimension and output a global feature vector. In marine environmental noise, non-wind-induced interference sources such as ship radiation noise and rain splash noise usually exhibit local transient characteristics in the time-frequency domain, that is, they generate energy anomalies within a limited frequency band and time range; while the energy distribution of wind-induced noise covers a wide frequency range within the dominant frequency band and changes slowly over time. The operation of global average pooling to calculate the mean of the feature map along the spatial dimension naturally dilutes local outliers during the mean calculation process, preventing them from having a dominant influence on the final feature output, thus making up for the shortcomings of traditional fixed threshold screening methods that rely on subjective experience and are difficult to accurately distinguish interference sources.

[0030] (3) Marine environmental noise data inevitably contains various non-wind-induced interference components. If conventional mean squared error loss is used during the training process for these interference samples, they will generate gradient values ​​much larger than those of normal samples, dominating the direction of parameter updates and disrupting the learning process of the wind speed mapping manifold. In the training phase, this invention uses the Huber loss function as the optimization objective for network training. When the prediction error is small, the Huber loss uses the form of L2 loss to ensure the fine-tuning accuracy of the network at the end of convergence. When encountering severe interference that causes the prediction error to exceed the threshold, it automatically degenerates into L1 loss to truncate the excessively large gradients generated by abnormal samples. At the same time, global average pooling dilutes the contribution of local abnormal responses to the feature vector during forward propagation, and the Huber loss limits the destructive impact of interference samples on parameter updates during backpropagation. Together, they improve the training stability of the model on noisy data.

[0031] (4) In this invention, a stratified sampling strategy is adopted when partitioning the dataset. All samples in the training dataset are divided into multiple wind speed intervals according to the magnitude of the true wind speed value. Within each wind speed interval, samples are drawn in the same proportion to be assigned to the training set and the test set, ensuring that the training set and the test set have the same statistical characteristics of wind speed distribution. This partitioning strategy ensures that the training set and the test set meet the independent and identically distributed condition. The inversion accuracy on the test set can truly reflect the model's learning ability across the entire wind speed range, avoiding evaluation bias caused by uneven wind speed distribution due to random partitioning. Attached Figure Description

[0032] Figure 1 This is a flowchart of the underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to the present invention;

[0033] Figure 2 This is a schematic diagram of the structure of the convolutional neural network of the present invention;

[0034] Figure 3 This is a schematic diagram of the fully connected regression layer of the present invention;

[0035] Figure 4 The example model and the comparison model are shown in the time series diagram comparing the wind speed inversion values ​​with the actual wind speed values ​​on the test set. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Example

[0038] refer to Figure 1 This embodiment provides an underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks, including the following steps:

[0039] S1. Preprocess the acquired underwater acoustic signals, extract the sound pressure level of each time frame in the wind-dominated noise frequency band from the preprocessed acoustic signals, and arrange the extracted sound pressure levels in the order of acquisition time into a two-dimensional time-frequency sound pressure level matrix. Specifically, this includes the following sub-steps:

[0040] S101, The acquired raw time-series acoustic signal As a signal to be processed This refers to the overall sampling point sequence number. , This represents the total number of sampling points. Based on the preset frame length. ( (Number of sampling points) to process the signal Cut into Each non-overlapping frame is divided into two time frames; when the total number of sampling points of the signal to be processed is... Not the preset frame length When the length is an integer multiple of the preset frame length, the last frame is padded with zeros to ensure that the length of each frame is consistent.

[0041] No. The acoustic signal of each time frame is represented as:

[0042]

[0043] In the formula, The intra-frame sampling point number, For the first Acoustic signals of time frames.

[0044] S102. Apply a Hanning window to the acoustic signal of each time frame to obtain the windowed signal:

[0045]

[0046]

[0047] For the first Windowed signal for each time frame This is the Hanning window function.

[0048] Perform a Fast Fourier Transform on the windowed signal of each frame to obtain the power spectral density of each time frame;

[0049]

[0050] In the formula, For frequency point number, For the first Power spectral density of each time frame, It is the imaginary unit.

[0051] S103. For the power spectral density of each time frame, according to the ANSI (American National Standards Institute) standard, extract the sound pressure level at the center frequency of 27 octave bands in the range of 50Hz to 20kHz to form the octave band sound pressure level matrix of that time frame.

[0052] The sound pressure level at the center frequency of each 1 / 3 octave band is calculated using the following formula:

[0053]

[0054] In the formula, The index of the center frequency of 1 / 3 octave band. , The total number of center frequencies in a 1 / 3 octave band, i.e. =27, For the first The time frame in the ... The sound pressure level at the center frequency of one-third octave band. and The first The frequency point numbers corresponding to the lower and upper limits of a 1 / 3 octave band. For reference sound pressure, in underwater acoustics, Take 1µPa as the unit of sound pressure level, dBre1µPa. For frequency resolution, , The sampling rate.

[0055] S104. Arrange the 1 / 3 octave band sound pressure level matrices of each time frame in chronological order of acquisition time to construct the two-dimensional time-frequency sound pressure level matrix. .

[0056]

[0057]

[0058] In the formula, For the first The 1 / 3 octave band sound pressure level matrix for each time frame The dimension is , This is the matrix transpose operation.

[0059] S2, Regarding the two-dimensional time-frequency sound pressure level matrix Tensor quantization is performed to obtain a shape of 1× × 3D tensor The first dimension is the channel dimension; the second dimension is the frequency dimension, corresponding to... Each is a 1 / 3 octave band, and the third dimension is the time dimension, corresponding to... Each time frame.

[0060] For three-dimensional tensors Perform global z-score normalization to obtain the normalized tensor. :

[0061]

[0062] In the formula, for The average of all sound pressure levels in the system. for The standard deviation of all sound pressure levels As a numerical stabilizing term used to avoid division by zero errors, in this embodiment, =1×10 -8 .

[0063] Normalized tensor Input into the trained convolutional neural network.

[0064] S3. The convolutional neural network performs multiple convolution operations on the input normalized tensor sequentially to output a three-dimensional feature map. The three-dimensional feature map As input to the global average pooling layer, the global average pooling layer processes the 3D feature map. Calculate the mean along the spatial dimension and output the global feature vector. .

[0065] refer to Figure 2 The convolutional neural network contains multiple convolutional blocks connected in series. Adjacent convolutional blocks are connected through a max pooling layer, and the last convolutional block is connected to a global average pooling layer.

[0066] Each convolutional block consists of a convolutional layer, a batch normalization layer, and a nonlinear activation layer connected sequentially; its propagation process is represented as follows:

[0067]

[0068] For the first The 3D feature map output by each convolutional block , The total number of convolutional blocks. For the first The convolution kernel of the convolutional layer in each convolutional block This represents a two-dimensional convolution operation. For the first The bias term of each convolutional block, For batch normalization function, For non-linear activation functions, when hour, The input is a normalized tensor.

[0069] The normalized tensor is sequentially passed through a series of convolutional blocks, with the last convolutional block outputting a 3D feature map. As input to the global average pooling layer, i.e., the 3D feature map .

[0070] The global average pooling operation can be represented as:

[0071]

[0072] In the formula, For the first The global average pooling output value of each channel. 3D feature map output by a convolutional neural network In the middle, the first The first channel, the first line, number The element values ​​of the column, and They are respectively Height and width.

[0073] The mean operation of global average pooling can weaken the weight of local outliers, preventing them from having a dominant influence on the final feature output, thereby achieving mean suppression of local transient interferences such as ship radiated noise and rain splash at the network structure level.

[0074] In this embodiment, the number of convolutional blocks in the convolutional neural network is 3, the kernel size of each convolutional layer is 3×3, the stride is 1, and the padding is 1. The first and second convolutional blocks are connected by a 2×2 max pooling layer, and the second and third convolutional blocks are connected by a max pooling layer, which is used to reduce the spatial dimension of the feature map. After the third convolutional block, no pooling operation is performed, and it is directly connected to a global average pooling layer.

[0075] The number of output channels for the three convolutional blocks are 32, 64, and 128, respectively. Specifically:

[0076] The first convolutional block has 1 input channel and 32 output channels.

[0077] The second convolutional block has 32 input channels and 64 output channels.

[0078] The third convolutional block has 64 input channels and 128 output channels.

[0079] S4. The global feature vector obtained in S3... The trained fully connected regression layer is input, and the global feature vector is mapped to a standardized wind speed prediction value through the fully connected regression layer. The standardized wind speed prediction value is then de-standardized to obtain the sea surface wind speed inversion value.

[0080] refer to Figure 3 The fully connected regression layer includes a first fully connected layer, a Dropout layer, and a second fully connected layer connected in sequence.

[0081] The first fully connected layer processes the global feature vector. A linear transformation is performed, and the result is then subjected to a nonlinear transformation using the ReLU activation function to obtain the first activated feature vector. :

[0082]

[0083] In the formula, This is the weight matrix of the first fully connected layer. This is the bias term for the first fully connected layer.

[0084] First activation feature vector After processing by the Dropout layer, the Dropout feature vector is obtained. During the application phase, the Dropout layer maintains the activation of all neurons, and its output equals its input. Input to the second fully connected layer, the second fully connected layer to Perform a linear transformation to map it to a 1-dimensional output, which is the standardized wind speed prediction value. :

[0085]

[0086] In the formula, This is the weight matrix of the second fully connected layer. This is the bias term for the second fully connected layer.

[0087] In this embodiment, the first fully connected layer maps the 128-dimensional global feature vector to a 64-dimensional feature vector, and the second fully connected layer maps the 64-dimensional feature vector to a 1-dimensional output.

[0088] For the standardized wind speed prediction value After inverse standardization, the sea surface wind speed inversion value is obtained. :

[0089] .

[0090] In the formula, , To train the mean and standard deviation of the true wind speed values.

[0091] In this embodiment, when training the convolutional neural network and the fully connected regression layer, all samples in the training dataset are stratified according to the magnitude of the true wind speed value, so that the proportion of training samples and test samples in each wind speed interval is the same, thereby dividing the training dataset into training sets and test sets with the same wind speed distribution statistical characteristics.

[0092] During training, the Huber loss function is used as the optimization objective to perform end-to-end optimization of the parameters of the convolutional neural network and the fully connected regression layer. For all samples within a batch, the loss is defined as the mean of the Huber losses for each sample:

[0093]

[0094]

[0095]

[0096] In the formula, This represents the average Huber loss within the batch. The number of samples within the batch. The sample number. For the first The prediction error for each sample. For the first Standardized wind speed prediction values ​​for each sample. For the first Standardized true wind speed values ​​for each sample. For the first Huber loss value for each sample. For the threshold hyperparameter, in this embodiment, we take... =1.0.

[0097] The AdamW optimizer was used for parameter optimization, with an initial learning rate of 1×10⁻⁶. -3 The weight decay factor is set to 1×10. -4 The batch size is set to 256. During the training phase, the Dropout layer randomly sets the neuron output to zero with a preset probability to suppress co-adaptation between neurons and improve the model's generalization ability. In this embodiment, the preset probability is 0.2, and the number of training rounds is set to 100. After training is completed, the model parameters are saved, resulting in the trained convolutional neural network and fully connected regression layer.

[0098] To better illustrate the beneficial effects of the present invention, a training dataset was obtained by long-term observation of continuous underwater passive acoustic signals in a certain sea area.

[0099] The data acquisition process was as follows: Underwater acoustic signals were recorded using a self-contained data acquisition device deployed 1600 meters below the sea surface. The device sensitivity was calibrated to -196 dBre 1µPa, the gain was set to 20 times the linear gain, and the sampling rate was 128 kHz. The data acquisition strategy involved acquiring 150 seconds of continuous acoustic signals every 20 minutes. The wind field component at 10 meters above the sea surface was obtained from the ERA5 reanalysis data for the corresponding time period and synthesized into a scalar wind speed. This synthesized scalar wind speed was then used as the true wind speed value.

[0100] For each 150-second continuous acoustic signal acquired, it is first divided into 5 segments of continuous acoustic signal, each segment lasting 30 seconds. Then, the process S1 in the above embodiment is performed on each segment of acoustic signal, dividing each segment of continuous acoustic signal into 6 non-overlapping time frames with a frame length of 5 seconds, i.e., frame length. =640,000 sampling points were used to obtain a two-dimensional time-frequency sound pressure level matrix corresponding to each continuous acoustic signal segment. The synthesized real wind speed values ​​were linearly interpolated so that each two-dimensional time-frequency sound pressure level matrix corresponds to a real wind speed value. Each two-dimensional time-frequency sound pressure level matrix and its corresponding real wind speed value together constitute a sample, resulting in a training dataset containing 100,146 samples.

[0101] The training dataset is divided into training and test sets using stratified sampling. Specifically, the samples are sorted by their actual wind speed values ​​from smallest to largest. Based on the wind speed range, all samples in the training dataset are divided into multiple wind speed intervals. Within each wind speed interval, samples are randomly selected in the same proportion and assigned to both the training and test sets, ensuring that the proportion of training and test samples is equal within each wind speed interval. This results in the training and test sets having the same statistical characteristics of wind speed distribution. The sample ratio between the training and test sets is 8:2.

[0102] The example model (i.e., a convolutional neural network and a fully connected regression layer) was trained using a training set. The Huber loss function was used as the optimization objective during training. After training, the model parameters were saved to obtain the trained example model. Simultaneously, a comparison model was trained using the same training set. The comparison model was an 8kHz single-frequency empirical model. This model extracted the sound pressure level of each acoustic signal at the 8kHz center frequency and constructed a wind speed inversion model by fitting a functional relationship between the sound pressure level and the true wind speed value. The fitting function was in logarithmic form.

[0103] The performance of the example model and the comparison model was validated on the test set. The performance evaluation metrics included: root mean square error (RMSE), mean absolute error (MAE), coefficient of determination (R²), and bias. The formulas for each metric are as follows:

[0104]

[0105]

[0106]

[0107]

[0108] In the formula, The number of samples in the test set. For the test set Wind speed inversion values ​​for each sample, For the test set The true wind speed value for each sample. The average of the actual wind speed values. The expected value of the wind speed inversion value. Expected value of the actual wind speed.

[0109] The two-dimensional time-frequency sound pressure level matrix in the test set was tensorized and globally normalized before being input into the implementation model. After forward propagation, standardized wind speed prediction values ​​were output, followed by inverse normalization to obtain sea surface wind speed inversion values. Simultaneously, the sound pressure level at the 8kHz center frequency of each sample in the test set was extracted and input into the comparison model to obtain sea surface wind speed inversion values. The evaluation metrics for each model are shown in Table 1.

[0110] Table 1 Evaluation metrics for each model

[0111]

[0112] As shown in Table 1, the embodiment model outperforms the comparison model in all four evaluation metrics. The RMSE and MAE of the embodiment model are both lower than those of the comparison model, indicating that the embodiment model has smaller inversion errors and higher prediction accuracy. The R² of the embodiment model is positive and significantly higher than that of the comparison model, indicating that the embodiment model has a better fitting ability to wind speed change trends. The absolute value of the bias of the embodiment model is smaller than that of the comparison model and closer to zero, indicating that the embodiment model does not have obvious systematic overestimation or underestimation. The comparison of the four metrics verifies the effectiveness and superiority of the method proposed in this invention.

[0113] Figure 4 The diagram presents a time-series comparison of the wind speed inversion results and actual wind speed values ​​on the test set for both the example model and the comparison model. To more clearly demonstrate the degree of agreement between the inversion curves and the actual wind speed trends, the wind speed values ​​for each hour have been averaged in the diagram. Figure 4 It can be seen that the inversion curve of the example model is in good agreement with the actual wind speed curve, and can accurately follow the evolution trend of wind speed; while the inversion curve of the comparison model is basically consistent with the actual value in the overall trend, it deviates significantly in many periods of rapid wind speed change, and the inversion error is large.

[0114] The specific embodiments of the present invention enable those skilled in the art to understand or implement the invention. Various modifications to the above embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention.

[0115] It should be understood that the present invention is not limited to the content already described above, and various modifications and changes can be made without departing from its scope. The scope of the present invention is limited only by the appended claims.

Claims

1. A method for underwater passive acoustic sea surface wind speed inversion based on convolutional neural networks, characterized in that, Includes the following steps: S1. Preprocess the collected underwater acoustic signals, extract the sound pressure level of each time frame in the wind noise-dominant frequency band from the preprocessed acoustic signals, and arrange the extracted sound pressure levels into a two-dimensional time-frequency sound pressure level matrix according to the acquisition time order. S2. Tensor quantization is performed on the two-dimensional time-frequency sound pressure level matrix to obtain a three-dimensional tensor. After global normalization of the three-dimensional tensor, a normalized tensor is obtained. The normalized tensor is then input into the trained convolutional neural network. S3. The convolutional neural network performs multiple convolution operations on the input normalized tensor in sequence to output a three-dimensional feature map; the three-dimensional feature map is processed by a global average pooling layer to output a global feature vector. S4. The global feature vector is mapped to a standardized wind speed prediction value through a trained fully connected regression layer. The standardized wind speed prediction value is then de-standardized to obtain the sea surface wind speed inversion value.

2. The underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to claim 1, characterized in that, S1 specifically includes: S101. Divide the acquired acoustic signal into non-overlapping frames according to the preset frame length, and treat each frame as a time frame. S102. Apply a Hanning window to the acoustic signal of each time frame and then perform a fast Fourier transform to obtain the power spectral density of each time frame. S103. For the power spectral density of each time frame, extract the sound pressure level at the center frequency of its 1 / 3 octave band in the range of 50 Hz to 20 kHz to form a 1 / 3 octave band sound pressure level matrix. S104. Arrange the 1 / 3 octave band sound pressure level matrix of each time frame in the order of acquisition time to construct the two-dimensional time-frequency sound pressure level matrix.

3. The underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to claim 1, characterized in that, During the training of the convolutional neural network and the fully connected regression layer, the samples in the training dataset are stratified according to the magnitude of the actual wind speed value to obtain the training set and the test set, so that the number of training samples and test samples in each wind speed range is equal. During the training process, the Huber loss function is used as the optimization objective to optimize the parameters of the convolutional neural network and the fully connected regression layer end-to-end.

4. The underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to claim 1, characterized in that, In S2, the three-dimensional tensor is subjected to global z-score normalization.

5. The underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to claim 1, characterized in that, In S3, the convolutional neural network contains multiple convolutional blocks connected in series, adjacent convolutional blocks are connected through a max pooling layer, and the last convolutional block is connected to the global average pooling layer. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a non-linear activation layer connected in sequence. The normalized tensor is sequentially passed through concatenated convolutional blocks to obtain the three-dimensional feature map. The global average pooling layer calculates the mean of the three-dimensional feature map along the spatial dimension and outputs the global feature vector.

6. The underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to claim 5, characterized in that, The number of convolutional blocks is three, and the number of output channels of the three convolutional blocks is 32, 64 and 128 respectively. The kernel size of the convolutional layer in each convolutional block is 3×3.

7. The underwater passive acoustic sea surface wind speed inversion method based on convolutional neural networks according to claim 1, characterized in that, In S4, the fully connected regression layer includes a first fully connected layer, a Dropout layer, and a second fully connected layer connected in sequence; The global feature vector is sequentially mapped to a standardized wind speed prediction value through a first fully connected layer, a Dropout layer, and a second fully connected layer.