Ship self-noise generation method based on neural network

Through the noise autoencoder and flow matching generation module based on neural network, the accuracy problem of traditional methods in generating ship noise signals in different navigation environments is solved, and high-precision ship noise signal generation is achieved, adapting to changing environments and providing design optimization basis.

CN120296411APending Publication Date: 2025-07-11INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302835.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing ship noise signal generation methods are difficult to generate high-precision noise signals under different navigation environments and operating conditions. The traditional methods rely on complex physical models and computing resources, and the neural network-based methods are not yet mature in water acoustic signal generation.

Method used

The noise autoencoder and stream matching generation module are adopted based on neural networks. The time resolution is reduced by pre-training the noise autoencoder, high-level features are extracted, meaningless noise signals are spliced, and dynamic changes are captured by the stream matching generation module to generate ship noise signals that meet the boat model and working conditions.

Benefits of technology

The generated ship noise signal conforms to the time-varying characteristics and working conditions information of real noise, improves the model's adaptability in a variety of environments, and provides a basis for noise control and design optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296411A_ABST
    Figure CN120296411A_ABST
Patent Text Reader

Abstract

The invention provides a ship self-noise generation method based on a neural network. The method comprises the following steps: reducing the time resolution of an original ship noise signal through a pre-trained noise auto-encoder, and extracting high-level features from the original ship noise signal to obtain a first hidden layer representation sequence; splicing a section of meaningless noise signal behind the first hidden layer representation sequence; the noise representation generation module based on pre-training and stream matching predicts a second hidden layer representation sequence according to the context of the first hidden layer representation sequence, and replaces the meaningless noise signal with the second hidden layer representation sequence to obtain a third hidden layer representation sequence; and improving the time resolution of the third implicit strata representation sequence through a pre-trained noise auto-encoder, and mapping the implicit strata representation sequence back to the original signal space to generate a new ship noise signal. The noise generated by the method provided by the invention not only can accord with information such as ship models and working conditions of reference noise, but also can reflect the time-varying characteristics of the noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater acoustic signal generation, and particularly to a ship self-noise generation method based on a neural network. Background Art

[0002] The ship self-noise signal is the data basis for research fields such as ship noise recognition, fault diagnosis, and ship design optimization. However, due to the time-frequency characteristics of the ship noise signal varying under different navigation environments and working conditions, traditional noise acquisition methods are limited by conditions such as equipment, environment, and time, making it difficult to cover all possible noise patterns, and the data acquisition volume under different working conditions is also very limited. Therefore, how to generate more ship noise signals through a ship self-noise generation method is a very important research direction. Through noise signal generation, noise signals under different working conditions can be simulated according to variables such as ship speed, load, and sea state, providing more diverse training data for the recognition model, thereby enhancing the adaptability of the recognition model to a changing environment. At the same time, the analysis of a large amount of generated data can also provide a basis for ship noise control and design optimization.

[0003] The ship noise signal belongs to a kind of underwater acoustic signal. And acoustic modeling is the basis of traditional underwater acoustic signal generation, which involves modeling the acoustic characteristics of the underwater sound field and targets, including the influence of the geometric shape, material properties, sound source characteristics, and underwater acoustic propagation environment (such as temperature, salinity, flow velocity distribution, etc.) of underwater targets. At present, researchers have proposed various models to describe these complex acoustic phenomena. In terms of the underwater acoustic signal propagation model, the underwater sound field is usually modeled as a multi-path propagation system. These models consider the influence of factors such as underwater terrain, water depth, water temperature, and salinity on sound wave propagation. Some models are based on numerical methods, such as the finite element method, finite difference method, etc., to numerically simulate the sound wave propagation in order to better understand the propagation law of underwater acoustic signals in different environments. These traditional data generation methods rely on complex physical models and numerical calculations, which not only result in high computational complexity but also require a large amount of computational resources and time.

[0004] Existing ship self-noise signal generation methods also adopt the above-mentioned underwater acoustic signal processing methods as the mainstream technical route, which mainly includes the synthesis of sine signals, the addition of noise, the filtering of signals, etc., to simulate the line spectrum signals, modulation signals, etc. in ship self-noise. However, the line spectrum and modulation spectrum features of the ship noise simulated by such traditional methods have a high distinction from the real signals and seriously rely on the noise source parameters and acoustic field parameters. In practical applications, it is difficult to accurately obtain these parameters, which further limits the authenticity and practicality of the generation effect. For example, the specific characteristics of mechanical noise and propeller noise are affected by various factors, including the speed and heading of the ship, the ocean environment, etc., and it is difficult for traditional methods to maintain high accuracy in these dynamic changes. On the other hand, generative models based on neural networks have been able to generate very realistic samples in the signal generation of air sound fields such as speech and natural environment audio. However, the application of such neural network-based signal generation methods in the underwater acoustic field is still in a blank stage, and due to the differences in the physical characteristics between underwater acoustic signals and air sound signals, a neural network generation model specifically for underwater acoustic signals needs to be studied. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a neural network-based ship self-noise generation method, which includes:

[0006] Collect the first ship noise signals in multiple consecutive time periods; reduce the time resolution of the first ship noise signals through a pre-trained noise autoencoder and extract high-level features therefrom to obtain a hidden layer representation sequence; splice a meaningless noise signal after the hidden layer representation sequence; based on a pre-trained noise representation generation module based on flow matching, capture the dynamic changes of the hidden layer representation sequence and predict a second hidden layer representation sequence according to the context of the hidden layer representation sequence, and then replace the meaningless noise signal with the second hidden layer representation sequence to obtain a third hidden layer representation sequence; improve the time resolution of the third hidden layer representation sequence through a pre-trained noise autoencoder and map the hidden layer representation sequence back to the original signal space to generate the second ship noise signals.

[0007] In some embodiments, the noise autoencoder includes a noise encoder and a noise decoder. The noise encoder is composed of multiple one-dimensional CNNs stacked, and the noise decoder is composed of multiple transposed CNNs stacked.

[0008] In some embodiments, the pre-trained noise autoencoder is obtained by training the noise autoencoder, and its training process includes:

[0009] Input the ship noise signal samples into a noise encoder to obtain the hidden layer representation sequence samples; input the hidden layer representation sequence samples into a noise decoder to obtain the reconstructed ship noise signal samples; use the ship noise signal samples, the hidden layer representation sequence samples, and the reconstructed ship noise signal samples as training samples to train an autoencoder, and obtain a pre-trained noise autoencoder.

[0010] In some embodiments, during the process of training the noise autoencoder, the spectral loss between the ship noise samples and the reconstructed ship noise signal samples is also calculated, and the noise autoencoder is iteratively adjusted according to the spectral loss.

[0011] In some embodiments, the method for calculating the spectral loss includes:

[0012] Select different Fourier transform window lengths and number of frequency band channels, calculate multiple Mel spectrograms between the ship noise signal samples and the reconstructed ship noise signal samples; sum the losses of the multiple Mel spectrograms, and use the sum of the losses as the final loss.

[0013] In some embodiments, during the process of training the noise autoencoder, multi-subband analysis and multi-period discrimination are performed on the ship noise signal samples and the reconstructed ship noise signal samples.

[0014] In some embodiments, the construction process of the noise representation generation module based on flow matching includes:

[0015] Use the meaningless standard Gaussian noise for replacing the hidden layer representation as the starting distribution of the continuous normalizing flow, and set the hidden layer representation of the ship noise signal as the final distribution of the continuous normalizing flow; build a neural network composed of 24 layers of Transformers to simulate the vector field in the continuous normalizing flow.

[0016] In some embodiments, training the constructed noise representation generation module based on flow matching to obtain a pre-trained noise representation generation module based on flow matching, the process specifically includes: sampling a batch of data from the training dataset, generating a hidden layer representation sequence through the pre-trained noise autoencoder; sampling standard Gaussian noise from the standard Gaussian distribution, selecting a sequence in the hidden layer representation sequence, and using the standard Gaussian noise to replace the selected sequence to generate context features; calculating intermediate features according to the hidden layer representation sequence and the standard Gaussian noise, and completing through the following formula:

[0017] x t =(1 - t)ζ + tx1

[0018] where, ζ represents standard Gaussian noise; x1 represents the hidden layer representation sequence; t is randomly sampled between [0, 1]; the Transformer neural network is trained with the intermediate feature and the context feature as the input and x1 - ζ as the fitting target;

[0019] In some embodiments, based on a pre-trained flow-matching-based noise representation generation module, the dynamic changes of the hidden layer representation sequence are captured, and the second hidden layer representation sequence is predicted according to the context of the hidden layer representation sequence. Then, the method of replacing the meaningless noise signal with the second hidden layer representation sequence to obtain the third hidden layer representation sequence is completed by the following formula:

[0020]

[0021] where, x1′ represents the third hidden layer representation sequence; x t represents the intermediate feature; x ctx represents the context feature.

[0022] In some embodiments, after generating the second ship noise signal, the reduction degree of the second ship noise signal is evaluated. The evaluation method includes:

[0023] The second ship noise signal and the first ship noise signal are respectively sliced with a window of 3 - 7 s, and the Z-score and the interquartile range are applied to extract the high-energy frequency points below 500 - 700 HZ to characterize the transient low-frequency line spectrum features of the noise signal; the density-based noise application spatial clustering is used for the frequency points within 20 - 40 windows to remove outliers, so as to obtain the fluctuation distribution of the low-frequency line spectrum features in the long term; the Wasserstein distance between the low-frequency line spectrum fluctuations of the second ship noise signal and the first ship noise signal is used to evaluate the reduction degree of the generated signal.

[0024] The present invention encodes ship noise into a hidden layer representation sequence by training a ship noise autoencoder, trains a flow-matching-based generation model in the hidden space of the autoencoder, combines context learning to capture the dynamic changes of the hidden layer representation, generates a new hidden layer encoding according to the hidden layer encoding of the reference audio, and finally uses the pre-trained ship noise decoder to decode the generated hidden layer encoding into a time-domain noise signal. The noise generated by the method provided by the present invention can not only conform to the information such as the boat model and working conditions of the reference noise, but also reflect its time-varying characteristics. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 Schematic flowchart showing a method for generating ship self-noise based on a neural network provided by an embodiment of this specification;

[0027] Figure 2 Schematic flowchart showing a training process of a noise autoencoder provided by an embodiment of this specification;

[0028] Figure 3 Schematic diagram showing a calculation of multi-scale and multi-resolution spectral loss provided by an embodiment of this specification;

[0029] Figure 4 Schematic diagram of the low-frequency line spectrum of the ship noise samples reconstructed by the noise autoencoder before and after the multi-scale and multi-resolution spectral loss provided by an embodiment of this specification;

[0030] Figure 5 Schematic flowchart showing a training process of a noise characterization generation module based on flow matching provided by an embodiment of this specification;

[0031] Figure 6 Schematic framework diagram showing a method for generating ship self-noise based on a neural network provided by an embodiment of this specification;

[0032] Figure 7 Schematic comparison diagram showing a generated ship noise signal and an original ship noise signal provided by an embodiment of this specification;

[0033] Figure 8 Schematic comparison diagram of the transient low-frequency line spectrum features of the original ship noise signal and the generated ship noise signal at different times provided by an embodiment of this specification;

[0034] Figure 9 Schematic structural diagram showing a device for generating ship self-noise based on a neural network provided by an embodiment of this specification. Detailed implementation manners

[0035] The following describes the solutions provided in this specification in conjunction with the drawings.

[0036] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described below in conjunction with the drawings.

[0037] In the description of the embodiments of the present application, words such as "exemplary", "for example" or "for instance" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example" or "for instance" is intended to present related concepts in a specific manner.

[0038] In the description of the embodiments of the present application, the term "and / or" is merely an associative relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and both A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more.

[0039] Furthermore, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0040] Figure 1 The schematic flowchart of a ship self-noise generation method based on a neural network provided by the embodiments of this specification is shown, as Figure 1 shown, the method includes the following steps:

[0041] Step S101, collect the first ship noise signals in multiple consecutive time periods. Specifically, an underwater microphone or an in-ship sensor can be used to obtain the first ship noise signals in multiple consecutive time periods. Specifically, the first ship noise signal refers to the original ship noise signal without being processed, and this signal carries information such as the model and the ship working conditions in the consecutive time periods. Before converting the first ship noise signal into the first hidden layer representation sequence, the first ship noise signal can be preprocessed first. For example, the collected original ship signal can be converted into a standard format, a filter (such as a band-pass filter) or a noise reduction algorithm (such as wavelet transform) can be used to remove environmental noise to improve the signal quality, and the signal recorded for a long time can be segmented into short time periods (such as 5 seconds per segment) for easy analysis. Step S102, reduce the time resolution of the first ship noise signal through a pre-trained noise autoencoder and extract high-level features from it to obtain the first hidden layer representation sequence.

[0042] By using the encoder part of the trained noise autoencoder to map the first ship noise signal to a low-dimensional hidden layer space, reducing the time resolution and extracting high-level features, a first hidden layer representation sequence corresponding to the first ship noise signal is obtained.

[0043] Step S103, concatenate a segment of meaningless noise signal after the first hidden layer representation sequence.

[0044] First, a segment of meaningless noise signal needs to be generated, which can be achieved by a random number generator. Specifically, an appropriate distribution (such as Gaussian distribution, uniform distribution, etc.) can be selected according to requirements, and the length of the noise signal can be determined according to the application scenario and experimental design. Finally, the first hidden layer representation sequence and the generated meaningless noise signal are concatenated along the feature dimension.

[0045] Step S104, based on the pre-trained noise representation generation module based on flow matching, capture the dynamic changes of the first hidden layer representation sequence, and predict the second hidden layer representation sequence according to the context of the first hidden layer representation sequence. Then, replace the meaningless noise signal with the second hidden layer representation sequence to obtain a third hidden layer representation sequence.

[0046] Specifically, after concatenating a segment of meaningless noise signal to the first hidden layer representation sequence, input the first hidden layer representation sequence into the pre-trained noise representation generation module based on flow matching. Utilize the characteristic that the noise representation generation module based on flow matching can capture the dynamic changes of the first hidden layer representation sequence, and predict the second hidden layer representation sequence through the way of context learning. Then use the second hidden layer representation sequence to replace its meaningless noise signal to obtain a third hidden layer representation sequence.

[0047] Step S105, improve the time resolution of the third hidden layer representation sequence through the pre-trained noise autoencoder, and map the hidden layer representation sequence back to the original signal space to generate the second ship noise signal.

[0048] Specifically, input the third hidden layer representation sequence into the trained noise autoencoder. Through the decoder part in the autoencoder, improve the time resolution of the third hidden layer representation sequence, reduce the dimension of its hidden layer features, map the hidden layer representation sequence back to the original signal space, and finally decode it into the newly generated ship noise signal, that is, the second ship noise signal.

[0049] In this embodiment, the trained noise autoencoder can extract useful information from ship noise while suppressing noise, and its hidden layer representation can capture the long-term dependencies and global patterns of ship noise signals. The trained noise representation generation module based on flow matching can generate hidden layer representations that conform to the target distribution and replace meaningless noise signals. The noise generated by this method can not only conform to the information such as the boat model and working conditions of the reference noise, but also reflect its time-varying characteristics.

[0050] In one implementation, the noise autoencoder includes a noise encoder and a noise decoder. The noise encoder is composed of multiple one-dimensional CNNs stacked, and the noise decoder is composed of multiple transposed CNNs stacked. Specifically, the noise encoder is used to reduce the time resolution of the first ship noise signal and extract high-level features from it to generate a sequence of hidden layer representations. The noise decoder is used to increase the time resolution of the sequence of hidden layer representations and map the sequence of hidden layer representations back to the original signal space to generate a second ship noise signal similar to the first ship noise signal.

[0051] In one implementation, the noise autoencoder is trained to obtain a pre-trained noise autoencoder, and its training process includes:

[0052] Input the ship noise signal samples into the noise encoder to obtain a sequence of hidden layer representation samples.

[0053] Input the sequence of hidden layer representation samples into the noise decoder to obtain reconstructed ship noise signal samples.

[0054] Use the ship noise signal samples, the sequence of hidden layer representation samples, and the reconstructed ship noise signal samples as training samples to train the autoencoder to obtain a pre-trained noise autoencoder.

[0055] In a specific example, to more intuitively describe the implementation of training the noise autoencoder, please refer to Figure 2 , collect ship noise signal samples and input them into the noise encoder to obtain a sequence of hidden layer representations, namely Z1, Z2, Z3, Z4, Z5, Z6, Z7. Input the above sequence of hidden layer representation samples into the noise decoder to obtain reconstructed ship noise signal samples.

[0056] In one implementation, continue to refer to Figure 2 and in combination with Figure 3 , during the process of training the noise autoencoder, the spectral loss between the ship noise signal samples and the reconstructed ship noise signal samples is also calculated, and the noise autoencoder is iteratively adjusted according to the spectral loss.

[0057] In one implementation, the method for calculating the spectral loss between the ship noise signal samples and the reconstructed ship noise signal samples includes:

[0058] Select different Fourier transform window lengths and number of frequency band channels, and calculate multiple Mel spectrograms (Mel spectrogram is a way to convert sound signals into spectrograms) between the ship noise samples and the reconstructed ship noise signal samples. Specifically, the Fourier transform window length refers to the time length of each frame of signal when performing short-time Fourier transform (STFT). The number of frequency band channels refers to the number of filters in the Mel filter bank during the generation of the Mel spectrogram, that is, the resolution in the frequency dimension of the Mel spectrogram. Specifically, for a noisy audio w, the formula for calculating its spectrogram Melspec is as follows:

[0059] Melspec = Filter nmel (spec)

[0060] spec = STFT n (w)

[0061] where STFT refers to short-time Fourier transform; n represents the number of points of the short-time Fourier transform, and at the same time determines the length of the time-domain signal for STFT. For example, if n = 2048, 2048 points are selected for short-time Fourier transform; Filter refers to the Mel filter, and the Mel filter is a group of triangular band-pass filters based on the Mel scale, which is used to convert the spectrum of the audio signal from the linear frequency axis to the Mel frequency axis; nmel represents the number of groups of Mel filters, and the larger the number of groups, the higher the frequency resolution of the spectrogram.

[0062] In a specific example, n is set to 7 values, which are [256, 512, 1024, 2048, 4096, 8192, 16384] respectively, and the corresponding nmel is set to [40, 80, 160, 320, 640, 1280, 2560]. Among them, the latter three larger scales n and higher resolutions nmel are designed to adapt to the long-time stability and fine low-frequency line spectrum of the boat noise signal. Seven spectrograms can be calculated for a noisy audio through the above 7 groups of parameters.

[0063] After calculating multiple Mel spectrograms, sum the losses of multiple Mel spectrograms, and take this loss sum as the final loss. It can be calculated by the following formula:

[0064]

[0065] where Melspec i refers to the spectrogram of the boat noise signal sample, refers to the spectrogram of the reconstructed ship noise signal sample, and n refers to the number of spectrograms.

[0066] In a specific example,Figure 4 The low-frequency line spectrum of the ship noise samples reconstructed by the noise autoencoder before and after applying the multi-scale and multi-resolution spectral loss proposed in this embodiment is shown. Figure 4 (a) is the low-frequency line spectrum of the ship noise signal sample, Figure 4 (b) is the low-frequency line spectrum of the reconstructed ship noise signal sample before applying the multi-scale and multi-resolution spectral loss, Figure 4 (c) is the low-frequency line spectrum of the reconstructed ship noise signal sample after applying the multi-scale and multi-resolution spectral loss. It can be seen that after introducing this spectral loss, the reconstruction of the low-frequency part of the ship noise signal is significantly improved, and the low-frequency part is exactly the most concerned part in underwater acoustic signal processing.

[0067] In the above embodiment, since each spectrum corresponds to different time resolutions and frequency resolutions, it can ensure that short-time pulse signals or subtle spectral changes will not be over-smoothed. This enables the model to learn more effective information, while paying attention to the detailed features without affecting the listening experience of the reconstructed audio.

[0068] In one implementation, during the training of the noise autoencoder, a multi-subband and multi-period discriminator is introduced to jointly train with the noise autoencoder.

[0069] Specifically, the multi-subband discriminator evaluates the reconstruction quality of each frequency band by decomposing the signal into multiple frequency bands. First, the input signal is decomposed into multiple subbands using a filter bank or short-time Fourier transform (STFT). For example, the signal is decomposed into low-frequency, middle-frequency, and high-frequency subbands. An independent discriminator is designed for each subband. Each discriminator receives the reconstructed signal and the real signal of the corresponding subband and outputs a scalar value representing the authenticity of the signal. Calculate the adversarial loss of each subband discriminator, and sum the losses of all subbands weighted to obtain the total loss of the multi-subband discriminator.

[0070] The multi-period discriminator evaluates the quality of the reconstructed signal by extracting periodic features from different time scales. Specifically, the input signal is segmented according to different time windows (such as short window, medium window, long window). An independent discriminator is designed for each time scale. Each discriminator receives the reconstructed signal and the real signal of the corresponding time scale and outputs a scalar value representing the authenticity of the signal. Calculate the adversarial loss of each period discriminator, and sum the losses of all time scales weighted to obtain the total loss of the multi-period discriminator. Finally, the multi-subband and multi-period discriminators are jointly trained with the noise autoencoder, and multi-subband and multi-period analyses are performed on the reconstructed signal and the real signal respectively.

[0071] In the above embodiments, introducing a multi-subband and multi-period discriminator to train the noise autoencoder is a strategy that combines multi-scale analysis and adversarial training, which can significantly improve the performance of the noise autoencoder, especially in terms of the quality of reconstructed waveforms and feature extraction.

[0072] In one implementation, the construction process of the noise representation generation module based on flow matching specifically includes:

[0073] Taking the meaningless standard Gaussian noise used to replace the hidden layer representation as the starting distribution of the initial continuous normalizing flow, and setting the hidden layer representation of the ship noise signal as the final distribution of the continuous normalizing flow. Specifically, the continuous normalizing flow is a generative model based on ordinary differential equations. By defining a continuous vector field, it transforms a simple distribution into a complex distribution. That is, mapping the standard Gaussian noise (starting distribution) to the hidden layer representation of the ship noise signal (target distribution) through the continuous normalizing flow.

[0074] Then, build a neural network composed of 24 layers of Transformers to simulate the vector field in the continuous normalizing flow.

[0075] In one implementation, training the constructed noise representation generation module based on flow matching to obtain a pre-trained noise representation generation module based on flow matching, and its process specifically includes:

[0076] Sampling a batch of data from the training dataset and generating a sequence of hidden layer representations through the pre-trained ship autoencoder.

[0077] Sampling standard Gaussian noise from the standard Gaussian distribution, selecting a sequence in the sequence of hidden layer representations, and using the standard Gaussian noise to replace the selected sequence to generate context features.

[0078] Calculating intermediate features based on the hidden layer representation and the standard Gaussian noise. Specifically, it can be completed through the following formula:

[0079] x t =(1 - t)ζ + tx1

[0080] where ζ represents the standard Gaussian noise; x1 represents the hidden layer representation; t is randomly sampled between [0, 1].

[0081] Using the intermediate features and context features as inputs and x1 - ζ as the fitting target to train the Transformer neural network.

[0082] In a specific example, to more intuitively describe the implementation scheme of training the noise representation generation module based on flow matching, please refer to Figure 5 :

[0083] First, sample a batch of original ship noise data, and generate a hidden layer representation sequence, namely Z1, Z2, Z3, Z4, Z5, Z6, Z7, through the encoder in the pre-trained ship autoencoder.

[0084] Sample standard Gaussian noise (meaningless noise, replaced by 0 in this example) from the standard Gaussian distribution, select a sequence in the hidden layer representation sequence, namely Z5, Z6, Z7, and use the standard Gaussian noise to replace the selected sequence to generate context features, namely Z1, Z2, Z3, Z4, 0, 0, 0.

[0085] Calculate intermediate features based on the hidden layer representation and the standard Gaussian noise. Use the intermediate features and the context features as inputs, and use x1-ζ as the fitting target to train the Transformer neural network (i.e., the noise representation generation module based on flow matching) to generate the final hidden layer representation sequence, namely Z1, Z2, Z3, Z4, Z5', Z6', Z7'.

[0086] Finally, decode the final hidden layer representation sequence into a new ship noise signal through the decoder in the pre-trained ship autoencoder.

[0087] In one implementation, based on the pre-trained noise representation generation module based on flow matching, capture the dynamic changes of the hidden layer representation sequence, and predict the second hidden layer representation sequence according to the context of the hidden layer representation sequence. Then, replace the meaningless noise signal with the second hidden layer representation sequence to obtain the third hidden layer representation sequence. The method is completed through the following formula:

[0088]

[0089] Among them, x1′ represents the third hidden layer representation sequence; x t represents the intermediate feature; x ctx represents the context feature; v t refers to the simulated vector field.

[0090] In a specific example, to more intuitively describe the implementation scheme of ship noise generation in actual work, please refer to Figure 6 , a method for generating ship self-noise is as follows:

[0091] Collect the first ship noise signal, that is, the original ship noise signal without processing. Input the first ship noise signal into the encoder in the pre-trained noise autoencoder to convert the first ship noise signal into a hidden layer representation sequence, namely Z1, Z2, Z3, Z4, Z5, Z6, Z7.

[0092] Subsequently, a meaningless noise signal is concatenated after the above first hidden layer representation sequence, that is, the concatenated first hidden layer representation sequence is Z1, Z2, Z3, Z4, Z5, Z6, Z7, 0, …, 0.

[0093] Then, the concatenated first hidden layer representation sequence is input into a pre-trained noise representation generation module based on flow matching. The pre-trained noise representation generation module based on flow matching will predict a second hidden layer representation sequence according to the hidden layer representation sequence, that is, Z8, …, Zn. And replace the meaningless noise signal with the second hidden layer representation sequence to obtain a third hidden layer representation sequence, that is, Z1, Z2, Z3, Z4, Z5, Z6, Z7, Z8, …, Zn.

[0094] Finally, the third hidden layer representation sequence is decoded into a second ship noise signal through a pre-trained noise autoencoder, that is, the newly generated ship noise signal.

[0095] It can be seen from the above embodiments that the noise representation generation module based on flow matching can capture the dynamic change characteristics of ship noise, and introduce information such as the working conditions and models of ships through the way of context learning (the original noise information contains information such as the working conditions and models of ships), and generate a brand-new ship self-noise according to the original noise. The generated noise conforms to the working conditions and ship models of the original noise, but is different from the existing noise. This data augmentation method can be used for multiple tasks such as ship noise analysis and ship classification.

[0096] Exemplarily, Figure 7 A comparison graph of the new ship noise signal generated by the ship self-noise generation method provided by this solution and the original ship noise signal is shown. Through comparison, it can be seen that the generated ship noise signal is similar to the original ship noise signal in both the time domain and the frequency domain, and there is also volatility.

[0097] In one implementation manner, after generating the second ship noise signal, the reduction degree of the second ship noise signal is evaluated. The evaluation method includes:

[0098] The second ship noise signal and the first ship noise signal are respectively sliced with a window of 3 - 7 s, and Z-score and interquartile range are applied to extract high-energy frequency points below 500 - 700 Hz to characterize the transient low-frequency line spectrum features of the noise signal. For example, high-energy frequency points below 500, 550, 600, 650, and 700 are extracted.

[0099] Specifically, the first ship noise signal refers to the original ship signal, and the second ship noise signal refers to the newly generated ship noise signal. The first ship noise signal and the newly generated second ship noise signal are respectively subjected to a sliding window slicing of 3 - 7 seconds, preferably 5 seconds. If the signal length is not an integer multiple of 3 - 7 seconds, zero-padding or truncation processing can be performed on the last window. Z-score normalization is used to eliminate the amplitude differences of the signals, making different signals comparable. The interquartile range (IQR) is used to identify high-energy frequency points. The extracted high-energy frequency points are used as the transient low-frequency line spectrum features of the ship noise signal.

[0100] For the frequency points within 20 - 40 windows, density-based spatial clustering of applications with noise (such as selecting 20, 25, 30, 35, 40 windows) is used to remove outliers, so as to obtain the fluctuation distribution of the low-frequency line spectrum features in the long term. Specifically, statistical analysis can be performed on the clustered frequency points to extract the fluctuation distribution features.

[0101] The Wasserstein distance between the low-frequency line spectrum fluctuations of the second ship noise signal and the first ship noise signal is used to evaluate the restoration degree of the generated signal. Specifically, the Wasserstein distance is a measure of the difference between two probability distributions, which represents the minimum cost required to move one distribution to another. The smaller the Wasserstein distance, the closer the low-frequency line spectrum fluctuation distribution of the newly generated ship noise signal is to the original ship noise signal, and the higher the restoration degree of the generated signal.

[0102] Exemplarily, Figure 8 The transient low-frequency line spectrum features of the original ship noise signal and the newly generated ship noise signal at different times are shown. In this embodiment, the original ship noise signal and the newly generated ship noise signal are visualized, and the fluctuations of the low-frequency line spectrum features clustered according to the evaluation index proposed in this embodiment can be seen. It can be seen that the newly generated ship noise signal (red) shows line spectrum features similar to the original ship noise signal (blue, light blue) at different times, which cannot be achieved by traditional signal simulation.

[0103] In addition, in this embodiment, an objective index evaluation is also performed on the generated samples (i.e., the newly generated ship noise signals), and the original samples (Real-DT) at different times are used as the baseline samples. Table 1 shows the Wasserstenin distance (WD-lofar) between the fluctuation distribution of the low-frequency line spectrum of the generated samples generated by the self-noise generation model proposed in this scheme and the fluctuation distribution of the original samples (i.e., the original ship noise signals). It can be seen from this that the performance of the generated samples is close to that of the original samples at different times. In addition, in this embodiment, the centroid of the fluctuation distribution is also used to simulate the static low-frequency line spectrum distribution (Static) of the signals generated according to experience and signal processing methods. Due to its lack of variation, there is an obvious difference from the original signal in terms of the WD-lofar index. In addition, for the comparison of the overall spectrum, the MCD of the generation model proposed by the present invention is close to the MCD of the original samples at different times.

[0104]

[0105] Table 1. Comparison of objective indexes between original samples (at different times) and generated samples

[0106] Corresponding to the above method provided by the present invention, the present invention also provides a device. Figure 9 The structure diagram of a ship self-noise generation device based on a neural network provided by an embodiment of this specification is shown. As Figure 9 shown, the device includes:

[0107] A ship noise acquisition module 201, configured to acquire first ship noise signals within a plurality of consecutive time periods.

[0108] A pre-trained noise encoding module 202, configured to convert the first ship noise signals into a first hidden layer representation sequence.

[0109] A hidden layer representation sequence splicing module 203, configured to splice a meaningless noise signal after the first hidden layer representation sequence.

[0110] A pre-trained noise representation generation module 204 based on flow matching, configured to capture the dynamic changes of the hidden layer representation sequence, predict a second hidden layer representation sequence according to the context of the first hidden layer representation sequence, and then replace the meaningless noise signal with the second hidden layer representation sequence to obtain a third hidden layer representation sequence.

[0111] A pre-trained noise decoding module 205, configured to decode the third hidden layer representation sequence into second ship noise signals.

[0112] It should be noted that Figure 9 the shown device corresponds to Figure 1 the method embodiment shown, and can be respectively applied to Figure 1The conversion of the first ship noise signal, the splicing of meaningless noise signals, the prediction of the second hidden layer representation sequence, and the decoding of the third hidden layer representation sequence in the illustrated method embodiments will not be elaborated herein. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0113] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor, where executable code is stored in the memory, and when the processor executes the executable code, the method described in combination with Figure 1 is implemented.

[0114] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in combination with Figure 1 is implemented.

[0115] Those skilled in the art should be able to realize that, in one or more of the above examples, the functions described in this invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0116] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of this invention. It should be understood that the above is only the specific embodiments of this invention and is not used to limit the protection scope of this invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this invention should be included in the protection scope of this invention.

Claims

1. A method for generating ship self-noise based on a neural network, characterized in that, The method includes: Collecting first ship noise signals within multiple consecutive time periods; Reducing the time resolution of the first ship noise signals through a pre-trained noise autoencoder, and extracting high-level features therefrom to obtain a first hidden layer representation sequence; Concatenating a segment of meaningless noise signal after the first hidden layer representation sequence; Based on a pre-trained noise representation generation module based on flow matching, capturing the dynamic changes of the first hidden layer representation sequence, and predicting a second hidden layer representation sequence according to the context of the first hidden layer representation sequence, and then replacing the meaningless noise signal with the second hidden layer representation sequence to obtain a third hidden layer representation sequence; Enhancing the time resolution of the third hidden layer representation sequence through a pre-trained noise autoencoder, and mapping the third hidden layer representation sequence back to the original signal space to generate a second ship noise signal.

2. The method according to claim 1, wherein The noise autoencoder includes a noise encoder and a noise decoder. The noise encoder is composed of multiple one-dimensional CNNs stacked, and the noise decoder is composed of multiple transposed CNNs stacked.

3. The method according to claim 2, wherein Training the noise autoencoder to obtain a pre-trained noise autoencoder, and its training process includes: Inputting ship noise signal samples into the noise encoder to obtain hidden layer representation sequence samples; Inputting the hidden layer representation sequence samples into the noise decoder to obtain reconstructed ship noise signal samples; Using the ship noise signal samples, the hidden layer representation sequence samples, and the reconstructed ship noise signal samples as training samples to train the autoencoder to obtain a pre-trained noise autoencoder.

4. The method according to claim 3, wherein During the process of training the noise autoencoder, the spectral loss between the ship noise samples and the reconstructed ship noise signal samples is also calculated, and the noise autoencoder is iteratively adjusted according to the spectral loss.

5. The method according to claim 4, wherein The method for calculating the spectral loss includes: Selecting different Fourier transform window lengths and number of frequency band channels, and calculating multiple Mel spectrograms between the ship noise signal samples and the reconstructed ship noise signal samples; Summing the losses of the multiple Mel spectrograms, and using the sum of the losses as the final loss.

6. The method according to any one of claims 3-5, characterized in that During the process of training the noise autoencoder, a multi-subband and multi-period discriminator is introduced for joint training with the noise autoencoder.

7. The method according to claim 1, characterized in that, The construction process of the noise representation generation module based on flow matching includes: Using the meaningless standard Gaussian noise for replacing the hidden layer representation as the starting distribution of the continuous normalizing flow, and setting the hidden layer representation of the ship noise signal as the final distribution of the continuous normalizing flow; Building a neural network composed of 24 layers of Transformers to simulate the vector field in the continuous normalizing flow.

8. The method according to claim 7, wherein Training the constructed noise representation generation module based on flow matching to obtain a pre-trained noise representation generation module based on flow matching, and its process specifically includes: Sampling a batch of data from the training dataset, and generating a hidden layer representation sequence through a pre-trained noise autoencoder; Sampling standard Gaussian noise from the standard Gaussian distribution, selecting a segment of sequence in the hidden layer representation sequence, and using the standard Gaussian noise to replace the selected sequence to generate context features; The intermediate features are calculated based on the hidden layer representation sequence and standard Gaussian noise, which is completed through the following formula: x t = (1 - t)ζ + tx1 where ζ represents the standard Gaussian noise; x1 represents the hidden layer representation sequence; t is randomly sampled between [0, 1]; Using the intermediate features and context features as inputs, and x1 - ζ as the fitting target, to train the Transformer neural network.

9. The method according to claim 1, wherein A method based on a pre-trained noise representation generation module based on flow matching, which captures the dynamic changes of the hidden layer representation sequence and predicts the second hidden layer representation sequence according to the context of the hidden layer representation sequence, and then replaces the meaningless noise signal with the second hidden layer representation sequence to obtain the third hidden layer representation sequence, is completed through the following formula: x′1 = ζ+∫0 1 v t (,x t ,x ctx )dt Among them, x1′ represents the third hidden layer representation sequence; x t represents the intermediate feature; x ctx represents the context feature.

10. The method according to claim 1, wherein After generating the second ship noise signal, evaluate the restoration degree of the second ship noise signal. The evaluation method includes: Slice the second ship noise signal and the first ship noise signal with a window of 3 - 7 s respectively, and apply Z-score and interquartile range to extract high-energy frequency points below 500 - 700 HZ to characterize the transient low-frequency line spectrum features of the noise signal; Use density-based spatial clustering of noise applications for the frequency points within 20 - 40 windows to remove outliers, so as to obtain the fluctuation distribution of the low-frequency line spectrum features in the long term; Use the Wasserstein distance between the low-frequency line spectrum fluctuations of the second ship noise signal and the first ship noise signal to evaluate the restoration degree of the generated signal.