A ship radiated noise generation method and system based on flow matching

By using a flow matching-based method, the low-frequency line spectrum and mechanical modulation features of ship radiated noise are extracted. Combined with a neural audio codec and a flow matching acoustic model, a high-fidelity, controllable time-varying ship radiated noise signal is generated, solving the problem of signal distortion in existing technologies and realizing the generation of ship radiated noise with natural time-varying characteristics.

CN122090853APending Publication Date: 2026-05-26INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF ACOUSTICS CHINESE ACAD OF SCI
Filing Date
2026-02-25
Publication Date
2026-05-26

Smart Images

  • Figure CN122090853A_ABST
    Figure CN122090853A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for generating ship radiated noise based on flow matching. The method includes: firstly, constructing a generative model comprising a feature extraction layer, a neural audio encoder / decoder layer, and a flow matching acoustic model layer; extracting low-frequency line spectrum features and mechanical modulation features from a reference signal using the feature extraction layer; encoding the signal into a continuous latent space using a pre-trained neural audio encoder; during the training phase, masking the latent layer representation, and using the target state and Gaussian noise in the masked region as a starting point, training a neural network direction estimator to learn the conditional probability flow in the latent space by combining the time step and the extracted features; during the inference phase, based on the new reference signal features and target duration, using Gaussian noise as the initial state, guiding the ordinary differential equation solver to integrate in the latent space through the trained direction estimator to generate a third latent layer representation sequence, which is finally output as a time-domain waveform by the decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater acoustic signal generation technology, specifically to a method and system for generating ship radiated noise based on flow matching. Background Technology

[0002] Ship radiated noise is a crucial data foundation for underwater target identification, acoustic countermeasures, fault diagnosis, and underwater acoustic simulation systems. High-quality and diverse ship radiated noise samples are essential for improving the robustness and generalization ability of relevant models. However, the acquisition of real ship radiated noise data is costly and time-consuming, and is constrained by multiple variables such as ship type, navigation status, and marine environment, making it difficult to cover the entire operating condition spectrum. More importantly, it is difficult to obtain a long-term, continuously varying noise sequence of the same ship under the same stable operating conditions. Furthermore, research has confirmed that even under constant speed and heading conditions, real ship radiated noise still exhibits significant time-varying characteristics—its line spectrum frequency, modulation depth, and broadband energy distribution naturally fluctuate over time. Therefore, developing high-fidelity, controllable, and scalable ship radiated noise generation technology has become an urgent technical problem to be solved in this field.

[0003] In recent years, neural network models have made breakthrough progress in the fields of speech and environmental audio. Some studies have attempted to transfer them to the field of underwater acoustic signal processing. For example, some researchers have used neural networks to predict the radiated noise at the passing frequency of the three blades of a propeller under specific operating conditions. However, the training objective of this technique still relies on static numerical simulation results and cannot characterize the time-varying characteristics of real noise. Other studies have directly used generative adversarial networks to generate ship radiated noise. However, the generated samples are mostly based on power spectrum constraints and fail to consider long-term dynamic evolution characteristics, resulting in distortion of the spectral details of the generated signal (especially in the narrowband line spectrum region). Furthermore, the time-varying realism of the generated signal has not been systematically modeled and quantitatively evaluated, making it difficult to meet the requirements of noise sample fidelity and dynamic consistency in practical applications.

[0004] Therefore, a method and system for generating ship radiated noise based on flow matching is needed. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for generating ship radiated noise based on flow matching. Without requiring complex physical parameters, it can generate a high-fidelity ship radiated noise signal with controllable ship type and operating conditions and natural time-varying fluctuation characteristics based only on short-time reference audio.

[0006] To achieve the above objectives, in a first aspect, the present invention provides a method for generating ship radiated noise based on flow matching, comprising:

[0007] Collect the first ship's radiated noise signal over multiple consecutive time periods; obtain the second ship's radiated noise warning signal and the duration of the second ship's radiated noise generation;

[0008] A ship radiated noise generation model is constructed, which includes a feature extraction layer, a neural audio codec layer, and a flow-matching acoustic model layer.

[0009] The feature extraction layer extracts the first low-frequency line spectrum features and the first mechanical modulation features of the first ship radiated noise signal, and extracts the second low-frequency line spectrum features and the second mechanical modulation features of the second ship radiated noise warning signal; wherein the second low-frequency line spectrum features are used to determine the line spectrum structure and ship type of the second ship radiated noise, and the second mechanical modulation features are used to determine the modulation characteristics of the second ship radiated noise.

[0010] The first ship radiated noise signal is encoded into a first hidden layer representation sequence by a pre-trained neural audio encoder in the neural audio encoder-decoder layer, and the second ship radiated noise cue signal is encoded into a second hidden layer representation sequence.

[0011] In the training phase of the flow-matching acoustic model layer: a first region is randomly masked in the first hidden layer representation sequence, and the masked sequence is used as a context constraint, wherein the original hidden layer representation sequence of the first region is used as the first target state; a standard Gaussian noise of the same length as the first region is generated as the first initial state; linear interpolation is performed on the first initial state and the first target state according to the time step parameters to obtain the current interpolated hidden layer state; the first low-frequency line spectrum feature, the first mechanical modulation feature, the time step parameters, the interpolated hidden layer state, and the context constraint are concatenated and input into the neural network direction estimator for training; the goal of the training is to enable the velocity field predicted by the neural network direction estimator to guide the ordinary differential equation solver to reconstruct the first target state from the first initial state;

[0012] In the inference phase of the flow-matched acoustic model layer: Gaussian noise with a length corresponding to the generation duration of the second ship radiated noise is used as the second initial state; the second initial state, the second hidden layer representation sequence, the second low-frequency line spectrum feature, and the second mechanical modulation feature are input to the trained neural network direction estimator; the neural network direction estimator guides the ordinary differential equation solver to perform step-by-step integration to obtain the third hidden layer representation sequence; the third hidden layer representation sequence is decoded into the second ship radiated noise by the pre-trained neural audio decoder in the neural audio codec layer.

[0013] Specifically, the feature extraction layer includes a low-frequency line spectrum feature extraction layer. The low-frequency line spectrum feature extraction layer performs a short-time Fourier transform on the ship radiated noise signal to obtain the time spectrum; standardizes each frame of spectrum and filters peak points with energy higher than the background; retains peak points within a preset low-frequency band; and performs density clustering on the peak points after filtering multiple frames to remove noise points, forming a LOFAR distribution map with time as the horizontal axis and frequency as the vertical axis, which serves as the low-frequency line spectrum feature.

[0014] Preferably, the preset low frequency band is from 1Hz to 600Hz.

[0015] Specifically, the feature extraction layer includes a mechanical modulation feature extraction layer, which performs bandpass filtering on the ship's radiated noise signal; performs full-wave rectification on the filtered signal to extract the amplitude envelope; performs low-pass filtering on the amplitude envelope to smooth high-frequency jitter; performs Fourier transform on the smoothed envelope and truncates a preset modulation frequency band to form a DEMON spectrum with time as the horizontal axis and frequency as the vertical axis, which serves as the mechanical modulation feature.

[0016] Preferably, the cutoff frequency of the bandpass filter is 50Hz to 5000Hz; the preset modulation frequency band is 0Hz to 50Hz.

[0017] Specifically, when randomly masking the first hidden layer representation sequence, the masking ratio is dynamically adjusted between 70% and 100%.

[0018] Preferably, the neural network orientation estimator adopts the Transformer model architecture.

[0019] Specifically, the training of the neural audio codec layer employs joint optimization of multi-scale Mel-spectrum loss and adversarial loss; wherein, the adversarial loss is calculated by introducing a multi-scale discriminator that makes judgments in the time domain and at multiple STFT scales, and the multi-scale discriminator adds a denser STFT resolution window in the frequency band below 200Hz.

[0020] Secondly, the present invention provides a ship radiated noise generation system based on flow matching, the system comprising:

[0021] The feature extraction module is configured to extract low-frequency line spectrum features and mechanical modulation features from the input ship radiated noise signal;

[0022] The neural audio codec module includes a pre-trained neural audio encoder and a neural audio decoder; the encoder is used to encode a ship radiated noise waveform into a hidden layer representation sequence, and the decoder is used to decode the hidden layer representation sequence into a ship radiated noise waveform.

[0023] The flow-matching acoustic model module includes a neural network direction estimator and an ordinary differential equation solver;

[0024] During the training phase, the system is used for:

[0025] The first ship radiated noise signal is processed by the feature extraction module and the neural audio encoder to obtain the first feature and the first hidden layer representation sequence.

[0026] The first hidden layer representation sequence is masked, and the neural network orientation estimator is trained based on the first starting state, the first target state, the first feature, and the time step parameters of the masked region.

[0027] During the inference phase, the system is used for:

[0028] The second ship radiated noise warning signal is processed by the feature extraction module and the neural audio encoder to obtain the second feature and the second hidden layer representation sequence.

[0029] Based on the initial state of Gaussian noise for the target duration, the second hidden layer representation sequence, and the second feature, the trained neural network direction estimator guides the ordinary differential equation solver to generate the third hidden layer representation sequence.

[0030] The neural audio decoder decodes the third hidden layer representation sequence into second ship radiated noise. Attached Figure Description

[0031] Figure 1 A schematic diagram of a method for generating ship radiated noise based on flow matching provided in an embodiment of the present invention;

[0032] Figure 2 A flowchart illustrating a method for generating ship radiated noise based on flow matching, provided in an embodiment of the present invention;

[0033] Figure 3 A feature extraction flowchart provided in an embodiment of the present invention;

[0034] Figure 4 This is a schematic diagram of the structure of a neural audio codec model provided in an embodiment of the present invention;

[0035] Figure 5 This is a schematic diagram comparing a reference sample and a generated sample, provided as an embodiment of the present invention. Detailed Implementation

[0036] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be described below with reference to the accompanying drawings. It should be noted that similar reference numerals and letters in the following drawings indicate similar items; therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0038] In the description of the embodiments of the present invention, the words "exemplary," "for example," or "for instance" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary," "for example," or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the words "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a specific manner.

[0039] In existing technologies, the main methods for generating ship radiated noise are divided into physical modeling methods and traditional signal synthesis methods. Physical modeling methods decompose ship radiated noise into typical components such as mechanical noise, propeller noise, and cavitation noise. Each component signal is synthesized separately through analytical models or high-fidelity numerical simulations, and then the components are superimposed to obtain the second ship radiated noise. Although this type of method has physical interpretability, it heavily relies on precise prior information such as ship structural parameters and operating state parameters, which are usually difficult to obtain accurately in actual engineering scenarios. Moreover, it generally assumes that the characteristics of each noise source are in a static and stable state, which cannot characterize the inherent non-stationary fluctuation characteristics of real ship radiated noise. This results in the generated signal exhibiting a rigid and singular spectral structure over long time scales, with significant time-varying differences from the measured signal. Traditional signal synthesis methods rely on conventional signal processing techniques such as sine wave superposition and filter bank modulation to construct the line spectrum and broadband components of ship radiated noise separately, and then integrate them to obtain the target noise signal. While these methods are computationally efficient, the line spectrum sharpness and modulation envelope shape of the generated signal deviate significantly from the actual ship radiated noise. Furthermore, they lack explicit control mechanisms for ship type and navigation conditions, making it difficult to generate ship radiated noise under specific types and conditions as needed. In addition, the line spectrum, continuous spectrum, and modulation spectrum characteristics of the simulated signal obtained by these methods have a low degree of matching with the actual ship radiated noise, failing to meet the requirements of real-world noise sample accuracy in practical applications.

[0040] To overcome the shortcomings of existing technologies, this invention provides a method for generating ship radiated noise based on flow matching. Without requiring complex physical parameters, it can generate high-fidelity ship radiated noise signals with controllable ship type and operating conditions, and natural time-varying fluctuation characteristics based solely on short-time reference audio.

[0041] Figure 1The figure shows a schematic diagram of a method for generating ship radiated noise based on flow matching, which is provided in an embodiment of the present invention. The core process of this method revolves around two stages: training and inference, and aims to generate long-term ship radiated noise with natural time-varying characteristics from short-time reference audio.

[0042] The goal of the training phase is to enable the model to complete masked representations based on acoustic cues. In this phase, the complete original audio is used. First, the audio is encoded into a temporal sequence of hidden layer representations using a neural audio codec, and its LOFAR and DEMON spectra are extracted simultaneously. Next, the hidden layer representation sequence is randomly masked, retaining only partial segments, and linear interpolation is performed on the masked regions. Subsequently, the model fuses these interpolated hidden layer segments with acoustic cues such as the DEMON spectrum and inputs them into the Transformer module for learning. The optimization goal of the model is to make its output predicted hidden layers approximate the true hidden layers of the masked regions as closely as possible, thereby learning to reconstruct coherent time-varying features based on contextual fragments and acoustic conditions.

[0043] The inference phase utilizes the learned capabilities to generate complete time-varying noise based on limited reference information. This phase requires only a short reference audio segment. Similarly, the corresponding reference hidden layer segment is first obtained, and reference acoustic cues are extracted. Simultaneously, LOFAR spectral cues defining the generation target (e.g., ship type, duration) are provided. Unlike training, the hidden layer region to be generated is initialized directly with Gaussian noise. Then, this initial sequence and acoustic cues (including the DEMON spectrum) are input into the trained model, and multi-step iterative optimization is performed through an ordinary differential equation solver to obtain a complete hidden layer representation sequence. The generated portion corresponding to the masked region is extracted and concatenated with the unmasked portion of the reference hidden layer representation to obtain the third hidden layer representation sequence. This sequence is input into the decoder to reconstruct a high-fidelity ship radiated noise waveform with natural time-varying characteristics.

[0044] Figure 2 A flowchart of a ship radiated noise generation method based on flow matching, provided for an embodiment of the present invention, is shown in the figure. The method includes:

[0045] Step S101: Collect the first ship radiated noise signal within multiple consecutive time periods; obtain the second ship radiated noise warning signal and the second ship radiated noise generation duration.

[0046] Step S102: Construct a ship radiated noise generation model, which includes a feature extraction layer, a neural audio codec layer, and a flow-matching acoustic model layer.

[0047] Step S103: Extract the first low-frequency line spectrum feature and the first mechanical modulation feature of the first ship radiated noise signal through the feature extraction layer; extract the second low-frequency line spectrum feature and the second mechanical modulation feature of the second ship radiated noise warning signal through the feature extraction layer; wherein, the second low-frequency line spectrum feature is used to determine the line spectrum structure and ship type of the second ship radiated noise, and the second mechanical modulation feature is used to determine the modulation characteristics of the second ship radiated noise.

[0048] For example, a short-time Fourier transform is performed on the reference ship radiated noise signal to obtain the time spectrum. The spectrum of each frame is normalized by Z-score and peak frequency points with higher energy than the background are screened by combining interquartile range. Peak frequency points that pass the screening within the 1~600 Hz frequency band are retained. A density-based noise application spatial clustering algorithm is used to remove isolated noise points to obtain the LOFAR distribution map and use it as a low-frequency line spectrum feature.

[0049] A 50-5000Hz bandpass filter is applied to the pre-acquired reference ship radiated noise signal. The amplitude envelope of the filtered signal is extracted and smoothed by low-pass filtering. The smoothed envelope is then subjected to Fourier transform, and the modulation frequency band of 0-50Hz is extracted to obtain the DEMON spectrum, which is used as the mechanical modulation feature.

[0050] For example, Figure 3 A feature extraction flowchart provided for an embodiment of the present invention is shown in the figure. Two types of low-dimensional acoustic cues are extracted from a segment of ship radiated noise signal to guide the generation process of ship radiated noise. The low-dimensional acoustic cues include low-frequency line spectrum distribution (LOFAR spectrum) and mechanical modulation spectrum (DEMON spectrum). The specific feature extraction process is as follows:

[0051] 1. Extraction of LOFAR spectrum:

[0052] A Short-Time Fourier Transform (STFT) operation is performed on the reference ship radiated noise signal: An STFT calculation is performed using a 2-second Hamming window and a 1-second time overlap to obtain the time spectrum of the reference signal. The LOFAR spectrum is calculated using the following formula:

[0053]

[0054] Represents a time-domain signal. Indicates a Hanming window;

[0055] Z-score normalization is performed on each frame of the time spectrum data, and then the frequency peaks with significantly higher energy than the background noise are selected by combining the interquartile range criterion.

[0056] Extract the low-frequency band of 1~600 Hz—this band contains the line spectrum components in ship radiated noise that have the ability to distinguish categories, and retain the frequency peak points in the band that have passed the above screening;

[0057] The set of frequency peak points after multi-frame filtering is input into the density clustering algorithm (DBSCAN) to remove isolated noise points, and finally form the LOFAR spectrum. The LOFAR spectrum is a two-dimensional time-frequency spectrum, with the horizontal axis corresponding to the time dimension and the vertical axis corresponding to the frequency dimension. Pixel intensity is used to characterize the probability of high-energy peaks.

[0058] 2. Extraction of DEMON spectrum:

[0059] A bandpass filter is applied to the original ship radiated noise signal: the cutoff frequency range of the bandpass filter is 50~5000 Hz, which is used to suppress extremely low frequency drift signals and high frequency interference signals, and retain the broadband components with significant modulation energy;

[0060] Perform full-wave rectification on the bandpass filtered signal and extract the amplitude envelope of the signal;

[0061] A low-pass filter is used to smooth the amplitude envelope in order to filter out high-frequency jitter signals that are independent of the mechanical cycle;

[0062] A Fourier transform is performed on the smoothed amplitude envelope, and the modulation band of 0~50 Hz is extracted—this band covers the modulation range of typical propeller speed and its harmonics, and finally a DEMON spectrum is formed: the DEMON spectrum is a two-dimensional time-frequency spectrum, with its horizontal axis corresponding to the time dimension and its vertical axis corresponding to the frequency dimension.

[0063] Both the LOFAR spectrum and the DEMON spectrum are output in the form of two-dimensional time-frequency spectra. The LOFAR spectrum is used to dominate the global line spectrum structure and ship type discrimination characteristics of the generated signal, while the DEMON spectrum is used to constrain the modulation depth and time-varying performance of the generated signal. The two types of acoustic cues serve as the conditional inputs to the flow-matching acoustic model to ensure that the generated ship radiated noise is controllable and realistic in terms of ship type consistency and dynamic performance.

[0064] Step S104: The first ship radiated noise signal is encoded into a first hidden layer representation sequence and the second ship radiated noise cue signal is encoded into a second hidden layer representation sequence using the pre-trained neural audio encoder in the neural audio encoder-decoder layer.

[0065] For example, the neural audio codec includes a temporal encoder and a symmetric decoder. The temporal encoder consists of a one-dimensional convolutional layer and convolutional blocks, and the symmetric decoder is a mirror-symmetric structure consisting of multiple transposed convolutional blocks. Ship radiated noise samples are input into the temporal encoder, and initial features are extracted through a one-dimensional convolutional layer. Then, the samples are sequentially passed through multiple convolutional blocks, each containing a cubic residual unit and a Snake activation function to obtain a hidden layer representation sequence. The hidden layer representation sequence is input into the symmetric decoder to obtain a reconstructed signal. Different Fourier transform window lengths and frequency band channels are selected to calculate multiple Mel spectra between the samples and the reconstructed signal. The L1 loss of the multiple Mel spectra is summed as the multi-scale Mel spectrum loss. The adversarial loss is calculated by a multi-scale discriminator, which adds higher resolution and an STFT window in the frequency band below 200Hz to jointly optimize the neural audio codec.

[0066] For example, Figure 4 The figure shows a schematic diagram of a neural audio codec model provided in an embodiment of the present invention. The neural audio codec maps the ship radiated noise waveform to a high-dimensional continuous latent space representation and reconstructs it into a time-domain waveform. The neural audio codec consists of a time-domain encoder and a symmetrically set decoder. The time-domain encoder and the symmetrical decoder are jointly trained to ensure that the latent space can compactly preserve key underwater acoustic feature details such as narrowband line spectrum sharpness and weak modulation structure.

[0067] The temporal encoder takes the original ship radiated noise waveform as input. First, it performs initial feature extraction through a weighted, one-dimensional convolutional layer. Then, it passes through four convolutional blocks, each containing three residual units and a Snake activation function. The residual units use dilated convolutions, and the dilation rates of each residual unit are set to 1, 3, and 9, respectively. Temporal downsampling is then achieved through stride convolutions with strides of 2, 5, 8, and 8, achieving a total compression ratio of 640 times, and the hidden layer feature dimension is set to 256. Finally, a continuous hidden vector sequence is output through a convolutional layer.

[0068] The symmetric decoder adopts a structure mirror-symmetric to the time-domain encoder, replacing the downsampling convolution in the time-domain encoder with transposed convolution, and the stride of the transposed convolution is 8, 8, 5, 2 in sequence; the configuration of the remaining residual units and activation functions in the decoder is consistent with that of the time-domain encoder to ensure stable reconstruction from the hidden layer representation sequence to the time-domain waveform.

[0069] During training, the neural audio codec employs a joint optimization approach using multi-scale Mel spectrum loss and adversarial loss. The multi-scale Mel spectrum loss is obtained by calculating the L1 distance between the Mel spectra of the generated and target signals under different FFT windows. The adversarial loss is achieved by introducing a multi-scale discriminator, which judges the authenticity of the reconstructed waveform in the time domain and at multiple STFT scales. In particular, to enhance the reconstruction accuracy of the low-frequency line spectrum of ship radiated noise, the multi-scale discriminator adds a denser STFT resolution window in the frequency band below 200Hz.

[0070] After the above training, the neural audio codec can compress a 1-second-long, 32kHz-sampling ship radiated noise into a continuous hidden layer representation sequence of only 50 frames in length and 256 dimensions while maintaining high-fidelity reconstruction performance. This continuous hidden layer representation sequence does not contain vector quantization, which can avoid information discretization loss and thus provide a differentiable and highly expressive generation space for the subsequent conditional flow matching model.

[0071] Step S105: In the training phase of the flow-matching acoustic model layer: a first region is randomly masked in the first hidden layer representation sequence, and the masked sequence is used as a context constraint, wherein the original hidden layer representation sequence of the first region is used as the first target state; a standard Gaussian noise of the same length as the first region is generated as the first initial state; according to the time step parameters, the first initial state and the first target state are linearly interpolated to obtain the current interpolated hidden layer state; the first low-frequency line spectrum feature, the first mechanical modulation feature, the time step parameters, the interpolated hidden layer state, and the context constraint are concatenated and input into the neural network direction estimator for training; the goal of the training is to enable the velocity field predicted by the neural network direction estimator to guide the ordinary differential equation solver to reconstruct the first target state from the first initial state.

[0072] For example, during the training phase, a continuous region is randomly masked from the hidden layer representation sequence of the first ship's radiated noise signal, with a masking ratio of 70% to 100%. Standard Gaussian noise is used as the initial distribution, and the masked hidden layer representation sequence is used as the target distribution. A Transformer neural network is constructed as the direction estimator. The input includes the interpolated hidden layer state, the flow time step, the masked context, low-frequency line spectrum features, and mechanical modulation features. The training aims to minimize the error between the predicted velocity direction and the actual transmission direction, resulting in a trained conditional flow matching acoustic model.

[0073] For example, the hidden layer representation obtained after processing the original audio using the aforementioned neural audio temporal encoder is first randomly masked by a continuous region. The masking ratio is dynamically set to 70%~100%, with the masked portion set to zero, and the remaining unmasked portion retained as context constraints. The true hidden layer value corresponding to the masked region is considered the target state, and a standard Gaussian noise segment of the same length as the masked region is generated as the initial state. During training, the model randomly selects a time step parameter between 0 and 1, and then linearly interpolates the initial state and the target state according to this time step to obtain the interpolated hidden layer state at the current time. This state is the main input that the direction estimator needs to process.

[0074] Besides the interpolation state, the direction estimator also receives several other types of conditional information: first, the remaining contextual hidden layer representation after masking, which tells the model what some known information looks like; second, the LOFAR spectrum extracted from the original audio, which describes the time-frequency distribution characteristics of the ship's line spectrum. During the training phase, the reference audio and target generation use the same LOFAR spectrum, constituting a self-reconstruction task; and third, the DEMON spectrum, which characterizes mechanical modulation characteristics, such as periodic fluctuations caused by propeller speed. These three types of conditional information are encoded and concatenated with the interpolation hidden layer state and time step parameters into a matrix (the Transformer can receive the input matrix dimension) and then fed into the Transformer-based direction estimator.

[0075] Among them, the low-frequency line spectrum hint is composed of LOFAR spectrum features extracted by the aforementioned feature extraction module, which is used to control the ship type and global spectrum structure of the generated noise; the mechanical modulation hint is composed of DEMON spectrum features extracted by the aforementioned feature extraction module, which is used to control the periodic fluctuation characteristics related to the propulsion system operating conditions.

[0076] The direction estimator is trained by minimizing the deviation between the predicted velocity direction and the actual transmission direction, that is, minimizing the error between the velocity field learned by the model and the actual velocity field. Ultimately, the direction estimator learns to construct a generation path in the hidden layer space that smoothly transitions from standard Gaussian noise to the target hidden layer representation.

[0077] Step S106: In the inference stage of the flow-matched acoustic model layer: Gaussian noise with a length corresponding to the generation duration of the second ship radiated noise is used as the second initial state; the second initial state, the second hidden layer representation sequence, the second low-frequency line spectrum feature, and the second mechanical modulation feature are input to the trained neural network direction estimator; the neural network direction estimator guides the ordinary differential equation solver to perform step-by-step integration to obtain the third hidden layer representation sequence; the third hidden layer representation sequence is decoded into the second ship radiated noise by the pre-trained neural audio decoder in the neural audio codec layer.

[0078] For example, during the inference phase, the low-frequency line spectrum features and mechanical modulation features of the second ship radiated noise warning signal are extracted. The target generation duration is determined based on the time frame length of the low-frequency line spectrum features. Starting with Gaussian noise, the hidden layer representation sequence of the reference ship radiated noise signal, the splicing result of the low-frequency line spectrum features of the reference and the target, and the mechanical modulation features are input into the trained direction estimator. An ordinary differential equation solver is used to integrate step by step along the generation path to obtain the complete hidden layer representation sequence. The generated part of the masked region of the hidden layer representation sequence of the reference ship radiated noise signal in the complete hidden layer representation sequence is spliced ​​with the unmasked region of the hidden layer representation sequence of the reference ship radiated noise signal to obtain the third hidden layer representation sequence.

[0079] For example, after training, the model has learned a set of "navigation rules," knowing how to evolve random noise step by step into a reasonable signal with specific ship acoustic characteristics in the latent space. In practice, the Gaussian noise corresponding to the target generation time is first used as the initial state. Simultaneously, the latent representation of the reference audio (providing local context), the LOFAR spectrum concatenation result of the reference and target (determining the line spectrum structure and ship type), and the DEMON spectrum (determining modulation characteristics) are all fed into the trained direction estimator. The direction estimator no longer predicts errors but, based on the current noise state, the current time point (from 0 to 1), and all cue features, outputs real-time guidance on "which direction to move next."

[0080] The ordinary differential equation solver follows this guidance to perform 16 iterations: starting from time 0, each step queries the direction estimator to obtain the direction of movement, then moves forward in small steps along that direction, updating the current hidden state; then it moves to the next time point, repeating the query and movement, and so on for 16 times until time 1 is reached. These 16 steps are equivalent to evenly dividing the entire generation process into 16 intermediate stages, each step being continuously guided by prompts, ensuring that the noise is gradually "shaped" into a hidden representation that conforms to the target ship type and operating conditions during the evolution process.

[0081] The generated portion of the masked region of the original audio hidden layer representation sequence from the complete hidden layer representation sequence during the training stage is extracted and concatenated with the unmasked original portion of the hidden layer representation sequence to form the final third hidden layer representation sequence. This third hidden layer representation sequence is then input into the previously trained neural audio decoder for conversion, thereby generating and outputting a high-fidelity time-varying ship radiated noise waveform signal with variable length, controllable ship type and operating conditions, and natural time fluctuation characteristics.

[0082] To verify the effectiveness and practicality of the method proposed in this invention, a verification experiment was conducted on the publicly available DeepShip ship radiated noise dataset. This dataset contains real underwater radiated noise recordings of four types of ships: cargo ships, passenger ships, tankers, and tugs, with a sampling rate of 32kHz.

[0083] The experimental setup is as follows: the reference audio is uniformly truncated into 10-second segments, and the target duration is synchronously set to 10 seconds; the low-frequency line spectrum cue (LOFAR spectrum) used in the generation process is consistent with the LOFAR spectrum cue extracted from the reference audio itself, in order to verify the model's performance in reproducing local acoustic features.

[0084] To evaluate the generated results, this invention employs three types of indicators for multi-dimensional verification, as detailed below:

[0085] The Wasserstein distance based on the low-frequency line spectrum peak distribution (WD-LOFAR) is used to measure the similarity between the generated signal and the real signal in the frequency evolution behavior of the line spectrum over a long time scale. The calculation process is as follows: the signal is processed in frames, the LOFAR energy peak of each frame is extracted, an empirical distribution is constructed based on the energy peak, and then the 1-Wasserstein distance between the two distributions is calculated. This index can effectively reflect the dynamic characteristics such as the drift characteristics and broadening law of the line spectrum center frequency over time.

[0086] Mel-Cepstral Distance (MCD) is used to evaluate the time-varying matching degree between the generated and reference signals across the overall spectral envelope. The calculation process involves extracting Mel-Cepstral coefficients from both the generated and reference signals, performing dynamic time warping, and then calculating the mean of the frame-level distance to reflect the fidelity of spectral details.

[0087] An underwater target recognition model with ResNet-18 as the backbone network is adopted, and the discriminative consistency of the generated signal is evaluated by the classification accuracy (ACC). That is, it is tested whether the generated sample can be identified as the same ship type by the classifier with a confidence level similar to that of the real sample, thereby verifying the ship type category feature preservation effect of the generated signal.

[0088] Figure 5 The figure illustrates a comparison between reference and generated samples provided in this embodiment of the invention. Since the LOFAR cue signal used in the inference process is directly extracted from the reference signal, an effective model should be able to generate ship radiated noise that retains the inherent spectral structure of the reference signal. As shown in the Mel spectrum comparison results in the figure, the main spectral modes of the generated signal and the reference ship radiated noise are highly consistent. This result indicates that the LOFAR cue signal can effectively control the overall spectral characteristics of the generated signal. Similarly, after extracting the LOFAR spectrum from the generated signal, it exhibits a spectral structure almost identical to the reference signal in terms of frequency position and energy distribution. This result further confirms that the model proposed in this invention can accurately capture the time-varying characteristics of ship radiated noise, ensuring that the dynamic evolution of the generated signal matches the real signal.

[0089] In addition to sample visualization analysis, this embodiment also objectively evaluated the generated samples (i.e., the generated ship radiated noise signals). The evaluation results are shown in Table 1. The LOFAR distribution evolution characteristics of the ship radiated noise generated by this invention are highly consistent with the real signal; the Mel-Cepstral Distance (MCD) index verifies that the spectral envelope of the generated signal has excellent time-varying restoration capability; at the same time, the classification accuracy (ACC) is close to the level of the real reference signal. The above results fully demonstrate that the signal generated by this invention not only completely preserves the unique line spectrum structure and modulation characteristics of ships, but also accurately reproduces the naturally existing time-domain fluctuations under the same operating conditions, achieving effective restoration of the dynamic characteristics of real ship radiated noise.

[0090] WD-LOFAR MCD (dB) ACC (%) Cargo 17.48 3.91 67.10 Passenger 48.65 3.90 59.75 Tanker 19.98 4.06 72.94 Tug 65.52 3.81 70.79 average 33.27 3.95 67.60

[0091] Table 1. Evaluation of objective indicators of generated ship radiated noise signals.

[0092] According to another embodiment, a ship radiated noise generation system based on flow matching is provided, wherein a method for generating ship radiated noise based on flow matching is implemented, the system comprising:

[0093] The feature extraction module is configured to extract low-frequency line spectrum features and mechanical modulation features from the input ship radiated noise signal;

[0094] The neural audio codec module includes a pre-trained neural audio encoder and a neural audio decoder; the encoder is used to encode a ship radiated noise waveform into a hidden layer representation sequence, and the decoder is used to decode the hidden layer representation sequence into a ship radiated noise waveform.

[0095] The flow-matching acoustic model module includes a neural network direction estimator and an ordinary differential equation solver;

[0096] During the training phase, the system is used for:

[0097] The first ship radiated noise signal is processed by the feature extraction module and the neural audio encoder to obtain the first feature and the first hidden layer representation sequence.

[0098] The first hidden layer representation sequence is masked, and the neural network orientation estimator is trained based on the first starting state, the first target state, the first feature, and the time step parameters of the masked region.

[0099] During the inference phase, the system is used for:

[0100] The second ship radiated noise warning signal is processed by the feature extraction module and the neural audio encoder to obtain the second feature and the second hidden layer representation sequence.

[0101] Based on the initial state of Gaussian noise for the target duration, the second hidden layer representation sequence, and the second feature, the trained neural network direction estimator guides the ordinary differential equation solver to generate the third hidden layer representation sequence.

[0102] The neural audio decoder decodes the third hidden layer representation sequence into second ship radiated noise.

[0103] It is understood that the method steps in the embodiments of the present invention can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0104] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0105] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating ship radiated noise based on flow matching, comprising: Collect the first ship radiated noise signal over multiple consecutive time periods; Acquire the second ship's radiated noise warning signal and the duration of the second ship's radiated noise generation; A ship radiated noise generation model is constructed, which includes a feature extraction layer, a neural audio codec layer, and a flow-matching acoustic model layer. The feature extraction layer extracts the first low-frequency line spectrum features and the first mechanical modulation features of the first ship radiated noise signal, and extracts the second low-frequency line spectrum features and the second mechanical modulation features of the second ship radiated noise warning signal; wherein, the second low-frequency line spectrum features are used to determine the line spectrum structure and ship type of the second ship radiated noise, and the second mechanical modulation features are used to determine the modulation characteristics of the second ship radiated noise. The first ship radiated noise signal is encoded into a first hidden layer representation sequence by a pre-trained neural audio encoder in the neural audio encoder-decoder layer, and the second ship radiated noise cue signal is encoded into a second hidden layer representation sequence. In the training phase of the flow-matching acoustic model layer: a first region is randomly masked in the first hidden layer representation sequence, and the masked sequence is used as a context constraint, wherein the original hidden layer representation sequence of the first region is used as the first target state; a standard Gaussian noise of the same length as the first region is generated as the first initial state; linear interpolation is performed on the first initial state and the first target state according to the time step parameters to obtain the current interpolated hidden layer state; the first low-frequency line spectrum feature, the first mechanical modulation feature, the time step parameters, the interpolated hidden layer state, and the context constraint are concatenated and input into the neural network direction estimator for training; the goal of the training is to enable the velocity field predicted by the neural network direction estimator to guide the ordinary differential equation solver to reconstruct the first target state from the first initial state; In the inference phase of the flow-matched acoustic model layer: Gaussian noise with a length corresponding to the generation duration of the second ship radiated noise is used as the second initial state; the second initial state, the second hidden layer representation sequence, the second low-frequency line spectrum feature, and the second mechanical modulation feature are input to the trained neural network direction estimator; the neural network direction estimator guides the ordinary differential equation solver to perform step-by-step integration to obtain the third hidden layer representation sequence; the third hidden layer representation sequence is decoded into the second ship radiated noise by the pre-trained neural audio decoder in the neural audio codec layer.

2. The method of claim 1, wherein, The feature extraction layer includes a low-frequency line spectrum feature extraction layer, which performs a short-time Fourier transform on the ship radiated noise signal to obtain the time spectrum. Each frame of spectrum is standardized and peak points with energy higher than the background are selected; peak points in the preset low frequency band are retained; density clustering is performed on the peak points after selection of multiple frames to remove noise points, forming a LOFAR distribution map with time as the horizontal axis and frequency as the vertical axis, which serves as the low frequency line spectrum feature.

3. The method according to claim 2, wherein, The preset low-frequency band is 1Hz to 600Hz.

4. The method according to claim 1, wherein, The feature extraction layer includes a mechanical modulation feature extraction layer, which performs bandpass filtering on the ship radiated noise signal. The filtered signal is rectified by full-wave rectification to extract the amplitude envelope; the amplitude envelope is low-pass filtered to smooth high-frequency jitter; the smoothed envelope is Fourier transformed and a preset modulation frequency band is extracted to form a DEMON spectrum with time as the horizontal axis and frequency as the vertical axis, which serves as the mechanical modulation feature.

5. The method according to claim 4, wherein, The cutoff frequency of the bandpass filter is 50Hz to 5000Hz; the preset modulation frequency band is 0Hz to 50Hz.

6. The method according to claim 1, wherein, When randomly masking the first hidden layer representation sequence, the masking ratio is dynamically adjusted between 70% and 100%.

7. The method according to claim 1, wherein, The neural network orientation estimator adopts the Transformer model architecture.

8. The method according to claim 1, wherein, The training of the neural audio codec layer employs joint optimization of multi-scale Mel-spectrum loss and adversarial loss; wherein, the adversarial loss is calculated by introducing a multi-scale discriminator that makes judgments in the time domain and at multiple STFT scales, and the multi-scale discriminator adds a denser STFT resolution window in the frequency band below 200Hz.

9. A ship radiated noise generation system based on flow matching, wherein, The system for implementing the method as described in any one of claims 1 to 8 comprises: The feature extraction module is configured to extract low-frequency line spectrum features and mechanical modulation features from the input ship radiated noise signal; The neural audio codec module includes a pre-trained neural audio encoder and a neural audio decoder; the encoder is used to encode a ship radiated noise waveform into a hidden layer representation sequence, and the decoder is used to decode the hidden layer representation sequence into a ship radiated noise waveform. The flow-matching acoustic model module includes a neural network direction estimator and an ordinary differential equation solver; During the training phase, the system is used for: The first ship radiated noise signal is processed by the feature extraction module and the neural audio encoder to obtain the first feature and the first hidden layer representation sequence. The first hidden layer representation sequence is masked, and the neural network orientation estimator is trained based on the first starting state, the first target state, the first feature, and the time step parameters of the masked region. During the inference phase, the system is used for: The second ship radiated noise warning signal is processed by the feature extraction module and the neural audio encoder to obtain the second feature and the second hidden layer representation sequence. Based on the initial state of Gaussian noise for the target duration, the second hidden layer representation sequence, and the second feature, the trained neural network direction estimator guides the ordinary differential equation solver to generate the third hidden layer representation sequence. The neural audio decoder decodes the third hidden layer representation sequence into second ship radiated noise.