A marine-biology-oriented sound extraction enhancement method, system, device and medium

By employing a denoising diffusion probability model and classifier-free guidance technology, the problem of extracting marine biological sounds was solved, enabling the extraction and enhancement of marine biological sounds in complex environments. This improved the application of new technologies in marine biological sound extraction and enhancement, and enhanced the potential for the generalization of marine biological research and conservation technologies in diverse environments. In particular, through the application of diffusion techniques in complex acoustic environments, the extraction and enhancement of marine biological sounds were achieved.

CN120089150BActive Publication Date: 2025-12-23HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510463281.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-12-23
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as poor handling of overlapping audio, scarcity of training data, and insufficient model generalization ability when extracting marine organism sounds in complex acoustic scenarios, resulting in unsatisfactory extraction accuracy and results.

Method used

The Denoising Diffusion Probability Model (DDPM) combined with classifier-free guided (CFG) technology is used to gradually add and remove noise through forward and backward processes, optimize audio feature extraction, and fine-tune the pre-trained model with a small amount of labeled data to improve the model's generalization ability in complex marine environments.

Benefits of technology

It achieves efficient extraction and enhancement of marine biological sounds in complex marine environments, reduces dependence on large-scale labeled data, and improves the model's generalization ability and extraction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089150B_ABST
    Figure CN120089150B_ABST
Patent Text Reader

Abstract

The application discloses a kind of marine biological sound extraction enhancement methods, systems, equipment and medium, belong to intelligent information processing technical field.Method first constructs training dataset by mixed natural background noise, pure biological call and ship noise, uses pre-training model to combine diffusion probability model and carries out iterative optimization, learns the mapping relationship of noise addition and removal.In inference phase, through conditional guidance technology dynamic adjustment model output, accurately separate target biological sound and background noise, and combine verification set and actual collection data to evaluate enhancement effect.The application significantly improves the extraction accuracy and purity of marine biological call in noisy environment, solves the problem of poor adaptability and low efficiency of traditional method, and can be widely used in marine ecological monitoring, species behavior research and underwater acoustic communication engineering, providing efficient and reliable technical support for marine environmental protection and biodiversity research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent information processing technology, specifically relating to a method, system, device, and medium for sound extraction and enhancement for marine organisms. Background Technology

[0002] In the field of Target Sound Extraction (TSE), existing techniques primarily rely on discriminative models, which achieve sound separation by minimizing the difference between the estimated audio and the target audio. However, these methods perform poorly when dealing with overlapping audio regions, especially in complex acoustic scenes where sound overlap is prevalent, leading to a significant decrease in separation efficiency. Furthermore, TSE tasks face the challenge of scarce training data, particularly clean, single-label audio data, which further complicates model training. Therefore, there is an urgent need for a TSE method that can effectively handle complex acoustic scenes and possesses a certain degree of generalization ability to meet diverse real-world needs.

[0003] Currently, the main existing methods for target sound extraction tasks are:

[0004] (1) Grruden, Pina, and Paul R. White. "Automated extraction of dolphinwhistles—A sequential Monte Carlo probability hypothesis density approach." The Journal of the Acoustical Society of America 148.5(2020):3014-3026. This study utilizes a sequential Monte Carlo probability hypothesis density (SMC-PHD) filter to model the extraction of dolphinwhistles as a multi-target tracking problem, estimating the trajectory of dolphinwhistle frequency changes through particle filtering. First, spectral analysis is performed on the acoustic data to extract peak frequencies as measurement data. Then, SMC-PHD is used to recursively track multiple overlapping dolphinwhistle contours, achieving automatic detection and continuous tracking.

[0005] In the aforementioned technique, echo detectors and pulse sounds may be misdetected as dolphin calls, leading to false detections in the presence of other sound interference, thus reducing extraction accuracy. Furthermore, the limited size and insufficient data of manually labeled datasets result in inadequate generalization ability of the model; the type and quality of training data directly affect the learning performance of the RBF motion model.

[0006] (2) Kershenbaum, Arik, and Marie A. Roch. "An image processing based paradigm for the extraction of tonal sounds in cetacean communications." The Journal of the Acoustical Society of America 134.6 (2013):4435-4445. This study proposes an automated algorithm for extracting tonal sounds from cetaceans. The method uses image processing techniques to extract frequency modulations from the spectrogram as "ridges," which represent the basic characteristics of cetacean calls. The algorithm is implemented as a MATLAB script (IPRiT) and can automatically extract frequency variations in cetacean sounds.

[0007] Interference from other sounds in the aforementioned technique may affect the accuracy of the spectrogram, especially under complex background noise, potentially leading to reduced accuracy in frequency ridge extraction. Furthermore, at high noise levels, the algorithm's recognition capability may be limited, and noise interference may prevent the accurate capture of frequency modulation.

[0008] (3) Wang, Xianquan, et al. "A method for enhancement and automated extraction and tracing of Odontoceti whistle signals based on time-frequency spectrogram." Applied Acoustics 176(2021):107698. This study extracts and enhances toothed whale whistle signals from background noise using a time-frequency spectrogram (TFS) image processing method combined with morphological algorithms. First, the sound data is converted into a time-frequency spectrogram using short-time Fourier transform (STFT), and then adaptive thresholding, region growing, and morphological screening are performed to finally separate and enhance the whistle signals.

[0009] Other types of sound interference in the described technique may affect the accuracy of signal extraction, especially in complex underwater environments. Furthermore, insufficient labeled data can lead to difficulties in model training, particularly when signals overlap or are mixed with other sounds, which can affect the model's generalization ability and extraction accuracy. Summary of the Invention

[0010] The purpose of this invention is to provide a method, system, device, and medium for sound extraction and enhancement of marine organisms. It explains the sound extraction process of marine organisms from the perspective of audio features, improves the credibility of audio enhancement results, and has the potential to enhance the generalization ability of the model in diverse marine environments, thus providing reliable technical support for marine biological research and conservation.

[0011] The objective of this invention is achieved through the following technical solution:

[0012] A sound extraction and enhancement method for marine organisms, the specific steps of which are as follows:

[0013] Step 1: Create a dataset by mixing audio data from different datasets, and divide it into a training set and a test set;

[0014] Step 2: Iteratively train the denoised diffusion probability model DDPM based on the pre-trained model, and optimize the parameters of the pre-trained model using the training dataset generated in Step 1;

[0015] Step 3: Construct the Denoising Diffusion Probability Model (DDPM) by gradually adding Gaussian noise to the data through a forward process, with the variance of the noise controlled by the time step sequence.

[0016] Step 4: Generate noise data x at any time step t by linearly combining the clean signal x0 and the noise ∈. t ;

[0017] Step 5: Predict the speed v using a neural network t This is used for gradual noise reduction during the reverse process;

[0018] Step 6: Reverse the process from the noisy data x t Reconstructing the original data x step by step t-1 ;

[0019] Step 7: During the inference process, the model output is adjusted using classifier-free CFG technology to enhance the extraction of the target signal;

[0020] Step 8: Validate the trained model. Use the portion of the dataset generated in Step 1 that was not used in training as the validation set to input the model for inference and obtain the enhanced audio. Then, perform a similarity score (MSE) between the enhanced audio and the pure Whale FM audio of marine life calls before mixing.

[0021] Step 9: Verify the trained model with actual collected data. Input the audio of marine life calls with noise collected in real-world situations into the model for inference. Compare the scores of the NISQA parameters of the obtained results with the scores of the original audio without the processing of this invention. Based on the scores, determine the effect of the actual collected data, and then extract the sounds of marine life.

[0022] Furthermore, the time step sequence followed by the noise variance in step 3 is β1,…,β T Its forward process is defined as:

[0023]

[0024] Where, x t This represents the noise data at time step t; x t-1 This represents the data at time step t; β t It is the noise variance at time step t, which controls the intensity of noise addition; Let I represent a Gaussian distribution; I is the identity matrix.

[0025] Furthermore, the x t for:

[0026]

[0027] Where x0 is the clean, raw signal; α t =1-β t Indicates the noise attenuation coefficient; This represents the cumulative noise attenuation coefficient; It is standard Gaussian noise;

[0028] To optimize the consistency of the training and inference processes, noise scheduling is adjusted to maintain... Unchanged, will Set it to zero, and for intermediate time steps t∈[2,…,T-1] Perform linear scaling.

[0029] Furthermore, the step v t for:

[0030]

[0031] Among them, v t x0 represents the speed of model prediction, used to improve the purity of sound extraction; x0 is the clean, raw signal. It is standard Gaussian noise.

[0032] Furthermore, the probability distribution of the reverse process in step 6 is as follows:

[0033] in, It is the mean of the model predictions, calculated using the following formula: It is the variance of the model predictions, calculated using the following formula: x0 is the clean, raw signal, which can be obtained through the formula From noise data x t In the middle estimation; θ represents the model parameters.

[0034] Furthermore, in step 7, the input is made more consistent with the given conditions:

[0035]

[0036] in, It is unconditional sampling prediction; It is a conditional sampling prediction; γ is the guiding scale used to control the strength of the conditions.

[0037] Furthermore, the similarity score MSE in step 8 is:

[0038]

[0039] Among them, y i These are the sounds of pure marine life. This is the enhanced audio from this invention, where n is the length of the audio and MSE is the similarity score.

[0040] A computer device / apparatus / system includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of a sound extraction and enhancement method for marine organisms.

[0041] A computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of a sound extraction and enhancement method for marine organisms.

[0042] An electronic device, characterized in that it comprises:

[0043] Memory, used to store computer programs;

[0044] A processor for executing the computer program to implement the instruction tracking method described in the sound extraction and enhancement method for marine organisms.

[0045] The beneficial effects of this invention are as follows:

[0046] This invention proposes an innovative method for sound extraction and enhancement targeting marine organisms. It utilizes a diffusion model architecture combined with language-guided audio feature extraction techniques to achieve efficient extraction and enhancement of marine biological sounds (such as whale calls). Unlike traditional discriminative model-based sound extraction methods, this method employs a denoised diffusion probability model (DDPM) operating within the latent space of the audio, avoiding the quality limitations of traditional spectrogram reconstruction and improving the model's generalization ability in complex marine acoustic environments.

[0047] Furthermore, this model employs a few-shot learning strategy, fine-tuning it under conditions of scarce labeled data to improve the extraction and enhancement of marine biological sounds. Addressing challenges such as weak signals, high noise levels, and difficult annotation in the marine environment, this method utilizes the prior knowledge of a pre-trained model and fine-tunes it based on a limited amount of labeled data, enabling the model to more accurately identify and enhance the sounds of target organisms. This strategy reduces reliance on large-scale labeled data, allowing the model to achieve efficient and accurate extraction and enhancement of marine biological sounds even with limited data resources. Attached Figure Description

[0048] Figure 1 A process for sound extraction and enhancement for marine organisms.

[0049] Figure 2 This is a histogram showing the interval distribution of MSE data. Detailed Implementation

[0050] The present invention will now be further described with reference to the accompanying drawings.

[0051] This invention presents a sound extraction and enhancement method for marine organisms, achieving efficient extraction and enhancement of marine organism sounds (such as whale calls) based on a diffusion model. First, a denoising diffusion probability model (DDPM) is constructed, adding noise through a forward pass and gradually reconstructing the original audio signal through a backward pass, utilizing velocity prediction to improve the purity of the extracted sound. Next, a SoloAudio model architecture is introduced, including a VAE encoder and decoder, a CLAP model, and a DiT model, used for extracting latent audio representations, generating reference embeddings, and processing latent features, respectively. Second, long skip connections and rotated position embeddings (RoPE) are introduced into the DiT model to enhance low-level feature transfer and positional encoding, simplifying the training process. Then, classifier-free guided inference (CFG) technology is used during inference, controlling the conditional strength of the model output by adjusting the guided scale. Finally, after multiple steps of denoising and sampling, a clean target sound is generated, achieving efficient sound extraction and enhancement in complex acoustic scenes.

[0052] according to Figure 1 The specific steps are as follows:

[0053] Step 1: Generate the required dataset. The dataset used is a random combination of audio from different datasets, including natural underwater noise at a depth of 7.5 meters as background noise (ShipsEarNO.82), pure marine life calls audio (Whale FM), and random audio from three other types of ships (DeepShip). The mixed audio is 10 seconds long. The mixing method is to superimpose pure marine life calls and ship audio noise at random time points within the 10-second audio, based on the background noise.

[0054] Step 2: Use the generated dataset for training. First, load the existing pre-trained model and load the weights. Randomly select a small amount of the generated dataset to input into the network. Train the original model for 20 rounds with a learning rate of 1E-05 and a weight decay of 1E-04.

[0055] Step 3: Construct a Denoising Diffusion Probability Model (DDPM) by progressively adding Gaussian noise to the data through a forward process. The variance of the noise follows the pattern β1,…,β T :

[0056]

[0057] Where, x t This represents the noise data at time step t; x t-1 This represents the data at time step t; β t It is the noise variance at time step t, which controls the intensity of noise addition; Let I represent a Gaussian distribution; I is the identity matrix.

[0058] The main purpose of the forward pass is to gradually transform the clean signal x0 into noisy data x. T This provides the foundation for the reverse process (denoising process). Through this process, the model can learn how to gradually recover a clean signal from noisy data during the reverse process.

[0059] Step 4: Sample noise data and generate noise data x at any time step t by linearly combining the clean signal x0 and the noise ∈. t The formula is:

[0060]

[0061] Where x0 is the clean, raw signal; α t =1-β t Indicates the noise attenuation coefficient; This represents the cumulative noise attenuation coefficient; It is standard Gaussian noise.

[0062] To optimize the consistency of the training and inference processes, noise scheduling is adjusted to maintain... Unchanged, will Set it to zero, and for intermediate time steps t∈[2,…,T-1] Linear scaling is applied. This adjustment helps prevent the introduction of additional noise during sampling and improves the model's denoising performance.

[0063] Step 5: Next, predict the speed v using a neural network. t Used for progressive noise reduction during the reverse process:

[0064]

[0065] Among them, v t x0 represents the speed of model prediction, used to improve the purity of sound extraction; x0 is the clean, raw signal. It is standard Gaussian noise.

[0066] Step 6: Reconstruct the data using the reverse process, from the noisy data x t Reconstructing the original data x step by step t-1 :

[0067]

[0068] in, It is the mean of the model predictions, calculated using the following formula:

[0069] It is the variance of the model predictions, calculated using the following formula:

[0070] x0 is the clean, raw signal, which can be obtained through the formula From noise data x t In the middle estimation; θ represents the model parameters.

[0071] Step 7: Build the model and adjust its output during inference using classifier-free bootstrapping (CFG) to better match the conditional input:

[0072]

[0073] in, It is unconditional sampling prediction; It is a conditional sampling prediction; γ is the guiding scale used to control the strength of the conditions.

[0074] Step 8: Validate the trained model on the validation set. Use the portion of the dataset generated in Step 1 that was not used in the training to input the model for inference and obtain the enhanced result obtained by the present invention. Evaluate the similarity (MSE) between this result and the pure marine life call audio (Whale FM) before mixing. Determine the effect of the present invention on the validation set based on the similarity with the corresponding pure marine life call.

[0075]

[0076] Among them, y i These are the sounds of pure marine life. This is the enhanced audio from this invention, where n is the length of the audio and MSE is the similarity score.

[0077] Step 9: Verify the trained model with actual collected data. Input the audio of marine life calls with noise collected in real-world situations into the model for inference. Compare the scores of the NISQA parameters of the obtained results with the scores of the original audio without the processing of this invention. Based on the scores, determine the effectiveness of this invention in actual collected data, and then extract the sounds of marine life.

[0078] Example 1:

[0079] The effectiveness of this invention is evaluated by assessing the noise index (NISQA) of the extracted audio and the similarity (MSE) between the extracted audio and the target audio. After sound extraction from the dataset, audio quality and naturalness are used for evaluation. This invention evaluated two types of datasets: one type consists of a 10-second dataset containing natural underwater current noise at a depth of 7.5 meters (ShipsEar NO.82) mixed with clean marine life calls (Whale FM) and random audio from three other types of vessels (DeepShip); the other type contains marine life calls with added noise (including ship sounds, water current sounds, etc.) (Whale FM).

[0080] First, mean square error (MSE) was evaluated on the target sound of the first audio and the mixed audio after the sound was extracted by this method. The results of the 101 audios after extraction were compared with the pure marine life call audio (WhaleFM) before mixing.

[0081] like Figure 2 As shown, after processing with the present invention, the mean square error of the 101 test data of type 1 with the original pure marine life call audio (Whale FM) was mostly (92%) less than 2.43E-04, indicating that the present invention is very effective in extracting marine life calls.

[0082] Table 1. Changes in scores after extraction of noisy marine life call audio (Whale FM)

[0083]

[0084] As shown in Table 1, after the extraction and enhancement processing of 194 noisy marine animal calls using this invention, there was a significant improvement in the NISQA noise index score, indicating that this invention has a certain generalization ability for noise distributed in its natural state. Other indicators are for reference only, as NISQA is not a specific evaluation model for marine animal sounds.

[0085] This invention was derived by fine-tuning the original model in 20 steps using 19 data samples with a learning rate of 1E-05 and a weight decay of 1E-04. Based on the analysis of the two test data, it can be seen that this invention can almost perfectly filter noise similar to that used in training. However, it may not be very effective for data with a distribution that is very different from that used in training. Therefore, if there is a large amount of relevant noise data to support it, the ability of this invention to deal with natural noise can be further enhanced.

[0086] All relevant content of each step involved in the aforementioned embodiment of a sound extraction and enhancement method for marine organisms can be referenced to the functional description of the corresponding functional module of the deepfake detection interpretable system based on causal analysis in the embodiments of the present invention, and will not be repeated here.

[0087] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0088] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used in the operation of a sound extraction and enhancement method for marine organisms.

[0089] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the operating system of the terminal. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the sound extraction and enhancement method for marine organisms described in the above embodiments.

[0090] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0091] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0092] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0093] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0094] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A sound extraction and enhancement method for marine organisms, characterized in that... The specific steps are as follows: Step 1: Construct a dataset by mixing audio from different datasets, including natural underwater noise as background noise, pure marine life calls, and random audio from three types of ships, and divide it into training set and test set. Step 2: Iteratively train the denoised diffusion probability model DDPM based on the pre-trained model, and optimize the parameters of the pre-trained model using the training dataset generated in Step 1; Step 3: Construct the Denoising Diffusion Probability Model (DDPM) by progressively adding standard Gaussian noise to the data through a forward process. The variance of the noise is controlled by the time step sequence. Step 4: Clean the original signal through linear combination and standard Gaussian noise Generate arbitrary time steps noise data ; Step 5: Predict speed using a neural network This is used for gradual noise reduction during the reverse process; ; in, It refers to the speed of model prediction, which is used to improve the purity of sound extraction; It is a clean, raw signal; It is standard Gaussian noise. It is the identity matrix; Indicates the noise attenuation coefficient; This represents the cumulative noise attenuation coefficient; Step 6: Reverse the process from noisy data Gradually reconstruct the original data ; Step 7: During the inference process, the model output is adjusted using classifier-free CFG technology to enhance the extraction of the target signal; The CFG technology adjustment must meet the following conditions: ; in, It is unconditional sampling prediction; It is a conditional sampling prediction; It is a guiding scale used to control the intensity of conditions; Step 8: Validate the trained model. Use the portion of the dataset generated in Step 1 that was not used in training as the validation set to input the model for inference and obtain the enhanced audio. Then, perform a similarity score (MSE) between the enhanced audio and the pure marine animal call audio before mixing. Step 9: Validate the trained model with real-world data collection. Input real-world audio recordings of marine life calls with noise into the model for inference. Compare the NISQA scores of the resulting data with the scores of the original, unprocessed audio recordings. Determine the effectiveness of the real-world data collection based on the scores, and then extract the sounds of marine life.

2. The sound extraction and enhancement method for marine organisms according to claim 1, characterized in that... The time step sequence followed by the noise variance in step 3 is as follows: Its forward process is defined as: ; in, Indicates at time step Noise data; It is a time step The noise variance is used to control the added intensity of noise; This indicates a Gaussian distribution.

3. The sound extraction and enhancement method for marine organisms according to claim 2, characterized in that... The above for: ; To optimize the consistency of the training and inference processes, noise scheduling is adjusted to maintain... Unchanged, will Set to zero, and set the intermediate time step. of Perform linear scaling.

4. The sound extraction and enhancement method for marine organisms according to claim 3, characterized in that... The probability distribution of the reverse process in step 6 is as follows: ; in, It is the mean of the model predictions, calculated using the following formula: ; It is the variance of the model predictions, calculated using the following formula: ; It is a clean, raw signal, which can be obtained through the formula. From noise data Medium estimate; Indicates model parameters.

5. The sound extraction and enhancement method for marine organisms according to claim 1, characterized in that: The similarity score MSE in step 8 is: ; in These are the sounds of pure marine life. This is the enhanced audio, where n is the length of the audio and MSE is the similarity score.

6. A computer system comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Audio generation acceleration method based on reasoning integrated into training process

    CN117765923A

  • Speech enhancement acceleration method based on de-noising probability diffusion model

    CN117975980A