Marine organism-oriented sound extraction enhancement method, system, equipment and medium
By applying the denoising diffusion probability model (DDPM) and language-oriented audio feature extraction technology in marine biological sound extraction, the problem of difficulty in extracting marine biological sounds in complex acoustic scenes is solved, efficient and accurate sound extraction and enhancement is achieved, and the generalization ability of the model is improved.
Patent Information
- Application Number
- CN202510463281.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art is difficult to effectively extract and separate the sounds of marine organisms when dealing with complex acoustic scenarios, especially in marine environments, and faces the problem of scarce training data, which leads to insufficient generalization capabilities of the model.
Denoising diffusion probability model (DDPM) combined with language-oriented audio feature extraction technology is used to improve the generalization ability of the model in complex marine acoustic environments through iterative training and small sample learning strategies, and the model output is adjusted through classifier-free boot (CFG) technology to enhance the extraction of target signals.
It realizes efficient extraction and enhancement of marine biological sounds, improves the generalization ability of the model in complex acoustic environments, reduces the dependence on large-scale annotation data, and enables the model to achieve efficient and accurate extraction and enhancement of marine biological sounds when data resources are limited.
Smart Images

Figure CN120089150A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent information processing, and particularly relates to a method, system, device and medium for sound extraction and enhancement for marine organisms. Background Art
[0002] In the field of Target Sound Extraction (TSE), existing technologies mainly rely on discriminant models, which achieve sound separation by minimizing the difference between the estimated audio and the target audio. However, these methods perform poorly in processing overlapping audio regions, especially in complex acoustic scenarios where sound overlap is prevalent, resulting in a significant decline in separation effectiveness. Additionally, the target sound extraction task faces the problem of scarce training data, especially clean single-label audio data, which further exacerbates the training difficulty of the model. Therefore, there is an urgent need for a target sound extraction method that can effectively handle complex acoustic scenarios and has a certain generalization ability to meet the diverse needs in the real world.
[0003] Currently, regarding the target sound extraction task, the existing related methods mainly include:
[0004] (1) Gruden, Pina, and Paul R. White. "Automated extraction of dolphin whistles—A sequential Monte Carlo probability hypothesis density approach." The Journal of the Acoustical Society of America 148.5 (2020): 3014 - 3026. This study uses the Sequential Monte Carlo Probability Hypothesis Density (SMC-PHD) filter to model the extraction of dolphin whistles as a multi-target tracking problem, and estimates the changing trajectory of the whistle frequency through a particle filtering method. First, perform spectral analysis on the acoustic data, extract the peak frequency as measurement data, and then use SMC-PHD to recursively track multiple overlapping whistle contours to achieve automatic detection and continuous tracking.
[0005] In the described technology, echo detectors and impulse sounds may be misdetected as dolphin whistles, resulting in false detections in the presence of other sound interferences, thereby reducing the extraction accuracy. In addition, the scale of the manually annotated dataset is limited and the data is insufficient, resulting in insufficient generalization ability of the model. The type and quality of the training data directly affect the learning effect of the RBF motion model.
[0006] (2) Kershenbaum, Arik, and Marie A. Roch. "An image processing based paradigm for the extraction of tonal sounds in cetacean communications." The Journal of the Acoustical Society of America 134.6 (2013): 4435-4445. This study proposed an automated algorithm for extracting tonal sound waves of cetaceans. Through image processing techniques, this method extracts the frequency modulation in spectrograms as "ridge" structures, and these ridges represent the basic characteristics of cetacean calls. The algorithm was implemented as a MATLAB script (IPRiT), which can automatically extract the frequency variations of cetacean sounds.
[0007] Interference from other sounds in the described technology may affect the accuracy of spectrograms. Especially in complex background noise, it may lead to a reduction in the accuracy of frequency ridge extraction. In addition, when the noise level is high, the recognition ability of the algorithm may be limited, and noise interference may prevent the frequency modulation from being accurately captured.
[0008] (3) Wang, Xianquan, et al. "A method for enhancement and automated extraction and tracing of Odontoceti whistle signals base on time-frequency spectrogram." Applied Acoustics 176 (2021): 107698. This study extracts and enhances Odontoceti whistle signals from background noise through a processing method based on time-frequency spectrogram (TFS) images, combined with morphological algorithms. First, the sound data is converted into a time-frequency spectrogram using the short-time Fourier transform (STFT), and then adaptive threshold processing, region growing, and morphological screening are performed. Finally, the whistle signals are separated and enhanced.
[0009] Interference from other types of sounds in the described technology may affect the accuracy of signal extraction, especially in complex underwater environments. In addition, insufficient labeled data may lead to difficulties in model training, especially when signals overlap or are mixed with other sounds, which may affect the generalization ability and extraction accuracy of the model. Summary of the Invention
[0010] The object of the present invention is to provide a method, system, device and medium for sound extraction and enhancement for marine organisms, to explain the process of marine organism sound extraction from the perspective of audio features, improve the credibility of audio enhancement results, and have the potential to improve the generalization ability of the model in diverse marine environments, providing reliable technical support for marine organism research and protection.
[0011] The object of the present invention is achieved by the following technical solutions:
[0012] A method for sound extraction and enhancement for marine organisms, the specific steps are as follows:
[0013] Step 1: Combine the audio of different data sets to form a data set, and divide it into a training set and a test set;
[0014] Step 2: Perform iterative training based on the pre-trained denoising diffusion probabilistic model DDPM, and use the training data set generated in Step 1 to optimize the parameters of the pre-trained model;
[0015] Step 3: Construct a denoising diffusion probabilistic model DDPM, and gradually add Gaussian noise to the data through the forward process, and the variance of the noise is controlled by the time step sequence;
[0016] Step 4: Generate the noise data x 0 at any time step t by linearly combining the clean signal x t and the noise ∈;
[0017] Step 5: Predict the velocity v t through a neural network for gradual denoising in the reverse process;
[0018] Step 6: Gradually reconstruct the original data x t from the noise data x t-1 through the reverse process;
[0019] Step 7: Adjust the model output through the classifier-free guidance CFG technique during the inference process to enhance the extraction of the target signal;
[0020] Step 8: Verify the trained model, use the part of the data set generated in Step 1 that is not involved in training as the verification set to input into the model for inference, obtain the enhanced audio, and perform a similarity score MSE on the enhanced audio and the pure marine organism call audio Whale FM before mixing;
[0021] Step 9: Validate the trained model with actual collected data. Input the noisy audio of marine creature calls collected in reality into the model for inference. Compare the scores of each parameter of NISQA in the obtained results with the scores of each parameter of NISQA of the original audio without being processed by the present invention. Based on the high and low scores, obtain the effect of the actual collected data, and further extract the sounds of marine creatures.
[0022] Further, the time step sequence followed by the noise variance in step 3 is β 1 ,…,β T , and its forward process is defined as:
[0023]
[0024] where x t represents the noise data at time step t; x t-1 represents the data at time step t; β t is the noise variance at time step t, controlling the intensity of noise addition; represents the Gaussian distribution; I is the identity matrix.
[0025] Further, the x t is:
[0026]
[0027] where x 0 is the clean original signal; α t = 1 - β t represents the noise attenuation coefficient; represents the cumulative noise attenuation coefficient; is the standard Gaussian noise;
[0028] To optimize the consistency of the training and inference processes, adjust the noise schedule, keep unchanged, set to zero, and linearly scale for the intermediate time steps t ∈ [2,…, T - 1].
[0029] Further, the step v t is:
[0030]
[0031] where v t is the speed predicted by the model, used to improve the purity of sound extraction; x 0 is the clean original signal; is the standard Gaussian noise.
[0032] Further, the probability distribution of the reverse process in step 6 is as follows:
[0033] where is the mean predicted by the model, and the calculation formula is: is the variance predicted by the model, and the calculation formula is: x 0 is the clean original signal, which can be obtained through the formula estimated from the noisy data x t ; θ represents the model parameters.
[0034] Further, in step 7, make it more conform to the conditional input:
[0035]
[0036] where is the unconditional sampling prediction; is the conditional sampling prediction; γ is the guidance scale for controlling the conditional intensity.
[0037] Further, the similarity score MSE in step 8 is:
[0038]
[0039] where y i is the corresponding pure marine creature call, is the enhanced audio of the present invention, n is the length of the audio, and MSE is the similarity score.
[0040] A computer device / system, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of a method for extracting and enhancing sounds of marine creatures.
[0041] A computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps of a method for extracting and enhancing sounds of marine creatures are implemented.
[0042] An electronic device, characterized by comprising:
[0043] A memory for storing a computer program;
[0044] A processor for executing the computer program to implement the instruction tracking method of a method for extracting and enhancing sounds of marine creatures.
[0045] The beneficial effects of the present invention are as follows:
[0046] The present invention proposes an innovative method for extracting and enhancing the sounds of marine organisms, which uses a diffusion model architecture and combines language-guided audio feature extraction techniques to achieve efficient extraction and enhancement of the sounds of marine organisms (such as whale calls). Different from traditional discriminative model-based sound extraction methods, this method uses the denoising diffusion probabilistic model (DDPM) to operate in the latent space of audio, avoiding the quality limitations of traditional spectrogram reconstruction and improving the generalization ability of the model in complex marine acoustic environments.
[0047] In addition, this model adopts a few-shot learning strategy and is fine-tuned in the case of scarce labeled data to improve the effect of extracting and enhancing the sounds of marine organisms. In response to challenges such as weak signals, strong noise, and difficult annotation in the marine environment, this method uses the prior knowledge of the pre-trained model and is fine-tuned based on a small amount of labeled data, enabling the model to more accurately identify and enhance the target organism sounds. This strategy reduces the dependence on large-scale labeled data and enables the model to still achieve efficient and accurate extraction and enhancement of the sounds of marine organisms under limited data resources. Brief Description of the Drawings
[0048] Figure 1 It is the process flow of the method for extracting and enhancing the sounds of marine organisms.
[0049] Figure 2 It is the histogram of the MSE data interval distribution. Detailed Embodiment
[0050] The following further describes the present invention in conjunction with the drawings.
[0051] A method for extracting and enhancing the sounds of marine organisms according to the present invention is based on a diffusion model to achieve efficient extraction and enhancement of the sounds of marine organisms (such as whale calls). First, a denoising diffusion probabilistic model (DDPM) is constructed, noise is added through the forward process, and the original audio signal is gradually reconstructed through the reverse process, and the purity of sound extraction is improved by using speed prediction; then, the SoloAudio model architecture is introduced, including a VAE encoder and decoder, a CLAP model, and a DiT model, which are used to extract audio latent representations, generate reference embeddings, and process latent features respectively; secondly, long skip connections and rotary position embeddings (RoPE) are introduced into the DiT model to enhance low-level feature transmission and position encoding and simplify the training process; then, the classifier-free guidance (CFG) technique is used in the inference process to control the conditional intensity of the model output by adjusting the guidance scale; finally, after multi-step denoising sampling, clean target sounds are generated to achieve efficient sound extraction and enhancement in complex acoustic scenarios.
[0052] According to Figure 1 , the specific steps are as follows:
[0053] Step 1: Generate the required dataset. The dataset used is a random combination of audio from different datasets, including the natural noise of water flow at a depth of 7.5 meters underwater (ShipsEar NO.82) as the background noise, pure marine biological call audio (Whale FM), and audio of 3 other types of ships randomly selected (DeepShip); the length of the mixed audio is 10 seconds, and the mixing method is based on the background noise, with pure marine biological calls and ship audio noise superimposed at random time points within the 10-second audio.
[0054] Step 2: Use the generated dataset for training. First, load the existing pre-trained model and load the weights. Randomly select a small amount of the generated dataset and input it into the network. The original model was trained for 20 epochs with a learning rate of 1E-05 and a weight decay of 1E-04.
[0055] Step 3: Construct a denoising diffusion probabilistic model (DDPM). Gradually add Gaussian noise to the data through the forward process, and the variance of the noise follows β 1 ,…,β T :
[0056]
[0057] where x t represents the noisy data at time step t; x t-1 represents the data at time step t; β t is the noise variance at time step t, controlling the intensity of noise addition; represents the Gaussian distribution; I is the identity matrix.
[0058] The main purpose of the forward process is to gradually transform the clean signal x 0 into the noisy data x T , thus providing a basis for the reverse process (denoising process). Through this process, the model can learn how to gradually recover the clean signal from the noisy data in the reverse process.
[0059] Step 4: Sample the noisy data. Generate the noisy data x 0 at any time step t through a linear combination of the clean signal x t and the noise ∈, and the formula is:
[0060]
[0061] where x 0 is the clean original signal; α t = 1 - β t represents the noise attenuation coefficient; represents the cumulative noise attenuation coefficient; is the standard Gaussian noise.
[0062] To optimize the consistency of the training and inference processes, the noise schedule is adjusted, keeping unchanged, setting to zero, and linearly scaling for intermediate time steps \(t\in[2,\ldots,T - 1]\). This adjustment helps prevent the introduction of additional noise during sampling and improves the denoising effect of the model.
[0063] Step 5: Then, the velocity \(v\) t is predicted through the neural network for progressive denoising during the reverse process:
[0064]
[0065] where \(v\) t is the velocity predicted by the model for improving the purity of sound extraction; \(x\) 0 is the clean original signal; is the standard Gaussian noise.
[0066] Step 6: The data is reconstructed in the reverse process, progressively reconstructing the original data \(x\) t from the noisy data \(x\) t-1 :
[0067]
[0068] where is the mean predicted by the model, calculated as:
[0069] is the variance predicted by the model, calculated as:
[0070] \(x\) 0 is the clean original signal, which can be estimated from the noisy data \(x\) using the formula t ; \(\theta\) represents the model parameters.
[0071] Step 7: The model is constructed, and during the inference process, the model output is adjusted through classifier-free guidance (CFG) technology to make it more consistent with the conditional input:
[0072]
[0073] where is the unconditional sampling prediction; is the conditional sampling prediction; \(\gamma\) is the guidance scale for controlling the conditional strength.
[0074] Step 8: Validate the trained model using the validation set. Use the part of the dataset generated in Step 1 that was not involved in training to input into the model for inference, obtaining the enhanced result of the present invention. Evaluate the similarity (MSE) between this result and the pure marine creature call audio (Whale FM) before mixing, and obtain the effect of the present invention on the validation set based on the similarity to the corresponding pure marine creature calls.
[0075]
[0076] Among them, y i is the corresponding pure marine creature call, is the audio enhanced by the present invention, n is the length of the audio, and MSE is the similarity score.
[0077] Step 9: Validate the trained model using the actually collected data. Input the audio of marine creature calls with noise in the real situation collected in reality into the model for inference, compare the scores of each parameter of NISQA of the obtained result with the scores of each parameter of NISQA of the original audio without being processed by the present invention, obtain the effect of the present invention on the actually collected data based on the score, and then extract the sound of marine creatures.
[0078] Example 1:
[0079] The effect of the present invention is evaluated by the noise index (NISQA) of the extracted audio and the similarity (MSE) between the extracted audio and the target audio. After extracting the sound from the audio of the dataset, audio quality and naturalness are evaluated. The present invention evaluated two types of datasets. One type is the pure marine creature call audio (Whale FM) mixed with the natural noise of the water flow at a depth of 7.5 meters underwater (ShipsEarNO.82) and the audio of 3 other random types of ships (DeepShip) with a length of 10 seconds. The other type is the audio of marine creature calls (Whale FM) with noise (including ship sounds, water flow sounds, etc.).
[0080] First, evaluate the mean square error (MSE) between the target sound of the first type of audio and the mixed audio after extracting the sound by this method. Compare the results of 101 audio after extraction with the pure marine creature call audio (WhaleFM) before mixing.
[0081] As Figure 2 shown, for 101 pieces of type 1 test data after being processed by the present invention, the mean square error with the original pure marine creature call audio (Whale FM) is less than 2.43E - 04 for the vast majority (92%), indicating that the present invention is very effective in extracting the calls of marine creatures.
[0082] Table 1 Changes in the Extracted Scores of Noisy Marine Creature Call Audio (Whale FM)
[0083]
[0084] As shown in Table 1, after the extraction and enhancement processing of the present invention on 194 noisy marine creature call audios, there is a relatively obvious improvement in the NISQA noise index score, indicating that the present invention also has a certain generalization ability for the noise distributed in the natural state. For other indicators, since NISQA is not an evaluation model exclusive to marine creature sounds, it is for reference only.
[0085] The present invention is obtained by fine-tuning the original model for 20 steps with 19 data samples at a learning rate of 1E-05 and a weight decay of 1E-04. Combining the analysis of the above two test data, it can be seen that the present invention can almost perfectly filter the noise similar to the training, while the effect may not be very obvious for the data with a very different distribution from the training noise. Therefore, if there is a large amount of relevant noise data support, the ability of the present invention to cope with natural noise can be further enhanced.
[0086] All relevant contents of each step involved in the foregoing embodiment of the method for enhancing sound extraction for marine creatures can be cited in the functional description of the functional modules corresponding to the explainable system for deepfake detection based on causal analysis in the embodiments of the present invention, and will not be elaborated here.
[0087] The division of modules in the embodiments of the present invention is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, the functional modules can be integrated in one processor, or can exist separately physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0088] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of a method for enhancing sound extraction for marine organisms.
[0089] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for enhancing sound extraction for marine organisms in the above embodiment.
[0090] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0091] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0094] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still, modifications can be made to the specific embodiments of the present invention or equivalent substitutions can be made, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A sound extraction and enhancement method for marine organisms, characterized in that The specific steps are as follows: Step 1: Mix audio from different datasets to form a dataset, and divide it into a training set and a test set; Step 2: Perform iterative training based on the pre-trained denoising diffusion probability model DDPM, and use the training data set generated in step 1 to optimize the parameters of the pre-trained model; Step 3: Construct a denoising diffusion probability model DDPM, and gradually add Gaussian noise to the data through a forward process. The variance of the noise is controlled by the time step sequence. Step 4: Generate noisy data x at any time step t by linearly combining the clean signal x0 and the noise ∈ t ; Step 5: Predict the speed v through the neural network t , used for progressive denoising in the reverse process; Step 6: Reverse the process from the noisy data x t Reconstruct the original data x step by step t-1 ; Step 7: Adjust the model output during inference by using the classifier-free guided CFG technique to enhance the extraction of target signals; Step 8: Verify the trained model, use the part of the data set generated in step 1 that was not involved in the training as the verification set to input the model for inference, and obtain the enhanced audio. The enhanced audio is then compared with the pure marine life call audio Whale FM before mixing to perform a similarity score MSE. Step 9: The trained model is verified by actual collected data. The audio of marine life calls with noise collected in real life is input into the model for inference. The scores of the NISQA parameters of the obtained results are compared with the scores of the NISQA parameters of the original audio that has not been processed by the present invention. The effect of the actual collected data is obtained according to the scores, and the sounds of marine life are extracted.
2. The method for sound extraction and enhancement for marine organisms according to claim 1, characterized in that : The time step sequence followed by the noise variance in step 3 is β1,…,β T , and its forward process is defined as: Among them, x t represents the noise data at time step t; x t-1 represents the data at time step t; β t is the noise variance at time step t, which controls the intensity of noise addition; represents Gaussian distribution; I is the identity matrix.
3. The method for sound extraction and enhancement for marine organisms according to claim 2, characterized in that : The x t for: Among them, x0 is the clean original signal; α t =1-β t Represents the noise attenuation coefficient; represents the cumulative noise attenuation coefficient; is standard Gaussian noise; In order to optimize the consistency of training and inference processes, the noise schedule is adjusted to maintain No change, will Set to zero and for the intermediate time steps t∈[2,…,T-1] Perform linear scaling.
4. The method for sound extraction and enhancement for marine organisms according to claim 1, characterized in that : The step v t for: Among them, v t is the speed predicted by the model, which is used to improve the purity of sound extraction; x0 is the clean original signal; is standard Gaussian noise.
5. A method for sound extraction and enhancement for marine organisms according to claim 1, characterized in that : The probability distribution of the reverse process in step 6 is: in, is the mean of the model predictions, calculated as: is the variance of the model prediction, calculated as: x0 is the clean original signal, which can be obtained by the formula From the noisy data x t is estimated in the middle; θ represents the model parameters.
6. A method for sound extraction and enhancement for marine organisms according to claim 1, characterized in that :In step 7, enter the following to make it more suitable: in, is an unconditional sample prediction; is the conditional sample prediction; γ is the bootstrap scale used to control the conditional strength.
7. A method for sound extraction and enhancement for marine organisms according to claim 1, characterized in that : The similarity score MSE in step 8 is: Among them, y i Is the corresponding pure sound of marine life, is the audio enhanced by the present invention, n is the length of the audio, and MSE is the similarity score.
8. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the instruction tracing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Audio generation acceleration method based on reasoning integrated into training process
CN117765923A
Speech enhancement acceleration method based on de-noising probability diffusion model
CN117975980A
Fingerprint image restoration method based on conditional diffusion probability model
CN119784628A
Cited By
Dolphin signal identification method and device based on diffusion model
CN119993169A
Propagation source positioning method and device
CN120342782A