Voice processing method, medium, electronic device and program product

By constructing a pre-defined linear relationship between clean speech and noisy speech in the frequency domain, and using a diffusion model for speech denoising, the problem of misidentifying valid speech as noise in existing technologies is solved, achieving efficient denoising without damaging speech data.

CN120913581AActive Publication Date: 2025-11-07HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202410544667.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-11-07
Estimated Expiration
2044-04-30

AI Technical Summary

Technical Problem

Existing speech denoising models are prone to misidentifying effective speech as noise during training, resulting in damage to effective speech during the denoising process, and their generalization performance and computational efficiency are insufficient.

Method used

A diffusion model based on linear trajectory is used for speech denoising. By constructing a preset linear relationship between clean speech and noisy speech in the frequency domain, the diffusion model is used for speech denoising to ensure the continuity and clarity of speech data during the denoising process.

Benefits of technology

It improves the generalization performance and computational efficiency of speech denoising without compromising speech data, and can remove noise quickly and effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913581A_ABST
    Figure CN120913581A_ABST
Patent Text Reader

Abstract

The invention relates to the field of voice processing, and discloses a voice processing method, a medium, electronic equipment and a program product, voice de-noising processing is carried out based on a diffusion model of ODE of a linear track, voice data cannot be damaged, the generalization performance is high, the reasoning process speed is high, and the calculation efficiency is high. Specifically, in the de-noising reasoning process, correction information of noisy voice data can be reasoned through a diffusion model, the correction information of the noisy voice data under each time step is specifically determined, and corresponding de-noised voice data is determined through the correction information. And then, through one or more time steps, taking the determined de-noised voice data as corresponding clean voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a speech processing method, medium, electronic device and program product. Background Technology

[0002] With the widespread adoption of smart devices such as smartphones and smart speakers, voice interaction scenarios, including voice calls and human-computer interaction, are becoming increasingly common. Typically, in voice interaction scenarios, smart devices use microphones to capture speech, but this voice is often interfered with by noise, affecting the call experience and the quality of human-computer interaction. For example, during a voice call, the microphone may capture not only valid speech but also environmental noise such as car horns, making it difficult for the user to clearly hear the other person's conversation. Similarly, in human-computer interaction scenarios, the presence of noise can reduce the device's voice wake-up rate and recognition rate.

[0003] Currently, such as Figure 1 As shown, traditional methods employ discriminative speech denoising models to denoise the speech captured by the microphone. Specifically, in Figure 1 In the training process of the illustrated speech denoising model, noisy training speech 1 can be input into the neural network for denoising processing, outputting the corresponding denoised training speech 3. Then, by comparing the differences between denoised training speech 3 and the corresponding clean training speech 2, the network weights of the neural network are adjusted to obtain the trained speech denoising model. Furthermore, in Figure 1 During the test shown, the noisy speech 4, which is collected in real time by the microphone, is input into the denoising model, and the denoised speech 5 is output.

[0004] In the training process of the discriminative denoising model, a large amount of labeled data is required for training data such as noisy speech 1 to label the speech signal and noise. If this labeled data contains errors or inaccuracies, the model may be unable to correctly distinguish between valid speech and noise, causing valid speech in noisy speech 1 to be misidentified as noise and removed. Similarly, during testing, the model may misidentify valid speech in noisy speech 4 as noise, especially if the model has not learned from noisy speech 4 during training. In such cases, the model may misidentify valid speech as noise, thus damaging valid speech during the denoising process. Summary of the Invention

[0005] This application provides a speech processing method, medium, electronic device, and program product. The speech denoising process is based on the ODE diffusion model of linear trajectory, which does not damage the speech data, has strong generalization performance, and has a fast inference process and high computational efficiency.

[0006] In a first aspect, an embodiment of the present application provides a speech processing method, which comprises: obtaining first time domain data of to-be-processed speech data; converting the first time domain data from the time domain to the frequency domain to obtain first frequency domain data; inputting the first frequency domain data into a diffusion model to obtain second frequency domain data, wherein the diffusion model obtains correction information corresponding to the first frequency domain data according to a preset linear relationship, and performs denoising processing on the first frequency domain data through the correction information to obtain the second frequency domain data, and the first frequency domain data and the second frequency domain data have a preset linear relationship; converting the second frequency domain data from the frequency domain to the time domain to obtain second time domain data of the to-be-processed speech data after denoising processing.

[0007] It can be understood that the first frequency domain data and the corresponding second frequency domain data conform to the above-mentioned preset linear relationship, and the data mapping trajectory between the first frequency domain data and the corresponding second frequency domain data in the diffusion model is a straight line. In this way, the second frequency domain data (i.e. denoised frequency domain data) can be inferred from the first frequency domain data (i.e. noisy frequency domain data) based on the straight line data mapping trajectory, and the inference process is fast. In addition, the diffusion model provided in the present application does not damage the speech data during speech denoising processing, and has strong generalization performance and strong suppression ability for noise data that has not been seen in the training stage. Moreover, since the present application performs denoising on noisy speech data in the frequency domain, it utilizes the distribution of speech data in the frequency domain, ensures the continuity of the speech spectrum of the denoised speech data, and ensures the clarity of the denoised speech data.

[0008] In a possible implementation of the above-mentioned first aspect, the correction information is used to indicate the data change speed of the first frequency domain data changing to the second frequency domain data.

[0009] At this time, the second frequency domain data (i.e. denoised frequency domain data) can be considered as clean speech data in the first frequency domain data (i.e. noisy frequency domain data), that is, the noisy frequency domain data and the corresponding denoised frequency domain data conform to the above-mentioned preset linear relationship.

[0010] In a possible implementation of the above-mentioned first aspect, the second frequency domain data is obtained by: inputting N sampling times into the diffusion model; for i

[0011] That is, in the reasoning process of the diffusion model, N pieces of correction information corresponding to the noisy speech data to be processed can be obtained at N sampling time points in the frequency domain, so as to perform N times of denoising processing to obtain the final denoised speech data. It can be understood that when the data mapping track of the diffusion model is a straight line, the value of N can be selected as a small data, so that the reasoning time of the diffusion model is short and the speed is fast.

[0012] In a possible implementation of the first aspect, the (i+1)th frequency domain data is obtained by: for the ith sampling time, deriving the ith frequency domain data with respect to time by the diffusion model to obtain the ith correction information, where the ith correction information is used to indicate the data change speed of the change of the ith frequency domain data to the (i+1)th frequency domain data; and performing denoising processing on the ith frequency domain data according to the ith correction information by using a correction algorithm to obtain the (i+1)th frequency domain data.

[0013] It can be understood that N represents the number of denoising reasoning of the diffusion model, and then in the denoising reasoning process of the diffusion model, the data change speed of the noisy speech data can be used as the correction information corresponding to the current time step.

[0014] In a possible implementation of the first aspect, the correction algorithm includes at least one of Euler algorithm and Chasen extrapolation RK algorithm.

[0015] In a possible implementation of the first aspect, the ith correction information is a ratio of the ith frequency domain data to N; and the (i+1)th frequency domain data is obtained by adding the ith correction information to the ith frequency domain data. Wherein, N represents the number of denoising reasoning of the diffusion model. It can be understood that since the data mapping track of the diffusion model can be a straight line, the ratio of the ith frequency domain data to N, that is, the ith correction information, can represent the difference between the noisy speech data at the current time step and the corresponding denoised speech data.

[0016] In a possible implementation of the first aspect, the diffusion model is trained by: obtaining first training time domain data and second training time domain data of training speech, where the second training time domain data is obtained by superimposing preset time domain noise data on the first training time domain data; converting the first training time domain data and the second training time domain data from time domain to frequency domain to obtain first training frequency domain data and second training frequency domain data; and training the diffusion model according to the first training frequency domain data and the second training frequency domain data to obtain a trained diffusion model.

[0017] The embodiment of the present application provides a speech processing method. In the process of training an ODE-based diffusion model, a preset linear relationship between training clean speech data (i.e., first training frequency domain data) and training noisy speech data (second training frequency domain data) can be constructed, so that the data mapping trajectory between the training clean speech data and the training noisy speech data is a straight line. Specifically, the training sample and the training target in the diffusion model are constructed to train the ability of the diffusion model to evolve the training clean speech data into the training noisy speech data along the data mapping trajectory of the straight line. Thus, the pre-trained diffusion model can perform de-noising processing on the to-be-processed speech data containing noise to obtain de-noised speech data.

[0018] In a possible implementation of the first aspect, the preset time domain noise data indicates noise including at least one of the following: wind sound, siren sound, keyboard sound, and human speech sound. In this way, the second training time domain data is a kind of noisy speech data constructed in the time domain, i.e., a kind of low-quality speech data.

[0019] In a possible implementation of the first aspect, the diffusion model is trained according to the first training frequency domain data and the second training frequency domain data, including: obtaining a training sampling time; performing linear interpolation on the first training frequency domain data and the second training frequency domain data based on the preset linear relationship according to the training sampling time to obtain third training frequency domain data; superimposing preset frequency domain noise data on the third training frequency domain data to obtain fourth training frequency domain data, wherein the dimension of the third training frequency domain data is the same as the dimension of the preset frequency domain noise data; inputting the fourth training frequency domain data into the diffusion model, and deriving the fourth training frequency domain data with respect to time through the diffusion model to obtain first training correction information; and adjusting network parameters of the diffusion model according to the difference between the first training correction information and preset correction information.

[0020] It can be understood that the first training frequency domain data and the second training frequency domain data can be time-frequency domain data, for example, a frequency spectrum diagram. At this time, the third training frequency domain data is a training sample of the diffusion model, and the fourth training frequency domain data is a noisy training sample of the diffusion model.

[0021] In a possible implementation of the first aspect, the third training frequency domain data x t is obtained by the following formula: x t =t*x+(1-t)*y, wherein x represents the first training frequency domain data, y represents the second training frequency domain data, t represents the training sampling time, and x and y in t*x+(1-t)*y have the preset linear relationship.

[0022] It can be understood that, in the diffusion model training process, x t is the evolution of the training clean speech data x (t=0, x0=x) to the training noisy speech data y (t=T, xT is linearly interpolated between x and y along the forward process, so that the backward process can be pushed to the training clean speech data x at t = 0.

[0023] In a possible implementation of the first aspect, the fourth training frequency domain data x t_noise is obtained by: x t_noise = μ + δz, where μ represents a mean value of the third training frequency domain data, z represents preset frequency domain noise data, and δ represents a standard deviation of the preset frequency domain noise data.

[0024] In a possible implementation of the first aspect, 0 < δ < 1, and the value of δ increases first and then decreases as the training sampling time increases. As an example, when the training sampling time t takes values of 0.003, 0.5, and 1 respectively, the value of δ takes values of 0, 0.5, and 0 respectively.

[0025] In a possible implementation of the first aspect, the preset frequency domain noise data is Gaussian noise data with a mean value of 0 and a variance of 1. In a diffusion model, a noise with a standard normal distribution is usually added to data to simulate the randomness of the data. In the training phase, the model learns how to restore the noisy speech data (such as X T ) into clean speech data (that is, X0). In the generation phase, the model uses the learned knowledge to generate new data instances.

[0026] In a possible implementation of the first aspect, the preset correction information is y - x.

[0027] In a possible implementation of the first aspect, the difference between the first training correction information and the preset correction information is determined by a loss function loss = mse(v, y - x), where mse() represents a mean square error.

[0028] In a possible implementation of the first aspect, the first frequency domain data, the first training frequency domain data, and the second training frequency domain data are obtained by using a short-time Fourier transform (STFT) technique.

[0029] In a possible implementation of the first aspect, the value of N is determined according to a place or an application program in which the to-be-processed speech data is located, and the value of N corresponding to the to-be-processed speech data is different in different places or different application programs.

[0030] In a second aspect, an embodiment of the present application provides a readable medium, and the readable medium stores instructions. When the instructions are executed on an electronic device, the electronic device performs the speech processing method in the first aspect and any possible implementation manner thereof.

[0031] In a third aspect, an electronic device is provided, including: a memory, configured to store instructions executed by one or more processors of the electronic device; and a processor, one of the processors of the electronic device, configured to execute the speech processing method in the first aspect and any possible implementation manner thereof.

[0032] In a fourth aspect, a computer program product is provided, which, when running on an electronic device, causes the electronic device to implement the speech processing method in the first aspect and any possible implementation manner thereof. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 A schematic diagram of a speech processing scenario for a traditional method using a discriminative speech denoising model;

[0034] Figure 2 A schematic diagram of a speech processing scenario for a diffusion model provided by an embodiment of the present application;

[0035] Figure 3 A schematic diagram of a data mapping track of speech processing for a diffusion model provided by an embodiment of the present application;

[0036] Figure 4 A schematic diagram of a data mapping track of a diffusion model based on ODE provided by an embodiment of the present application;

[0037] Figure 5 A schematic diagram of time domain data and frequency domain data of speech in a speech denoising process provided by an embodiment of the present application;

[0038] Figure 6 A schematic diagram of an application scenario of a speech processing method provided by an embodiment of the present application;

[0039] Figure 7 A schematic diagram of a flow of a speech processing method of a diffusion model provided by an embodiment of the present application;

[0040] Figure 8 A schematic diagram of a flow of a training method of a diffusion model provided by an embodiment of the present application;

[0041] Figure 9 A schematic diagram of a structure of a mobile phone provided by an embodiment of the present application. DETAILED DESCRIPTION

[0042] Illustrative embodiments of the present application include, but are not limited to, a speech processing method, a medium, an electronic device, and a program product.

[0043] First, some terms in the speech processing method provided by the present application are explained.

[0044] Diffusion model: A generative model that learns the distribution of data by simulating the diffusion process of data, so as to generate new data samples.

[0045] Time step: Time step refers to the time unit of simulating the diffusion process of data in the diffusion model. At each time step, the model will add some noise to the current state, or reduce the noise to restore the original data.

[0046] Sampling time: refers to the sampling time of the model at each time step in the training process of the diffusion model.

[0047] In some embodiments, in order to solve the problem of damaging effective speech in the speech denoising process in the background art, a pre-trained diffusion model is provided, which is used to denoise the noisy speech data to obtain denoised speech data.

[0048] In the field of speech denoising, the diffusion model can be used to restore the speech data damaged by noise. Specifically, the diffusion model learns the distribution of data by gradually increasing noise on the initial clean speech data through a neural network and then trying to reduce the noise, so that the neural network can generate denoised speech data close to or the same as the initial clean speech data, that is, the neural network restores the initial clean speech data as much as possible. Therefore, the diffusion model denoises the noisy speech data by generating new data, rather than by labeling noise and directly removing the labeled noise from the noisy speech data, so it will not damage the effective speech data in the noisy speech data due to mislabeling, etc.

[0049] As shown in Figure 2 , the diffusion model usually includes two processes: a forward process (i.e., a forward diffusion process) and a reverse process (i.e., a reverse diffusion process). Figure 2 The forward process of the diffusion model gradually adds Gaussian noise to the clean speech data x0 at T (e.g., T = 4) sampling time points (i.e., T time steps) to obtain noisy speech data x T The reverse process is the process of gradually denoising the noisy speech data x T to restore the clean speech data x0 at T sampling time points.

[0050] As shown in Figure 2As shown, at the first sampling time point, Gaussian noise n0 can be added to clean speech data x0 in the forward process to obtain noisy speech data x1, and Gaussian noise n0 can be removed from noisy speech data x1 in the corresponding reverse process to obtain clean speech data x0. Similarly, Gaussian noise n0, n1, n2, and n3 can be added to speech data x1 in the forward process in turn through four sampling time points to obtain noisy speech signal x4, and conversely, Gaussian noise n0, n1, n2, and n3 can be removed from noisy speech data X4 in the reverse process for multiple iterations to obtain restored clean speech data x0.

[0051] It can be understood that, in the actual process of speech denoising, the diffusion model removes noise from the to-be-processed noisy speech data through the reverse process to obtain denoised speech data, i.e., clean speech data.

[0052] The diffusion model can process speech data in the time-frequency domain, i.e., the speech data at each sampling time point in the diffusion model is time-frequency domain data, such as a spectrogram. For example, Figure 2 The noisy speech data x T may be a speech spectrogram. Specifically, the horizontal coordinate of the speech spectrogram is time, the vertical coordinate is frequency, and the coordinate point value is speech data energy. Moreover, the energy value in the spectrogram is represented by color, and the deeper the color, the stronger the energy of the point.

[0053] In some embodiments, the diffusion model in the present application can be based on stochastic differential equations (SDE) or ordinary differential equations (ODE) to construct a data mapping relationship between clean speech data x(0) and noisy speech data x(T). As shown in Figure 3 The data mapping relationship of the diffusion model is used to transform the data distribution, specifically to convert the data distribution p0(x) of clean speech data x(0) to the data distribution p T (x) of noisy speech data x(T). In practical applications, the diffusion model converts clean speech data x(0) to noisy speech data x(T) through T sampling time points. The data distribution of the speech data can be visualized by the spectrogram in the time-frequency domain, i.e., the data distribution of the speech data can be represented by the spectrogram.

[0054] It can be understood that, since SDE describes the evolution of data distribution over time in the diffusion model, and the noise term is included in SDE, the solution of SDE is random, so the trajectory of data mapping relationship is usually complex and uncertain. In contrast, ODE describes the deterministic dynamic behavior of the system, and the solution of ODE is deterministic, so the trajectory of data mapping relationship is relatively smooth and stable. As shown in Figure 3 , L1 and L1' are respectively the data mapping trajectories of the forward process and the reverse process of SDE, and L2 and L2' are respectively the data mapping trajectories of the forward process and the reverse process of the conventional ODE. The trajectories L1 and L1' of SDE fluctuate greatly and are relatively complex. While the trajectories L2 and L2' of the conventional ODE fluctuate less and are relatively simple.

[0055] However, during the training process, the shape of the data mapping trajectory of SDE and ODE affects the convergence speed and inference speed of the model. Specifically, the data mapping trajectory of SDE is too complex, and the solution of SDE is random, which may cause the model to be difficult to capture the essential characteristics of the data, thereby prolonging the training time and inference time. In addition, the training process of SDE may require complex numerical methods, such as Monte Carlo simulation or numerical solver, to solve the solution of SDE based on the random and continuous data mapping trajectory, resulting in a longer training time. For ODE, since ODE is a discrete-time model, its training time is usually limited, and can be controlled by adjusting the time step. Moreover, the training process of ODE is relatively simple, and usually only a simple numerical solver is needed to solve it, so the training process is relatively simple, which is beneficial to reduce the training time and inference time.

[0056] In addition, in the numerical solution process of ODE, in order to realize that the data mapping trajectory is as close as possible to the real physical process, the curved trajectory is adjusted to a more smooth and direct straight line trajectory, so as to reduce the numerical error of solving ODE and improve the calculation efficiency, thereby facilitating the reduction of the training time of the diffusion model.

[0057] In some embodiments, the present application can straighten the ODE data mapping trajectory of the diffusion model through a reflow technique. Continuing to refer to Figure 3 , L3 and L3' are respectively the ODE data mapping trajectories straightened by the reflow technique. As an example, the speech processing method in the present application can specifically use the straight data mapping trajectory of the diffusion model to perform denoising processing on noisy speech data, to obtain clean speech data after denoising.

[0058] Further, referring to Figure 4 , the data mapping trajectory of ODE is described. Figure 4The data mapping trajectory of the clean speech data x and the noisy speech data y shown in (a) in FIG. 4 is curved in continuous time, which is L4. The trajectory L4 can be realized by four time steps L41, L42, L43 and L44 in the sampling process. In addition, Figure 4 The data mapping trajectory of the clean speech data x and the noisy speech data y shown in (a) in FIG. 4 is straight in continuous time, which is L5. The straight trajectory L5 can be realized by one time step L51 in the sampling process. Therefore, the straight ODE data mapping trajectory constructed in the diffusion model has a faster training and inference speed.

[0059] In some embodiments, the speech processing method provided in the present application sets the data mapping relationship between the clean speech data and the noisy speech data in the diffusion model as a preset linear relationship, so that the trajectory of the data mapping relationship is a straight line. In this way, based on the straight line data mapping trajectory, the training and inference speed of the speech denoising process of the diffusion model is faster.

[0060] In some embodiments, the pre-trained diffusion model in the present application can be constructed based on a linear ordinary differential equation, wherein the linear ordinary differential equation is constructed based on the noisy speech data and the corresponding clean speech data according to a preset linear relationship. Specifically, in the denoising inference process of the diffusion model on the noisy speech data, the ordinary differential equation is used to indicate the data transformation speed of the noisy speech data over time, and specifically to indicate the difference between the noisy speech data and the corresponding clean speech data. For example, the differential equation is x t = t * x + (1-t) * y, where y represents the noisy speech data, x represents the corresponding denoised speech data, and t represents the sampling time at one time step. For example, in one time step, the correction information of the noisy speech data to be processed can represent the difference between the denoised speech signal (i.e., clean speech data) corresponding to the noisy speech data and the noisy speech data. At this time, adding the correction information to the noisy speech data can obtain the corresponding denoised speech data.

[0061] Therefore, in the denoising inference process, the diffusion model can infer the correction information of the noisy speech data, specifically determine the correction information of the noisy speech data at each time step, and determine the corresponding denoised speech data through the correction information. Further, after one or more time steps, the determined denoised speech data is the corresponding clean speech data.

[0062] In some embodiments, the training process of the diffusion model in this application includes: establishing a preset linear relationship between the clean training frequency domain data, the corresponding noisy training frequency domain data, and the sampling time in the frequency domain to construct sample data. Then, adding Gaussian noise or other noise to the training sample data in the frequency domain to obtain updated sample data. Next, inputting the updated sample data into the neural network of the diffusion model, and outputting corresponding training correction information, which indicates the difference between the noisy training frequency domain data and the clean training frequency domain data. Finally, adjusting the network parameters of the neural network based on the difference between the training correction information and the preset correction information to obtain a pre-trained diffusion model.

[0063] Specifically, the speech processing method of this application, during the speech denoising inference process of the diffusion model, can acquire the noisy frequency domain data of the noisy speech to be processed in the time-frequency domain, such as a spectrogram. Then, the correction information corresponding to the noisy frequency domain data is obtained through the pre-trained diffusion model. Thus, the noisy frequency domain data is denoised using the correction information to obtain the corresponding denoised frequency domain data, thereby achieving the denoising effect. At this time, the denoised frequency domain data can be considered as the clean speech data in the noisy frequency domain data, that is, the noisy frequency domain data and the corresponding denoised frequency domain data conform to the aforementioned preset linear relationship. In this way, denoised frequency domain data can be inferred from the noisy frequency domain data based on a straight-line data mapping trajectory, and the inference process is relatively fast. In addition, the diffusion model provided in this application does not damage the speech data during the speech denoising process, and has strong generalization performance and strong suppression ability for noise data not seen in the training stage.

[0064] It is understandable that the diffusion model performs speech denoising in the time-frequency domain. After an electronic device collects noisy speech data through a microphone, it first converts the noisy speech data from the time domain to the frequency domain, obtaining noisy frequency domain data, such as a spectrogram. Then, a pre-trained diffusion model is used to denoise the noisy frequency domain data, obtaining denoised frequency domain data. Furthermore, the denoised frequency domain data is converted back from the frequency domain to the time domain to obtain denoised speech data in the time domain. Thus, because this application denoises noisy speech data in the frequency domain, it utilizes the distribution of speech data in the frequency domain, ensuring the continuity of the speech spectrum in the denoised speech data and guaranteeing the clarity of the denoised speech data.

[0065] like Figure 5 The diagram shows the time-domain and frequency-domain data of speech during the speech denoising process. Figure 5 The noisy speech data A1 shown in (a) is noisy speech collected in a supermarket setting, and Figure 5 Figure (a) shows the time-domain waveform A11' and the corresponding spectrogram B1 of a segment of noisy speech A11 from the noisy speech data A1.Figure 5 (b) in FIG. 1 shows the time-domain data waveform A12 and the corresponding spectrum B2 of the noisy speech data A11 after the de-noising of the technical in the background. In the time-domain data waveform A12, although no noise data is contained, some valid speech data can also be removed. In addition, Figure 5 (c) in FIG. 1 shows the time-domain data waveform A13 and the corresponding spectrum B3 of the noisy speech A11 after the de-noising of the diffusion model based on the straight-line data mapping trajectory in the present application. At this time, the time-domain data waveform A13 contains no noise data and also contains relatively complete valid speech data, and the de-noising effect is good.

[0066] In addition, in some embodiments, in the inference process of the diffusion model, N pieces of correction information corresponding to the noisy speech data to be processed can be obtained at N sampling time points in the frequency domain, so as to perform N times of de-noising processing to obtain the final de-noised speech data. Wherein, N is a positive integer. It can be understood that when the data mapping trajectory of the diffusion model is a straight line, the value of N can be selected as a small data, so that the inference time of the diffusion model is short and the speed is fast.

[0067] The application scenarios of the speech processing method of the present application include but are not limited to: recording (such as classroom recording), human-computer interaction scenarios (such as human-computer interaction based on voice assistants), voice calls, video calls, video conferences, etc.

[0068] In some embodiments, the electronic device can perform de-noising processing on the complete noisy speech data after the microphone completes the collection of the complete noisy speech data to obtain de-noised speech data. As an example, in the recording scene or the human-computer interaction scene, the electronic device can perform de-noising processing on the complete noisy speech data. For example, Figure 6 (a) in FIG. 1 shows a human-computer interaction scenario. Specifically, when a user uses a mobile phone 10 to listen to the radio in a park, the mobile phone 10 displays an input control 12 of a voice assistant on a desktop interface 11, and collects noisy speech data containing the voice "play the morning news" and the park's human voice and bird chirping sound through the microphone. At this time, the mobile phone 10 can perform de-noising processing on the noisy speech data to obtain de-noised speech data including only the voice "play the morning news", and call the audio software through the voice assistant to play the audio of the morning news.

[0069] In other embodiments, the electronic device can perform segmented de-noising processing on the process of collecting noisy speech data by the microphone, for example, performing de-noising processing on a segment of noisy speech data collected in real time for tens of milliseconds (such as 40 ms), and then performing de-noising processing on another segment of noisy speech data collected in real time. For example, in the scenarios of voice calls, video calls, video conferences, etc., the real-time collected speech is segmented and de-noised by 40 ms. For example,Figure 6 (b) in FIG. 13 shows a video call scene. Specifically, when a user uses the mobile phone 10 to video call with family at the seaside, the mobile phone 10 displays the call interface 13 and collects noisy speech data containing the call speech "look at the big ship on the sea" and the sound of the ship's whistle and the sound of the sea waves. At this time, the mobile phone 10 can perform segment denoising processing on the noisy speech data to obtain denoised segment denoising speech data, and then transmit the segment denoising speech data to the friend in the video call in chronological order. In this way, it is ensured that the friend can hear and clearly hear the speech "look at the big ship on the sea" in time.

[0070] The speech processing method provided by the embodiments of the present application can be applied to various electronic devices. For example, the electronic device is not limited to the mobile phone 10, but can also be a smart screen, a tablet computer, a wearable electronic device, a vehicle-mounted electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, etc. The embodiments of the present application do not make any limitation on the specific type of the electronic device.

[0071] Optionally, the subject performing the speech processing method in the present application can be an electronic device, or a device (denoted as a speech processing device) in the electronic device for performing the speech processing method, and the embodiments of the present application do not make any specific limitation thereon.

[0072] Next, the training and inference process of the diffusion model for speech denoising provided by the embodiments of the present application will be introduced.

[0073] Diffusion model inference process

[0074] Referring to Figure 7 The flowchart of the speech processing method provided by the embodiments of the present application is realized by the backward process of the diffusion model to realize denoising inference. Specifically, the flowchart includes the following steps:

[0075] S701: Obtain the inference number N of the to-be-processed speech data.

[0076] The to-be-processed speech data can be speech data containing noise collected by the electronic device in real time, for example, recording data collected by a microphone in a recording scene, or a segment of call speech data collected by a microphone in a voice call scene.

[0077] In some embodiments, the inference number N of the to-be-processed voice data can be a preset fixed value, or a dynamic value preset according to factors such as an application or a location where the to-be-processed voice data is located.

[0078] As an example, for the application where the to-be-processed noisy voice data is located, the N value corresponding to the recording application is 3, the N value corresponding to the call application is 2, and the N value corresponding to the voice assistant application is 1. In some embodiments, when the electronic device collects the to-be-processed noisy voice data through the microphone, the application that invokes the microphone is detected, such as the application that is currently running in the foreground.

[0079] As another example, for the location where the to-be-processed noisy voice data is located, for example, the N value corresponding to a relatively quiet location such as a home, a conference room, a small street, and a library is 2, the N value corresponding to a relatively noisy location such as a restaurant, an office, and a shopping mall is 5, and the N value corresponding to a more noisy location such as a station, a market, and a factory is 10. In some embodiments, when the electronic device collects the to-be-processed noisy voice data through the microphone, the current location can be located to determine the location where the current location is located.

[0080] S702: converting the first time domain data of the to-be-processed voice data from the time domain to the frequency domain to obtain first frequency domain data.

[0081] In some embodiments, the present application can use a Fourier transform method such as a short-time Fourier transform (STFT) to convert time domain data into frequency domain data, specifically time-frequency domain data, through feature transformation. For example, the first frequency domain data after feature transformation is time-frequency domain data, such as a spectrogram.

[0082] S703: obtaining N sampling times corresponding to the N inference numbers.

[0083] Optionally, the present application can divide the preset time range t min ~ t max into equal intervals to obtain N sampling times, that is, N time steps. Among them, t min and t max can be set according to actual needs, for example, t min is 0, and t max is 1. In actual applications, in order to realize normal sampling, the sampling time t is a value greater than 0, for example, t min can be 0.003.

[0084] As an example, the application can determine 4 sampling times from 0.003 to 1, i.e., divide 0.003 to 1 into 4 time steps, and the 4 sampling times are 0.25, 0.5, 0.75, and 1 respectively. Among them, the sampling times 0.25, 0.5, 0.75, and 1 can be denoted as t1, t2, t3, and t4 respectively.

[0085] It can be understood that each sampling time in the N sampling times is applied to a reasoning process of the diffusion model respectively.

[0086] S704: Taking the first frequency domain data as the 1st frequency domain data, obtaining the i-th correction information corresponding to the i-th frequency domain data through the pre-trained diffusion model.

[0087] Specifically, i takes a value of 1≤i≤N. Among them, in the 1st reasoning process of the diffusion model, the 1st frequency domain data is the first frequency domain data, and i is a positive integer. At this time, in the i-th denoising reasoning process, the diffusion model obtains the i-th correction information corresponding to the i-th frequency domain data.

[0088] In some embodiments, the pre-trained diffusion model evolves the clean speech data x into the corresponding noisy speech data y in the time-frequency domain, i.e., the evolution process of the diffusion model is x→y. Specifically, the diffusion model can evolve the speech data x0 into x T through T time steps, i.e., the evolution process is x0→x T . Among them, x0=x at the sampling time t=0, and x T =y at the sampling time t=T.

[0089] In some embodiments, the evolution process of the ODE-based diffusion model in the application is as follows: dx t / dt=f(x t ). That is, the diffusion model defines the data change rate in the form of a differential equation, indicating how to evolve x0→x T .

[0090] It can be understood that x t in the application is obtained by linearly interpolating the known clean speech data x(t=0, x0=x) to the known noisy speech data y(t=T, x T =y) along the forward process, so that the reverse process can derive the clean speech data x at t=0. Among them, the specific construction of x t will be described in detail below, which will not be repeated here.

[0091] Specifically, the ODE-based diffusion model can calculate the data change rate of the speech data at each time step in the process of converting clean speech data into noisy speech data by gradually increasing noise at T time steps. Correspondingly, the value of the time step N in the inference process is less than or equal to T, and at this time, the N sampling times corresponding to the N time steps can be selected from the T sampling times corresponding to the T time steps.

[0092] It can be understood that, in the denoising inference process of the diffusion model, x T (i.e., y) is the first frequency domain data of the speech data to be processed. At this time, in the denoising inference process of the diffusion model, the data change rate of the speech data can be used as the correction information corresponding to the current time step.

[0093] In some embodiments, the i-th correction information in the present application can be obtained by deriving the i-th sampling time t with respect to the i-th frequency domain data according to the following formula (1).

[0094]

[0095] In the formula (1), x ti represents the i-th frequency domain data, and specifically represents the frequency domain data of the speech data to be processed before denoising in the i-th denoising inference process. At this time, x ti is derived with respect to the sampling time t to obtain v i , v i represents the change rate of the i-th frequency domain data and the speech data after denoising with respect to time.

[0096] S705: Denoising the i-th frequency domain data by using the i-th correction information to obtain the i+1-th frequency domain data.

[0097] Specifically, the present application can denoise the i-th frequency domain data by using the i-th correction information and N to obtain the i+1-th frequency domain data.

[0098] In some embodiments, the present application can use a correction algorithm to denoise the i-th frequency domain data by using the i-th correction information. For example, the correction algorithm can be Euler algorithm or richardson extrapolation (RK) (for example, RK-45 algorithm).

[0099] In some embodiments, the i-th correction information of the i-th frequency domain data is calculated by using the following formula (2) to obtain the i+1-th frequency domain data:

[0100]

[0101] It can be understood that v idenotes the difference between the first frequency domain data and the corresponding denoised speech data, such as the data change speed of the first frequency domain data changing over time to the second frequency domain data. Wherein, v i is the i-th correction information, y i refers to the i-th frequency domain data (i.e. the noisy speech data before the i-th denoising processing), y i+1 refers to the i+1-th frequency domain data (i.e. the denoised speech data after the i-th denoising processing). At this time, y N+1 refers to the N+1-th frequency domain data (denoted as the second frequency domain data), i.e. the frequency domain data after the N-th denoising processing of the to-be-processed speech data.

[0102] In some embodiments, the first frequency domain data of the to-be-processed speech data and the second frequency domain data after denoising processing have a preset linear relationship. At this time, the first frequency domain data can evolve into the second frequency domain data through the denoising processing along the data mapping track of the straight line in the reverse process through the diffusion model.

[0103] S706: Determine whether the inference number i is less than N.

[0104] If i < N, re-enter S704 to continue denoising inference, and if i = N, enter S707 to output the N+1-th frequency domain data through the diffusion model.

[0105] S707: Corresponding to i = N, convert the i+1-th frequency domain data from the frequency domain to the time domain to obtain the second time domain data corresponding to the denoised speech data.

[0106] At this time, the second time domain data is the time domain data of the denoised speech signal after denoising of the to-be-processed speech data.

[0107] In this way, the present application can obtain the frequency domain data of the to-be-processed speech signal through the diffusion model, and perform denoising processing on the frequency domain data through the diffusion model along the data mapping track of the straight line according to the preset linear relationship to obtain the frequency domain data after denoising. Wherein, for different places and different application programs, the diffusion model can obtain the denoised speech data through corresponding N times of denoising inference, which is conducive to improving the speech denoising effect.

[0108] Diffusion model training process

[0109] Referring to Figure 8 , the training process of the diffusion model provided by the embodiments of the present application is introduced. As Figure 8 shown, the execution subject of the training process can be an electronic device, and the training process includes the following steps:

[0110] S801: Obtain the first training time domain data of the training speech.

[0111] The first training time domain data can include time domain data of training speech, i.e., clean and undamaged speech data in the time domain.

[0112] S802: Obtain preset time domain noise data, add the preset time domain noise data to the first training time domain data, and obtain second training time domain data of training noisy speech.

[0113] The preset time domain noise data can include noise of multiple sound sources, such as one or more of wind noise, siren noise, keyboard noise, and speaking noise. Correspondingly, the second training time domain data is a kind of noisy speech data constructed in the time domain, i.e., a kind of low-quality speech data.

[0114] S803: Convert the first training time domain data and the second training time domain data from the time domain to the frequency domain to obtain first training frequency domain data and second training frequency domain data, respectively.

[0115] In some embodiments, the Fourier transform method such as short-time Fourier transform can be used to convert the time domain data into frequency domain data through feature transformation, specifically, time-frequency domain data.

[0116] It can be understood that the first training frequency domain data and the second training frequency domain data can be time-frequency domain data, such as a frequency spectrum diagram.

[0117] S804: Obtain a training sampling time.

[0118] In some embodiments, the training sampling time t is randomly sampled from a preset time range t min ~ t max . The values of t min and t max may be set according to actual needs, for example, t min is 0 and t max is 1. In actual applications, the sampling time t is a value greater than 0, and in order to normally implement sampling, t min may be set to 0.003 and t max may be set to 1.

[0119] Optionally, the preset time range t min ~ t max may be divided into T time steps at equal intervals to obtain T sampling times.

[0120] For example, four sampling times can be determined from 0.003 to 1, i.e., 0.003 to 1 is divided into four time steps, and the four sampling times are 0.25, 0.5, 0.75, and 1.

[0121] As an example, in a training process of the diffusion model, the training sampling time t can take a value of t2 (such as 0.5).

[0122] S805: Linearly interpolating the first training frequency domain data and the second training frequency domain data according to the training sampling time to obtain third training frequency domain data.

[0123] In some embodiments, the first training frequency domain data and the second training frequency domain data can be linearly interpolated according to the following formula (3) to obtain the third training frequency domain data. At this time, the third training frequency domain data is a training sample of the diffusion model.

[0124] x t = t*x + (1-t)*y (3)

[0125] Wherein, in the training scene of the model, x in formula (1) is the first training frequency domain data of the training clean speech, y is the second training frequency domain data of the training noisy speech data, t represents the training sampling time, then x t represents the third training frequency domain data of the training sampling time t. That is, x t represents the data point of the data distribution on the data mapping track between the first training frequency domain data x and the second training frequency domain data y at the sampling time t in the diffusion model.

[0126] For example, when the training sampling time t takes a value of t2, x t may be At this time, x and y in formula (3) have a preset linear relationship.

[0127] In the training process of the diffusion model, x t is the linear interpolation along the forward process from the training clean speech data x (t=0, x0=x) to the training noisy speech data y (t=T, x T =y) so that the reverse process can reach the training clean speech data x at t=0.

[0128] S806: Superimposing the third training frequency domain data with preset frequency domain noise data to obtain fourth training frequency domain data.

[0129] The above-mentioned preset frequency domain noise data can be Gaussian noise, for example, Gaussian noise with a mean of 0 and a variance of 1.

[0130] It can be understood that the Gaussian noise with a mean of 0 and a variance of 1 obeys the standard normal distribution, also known as Gaussian distribution. The standard normal distribution is a symmetrical bell-shaped curve, and its graph is symmetrical about the mean (such as 0).

[0131] The dimension of the third training frequency domain data is the same as that of the preset frequency domain noise data. At this time, the fourth training frequency domain data is a noise-added training sample of the diffusion model.

[0132] In the diffusion model, a noise with a standard normal distribution is usually added to the data (such as clean and undamaged speech data) to simulate the randomness of the data. In the training stage, the model learns how to restore the noisy speech data (such as X T ) to clean speech data (i.e., X0). In the generation stage, the model uses the learned knowledge to generate new data instances.

[0133] In combination with formula (3), the following formula (4) shows a way of constructing the fourth training frequency domain data based on the preset frequency domain noise data.

[0134] x t_noise =μ+δ*z=x t +δ*z (4)

[0135] In the training scenario of the model, x t_noise is the fourth training frequency domain data corresponding to the training speech. Correspondingly, μ represents the mean of the third training frequency domain data x t , z represents the preset frequency domain noise data, and δ is the variance of the preset frequency domain noise data, for example, Gaussian noise.

[0136] Optionally, δ changes with the change of the training sampling time t, for example, δ can first increase and then decrease with the increase of the sampling time t. And δ takes a value in the range of 0≤δ≤1. As an example, when the training sampling time t takes values of 0.003, 0.5 and 1, the value of δ is 0, 0.5 and 0 respectively. Of course, the value of δ in the present application can be set according to actual needs, and is not limited to the above examples.

[0137] For example, when the training sampling time t takes the value t2, δ=δ2=0.5.

[0138] S807: input the fourth training frequency domain data and the training sampling time into the diffusion model, and derive the fourth training frequency domain data with respect to the training sampling time through the diffusion model to obtain first training correction information.

[0139] It can be understood that the diffusion model is used to learn the speed of change between two distribution data points in the present application, such as the above-mentioned first training correction information.

[0140] Optionally, the diffusion model in the present application can be implemented by a neural network model such as noise conditional score networks (NCSN) and Unet, but is not limited thereto.

[0141] In some embodiments, the application can implement derivation of the fourth training frequency domain data pair with respect to the sampling time t by using the following formula (5), i.e., to realize x t_noise Derivation of t.

[0142]

[0143] wherein v is the first training correction information mentioned above. And since x t_noise Gaussian noise is introduced in the formula (4), the diffusion model will increase x t_noise the robustness of derivation of the sampling time t. In an ideal case, the first training correction information v = x-y. It can be understood that v in the formula (1) above i can be determined according to the formula (5).

[0144] For example, when the training sampling time t takes the value t2, At this time, the first training correction information represented by v2 is

[0145] It can be understood that the first training correction information v is used to represent the speed of change of x t with time, i.e., to represent the difference between the training clean speech data and the training noisy speech data in the frequency domain at the training sampling time t, i.e., the data change speed of the training clean speech data changing to the training noisy speech data.

[0146] S808: Adjust the network parameters of the diffusion model according to the difference between the first training correction information and the preset correction information.

[0147] Optionally, the preset correction information can be y-x, i.e., the difference between the training noisy speech data and the training clean speech data in the time-frequency domain. At this time, the preset correction information can represent the noise data in the training noisy speech data other than the training clean speech data.

[0148] In some embodiments, the application can calculate the loss function loss according to the difference between the first training correction information and the preset correction information, to perform iterative training of the diffusion model. When the preset number of training is reached, or the loss function value loss is less than the first threshold, the model training is ended. It can be understood that as the loss function value gradually decreases, the diffusion model used in the reverse diffusion process can acquire the ability to infer x from y in N steps.

[0149] It can be understood that the first training correction information is the true value of the diffusion model, and the preset correction information is the predicted value of the training target. Optionally, the loss function includes but is not limited to a regression loss function, such as a mean-square error (MSE) and a mean absolute error (MAE).

[0150] For example, the loss function of the present application can be calculated by loss=mse(v, y-x). Wherein, mse() is the mean-square error, v represents the first training correction information, and y-x represents the preset correction information.

[0151] The speech processing method provided by the embodiments of the present application can construct a preset linear relationship between the training clean speech data and the training noisy speech data in the process of training the ODE-based diffusion model, so that the data mapping trajectory between the training clean speech data and the training noisy speech data is a straight line. Specifically, the embodiments of the present application train the diffusion model to have the ability to evolve the training clean speech data into the training noisy speech data along the data mapping trajectory of the straight line by constructing the training sample and the training target in the diffusion model. Thus, the pre-trained diffusion model can perform denoising processing on the to-be-processed speech data containing noise to obtain denoised speech data.

[0152] In some embodiments, the speech processing method provided by the embodiments of the present application can be executed by various functional modules in the speech processing device. For example, the speech processing device can include the following modules: an inference number acquisition module configured to acquire an inference number N of to-be-processed speech data; a feature transformation module configured to convert first time-domain data of the to-be-processed data from a time domain to a frequency domain to obtain first frequency-domain data; a sampling time generation module configured to acquire N sampling times corresponding to the N inference numbers; a neural network calculation module configured to acquire an i-th correction information corresponding to an i-th frequency-domain data in an i-th inference process; a signal enhancement module configured to perform denoising on the i-th frequency-domain data by using the i-th correction information to obtain an (i+1)-th frequency-domain data; an inference number judgment module configured to judge whether the inference number i is less than N; and a feature inverse transformation module configured to convert the (N+1)-th frequency-domain signal into time-domain data as denoised speech data.

[0153] In addition, in some other embodiments, the voice processing apparatus can further include the following modules: a noise simulation module configured to generate preset time-domain noise data such as wind sound, siren sound, keyboard sound, and human speech sound; a Gaussian noise generation module configured to obtain preset frequency-domain noise data such as Gaussian noise; a training sample generation module configured to perform linear interpolation on the first training frequency-domain data and the second training frequency-domain data according to a training sampling time to obtain third training frequency-domain data; a training sample noise addition module configured to superimpose the third training frequency-domain data with preset frequency-domain noise data to obtain fourth training frequency-domain data; and a loss calculation module configured to adjust network parameters of the diffusion model according to a difference between the first training correction information corresponding to the fourth training frequency-domain data and the preset correction information. In addition, the feature change module, the first training time-domain data and the second training time-domain data are converted from the time domain to the frequency domain, and the first training frequency-domain data and the second training frequency-domain data are obtained. The neural network calculation module is further configured to calculate the first training correction information corresponding to the fourth training frequency-domain data.

[0154] Of course, the various functional modules of the voice processing apparatus provided by the embodiments of the present application include but are not limited to the above examples, and can also include other functional modules.

[0155] Next, the hardware structure of the electronic device to which the voice processing method provided by the embodiments of the present application is applied will be described. Specifically, with reference to Figure 9 The hardware structure of the electronic device will be described taking the electronic device as a mobile phone as an example.

[0156] As shown in Figure 9 The mobile phone 10 can include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, a key 101, and a display screen 102, etc.

[0157] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the mobile phone 10. In some other embodiments of the present application, the mobile phone 10 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0158] The processor 110 can include one or more processing units such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro programmed control unit (MCU), an artificial intelligence (AI) processor, or a field programmable gate array (FPGA), and the like. The different processing units can be independent devices or integrated in one or more processors. The processor 110 can be provided with a storage unit for storing instructions and data. In some embodiments, the storage unit in the processor 110 is a cache memory 180. For example, the cache memory 180 described above can be used to store the pre-trained ODE-based diffusion model described above. Furthermore, the processor 110 can construct training samples according to the preset linear relationship based on the training clean speech data and the training noisy speech data, and train the ODE-based diffusion model based on the training samples and the corresponding training target, so that the data mapping trajectory of the diffusion model is a straight line. In turn, the processor 110 can perform de-noising processing on the to-be-processed speech data containing noise based on the straight-line data mapping trajectory, and quickly obtain the de-noised speech data that is not damaged.

[0159] The power module 140 can include a power source, a power management component, and the like. The power source can be a battery. The power management component is used to manage the charging of the power source and the power supply of the power source to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module is used to receive charging input from a charger; the power management module is used to connect the power source, the charging management module, and the processor 110. The power management module receives the input of the power source and / or the charging management module to power the processor 110, the display screen 102, the camera 170, and the wireless communication module 120, and the like.

[0160] The mobile communication module 130 can include, but is not limited to, an antenna, a power amplifier, a filter, a low noise amplifier (LNA), etc. The mobile communication module 130 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the mobile phone 10. The mobile communication module 130 can receive electromagnetic waves by the antenna, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer to the modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor, and radiate as electromagnetic waves by the antenna. In some embodiments, at least part of the function modules of the mobile communication module 130 can be disposed in the processor 110. In some embodiments, at least part of the function modules of the mobile communication module 130 can be disposed in the same device as at least part of the modules of the processor 110. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), longer evolution (LTE), bluetooth (BT), global navigation satellite system (GNSS), wireless local area network (WLAN), near field communication (NFC), frequency modulation (FM), and / or field communication, NFC), infrared (IR) technology, etc.The GNSS can include a global positioning system (GPS), a global navigation satellite system (GLONASS), a beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS), and / or a satellite based augmentation systems (SBAS).

[0161] The wireless communication module 120 can include an antenna and implement the transceiving of electromagnetic waves via the antenna. The wireless communication module 120 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., a wireless fidelity (Wi-Fi) network), Bluetooth (BT), a global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, and the like, which is applied to the mobile phone 10. The mobile phone 10 can communicate with a network and other devices through the wireless communication technology.

[0162] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 10 can also be located in the same module.

[0163] The display screen 102 is configured to display a human-computer interaction interface, an image, a video, and the like. The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), or the like.

[0164] The sensor module 190 can include a proximity light sensor, a pressure sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0165] The audio module 150 is configured to convert digital audio information into an analog audio signal output, or convert an analog audio input into a digital audio signal. The audio module 150 can also be configured to encode and decode audio signals. In some embodiments, the audio module 150 can be disposed in the processor 110, or some functional modules of the audio module 150 can be disposed in the processor 110. In some embodiments, the audio module 150 can include a speaker, a receiver, a microphone, and a headset jack. For example, the microphone in the audio module 150 can be configured to receive noisy voice data, and the speaker can be configured to output de-noised voice data after de-noising processing of the noisy voice data.

[0166] The camera 170 is configured to capture still images or videos. An object generates an optical image through a lens and projects the optical image onto a photosensitive element. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to an image signal processor (ISP) to convert the electrical signal into a digital image signal. The mobile phone 10 can implement a shooting function through the ISP, the camera 170, a video codec, a graphic processing unit (GPU), the display screen 102, and an application processor, etc.

[0167] The interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface, etc. The external memory interface can be configured to connect an external memory card, such as a MicroSD card, to implement an extension of the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 through the external memory interface to implement a data storage function. The universal serial bus interface is configured to enable communication between the mobile phone 10 and other electronic devices. The subscriber identification module card interface is configured to communicate with a SIM card installed in the mobile phone 10, such as reading a phone number stored in the SIM card, or writing a phone number into the SIM card.

[0168] In some embodiments, the mobile phone 10 further comprises a key 101, a motor, an indicator, and the like. The key 101 can include a volume key, a power on / off key, and the like. The motor is used to generate a vibration effect of the mobile phone 10, for example, to generate a vibration when the mobile phone 10 of the user is called, so as to prompt the user to answer the call of the mobile phone 10. The indicator can include a laser indicator, a radio frequency indicator, an LED indicator, and the like.

[0169] In some embodiments, the present application provides a readable medium, which stores instructions, when executed on an electronic device, causes the electronic device to perform the voice processing method in the foregoing.

[0170] In some embodiments, the present application provides an electronic device, comprising: a memory, configured to store instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, configured to perform the voice processing method in the foregoing.

[0171] In some embodiments, the present application provides a computer program product, which comprises instructions for implementing the voice processing method in the foregoing.

[0172] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. Embodiments of the application can be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0173] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices, in known fashion. For purposes of this application, a processing system includes any system that has a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0174] The program code can be implemented in a high level procedural or object oriented programming language to communicate with a processing system. The program code can be implemented in assembly or machine language, if desired. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0175] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer-readable) medium, which can be read and executed by one or more processors. For example, the instructions can be downloaded from a network or by way of another computer readable medium. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation floppy disks, optical disks, optical disks, compact discs, read-only memory (CD-ROMs), magnetic disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or a tangible, machine-readable storage used in the transmission of information over the Internet with a propagated signal in electronic, electromagnetic, or optical form, such as carrier waves, infrared signals digital signals, etc. Accordingly, a machine-readable medium includes any type of medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0176] In the drawings, some of the structural or methodological features can be shown in particular arrangements and / or orders. However, it should be understood that such particular arrangements and / or orders can not be required. Instead, these features can be arranged in a different manner and / or order than shown in the illustrative figures, in some embodiments. Additionally, the inclusion of a structural or methodological feature in a particular figure is not meant to imply that such feature is required in all embodiments, and in some embodiments, these features can not be included or can be combined with other features.

[0177] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module, in physical, one logical unit / module can be one physical unit / module, also can be a part of one physical unit / module, also can be realized in combination of multiple physical unit / modules, the physical realization of these logical units / modules is not the most important, the combination of the functions realized by these logical units / modules is the key to solve the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce the units / modules which are not closely related to solving the technical problems proposed in the present application, which does not mean that the above-mentioned device embodiments do not have other units / modules.

[0178] It has to be noted that, in the description of the application and in the claims the terms "including" and "having" and the like are used in the sense of "including at least the recited entity but not excluding others". Furthermore, the terms "first", "second" and the like do not imply any ordering, quantity or importance, but are used to identify a first and second entity, respectively. Moreover, the terms "comprises", "comprising", or the like can be regarded as encompassing a non-exclusive inclusion, such that processes, methods, articles, or apparatuses comprising a list of elements are not limited to those elements, but can include other elements not expressly listed or even inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.

[0179] While the application has been illustrated and described in detail in the drawings and foregoing description, the same is to be considered as illustrative and not restrictive in character, it being understood that only the preferred embodiments have been shown and described and that all changes and modifications that come within the spirit of the application are desired to be protected.

Claims

1. A voice processing method, characterized by, The method comprises: obtaining first time domain data of to-be-processed voice data; converting the first time domain data from the time domain to the frequency domain to obtain first frequency domain data; inputting the first frequency domain data into a diffusion model to obtain second frequency domain data, wherein the diffusion model obtains correction information corresponding to the first frequency domain data according to a preset linear relationship, and performs denoising processing on the first frequency domain data through the correction information to obtain the second frequency domain data, and the first frequency domain data and the second frequency domain data have a preset linear relationship; converting the second frequency domain data from the frequency domain to the time domain to obtain second time domain data of the to-be-processed voice data after denoising processing.

2. The method of claim 1, wherein, The correction information is used to indicate a data change speed of the first frequency domain data changing into the second frequency domain data.

3. The method of claim 1, wherein, The second frequency domain data is obtained by the following method: inputting N sampling times into the diffusion model; corresponding to i corresponding to i = N, taking the i+1th frequency domain data as the second frequency domain data.

4. The method of claim 3, wherein, The i+1th frequency domain data is obtained by the following method: for the i th sampling time, the diffusion model derives the i th frequency domain data with respect to time to obtain the i th correction information, wherein the i th correction information is used to indicate a data change speed of the i th frequency domain data changing into the i+1th frequency domain data; using a correction algorithm to perform denoising processing on the i th frequency domain data according to the i th correction information to obtain the i+1th frequency domain data.

5. The method of claim 4, wherein, The correction algorithm comprises at least one of Euler algorithm and Runge-Kutta (RK) algorithm.

6. The method of claim 4, wherein, The i th correction information is a ratio of the i th frequency domain data to N; and the i+1th frequency domain data is obtained by adding the i th correction information to the i th frequency domain data.

7. The method according to any one of claims 1 to 6, characterized in that, The diffusion model is trained by the following method: obtaining first training time domain data and second training time domain data of training voice, wherein the second training time domain data is obtained by superimposing preset time domain noise data on the first training time domain data; converting the first training time domain data and the second training time domain data from the time domain to the frequency domain to obtain first training frequency domain data and second training frequency domain data, respectively; training the diffusion model according to the first training frequency domain data and the second training frequency domain data to obtain the trained diffusion model.

8. The method of claim 7, wherein, The preset time domain noise data indicates noise including at least one of the following: wind sound, siren sound, keyboard sound, and human speech sound.

9. The method according to claim 7 or 8, characterized in that, The training of the diffusion model according to the first training frequency domain data and the second training frequency domain data comprises: obtaining training sampling times; The first training frequency domain data and the second training frequency domain data are linearly interpolated based on a preset linear relationship according to the training sampling time, to obtain third training frequency domain data; The third training frequency domain data is superimposed with preset frequency domain noise data to obtain fourth training frequency domain data, wherein the dimension of the third training frequency domain data is the same as the dimension of the preset frequency domain noise data; The fourth training frequency domain data is input into the diffusion model, and the fourth training frequency domain data is differentiated with respect to time by the diffusion model to obtain first training correction information; The network parameters of the diffusion model are adjusted according to the difference between the first training correction information and preset correction information.

10. The method of claim 9, wherein, The third training frequency domain data x t is obtained by the following formula: x t = t * x + (1 - t) * y, Wherein, x represents the first training frequency domain data, y represents the second training frequency domain data, t represents the training sampling time, and x and y in t*x+(1-t)*y have the preset linear relationship.

11. The method according to claim 9 or 10, characterized in that, The fourth training frequency domain data x t_noise is obtained by x t_noise = μ + δz, Wherein, μ represents the mean of the third training frequency domain data, z represents the preset frequency domain noise data, and δ represents the standard deviation of the preset frequency domain noise data.

12. The method of claim 11, wherein, 0≤δ≤1, and the value of δ increases first and then decreases with the increase of the training sampling time.

13. The method of claim 9, wherein, The preset frequency domain noise data is Gaussian noise data with a mean of 0 and a variance of 1.

14. The method of claim 10, wherein, The preset correction information is y-x.

15. The method of claim 14, wherein, The difference between the first training correction information and the preset correction information is determined by a loss function loss=mse(v,y-x), wherein mse() represents the mean square error.

16. The method of claim 7, wherein, The first frequency domain data, the first training frequency domain data and the second training frequency domain data are obtained by using short-time Fourier transform (STFT) technology.

17. The method of claim 3, wherein, The value of N is determined according to the place or application program where the to-be-processed voice data is located, and the value of N corresponding to the to-be-processed voice data is different in different places and different application programs.

18. A readable medium characterized by The readable medium stores instructions, which, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1-17.

19. An electronic device, comprising: Including: The memory is used to store instructions executed by one or more processors of the electronic device, and the processor is one of the processors of the electronic device, used to execute the method of any one of claims 1-17.

20. A computer program product, characterised in that, The computer program product, when running on an electronic device, causes the electronic device to implement the method of any one of claims 1-17.

Citation Information

Patent Citations

  • Generative model training method and device, sample generation method and computing equipment

    CN113822320A

  • Diffusion model with improved accuracy and reduced computing resource consumption

    CN117296061A

  • Voice noise reduction method and device based on model fusion and storage medium

    CN117789744A

  • Electronic equipment and voice signal processing method thereof

    CN117809668A

  • Sound source separating device, sound source separating method, sound source separating program, and recording medium

    JP2008131183A