Voice noise reduction model training method and device, voice noise reduction method and device, equipment and medium

By performing nonlinear processing and spectrum analysis on the voice signal, the ideal masking value is enhanced to improve the learning effect of the speech noise reduction model, the speech information retention problem of low signal-to-noise frequency points is solved, and high-precision speech noise reduction is achieved.

CN120375845APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411479780.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The ideal masking value of the existing speech noise reduction model at low signal-to-noise frequency points is too low, resulting in the inability to effectively retain voice information and reduce noise reduction effect.

Method used

By nonlinear processing of the first speech signal, the second speech signal is obtained, and the ideal masking value is enhanced in combination with the first and second spectrums, and the label masking value is obtained as the learning goal of the speech noise reduction model.

Benefits of technology

The speech noise reduction model retains the speech harmonic part, improves the prediction accuracy of the speech noise reduction model, and meets the high-precision speech noise reduction needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375845A_ABST
    Figure CN120375845A_ABST
Patent Text Reader

Abstract

The invention provides a voice noise reduction model training method and device, a voice noise reduction method and device, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a first frequency spectrum of a first voice signal and a second frequency spectrum of a second voice signal; wherein the second voice signal is obtained by performing nonlinear processing on the first voice signal; obtaining an ideal masking value corresponding to the noisy voice signal; processing the ideal masking value according to the first frequency spectrum and the second frequency spectrum to obtain a labeled masking value; and training the voice noise reduction model based on the annotation masking value. Therefore, fundamental frequency enhancement processing is performed on the ideal masking value based on the first frequency spectrum and the second frequency spectrum carrying the periodic characteristics at the same time, the retention degree of the marked masking value obtained through processing on the voice harmonic part can be improved, the marked masking value serves as the learning target of the voice noise reduction model, the learning effect of the voice noise reduction model can be improved, and the voice noise reduction efficiency is improved. And the high-precision voice noise reduction requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and medium for training a voice noise reduction model and voice noise reduction. Background Art

[0002] In related technologies, the training method of a voice noise reduction model is as follows: Based on the energy ratio of a clean voice signal to a noisy voice signal, calculate the ideal mask value of the noisy voice signal, and use this ideal mask value as the learning target or training target of the voice noise reduction model.

[0003] However, the above method will result in too low an ideal mask value for low signal-to-noise ratio frequency points, and then the voice noise reduction model cannot retain the voice information of low signal-to-noise ratio frequency points during the inference stage, thereby reducing the noise reduction effect of the voice signal. Summary of the Invention

[0004] This application aims to solve at least one of the technical problems in related technologies to some extent.

[0005] To this end, this application proposes a method, apparatus, device, and medium for training a voice noise reduction model and voice noise reduction, so as to perform fundamental frequency enhancement processing on the ideal mask value, improve the retention degree of the processed labeled mask value for the harmonic part of the voice, and thus use this labeled mask value as the learning target of the voice noise reduction model, which can improve the learning effect of the voice noise reduction model, that is, improve the prediction accuracy of the voice noise reduction model and meet the high-precision voice noise reduction requirements.

[0006] The first aspect embodiment of this application proposes a method for training a voice noise reduction model, including:

[0007] Obtain the first spectrum of the first voice signal and the second spectrum of the second voice signal; wherein, the second voice signal is obtained by performing non-linear processing on the first voice signal;

[0008] Obtain the ideal mask value corresponding to the noisy voice signal; wherein, the noisy voice signal is generated according to the first voice signal and the noise signal;

[0009] Process the ideal mask value according to the first spectrum and the second spectrum to obtain a labeled mask value;

[0010] Train the voice noise reduction model based on the labeled mask value.

[0011] The second aspect embodiment of this application proposes a voice noise reduction method, including:

[0012] Obtain the target time-frequency feature of the voice signal to be processed;

[0013] Predict a masking value for the target time-frequency feature by using a trained speech noise reduction model to obtain a target masking value;

[0014] Perform noise reduction processing on the to-be-processed speech signal according to the target masking value to obtain a target speech signal.

[0015] An embodiment of the third aspect of the present application provides a training device for a speech noise reduction model, including:

[0016] A first acquisition module, configured to acquire a first spectrum of a first speech signal and a second spectrum of a second speech signal; wherein, the second speech signal is obtained by performing non-linear processing on the first speech signal;

[0017] A second acquisition module, configured to acquire an ideal masking value corresponding to a noisy speech signal; wherein, the noisy speech signal is generated according to the first speech signal and a noise signal;

[0018] A processing module, configured to process the ideal masking value according to the first spectrum and the second spectrum to obtain a labeled masking value;

[0019] A training module, configured to train a speech noise reduction model based on the labeled masking value.

[0020] An embodiment of the fourth aspect of the present application provides a speech noise reduction device, including:

[0021] An acquisition module, configured to acquire a target time-frequency feature of a to-be-processed speech signal;

[0022] A prediction module, configured to predict a masking value for the target time-frequency feature by using a trained speech noise reduction model to obtain a target masking value;

[0023] A noise reduction module, configured to perform noise reduction processing on the to-be-processed speech signal according to the target masking value to obtain a target speech signal.

[0024] An embodiment of the fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the training method of the speech noise reduction model as described in the first aspect, or implements the speech noise reduction method as described in the second aspect.

[0025] An embodiment of the sixth aspect of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the speech noise reduction model as described in the first aspect, or implements the speech noise reduction method as described in the second aspect.

[0026] A computer program product is provided in an embodiment of the seventh aspect of the present application. A computer program is stored thereon, and when the computer program is executed by a processor, it implements the method for training a voice noise reduction model as described in the foregoing first aspect, or implements the voice noise reduction method as described in the foregoing second aspect.

[0027] In another aspect of the present disclosure, a chip is provided, including an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is used to implement the method for training a voice noise reduction model according to any one of the first aspect, or implement the voice noise reduction method described in the second aspect.

[0028] For the method, device, equipment, and medium for training a voice noise reduction model and voice noise reduction proposed in the present application, since the second voice signal obtained by non-linear processing carries the periodic information of the clean first voice signal, and the second spectrum of the second voice signal carries the periodic characteristics of the first voice signal. At the same time, based on the first spectrum of the first voice signal and the second spectrum carrying the periodic characteristics, fundamental frequency enhancement processing is performed on the ideal mask value, which can improve the retention degree of the processed labeled mask value for the harmonic part of the voice. Therefore, using the labeled mask value as the learning target of the voice noise reduction model can improve the learning effect of the voice noise reduction model, that is, improve the prediction accuracy of the voice noise reduction model, and meet the high-precision voice noise reduction requirements.

[0029] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings

[0030] The above and / or additional aspects and advantages of the present application will become apparent and be easily understood from the following description of the embodiments in conjunction with the drawings, where:

[0031] Figure 1 It is a schematic diagram of the production process of the existing ideal mask value;

[0032] Figure 2 It is a schematic diagram of the flow of the first method for training a voice noise reduction model provided by an embodiment of the present application;

[0033] Figure 3 It is a schematic diagram of the flow of the second method for training a voice noise reduction model provided by an embodiment of the present application;

[0034] Figure 4 It is a schematic diagram of the flow of the third method for training a voice noise reduction model provided by an embodiment of the present application;

[0035] Figure 5 It is a schematic diagram of the flow of the fourth method for training a voice noise reduction model provided by an embodiment of the present application;

[0036] Figure 6 It is a schematic flowchart of the training method of the fifth voice noise reduction model provided by the embodiments of the present application;

[0037] Figure 7 It is a schematic flowchart of the production process of the ideal masking value combined with post-processing provided by the embodiments of the present application;

[0038] Figure 8 It is a schematic flowchart of the production process of the ideal masking value combined with voice harmonic regeneration post-processing provided by the embodiments of the present application;

[0039] Figure 9 It is a schematic flowchart of a voice noise reduction method provided by the embodiments of the present application;

[0040] Figure 10 It is a schematic structural diagram of a training device for a voice noise reduction model provided by the embodiments of the present application;

[0041] Figure 11 It is a schematic structural diagram of a voice noise reduction device provided by the embodiments of the present application;

[0042] Figure 12 It is a block diagram of an electronic device shown according to an exemplary embodiment;

[0043] Figure 13 It is a schematic structural diagram of a chip proposed by the embodiments of the present application. Detailed implementation manners

[0044] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0045] In the related art, the training principle of the voice noise reduction model is as shown in 1, mainly including the following steps:

[0046] Step 1: Add the clean voice signal and the noise signal to obtain a noisy voice signal.

[0047] Step 2: Perform FFT (Fast Fourier Transform) on the clean voice signal to obtain the spectrum of the clean voice signal ( Figure 1 denoted as the clean voice spectrum in ) S(t,f); where t represents the index of the time frame, that is, which frame; f represents the index of the frequency point (i.e., the frequency bin).

[0048] Step 3: Perform FFT transformation on the noisy voice signal to obtain the spectrum of the noisy voice signal (Figure 1 Denoted as the noisy speech spectrum) Y(t, f) in

[0049] Step 4: Perform FFT transformation on the noise signal to obtain the spectrum of the noise signal ( Figure 1 Denoted as the noise spectrum) N(t, f) in

[0050] Step 5: Calculate the ideal masking value according to S(t, f), the spectrum Y(t, f), and N(t, f).

[0051] Step 6: Use the ideal masking value as the learning target of the speech noise reduction model to train the speech noise reduction model.

[0052] However, the above method will result in too low ideal masking values for low signal-to-noise ratio frequency points, which in turn causes the speech noise reduction model to not retain the speech information of low signal-to-noise ratio frequency points during the inference stage, reducing the noise reduction effect of the speech signal.

[0053] In view of at least one of the above problems, the present application proposes a training method, a speech noise reduction method, device, equipment, and medium for a speech noise reduction model.

[0054] The following describes the training method, speech noise reduction method, device, equipment, and medium for the speech noise reduction model according to the embodiments of the present application with reference to the accompanying drawings. Before specifically describing the embodiments of the present application, for the sake of easy understanding, common technical terms are first introduced:

[0055] Nonlinear processing refers to performing a nonlinear transformation or operation on a speech signal to improve the quality of the speech signal, extract features, or achieve specific application goals.

[0056] ABS (Absolute Value) processing is used to convert all negative value parts of a speech signal into positive values, thereby only retaining the amplitude information of the speech signal while ignoring its sign. In some cases, ABS processing can be used to reduce certain types of noise, especially when the noise has positive and negative symmetry characteristics.

[0057] Logarithmic Compression processing refers to compressing the dynamic range by taking the logarithm of the amplitude of a speech signal.

[0058] Thresholding processing refers to setting the part of a speech signal that is lower or higher than a certain threshold (denoted as TH) to TH. Exemplarily, denoting the speech signal as s, the mathematical expression of thresholding processing can be: max(s, TH) or min(s, TH).

[0059] Figure 2 It is a schematic flowchart of the first training method for the speech noise reduction model provided by the embodiments of the present application.

[0060] In the embodiments of the present application, the training method of the speech noise reduction model is exemplified by being configured in a training device for the speech noise reduction model. The training device can be applied to any electronic device so that the electronic device can perform the training function of the speech noise reduction model.

[0061] Among them, the electronic device can be any device with computing power, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.

[0062] In an embodiment of the present application, the training method of the speech noise reduction model can be executed by a chip, such as an audio processing chip.

[0063] In another embodiment of the present application, the chip can be integrated into an electronic device such as a mobile phone or an AI (Artificial Intelligence) server.

[0064] As Figure 2 shown, the training method of the speech noise reduction model may include the following steps S201 to S204:

[0065] Step S201, obtaining a first spectrum of a first speech signal and a second spectrum of a second speech signal; wherein, the second speech signal is obtained by performing non-linear processing on the first speech signal.

[0066] Among them, the first speech signal includes a clean speech signal without noise. It should be noted that the present application does not limit the acquisition method of the first speech signal. For example, the first speech signal can be a clean speech signal obtained from an existing training set, or the first speech signal can be a clean speech signal provided manually by relevant personnel, or the first speech signal can be an artificially synthesized clean speech signal, and so on.

[0067] Among them, the second voice signal includes a voice signal obtained by performing non-linear processing on the first voice signal. The non-linear processing includes, but is not limited to, ABS processing, logarithmic compression processing, threshold processing, etc. Exemplarily, in order to obtain the periodic information in the first voice signal, the ABS processing can be performed on the first voice signal to obtain the second voice signal, or the ABS processing and logarithmic compression processing can be sequentially performed on the first voice signal to obtain the second voice signal, or the ABS processing and threshold processing can be sequentially performed on the first voice signal to obtain the second voice signal. In the embodiments of the present application, after the first voice signal is obtained, a relevant transformation algorithm (such as Fourier transform (such as FFT transform, STFT (Short-Time Fourier Transform)), MFCCs (Mel-Frequency Cepstral Coefficients), Gammatone Filterbank, Bark algorithm, etc.) can be used to perform transformation processing on the first voice signal to obtain the spectrum of the first voice signal (denoted as the first spectrum in the present application).

[0068] In the embodiments of the present application, after the first voice signal is obtained, the first voice signal can also be subjected to non-linear processing to obtain a second voice signal, and a relevant transformation algorithm can be used to perform transformation processing on the second voice signal to obtain the spectrum of the second voice signal (denoted as the second spectrum in the present application).

[0069] Step S202, obtain the ideal masking value corresponding to the noisy voice signal; among them, the noisy voice signal is generated according to the first voice signal and the noise signal.

[0070] Among them, the ideal masking value is used to recover the clean voice signal from the noisy voice signal.

[0071] It should be noted that the present application does not limit the acquisition method of the noise signal either. For example, the noise signal can be a noise signal provided manually by relevant personnel, or a synthetic noise signal, or a noise signal collected online. For example, the network crawler technology can be used to collect the noise signal online, and so on.

[0072] In the embodiments of the present application, after the noise signal is obtained, the first voice signal and the noise signal can be added to obtain the noisy voice signal, and the ideal masking value corresponding to the noisy voice signal can be calculated.

[0073] Step S203, process the ideal masking value according to the first spectrum and the second spectrum to obtain the labeled masking value.

[0074] In an embodiment of the present application, the first spectrum and the second spectrum can be used to perform fundamental frequency enhancement processing on the ideal masking value to obtain the labeled masking value corresponding to the noisy speech signal.

[0075] Step S204, train the speech noise reduction model based on the labeled masking value.

[0076] In an embodiment of the present application, the labeled masking value can be used as the learning target or training target of the speech noise reduction model to train the speech noise reduction model.

[0077] In the training method of the speech noise reduction model according to the embodiment of the present application, since the second speech signal obtained by non-linear processing carries the periodic information of the clean first speech signal, and the second spectrum of the second speech signal carries the periodic characteristics of the first speech signal, and at the same time, the fundamental frequency enhancement processing is performed on the ideal masking value based on the first spectrum of the first speech signal and the second spectrum carrying the periodic characteristics, the retention degree of the speech harmonic part of the obtained labeled masking value can be improved. Therefore, using the labeled masking value as the learning target of the speech noise reduction model can improve the learning effect of the speech noise reduction model, that is, improve the prediction accuracy of the speech noise reduction model, and meet the high-precision speech noise reduction requirements.

[0078] The embodiment of the present application provides another training method for the speech noise reduction model. Figure 3 It is a schematic flowchart of the second training method for the speech noise reduction model provided by the embodiment of the present application.

[0079] It should be noted that the training method of this speech noise reduction model can be executed alone, or it can be executed in combination with any one of the embodiments in the present application or the possible implementation manners in the embodiments, or it can also be executed in combination with any one of the technical solutions in the related technologies. The embodiments of the present application do not limit this.

[0080] As Figure 3 shown, the training method of this speech noise reduction model may include the following steps S301 to S306:

[0081] Step S301, obtain the first spectrum of the first speech signal and the second spectrum of the second speech signal.

[0082] Among them, the second speech signal is obtained by performing non-linear processing on the first speech signal.

[0083] Step S302, obtain the ideal masking value corresponding to the noisy speech signal; among them, the noisy speech signal is generated according to the first speech signal and the noise signal.

[0084] Step S303, process the ideal masking value according to the first spectrum and the second spectrum to obtain the labeled masking value.

[0085] For the explanatory notes of steps S301 to S303, reference can be made to the relevant descriptions in any embodiment of this application, which will not be elaborated here.

[0086] Step S304: Obtain the time-frequency features of the noisy speech signal.

[0087] Among them, the time-frequency features include but are not limited to: amplitude, spectrum, compressed amplitude, compressed spectrum, etc.

[0088] In the embodiments of this application, feature extraction can be performed on the noisy speech signal to obtain the time-frequency features of the noisy speech signal.

[0089] As an example, taking the spectrum of the noisy speech signal (denoted as the fourth spectrum in this application) as the time-frequency feature for example, relevant transformation algorithms can be used to perform transformation processing on the noisy speech signal to obtain the fourth spectrum of the noisy speech signal.

[0090] Step S305: Use the speech denoising model to predict the masking value of the time-frequency features to obtain the predicted masking value.

[0091] In the embodiments of this application, the time-frequency features of the noisy speech signal can be input into the speech denoising model to use the speech denoising model to predict the masking value of the time-frequency features and obtain the predicted masking value output by the speech denoising model.

[0092] Step S306: Train the speech denoising model according to the difference between the labeled masking value and the predicted masking value.

[0093] In the embodiments of this application, the value of the loss function (denoted as the loss value in this application) can be determined according to the difference between the labeled masking value and the predicted masking value. Among them, the loss value is positively correlated with the above difference, that is, the greater the difference, the greater the loss value, and vice versa, the smaller the difference, the smaller the loss value. Therefore, in this application, the speech denoising model can be trained according to the above loss value to minimize the loss value.

[0094] It should be noted that the above only takes the termination condition of model training as the minimization of the loss value for example. In actual applications, other termination conditions can also be set. For example, the termination conditions can also include: the training duration reaches the set duration, the number of training rounds reaches the set number of rounds, and so on.

[0095] The training method of the speech denoising model in the embodiments of this application takes the labeled masking value as the learning target or training target of the speech denoising model, which can improve the learning effect of the speech denoising model and further improve the prediction accuracy of the speech denoising model.

[0096] The embodiments of this application provide another training method for the speech denoising model. Figure 4Schematic flowchart of the training method for the third voice noise reduction model provided by the embodiments of the present application.

[0097] It should be noted that the training method of the voice noise reduction model can be executed alone, or it can also be executed in combination with any one of the embodiments or possible implementation manners in the present application, or it can also be executed in combination with any one of the technical solutions in the related art. The embodiments of the present application do not limit this.

[0098] As Figure 4 shown, the training method of the voice noise reduction model may include the following steps S401 to S406:

[0099] Step S401, obtain the first spectrum of the first voice signal and the second spectrum of the second voice signal.

[0100] Wherein, the second voice signal is obtained by non-linearly processing the first voice signal.

[0101] Step S402, obtain the ideal masking value corresponding to the noisy voice signal; wherein, the noisy voice signal is generated according to the first voice signal and the noise signal.

[0102] For the explanatory description of steps S401 to S402, reference can be made to the relevant descriptions in any embodiment of the present application, and details are not described herein.

[0103] Step S403, determine the first masking component according to the product of the first spectrum and the ideal masking value.

[0104] As an example, the product of the first spectrum and the ideal masking value can be used as the first masking component. For example, mark the ideal masking value as M(t,f) (abbreviated as M), the first voice signal as s(n), and the spectrum of s(n) (i.e., the first spectrum) as S(t,f), then the first masking component can be: S(t,f)*M.

[0105] Wherein, n represents the time index, t represents the index of the time frame, that is, which frame, and f represents the index of the frequency point.

[0106] Step S404, determine the second masking component according to the product of the second spectrum and the target difference; wherein, the target difference is the difference between the specified value and the ideal masking value.

[0107] Wherein, the specified value is a preset value. Exemplarily, the specified value can be 1.

[0108] As an example, the product of the second spectrum and the target difference can be used as the second masking component. For example, taking the specified value as 1 for illustration, marking the second speech signal as NL(s(n)), the spectrum of NL(s(n)) (i.e., the second spectrum) is NL(S(t,f)), then the second masking component can be: NL(S(t,f)) * (1 - M).

[0109] Where NL refers to the abbreviation of Nonlinear processing.

[0110] Step S405: Determine the labeled masking value according to the first masking component and the second masking component.

[0111] In the embodiments of the present application, the first masking component and the second masking component can be combined to calculate the labeled masking value.

[0112] In any embodiment of the present application, the labeled masking value can be calculated by the following steps A to C:

[0113] Step A: Obtain the third spectrum of the noise signal.

[0114] As an example, a correlation transformation algorithm can be used to perform transformation processing on the noise signal to obtain the spectrum of the noise signal (denoted as the third spectrum in this application).

[0115] Step B: Determine the third masking component according to the sum of the first masking component and the second masking component.

[0116] Step C: Determine the labeled masking value according to the ratio of the third masking component to the third spectrum.

[0117] As an example, marking the noise signal as n(n), the spectrum of n(n) (i.e., the third spectrum) is N(t,f), then the labeled masking value can be calculated by the following formula:

[0118]

[0119] It should be noted that the above calculation method is only exemplary, and those skilled in the art can also set other calculation formulas according to the actual situation. For example, those skilled in the art can also add some correction factors to the above calculation formula. Such changes in the specific calculation method do not deviate from the basic principle of the present application and belong to the protection scope of the present application.

[0120] Step S406: Train the speech denoising model based on the labeled masking value.

[0121] For the explanation of Step S406, reference can be made to the relevant descriptions in any embodiment of the present application, and details will not be elaborated here.

[0122] The training method of the voice noise reduction model according to the embodiments of the present application can effectively perform fundamental frequency enhancement processing on the ideal masking value by using the first spectrum and the second spectrum, and improve the retention degree of the enhanced masking value for the harmonic part of the voice.

[0123] The embodiments of the present application provide another training method of the voice noise reduction model. Figure 5 It is a schematic flowchart of the fourth training method of the voice noise reduction model provided by the embodiments of the present application.

[0124] It should be noted that the training method of the voice noise reduction model can be executed alone, or can be executed in combination with any one of the embodiments or possible implementation manners in the present application, or can also be executed in combination with any one of the technical solutions in the related art. The embodiments of the present application do not limit this.

[0125] As shown in Figure 5 , the training method of the voice noise reduction model may include the following steps S501 to S506:

[0126] Step S501: Obtain the first spectrum of the first voice signal and the second spectrum of the second voice signal.

[0127] Among them, the second voice signal is obtained by performing non-linear processing on the first voice signal.

[0128] For the explanation of step S501, reference can be made to the relevant descriptions in any embodiment of the present application, and details are not described herein again.

[0129] Step S502: Calculate the first energy distribution of the first voice signal according to the first spectrum.

[0130] As an example, the first energy distribution may be: |S(t,f) 2 .

[0131] Step S503: Calculate the second energy distribution of the noise signal according to the third spectrum of the noise signal.

[0132] As an example, the second energy distribution may be: |N(t,f) 2 .

[0133] Step S504: Determine the ideal masking value corresponding to the noisy voice signal according to the first energy distribution and the second energy distribution.

[0134] Among them, the noisy voice signal is generated according to the first voice signal and the noise signal.

[0135] In the embodiments of the present application, the ideal masking value corresponding to the noisy voice signal can be calculated according to the first energy distribution and the second energy distribution.

[0136] In any embodiment of the present application, first, a first coefficient can be calculated according to the first energy distribution and the second energy distribution, where the first coefficient is used to characterize the quality of the noisy speech signal. For example, the difference between the first energy distribution and the second energy distribution can be used as the first coefficient. Then, according to the magnitude relationship between the first coefficient and the set energy threshold, the ideal masking value corresponding to the noisy speech signal can be determined.

[0137] Among them, the set energy threshold refers to the pre-set energy threshold. Exemplarily, the set energy threshold can be 0.

[0138] As an example, marking the set energy threshold as θ1, the following formula can be used to calculate the ideal masking value corresponding to the noisy speech signal:

[0139]

[0140] That is, when the energy of the speech signal minus the energy of the noise is greater than the threshold θ1, the value of M(t,f) is 1; otherwise, the value of M(t,f) is 0. M(t,f) in formula (2) can be called IBM (Ideal binary mask).

[0141] In any embodiment of the present application, the third energy distribution can also be determined according to the sum of the first energy distribution and the second energy distribution. Then, according to the first energy distribution and the third energy distribution, a second coefficient is determined, where the second coefficient is used to characterize the relative difference between the first energy distribution and the third energy distribution. For example, the ratio of the first energy distribution to the third energy distribution can be used as the second coefficient. Furthermore, based on the adjustment factor, the second coefficient can be adjusted to obtain the ideal masking value corresponding to the noisy speech signal.

[0142] Among them, the adjustment factor is an adjustable parameter. Exemplarily, the adjustment factor can be set to 0.5.

[0143] As an example, marking the adjustment factor as β, the following formula can be used to calculate the ideal masking value corresponding to the noisy speech signal:

[0144]

[0145] Among them, the value of M(t,f) reflects the power of the ratio of the energy of the speech signal to the energy of the noise, and its value is a real number between 0 and 1. When the energy of the speech signal is much greater than the energy of the noise, M(t,f) approaches 1; when the energy of the speech signal is close to the energy of the noise, M(t,f) approaches 0. M(t,f) in formula (3) can be called IRM (Ideal ratio mask).

[0146] It should be noted that the ideal masking value in formula (3) provides continuous weights in the time-frequency domain instead of hard segmentation. This soft decision-making method can better maintain the continuity and integrity of the speech signal while reducing the influence of noise.

[0147] Step S505: Process the ideal masking value according to the first spectrum and the second spectrum to obtain a labeled masking value.

[0148] Step S506: Train the speech noise reduction model based on the labeled masking value.

[0149] For the explanations of steps S505 to S506, reference can be made to the relevant descriptions in any embodiment of this application, which will not be elaborated here.

[0150] The training method of the speech noise reduction model in the embodiments of this application can effectively calculate the ideal masking value corresponding to the noisy speech signal based on the energy distribution of the clean speech signal and the energy distribution of the noise signal.

[0151] The embodiments of this application provide another training method for the speech noise reduction model. Figure 6 It is a schematic flowchart of the fifth training method for the speech noise reduction model provided by the embodiments of this application.

[0152] It should be noted that the training method of this speech noise reduction model can be executed alone, or it can also be executed in combination with any one embodiment in this application or the possible implementation manners in the embodiments, or it can also be executed in combination with any one technical solution in the related technologies. The embodiments of this application do not limit this.

[0153] As Figure 6 shown, the training method of this speech noise reduction model may include the following steps S601 to S605:

[0154] Step S601: Obtain the first spectrum of the first speech signal and the second spectrum of the second speech signal.

[0155] Among them, the second speech signal is obtained by performing non-linear processing on the first speech signal.

[0156] Step S602: Obtain the fourth spectrum of the noisy speech signal.

[0157] Among them, the noisy speech signal is generated according to the first speech signal and the noise signal.

[0158] For the explanations of steps S601 to S602, reference can be made to the relevant descriptions in any embodiment of this application, which will not be elaborated here.

[0159] Step S603: Determine the ideal masking value corresponding to the noisy speech signal according to the first spectrum and the fourth spectrum.

[0160] In the embodiments of the present application, the ideal masking value corresponding to the noisy speech signal can be calculated by integrating the first spectrum and the fourth spectrum.

[0161] In any embodiment of the present application, the ideal masking value can be calculated by the following steps a to d:

[0162] Step a: Take the sum of the first product and the second product as the third coefficient; where the first product is the product of the first real part of the first spectrum and the second real part of the fourth spectrum; the second product is the product of the first imaginary part of the first spectrum and the second imaginary part of the fourth spectrum.

[0163] As an example, mark the noisy speech signal as y(n), the spectrum of y(n) (i.e., the fourth spectrum) as Y(t,f), the real part of Y(t,f) (denoted as the second real part in this application) as Y r , the imaginary part of Y(t,f) (denoted as the second imaginary part in this application) as Y i , the real part of the first spectrum S(t,f) (denoted as the first real part in this application) as S r , the imaginary part of S(t,f) (denoted as the first imaginary part in this application) as S i , then the third coefficient can be: Y r S r +Y i S i .

[0164] Step b: Take the difference between the third product and the second real part as the fourth coefficient; where the third product is the product of the second real part and the first imaginary part.

[0165] As an example, the fourth coefficient can be: Y r S i -Y r .

[0166] Step c: Determine the fifth coefficient according to the fourth spectrum; where the fifth coefficient is used to characterize the first modulus of the fourth spectrum.

[0167] As an example, the first modulus can be: The fifth coefficient can be:

[0168] Step d: Determine the real part of the ideal masking value according to the ratio of the third coefficient to the fifth coefficient, and determine the imaginary part of the ideal masking value according to the ratio of the fourth coefficient to the fifth coefficient.

[0169] As an example, the following formula can be used to calculate the ideal masking value M:

[0170]

[0171] The M in formula (4) can be called cIRM (Complex Ideal ratio mask). cIRM combines the amplitude and phase information of the speech signal, making the separated speech signal not only have a better signal-to-noise ratio, but also can retain the natural characteristics of the original clean speech signal as much as possible. Compared with the mask based only on amplitude, cIRM can better recover high-quality speech signals. In practical applications, cIRM can be applied to multiple scenarios such as speech enhancement and speech separation, especially to improve speech quality in noisy environments.

[0172] In any embodiment of the present application, the ideal mask value can also be calculated by the following steps a' to c':

[0173] Step a': Obtain the phase difference between the first spectrum and the fourth spectrum. For example, mark this phase difference as θ2, then there is: θ2 = θ S + θ Y , where θ S is the phase of S(t,f), and θ Y is the phase of Y(t,f).

[0174] Step b': Determine the sixth coefficient according to the second modulus length of the first spectrum and the third modulus length of the fourth spectrum; where the sixth coefficient is used to characterize the relative difference between the second modulus length and the third modulus length.

[0175] As an example, the ratio of the second modulus length to the third modulus length can be used as the sixth coefficient.

[0176] Step c': Determine the ideal mask value corresponding to the noisy speech signal according to the sixth coefficient and the phase difference.

[0177] As an example, the following formula can be used to calculate the ideal mask value M(t,f):

[0178]

[0179] Among them, M(t,f) reflects the ratio of the amplitude of the clean speech signal to the amplitude of the noisy speech signal, as well as the phase relationship between the two. M(t,f) in formula (4) can be called PSM (Phase sensitive mask), and PSM directly utilizes the phase information of the speech signal, which can more accurately separate the clean speech signal. Since the phase information plays an important role in speech signals, especially in acoustic modeling and speech synthesis, PSM can provide a better speech restoration effect. PSM takes into account the phase difference, making the separated speech signal closer to the original clean speech signal, thereby improving the naturalness and intelligibility of the speech.

[0180] Step S604: Process the ideal mask value according to the first spectrum and the second spectrum to obtain the labeled mask value.

[0181] Step S605: Train the speech noise reduction model based on the labeled mask value.

[0182] For the explanatory descriptions of steps S604 to S605, reference can be made to the relevant descriptions in any embodiment of this application, and details will not be elaborated here.

[0183] The training method of the speech noise reduction model in the embodiments of this application can effectively calculate the ideal mask value corresponding to the noisy speech signal based on the spectrum of the clean speech signal and the spectrum of the noisy speech signal.

[0184] In any embodiment of this application, different from Figure 1 the production process of the ideal mask value shown, this application proposes a production process of the ideal mask value combined with post-processing as shown in Figure 7 and a production process of the ideal mask value combined with speech harmonic regeneration post-processing as shown in Figure 8 to improve the retention degree of the speech harmonic part of the produced labeled mask value, thereby improving the learning effect of the speech noise reduction model. Among them, the production process of the ideal mask value combined with speech harmonic regeneration post-processing mainly includes the following parts:

[0185] The first part: Produce the ideal mask value M of the noisy speech signal according to formula (2), formula (3), formula (4), or formula (5).

[0186] The second part: Perform post-processing on the ideal mask value M to obtain the labeled mask value M new .

[0187] The third part: Use the labeled mask value M new as the learning target of the speech noise reduction model.

[0188] As an example, the entire processing flow is:

[0189] 1) Perform FFT transformation on the clean speech signal, noise signal, and noisy speech signal respectively to obtain the spectrum of the clean speech signal, the spectrum of the noise signal, and the spectrum of the noisy speech signal;

[0190] 2) According to the above spectra, calculate the ideal masking value M of the noisy speech signal according to formula (2), formula (3), formula (4), or formula (5);

[0191] 3) Perform non-linear processing (such as ABS) on the clean speech signal and calculate the spectrum of the clean speech signal after non-linear processing;

[0192] 4) Calculate the updated ideal masking value according to formula (1) to obtain the labeled masking value;

[0193] 5) Use the labeled masking value M new as the learning target or training target of the speech noise reduction model.

[0194] The above are the respective embodiments corresponding to the training method of the speech noise reduction model. The present application also proposes an application method of the speech noise reduction model, that is, a speech noise reduction method.

[0195] Figure 9 It is a schematic flowchart of a speech noise reduction method provided by an embodiment of the present application.

[0196] As Figure 9 shown, the speech noise reduction method may include the following steps S901 to S903:

[0197] Step S901, obtain the target time-frequency feature of the speech signal to be processed.

[0198] Among them, there is no limitation on the acquisition method of the speech signal to be processed. For example, the speech signal to be processed may be a received noisy speech signal, a real-time collected noisy speech signal, an artificially synthesized noisy speech signal, and so on.

[0199] Among them, the target time-frequency feature includes but is not limited to: amplitude, spectrum, compressed amplitude, compressed spectrum, etc.

[0200] In the embodiment of the present application, after the speech signal to be processed is obtained, the speech signal to be processed may be subjected to feature extraction to obtain the target time-frequency feature.

[0201] As an example, taking the target video feature as the spectrum of the speech signal to be processed (denoted as the first spectrum in the present application) for example, a relevant transformation algorithm may be used to perform transformation processing on the speech signal to be processed to obtain the first spectrum.

[0202] Step S902: Use the trained voice noise reduction model to predict the masking value of the target time-frequency feature, and obtain the target masking value.

[0203] Among them, the voice noise reduction model is a pre-trained model. Exemplarily, the training method of the voice noise reduction model provided in any of the above embodiments can be used to pre-train the voice noise reduction model.

[0204] In the embodiment of the present application, the target time-frequency feature of the voice signal to be processed can be input into the voice noise reduction model, so as to use the voice noise reduction model to predict the masking value of the target time-frequency feature, and obtain the target masking value output by the voice noise reduction model.

[0205] Step S903: Perform noise reduction processing on the voice signal to be processed according to the target masking value, and obtain the target voice signal.

[0206] In the embodiment of the present application, the target masking value can be used to perform noise reduction processing on the voice signal to be processed, and obtain the target voice signal.

[0207] In any embodiment of the present application, first, the target masking value can be used to perform noise reduction and enhancement processing on the first spectrum of the voice signal to be processed, and obtain the second spectrum. For example, the target masking value can be multiplied element by element with the first spectrum to obtain the enhanced voice spectrum, which is denoted as the second spectrum in the present application. After that, inverse transformation processing (such as inverse Fourier transform processing, etc.) can be performed on the second spectrum to obtain the target voice signal.

[0208] The voice noise reduction method in the embodiment of the present application uses deep learning technology to predict the target masking value of the voice signal to be processed, and based on this target masking value, a clean voice signal is recovered from the voice signal to be processed, which can improve the quality of the voice signal.

[0209] Corresponding to the Figures 2 to 8 training method of the voice noise reduction model provided in the above embodiment, the present application also provides a training device for the voice noise reduction model. Since the training device for the voice noise reduction model provided in the embodiment of the present application corresponds to the Figures 2 to 8 training method of the voice noise reduction model provided in the above embodiment, the implementation manner of the training method of the voice noise reduction model is also applicable to the training device for the voice noise reduction model provided in the embodiment of the present application, and will not be described in detail in the embodiment of the present application.

[0210] Figure 10 It is a structural schematic diagram of a training device for a voice noise reduction model provided in an embodiment of the present application.

[0211] As Figure 10As shown in the figure, the training device 1000 of the voice noise reduction model may include: a first acquisition module 1010, a second acquisition module 1020, a processing module 1030, and a training module 1040.

[0212] Among them, the first acquisition module 1010 is used to acquire the first spectrum of the first voice signal and the second spectrum of the second voice signal; wherein, the second voice signal is obtained by performing non-linear processing on the first voice signal;

[0213] The second acquisition module 1020 is used to acquire the ideal masking value corresponding to the noisy voice signal; wherein, the noisy voice signal is generated according to the first voice signal and the noise signal;

[0214] The processing module 1030 is used to process the ideal masking value according to the first spectrum and the second spectrum to obtain the labeled masking value;

[0215] The training module 1040 is used to train the voice noise reduction model based on the labeled masking value.

[0216] Further, in an implementation manner of the embodiment of the present application, the processing module 1030 is used to: determine the first masking component according to the product of the first spectrum and the ideal masking value; determine the second masking component according to the product of the second spectrum and the target difference; wherein, the target difference is the difference between the specified value and the ideal masking value; determine the labeled masking value according to the first masking component and the second masking component.

[0217] In an implementation manner of the embodiment of the present application, the processing module 1030 is used to: acquire the third spectrum of the noise signal; determine the third masking component according to the sum of the first masking component and the second masking component; determine the labeled masking value according to the ratio of the third masking component to the third spectrum.

[0218] In an implementation manner of the embodiment of the present application, the training module 1040 is used to: acquire the time-frequency feature of the noisy voice signal; predict the masking value of the time-frequency feature by using the voice noise reduction model to obtain the predicted masking value; train the voice noise reduction model according to the difference between the labeled masking value and the predicted masking value.

[0219] In an implementation manner of the embodiment of the present application, the non-linear processing includes at least one of the following: absolute value processing; logarithmic compression processing; threshold processing.

[0220] In an implementation manner of the embodiment of the present application, the second acquisition module 1020 is used to: calculate the first energy distribution of the first voice signal according to the first spectrum; calculate the second energy distribution of the noise signal according to the third spectrum of the noise signal; determine the ideal masking value corresponding to the noisy voice signal according to the first energy distribution and the second energy distribution.

[0221] In an implementation manner of the embodiment of the present application, the second acquisition module 1020 is configured to: determine a first coefficient according to a first energy distribution and a second energy distribution; wherein the first coefficient is used to characterize the quality of the noisy speech signal; determine an ideal masking value corresponding to the noisy speech signal according to the magnitude relationship between the first coefficient and a set energy threshold.

[0222] In an implementation manner of the embodiment of the present application, the second acquisition module 1020 is configured to: determine a third energy distribution according to the sum of the first energy distribution and the second energy distribution; determine a second coefficient according to the first energy distribution and the third energy distribution; wherein the second coefficient is used to characterize the relative difference between the first energy distribution and the third energy distribution; adjust the second coefficient based on an adjustment factor to obtain an ideal masking value corresponding to the noisy speech signal.

[0223] In an implementation manner of the embodiment of the present application, the second acquisition module 1020 is configured to: acquire a fourth spectrum of the noisy speech signal; determine an ideal masking value corresponding to the noisy speech signal according to the first spectrum and the fourth spectrum.

[0224] In an implementation manner of the embodiment of the present application, the second acquisition module 1020 is configured to: use the sum of a first product and a second product as a third coefficient; wherein the first product is the product of the first real part of the first spectrum and the second real part of the fourth spectrum; the second product is the product of the first imaginary part of the first spectrum and the second imaginary part of the fourth spectrum; use the difference between a third product and the second real part as a fourth coefficient; wherein the third product is the product of the second real part and the first imaginary part; determine a fifth coefficient according to the fourth spectrum; wherein the fifth coefficient is used to characterize the first modulus length of the fourth spectrum; determine the real part of the ideal masking value according to the ratio of the third coefficient to the fifth coefficient, and determine the imaginary part of the ideal masking value according to the ratio of the fourth coefficient to the fifth coefficient.

[0225] In an implementation manner of the embodiment of the present application, the second acquisition module 1020 is configured to: acquire a phase difference between the first spectrum and the fourth spectrum; determine a sixth coefficient according to the second modulus length of the first spectrum and the third modulus length of the fourth spectrum; wherein the sixth coefficient is used to characterize the relative difference between the second modulus length and the third modulus length; determine an ideal masking value corresponding to the noisy speech signal according to the sixth coefficient and the phase difference.

[0226] In the training device of the voice noise reduction model according to the embodiment of the present application, since the second voice signal obtained by non-linear processing carries the periodic information of the clean first voice signal, the second spectrum of the second voice signal carries the periodic characteristics of the first voice signal. At the same time, based on the first spectrum of the first voice signal and the second spectrum carrying the periodic characteristics, the fundamental frequency enhancement processing is performed on the ideal masking value, which can improve the retention degree of the processed labeled masking value for the voice harmonic part. Therefore, taking the labeled masking value as the learning target of the voice noise reduction model can improve the learning effect of the voice noise reduction model, that is, improve the prediction accuracy of the voice noise reduction model and meet the high-precision voice noise reduction requirements.

[0227] Corresponding to Figure 9 the voice noise reduction method provided in the above Figure 9 embodiment, the present application also provides a voice noise reduction device. Since the voice noise reduction device provided in the embodiment of the present application corresponds to

[0228] Figure 11 the voice noise reduction method provided in the above

[0229] As Figure 11 shown, the voice noise reduction device 1100 may include: an acquisition module 1110, a prediction module 1120, and a noise reduction module 1130.

[0230] Among them, the acquisition module 1110 is configured to acquire the target time-frequency feature of the voice signal to be processed;

[0231] The prediction module 1120 is configured to predict the masking value of the target time-frequency feature by using the trained voice noise reduction model to obtain the target masking value;

[0232] The noise reduction module 1130 is configured to perform noise reduction processing on the voice signal to be processed according to the target masking value to obtain the target voice signal.

[0233] In an implementation manner of the embodiment of the present application, the noise reduction module 1130 is configured to: acquire the first spectrum of the voice signal to be processed; perform noise reduction and enhancement processing on the first spectrum by using the target masking value to obtain the second spectrum; and perform inverse transformation processing on the second spectrum to obtain the target voice signal.

[0234] The voice noise reduction device according to the embodiment of the present application uses deep learning technology to predict the target masking value of the voice signal to be processed. Based on the target masking value, a clean voice signal is recovered from the voice signal to be processed, which can improve the quality of the voice signal.

[0235] To implement the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the training method or the voice noise reduction method of the voice noise reduction model as described in any of the foregoing embodiments is implemented.

[0236] To implement the above embodiments, the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the training method or the voice noise reduction method of the voice noise reduction model as described in any of the foregoing embodiments is implemented.

[0237] To implement the above embodiments, the present application further provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the training method or the voice noise reduction method of the voice noise reduction model as described in any of the foregoing embodiments is implemented.

[0238] Figure 12 It is a block diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 1200 may be a mobile phone, a computer, a digital broadcast device, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0239] Referring to Figure 12 , the electronic device 1200 may include one or more of the following components: a processing component 1202, a memory 1204, a power component 1206, a multimedia component 1208, an audio component 1210, an input / output (I / O) interface 1212, a sensor component 1214, and a communication component 1216.

[0240] The processing component 1202 generally controls the overall operation of the electronic device 1200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1202 may include one or more modules to facilitate the interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate the interaction between the multimedia component 1208 and the processing component 1202.

[0241] The memory 1204 is configured to store various types of data to support the operation of the electronic device 1200. Examples of such data include instructions for any application or method operating on the electronic device 1200, contact data, phone book data, messages, pictures, videos, and the like. The memory 1204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0242] The power component 1206 provides power to various components of the electronic device 1200. The power component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 1200.

[0243] The multimedia component 1208 includes a screen that provides an output interface between the electronic device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the electronic device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0244] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 further includes a speaker for outputting audio signals.

[0245] The I / O interface 1212 provides an interface between the processing component 1202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0246] The sensor assembly 1214 includes one or more sensors for providing status assessment of various aspects for the electronic device 1200. For example, the sensor assembly 1214 can detect the on / off state of the electronic device 1200, the relative positioning of components, such as the display and keypad of the electronic device 1200. The sensor assembly 1214 can also detect a change in the position of the electronic device 1200 or a component of the electronic device 1200, the presence or absence of user contact with the electronic device 1200, the orientation or acceleration / deceleration of the electronic device 1200, and the temperature change of the electronic device 1200. The sensor assembly 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1214 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0247] The communication component 1216 is configured to facilitate communication between the electronic device 1200 and other devices in a wired or wireless manner. The electronic device 1200 can access a wireless network based on communication standards, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0248] In an exemplary embodiment, the electronic device 1200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.

[0249] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 1204 including instructions, is also provided. The above instructions can be executed by the processor 1220 of the electronic device 1200 to complete the above methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0250] To implement the above embodiments, the present application further provides a chip, which includes an interface circuit and a processing circuit that are coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is used to implement the training method of the voice noise reduction model proposed in any of the above embodiments, or to implement the voice noise reduction method proposed in any of the above embodiments.

[0251] As an example, Figure 13 is a schematic structural diagram of a chip proposed in an embodiment of the present application. Reference can be made to Figure 13 the schematic structural diagram of the chip 1300 shown, but not limited thereto.

[0252] The chip 1300 includes a processing circuit 1301, and the processing circuit 1301 is configured to execute the method proposed in any of the above embodiments.

[0253] In some embodiments, the chip 1300 further includes one or more interface circuits 1302. Optionally, the interface circuit 1302 is connected to the memory 1303. The interface circuit 1302 can be used to receive signals from the memory 1303 or other devices, and the interface circuit 1302 can be used to send signals to the memory 1303 or other devices. For example, the interface circuit 1302 can read the instructions stored in the memory 1303 and send the instructions to the processing circuit 1301.

[0254] In some embodiments, the interface circuit 1302 executes at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 1301 executes other steps.

[0255] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. can be replaced with each other.

[0256] In some embodiments, the chip 1300 further includes one or more memories 1303 for storing instructions. Optionally, all or part of the memories 1303 can be outside the chip 1300.

[0257] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0258] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0259] Any process or method description represented in a flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0260] The logic and / or steps represented in a flowchart or described otherwise herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function and may be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary to obtain the program in electronic form and then storing it in a computer memory.

[0261] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques well known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0262] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0263] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0264] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A training method for a voice noise reduction model, characterized in that, Including: Obtaining a first spectrum of a first speech signal and a second spectrum of a second speech signal; wherein, the second speech signal is obtained by performing non-linear processing on the first speech signal; Obtaining an ideal masking value corresponding to a noisy speech signal; wherein, the noisy speech signal is generated according to the first speech signal and a noise signal; Processing the ideal masking value according to the first spectrum and the second spectrum to obtain a labeled masking value; Training a speech noise reduction model based on the labeled masking value.

2. The method according to claim 1, characterized in that, The processing the ideal masking value according to the first spectrum and the second spectrum to obtain a labeled masking value includes: Determining a first masking component according to the product of the first spectrum and the ideal masking value; Determining a second masking component according to the product of the second spectrum and a target difference; wherein, the target difference is the difference between a specified value and the ideal masking value; Determining the labeled masking value according to the first masking component and the second masking component.

3. The method according to claim 2, wherein The determining the labeled masking value according to the first masking component and the second masking component includes: Obtaining a third spectrum of the noise signal; Determining a third masking component according to the sum of the first masking component and the second masking component; Determining the labeled masking value according to the ratio of the third masking component to the third spectrum.

4. The method according to claim 1, wherein The training the speech noise reduction model based on the labeled masking value includes: Obtaining the time-frequency features of the noisy speech signal; Using the speech noise reduction model to predict a masking value for the time-frequency features to obtain a predicted masking value; Training the speech noise reduction model according to the difference between the labeled masking value and the predicted masking value.

5. The method according to claim 1, wherein The non-linear processing includes at least one of the following: absolute value processing; logarithmic compression processing; threshold processing.

6. The method according to any one of claims 1-5, characterized in that, The obtaining the ideal masking value corresponding to the noisy speech signal includes: Calculating a first energy distribution of the first speech signal according to the first spectrum; Calculating a second energy distribution of the noise signal according to the third spectrum of the noise signal; Determining the ideal masking value corresponding to the noisy speech signal according to the first energy distribution and the second energy distribution.

7. The method according to claim 6, characterized in that, The determining the ideal masking value corresponding to the noisy speech signal according to the first energy distribution and the second energy distribution includes: Determining a first coefficient according to the first energy distribution and the second energy distribution; wherein, the first coefficient is used to characterize the quality of the noisy speech signal; Determining the ideal masking value corresponding to the noisy speech signal according to the magnitude relationship between the first coefficient and a set energy threshold.

8. The method according to claim 6, characterized in that, The determining the ideal masking value corresponding to the noisy speech signal according to the first energy distribution and the second energy distribution includes: Determining a third energy distribution according to the sum of the first energy distribution and the second energy distribution; Determining a second coefficient according to the first energy distribution and the third energy distribution; wherein, the second coefficient is used to characterize the relative difference between the first energy distribution and the third energy distribution; Adjust the second coefficient based on the adjustment factor to obtain the ideal masking value corresponding to the noisy speech signal.

9. The method according to any one of claims 1-5, characterized in that, The obtaining of the ideal masking value corresponding to the noisy speech signal includes: Obtain the fourth spectrum of the noisy speech signal; Determine the ideal masking value corresponding to the noisy speech signal according to the first spectrum and the fourth spectrum.

10. The method according to claim 9, wherein The determining of the ideal masking value corresponding to the noisy speech signal according to the first spectrum and the fourth spectrum includes: Take the sum of the first product and the second product as the third coefficient; wherein, the first product is the product of the first real part of the first spectrum and the second real part of the fourth spectrum; the second product is the product of the first imaginary part of the first spectrum and the second imaginary part of the fourth spectrum; Take the difference between the third product and the second real part as the fourth coefficient; wherein, the third product is the product of the second real part and the first imaginary part; Determine the fifth coefficient according to the fourth spectrum; wherein, the fifth coefficient is used to characterize the first modulus length of the fourth spectrum; Determine the real part of the ideal masking value according to the ratio of the third coefficient to the fifth coefficient, and determine the imaginary part of the ideal masking value according to the ratio of the fourth coefficient to the fifth coefficient.

11. The method according to claim 9, wherein The determining of the ideal masking value corresponding to the noisy speech signal according to the first spectrum and the fourth spectrum includes: Obtain the phase difference between the first spectrum and the fourth spectrum; Determine the sixth coefficient according to the second modulus length of the first spectrum and the third modulus length of the fourth spectrum; wherein, the sixth coefficient is used to characterize the relative difference between the second modulus length and the third modulus length; Determine the ideal masking value corresponding to the noisy speech signal according to the sixth coefficient and the phase difference.

12. A method for voice noise reduction, characterized in that, Includes: Obtain the target time-frequency feature of the speech signal to be processed; Use the trained speech noise reduction model to predict the masking value of the target time-frequency feature to obtain the target masking value; Perform noise reduction processing on the speech signal to be processed according to the target masking value to obtain the target speech signal.

13. The method according to claim 12, characterized in that, The performing of noise reduction processing on the speech signal to be processed according to the target masking value to obtain the target speech signal includes: Obtain the first spectrum of the speech signal to be processed; Use the target masking value to perform noise reduction and enhancement processing on the first spectrum to obtain the second spectrum; Perform inverse transformation processing on the second spectrum to obtain the target speech signal.

14. A training device for a voice noise reduction model, characterized in that, Includes: The first obtaining module is used to obtain the first spectrum of the first speech signal and the second spectrum of the second speech signal; wherein, the second speech signal is obtained by performing non-linear processing on the first speech signal; The second obtaining module is used to obtain the ideal masking value corresponding to the noisy speech signal; wherein, the noisy speech signal is generated according to the first speech signal and the noise signal; The processing module is used to process the ideal masking value according to the first spectrum and the second spectrum to obtain the labeled masking value; The training module is used to train the speech noise reduction model based on the labeled masking value.

15. A voice noise reduction device, characterized in that, Includes: The obtaining module is used to obtain the target time-frequency feature of the speech signal to be processed; A prediction module, configured to predict a masking value for the target time-frequency feature by using a trained speech noise reduction model, so as to obtain a target masking value; A noise reduction module, configured to perform noise reduction processing on the to-be-processed speech signal according to the target masking value, so as to obtain a target speech signal.

16. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1 to 11, or implement the steps of the method according to any one of claims 12 to 13.

17. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the processor to execute the steps of the method according to any one of claims 1 to 11, or execute the steps of the method according to any one of claims 12 to 13.

18. A computer program product, characterized in that, Comprising a computer program, when the computer program is executed by a processor, implementing the method according to any one of claims 1 to 11, or implementing the method according to any one of claims 12 to 13.

19. A chip, comprising an interface circuit and a processing circuit which are coupled to each other, the interface circuit is used for inputting or outputting signals, and the processing circuit is used for implementing the method according to any one of claims 1 to 11, or implementing the method according to any one of claims 12 to 13.