A speech filtering method, device, storage medium and equipment

By using the encoding without post-filtering module and the decoding terminal filtering of the pre-trained neural network model in Bluetooth devices, the problem of high complexity of LTPF operations in LC3 is solved, and the application of high-definition audio on low-power devices is realized, extending the service life of the device and maintaining the sound quality.

CN115497488BActive Publication Date: 2025-08-26BEIJING BAIRUI INTERNET TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211199937.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-08-26
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

In the prior art, the long-term post-filter module (LTPF) of LC3 has high computational complexity, resulting in excessive computing resources occupied by Bluetooth devices, limiting the application of high-definition audio on low-power devices.

Method used

The standard Bluetooth encoder without post-filtering module is used for encoding, and the pre-trained neural network model is used to filter at the decoding end. After obtaining the target spectrum coefficient, the target voice signal is obtained through the low-latency improved discrete cosine inverse transformation module, and the complex filtering steps at the encoding end are omitted.

Benefits of technology

It reduces the complexity of the codec, reduces the system's computing volume, improves the computing efficiency, extends the service life of the equipment, and maintains the effect of sound quality close to standard decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497488B_ABST
    Figure CN115497488B_ABST
Patent Text Reader

Abstract

The present application discloses a speech filtering method, device, storage medium and equipment, which belongs to the field of speech coding and decoding technology. The method mainly includes: encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain speech spectrum coefficients; inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and according to the remaining decoding steps of the standard decoder without a post-filtering module, inputting the target spectrum coefficients into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain the target speech signal corresponding to the target spectrum coefficients. The present application omits the complex post-filtering operation in the Bluetooth encoding process, and only uses the pre-trained neural network model for filtering in the Bluetooth decoding process, so that it achieves a sound quality close to that of standard decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech coding and decoding technology, and in particular to a speech filtering method, apparatus, storage medium and device. Background Art

[0002] In the existing technology, in order to enhance the sound quality of voice data, LC3 introduces a long-term post-filter (LTPF) module based on time domain signal processing, which sharpens the harmonic structure of the signal by attenuating the quantization noise in the spectral valley. The specific operation steps are as follows: at the encoding end: determine whether the LTPF needs to be activated, and extract the relevant fundamental frequency parameters at the same time. The encoding end mainly includes resampling, high-pass filtering, downsampling, fundamental frequency detection, fundamental frequency delay estimation and activation judgment; at the decoding end, according to the parameters extracted by the encoding end, when the LTPF is activated, an IIR filter is used to implement filtering.

[0003] However, in the above-mentioned filtering steps, the computational complexity of resampling, pitch detection (based on autocorrelation), and pitch delay (based on autocorrelation) at the encoding end is very large, making the LTPF (long-term post-filter module) one of the modules with the highest computational complexity in LC3, affecting its application in low-power Bluetooth devices; at the decoding end, filtering is achieved only by using IIR filters when the LTPF (long-term post-filter module) is activated based on the parameters extracted from the encoding end.

[0004] For example, as user experience requirements become increasingly demanding, TWS Bluetooth headsets and Bluetooth microphones used by anchors tend to use high-definition audio mode to capture voice to enhance the user experience. Taking a 48kHz sampling rate configuration as an example, the computing power required by the LC3 encoder is proportional to the sampling rate. The LTPF (long-term post-filter module) not only takes up more computing power but also requires more memory resources. However, the size limitations of low-power devices such as TWS Bluetooth headsets result in extremely limited battery capacity and small memory capacity. This contradiction restricts the application of high-definition audio applications in low-power devices. Summary of the Invention

[0005] In response to the problem of high complexity of LTPF (long-term post-filter module) in the existing technology, this application mainly provides a speech filtering method, device, storage medium and equipment.

[0006] In order to achieve the above-mentioned purpose, a technical solution adopted in the present application is: providing a speech filtering method, which includes: encoding a speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain speech spectrum coefficients corresponding to the speech signal; inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and according to the remaining decoding steps of the standard decoder without a post-filtering module, inputting the target spectrum coefficients into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain a target speech signal corresponding to the target spectrum coefficients.

[0007] Another technical solution adopted in the present application is: providing a speech filtering device, which includes: a module for encoding a speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain a speech spectrum coefficient corresponding to the speech signal; a module for inputting the speech spectrum coefficient into a pre-trained neural network model to obtain a target spectrum coefficient corresponding to the speech spectrum coefficient; and a module for inputting the target spectrum coefficient into a low-latency improved inverse discrete cosine transform module of a standard decoder according to the remaining decoding steps of the standard decoder without a post-filtering module to obtain a target speech signal corresponding to the target spectrum coefficient.

[0008] Another technical solution adopted in the present application is: providing a computer-readable storage medium storing computer instructions, which are operated to execute the speech filtering method in solution one.

[0009] Another technical solution adopted in this application is: providing a computer device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores computer instructions that can be executed by the at least one processor, and the at least one processor operates the computer instructions to execute the speech filtering method in solution one.

[0010] The beneficial effects that can be achieved by the technical solution of the present application are: omitting the complex post-filtering operation in the Bluetooth encoding process, and only using the pre-trained neural network model to filter the speech spectrum coefficients in the Bluetooth decoding process, so that it can achieve sound quality close to that of standard decoding, reducing the complexity of the codec, while reducing the system calculation amount and improving calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0012] Figure 1 This is a schematic diagram of an optional implementation of a speech filtering method of the present application;

[0013] Figure 2 It is a schematic diagram of an optional example of standard encoding and decoding steps in the prior art;

[0014] Figure 3 1 is a schematic diagram of an optional example of a corresponding relationship between parameter configurations of a transmitting end and a receiving end in a speech filtering method of the present application;

[0015] Figure 4 It is a schematic diagram of an example of the encoding and decoding steps in this application;

[0016] Figure 5 This is a schematic diagram of an optional implementation of a speech filtering device of the present application.

[0017] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0018] The preferred embodiments of the present application are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present application can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present application.

[0019] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0020] In the existing technology, in order to enhance the sound quality of voice data, LC3 introduces a long-term post-filter (LTPF) module based on time domain signal processing, which sharpens the harmonic structure of the signal by attenuating the quantization noise in the spectral valley. The specific operation steps are as follows: at the encoding end: determine whether the LTPF needs to be activated, and extract the relevant fundamental frequency parameters at the same time. The encoding end mainly includes resampling, high-pass filtering, downsampling, fundamental frequency detection, fundamental frequency delay estimation and activation judgment; at the decoding end, according to the parameters extracted by the encoding end, when the LTPF is activated, an IIR filter is used to implement filtering.

[0021] However, in the above-mentioned filtering steps, the computational complexity of resampling, pitch detection (based on autocorrelation), and pitch delay (based on autocorrelation) at the encoding end is very large, making the LTPF (long-term post-filter module) one of the modules with the highest computational complexity in LC3, affecting its application in low-power Bluetooth devices; at the decoding end, filtering is achieved only by using IIR filters when the LTPF (long-term post-filter module) is activated based on the parameters extracted from the encoding end.

[0022] For example, as user experience requirements become increasingly demanding, TWS Bluetooth headsets and Bluetooth microphones used by anchors tend to use high-definition audio mode to capture voice to enhance the user experience. Taking a 48kHz sampling rate configuration as an example, the computing power required by the LC3 encoder is proportional to the sampling rate. The LTPF (long-term post-filter module) not only takes up more computing power but also requires more memory resources. However, the size limitations of low-power devices such as TWS Bluetooth headsets result in extremely limited battery capacity and small memory capacity. This contradiction restricts the application of high-definition audio applications in low-power devices.

[0023] In response to the problem of high complexity of the LTPF (long-term post-filter module) in the prior art, the present application mainly provides a speech filtering method, apparatus, storage medium and device. The speech filtering method includes: encoding a speech signal according to a standard Bluetooth encoder without a post-filter module, decoding the encoded speech signal into a transform domain noise shaping decoding module according to a standard decoder without a post-filter module, and obtaining speech spectrum coefficients corresponding to the speech signal; inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and according to the remaining decoding steps of the standard decoder without a post-filter module, inputting the target spectrum coefficients into a low-latency improved inverse discrete cosine transform module of the standard decoder to obtain a target speech signal corresponding to the target spectrum coefficients.

[0024] The complex post-filtering module is omitted during the Bluetooth encoding process, and only the remaining decoding steps except the post-filtering module are performed on the voice signal. During the Bluetooth decoding process, the encoded voice signal is decoded into the transform domain noise shaping decoding module to obtain the voice spectrum coefficients corresponding to the voice signal. A pre-trained neural network model is used to replace the post-filtering module in the standard Bluetooth decoder to filter the voice spectrum coefficients to obtain the target spectrum coefficients. The target spectrum coefficients are then input into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain the target voice signal corresponding to the target spectrum coefficients. This ensures that the target voice signal achieves sound quality close to that of standard decoding, reduces the complexity of the codec, and at the same time reduces the system computation load, improves budget efficiency, and extends the life of the codec.

[0025] The following describes in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems using specific embodiments. The specific embodiments described below can be combined with each other to form new embodiments. The same or similar ideas or processes described in one embodiment may not be repeated in other embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0026] Figure 1 An optional implementation of a speech filtering method of the present application is shown.

[0027] exist Figure 1 In the optional implementation shown, the speech filtering method mainly includes step S101, encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain the speech spectrum coefficients corresponding to the speech signal.

[0028] In this optional embodiment, since the standard Bluetooth encoding steps in the prior art are as follows Figure 2 As shown in the left figure, a long-term post-filter is used to filter the input voice signal to ensure the sound quality of the voice signal; however, the long-term post-filter used in the above filtering method not only takes up more computing power, but also requires more memory resources, which is not conducive to use on small and delicate devices, and will shorten the service life of the encoder. Therefore, the present application provides a voice filtering method, which first omits the filtering step in the Bluetooth encoding process, that is, deletes Figure 2 The long-term post-filter module in the left figure uses Figure 2 The other modules in the left figure encode the voice signal and obtain the code stream corresponding to the voice signal. In the decoding process, according to the standard Bluetooth decoder, Figure 2The decoding step in the right figure partially decodes the code stream, and decodes it to the transform domain noise shaping decoding module to output the speech spectrum coefficients corresponding to the code stream; so that the subsequent filtering steps are all performed in the frequency domain state, providing the necessary basis for the subsequent filtering steps.

[0029] exist Figure 1 In the optional implementation shown, the speech filtering method further includes step S102, inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients.

[0030] In this optional embodiment, since the speech spectrum coefficients obtained above are not filtered during the encoding and decoding processes, the speech spectrum coefficients at this time are noisy speech spectrum coefficients; the pre-trained neural network model is used as the filter of this application to filter the speech spectrum coefficients to obtain the target spectrum coefficients for denoising, thereby achieving the filtering purpose; the speech filtering method provided by this application only performs one filtering step at the decoding end to achieve sound quality similar to that of the speech signal obtained after performing the filtering step in the standard encoding and decoding process.

[0031] It should be noted that the neural network models used in this application include but are not limited to autoencoders, CNN, RNN, CRNN, and LSTM. This application does not limit the type of neural network model, as long as it can achieve the filtering effect of this application.

[0032] In an optional embodiment of the present application, the pre-training process of the neural network model includes: encoding and decoding the training speech signal to a transform domain noise shaping decoding module according to a standard codec to obtain clean speech spectrum coefficients corresponding to the training speech signal; encoding the training speech signal according to a standard encoder without a post-filtering module, and decoding the encoded training speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain noisy speech spectrum coefficients corresponding to the training speech signal; performing feature extraction on the clean speech spectrum coefficients and the noisy speech spectrum coefficients respectively to obtain a clean amplitude spectrum corresponding to the clean speech spectrum coefficients and a noisy amplitude spectrum corresponding to the noisy speech spectrum coefficients; inputting the absolute value of the noisy speech spectrum coefficients into a preset neural network model to obtain the gain of the noisy speech spectrum coefficients; and adjusting the relevant parameters of the preset neural network model accordingly based on the relationship between the gain and the noisy amplitude spectrum and the clean amplitude spectrum to obtain a pre-trained neural network model.

[0033] In this optional embodiment, first the encoding and decoding steps according to the standard codec, i.e. Figure 2In the encoding step of the left figure and the decoding step of the right figure, the training speech signal is filtered according to the long-term post-filtering module in the standard codec to ensure the sound quality of the training speech signal; and the clean speech spectrum coefficient corresponding to the training speech signal after post-filtering is used as the control group. At the same time, the training speech signal is encoded and decoded according to the standard codec without a post-filtering module to obtain the unfiltered noisy speech spectrum coefficient, and used as the experimental group. Feature extraction is performed on the clean speech spectrum coefficient of the control group and the noisy speech spectrum coefficient of the experimental group respectively, and the pure amplitude spectrum corresponding to the pure speech spectrum coefficient and the noisy amplitude spectrum corresponding to the noisy speech spectrum coefficient are obtained; the absolute value of the noisy speech spectrum coefficient of the experimental group is input into the preset neural network model, and the neural network model is obtained. The gain of the noisy speech spectrum coefficient is calculated according to its own current relevant parameters, and the relevant parameters in the neural network model are adjusted accordingly according to the relationship between the gain and the noisy amplitude spectrum and the pure amplitude spectrum to obtain a pre-trained neural network model. Provide the necessary basis for obtaining the target spectrum coefficient in this application.

[0034] In an optional embodiment of the present application, feature extraction is performed on the clean speech spectrum coefficients and the noisy speech spectrum coefficients respectively to obtain the clean amplitude spectrum corresponding to the clean speech spectrum coefficients and the noisy amplitude spectrum corresponding to the noisy speech spectrum coefficients, further comprising: performing a discrete sine transform on the clean speech spectrum coefficients to obtain the clean sine amplitude spectrum corresponding to the clean speech spectrum coefficients; and taking the sum of the clean speech spectrum coefficients and the clean sine amplitude spectrum as the clean amplitude spectrum.

[0035] In this optional embodiment, the feature extraction method for the pure speech spectrum coefficient is: first, the pure speech spectrum coefficient is discrete sine transformed to obtain the pure sine amplitude spectrum of the pure speech spectrum coefficient, and the sum of the pure sine amplitude spectrum and the pure speech spectrum coefficient is the aforementioned pure amplitude spectrum.

[0036] In an optional embodiment of the present application, the calculation formula for calculating the pure amplitude spectrum is as follows:

[0037]

[0038] Among them, This is the aforementioned pure speech spectrum coefficient; That is, the pure sine amplitude spectrum obtained by performing improved discrete sine transform on the pure speech spectrum coefficients; That is the sum of the pure sinusoidal amplitude spectrum and the pure speech spectrum coefficient, the pure amplitude spectrum.

[0039] In an optional embodiment of the present application, feature extraction is performed on the clean speech spectrum coefficients and the noisy speech spectrum coefficients respectively to obtain the clean amplitude spectrum corresponding to the clean speech spectrum coefficients and the noisy amplitude spectrum corresponding to the noisy speech spectrum coefficients, and the method further includes: performing a discrete sine transform on the noisy speech spectrum coefficients to obtain the noisy sine amplitude spectrum corresponding to the noisy speech spectrum coefficients; and taking the sum of the noisy speech spectrum coefficients and the noisy sine amplitude spectrum as the noisy amplitude spectrum.

[0040] In this optional embodiment, the method for extracting features of the noisy speech spectrum coefficients is: first, a discrete sine transform is performed on the noisy speech spectrum coefficients to obtain a pure sine amplitude spectrum of the noisy speech spectrum coefficients, and the sum of the noisy sine amplitude spectrum and the noisy speech spectrum coefficients is the aforementioned noisy amplitude spectrum.

[0041] In an optional embodiment of the present application, the calculation formula for calculating the noisy amplitude spectrum is as follows:

[0042]

[0043] Among them, That is, the sum of the noisy sine amplitude spectrum and the noisy speech spectrum coefficient, the noisy amplitude spectrum; This is the aforementioned noisy speech spectrum coefficient; That is, the noisy sine amplitude spectrum obtained by performing an improved discrete sine transform on the noisy speech spectrum coefficients.

[0044] In an optional embodiment of the present application, in the process of calculating the clean amplitude spectrum and the noisy amplitude spectrum, the following formulas are used to respectively calculate and obtain the clean speech spectrum coefficients and the noisy speech spectrum coefficients:

[0045]

[0046] Where X in the above formula mdct (k) is the speech spectrum coefficient; in the process of calculating the clean speech spectrum coefficient and the noisy speech spectrum coefficient, the relevant data corresponding to the clean speech spectrum coefficient and the noisy speech spectrum coefficient are respectively substituted into the formula to obtain the above-mentioned clean speech spectrum coefficient and the noisy speech spectrum coefficient.

[0047] The pure sine amplitude spectrum and the noisy sine amplitude spectrum are calculated according to the following formulas:

[0048]

[0049] t(n)=x(ZN F +n), for n=0...2·N F -1-Z

[0050] t(2N F-Z+n)=0,forn=0...Z-1

[0051] Where X in the above formula mdst (k) is the sine amplitude spectrum. In the process of calculating the pure sine amplitude spectrum and the noisy sine amplitude spectrum, the relevant data corresponding to the pure sine amplitude spectrum and the noisy sine amplitude spectrum are respectively substituted into the formula to obtain the above-mentioned pure sine amplitude spectrum and the noisy sine amplitude spectrum.

[0052] In an optional embodiment of the present application, the relevant parameters of the preset neural network model are adjusted accordingly according to the gain and the noisy amplitude spectrum to obtain a pre-trained neural network model, further including: calculating and obtaining a first updated amplitude spectrum corresponding to the noisy speech spectrum coefficient according to the product of the gain and the noisy amplitude spectrum; calculating a first error between the first updated amplitude spectrum and the pure amplitude spectrum; when the first error is greater than a preset error threshold, the relevant parameters of the preset neural network model are adjusted accordingly according to the first error to obtain a pre-trained neural network model.

[0053] In this optional embodiment, the gain and the noisy amplitude spectrum are used as the first updated amplitude spectrum; a first error between the clean amplitude spectrum and the first updated amplitude spectrum is calculated. Since the clean amplitude spectrum is the amplitude spectrum corresponding to the clean speech spectrum coefficients of the control group of this application, and the noisy amplitude spectrum is the amplitude spectrum corresponding to the unfiltered noisy speech spectrum coefficients of this application, the goal of this solution is to ensure that after the absolute value of the noisy speech spectrum coefficients is input into the neural network model, the error between the output gain and the noisy amplitude spectrum, that is, the error between the obtained updated amplitude spectrum and the clean amplitude spectrum, is less than a preset error threshold. This ensures the sound quality of the speech signal. Therefore, the first updated amplitude spectrum is first calculated and the first error between the first updated amplitude spectrum and the clean amplitude spectrum is used. When the first error is less than or equal to the preset error threshold, it indicates that the neural network model at this time can meet the filtering effect required by this solution, ensuring sound quality, and the neural network model at this time is used as the pre-trained neural network model. When the first error is greater than the preset error threshold, it indicates that the neural network model at this time cannot meet the filtering effect required by this solution, and therefore the bias and weights in the neural network model need to be adjusted to achieve the required filtering effect.

[0054] In an optional embodiment of the present application, adjusting relevant parameters of a preset neural network model according to a first error to obtain a pre-trained neural network model further includes: adjusting relevant parameters according to an Nth error to obtain an N+1th updated neural network model, where N is a natural number not equal to 0; inputting the absolute value of the noisy speech spectrum coefficient into the N+1th updated neural network model to obtain an N+1th updated amplitude spectrum corresponding to the noisy speech spectrum coefficient; calculating the N+1th error between the N+1th updated amplitude spectrum and the clean amplitude spectrum; when the N+1th error is greater than a preset error threshold, adjusting relevant parameters of the N+1th updated neural network model to obtain a pre-trained neural network model. When the N+1th error is less than or equal to the preset error threshold, using the N+1th updated neural network model as the pre-trained neural network model.

[0055] In this optional embodiment, after updating the neural network model, it is determined whether the N+1th updated neural network model can achieve the filtering effect, that is, the gain output by the updated N+1th neural network model is multiplied by the noisy amplitude spectrum, the product is compared with the pure amplitude spectrum, and the N+1th error between the two is calculated. When the N+1th error is less than the preset error threshold, it means that the N+1th neural network model can now meet the filtering effect required by this scheme, and the N+1th updated neural network model can be used as a pre-trained neural network model; when the N+1th error is greater than or equal to the preset error threshold, it means that the N+1th neural network model cannot meet the filtering effect required by this scheme, and the relevant parameters of the N+1th updated neural network model are adjusted until the gain output by the updated neural network model can achieve the filtering effect required by this scheme after calculation.

[0056] In an optional example of the present application, taking a 48kHz sampling rate and a 10ms frame length as an example, the autoencoder is used as the neural network model of the present application; at this time, the configuration of the neural network model can be: the input layer size is 5x400, where 5 represents the current frame and the 4 frames before the current frame, the first convolution layer input is 1x5x400, and the output is 40x5x199; the second convolution layer input is 40x5x199, and the output is 80x5x99; the third convolution layer input is 80x5x99, and the output is 160x5x49; the fourth deconvolution layer input is 160x5x49, and the output is 80x5x99; the fifth deconvolution layer input is 80x5x99, and the output is 40x5x199; the sixth deconvolution layer input is 40x5x199, and the output is 1x5x399. The output layer is a fully connected layer with a size of 400, which corresponds to the gain of the spectral coefficients of one frame. This gain is applied to the spectral coefficients of the current frame to obtain the new spectral coefficients. Then, 80 zeros are added to the 400 spectral coefficients, and then IMDCT and overlap-add are performed to output the time domain audio signal.

[0057] In addition, there is a skip connection between the output of the first convolutional layer and the sixth deconvolutional layer, and a skip connection between the output of the second convolutional layer and the input of the fifth deconvolutional layer.

[0058] Among them, the forward propagation function is as follows:

[0059]

[0060] X in the above formula noise,mdct That is, the spectral coefficients of the decoded output without the LTPF decoder, Gain(j) is the output spectral coefficient gain, and f() is the activation function; among them, the Softplus function can be used as the activation function of this application, and its expression is as follows:

[0061] f(x)=log(1+exp(x))

[0062] During the training process, the weights and bias of the hidden layer of the neural network can be updated based on backpropagation. The specific formula is as follows:

[0063]

[0064]

[0065] In the above formula, μ is the learning rate, which affects the speed of convergence, and E is the loss function. The difference between the new amplitude spectrum and the reference amplitude spectrum is calculated as follows:

[0066]

[0067] Wherein k in the above formula is the number of output spectrum coefficients. When configured with a 48kHz sampling rate and a 10ms frame length, k=400.

[0068] The significance of the above neural network in the training phase is to convert the spectrum coefficients X output by the LTPF decoder into noise,mdct The neural network is input, and after nonlinear processing by the neural network, a gain is output. Using a large number of training samples, the weights and offsets are adjusted to minimize the mean square error between the new MDFT magnitude spectrum with this gain applied and the pure MDFT magnitude spectrum (i.e., the reference magnitude spectrum). During the inference phase, the spectral coefficients decoded by the non-LTPF decoder are input, and the gain is output. The gain is applied to the spectral coefficients to obtain new spectral coefficients. IMDCT and overlap-add are then performed to output the time-domain audio signal.

[0069] In an optional embodiment of the present application, the relevant parameters of the preset neural network model are adjusted accordingly according to the error to obtain a pre-trained neural network model, and the method also includes: recording the number of training times M of the preset neural network model; if the number of training times M is less than or equal to the preset training times threshold, then continuing to train the N+1th updated neural network model; if the number of training times M is greater than the training times threshold, then determining the N+1th updated neural network model as the pre-trained neural network model.

[0070] In this optional embodiment, when the first updated neural network model is obtained, the number of training times of the neural network model is recorded as 1, and so on, when the N+1th updated neural network model is obtained, the number of training times of the neural network model is recorded as M; when the N+1th error is greater than the preset error threshold, the number of training times M is compared with the training times threshold. When the number of training times M is greater than or equal to the training times threshold, the N+1th updated neural network model is no longer trained, and the N+1th updated neural network model is determined as the pre-trained neural network model; when the number of training times M is less than the training times threshold, the N+1th updated neural network model is trained, that is, the relevant parameters of the N+1th updated neural network model are adjusted to obtain the N+2th updated neural network model, to provide a basis for the next cycle and training.

[0071] exist Figure 1 In the optional implementation shown, the speech filtering method also includes step S103, according to the remaining decoding steps of the standard codec without a post-filtering module, inputting the target spectral coefficient into the low-delay improved inverse discrete cosine transform module of the standard codec to obtain the target speech signal corresponding to the target spectral coefficient.

[0072] In this optional implementation, the remaining decoding steps are performed on the target spectrum coefficients determined above, that is, the target spectrum coefficients are input into a low-delay improved inverse discrete cosine transform module to obtain the target speech signal corresponding to the target spectrum coefficients.

[0073] In an optional example of the present application, parameter negotiation and configuration are performed on both the Bluetooth transmitter and the Bluetooth receiver, that is, when the application is started, the Bluetooth transmitter and the Bluetooth receiver perform the step of negotiating parameters, that is, judging whether the Bluetooth transmitter and the Bluetooth receiver can support filtering only at the decoding end based on the parameters of the Bluetooth transmitter and the Bluetooth receiver; when the parameters of the Bluetooth transmitter and the Bluetooth receiver both meet the preset decoding end filtering standards, it indicates that the Bluetooth transmitter and the Bluetooth receiver support filtering only at the decoding end.

[0074] Figure 3 An optional example of the corresponding relationship between the parameter configurations of the transmitting end and the receiving end in a speech filtering method of the present application is shown.

[0075] according to Figure 3 In the example shown, when a voice call starts, the parameters are first negotiated between the Bluetooth transmitter and the Bluetooth receiver, that is, the audio format, sampling rate, and bit rate range are compared with the preset decoding end filtering standards to determine whether the above parameters meet the preset decoding end filtering standards, so as to know whether the Bluetooth transmitter and the Bluetooth receiver support filtering only at the decoding end; if the Bluetooth transmitter and the Bluetooth receiver both support filtering only at the decoding end, the Bluetooth transmitter selects the LTPF-free encoding mode and completely skips the LTPF-related operations during the encoding process; the Bluetooth receiver selects the decoding based on the autoencoder LTPF; otherwise, the standard mode of encoding and decoding is selected.

[0076] The complex post-filtering module is omitted during the Bluetooth encoding process, and only the remaining decoding steps except the post-filtering module are performed on the voice signal. During the Bluetooth decoding process, the encoded voice signal is decoded into the transform domain noise shaping decoding module to obtain the voice spectrum coefficients corresponding to the voice signal. A pre-trained neural network model is used to replace the post-filtering module in the standard Bluetooth decoder to filter the voice spectrum coefficients to obtain the target spectrum coefficients. The target spectrum coefficients are then input into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain the target voice signal corresponding to the target spectrum coefficients. This ensures that the target voice signal achieves sound quality close to that of standard decoding, reduces the complexity of the codec, and at the same time reduces the system computation load, improves budget efficiency, and extends the life of the codec.

[0077] In an optional embodiment of the present application, when global filtering is enabled only at the decoding end, a bit indication is added to the output code stream of each frame, located after the code stream of time domain noise shaping, 1: indicates that the current frame is enabled, 0: indicates that the current frame is not enabled; the above bit can be written to the end of the side information (auxiliary information); where the side information (auxiliary information) is part of the Bluetooth encoding output code stream, mainly used to store some frame-level information, such as bandwidth, global gain and TNS activation flag.

[0078] Figure 4 A schematic diagram showing an example of the encoding and decoding steps of the present application.

[0079] exist Figure 4 In the example shown, Figure 2 Compared with the standard encoding and decoding steps shown in , the long-term post-filter processing is omitted during the encoding of the audio data; during the decoding process, as shown in Figure 4 As shown in the figure, the post-filter processing part of the pre-trained neural network model is newly added to process the speech spectrum coefficients to obtain the corresponding target spectrum coefficients, and then the target spectrum coefficients are input into the low-latency improved inverse discrete cosine transform to obtain the final decoding result. Figure 2 Compared with the standard encoding and decoding steps in the decoding step, the long-term post-filter decoding step is omitted.

[0080] The speech filtering method provided by this solution can achieve a sound quality similar to that of the existing technology by filtering at the encoding and decoding ends only by filtering at the decoding end, omitting the complex post-filtering operation in the LC3 encoding process, and can extend the usage time of Bluetooth devices with limited power consumption; this application provides two filtering methods: for standard LC3 code streams, either standard LTPF (relevant modules need to be enabled) or a new frequency domain post-filtering module can be used, or both can be used; for code streams without LTPF, a new frequency domain post-filtering module can be used to achieve sound quality close to that of standard codecs.

[0081] In the standard LC3 encoder, the LTPF function is usually disabled if the bit rate is high. However, if the post-filtering of the present invention is applied at the decoding end, the sound quality can still be enhanced to a certain extent. In the standard LC3 encoder, the calculation of LTPF-related parameters is relatively strict. In certain critical situations, such as when the pitch is detected but it is in the initial stage, the encoder is likely to output pitch_present = 0, so the decoding end will not use LTPF to enhance the sound quality. However, the post-filtering of the present invention can still enhance the sound quality. This improves the flexibility of the codec in filtering methods.

[0082] Figure 5 An optional implementation of a speech filtering device of the present application is shown.

[0083] exist Figure 5 In the optional embodiment shown, the speech filtering device mainly includes: a module 501 for encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard codec without a post-filtering module to obtain speech spectrum coefficients corresponding to the speech signal; a module 502 for inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and a module 503 for inputting the target spectrum coefficients into a low-latency improved inverse discrete cosine transform module of the standard decoder according to the remaining decoding steps of the standard decoder without a post-filtering module to obtain a target speech signal corresponding to the target spectrum coefficients.

[0084] In an optional embodiment of the present application, each functional module in a speech filtering device of the present application may be directly in hardware, in a software module executed by a processor, or in a combination of the two.

[0085] The software modules may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from and write information to the storage medium.

[0086] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In the alternative, the storage medium may be integral to the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and storage medium may reside as discrete components in the user terminal.

[0087] The speech filtering device provided in this application can be used to execute the speech filtering method described in any of the above embodiments. Its implementation principles and technical effects are similar and will not be repeated here.

[0088] In another optional embodiment of the present application, a computer-readable storage medium stores computer instructions, and the computer instructions are operated to execute the speech filtering method described in the above embodiment.

[0089] In an optional embodiment of the present application, a computer device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores computer instructions that can be executed by the at least one processor, and the at least one processor operates the computer instructions to execute the speech filtering method described in the above embodiment.

[0090] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0091] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0092] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structural transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A speech filtering method, characterized in that: include: Encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without the post-filtering module to obtain speech spectral coefficients corresponding to the speech signal; Inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; as well as According to the remaining decoding steps of the standard decoder without the post-filtering module, the target spectral coefficient is input into the low-delay improved inverse discrete cosine transform module of the standard decoder without the post-filtering module to obtain a target speech signal corresponding to the target spectral coefficient; The pre-training process of the neural network model includes: Encode and decode the training speech signal to a transform domain noise shaping decoding module according to a standard codec to obtain clean speech spectrum coefficients corresponding to the training speech signal; Encoding the training speech signal according to the standard encoder without the post-filtering module, and decoding the encoded training speech signal to the transform domain noise shaping decoding module according to the standard decoder without the post-filtering module to obtain noisy speech spectrum coefficients corresponding to the training speech signal; Performing feature extraction on the clean speech spectrum coefficients and the noisy speech spectrum coefficients respectively to obtain a clean amplitude spectrum corresponding to the clean speech spectrum coefficients and a noisy amplitude spectrum corresponding to the noisy speech spectrum coefficients; Inputting the absolute value of the noisy speech spectrum coefficient into a preset neural network model to obtain the gain of the noisy speech spectrum coefficient; and According to the relationship between the gain and the noisy amplitude spectrum and the clean amplitude spectrum, relevant parameters of the preset neural network model are adjusted accordingly to obtain the pre-trained neural network model.

2. The speech filtering method according to claim 1, wherein: The feature extraction of the clean speech spectrum coefficient and the noisy speech spectrum coefficient respectively to obtain the clean amplitude spectrum corresponding to the clean speech spectrum coefficient and the noisy amplitude spectrum corresponding to the noisy speech spectrum coefficient further includes: Performing discrete sine transform on the clean speech spectrum coefficients to obtain a clean sine amplitude spectrum corresponding to the clean speech spectrum coefficients; The sum of the clean speech spectrum coefficient and the clean sinusoidal amplitude spectrum is used as the clean amplitude spectrum.

3. The speech filtering method according to claim 1, wherein: The feature extraction of the clean speech spectrum coefficients and the noisy speech spectrum coefficients is performed respectively to obtain the clean amplitude spectrum corresponding to the clean speech spectrum coefficients and the noisy amplitude spectrum corresponding to the noisy speech spectrum coefficients, further comprising: Performing discrete sine transform on the noisy speech spectrum coefficients to obtain a noisy sine amplitude spectrum corresponding to the noisy speech spectrum coefficients; The sum of the noisy speech spectrum coefficient and the noisy sinusoidal amplitude spectrum is used as the noisy amplitude spectrum.

4. The speech filtering method according to claim 1, wherein: The step of adjusting relevant parameters of the preset neural network model according to the relationship between the gain and the noisy amplitude spectrum and the clean amplitude spectrum to obtain the pre-trained neural network model further includes: Calculating and obtaining a first updated amplitude spectrum corresponding to the noisy speech spectrum coefficient according to the product of the gain and the noisy amplitude spectrum; calculating a first error between the first updated amplitude spectrum and the pure amplitude spectrum; When the first error is greater than a preset error threshold, relevant parameters of the preset neural network model are adjusted accordingly according to the first error to obtain the pre-trained neural network model.

5. The speech filtering method according to claim 4, characterized in that: The step of adjusting relevant parameters of the preset neural network model according to the first error to obtain the pre-trained neural network model further includes: Adjust the relevant parameters according to the Nth error to obtain the N+1th updated neural network model, where N is a natural number not equal to 0; Inputting the absolute value of the noisy speech spectrum coefficient into the N+1th updated neural network model to obtain the N+1th updated amplitude spectrum corresponding to the noisy speech spectrum coefficient; Calculating an N+1th error between the N+1th updated amplitude spectrum and the pure amplitude spectrum; When the N+1th error is greater than the preset error threshold, adjusting the relevant parameters of the N+1th updated neural network model accordingly to obtain the pre-trained neural network model; When the N+1th error is less than or equal to the preset error threshold, the N+1th updated neural network model is used as the pre-trained neural network model.

6. The speech filtering method according to claim 5, characterized in that: The step of adjusting relevant parameters of the preset neural network model according to the error to obtain the pre-trained neural network model further includes: Recording the number of training times M for the preset neural network model; If the training number M is less than or equal to the preset training number threshold, then continue training the N+1th updated neural network model; If the training number M is greater than the training number threshold, the N+1th updated neural network model is determined as the pre-trained neural network model.

7. A speech filtering device, characterized in that: include: A module for encoding a speech signal according to a standard Bluetooth encoder without a post-filtering module, decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without the post-filtering module, and obtaining a speech spectral coefficient module corresponding to the speech signal; A module for inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; as well as a module configured to input the target spectral coefficients into a low-delay inverse discrete cosine transform module of the standard decoder without the post-filtering module according to the remaining decoding steps of the standard decoder without the post-filtering module to obtain a target speech signal corresponding to the target spectral coefficients; The pre-training process of the neural network model includes: Encode and decode the training speech signal to a transform domain noise shaping decoding module according to a standard codec to obtain clean speech spectrum coefficients corresponding to the training speech signal; Encoding the training speech signal according to the standard encoder without the post-filtering module, and decoding the encoded training speech signal to the transform domain noise shaping decoding module according to the standard decoder without the post-filtering module to obtain noisy speech spectrum coefficients corresponding to the training speech signal; Performing feature extraction on the clean speech spectrum coefficients and the noisy speech spectrum coefficients respectively to obtain a clean amplitude spectrum corresponding to the clean speech spectrum coefficients and a noisy amplitude spectrum corresponding to the noisy speech spectrum coefficients; Inputting the absolute value of the noisy speech spectrum coefficient into a preset neural network model to obtain the gain of the noisy speech spectrum coefficient; and According to the relationship between the gain and the noisy amplitude spectrum and the clean amplitude spectrum, relevant parameters of the preset neural network model are adjusted accordingly to obtain the pre-trained neural network model.

8. A computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are operated to perform the speech filtering method according to any one of claims 1 to 6.

9. A computer device, characterized in that: include: at least one processor; as well as a memory in communicative connection with the at least one processor; The memory stores computer instructions that can be executed by the at least one processor, and the at least one processor operates the computer instructions to execute the speech filtering method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Keyword recognition method and system based on deep learning, medium and equipment

    CN113823277A

  • Model training method for tone quality conversion, and method and device for improving voice quality

    CN114863942A

  • Voice noise reduction model training method, voice noise reduction method, device and medium

    CN115083429A

  • Hearing device comprising a recurrent neural network and a method of processing an audio signal

    US20220232331A1