Voice wake-up processing method and apparatus, storage medium, and electronic device

CN117690448BActive Publication Date: 2026-09-11SHENZHEN TCL NEW-TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311719076.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2026-09-11
Estimated Expiration
2043-12-13

AI Technical Summary

Technical Problem

[0003]但是,相关技术中,通常非线性后处理和唤醒检测处理两个步骤通常时相互独立的,容易把残余回声中接近唤醒语音信号的非唤醒语音进行增强,使得唤醒算法把它误认为是唤醒语音信号,继而产生误唤醒,虽然唤醒率得以提升,但是会使得误唤醒率较高

Benefits of technology

[0018] In this way, the nonlinear post-processing network and the wake-up network obtained by joint adversarial training are used for wake-up processing, which effectively combines nonlinear post-processing and wake-up detection processing. Furthermore, in adversarial training, the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator. In contrast to the conventional adversarial training objective, the objective of adversarial training in this application is to make the discriminator more and more accurate in distinguishing between real and fake data generated by the generator, so that the data generated by the generator becomes easier for the discriminator to distinguish between real and fake data. This allows the voice wake-up processing result to enhance the wake-up rate while reducing the false wake-up rate caused by the wake-up processing process itself.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117690448B_ABST
    Figure CN117690448B_ABST
Patent Text Reader

Abstract

The application discloses a voice wake-up processing method and device, a storage medium and electronic equipment, and relates to the technical field of audio processing. The method comprises the following steps: obtaining a to-be-processed voice signal; performing adaptive filtering echo cancellation processing on the to-be-processed voice signal to obtain an error signal; performing echo cancellation nonlinear post-processing on the basis of a post-processing input signal and the to-be-processed voice signal by using a nonlinear post-processing network to obtain an echo cancellation signal; and performing wake-up detection processing on the basis of the echo cancellation signal by using a wake-up network. The nonlinear post-processing network and the wake-up network are obtained through joint adversarial training. In the adversarial training, the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator. The goal of the adversarial training is to make the accuracy of the discriminator in identifying the authenticity of the data generated by the generator higher and higher. The application can enhance the wake-up rate while reducing the false wake-up rate caused by the wake-up processing procedure itself.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, specifically to a voice wake-up processing method, apparatus, storage medium, and electronic device. Background Technology

[0002] Voice wake-up processing is the process of detecting and analyzing whether a specific wake-up word is contained in a speech signal. In related technologies, voice wake-up processing usually involves first performing adaptive filtering and echo cancellation on the speech signal to obtain an error signal, and then performing nonlinear post-processing on the error signal to further eliminate residual echoes. The resulting signal is finally processed by a wake-up algorithm for wake-up detection.

[0003] However, in related technologies, the two steps of nonlinear post-processing and wake-up detection processing are usually independent of each other. This can easily lead to the enhancement of non-wake-up speech in the residual echo that is close to the wake-up speech signal, causing the wake-up algorithm to mistakenly identify it as a wake-up speech signal, thus resulting in false wake-ups. Although the wake-up rate is improved, the false wake-up rate is high. Summary of the Invention

[0004] This application provides a voice wake-up processing scheme that can enhance the voice wake-up rate while reducing the false wake-up rate.

[0005] The embodiments of this application provide the following technical solutions:

[0006] According to one embodiment of this application, a voice wake-up processing method includes: acquiring a voice signal to be processed; performing adaptive filtering echo cancellation processing on the voice signal to be processed to obtain an error signal; employing a nonlinear post-processing network to perform nonlinear echo cancellation post-processing based on a post-processing input signal and the voice signal to be processed to obtain an echo cancellation signal, wherein the post-processing input signal is obtained based on the error signal; employing a wake-up network to perform wake-up detection processing based on the echo cancellation signal, wherein the nonlinear post-processing network and the wake-up network are obtained through joint adversarial training, wherein the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator in the adversarial training, and the goal of the adversarial training is to make the discriminator increasingly more accurate in identifying whether the data generated by the generator is real or fake.

[0007] In some embodiments of this application, the nonlinear post-processing network is a real number network; the step of using a nonlinear post-processing network to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed to obtain an echo-cancelled signal includes: splitting the frequency domain signals of the post-processing input signal and the speech signal to be processed into real part spectra and imaginary part spectra respectively; inputting the real part spectra and the imaginary part spectra into the nonlinear post-processing network to perform echo cancellation nonlinear post-processing to obtain an echo-cancelled signal.

[0008] In some embodiments of this application, the step of inputting the real part spectrum and the imaginary part spectrum into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain an echo-cancelled signal includes: performing a first convolution process on the real part spectrum and the imaginary part spectrum to obtain a first convolution result; performing a first pooling process on the first convolution result to obtain a first pooling result; performing a second convolution process on the first pooling result to obtain a second convolution result; performing a second pooling process on the second convolution result to obtain a second pooling result; performing a third convolution process on the second pooling result to obtain a third convolution result; and performing a recurrent neural network on the third convolution result. Encoding processing is performed to obtain an encoding result; the encoding result and the third convolution result are concatenated and then subjected to a first deconvolution process to obtain a first deconvolution result; the first deconvolution result and the second convolution result are concatenated and then subjected to a second deconvolution process to obtain a second deconvolution result; the second deconvolution result and the first convolution result are concatenated and then subjected to a third deconvolution process to obtain a third deconvolution result; the third deconvolution result is subjected to a fourth convolution process to obtain a fourth convolution result; the fourth convolution result is multiplied by the real part spectrum and the imaginary part spectrum to obtain an output result; the echo cancellation signal is obtained based on the output result.

[0009] In some embodiments of this application, the use of a wake-up network to perform wake-up detection processing based on the echo cancellation signal includes: performing Mel filtering on the echo cancellation signal to obtain Mel cepstral features; inputting the Mel cepstral features into the wake-up network to perform wake-up detection processing to obtain a wake-up detection result.

[0010] In some embodiments of this application, the wake-up network is obtained by combining a temporal convolutional neural network with a recurrent neural network, wherein the recurrent neural network is located between the fully connected layer and the activation layer in the temporal convolutional neural network; the step of inputting the Mel-Cepstral features into the wake-up network for wake-up detection processing to obtain a wake-up detection result includes: inputting the Mel-Cepstral features into the temporal convolutional neural network for temporal convolution processing to obtain the fully connected operation result output by the fully connected layer; inputting the fully connected operation result into the recurrent neural network for recurrent neural network encoding processing to obtain a recurrent encoding result; and inputting the recurrent encoding result into the activation layer for activation processing to obtain the wake-up detection result.

[0011] In some embodiments of this application, before performing echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed using a nonlinear post-processing network to obtain an echo cancellation signal, the method further includes: buffering the error signal using a buffer; and obtaining the echo cancellation signal based on a predetermined number of buffered error signals.

[0012] In some embodiments of this application, the nonlinear post-processing network and the wake-up network are obtained by jointly performing adversarial training in the following manner: acquiring training data, the training data including sample speech signals and corresponding tag echo cancellation signals and tag wake-up detection results; performing iterative adversarial training on a preset nonlinear post-processing network and a preset wake-up network based on the training data to obtain the trained nonlinear post-processing network and the wake-up network, wherein each round of iterative adversarial training includes the following steps: performing adaptive filtering echo cancellation processing on the sample speech signals to obtain a sample error signal; using the preset nonlinear post-processing network, performing echo cancellation nonlinear post-processing based on the sample error signal and the sample speech signals to obtain a first predicted echo cancellation signal; and performing echo cancellation based on the first predicted echo cancellation signal and the tag echo cancellation signal. A first loss is calculated; using the preset wake-up network, wake-up detection processing is performed based on the first predicted echo cancellation signal to obtain a first wake-up detection result; a second loss is calculated based on the first wake-up detection result and the tag wake-up detection result; the parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss to obtain an updated nonlinear post-processing network; using the updated nonlinear post-processing network, echo cancellation nonlinear post-processing is performed based on the sample error signal and the sample speech signal to obtain a second predicted echo cancellation signal; using the preset wake-up network, wake-up detection processing is performed based on the second predicted echo cancellation signal to obtain a second wake-up detection result; a third loss is calculated based on the second wake-up detection result and the tag wake-up detection result; the parameters in the preset wake-up network are updated based on the third loss.

[0013] According to one embodiment of this application, a voice wake-up processing device includes: an acquisition module for acquiring a voice signal to be processed; an adaptive cancellation module for performing adaptive filtering echo cancellation processing on the voice signal to be processed to obtain an error signal; a nonlinear post-processing module for using a nonlinear post-processing network to perform echo cancellation nonlinear post-processing based on a post-processing input signal and the voice signal to be processed to obtain an echo cancellation signal, wherein the post-processing input signal is obtained based on the error signal; and a wake-up module for using a wake-up network to perform wake-up detection processing based on the echo cancellation signal, wherein the nonlinear post-processing network and the wake-up network are obtained through joint adversarial training, wherein the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator in the adversarial training, and the goal of the adversarial training is to make the discriminator increasingly accurate in identifying whether the data generated by the generator is real or fake.

[0014] According to another embodiment of this application, a storage medium stores a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the methods described in the embodiments of this application.

[0015] According to another embodiment of this application, an electronic device may include: a memory storing a computer program; and a processor reading the computer program stored in the memory to execute the methods described in the embodiments of this application.

[0016] According to another embodiment of this application, a computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of this application.

[0017] In this embodiment, a speech signal to be processed is acquired; adaptive filtering echo cancellation processing is performed on the speech signal to be processed to obtain an error signal; a nonlinear post-processing network is used to perform nonlinear echo cancellation post-processing based on the post-processing input signal and the speech signal to be processed to obtain an echo cancellation signal, wherein the post-processing input signal is obtained based on the error signal; a wake-up network is used to perform wake-up detection processing based on the echo cancellation signal, wherein the nonlinear post-processing network and the wake-up network are obtained by joint adversarial training, wherein the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator in the adversarial training, and the goal of the adversarial training is to make the discriminator's accuracy in identifying whether the data generated by the generator is real or fake increasingly higher.

[0018] In this way, the nonlinear post-processing network and the wake-up network obtained by joint adversarial training are used for wake-up processing, which effectively combines nonlinear post-processing and wake-up detection processing. Furthermore, in adversarial training, the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator. In contrast to the conventional adversarial training objective, the objective of adversarial training in this application is to make the discriminator more and more accurate in distinguishing between real and fake data generated by the generator, so that the data generated by the generator becomes easier for the discriminator to distinguish between real and fake data. This allows the voice wake-up processing result to enhance the wake-up rate while reducing the false wake-up rate caused by the wake-up processing process itself. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart of a voice wake-up processing method according to an embodiment of this application is shown.

[0021] Figure 2 A flowchart of signal post-processing according to an embodiment of this application is shown.

[0022] Figure 3 A flowchart of signal post-processing according to an embodiment of this application is shown.

[0023] Figure 4 It shows Figure 3 A schematic diagram of the structural elements of the generator in the embodiment.

[0024] Figure 5 A flowchart of a wake-up process according to an embodiment of this application is shown.

[0025] Figure 6 A flowchart of a wake-up process according to an embodiment of this application is shown.

[0026] Figure 7 A block diagram of a voice wake-up processing apparatus according to an embodiment of this application is shown.

[0027] Figure 8 A block diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0028] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Furthermore, the embodiments provided below are some embodiments for implementing the present disclosure, and not all embodiments for implementing the present disclosure. Unless otherwise specified, the technical solutions described in the embodiments of the present disclosure can be implemented in any combination.

[0029] It should be noted that, in the embodiments of this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or apparatus that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the apparatus, such as portions of circuitry, processors, programs, or software, etc.) in the method or apparatus that includes that element.

[0030] For example, the voice wake-up processing method provided in this embodiment includes a series of steps, but the voice wake-up processing method provided in this embodiment is not limited to the steps described. Similarly, the voice wake-up processing device provided in this embodiment includes a series of units, but the device provided in this embodiment is not limited to the units explicitly described, but may also include units that need to be set up for obtaining relevant information or processing based on information.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure.

[0032] Figure 1 A flowchart illustrating a voice wake-up processing method according to an embodiment of this application is shown. The device executing this voice wake-up processing method can be any device with processing capabilities, such as a television, speaker, computer, mobile phone, smartwatch, and home appliance.

[0033] like Figure 1 As shown, the voice wake-up processing method may include steps S110 to S140.

[0034] Step S110: Acquire the speech signal to be processed; Step S120: Perform adaptive filtering echo cancellation processing on the speech signal to be processed to obtain an error signal; Step S130: Use a nonlinear post-processing network to perform nonlinear echo cancellation post-processing based on the post-processing input signal and the speech signal to be processed to obtain an echo cancellation signal, wherein the post-processing input signal is obtained based on the error signal; Step S140: Use a wake-up network to perform wake-up detection processing based on the echo cancellation signal, wherein the nonlinear post-processing network and the wake-up network are obtained through joint adversarial training, wherein the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator in the adversarial training, and the goal of the adversarial training is to make the discriminator's accuracy in identifying whether the data generated by the generator is real or fake increasingly higher.

[0035] The nonlinear post-processing network and the wake-up network are obtained in advance through joint adversarial training. That is, the training method of adversarial network GAN is adopted to jointly train the nonlinear post-processing network and the wake-up network. In the adversarial training, the nonlinear post-processing network acts as the generator and the wake-up network acts as the discriminator.

[0036] In contrast to the conventional adversarial training objective (making the generator's generated data increasingly difficult for the discriminator to distinguish between real and fake data), the objective of adversarial training in this application is to make the discriminator more and more accurate in identifying the real and fake data generated by the generator, that is, to make the generator's generated data increasingly easier for the discriminator to distinguish between real and fake data.

[0037] The voice signal to be processed is the voice signal used for voice wake-up processing. Specifically, the voice signal to be processed can be the voice signal received by the device's microphone.

[0038] An adaptive filter can be used to perform adaptive filtering echo cancellation on the speech signal to be processed, and an error signal is obtained. The error signal is the canceled signal containing residual echoes.

[0039] The post-processing input signal, used to input the nonlinear post-processing network, can be obtained from the error signal. Using the nonlinear post-processing network, echo cancellation nonlinear post-processing is performed based on the post-processing input signal and the speech signal to be processed, yielding the echo-cancelled signal, which is the signal that eliminates residual echoes.

[0040] A wake-up network is used to perform wake-up detection processing based on echo cancellation signals. The wake-up detection results can be used to reflect whether the speech signal to be processed contains a specific wake-up word. The wake-up detection result can be 1 or 0. 1 indicates that the speech signal to be processed contains a specific wake-up word, and 0 indicates that the speech signal to be processed does not contain a specific wake-up word.

[0041] In this way, based on steps S110 to S140, wake-up processing is performed by using a nonlinear post-processing network and a wake-up network obtained through joint adversarial training. This effectively combines nonlinear post-processing and wake-up detection processing. Furthermore, in adversarial training, the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator. Unlike conventional adversarial training objectives, the objective of adversarial training in this application is to make the discriminator's accuracy in identifying whether the data generated by the generator is real or fake increasingly higher. This makes it easier for the discriminator to distinguish between real and fake data generated by the generator, thereby enhancing the wake-up rate while reducing the false wake-up rate caused by the wake-up processing process itself.

[0042] The following description Figure 1 Further optional specific embodiments for each step performed during voice wake-up processing in the example.

[0043] In one embodiment, the speech signal to be processed is subjected to adaptive filtering echo cancellation processing to obtain an error signal. Specifically, this may involve: performing a short-time Fourier transform on the speech signal to be processed to obtain a first transformed signal; performing a short-time Fourier transform on a reference signal to obtain a second transformed signal; and using an adaptive filter to perform adaptive filtering processing based on the first transformed signal and the second transformed signal to obtain the error signal.

[0044] Specifically, the speech signal to be processed, d(n), and the reference signal, x(n), can be input and short-time Fourier transforms can be performed respectively to obtain the first transformed signal D(l,k) and the second transformed signal X(l,k), where l is the frame index, k is the frequency index, and k = 1, 2, ..., K, and K is the number of points of the short-time Fourier transform (FFT).

[0045] Then, an adaptive filter can be used to perform adaptive filtering based on the first and second transformed signals to obtain the error signal. The adaptive algorithm used for adaptive filtering in the adaptive filter can be implemented using traditional adaptive algorithms such as NLMS, LMS, or RLS.

[0046] For example, taking frequency domain NLMS as the method for echo cancellation, the adaptive filtering process is implemented as follows:

[0047] E(l,k)=D(l,k)-Y(l,k),

[0048] Y(l,k)=X h (l,k)W(l,k),

[0049] E(l,k) is the signal after echo cancellation, called the error signal, X h (l,k) is the historical cached value of X(l,k): X h(l,k)=[X(l,k),X(l-1,k),...,X(l-ORD,k)], where ORD is the order of NLMS.

[0050] W(l,k) represents the filter coefficients: Where μ is the step size adjustment factor, · * This indicates the search for conjugate.

[0051] In one embodiment, the nonlinear post-processing network is a real number network; the step of using the nonlinear post-processing network to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed to obtain an echo-cancelled signal includes: splitting the frequency domain signals of the post-processing input signal and the speech signal to be processed into real part spectra and imaginary part spectra respectively; inputting the real part spectra and the imaginary part spectra into the nonlinear post-processing network to perform echo cancellation nonlinear post-processing to obtain an echo-cancelled signal.

[0052] See Figure 2 Step S210: The frequency domain signal of the speech signal to be processed can be decomposed into the corresponding real part spectrum and imaginary part spectrum; Step S220: The frequency domain signal of the post-processing input signal can be decomposed into the corresponding real part spectrum and imaginary part spectrum; Step S230: The real part spectrum and the imaginary part spectrum are input into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain the output result (including the enhanced real part spectrum and the enhanced imaginary part spectrum); Step S240: The output result (including the enhanced real part spectrum and the enhanced imaginary part spectrum) is combined into a frequency domain signal to obtain the echo cancellation signal.

[0053] The nonlinear post-processing network in this embodiment uses a real number network to decompose the frequency domain signals of the post-processing input signal and the speech signal to be processed into real part spectrum and imaginary part spectrum for processing. This can make full use of real part information and imaginary part information, and further reduce the false wake-up rate.

[0054] In one embodiment, the step of inputting the real part spectrum and the imaginary part spectrum into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain the echo-cancelled signal may specifically include:

[0055] The real and imaginary spectra are subjected to a first convolution process to obtain a first convolution result; the first convolution result is subjected to a first pooling process to obtain a first pooling result; the first pooling result is subjected to a second convolution process to obtain a second convolution result; the second convolution result is subjected to a second pooling process to obtain a second pooling result; the second pooling result is subjected to a third convolution process to obtain a third convolution result; the third convolution result is subjected to recurrent neural network encoding to obtain an encoding result; the encoding result and the third convolution result are concatenated and then subjected to a first deconvolution process to obtain a first deconvolution result; the first deconvolution result and the second convolution result are concatenated and then subjected to a second deconvolution process to obtain a second deconvolution result; the second deconvolution result and the first convolution result are concatenated and then subjected to a third deconvolution process to obtain a third deconvolution result; the third deconvolution result is subjected to a fourth convolution process to obtain a fourth convolution result; the fourth convolution result is multiplied by the real and imaginary spectra to obtain an output result; the echo cancellation signal is obtained based on the output result.

[0056] In this embodiment, the fourth convolution result obtained through multiple rounds of convolution processing, pooling processing, and deconvolution processing is multiplied with the real and imaginary part spectra of the frequency domain signals of the post-processing input signal and the speech signal to be processed to obtain the output result (including the enhanced real part spectrum and the enhanced imaginary part spectrum), which can effectively eliminate residual echoes.

[0057] In one specific implementation method, see [reference] Figure 3 and Figure 4 The specific network structure of the nonlinear post-processing network is as follows: Figure 3 As shown. Figure 3 The units or elements in the nonlinear post-processing network shown are as follows: Figure 4 The details are as follows:

[0058] ①Input layer: It is a combination of the real and imaginary part spectra of the frequency domain signals of the post-processing input signal and the speech signal to be processed, and has a (4×T×K) structure.

[0059] ②Output layer: The output results (including the enhanced real part spectrum and the enhanced imaginary part spectrum) have a (2×T×K) structure.

[0060] ③ Convolution block: Figure 3 and Figure 4A gradient rectangle is used in a convolutional module, which may further include one or more cascaded sub-convolutional modules. Each sub-convolutional module contains a batch normalization (BN) operation, a convolution (Conv) operation, and an activation function (such as the ReLU activation function).

[0061] ④pool block: Figure 3 and Figure 4 The gray rectangle represents a pooling module that contains a max pooling layer.

[0062] ⑤ RNN block (Recurrent Neural Network Module): Figure 3 and Figure 4 The white rectangle in the middle contains one or more recurrent neural network (RNN) layers connected in series, with the output of the previous layer serving as the input to the next layer. The RNN can be GRU, LSTM, BLSTM, etc.

[0063] ⑥ Deconvolution block: Figure 3 and Figure 4 The black rectangle in the middle is similar to the convolution block in ③. A deconvolution block can further include one or more cascaded sub-deconvolution blocks. Each sub-deconvolution block contains a batch normalization (BN) operation, a convolution (Conv) operation, and an activation function (such as the ReLU activation function). However, as... Figure 4 As shown, the last layer in the last sub-deconvolution module is the deconvolution operation (Deconv).

[0064] ⑦ Concatenate: This simply means to concatenate. Figure 3 and Figure 4 The dashed arrow in the middle indicates that the data can be concatenated according to the first dimension of the input audio signal (i.e., the dimension corresponding to the number of convolution kernels).

[0065] ⑧Multiplication: Figure 3 and Figure 4 The multiplication sign (or multiplication sign) specifically multiplies elements at the same position, such as... Figure 3 The output of the multiplication module in the structure diagram is Y.

[0066] Specifically, such as Figure 3As shown, in this embodiment, the network structure of the nonlinear post-processing network includes a first convolution module 301, a first pooling module 302, a second convolution module 303, a second pooling module 304, a third convolution module 305, a recurrent neural network module 306, a first deconvolution module 307, a second deconvolution module 308, a third deconvolution module 309, a fourth convolution module 310, and a multiplication module 311.

[0067] Specifically, in the first convolution module 301, a first convolution process is performed on the input real and imaginary spectra to obtain a first convolution result; in the first pooling module 302, a first pooling process is performed on the first convolution result to obtain a first pooling result; in the second convolution module 303, a second convolution process is performed on the first pooling result to obtain a second convolution result; in the second pooling module 304, a second pooling process is performed on the second convolution result to obtain a second pooling result; in the third convolution module 305, a third convolution process is performed on the second pooling result to obtain a third convolution result; in the recurrent neural network module 306, the third convolution result is subjected to recurrent neural network encoding processing to obtain an encoding result; and in the first deconvolution module 3... In module 07, the encoding result and the third convolution result can be concatenated and then subjected to a first deconvolution process to obtain a first deconvolution result; in module 308, the first deconvolution result and the second convolution result can be concatenated and then subjected to a second deconvolution process to obtain a second deconvolution result; in module 309, the second deconvolution result and the first convolution result can be concatenated and then subjected to a third deconvolution process to obtain a third deconvolution result; in module 310, the third deconvolution result can be subjected to a fourth convolution process to obtain a fourth convolution result; in module 311, the fourth convolution result can be multiplied with the real part spectrum and the imaginary part spectrum to obtain an output result (including the enhanced real part spectrum and the enhanced imaginary part spectrum).

[0068] Based on the output results, the output results (including the enhanced real part spectrum and the enhanced imaginary part spectrum) are combined into a frequency domain signal to obtain the echo cancellation signal.

[0069] Furthermore, in one embodiment, the wake-up network is used to perform wake-up detection processing based on the echo cancellation signal, which includes: performing Mel filtering on the echo cancellation signal to obtain Mel cepstral features; inputting the Mel cepstral features into the wake-up network to perform wake-up detection processing to obtain a wake-up detection result.

[0070] See Figure 5In step S410, the echo cancellation signal is processed by Mel filtering to obtain Mel cepstral features (i.e., MFCC features). Then, in step S420, the Mel cepstral features are input into the wake-up network for wake-up detection processing, which can accurately obtain the wake-up detection results.

[0071] In one specific embodiment, the wake-up network is obtained by combining a temporal convolutional neural network with a recurrent neural network, wherein the recurrent neural network is located between the fully connected layer and the activation layer in the temporal convolutional neural network; the step of inputting the Mel-Cepstral features into the wake-up network for wake-up detection processing to obtain a wake-up detection result includes: inputting the Mel-Cepstral features into the temporal convolutional neural network for temporal convolution processing to obtain the fully connected operation result output by the fully connected layer; inputting the fully connected operation result into the recurrent neural network for recurrent neural network encoding processing to obtain a recurrent encoding result; and inputting the recurrent encoding result into the activation layer for activation processing to obtain the wake-up detection result.

[0072] In this embodiment, the wake-up network is obtained by combining a temporal convolutional neural network (TC-ResNet) with a recurrent neural network (RNN). The RNN is located between the fully connected (FC) layers and activation layers in the TC-ResNet. Specifically, the RNN can be an LSTM or a GRU, etc.

[0073] See Figure 6 In one approach, the Temporal Convolutional Neural Network (TC-ResNet) specifically uses TC-ResNet8, which includes a convolutional layer (conv) 510, a first residual block (Blocks) 520, a second residual block (Blocks) 530, a third residual block (Blocks) 540, a pooling layer (Pool) 550, a fully connected layer (FC) 560, and an activation layer (Softmax) 570. The recurrent neural network (RNN Block) 580 is located between the fully connected layer (FC) 560 and the activation layer (Softmax) 570.

[0074] Mel-frequency cepstral features (MFCC) are input into a temporal convolutional neural network for temporal convolution processing. Each layer is processed sequentially to obtain the fully connected operation result output by the fully connected layer. The fully connected operation result is input into a recurrent neural network for recurrent neural network encoding processing to obtain the recurrent encoding result. The recurrent encoding result is input into an activation layer for activation processing to obtain the wake-up detection result B.

[0075] Combining a temporal convolutional neural network with a recurrent neural network (RNN) yields a wake-up network for wake-up detection processing. The wake-up network can leverage the memory property of RNNs to perform wake-up detection processing on the Mel-Cepstral features of streaming speech signals.

[0076] Furthermore, in one embodiment, before performing echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed using a nonlinear post-processing network to obtain an echo cancellation signal, the method further includes: buffering the error signal using a buffer; and obtaining the echo cancellation signal based on a predetermined number of buffered error signals.

[0077] The streaming data is processed by the wake-up network based on the nonlinear post-processing network. The predetermined number of frames can be set according to the actual situation. The error signal of the predetermined number of frames T is buffered by the buffer as the echo cancellation signal, and then sent to the nonlinear post-processing network to obtain the output of the nonlinear post-processing network. After calculating the MFCC feature, it is sent to the wake-up network to obtain a wake-up result. Only the last piece of data of the actual wake-up speech can get the result of 1, and all other cases can only get the result of 0.

[0078] Furthermore, in one embodiment of this application, the nonlinear post-processing network and the wake-up network in the foregoing embodiments can specifically be obtained by jointly performing adversarial training in the following manner:

[0079] Acquire training data, which includes sample speech signals and corresponding tag echo cancellation signals and tag wake-up detection results; perform iterative adversarial training on a preset nonlinear post-processing network and a preset wake-up network based on the training data to obtain the trained nonlinear post-processing network and the wake-up network, wherein each round of iterative adversarial training includes the following steps:

[0080] The sample speech signal is subjected to adaptive filtering echo cancellation processing to obtain a sample error signal; the preset nonlinear post-processing network is used to perform echo cancellation nonlinear post-processing based on the sample error signal and the sample speech signal to obtain a first predicted echo cancellation signal; a first loss is calculated based on the first predicted echo cancellation signal and the tag echo cancellation signal; the preset wake-up network is used to perform wake-up detection processing based on the first predicted echo cancellation signal to obtain a first wake-up detection result; a second loss is calculated based on the first wake-up detection result and the tag wake-up detection result; the parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss to obtain an updated nonlinear post-processing network; the updated nonlinear post-processing network is used to perform echo cancellation nonlinear post-processing based on the sample error signal and the sample speech signal to obtain a second predicted echo cancellation signal; the preset wake-up network is used to perform wake-up detection processing based on the second predicted echo cancellation signal to obtain a second wake-up detection result; a third loss is calculated based on the second wake-up detection result and the tag wake-up detection result; the parameters in the preset wake-up network are updated based on the third loss.

[0081] Based on the training data, iterative adversarial training is performed on the preset nonlinear post-processing network and the preset wake-up network. The process is repeated until both networks converge, resulting in the trained nonlinear post-processing network and wake-up network. If it is not possible to guarantee that both networks converge to the best result, then the wake-up network can be guaranteed to converge to the best result because it has a higher priority than the nonlinear post-processing network.

[0082] The specific steps in each round of iterative adversarial training include: (1) iteratively updating the parameters of the nonlinear post-processing network; (2) using the wake-up network to calculate the loss value; and (3) iteratively updating the parameters of the wake-up network based on the loss value.

[0083] Specifically, (1) the first loss LOSS1 is calculated based on the first predicted echo cancellation signal and the tag echo cancellation signal, and the second loss LOSS2 is calculated based on the first wake-up detection result B1 and the tag wake-up detection result. The parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss to obtain the updated nonlinear post-processing network. Specifically, the LOSS2 can be calculated. G = LOSS1 + LOSS2, based on the combined loss LOSS G Update the parameters of the preset nonlinear post-processing network.

[0084] (2) In step (1), the parameters of the nonlinear post-processing network were updated. Now, the training corpus (the sample error signal and the sample speech signal) is fed into this newly updated nonlinear post-processing network to obtain a new output second predicted echo cancellation signal Z. Further, a preset wake-up network is used to perform wake-up detection processing based on the second predicted echo cancellation signal to obtain the second wake-up detection result B2. The third loss LOSS is calculated based on the second wake-up detection result and the label wake-up detection result. D .

[0085] (3) According to the third loss LOSS D Update the parameters in the preset wake-up network.

[0086] Furthermore, in each round of adversarial training, the loss value of the wake-up network is used to participate in the parameter iteration update of the nonlinear post-processing network, so that the output of the post-processing network can be adapted to the wake-up network. This further ensures that when the trained nonlinear post-processing network and the wake-up network are used for voice wake-up processing, they can improve the wake-up rate in harsh environments without increasing false wake-ups.

[0087] Specifically, a preset nonlinear post-processing network is used to perform echo cancellation nonlinear post-processing based on the sample error signal and the sample speech signal. Specifically, the sample error signal and the sample speech signal are decomposed into real part spectrum and imaginary part spectrum respectively and then input into the preset nonlinear post-processing network for echo cancellation nonlinear post-processing.

[0088] An updated nonlinear post-processing network is used to perform echo cancellation nonlinear post-processing based on the sample error signal and the sample speech signal. Specifically, the sample error signal and the sample speech signal are decomposed into real part spectrum and imaginary part spectrum, respectively, and then input into the updated nonlinear post-processing network for echo cancellation nonlinear post-processing.

[0089] In one specific implementation, a first loss is calculated based on the first predicted echo cancellation signal and the tag echo cancellation signal, specifically using KL divergence to calculate the first loss LOSS1; a second loss is calculated based on the first wake-up detection result and the tag wake-up detection result, specifically using MAE (Mean Absolute Error) to calculate the second loss LOSS2; and a third loss is calculated based on the second wake-up detection result and the tag wake-up detection result, specifically using MAE (Mean Absolute Error) to calculate the third loss LOSS. D In this way, the effectiveness of joint combat training can be effectively guaranteed.

[0090] To facilitate better implementation of the voice wake-up processing method provided in this application, this application also provides a voice wake-up processing device based on the above-described voice wake-up processing method. The meanings of the terms used are the same as in the above-described voice wake-up processing method, and specific implementation details can be found in the descriptions within the method embodiments. Figure 7 A block diagram of a voice wake-up processing apparatus according to an embodiment of this application is shown.

[0091] like Figure 7 As shown, the voice wake-up processing device 600 may include: an acquisition module 610 for acquiring a voice signal to be processed; an adaptive cancellation module 620 for performing adaptive filtering echo cancellation processing on the voice signal to be processed to obtain an error signal; a nonlinear post-processing module 630 for using a nonlinear post-processing network to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the voice signal to be processed to obtain an echo cancellation signal, wherein the post-processing input signal is obtained based on the error signal; and a wake-up module 640 for using a wake-up network to perform wake-up detection processing based on the echo cancellation signal, wherein the nonlinear post-processing network and the wake-up network are obtained through joint adversarial training, wherein the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator in the adversarial training, and the goal of the adversarial training is to make the discriminator's accuracy in identifying whether the data generated by the generator is true or false increasingly higher.

[0092] In some embodiments of this application, the nonlinear post-processing network is a real number network; the nonlinear post-processing module is used to: decompose the frequency domain signals of the post-processing input signal and the speech signal to be processed into real part spectrum and imaginary part spectrum respectively; input the real part spectrum and the imaginary part spectrum into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain an echo cancellation signal.

[0093] In some embodiments of this application, the nonlinear post-processing module is configured to: perform a first convolution process on the real part spectrum and the imaginary part spectrum to obtain a first convolution result; perform a first pooling process on the first convolution result to obtain a first pooling result; perform a second convolution process on the first pooling result to obtain a second convolution result; perform a second pooling process on the second convolution result to obtain a second pooling result; perform a third convolution process on the second pooling result to obtain a third convolution result; perform recurrent neural network encoding processing on the third convolution result to obtain an encoding result; and convert the encoding result into a digital representation of the result. The first deconvolution result is obtained by concatenating the first deconvolution result with the second convolution result and performing a first deconvolution process. The second deconvolution result is then concatenated with the first convolution result and performed a second deconvolution process. The third deconvolution result is then concatenated with the first convolution result and performed a third deconvolution process. The third deconvolution result is then subjected to a fourth convolution process. The fourth convolution result is multiplied by the real part spectrum and the imaginary part spectrum to obtain an output result. Based on the output result, the echo cancellation signal is obtained.

[0094] In some embodiments of this application, the wake-up module 640 can be used to: perform Mel filtering on the echo cancellation signal to obtain Mel cepstral features; input the Mel cepstral features into the wake-up network for wake-up detection processing to obtain a wake-up detection result.

[0095] In some embodiments of this application, the wake-up network is obtained by combining a temporal convolutional neural network with a recurrent neural network, wherein the recurrent neural network is located between the fully connected layer and the activation layer in the temporal convolutional neural network; the wake-up module 640 can be used to: input the Mel-Cepstral features into the temporal convolutional neural network for temporal convolution processing to obtain the fully connected operation result output by the fully connected layer; input the fully connected operation result into the recurrent neural network for recurrent neural network encoding processing to obtain a recurrent encoding result; input the recurrent encoding result into the activation layer for activation processing to obtain the wake-up detection result.

[0096] In some embodiments of this application, before the nonlinear post-processing network is used to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed to obtain the echo cancellation signal, the device further includes a buffer module for: buffering the error signal using a buffer; and obtaining the echo cancellation signal based on a predetermined number of buffered error signals.

[0097] In some embodiments of this application, the apparatus further includes a training module, configured to: acquire training data, the training data including sample speech signals and corresponding tag echo cancellation signals and tag wake-up detection results; perform iterative adversarial training on a preset nonlinear post-processing network and a preset wake-up network based on the training data to obtain the trained nonlinear post-processing network and the wake-up network, wherein each round of iterative adversarial training includes the following steps: performing adaptive filtering echo cancellation processing on the sample speech signals to obtain a sample error signal; using the preset nonlinear post-processing network, performing nonlinear echo cancellation post-processing based on the sample error signal and the sample speech signals to obtain a first predicted echo cancellation signal; calculating a first loss based on the first predicted echo cancellation signal and the tag echo cancellation signal; and so on. Using the preset wake-up network, wake-up detection processing is performed based on the first predicted echo cancellation signal to obtain a first wake-up detection result; a second loss is calculated based on the first wake-up detection result and the tag wake-up detection result; the parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss to obtain an updated nonlinear post-processing network; using the updated nonlinear post-processing network, echo cancellation nonlinear post-processing is performed based on the sample error signal and the sample speech signal to obtain a second predicted echo cancellation signal; using the preset wake-up network, wake-up detection processing is performed based on the second predicted echo cancellation signal to obtain a second wake-up detection result; a third loss is calculated based on the second wake-up detection result and the tag wake-up detection result; the parameters in the preset wake-up network are updated based on the third loss.

[0098] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0099] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 8 As shown, Figure 8 A block diagram of an electronic device according to an embodiment of this application is shown, specifically:

[0100] The electronic device may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, and an input unit 704. Those skilled in the art will understand that... Figure 8The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0101] in:

[0102] The processor 701 is the control center of the electronic device. It connects to various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and application programs, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 701.

[0103] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0104] The electronic device also includes a power supply 703 that supplies power to the various components. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0105] The electronic device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0106] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 702 according to the following instructions, and the processor 701 runs the computer programs stored in the memory 702, thereby realizing the various functions in the foregoing embodiments of this application. For example, the processor 701 can perform the following steps:

[0107] The process involves: acquiring a speech signal to be processed; performing adaptive filtering echo cancellation processing on the speech signal to obtain an error signal; employing a nonlinear post-processing network to perform nonlinear echo cancellation post-processing based on the post-processing input signal and the speech signal to be processed to obtain an echo cancellation signal, wherein the post-processing input signal is obtained based on the error signal; and employing a wake-up network to perform wake-up detection processing based on the echo cancellation signal. The nonlinear post-processing network and the wake-up network are jointly trained adversarially, wherein the nonlinear post-processing network acts as a generator and the wake-up network acts as a discriminator in the adversarial training, and the goal of the adversarial training is to increase the accuracy of the discriminator in identifying whether the data generated by the generator is real or fake.

[0108] In some embodiments of this application, the nonlinear post-processing network is a real number network; the step of using a nonlinear post-processing network to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed to obtain an echo-cancelled signal includes: splitting the frequency domain signals of the post-processing input signal and the speech signal to be processed into real part spectra and imaginary part spectra respectively; inputting the real part spectra and the imaginary part spectra into the nonlinear post-processing network to perform echo cancellation nonlinear post-processing to obtain an echo-cancelled signal.

[0109] In some embodiments of this application, the step of inputting the real part spectrum and the imaginary part spectrum into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain an echo-cancelled signal includes: performing a first convolution process on the real part spectrum and the imaginary part spectrum to obtain a first convolution result; performing a first pooling process on the first convolution result to obtain a first pooling result; performing a second convolution process on the first pooling result to obtain a second convolution result; performing a second pooling process on the second convolution result to obtain a second pooling result; performing a third convolution process on the second pooling result to obtain a third convolution result; and performing a recurrent neural network on the third convolution result. Encoding processing is performed to obtain an encoding result; the encoding result and the third convolution result are concatenated and then subjected to a first deconvolution process to obtain a first deconvolution result; the first deconvolution result and the second convolution result are concatenated and then subjected to a second deconvolution process to obtain a second deconvolution result; the second deconvolution result and the first convolution result are concatenated and then subjected to a third deconvolution process to obtain a third deconvolution result; the third deconvolution result is subjected to a fourth convolution process to obtain a fourth convolution result; the fourth convolution result is multiplied by the real part spectrum and the imaginary part spectrum to obtain an output result; the echo cancellation signal is obtained based on the output result.

[0110] In some embodiments of this application, the use of a wake-up network to perform wake-up detection processing based on the echo cancellation signal includes: performing Mel filtering on the echo cancellation signal to obtain Mel cepstral features; inputting the Mel cepstral features into the wake-up network to perform wake-up detection processing to obtain a wake-up detection result.

[0111] In some embodiments of this application, the wake-up network is obtained by combining a temporal convolutional neural network with a recurrent neural network, wherein the recurrent neural network is located between the fully connected layer and the activation layer in the temporal convolutional neural network; the step of inputting the Mel-Cepstral features into the wake-up network for wake-up detection processing to obtain a wake-up detection result includes: inputting the Mel-Cepstral features into the temporal convolutional neural network for temporal convolution processing to obtain the fully connected operation result output by the fully connected layer; inputting the fully connected operation result into the recurrent neural network for recurrent neural network encoding processing to obtain a recurrent encoding result; and inputting the recurrent encoding result into the activation layer for activation processing to obtain the wake-up detection result.

[0112] In some embodiments of this application, before performing echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed using a nonlinear post-processing network to obtain an echo cancellation signal, the method further includes: buffering the error signal using a buffer; and obtaining the echo cancellation signal based on a predetermined number of buffered error signals.

[0113] In some embodiments of this application, the nonlinear post-processing network and the wake-up network are obtained by jointly performing adversarial training in the following manner: acquiring training data, the training data including sample speech signals and corresponding tag echo cancellation signals and tag wake-up detection results; performing iterative adversarial training on a preset nonlinear post-processing network and a preset wake-up network based on the training data to obtain the trained nonlinear post-processing network and the wake-up network, wherein each round of iterative adversarial training includes the following steps: performing adaptive filtering echo cancellation processing on the sample speech signals to obtain a sample error signal; using the preset nonlinear post-processing network, performing echo cancellation nonlinear post-processing based on the sample error signal and the sample speech signals to obtain a first predicted echo cancellation signal; and performing echo cancellation based on the first predicted echo cancellation signal and the tag echo cancellation signal. A first loss is calculated; using the preset wake-up network, wake-up detection processing is performed based on the first predicted echo cancellation signal to obtain a first wake-up detection result; a second loss is calculated based on the first wake-up detection result and the tag wake-up detection result; the parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss to obtain an updated nonlinear post-processing network; using the updated nonlinear post-processing network, echo cancellation nonlinear post-processing is performed based on the sample error signal and the sample speech signal to obtain a second predicted echo cancellation signal; using the preset wake-up network, wake-up detection processing is performed based on the second predicted echo cancellation signal to obtain a second wake-up detection result; a third loss is calculated based on the second wake-up detection result and the tag wake-up detection result; the parameters in the preset wake-up network are updated based on the third loss.

[0114] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0115] Therefore, embodiments of this application also provide a storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the methods provided in embodiments of this application.

[0116] The storage medium can be a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0117] Since the computer program stored in the storage medium can execute the steps of any of the methods provided in the embodiments of this application, the beneficial effects that the methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0118] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0119] It should be understood that this application is not limited to the embodiments described above and shown in the accompanying drawings, but various modifications and changes can be made without departing from its scope.

Claims

1. A voice wake-up processing method, characterized in that, include: Acquire the speech signal to be processed; The speech signal to be processed is subjected to adaptive filtering and echo cancellation processing to obtain an error signal; A nonlinear post-processing network is used to perform nonlinear echo cancellation post-processing based on the post-processing input signal and the speech signal to be processed, to obtain an echo cancellation signal. The post-processing input signal is obtained based on the error signal. A wake-up network is used to perform wake-up detection processing based on the echo cancellation signal; The nonlinear post-processing network and the wake-up network are jointly trained in the following manner: Acquire training data, which includes sample speech signals and corresponding tag echo cancellation signals and tag wake-up detection results for the sample speech signals; Based on the training data, a preset nonlinear post-processing network and a preset wake-up network are iteratively trained to obtain the trained nonlinear post-processing network and the wake-up network. Each round of iterative training includes the following steps: The sample speech signal is subjected to adaptive filtering and echo cancellation processing to obtain the sample error signal; Using the preset nonlinear post-processing network, echo cancellation nonlinear post-processing is performed based on the sample error signal and the sample speech signal to obtain the first predicted echo cancellation signal; The first loss is calculated based on the first predicted echo cancellation signal and the tag echo cancellation signal; Using the preset wake-up network, wake-up detection processing is performed based on the first predicted echo cancellation signal to obtain the first wake-up detection result; The second loss is calculated based on the first wake-up detection result and the tag wake-up detection result; The parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss to obtain the updated nonlinear post-processing network; Using the updated nonlinear post-processing network, echo cancellation nonlinear post-processing is performed based on the sample error signal and the sample speech signal to obtain the second predicted echo cancellation signal. Using the preset wake-up network, wake-up detection processing is performed based on the second predicted echo cancellation signal to obtain the second wake-up detection result; The third loss is calculated based on the second wake-up detection result and the tag wake-up detection result; The parameters in the preset wake-up network are updated based on the third loss.

2. The method according to claim 1, characterized in that, The nonlinear post-processing network is a real-number network; the nonlinear post-processing network is used to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed, to obtain an echo-cancelled signal, including: The frequency domain signals of the post-processing input signal and the speech signal to be processed are respectively decomposed into real part spectrum and imaginary part spectrum; The real and imaginary spectra are input into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain the echo-cancelled signal.

3. The method according to claim 2, characterized in that, The step of inputting the real part spectrum and the imaginary part spectrum into the nonlinear post-processing network for echo cancellation nonlinear post-processing to obtain the echo-cancelled signal includes: The real part spectrum and the imaginary part spectrum are subjected to a first convolution process to obtain the first convolution result; The first convolution result is subjected to the first pooling process to obtain the first pooling result; The first pooling result is subjected to a second convolution process to obtain the second convolution result; The second convolution result is subjected to a second pooling process to obtain the second pooling result; The second pooling result is subjected to a third convolution to obtain the third convolution result; The third convolution result is processed by a recurrent neural network to obtain the encoded result; The encoding result and the third convolution result are concatenated and then subjected to a first deconvolution process to obtain the first deconvolution result. The first deconvolution result and the second convolution result are concatenated and then subjected to a second deconvolution process to obtain the second deconvolution result. The second deconvolution result and the first convolution result are concatenated and then subjected to a third deconvolution process to obtain the third deconvolution result. The third deconvolution result is subjected to a fourth convolution to obtain the fourth convolution result; The fourth convolution result is multiplied by the real part spectrum and the imaginary part spectrum to obtain the output result; The echo cancellation signal is obtained based on the output result.

4. The method according to claim 1, characterized in that, The wake-up network is used to perform wake-up detection processing based on the echo cancellation signal, including: The echo-cancelled signal is subjected to Mel filtering to obtain Mel cepstral features; The Mel-Cepstral features are input into the wake-up network for wake-up detection processing to obtain the wake-up detection result.

5. The method according to claim 4, characterized in that, The wake-up network is obtained by combining a temporal convolutional neural network with a recurrent neural network, wherein the recurrent neural network is located between the fully connected layer and the activation layer in the temporal convolutional neural network; The step of inputting the Mel-Cepstral features into the wake-up network for wake-up detection processing to obtain the wake-up detection result includes: The Mel-Cepstral features are input into the temporal convolutional neural network for temporal convolution processing to obtain the fully connected operation result output by the fully connected layer; The result of the fully connected operation is input into the recurrent neural network for recurrent neural network encoding processing to obtain the recurrent encoding result; The cyclic encoding result is input into the activation layer for activation processing to obtain the wake-up detection result.

6. The method according to claim 1, characterized in that, Before performing echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed using a nonlinear post-processing network to obtain the echo-cancelled signal, the method further includes: The error signal is buffered using a buffer. The echo cancellation signal is obtained based on a predetermined number of error signals in the buffer.

7. A voice wake-up processing device, characterized in that, include: The acquisition module is used to acquire the voice signal to be processed; An adaptive echo cancellation module is used to perform adaptive filtering echo cancellation processing on the speech signal to be processed to obtain an error signal. A nonlinear post-processing module is used to perform echo cancellation nonlinear post-processing based on the post-processing input signal and the speech signal to be processed using a nonlinear post-processing network to obtain an echo cancellation signal. The post-processing input signal is obtained based on the error signal. A wake-up module is used to perform wake-up detection processing based on the echo cancellation signal using a wake-up network. The nonlinear post-processing network and the wake-up network are jointly trained in the following manner: acquiring training data, the training data including sample speech signals and corresponding label echo cancellation signals and label wake-up detection results; iteratively training a preset nonlinear post-processing network and a preset wake-up network based on the training data to obtain the trained nonlinear post-processing network and the wake-up network, wherein each iteration includes the following steps: performing adaptive filtering echo cancellation processing on the sample speech signal to obtain a sample error signal; using the preset nonlinear post-processing network, performing nonlinear echo cancellation post-processing based on the sample error signal and the sample speech signal to obtain a first predicted echo cancellation signal; and based on the first predicted echo cancellation signal and the label... A first loss is calculated from the echo cancellation signal; a first wake-up detection result is obtained by using the preset wake-up network based on the first predicted echo cancellation signal; a second loss is calculated based on the first wake-up detection result and the tag wake-up detection result; the parameters of the preset nonlinear post-processing network are updated by combining the first loss and the second loss; an updated nonlinear post-processing network is obtained by using the updated nonlinear post-processing network based on the sample error signal and the sample speech signal to obtain a second predicted echo cancellation signal; a second wake-up detection result is obtained by using the preset wake-up network based on the second predicted echo cancellation signal; a third loss is calculated based on the second wake-up detection result and the tag wake-up detection result; and the parameters in the preset wake-up network are updated based on the third loss.

8. A storage medium, characterized in that, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the method described in any one of claims 1 to 6.

9. An electronic device, characterized in that, include: Memory, which stores computer programs; A processor reads a computer program stored in memory to perform the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Application awakening method and device, storage medium and electronic equipment

    CN110211599A

  • Model training method and device, audio processing method and device, equipment, storage medium and program

    CN114512136A