Signal processing method, device, storage medium and system
By jointly optimizing the sensor array signals through signal model and Kalman filtering method, the echo cancellation and reverberation suppression problems in extremely low signal-to-echo ratio and large reverberation scenarios are solved, and high-quality speech signal recovery is achieved.
Patent Information
- Application Number
- CN202211640661.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-12-20
AI Technical Summary
In scenarios with extremely low signal-to-echo ratios, nonlinear echoes, and large reverberations, existing technologies have difficulty effectively handling echo cancellation and reverberation suppression, resulting in the target voice signal quality being unable to meet the requirements of voice interaction scenarios.
The signal model is used to perform signal conversion processing on the observation signal captured by the sensor array to obtain the echo path parameters, multi-channel autoregressive coefficients, playback signal and target microphone observation signal. The Kalman filtering method is used for joint optimization processing to determine the target speech signal to be restored.
The echo cancellation and reverberation suppression effects of the target voice signal are improved, thereby enhancing the voice signal quality in voice interaction scenarios.
Smart Images

Figure CN116092510B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a signal processing method, device, storage medium and system. Background Art
[0002] Currently, far-field voice interaction technology is an artificial intelligence (AI) voice interaction technology widely used in smart devices such as televisions and speakers. Echo cancellation and reverberation suppression are two key issues in far-field voice interaction technology. Echo cancellation eliminates the echo generated by the playback signal in far-field voice interaction systems, while reverberation suppression reduces the reverberation effect of the observation signal to preserve the direct sound and early reflections of the speech.
[0003] As voice interaction scenarios become increasingly diverse, voice signal processing faces greater challenges. For example, voice signal processing (particularly echo cancellation and reverberation suppression) is challenging in scenarios with extremely low signal-to-response ratios, nonlinear echoes, and high reverberation. Therefore, expanding the performance boundaries of signal processing methods to enable their application in a wider range of scenarios has become a key issue in this field.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] Embodiments of the present invention provide a signal processing method, device, storage medium, and system to at least solve the technical problem in the related art that the quality of the target voice signal is difficult to meet the requirements of the voice interaction scenario due to independent echo cancellation processing and reverberation suppression processing of the signal.
[0006] According to one aspect of an embodiment of the present invention, a signal processing method is provided, comprising: performing signal conversion processing on a first signal to obtain a second signal, wherein the first signal is an observation signal captured by a sensor array, and the second signal is a reverberation speech signal described by a signal model; obtaining a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band; and determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal, and the fourth signal, wherein the fifth signal is the target speech signal to be recovered.
[0007] According to another aspect of an embodiment of the present invention, a signal processing method is also provided, including: receiving a first signal from a client, wherein the first signal is an observation signal captured by a sensor array; performing signal conversion processing on the first signal to obtain a second signal, obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, and determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, wherein the second signal is a reverberation speech signal described by a signal model, the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band, and the fifth signal is a target speech signal to be restored; and feeding the fifth signal back to the client.
[0008] According to another aspect of an embodiment of the present invention, a signal processing device is also provided, including: a processing module for performing signal conversion processing on a first signal to obtain a second signal, wherein the first signal is an observation signal captured by a sensor array, and the second signal is a reverberation speech signal described by a signal model; an acquisition module for acquiring a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band; a determination module for determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, wherein the fifth signal is the target speech signal to be restored.
[0009] According to another aspect of an embodiment of the present invention, a storage medium is further provided. The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute any one of the above-mentioned signal processing methods.
[0010] According to another aspect of an embodiment of the present invention, a signal processing system is also provided, including: a processor; and a memory, connected to the above-mentioned processor, for providing the above-mentioned processor with instructions for processing the following processing steps: performing signal conversion processing on a first signal to obtain a second signal, wherein the first signal is an observation signal captured by a sensor array, and the second signal is a reverberation speech signal described by a signal model; obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band; using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, determining a fifth signal, wherein the fifth signal is the target speech signal to be restored.
[0011] In an embodiment of the present invention, a second signal is obtained by performing signal conversion processing on a first signal, wherein the first signal is an observation signal captured by a sensor array, the second signal is a reverberation speech signal described by a signal model, and a first parameter, a second parameter, a third signal and a fourth signal are obtained based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band. The second signal, the first parameter, the second parameter, the third signal and the fourth signal are further used to determine a fifth signal, wherein the fifth signal is a target speech signal to be recovered. The purpose of processing the observation signal to obtain a high-quality target speech signal is achieved, thereby achieving the technical effect of improving the echo cancellation amount and reverberation suppression effect of the obtained target speech signal, and further solving the technical problem in the related art that the quality of the target speech signal is difficult to meet the requirements of the voice interaction scenario due to independent echo cancellation processing and reverberation suppression processing of the signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0013] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a signal processing method is shown;
[0014] Figure 2 is a flowchart of a signal processing method according to an embodiment of the present invention;
[0015] Figure 3 is a schematic diagram of an optional semantic signal processing scenario according to an embodiment of the present invention;
[0016] Figure 4 is a flowchart of another signal processing method according to an embodiment of the present invention;
[0017] Figure 5 is a schematic diagram of signal processing on a cloud server according to an embodiment of the present invention;
[0018] Figure 6 is a structural diagram of a signal processing device according to an embodiment of the present invention;
[0019] Figure 7 is a structural diagram of another signal processing device according to an embodiment of the present invention;
[0020] Figure 8is a structural block diagram of another computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] First, some nouns or terms that appear in the description of the embodiments of the present invention are subject to the following explanations:
[0024] Kalman filtering: An algorithm that uses the linear system state equation to optimally estimate the system state using observation data from the system's input and output. Because the observation data includes the effects of noise and interference in the system, optimal estimation can also be viewed as a filtering process.
[0025] Echo path: refers to the acoustic transfer function from the speaker to the microphone.
[0026] Example 1
[0027] According to an embodiment of the present invention, a signal processing method embodiment is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0028] The method embodiment provided in Embodiment 1 of the present invention may be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a signal processing method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, a keyboard, a cursor control device (such as a mouse), an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0029] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present invention, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the signal processing method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned signal processing method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0031] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0033] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the aforementioned computer device (or mobile device).
[0034] Under the above operating environment, the present invention provides Figure 2 A signal processing method shown. Figure 2 is a flow chart of a signal processing method according to an embodiment of the present invention. Figure 2 As shown, the signal processing method includes:
[0035] Step S21, performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by the sensor array, and the second signal is a reverberated speech signal described by a signal model;
[0036] Step S22: acquiring a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band;
[0037] Step S23: Determine a fifth signal using the second signal, the first parameter, the second parameter, the third signal, and the fourth signal, wherein the fifth signal is the target speech signal to be restored.
[0038] The sensor array includes multiple microphones. The first signal is an observation signal captured by the sensor array, and the first signal is a time domain signal. Signal conversion processing is performed on the first signal to obtain the second signal. The second signal may be a reverberated speech signal described by a signal model, and the second signal is a frequency domain signal.
[0039] The first parameter is the echo path parameter during the speech signal interaction process, and the second parameter is the multichannel autoregressive coefficient during the speech signal interaction process. Based on the reverberant speech signal, the echo path parameter and the multichannel autoregressive coefficient can be obtained. Furthermore, the playback signal corresponding to the reverberant speech signal and the observation signal of the target microphone in the target frequency band can also be obtained. The target microphone is any sensor selected from the sensor array.
[0040] The target speech signal to be recovered can be determined using the reverberant speech signal, the echo path parameters, the multi-channel autoregressive coefficients, the playback signal, and the observation signal of the target microphone in the target frequency band. The target speech signal is the result of echo cancellation and reverberation suppression processing on the observation signal.
[0041] Optionally, the above-mentioned signal processing method provided by the present invention can be applied to, but is not limited to, far-field voice interaction scenarios in large-screen full-duplex interactive voice systems related to artificial intelligence (such as artificial intelligence Internet of Things systems, smart speakers, smart voice assistants, etc.).
[0042] In an embodiment of the present invention, a second signal is obtained by performing signal conversion processing on a first signal, wherein the first signal is an observation signal captured by a sensor array, the second signal is a reverberation speech signal described by a signal model, and a first parameter, a second parameter, a third signal and a fourth signal are obtained based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band. The second signal, the first parameter, the second parameter, the third signal and the fourth signal are further used to determine a fifth signal, wherein the fifth signal is a target speech signal to be recovered. The purpose of processing the observation signal to obtain a high-quality target speech signal is achieved, thereby achieving the technical effect of improving the echo cancellation amount and reverberation suppression effect of the obtained target speech signal, and further solving the technical problem in the related art that the quality of the target speech signal is difficult to meet the requirements of the voice interaction scenario due to independent echo cancellation processing and reverberation suppression processing of the signals.
[0043] The above method of the embodiment of the present invention is further introduced below.
[0044] In an optional embodiment, the first parameter is used to adjust the echo generated by the third signal, and the second parameter is used to adjust the reverberation of the fourth signal.
[0045] The first parameter, i.e., the echo path parameter during the speech signal interaction process, is used to adjust the echo generated by the playback signal corresponding to the reverberant speech signal. Generally, the first parameter can be used to perform echo cancellation on the signal.
[0046] The second parameter, i.e., the multi-channel autoregressive coefficient in the speech signal interaction process, is used to adjust the reverberation of the observation signal of the target microphone in the target frequency band. Generally, the second parameter can be used to perform reverberation suppression processing on the signal.
[0047] In an optional embodiment, the signal model is an autoregressive model. In step S21, signal conversion processing is performed on the first signal to obtain a second signal, including the following method steps:
[0048] Step S211, performing Fourier transform processing on the first signal to obtain a sixth signal;
[0049] Step S212: Use an autoregressive model to perform approximate processing on the sixth signal to obtain a second signal.
[0050] The signal model used to describe the second signal may be an autoregressive model. Fourier transform processing is performed on the first signal to obtain the sixth signal. This sixth signal is a frequency domain signal and may correspond to multiple frequency bands. Approximate processing of the frequency domain signal using the autoregressive model yields the second signal. In other words, the second signal is represented in the form of an autoregressive model.
[0051] The signal processing method provided by the embodiments of the present invention can be applied, but is not limited to, to scenarios involving far-field voice interaction in the field of artificial intelligence, such as smart speakers. The following uses the voice interaction scenario of a smart speaker as an example to specifically illustrate the technical solution of the present invention.
[0052] The signal model used in processing the far-field voice signal of the smart speaker can be a convolution frequency domain model. Using this convolution frequency domain model, the observed time domain signal (equivalent to the first signal) is Fourier transformed and approximated to obtain the corresponding frequency domain signal (equivalent to the second signal). This frequency domain signal corresponds to multiple frequency bands. For example, a voice signal with a sampling rate of 16k (i.e., a time domain signal) undergoes a 320-point Fourier transform, and the resulting frequency domain signal corresponds to 161 valid frequency bands.
[0053] Specifically, the specific implementation process of performing Fourier transform processing on the first signal to obtain the sixth signal can be shown as the following formula (1):
[0054]
[0055] In the above formula (1), m represents the microphone number, t represents the time number, and A m,l represents the transfer function from the sound source to the microphone, B m,l represents the transfer function from the speaker to the microphone, S represents the observation signal captured by the sensor array (equivalent to the first signal mentioned above), X represents the system playback signal, V m represents noise, represents the frequency domain signal obtained by transformation (equivalent to the sixth signal mentioned above).
[0056] Specifically, the above-mentioned autoregressive model is used to analyze the sixth signal Perform approximate processing to obtain the second signal Y m The specific implementation process can be shown as the following formula (2):
[0057]
[0058] In the above formula (2), m represents the microphone number, M represents the total number of microphones, t represents the time sequence number, Y represents the observation signal, S represents the target speech, and C m,n,l represents the multi-channel autoregressive coefficient, B represents the echo path parameter, X represents the system playback signal, represents noise, L Y Represents the order of the autoregressive model, L X represents the order of echo path modeling, and Δ represents the dividing point between early reflections and late reverberation.
[0059] In an optional embodiment, in step S22, obtaining the first parameter, the second parameter, the third signal, and the fourth signal based on the second signal includes the following method steps:
[0060] Step S221 : Based on the second signal, obtain the first parameter, the second parameter, the third signal, and the fourth signal from the approximate expression corresponding to the autoregressive model.
[0061] Still taking the processing of the far-field voice signal of the smart speaker as an example, the approximate expression of the form of the autoregressive model corresponding to the frequency domain signal (equivalent to the above formula (2)) can determine the above echo path parameter B (equivalent to the above first parameter), the multi-channel autoregressive coefficient C (equivalent to the above second parameter), the system playback signal X (equivalent to the above third signal) and the observation signal Y (equivalent to the above fourth signal).
[0062] In an optional embodiment, in step S23, determining the fifth signal using the second signal, the first parameter, the second parameter, the third signal, and the fourth signal includes the following method steps:
[0063] Step S231, determining a first vector using a first parameter and a second parameter, wherein the first vector is determined by an echo cancellation correlation filter of a first length corresponding to the first parameter and a reverberation suppression correlation filter of a second length corresponding to the second parameter;
[0064] Step S232, determining a second vector using the third signal and the fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal in the target frequency band;
[0065] Step S233: Determine a fifth signal based on the second signal, the first vector, and the second vector.
[0066] The echo cancellation associated filter of the first length corresponding to the first parameter may be a filter of the order of the echo path modeling. The reverberation suppression associated filter of the second length corresponding to the second parameter may be a filter of the order of the autoregressive model. The first vector may be determined using the echo path parameter (equivalent to the first parameter) and the multi-channel autoregressive coefficient (equivalent to the second parameter). The first vector is used to determine the target speech signal to be recovered.
[0067] A specific implementation method for determining the second vector using the playback signal (equivalent to the third signal) and the observation signal of the target microphone in the target frequency band (equivalent to the fourth signal) may be to concatenate the playback signal and the observation signal in the target frequency band to obtain a second vector corresponding to the current target frequency band. This second vector is used to determine the target speech signal to be recovered.
[0068] Based on the reverberant speech signal described by the autoregressive model (equivalent to the second signal), the target speech signal to be restored (equivalent to the fifth signal) can be determined using the first vector and the second vector.
[0069] Still taking the processing of the far-field voice signal of the smart speaker as an example, the specific implementation method of determining the first vector w using the above-mentioned echo path parameter B (equivalent to the above-mentioned first parameter) and the above-mentioned multi-channel autoregressive coefficient C (equivalent to the above-mentioned second parameter) can be shown as the following formula (3):
[0070]
[0071] In the above formula (3), M represents the total number of microphones, C represents the multi-channel autoregressive coefficient, B represents the echo path parameter, and L ... multi-channel autoregressive coefficient, B represents the echo path parameter, and L represents the multi-channel autoregressive coefficient, B X represents the echo path modeling order (equivalent to the first length mentioned above), L Y Represents the order of the autoregressive model (equivalent to the second length mentioned above).
[0072] Still taking the processing of the far-field voice signal of the smart speaker as an example, the specific implementation method of determining the second vector z using the above-mentioned system playback signal X (equivalent to the above-mentioned third signal) and the above-mentioned observation signal Y (equivalent to the above-mentioned fourth signal) can be shown as the following formula (4):
[0073]
[0074] In the above formula (4), M represents the total number of microphones, t represents the time sequence number, Y represents the observation signal, X represents the system playback signal, and L Y Represents the order of the autoregressive model, L X represents the order of echo path modeling, and Δ represents the dividing point between early reflections and late reverberation.
[0075] In an optional embodiment, in step S233, determining the fifth signal based on the second signal, the first vector, and the second vector includes the following method steps:
[0076] Step S2331, performing transposition processing on the first vector to obtain a first calculation result;
[0077] Step S2332, obtaining a second calculation result by using the first calculation result and the second vector;
[0078] Step S2333: Determine a fifth signal based on the second signal and the second calculation result.
[0079] Still taking the processing of the far-field voice signal of the smart speaker as an example, the Hermitian transpose of the first vector w is denoted as w H (equivalent to the first calculation result above), w H The second calculation result is obtained by multiplying the second vector z. Based on the observation signal Y and the second calculation result, the specific implementation method of using the Kalman filter method to determine the target speech signal S to be restored is shown in the following formula (5):
[0080]
[0081] It is easy to understand that, in this example, the Kalman filter is used to optimize and estimate the echo cancellation coefficients and reverberation suppression coefficients involved in the signal processing process, thereby achieving higher echo cancellation performance and reverberation suppression performance.
[0082] In an optional embodiment, the signal processing method further includes the following method steps:
[0083] Step S241, performing variable estimation processing on the first vector to obtain an estimation result;
[0084] In step S242 , the expectation between the first vector and the estimation result is minimized using a square error loss function to obtain an optimal estimation value corresponding to the first vector, wherein the optimal estimation value is used to jointly solve the first parameter and the second parameter.
[0085] Still taking the processing of the far-field voice signal of the smart speaker as an example, the first vector w is subjected to variable estimation processing, and the obtained estimation result is recorded as The above uses the square error loss function (or optimization cost function) to estimate the first vector w and the result The specific implementation process of minimizing the expectation between can be shown as the following formula (6):
[0086]
[0087] In the above formula (6), m represents the microphone number, M represents the total number of microphones, and J w Represents the optimal estimate corresponding to the first vector w. Using this optimal estimate J w , parameters B and C can be solved jointly.
[0088] It is easy to understand that the signal processing method provided by the present invention uses a unified cost function and Kalman filtering method to solve the echo cancellation parameters and reverberation suppression parameters, which can avoid the situation where a local optimal solution is obtained due to solving the echo cancellation parameters and reverberation suppression parameters separately.
[0089] Figure 3 is a schematic diagram of an optional semantic signal processing scenario according to an embodiment of the present invention, such as Figure 3 As shown, according to the signal processing method provided by an embodiment of the present invention, when processing the far-field voice signal of the smart speaker, the microphone array captures signals such as stationary noise, scattered noise, and room reverberation in the interactive environment. The enhanced system then performs echo cancellation and reverberation suppression on the collected signals to recover the target voice signal. Furthermore, the system playback signal is provided to the user through the speakers in the interactive environment.
[0090] It is easy to understand that an embodiment of the present invention provides a joint signal processing method for acoustic echo cancellation (AEC) and speech reverberation suppression (DR) in the Fourier transform domain. An autoregressive (AR) model is used to describe the reverberant microphone signal, and a first-order Markov process is used for signal processing modeling. During the signal processing process, the AR coefficient and the acoustic transfer function from the speaker to the microphone are time-varying. In addition, a Kalman filter is used to optimize and estimate the echo cancellation coefficients and reverberation suppression coefficients involved in the above signal processing process, which can achieve higher echo cancellation performance and reverberation suppression performance.
[0091] The signal processing method provided by this invention focuses on the combined optimization of echo cancellation and reverberation suppression based on the Kalman filter method. Compared with the cascaded echo cancellation and reverberation suppression methods used in related technologies, the signal processing method provided by this invention achieves superior performance and is more suitable for application in artificial intelligence (AI) Internet of Things (IoT) scenarios.
[0092] It's easy to note that in related technologies, echo cancellation and reverberation suppression are typically studied as two separate problems. Using Kalman filtering to address these two issues separately is one approach employed by those skilled in the art. Furthermore, some researchers have studied methods for combining echo cancellation and reverberation suppression from the perspective of sound source separation, but have not employed Kalman filtering. However, the signal processing method provided by the present invention incorporates a technical approach to speech signal processing that combines and optimizes echo cancellation and reverberation suppression.
[0093] Furthermore, related technologies have studied acoustic echo cancellation and speech reverberation suppression from the perspective of semi-blind source separation, resulting in a signal processing solution based on recursive least squares. Compared to related art signal processing methods that use blind signal separation or recursive least squares, the present invention utilizes a Kalman filter for echo cancellation and reverberation suppression, achieving higher echo cancellation and improved signal quality.
[0094] One embodiment of the present invention further provides a signal processing method, which is run on a cloud server. Figure 4 is a flow chart of another signal processing method according to an embodiment of the present invention. Figure 4 As shown, the signal processing method includes:
[0095] Step S41, receiving a first signal from a client, wherein the first signal is an observation signal captured by a sensor array;
[0096] Step S42: performing signal conversion processing on the first signal to obtain a second signal, obtaining a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, and determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal, and the fourth signal, wherein the second signal is a reverberated speech signal described by a signal model, the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band, and the fifth signal is a target speech signal to be recovered;
[0097] Step S43: Feedback the fifth signal to the client.
[0098] Optionally, Figure 5 is a schematic diagram of signal processing on a cloud server according to an embodiment of the present invention. Figure 5 As shown, the client uploads a first signal to a cloud server, where the first signal is an observation signal captured by a sensor array; the cloud server performs signal conversion processing on the first signal to obtain a second signal, obtains a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, and uses the second signal, the first parameter, the second parameter, the third signal, and the fourth signal to determine a fifth signal, where the second signal is a reverberated speech signal described by a signal model, the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band, and the fifth signal is the target speech signal to be recovered. The cloud server then feeds back the target speech signal to be recovered to the client, and the final target speech signal to be recovered is provided to the user through the client's graphical user interface.
[0099] It should be noted that the above-mentioned signal processing method provided in the embodiment of the present invention can be applicable to, but not limited to, far-field voice interaction scenarios in large-screen full-duplex interactive voice systems related to artificial intelligence (such as artificial intelligence Internet of Things systems, smart speakers, smart voice assistants, etc.), through the interaction between the SaaS server and the client, and adopts the method of performing signal conversion processing on the first signal to obtain the second signal, obtaining the first parameter, the second parameter, the third signal and the fourth signal based on the second signal, and using the second signal, the first parameter, the second parameter, the third signal and the fourth signal to determine the fifth signal to obtain the target voice signal to be recovered, and provide the returned target voice signal to the user through the client.
[0100] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0101] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0102] Example 2
[0103] According to an embodiment of the present invention, a device for implementing the above signal processing method is also provided. Figure 6 is a structural diagram of a signal processing device according to an embodiment of the present invention. Figure 6 As shown, the device includes: a processing module 61, an acquisition module 62 and a determination module 63, wherein:
[0104] The processing module 61 is used to perform signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by the sensor array, and the second signal is a reverberation speech signal described by a signal model; the acquisition module 62 is used to obtain a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band; the determination module 63 is used to determine a fifth signal using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, wherein the fifth signal is the target speech signal to be restored.
[0105] Optionally, in the above signal processing device, the first parameter is used to adjust the echo generated by the third signal, and the second parameter is used to adjust the reverberation of the fourth signal.
[0106] Optionally, the signal model is an autoregressive model, and the processing module 61 is further configured to: perform Fourier transform processing on the first signal to obtain a sixth signal; and perform approximate processing on the sixth signal using the autoregressive model to obtain a second signal.
[0107] Optionally, the acquisition module 62 is further configured to: acquire the first parameter, the second parameter, the third signal, and the fourth signal from an approximate expression corresponding to the autoregressive model based on the second signal.
[0108] Optionally, the above-mentioned determination module 63 is also used to: determine a first vector using a first parameter and a second parameter, wherein the first vector is determined by an echo cancellation associated filter of a first length corresponding to the first parameter and a reverberation suppression associated filter of a second length corresponding to the second parameter; determine a second vector using a third signal and a fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal on the target frequency band; and determine a fifth signal based on the second signal, the first vector and the second vector.
[0109] Optionally, the determination module 63 is further configured to: perform transposition processing on the first vector to obtain a first calculation result; obtain a second calculation result by using the first calculation result and the second vector; and determine a fifth signal based on the second signal and the second calculation result.
[0110] Optionally, Figure 7 is a structural diagram of another signal processing device according to an embodiment of the present invention. Figure 7 As shown, the device includes Figure 6 In addition to all the modules shown, it also includes: an estimation module 64, which is used to perform variable estimation processing on the first vector to obtain an estimation result; and uses a square error loss function to minimize the expectation between the first vector and the estimation result to obtain an optimal estimation value corresponding to the first vector, wherein the optimal estimation value is used to jointly solve the first parameter and the second parameter.
[0111] It should be noted that the processing module 61, acquisition module 62, and determination module 63 described above correspond to steps S21 to S23 in Example 1. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0112] In an embodiment of the present invention, a second signal is obtained by performing signal conversion processing on a first signal, wherein the first signal is an observation signal captured by a sensor array, the second signal is a reverberation speech signal described by a signal model, and a first parameter, a second parameter, a third signal and a fourth signal are obtained based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band. The second signal, the first parameter, the second parameter, the third signal and the fourth signal are further used to determine a fifth signal, wherein the fifth signal is a target speech signal to be recovered. The purpose of processing the observation signal to obtain a high-quality target speech signal is achieved, thereby achieving the technical effect of improving the echo cancellation amount and reverberation suppression effect of the obtained target speech signal, and further solving the technical problem in the related art that the quality of the target speech signal is difficult to meet the requirements of the voice interaction scenario due to independent echo cancellation processing and reverberation suppression processing of the signals.
[0113] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.
[0114] Example 3
[0115] According to an embodiment of the present invention, an embodiment of an electronic device is also provided. The electronic device can be any computing device in a computing device group. The electronic device includes: a processor and a memory, wherein:
[0116] A memory is connected to the above-mentioned processor and is used to provide the above-mentioned processor with instructions for processing the following processing steps: performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by the sensor array, and the second signal is a reverberation speech signal described by a signal model; obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band; and using the second signal, the first parameter, the second parameter, the third signal and the fourth signal to determine a fifth signal, wherein the fifth signal is the target speech signal to be recovered.
[0117] In an embodiment of the present invention, a second signal is obtained by performing signal conversion processing on a first signal, wherein the first signal is an observation signal captured by a sensor array, the second signal is a reverberation speech signal described by a signal model, and a first parameter, a second parameter, a third signal and a fourth signal are obtained based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band. The second signal, the first parameter, the second parameter, the third signal and the fourth signal are further used to determine a fifth signal, wherein the fifth signal is a target speech signal to be recovered. The purpose of processing the observation signal to obtain a high-quality target speech signal is achieved, thereby achieving the technical effect of improving the echo cancellation amount and reverberation suppression effect of the obtained target speech signal, and further solving the technical problem in the related art that the quality of the target speech signal is difficult to meet the requirements of the voice interaction scenario due to independent echo cancellation processing and reverberation suppression processing of the signals.
[0118] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.
[0119] Example 4
[0120] The embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0121] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0122] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the signal processing method: performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by the sensor array, and the second signal is a reverberation speech signal described by a signal model; obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band; using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, determining a fifth signal, wherein the fifth signal is the target speech signal to be restored.
[0123] Optionally, Figure 8 is a structural block diagram of another computer terminal according to an embodiment of the present invention. Figure 8As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 122 , a memory 124 , and a peripheral interface 126 .
[0124] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the signal processing method and device in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned signal processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0125] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by the sensor array, and the second signal is a reverberation speech signal described by a signal model; obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band; using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, determining a fifth signal, wherein the fifth signal is the target speech signal to be restored.
[0126] Optionally, the processor may further execute program code of the following steps: the first parameter is used to adjust the echo generated by the third signal, and the second parameter is used to adjust the reverberation of the fourth signal.
[0127] Optionally, the processor may further execute program codes of the following steps: performing Fourier transform processing on the first signal to obtain a sixth signal; and performing approximate processing on the sixth signal using an autoregressive model to obtain a second signal.
[0128] Optionally, the processor may further execute program code of the following steps: based on the second signal, obtaining the first parameter, the second parameter, the third signal and the fourth signal from an approximate expression corresponding to the autoregressive model.
[0129] Optionally, the processor may also execute the program code of the following steps: determining a first vector using a first parameter and a second parameter, wherein the first vector is determined by an echo cancellation associated filter of a first length corresponding to the first parameter and a reverberation suppression associated filter of a second length corresponding to the second parameter; determining a second vector using a third signal and a fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal on the target frequency band; and determining a fifth signal based on the second signal, the first vector, and the second vector.
[0130] Optionally, the processor may further execute program code of the following steps: transpose the first vector to obtain a first calculation result; obtain a second calculation result using the first calculation result and the second vector; and determine a fifth signal based on the second signal and the second calculation result.
[0131] Optionally, the processor may also execute the program code of the following steps: performing variable estimation processing on the first vector to obtain an estimation result; minimizing the expectation between the first vector and the estimation result using the square error loss function to obtain the optimal estimation value corresponding to the first vector, wherein the optimal estimation value is used to jointly solve the first parameter and the second parameter.
[0132] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receiving a first signal from the client, wherein the first signal is an observation signal captured by the sensor array; performing signal conversion processing on the first signal to obtain a second signal, obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, and determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, wherein the second signal is a reverberated speech signal described by a signal model, the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band, and the fifth signal is a target speech signal to be restored; and feeding the fifth signal back to the client.
[0133] In an embodiment of the present invention, a second signal is obtained by performing signal conversion processing on a first signal, wherein the first signal is an observation signal captured by a sensor array, the second signal is a reverberation speech signal described by a signal model, and a first parameter, a second parameter, a third signal and a fourth signal are obtained based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band. The second signal, the first parameter, the second parameter, the third signal and the fourth signal are further used to determine a fifth signal, wherein the fifth signal is a target speech signal to be recovered. The purpose of processing the observation signal to obtain a high-quality target speech signal is achieved, thereby achieving the technical effect of improving the echo cancellation amount and reverberation suppression effect of the obtained target speech signal, and further solving the technical problem in the related art that the quality of the target speech signal is difficult to meet the requirements of the voice interaction scenario due to independent echo cancellation processing and reverberation suppression processing of the signals.
[0134] It can be understood by those skilled in the art that Figure 8 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 8 The structure of the electronic device is not limited. For example, the computer terminal may also include Figure 8 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 8 Different configurations shown.
[0135] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0136] According to an embodiment of the present invention, an embodiment of a storage medium is further provided. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the signal processing method provided in the first embodiment.
[0137] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0138] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by the sensor array, and the second signal is a reverberation speech signal described by a signal model; obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band; and determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, wherein the fifth signal is the target speech signal to be restored.
[0139] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: the first parameter is used to adjust the echo generated by the third signal, and the second parameter is used to adjust the reverberation of the fourth signal.
[0140] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: performing Fourier transform processing on the first signal to obtain a sixth signal; and performing approximation processing on the sixth signal using an autoregressive model to obtain a second signal.
[0141] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: based on the second signal, obtaining the first parameter, the second parameter, the third signal and the fourth signal from an approximate expression corresponding to the autoregressive model.
[0142] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: determining a first vector using a first parameter and a second parameter, wherein the first vector is determined by an echo cancellation associated filter of a first length corresponding to the first parameter and a reverberation suppression associated filter of a second length corresponding to the second parameter; determining a second vector using a third signal and a fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal on the target frequency band; and determining a fifth signal based on the second signal, the first vector, and the second vector.
[0143] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: transposing the first vector to obtain a first calculation result; using the first calculation result and the second vector to obtain a second calculation result; and determining a fifth signal based on the second signal and the second calculation result.
[0144] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: performing variable estimation processing on the first vector to obtain an estimation result; minimizing the expectation between the first vector and the estimation result using a square error loss function to obtain an optimal estimation value corresponding to the first vector, wherein the optimal estimation value is used to jointly solve the first parameter and the second parameter.
[0145] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving a first signal from the client, wherein the first signal is an observation signal captured by the sensor array; performing signal conversion processing on the first signal to obtain a second signal, obtaining a first parameter, a second parameter, a third signal and a fourth signal based on the second signal, and determining a fifth signal using the second signal, the first parameter, the second parameter, the third signal and the fourth signal, wherein the second signal is a reverberation speech signal described by a signal model, the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, the fourth signal is an observation signal of a target microphone selected from the sensor array on a target frequency band, and the fifth signal is a target speech signal to be restored; and feeding the fifth signal back to the client.
[0146] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0147] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0148] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0149] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0150] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0152] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A signal processing method, characterized in that: include: Performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by a sensor array, and the second signal is a reverberated speech signal described by a signal model; acquiring a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array at a target frequency band; Determine a first vector using the first parameter and the second parameter, wherein the first vector is determined by an echo cancellation correlation filter of a first length corresponding to the first parameter and a reverberation suppression correlation filter of a second length corresponding to the second parameter; determining a second vector using the third signal and the fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal in the target frequency band; A fifth signal is determined based on the second signal, the first vector, and the second vector, wherein the fifth signal is a target speech signal to be restored.
2. The signal processing method according to claim 1, wherein: The first parameter is used to adjust the echo generated by the third signal, and the second parameter is used to adjust the reverberation of the fourth signal.
3. The signal processing method according to claim 1, wherein: The signal model is an autoregressive model, and performing signal conversion processing on the first signal to obtain the second signal includes: Performing Fourier transform processing on the first signal to obtain a sixth signal; The autoregressive model is used to perform approximate processing on the sixth signal to obtain the second signal.
4. The signal processing method according to claim 1, wherein: Acquiring the first parameter, the second parameter, the third signal, and the fourth signal based on the second signal includes: Based on the second signal, the first parameter, the second parameter, the third signal and the fourth signal are obtained from an approximate expression corresponding to an autoregressive model.
5. The signal processing method according to claim 1, wherein: Determining the fifth signal based on the second signal, the first vector, and the second vector includes: performing a transposition process on the first vector to obtain a first calculation result; Obtaining a second calculation result by calculating the first calculation result and the second vector; The fifth signal is determined based on the second signal and the second calculation result. The signal processing method according to claim 1 , wherein: The signal processing method further includes: performing variable estimation processing on the first vector to obtain an estimation result; A square error loss function is used to minimize the expectation between the first vector and the estimation result to obtain an optimal estimate corresponding to the first vector, wherein the optimal estimate is used to jointly solve the first parameter and the second parameter.
7. A signal processing method, characterized in that: include: receiving a first signal from a client, wherein the first signal is an observation signal captured by a sensor array; Performing signal conversion processing on the first signal to obtain a second signal, and obtaining a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, wherein the second signal is a reverberated speech signal described by a signal model, the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band; and determining a first vector using the first parameter and the second parameter, wherein the first vector is determined by an echo cancellation correlation filter of a first length corresponding to the first parameter and a reverberation suppression correlation filter of a second length corresponding to the second parameter; determining a second vector using the third signal and the fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal in the target frequency band; and determining a fifth signal based on the second signal, the first vector, and the second vector, wherein the fifth signal is the target speech signal to be recovered; Feedback the fifth signal to the client.
8. A signal processing device, characterized in that: include: a processing module, configured to perform signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by a sensor array, and the second signal is a reverberated speech signal described by a signal model; an acquisition module, configured to acquire a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array in a target frequency band; A determination module is used to determine a first vector using the first parameter and the second parameter, wherein the first vector is determined by an echo cancellation associated filter of a first length corresponding to the first parameter and a reverberation suppression associated filter of a second length corresponding to the second parameter; determine a second vector using the third signal and the fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal on the target frequency band; and determine a fifth signal based on the second signal, the first vector, and the second vector, wherein the fifth signal is the target speech signal to be recovered.
9. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the signal processing method according to any one of claims 1 to 7.
10. A signal processing system, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Performing signal conversion processing on the first signal to obtain a second signal, wherein the first signal is an observation signal captured by a sensor array, and the second signal is a reverberated speech signal described by a signal model; acquiring a first parameter, a second parameter, a third signal, and a fourth signal based on the second signal, wherein the first parameter is an echo path parameter, the second parameter is a multi-channel autoregressive coefficient, the third signal is a playback signal, and the fourth signal is an observation signal of a target microphone selected from the sensor array at a target frequency band; A first vector is determined using the first parameter and the second parameter, wherein the first vector is determined by an echo cancellation associated filter of a first length corresponding to the first parameter and a reverberation suppression associated filter of a second length corresponding to the second parameter; a second vector is determined using the third signal and the fourth signal, wherein the second vector is obtained by concatenating the playback signal and the observation signal on the target frequency band; a fifth signal is determined based on the second signal, the first vector and the second vector, wherein the fifth signal is the target speech signal to be recovered.
Citation Information
Patent Citations
Control of Acoustic Echo Canceller Adaptive Filter for Speech Enhancement
US20150371658A1
Signal processing method and apparatus
WO2022227625A1