Vehicle-mounted echo cancellation method, device, equipment and storage medium
By combining adaptive filters and neural networks, the echo residue problem in vehicle echo cancellation was solved, achieving efficient echo removal of multi-channel and multi-reference signals and improving voice call quality.
Patent Information
- Application Number
- CN202411175080.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2044-08-26
AI Technical Summary
Existing in-vehicle echo cancellation methods cannot effectively eliminate all echoes, resulting in a decrease in the quality of remote voice calls.
An adaptive filter and neural network approach is adopted. By acquiring the speech signal and reference signal, frame windowing, Fourier transform and adaptive filtering are performed. Then, the neural network is used to generate a mask vector for echo removal. Finally, the echo-removed speech signal is obtained through inverse Fourier transform and signal frame reconstruction.
It effectively eliminates the echo received by the vehicle microphone, improves the quality of voice calls, solves the problem of echo cancellation for multi-channel and multi-reference signals, and reduces computational complexity.
Smart Images

Figure CN119028307B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of echo cancellation technology, and more particularly to vehicle-mounted echo cancellation methods, apparatus, equipment, and storage media. Background Technology
[0002] In automotive scenarios, there are various types of echoes, including music, navigation, and voice announcements. These echoes have multiple paths and are highly variable, requiring excellent tracking performance. Furthermore, real-world vehicles typically use multi-zone, multi-channel signals, necessitating simultaneous echo cancellation processing for multiple signals.
[0003] Current echo cancellation methods cannot eliminate all echoes; the remote end can still receive residual echoes that have not been completely eliminated, thus reducing the quality of voice calls.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this application is to provide a vehicle-mounted echo cancellation method, apparatus, device, and storage medium, which aims to solve the technical problem of large echo signals received by vehicle-mounted microphones.
[0006] To achieve the above objectives, this application proposes a vehicle-mounted echo cancellation method, which includes:
[0007] Acquire speech signals and reference signals;
[0008] The speech signal and the reference signal are processed by an adaptive filter to obtain a residual signal;
[0009] The speech signal, the reference signal, and the residual signal are processed by a neural network to obtain a de-echoed speech signal.
[0010] In one embodiment, after the steps of acquiring the speech signal and the reference signal, the method further includes:
[0011] The speech signal and the reference signal are respectively subjected to frame-segmented windowing processing to obtain the frame-segmented windowed speech signal and the reference signal;
[0012] Fourier transforms are performed on the framed and windowed speech signal and the reference signal respectively to obtain the time-frequency domain speech signal and the time-frequency domain reference signal;
[0013] The step of processing the speech signal and the reference signal using an adaptive filter to obtain the residual signal includes:
[0014] The time-frequency domain speech signal and the time-frequency domain reference signal are processed by an adaptive filter to obtain a time-frequency domain residual signal;
[0015] The step of processing the speech signal, the reference signal, and the time-frequency domain residual signal using a neural network to obtain a de-echoed speech signal includes:
[0016] The time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal are processed by a neural network to obtain a de-echoed speech signal.
[0017] In one embodiment, the step of processing the time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal using a neural network to obtain a de-echoed speech signal includes:
[0018] The time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal are processed by a neural network to obtain a mask vector;
[0019] By processing the mask vector and the time-frequency domain speech signal, a de-echoed speech signal is obtained.
[0020] In one embodiment, the step of processing the time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal using a neural network to obtain a mask vector includes:
[0021] The time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal are concatenated into a vector through the normalization layer of the neural network.
[0022] The vector is processed by the core layer of the neural network to obtain the mask vector.
[0023] In one embodiment, the step of concatenating the time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal into a vector through the normalization layer of the neural network includes:
[0024] The time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal are normalized through the normalization layer of the neural network to obtain normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal.
[0025] The normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal are concatenated into a vector along the frequency axis.
[0026] In one embodiment, the step of processing the mask vector and the time-frequency domain speech signal to obtain the de-echo speech signal includes:
[0027] The mask vector and the time-frequency domain speech signal are processed to obtain the time-frequency domain signal after echo removal.
[0028] Perform an inverse Fourier transform on the de-echoed time-frequency domain signal to obtain the de-echoed time-domain speech signal;
[0029] The de-echoed time-domain speech signal is reconstructed into a signal frame to obtain the de-echoed speech signal.
[0030] In one embodiment, the step of acquiring the reference signal includes:
[0031] Obtain the reference audio played by each vehicle speaker;
[0032] The reference audio is evenly divided and mixed according to a preset ratio to obtain the reference signal.
[0033] Furthermore, to achieve the above objectives, this application also proposes a vehicle-mounted echo cancellation device, the device comprising:
[0034] The signal acquisition module is used to acquire voice signals and reference signals;
[0035] The first processing module is used to process the speech signal and the reference signal through an adaptive filter to obtain a residual signal;
[0036] The second processing module is used to process the speech signal, the reference signal and the residual signal through a neural network to obtain a de-echoed speech signal.
[0037] In addition, to achieve the above objectives, this application also proposes an in-vehicle echo cancellation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the in-vehicle echo cancellation method as described above.
[0038] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the vehicle echo cancellation method described above.
[0039] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the vehicle echo cancellation method described above.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] The system acquires speech signals and reference signals; processes the speech signals and reference signals using an adaptive filter to obtain residual signals; processes the speech signals, reference signals, and residual signals using a neural network to obtain de-echoed speech signals; and performs echo cancellation processing in two levels based on the combination of adaptive filtering and neural networks, efficiently processing multi-channel, multi-reference signals. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating an embodiment of the vehicle-mounted echo cancellation method of this application.
[0045] Figure 2 This is a diagram of the vehicle echo audio link provided in Embodiment 1 of the vehicle echo cancellation method of this application;
[0046] Figure 3 This is a structural block diagram provided for Embodiment 1 of the vehicle-mounted echo cancellation method of this application;
[0047] Figure 4 This is a flowchart illustrating Embodiment 2 of the vehicle-mounted echo cancellation method of this application;
[0048] Figure 5 This is a flowchart illustrating Embodiment 3 of the vehicle-mounted echo cancellation method of this application;
[0049] Figure 6 This is a schematic diagram of the module structure of the vehicle-mounted echo cancellation device according to an embodiment of this application;
[0050] Figure 7 This is a schematic diagram of the hardware operating environment involved in the vehicle-mounted echo cancellation method in the embodiments of this application.
[0051] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0052] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0053] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0054] This application proposes an echo cancellation algorithm based on a combination of adaptive filtering and neural networks, which can handle multi-channel and multi-reference signals and effectively solves the problems of slow convergence, easy divergence and large computational load.
[0055] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or vehicle-mounted echo cancellation device capable of performing the above functions. The following description uses a vehicle-mounted echo cancellation device as an example to illustrate this embodiment and the subsequent embodiments.
[0056] Based on this, the embodiments of this application provide a vehicle-mounted echo cancellation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle-mounted echo cancellation method of this application.
[0057] In this embodiment, the vehicle-mounted echo cancellation method includes steps S10 to S30:
[0058] Step S10: Acquire the speech signal and the reference signal;
[0059] It should be noted that the voice signal is the sound signal received by the microphone or other in-vehicle voice receiving device. The received voice signal includes noise such as echoes. The reference signal refers to the reference audio signal played by the speakers in different positions in the vehicle (such as front, rear, left, and right).
[0060] like Figure 2 As shown, Figure 2 The image shows the audio link diagram for vehicle-mounted echo, where microphones and speakers are distributed. The echo signal received by microphone i (i = 1, 2, 3...) can be modeled as follows:
[0061]
[0062] n is the number of speakers in the vehicle. h is the reference signal for the loudspeaker j. ji for The signal propagation path from speaker j to microphone i, where * represents the convolution operation. Assume the near-end speech signal of microphone i is... Then the signal collected by microphone i is: Since echoes significantly impact the performance of speech recognition at the backend, we need to perform echo cancellation on the signal captured by the microphone to separate the speech signal. Therefore, this strategy for vehicle-mounted echo cancellation is proposed.
[0063] Further, the step of acquiring the reference signal includes:
[0064] Obtain the reference audio played by each vehicle speaker;
[0065] The reference audio is evenly mixed according to a preset ratio to obtain the reference signal.
[0066] In practice, the reference signal emitted by the speaker at each location in the vehicle is divided proportionally and mixed into a smaller number of reference audio signals (in real vehicles, there are usually 1 or 2 channels, and in some cases, 4 channels).
[0067] It should be noted that the reference signals of multiple speakers are adjusted and distributed according to a preset ratio to ensure that the volume and timbre are balanced in all positions in the car.
[0068] The process of combining audio signals into a smaller number of reference channels involves merging the signals from multiple speakers into fewer signal channels. For example, if there are eight speakers in the car, their audio signals can be merged into two main channels (typically the signals from the front left and front right speakers), or in some cases, four channels (front left, front right, rear left, and rear right) can be used.
[0069] Step S20: The speech signal and the reference signal are processed by an adaptive filter to obtain the residual signal;
[0070] It should be noted that the residual signal refers to the original speech signal minus the signal processed by the adaptive filter. This difference represents the noise or echo that the filter failed to completely eliminate.
[0071] Adaptive filters automatically adjust their parameters based on the reference signal to optimally reduce or eliminate unwanted sound components such as echoes or noise. The adaptive filter continuously monitors the residual signal and adjusts the filter parameters according to its characteristics using the Least Mean Square Error (LMS) algorithm or other adaptive algorithms to further reduce noise or echo components in the residual signal.
[0072] It should be understood that using a linear adaptive filter for the first echo cancellation process yields the speech time-frequency domain signal E, which has a higher signal-to-noise ratio than the speech signal and can provide some guidance for neural network training.
[0073] In the specific implementation, the mixed reference signal and the signal collected by the microphone are input into the adaptive filter module to obtain the linearly filtered residual signal.
[0074] Step S30: The speech signal, reference signal and residual signal are processed by a neural network to obtain the de-echoed speech signal.
[0075] In the specific implementation, the original microphone signal, the mixed reference signal, and the linearly filtered residual signal are input together into the subsequent neural network to obtain the final near-end speech signal after echo cancellation processing.
[0076] like Figure 3 As shown, Figure 3 As shown in the block diagram, the echo-bearing speech signal d and the reference signal x are first processed by frame-by-frame FFT to obtain the echo-bearing speech time-frequency domain signal D and the reference time-frequency domain signal X. Then, they are input into an adaptive filter to obtain the linearly filtered speech time-frequency domain signal E. The echo-bearing speech time-frequency domain signal D, the reference time-frequency domain signal X, and the linearly filtered speech time-frequency domain signal E are input into the LayerNorm layer for normalization processing to obtain their respective normalized signals. Then, they are input into the core layer of the neural network to obtain the time-frequency mask value. The time-frequency mask value and the echo-bearing speech time-frequency domain signal D are processed to obtain the enhanced near-end speech signal s.
[0077] It should be noted that an adaptive filter is a filter that can automatically adjust its parameters according to the characteristics of the input signal in order to improve the quality of the signal.
[0078] Echoing speech time-frequency: This is the original speech signal, which includes noise such as echoes.
[0079] Linear filter-enhanced domain signal D: By enhancing the signal through a linear filter, an improved time-frequency domain signal is obtained.
[0080] Speech time-frequency domain signal E: The speech signal after being processed by a linear filter is located in the time-frequency domain.
[0081] Reference time-frequency domain signal X: The time-frequency domain signal used as a reference.
[0082] LayerNorm is a normalization layer used to adjust the output of layers in a neural network to have a distribution with a mean of 0 and a standard deviation of 1, which helps improve the stability and performance of the network.
[0083] Core layer (RNN + fully connected + sigmoid): The core layer consists of a recurrent neural network (RNN), a fully connected layer, and a sigmoid activation function, which are used to further process the signal.
[0084] Time-frequency mask value: A mask applied in the time-frequency domain to mask or emphasize certain parts of a signal.
[0085] multiply: indicates a multiplication operation, used to apply a mask to a signal.
[0086] ifft & overlapadd: refers to the inverse fast Fourier transform (IFFT) and overlap-add technique, used to convert time-frequency domain signals back to time-domain signals.
[0087] This embodiment provides an in-vehicle echo cancellation method, which acquires a speech signal and a reference signal; processes the speech signal and the reference signal through an adaptive filter to obtain a residual signal; processes the speech signal, the reference signal and the residual signal through a neural network to obtain an echo-free speech signal. Based on the combination of adaptive filtering and neural network, echo cancellation processing is performed in two levels, which efficiently processes multi-channel and multi-reference signals.
[0088] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 After step S10, the vehicle-mounted echo cancellation method further includes steps S101-S102:
[0089] Step S101: Perform frame-segmentation and windowing processing on the speech signal and the reference signal respectively to obtain the frame-segmented and windowed speech signal and the reference signal.
[0090] It should be noted that framing involves dividing a continuous signal into shorter segments, which are called frames. The purpose of framing is to convert a continuous signal into a format suitable for Discrete Fourier Transform (DFT) processing.
[0091] Windowing involves multiplying each frame by a window function, such as a Hanning window, a Hamming window, or a rectangular window. The purpose of the window function is to reduce discontinuities at frame boundaries, thereby reducing spectral leakage.
[0092] Step S102: Perform Fourier transform on the framed and windowed speech signal and the reference signal respectively to obtain the time-frequency domain speech signal and the time-frequency domain reference signal.
[0093] Then, a Fast Fourier Transform (FFT) is performed on each frame processed by the window function. The FFT is an efficient algorithm that can quickly calculate the spectrum of a signal, i.e., the amplitude and phase of the signal at different frequencies.
[0094] In the specific implementation, the echo-bearing speech signal d and the reference signal x are framed and windowed, and Fourier transforms are performed to obtain the time-frequency domain signals D and X, respectively.
[0095] Furthermore, the step of processing the speech signal and the reference signal using an adaptive filter to obtain the residual signal includes:
[0096] The time-frequency domain speech signal and the time-frequency domain reference signal are processed by an adaptive filter to obtain the time-frequency domain residual signal;
[0097] The steps for processing the speech signal, reference signal, and time-frequency domain residual signal using a neural network to obtain the de-echoed speech signal include:
[0098] The time-frequency domain speech signal, time-frequency domain reference signal, and time-frequency domain residual signal are processed by a neural network to obtain a de-echoed speech signal.
[0099] This embodiment provides a vehicle-mounted echo cancellation method. First, the speech signal and the reference signal are processed by frame segmentation. Then, an echo cancellation algorithm combining adaptive filtering and neural networks is used for processing. It can handle multi-channel and multi-reference signals and effectively solves the problems of slow convergence, easy divergence and large computational load.
[0100] Based on the first and second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 Step S30 includes steps S301 to S302:
[0101] Step S301: The time-frequency domain speech signal, time-frequency domain reference signal and time-frequency domain residual signal are processed by a neural network to obtain a mask vector;
[0102] It should be noted that the mask vector can be a boolean array or a binary array, where the elements can be 0 or 1, indicating whether the corresponding data point is selected or masked.
[0103] In the specific implementation, the time-frequency domain speech signal, time-frequency domain reference signal and time-frequency domain residual signal are processed by the normalization layer and core layer of the neural network to obtain a mask vector with the same dimension as the time-frequency domain speech signal.
[0104] Further, step S301 includes:
[0105] The time-frequency domain speech signal, time-frequency domain reference signal, and residual signal are concatenated into a vector through the normalization layer of the neural network.
[0106] The mask vector is obtained by processing the vector through the core layer of the neural network.
[0107] In the specific implementation, the three time-frequency vectors are normalized through the LayerNorm layer, and then the three normalized time-frequency vectors are concatenated along the frequency axis into a single vector. This concatenated vector is then input into the core layer of the network, which includes n layers of RNN (GRU, LSTM), one fully connected layer, and finally a mask vector with the same dimension as D is obtained by non-linear activation of sigmoid.
[0108] Furthermore, the step of concatenating the time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal into a vector through the normalization layer of the neural network includes:
[0109] The time-frequency domain speech signal, time-frequency domain reference signal, and residual signal are normalized by the normalization layer of the neural network to obtain the normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal.
[0110] The normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal are concatenated into a vector along the frequency axis.
[0111] It should be noted that the role of the normalization layer is to adjust the distribution of the signal so that it has a stable mean and standard deviation; after normalization, the feature distribution of each signal is more standardized, which is helpful for subsequent neural network processing.
[0112] The normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal are concatenated along the frequency axis. In the frequency dimension, the values of these three signals are arranged sequentially to form a longer vector.
[0113] In the specific implementation, the three time-frequency vectors are normalized through the LayerNorm layer, and then the normalized vectors are concatenated along the frequency axis into a single vector.
[0114] Step S302: By processing the mask vector and the time-frequency domain speech signal, the echo-de-echoed speech signal is obtained.
[0115] Further, step S302 includes:
[0116] The mask vector and the time-frequency domain speech signal are processed to obtain the de-echoed time-frequency domain signal;
[0117] Perform an inverse Fourier transform on the de-echoed time-frequency domain signal to obtain the de-echoed time-domain speech signal;
[0118] The time-domain speech signal after echo removal is reconstructed into a signal frame to obtain the echo-removed speech signal.
[0119] In practice, the enhanced, de-echoed near-end speech signal can be obtained by multiplying the time-frequency domain speech signal with the mask vector, performing an inverse Fourier transform, and applying an overlap-add transformation.
[0120] It should be noted that performing an inverse Fourier transform (IFFT) on the de-echoed time-frequency domain signal converts the signal back from the frequency domain to the time domain, resulting in the de-echoed time-domain speech signal.
[0121] After performing the inverse Fourier transform, a series of short-time signal frames are obtained. Signal frame recombination refers to recombining these short-time signal frames into a continuous time-domain signal. The overlap-add method is used to ensure signal continuity and reduce boundary effects.
[0122] This embodiment provides a vehicle-mounted echo cancellation method. The original microphone signal, the mixed reference signal, and the linearly filtered residual signal are input together into the subsequent neural network to obtain the final near-end speech signal after echo cancellation processing. At this time, multiple microphone signals can be combined into a batch and sent to the network to save computational overhead. This allows the network to achieve good results without very complex RNN layers, resulting in excellent performance and low computational requirements.
[0123] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the vehicle-mounted echo cancellation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0124] This application also provides a vehicle-mounted echo cancellation device; please refer to... Figure 6 The vehicle-mounted echo cancellation device includes:
[0125] Signal acquisition module 10 is used to acquire speech signals and reference signals;
[0126] The first processing module 20 is used to process the speech signal and the reference signal through an adaptive filter to obtain the residual signal;
[0127] The second processing module 30 is used to process the speech signal, reference signal and residual signal through a neural network to obtain the echo-de-echoed speech signal.
[0128] The vehicle-mounted echo cancellation device provided in this application, employing the vehicle-mounted echo cancellation method in the above embodiments, can solve the technical problem of large voice echo signals received by vehicle-mounted microphones. Compared with the prior art, the beneficial effects of the vehicle-mounted echo cancellation device provided in this application are the same as those of the vehicle-mounted echo cancellation method provided in the above embodiments, and other technical features in the vehicle-mounted echo cancellation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0129] This application provides an in-vehicle echo cancellation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the in-vehicle echo cancellation method in Embodiment 1 above.
[0130] The following is for reference. Figure 7The diagram illustrates a structural schematic suitable for implementing the vehicle-mounted echo cancellation device of the embodiments of this application. The vehicle-mounted echo cancellation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), vehicle-mounted terminals (e.g., vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The vehicle-mounted echo cancellation device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.
[0131] like Figure 7 As shown, the vehicle-mounted echo cancellation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the vehicle-mounted echo cancellation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the vehicle-mounted echo cancellation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show vehicle-mounted echo cancellation devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0132] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0133] The vehicle-mounted echo cancellation device provided in this application, employing the vehicle-mounted echo cancellation method in the above embodiments, can solve the technical problem of large voice echo signals received by vehicle-mounted microphones. Compared with the prior art, the beneficial effects of the vehicle-mounted echo cancellation device provided in this application are the same as those of the vehicle-mounted echo cancellation method provided in the above embodiments, and other technical features of this vehicle-mounted echo cancellation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0134] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0135] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0136] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the vehicle echo cancellation method in the above embodiments.
[0137] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0138] The aforementioned computer-readable storage medium may be included in the vehicle-mounted echo cancellation device; or it may exist independently and not be installed in the vehicle-mounted echo cancellation device.
[0139] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the vehicle-mounted echo cancellation device, cause the vehicle-mounted echo cancellation device to: acquire a speech signal and a reference signal; process the speech signal and the reference signal using an adaptive filter to obtain a residual signal; and process the speech signal, the reference signal, and the residual signal using a neural network to obtain a de-echoed speech signal.
[0140] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0143] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described vehicle echo cancellation method, which can solve the technical problem of large voice echo signals received by vehicle microphones. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the vehicle echo cancellation method provided in the above embodiments, and will not be repeated here.
[0144] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the vehicle echo cancellation method described above.
[0145] The computer program product provided in this application can solve the technical problem of large echo signals received by vehicle microphones. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the vehicle echo cancellation method provided in the above embodiments, and will not be repeated here.
[0146] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A vehicle-mounted echo cancellation method, characterized in that, The vehicle-mounted echo cancellation method includes: Acquire speech signals and reference signals; The speech signal and the reference signal are processed by an adaptive filter to obtain a residual signal; The speech signal, the reference signal, and the residual signal are processed by a neural network to obtain a de-echoed speech signal; After the steps of acquiring the speech signal and the reference signal, the method further includes: The speech signal and the reference signal are respectively subjected to frame-segmented windowing processing to obtain the frame-segmented windowed speech signal and the reference signal; Fourier transforms are performed on the framed and windowed speech signal and the reference signal respectively to obtain the time-frequency domain speech signal and the time-frequency domain reference signal; The step of processing the speech signal and the reference signal using an adaptive filter to obtain the residual signal includes: The time-frequency domain speech signal and the time-frequency domain reference signal are processed by an adaptive filter to obtain a time-frequency domain residual signal; The step of processing the speech signal, the reference signal, and the time-frequency domain residual signal using a neural network to obtain a de-echoed speech signal includes: The time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal are processed by a neural network to obtain a de-echoed speech signal.
2. The vehicle-mounted echo cancellation method as described in claim 1, characterized in that, The step of processing the time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal through a neural network to obtain a de-echoed speech signal includes: The time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal are processed by a neural network to obtain a mask vector; By processing the mask vector and the time-frequency domain speech signal, a de-echoed speech signal is obtained.
3. The vehicle-mounted echo cancellation method as described in claim 2, characterized in that, The step of processing the time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal through a neural network to obtain a mask vector includes: The time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal are concatenated into a vector through the normalization layer of the neural network. The vector is processed by the core layer of the neural network to obtain the mask vector.
4. The vehicle-mounted echo cancellation method as described in claim 3, characterized in that, The step of concatenating the time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal into a vector through the normalization layer of the neural network includes: The time-frequency domain speech signal, the time-frequency domain reference signal, and the residual signal are normalized through the normalization layer of the neural network to obtain normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal. The normalized time-frequency domain speech signal, time-frequency domain reference signal, and residual signal are concatenated into a vector along the frequency axis.
5. The vehicle-mounted echo cancellation method as described in claim 2, characterized in that, The step of processing the mask vector and the time-frequency domain speech signal to obtain the de-echoed speech signal includes: The mask vector and the time-frequency domain speech signal are processed to obtain the time-frequency domain signal after echo removal. Perform an inverse Fourier transform on the de-echoed time-frequency domain signal to obtain the de-echoed time-domain speech signal; The de-echoed time-domain speech signal is reconstructed into a signal frame to obtain the de-echoed speech signal.
6. The vehicle-mounted echo cancellation method as described in claim 1, characterized in that, The step of acquiring the reference signal includes: Obtain the reference audio played by each vehicle speaker; The reference audio is evenly divided and mixed according to a preset ratio to obtain the reference signal.
7. A vehicle-mounted echo cancellation device, characterized in that, The device includes: The signal acquisition module is used to acquire voice signals and reference signals; The signal acquisition module is also used to perform frame-segmentation and windowing processing on the speech signal and the reference signal respectively to obtain the frame-segmented and windowed speech signal and reference signal; Fourier transforms are performed on the framed and windowed speech signal and the reference signal respectively to obtain the time-frequency domain speech signal and the time-frequency domain reference signal; The step of processing the speech signal and the reference signal using an adaptive filter to obtain the residual signal includes: The time-frequency domain speech signal and the time-frequency domain reference signal are processed by an adaptive filter to obtain a time-frequency domain residual signal; The step of processing the speech signal, the reference signal, and the time-frequency domain residual signal using a neural network to obtain a de-echoed speech signal includes: The time-frequency domain speech signal, the time-frequency domain reference signal, and the time-frequency domain residual signal are processed by a neural network to obtain a de-echoed speech signal. The first processing module is used to process the speech signal and the reference signal through an adaptive filter to obtain a residual signal; The second processing module is used to process the speech signal, the reference signal and the residual signal through a neural network to obtain a de-echoed speech signal.
8. A vehicle-mounted echo cancellation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vehicle echo cancellation method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the vehicle echo cancellation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Acoustic echo cancellation method and system based on adaptive filter and neural network
CN113436636A
Echo cancellation method and device, audio equipment and storage medium
CN117727317A