Method, device and electronic device for estimating time delay in echo cancellation
By generating ultrasonic signals and superimposing downlink signals and separating the effective signals for power spectrum analysis, the problems of large amount of delay estimation calculation and poor accuracy in echo cancellation are solved, and efficient and accurate delay estimation is achieved.
Patent Information
- Application Number
- CN202111070181.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-13
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-09-13
AI Technical Summary
The prior art has large amount of time delay estimation calculations and poor accuracy in echo cancellation, and is greatly affected by near-end speech and noise.
Generate ultrasonic signals and superimpose downlink signals to form reference signals, collect mixed signals and separate valid signals containing ultrasonic frequency bands, and determine the delay value through power spectrum analysis.
It realizes efficient and accurate delay estimation, has small calculation volume and strong robustness, and is adapted to noise interference in different ambients.
Smart Images

Figure CN113870889B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of Internet technology, and in particular to a method, device, and electronic device for delay estimation in echo cancellation. Background Art
[0002] With the development of the internet, video and voice calls have become widely used. In two-way voice conversations, the mixed signal picked up by the local microphone includes both the local user's near-end voice signal and the far-end voice signal played by the playback device. This inevitably generates an echo signal for the far-end user, necessitating echo cancellation. The most critical step in echo cancellation is delay estimation, which facilitates subsequent linear and nonlinear echo cancellation. Current delay estimation typically requires comparing the full mixed signal with the full downlink signal, resulting in high computational complexity. Furthermore, near-end voice and noise significantly impact delay estimation, leading to poor accuracy.
[0003] Based on this, a more efficient delay estimation solution in echo cancellation is needed. Summary of the Invention
[0004] One or more embodiments of this specification provide a method, apparatus, electronic device, and storage medium for delay estimation in echo cancellation, to solve the following technical problem: a more effective delay estimation solution in echo cancellation is needed.
[0005] To solve the above technical problems, in a first aspect, an embodiment of this specification provides a method for delay estimation in echo cancellation, including:
[0006] generating an ultrasonic signal, and superimposing the ultrasonic signal and the downlink signal to generate a reference signal;
[0007] collecting a mixed signal, wherein the mixed signal includes an echo signal generated by playing the reference signal;
[0008] Separating a first effective signal from the mixed signal, and separating a second effective signal from the reference signal, wherein frequency bands of the first effective signal and the second effective signal include a frequency band of the ultrasonic signal;
[0009] For any target frame signal in the first valid signal, determine a first power spectrum of the target frame signal, and determine a plurality of second power spectra corresponding to the plurality of frame signals included in the second valid signal;
[0010] A time delay value of the target frame signal relative to the reference signal is determined according to a distance between the first power spectrum and the second power spectrum.
[0011] In a second aspect, an embodiment of this specification provides a delay estimation device for echo cancellation, including:
[0012] A signal generating module, generating an ultrasonic signal, and superimposing the ultrasonic signal and the downlink signal to generate a reference signal;
[0013] a signal acquisition module for acquiring a mixed signal, wherein the mixed signal includes an echo signal generated by playing the reference signal;
[0014] a signal separation module, separating a first valid signal from the mixed signal, and separating a second valid signal from the reference signal, wherein the frequency bands of the first valid signal and the second valid signal include the frequency band of the ultrasonic signal;
[0015] a power spectrum determination module, for determining, for any target frame signal in the first valid signal, a first power spectrum of the target frame signal, and determining a plurality of second power spectra corresponding to the plurality of frame signals included in the second valid signal;
[0016] A delay estimation module determines a delay value of the target frame signal relative to the reference signal according to a distance between the first power spectrum and the second power spectrum.
[0017] In a third aspect, an embodiment of this specification provides an electronic device, including:
[0018] at least one processor; and,
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect.
[0021] In a fourth aspect, an embodiment of this specification provides a non-volatile computer storage medium storing computer-executable instructions. When a computer reads the computer instructions in the storage medium, the instructions enable one or more processors to execute the method described in the first aspect.
[0022] At least one of the above technical solutions adopted in one or more embodiments of this specification can achieve the following beneficial effects: generating an ultrasonic signal, superimposing the ultrasonic signal and the downlink signal to generate a reference signal; collecting a mixed signal; separating a first valid signal from the mixed signal, and separating a second valid signal from the reference signal, wherein the frequency bands of the first valid signal and the second valid signal include the frequency band of the ultrasonic signal; for any target frame signal in the first valid signal, determining the first power spectrum of the target frame signal, and determining multiple second power spectra corresponding to the multiple frame signals included in the second valid signal; determining the time delay value of the target frame signal relative to the reference signal based on the distance between the first power spectrum and the second power spectrum. Therefore, it is only necessary to perform frequency spectrum distance analysis based on the first valid signal and the second valid signal in the frequency band containing the ultrasonic signal, that is, accurate estimation of the time delay value can be achieved, with low computational complexity, higher efficiency, and strong robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0024] Figure 1 A schematic diagram of the system architecture provided in the embodiments of this specification;
[0025] Figure 2 A flowchart of a method for delay estimation in echo cancellation provided in an embodiment of this specification;
[0026] Figure 3 A schematic diagram of a power spectrum involved in the embodiments of this specification;
[0027] Figure 4 A schematic diagram of the logical structure of a delay value estimation provided in an embodiment of this specification;
[0028] Figure 5 A schematic diagram of the structure of a delay estimation device in echo cancellation provided in an embodiment of this specification;
[0029] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0030] The embodiments of this specification provide a method, apparatus, device, and storage medium for delay estimation in echo cancellation.
[0031] In order to help those skilled in the art better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0032] In two-way voice communication scenarios (e.g., mobile phone calls, third-party app audio and video calls, multi-party audio and video calls using conferencing equipment, etc.), the far-end voice is processed into a downlink signal, which is then sent to a playback device (e.g., a speaker or receiver) for playback. The local microphone picks up the local mixed signal (including near-end voice, noise, and the echo generated by the downlink signal) and performs uplink processing (including echo cancellation, noise reduction, volume adjustment, etc.) to form an uplink signal, which is then sent to the far-end. Echo cancellation is a crucial step in this process, and echo cancellation requires a quick and accurate estimation of the delay of the echo relative to the downlink signal.
[0033] Based on this, the embodiment of the present application provides a more efficient delay estimation solution. Figure 1 As shown, Figure 1 This is a schematic diagram of the system architecture provided in the embodiments of this specification. This system architecture includes an ultrasonic signal generator that generates an ultrasonic signal, which is then superimposed on the downlink signal to generate a reference signal. Subsequently, a portion of the valid signal containing the ultrasonic signal is extracted from the full signal for processing, thereby achieving more efficient delay estimation.
[0034] like Figure 2 As shown, Figure 2 A flowchart of a method for delay estimation in echo cancellation provided in an embodiment of this specification includes:
[0035] S201: Generate an ultrasonic signal, and superimpose the ultrasonic signal and a downlink signal to generate a reference signal.
[0036] When an ultrasonic signal is played through a speaker, it generates high-frequency ultrasonic waves that cannot be directly perceived by the human ear. High-frequency ultrasonic waves refer to sound waves with a frequency exceeding 20 Hz. In the embodiments of this specification, the frequency range of the ultrasonic signal generated can be preset to a smaller range for greater convenience during subsequent processing. For example, the frequency band (i.e., frequency range) of the ultrasonic signal generated during playback can be preset to [20 kHz, 22 kHz].
[0037] Furthermore, the ultrasonic signal and the downlink signal can be directly superimposed to generate a reference signal. The generated reference signal is played through the speaker, thereby generating corresponding far-end voice and ultrasonic waves at the near end; at the same time, the generated reference signal will also enter the acoustic echo cancellation (AEC) device for comparison with the echo signal, such as Figure 1 As shown in .
[0038] Because ultrasonic waves are imperceptible to the human ear, the generation of ultrasonic waves during the playback of the reference signal does not affect near-end listening or actual communication. Due to the high frequency of ultrasonic waves, even in a reverberant room, when they reflect off walls or other obstacles and are picked up by the microphone, the resulting reverberation echo is minimal, facilitating subsequent delay estimation.
[0039] S203: Collect a mixed signal, where the mixed signal includes an echo signal generated by playing the reference signal.
[0040] The collected mixed signal is a signal stream consisting of multiple frames. This mixed signal includes the near-end speech signal, noise, and echo signal. The echo signal is the reference signal generated by the near-end playback device and then captured by the near-end microphone. While the echo signal may have energy variations compared to the reference signal (due to attenuation by the environment or amplification by the device), it still corresponds to the reference signal in terms of frequency, power distribution, and phase.
[0041] S205 , separating a first valid signal from the mixed signal, and separating a second valid signal from the reference signal, wherein frequency bands of the first valid signal and the second valid signal include a frequency band of the ultrasonic signal.
[0042] Since the frequency of ultrasonic signals is relatively high, it is no longer necessary to perform Fast Fourier Transform (FFT) on the entire collected mixed signal during the processing process. Instead, only the portion of the mixed signal containing the frequency band of the ultrasonic signal is required. Therefore, separation filtering can be used to directly separate a first valid signal containing the frequency band of the ultrasonic signal from the collected mixed signal, and to separate a second valid signal containing the frequency band of the ultrasonic signal from the reference signal. Generally speaking, the frequency bands of the first and second valid signals can also be ultrasonic frequency bands that are inaudible to the human ear.
[0043] For example, assuming the frequency range of the generated ultrasonic wave is [20 kHz, 22 kHz], and the sampling frequency range of the sampler is 0 to 48 kHz, then the valid signals between 18 kHz and 24 kHz can be directly filtered out, and the sampling rate can be set to 12 kHz. This allows the first valid signal, which consists of multiple frames (each containing 12,000 sampling points), to be separated from the mixed signal, and the second valid signal, which consists of multiple frames, to be separated from the reference signal.
[0044] S207 : For any target frame signal in the first valid signal, determine a first power spectrum of the target frame signal, and determine a plurality of second power spectra corresponding to the plurality of frame signals included in the second valid signal.
[0045] Furthermore, FFT can be performed on the separated first effective signal and the second effective signal to obtain the corresponding first complex spectrum signal and the corresponding second complex spectrum signal.
[0046] A power spectrum is calculated based on the first complex spectrum signal and the second complex spectrum signal to obtain a power spectrum corresponding to each frame. That is, for any target frame signal in the first valid signal, a first power spectrum of the target frame signal is determined, and multiple second power spectra corresponding to multiple frame signals included in the second valid signal are determined.
[0047] At the same time, in practical applications, the current frame signal D(n) can be sequentially regarded as the target frame signal according to the frame order in the first valid signal to calculate the first power spectrum, and multiple second power spectra can be pre-cached in the buffer. The current frame in the first valid signal is recorded as the target frame signal D(n). Then, the second power spectra X(nm)...X(n) corresponding to the m frame signals before D(n) (i.e., the nmth frame to the nth frame) can be cached in the buffer.
[0048] Since each frame actually contains multiple sampling points, taking a sampling rate of 12kHz as an example (i.e. 12,000 samples per second), if the frame shift between each frame is set to 10ms, then each frame contains 120 sampling points. In the power spectrum, we can know that each sampling point has its own corresponding power value. The power spectrum of each frame represents the distribution of signal power in the frequency domain at the moment corresponding to the frame. Specifically, the power spectrum contains some amplitude information in the spectrum, while the phase information is discarded. Figure 3 Place, Figure 3This is a schematic diagram of the power spectrum involved in the embodiments of this specification. In practical applications, multiple frequency points can be taken from the frequency band of the ultrasonic wave contained in the effective signal. For example, if the frequency band of the effective signal is 18 kHz to 24 kHz and the frequency band of the ultrasonic wave is 20 kHz to 22 kHz, then 128 frequency points can be evenly distributed within the range of 20 kHz to 22 kHz, and the power spectrum can be characterized in the form of a discrete array or vector. For example, for the first power spectrum of the target frame signal, it can be illustrated as (power value 1, power value 2..., power value 128), each power value corresponds to a frequency point, and a vector with a dimension of 128 is obtained. Similarly, a similar method can be used to vectorize any second power spectrum.
[0049] S209: Determine a delay value of the target frame signal relative to the reference signal according to a distance between the first power spectrum and the second power spectrum.
[0050] It should be noted that, in the embodiment of this specification, the first effective signal and the second effective signal obtained by separation are both high-frequency signals, and the human voice has been filtered out. Therefore, under normal circumstances, if the echo attenuation effect is small, for the target frame in the first effective signal and the corresponding frame to be searched in the second effective signal, if there is a corresponding relationship between the two frame signals, then obviously, their corresponding first power spectrum and second power spectrum should be relatively close. Therefore, the delay value of the target frame signal relative to the reference signal can be determined based on the distance between the first power spectrum and the second power spectrum.
[0051] For example, after determining the current frame as the target frame signal and determining the first power spectrum of the target frame signal, a search can be performed from the buffer and a second power spectrum whose distance from the first power spectrum meets preset conditions (which may include a distance lower than a preset threshold, and / or, being ranked at the top of the distance order from small to large, etc.) can be determined as the corresponding frame, so that the delay value of the target frame signal relative to the reference signal can be determined based on the time difference between the target frame signal and the corresponding frame.
[0052] At least one of the above technical solutions adopted in one or more embodiments of this specification can achieve the following beneficial effects: generating an ultrasonic signal, superimposing the ultrasonic signal and the downlink signal to generate a reference signal; collecting a mixed signal; separating a first valid signal from the mixed signal, and separating a second valid signal from the reference signal, wherein the frequency bands of the first valid signal and the second valid signal include the frequency band of the ultrasonic signal; for any target frame signal in the first valid signal, determining the first power spectrum of the target frame signal, and determining multiple second power spectra corresponding to the multiple frame signals included in the second valid signal; determining the time delay value of the target frame signal relative to the reference signal based on the distance between the first power spectrum and the second power spectrum. Therefore, it is only necessary to perform frequency spectrum distance analysis based on the first valid signal and the second valid signal in the frequency band containing the ultrasonic signal, that is, accurate estimation of the time delay value can be achieved, with low computational complexity, higher efficiency, and strong robustness.
[0053] In one embodiment, when generating ultrasound, an ultrasound signal having different frame frequencies within the period can be generated based on a preset window duration as a period (i.e., a variable frequency ultrasound signal). Specifically, the preset window duration can be related to the frame duration setting, for example, the preset window duration can be 100 frames; and, within a period, the frequency of each frame is randomly extracted within a preset range, for example, from 20kHz to 22kHz with an interval of 0.2kHz, generating 100 frequencies from 20kHz, 20.2kHz, 20.4kHz... to 22kHz, and randomly extracting the frequency of the ultrasound signal generated for each frame from these 100 frequencies to ensure that the absolute value of the frequency difference between any two frames of ultrasound signals within a period is not less than 0.2kHz, thereby generating an ultrasound signal with different frame frequencies within the period, highlighting the frequency difference between each frame signal within a period. Since the difference between each frame is large, the distance between the first power spectrum and each second power spectrum will also be large, thereby improving the discrimination of the distance from the second power spectrum, which is conducive to the accurate calculation of the delay value.
[0054] In one embodiment, the delay value can be determined by using the following coarse delay search method, that is, obtaining m second power spectra corresponding to the m frame signals before the target signal frame from the buffer, and calculating the distance between the first power spectrum and each second power spectrum respectively, and then determining the second power spectrum with the smallest distance from the first power spectrum as the target power spectrum; determining the first time point corresponding to the first power spectrum, and determining the second time point corresponding to the target power spectrum; determining the time difference between the first time point and the second time point as the delay value. In this way, since the first time point and the second time point both correspond to a frame signal, the delay value determined is an integer multiple of the length of a frame (for example, the delay value is 10 frames). For example, for a sampling rate of 12kHz, if the frame shift selected during FFT is 120 sampling points (i.e., one frame contains 120 sampling points), then the duration of one frame is 10ms. In this case, the calculated delay value can be accurate to the order of 10ms. Since the frame shift can be selected based on actual needs, the accuracy of the delay value estimated in this way can be adjusted as needed to meet the actual needs of users.
[0055] In one embodiment, a more accurate delay value can be obtained by further employing the following fine delay search method based on the aforementioned coarse delay search. Specifically, the first time domain characteristics of multiple sampling point signals in the target frame signal can be determined; a delay signal of multiple consecutive frames including a frame signal corresponding to a second time point can be obtained to determine the second time domain characteristics of multiple sampling point signals in the delay signal; the first time domain characteristics and the second time domain characteristics can be cross-correlated to generate a cross-correlation result; and the time difference corresponding to the maximum cross-correlation result can be determined as the delay value.
[0056] For example, assuming that the current frame is the 100th frame, when the current frame is used as the target frame signal, the delay value is determined to be 10 frames at the same time. Then, the delay signals of multiple consecutive frames (for example, from the 89th frame to the 91st frame) including the 90th frame signal can be selected, and the second time domain characteristics of the multiple sampling point signals in the delay signal can be determined respectively, so as to perform cross-correlation and obtain a finer-grained delay value based on the cross-correlation result.
[0057] The time domain feature (the first time domain feature or the second time domain feature) characterizes the change in frequency of a signal over time. The cross-correlation result characterizes the relationship between the two signals after a time difference. If the correlation between the two is large, the cross-correlation result is larger, and if the correlation between the two is small, the cross-correlation result is smaller. Therefore, a plurality of different time differences can be selected (for example, an integer multiple of the time interval of the sampling points is used as a time difference sequence), and the first time domain features of the multiple sampling point signals in the target frame signal and the first time domain features of the multiple sampling point signals in the delay signal are cross-correlated based on the selected multiple different time differences, and a plurality of different cross-correlation results are obtained. When a maximum value appears in the cross-correlation, it means that the shapes of the two signals are closest at this time, that is, the time difference at this time is the delay value of the target frame signal relative to the reference signal. In other words, in this way, the delay value with the time interval of the sampling points as the granularity can be obtained from the cross-correlation result.
[0058] For example, when the sampling frequency is 12 kHz, the time interval between two sampling points is 0.083 ms. The delay value obtained at this time can achieve an accuracy of an integer multiple of 0.083 ms, which is beneficial for subsequent filtering and nonlinear processing.
[0059] like Figure 4 As shown, Figure 4 A schematic diagram of the logical structure of a delay value estimation provided by an embodiment of this specification. In this diagram, Delay0 is the delay value given by a coarse delay search, and Delay1 is a more accurate delay value obtained by a fine delay search based on the coarse delay search.
[0060] In one embodiment, the time delay signal can be obtained in the following manner: obtaining a specified number of frame signals before and after the second time point to form the time delay signal; or, taking the second time point as the end point, obtaining a specified number of frame signals before the second time point to form the time delay signal. For example, assuming that the frame number corresponding to the second time point is the 90th frame, then the frame signals of one frame before and after the 90th frame, i.e., the 89th to 91st frames, can be obtained to form the time delay signal; or, taking the 90th frame as the end point, obtaining signals of two consecutive frames before the 90th frame, i.e., the 88th to 90th frames, can be obtained to form the time delay signal. In this way, it can be ensured that the signal corresponding to the current frame always exists in the obtained time delay signal, thereby improving the accuracy of the time delay fine search.
[0061] In one embodiment, since the ultrasonic waves in the reference signal will undergo echo attenuation after propagation, the first power spectrum and the second power spectrum are similar in shape. At this time, the distance between the two can be calculated in the following way: binarize the first power spectrum and the second power spectrum based on a preset power threshold and / or frequency threshold to generate a binarized first power spectrum and a second power spectrum; determine the delay value of the target frame signal relative to the reference signal based on the distance between the binarized first power spectrum and the second power spectrum.
[0062] That is, the first power spectrum is binarized using the first power threshold (a power value greater than the power threshold is set to 1, and a power value not exceeding the power threshold is set to 0), and the second power spectrum is binarized using the second power threshold. The first power threshold and the second power threshold can be set based on experience, so that the first power spectrum and the second power spectrum that have a corresponding relationship after binarization have the same value in the corresponding frequency band, and then the delay value of the target frame signal relative to the reference signal can be determined based on the aforementioned coarse delay search method.
[0063] For example, for a first power spectrum, assuming that its corresponding vector is (power value 1, power value 2..., power value 128), then after binarization based on the first power threshold, its possible values are (0, 1..., 1). In this way, the cache also contains multiple vectors corresponding to the corresponding second power spectrum. Obviously, if the target signal corresponds to a frame in the reference signal, the vector corresponding to the second power spectrum in the reference signal will be similar in shape to the first power spectrum, that is, theoretically, the value of the vector corresponding to the second power spectrum is x times the first power spectrum (x is greater than 1), that is, its theoretical value may be (x*power value 1, x*power value 2..., x*power value 128). Based on this, another larger second power threshold (for example, the second power threshold = x*first power threshold) can be used to binarize the second power spectrum, thereby obtaining similar values (0, 1..., 1).
[0064] Of course, in actual applications, due to the presence of noise and equipment, the two binarized power spectra cannot be exactly the same (i.e., the distance cannot be zero). However, based on this idea, the distance between each binarized second power spectrum and the binarized first power spectrum in the buffer can be calculated separately. And as in the aforementioned coarse delay search method, the delay value of the target frame signal relative to the reference signal is determined based on the binarized second power spectrum with the smallest distance. In this way, the influence of echo attenuation on the power spectrum of the valid signal can be avoided, and the accuracy of the delay value estimation can be improved.
[0065] In another embodiment, in order to eliminate the interference of signals in irrelevant frequency bands in the detection of effective signals, the system can also perform tests on multiple preset frequency bands within the frequency band of the effective signal when it is started. For example, when the frequency band of the effective signal is 20kHz to 22kHz, the energy of the four frequency bands of 20kHz to 20.5kHz, 20.5kHz to 21kHz, 21kHz to 21.5kHz, and 21.5kHz to 22kHz are tested respectively. At this time, it is first detected whether there is interference from other equipment on the scene, and the frequency band with less noise is selected as the frequency band for subsequent ultrasound generation. The ultrasonic generator periodically generates signals within this frequency band; when the power spectrum is calculated for the subsequent time delay estimation, it is also calculated only on this frequency band. The values of the binary power spectrum in other frequency bands can be directly set to 0 to reduce interference. For example, if it is discovered in advance that there is less interference and noise in the frequency band of 20.5kHz to 21kHz, then only ultrasonic waves of 20.5kHz to 21kHz can be generated, and when the power spectrum of the extracted effective signal is subsequently binarized, the values of the frequency bands outside 20.5kHz to 21kHz (i.e., the frequency threshold) in the first power spectrum and the second power spectrum are directly set to 0, thereby reducing the impact of environmental noise on the delay value estimation.
[0066] It should be noted that binarization based on a preset power threshold or frequency threshold can be performed in one or both methods, and the two methods do not conflict with each other. In addition, the system can monitor the environment at any time during startup to determine the optimal frequency threshold. That is, the frequency threshold varies with the environment within the frequency band of the valid signal, improving the frequency threshold's adaptability to the environment.
[0067] In one embodiment, after determining the delay value, linear or nonlinear echo cancellation can be performed based on the delay value. For example, assuming no echo attenuation, the mixed signal and the reference signal can be time-aligned based on the delay value. The aligned mixed signal and reference signal are then sent to a linear filtering and nonlinear processing module for echo cancellation, generating an uplink signal. If echo attenuation is present, the mixed signal can be scaled based on an attenuation coefficient, aligned with the reference signal, and sent to a linear filtering and nonlinear processing module for echo cancellation, generating an uplink signal, thereby achieving accurate echo cancellation.
[0068] Based on the same idea, the embodiments of this specification also provide devices and electronic devices corresponding to the above methods, such as Figure 5 、 Figure 6 shown.
[0069] Figure 5This is a schematic diagram of the structure of a delay estimation device in echo cancellation provided in an embodiment of this specification, the device comprising:
[0070] The signal generating module 501 generates an ultrasonic signal and superimposes the ultrasonic signal and the downlink signal to generate a reference signal;
[0071] The signal acquisition module 503 acquires a mixed signal, wherein the mixed signal includes an echo signal generated by playing the reference signal;
[0072] a signal separation module 505 for separating a first valid signal from the mixed signal and a second valid signal from the reference signal, wherein the frequency bands of the first valid signal and the second valid signal include the frequency band of the ultrasonic signal;
[0073] The power spectrum determining module 507 determines, for any target frame signal in the first valid signal, a first power spectrum of the target frame signal, and determines a plurality of second power spectra corresponding to the plurality of frame signals included in the second valid signal;
[0074] The delay estimation module 509 determines a delay value of the target frame signal relative to the reference signal according to a distance between the first power spectrum and the second power spectrum.
[0075] Furthermore, the signal generating module 501 takes a preset window duration as a period and generates ultrasonic signals with different frame frequencies within the period.
[0076] Furthermore, the delay estimation module 509 determines the second power spectrum with the smallest distance from the first power spectrum as the target power spectrum; determines the first time point corresponding to the first power spectrum, and determines the second time point corresponding to the target power spectrum; and determines the time difference between the first time point and the second time point as the delay value.
[0077] Furthermore, when a frame signal contains multiple sampling point signals, the delay estimation module 509 determines the first time domain characteristics of the multiple sampling point signals in the target frame signal; obtains the delayed signals of multiple consecutive frames including the frame signal corresponding to the second time point, and determines the second time domain characteristics of the multiple sampling point signals in the delayed signal; cross-correlates the first time domain characteristics and the second time domain characteristics to generate a cross-correlation result; and determines the time difference corresponding to the maximum cross-correlation result as the delay value.
[0078] Furthermore, the delay estimation module 509 obtains a specified number of continuous frame signals before and after the second time point to form the delay signal; or, with the second time point as the end point, obtains a specified number of continuous frame signals before the second time point to form the delay signal.
[0079] Furthermore, the delay estimation module 509 binarizes the first power spectrum and the second power spectrum based on a preset power threshold to generate a binarized first power spectrum and a second power spectrum; and determines the delay value of the target frame signal relative to the reference signal based on the distance between the binarized first power spectrum and the second power spectrum.
[0080] Furthermore, the apparatus further includes an echo cancellation module 511 for time-aligning the mixed signal and the reference signal according to the delay value; performing echo cancellation based on the aligned mixed signal and the reference signal to generate an uplink signal.
[0081] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification, the device comprising:
[0082] at least one processor; and,
[0083] a memory communicatively connected to the at least one processor; wherein,
[0084] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the following steps: Figure 2 The method described.
[0085] Based on the same idea, the embodiment of this specification also provides a non-volatile computer storage medium corresponding to the above method, which stores computer executable instructions. After the computer reads the computer instructions in the storage medium, the instructions enable one or more processors to execute the following Figure 2 The method described.
[0086] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0087] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0088] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0089] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0090] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0091] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0092] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0094] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0095] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0096] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0097] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0098] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0099] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.
[0100] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0101] The foregoing description is merely one or more embodiments of this specification and is not intended to limit this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of one or more embodiments of this specification are intended to be within the scope of the claims of this specification.
Claims
1. A method for delay estimation in echo cancellation, comprising: generating an ultrasonic signal, and superimposing the ultrasonic signal and the downlink signal to generate a reference signal; collecting a mixed signal, wherein the mixed signal includes an echo signal generated by playing the reference signal; Separating a first effective signal from the mixed signal, and separating a second effective signal from the reference signal, wherein frequency bands of the first effective signal and the second effective signal include a frequency band of the ultrasonic signal; For any target frame signal in the first valid signal, determine a first power spectrum of the target frame signal, and determine a plurality of second power spectra corresponding to the plurality of frame signals included in the second valid signal; Determining the time delay value of the target frame signal relative to the reference signal according to the distance between the first power spectrum and the second power spectrum, including: determining the second power spectrum having the smallest distance from the first power spectrum as the target power spectrum; determining a first time point corresponding to the first power spectrum, and determining a second time point corresponding to the target power spectrum; and determining a time difference between the first time point and the second time point as the time delay value; The delay value is determined by using a coarse delay search method; When a frame signal includes multiple sampling point signals, the method further includes: Determining first time domain features of multiple sampling point signals in the target frame signal; obtaining a delayed signal of multiple consecutive frames including a frame signal corresponding to a second time point, and determining second time domain features of multiple sampling point signals in the delayed signal; cross-correlating the first time domain feature with the second time domain feature to generate a cross-correlation result; and determining the time difference corresponding to a maximum cross-correlation result as the delay value; The delay value is a more accurate delay value obtained by further adopting a delay fine search method based on the aforementioned delay coarse search.
2. The method according to claim 1, wherein Generates ultrasonic signals, including: The preset window duration is used as a period, and ultrasonic signals with different frame frequencies within the period are generated.
3. The method according to claim 1, wherein Acquiring a delay signal of a plurality of consecutive frames including a frame signal corresponding to a second time point, comprising: Acquire a specified number of consecutive frame signals before and after the second time point to form the time-delay signal; Alternatively, taking the second time point as the end point, frame signals of a specified number of frames continuously before the second time point are acquired to form the time-delay signal.
4. The method according to claim 1, wherein Determining a delay value of the target frame signal relative to the reference signal according to a distance between the first power spectrum and the second power spectrum includes: Binarizing the first power spectrum and the second power spectrum based on a preset power threshold and / or frequency threshold to generate a binarized first power spectrum and a second power spectrum; The time delay value of the target frame signal relative to the reference signal is determined according to the distance between the binarized first power spectrum and the second power spectrum.
5. The method according to claim 1, wherein The method further comprises: Time-aligning the mixed signal and the reference signal according to the delay value; Echo cancellation is performed according to the aligned mixed signal and the reference signal to generate an uplink signal.
6. A delay estimation device for echo cancellation, comprising: A signal generating module, generating an ultrasonic signal, and superimposing the ultrasonic signal and the downlink signal to generate a reference signal; a signal acquisition module for acquiring a mixed signal, wherein the mixed signal includes an echo signal generated by playing the reference signal; a signal separation module, separating a first valid signal from the mixed signal, and separating a second valid signal from the reference signal, wherein the frequency bands of the first valid signal and the second valid signal include the frequency band of the ultrasonic signal; a power spectrum determination module, for determining, for any target frame signal in the first valid signal, a first power spectrum of the target frame signal, and determining a plurality of second power spectra corresponding to the plurality of frame signals included in the second valid signal; a delay estimation module, determining a delay value of the target frame signal relative to the reference signal according to a distance between the first power spectrum and the second power spectrum; The delay estimation module determines the second power spectrum having the smallest distance from the first power spectrum as the target power spectrum; determines a first time point corresponding to the first power spectrum, and determines a second time point corresponding to the target power spectrum; and determines a time difference between the first time point and the second time point as the delay value; The delay value is determined by using a coarse delay search method; When a frame signal includes multiple sampling point signals, the delay estimation module determines a first time domain feature of the multiple sampling point signals in the target frame signal; obtains a delay signal of multiple consecutive frames including a frame signal corresponding to a second time point, and determines a second time domain feature of the multiple sampling point signals in the delay signal; cross-correlates the first time domain feature with the second time domain feature to generate a cross-correlation result; and determines the time difference corresponding to the maximum cross-correlation result as the delay value; The delay value is a more accurate delay value obtained by further adopting a delay fine search method based on the aforementioned delay coarse search.
7. The device according to claim 6, wherein the signal generating module uses a preset window duration as a period and generates ultrasonic signals with different frame frequencies within the period.
8. In the device as described in claim 6, the delay estimation module obtains a specified number of continuous frame signals before and after the second time point to form the delay signal; or, with the second time point as the end point, obtains a specified number of continuous frame signals before the second time point to form the delay signal.
9. In the device as described in claim 6, the delay estimation module binarizes the first power spectrum and the second power spectrum based on a preset power threshold to generate a binarized first power spectrum and a second power spectrum; and determines the delay value of the target frame signal relative to the reference signal based on the distance between the binarized first power spectrum and the second power spectrum.
10. The device according to claim 6, further comprising an echo cancellation module, which performs time alignment on the mixed signal and the reference signal according to the delay value; performs echo cancellation on the aligned mixed signal and the reference signal to generate an uplink signal.
11. An electronic device comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Information processing method and terminal
CN107689228A
Far-field pickup device and method for collecting human voice signals in far-field pickup device
CN110166882A