Abnormal Echo Delay Identification Method, Device, Terminal and Storage Medium
By matching audio features after delaying the target delay in the echo cancellation system, detecting and correcting abnormal echo delay, the problem of negative delay in the echo cancellation process is solved, and the accuracy of echo cancellation and voice call quality are improved.
Patent Information
- Application Number
- CN202110936165.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-08-16
AI Technical Summary
In the prior art, negative delays are prone to occur during the echo cancellation process, which causes the echo cancellation module to be unable to accurately estimate the echo delay, affecting the quality of voice calls.
By extracting audio features of the audio frames collected by the microphone, matching features after delaying the target, detecting and correcting abnormal echo delays, ensuring that the echo cancellation module does not continue to work under error delays.
Improves the accuracy of echo cancellation, avoids echo cancellation failure caused by negative delay, and improves voice call quality.
Smart Images

Figure CN115706756B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of call technologies, and particularly to an abnormal echo delay identification method, device, terminal, and storage medium. Background Art
[0002] During a voice call through an audio terminal device, the sound played by the speaker, especially the sound played through the speaker in the hands-free mode, is relatively loud and is easily collected by the microphone; as a result, the sound played by the speaker is collected by the microphone and then fed back to the far end, and the person speaking at the far end will hear their own voice, forming an echo and seriously affecting the quality of the voice call.
[0003] In the related art, software or hardware echo cancellation modules are deployed in audio terminal devices to cancel the echo collected by the microphone. The echo cancellation process generally uses the echo cancellation method. By comparing the reference point signal and the receiving point signal, the echo delay from the reference point signal to the receiving point signal is estimated, and after delaying the reference point signal according to the echo delay, the transfer function is calculated together with the receiving point signal. The transfer function is used to predict the echo replica, so that the receiving point signal can be echo-cancelled using the echo replica.
[0004] As can be seen from the above echo cancellation process, accurate echo delay estimation is a prerequisite for improving the echo cancellation effect. Summary of the Invention
[0005] The embodiments of the present application provide an abnormal echo delay identification method, device, terminal, and storage medium. The technical solutions are as follows:
[0006] According to one aspect of the present application, an abnormal echo delay identification method is provided. The method includes:
[0007] Extract audio features from the input audio frame collected by the microphone to obtain the first audio feature;
[0008] In response to reaching the target delay, based on the first audio feature, determine the second audio feature from the candidate audio features. The candidate audio features are the audio features corresponding to the output audio frames for playing by the speaker, and the second audio feature matches the first audio feature;
[0009] Determine the echo delay of the output audio frame corresponding to the second audio feature;
[0010] In response to the echo delay being less than the target delay, determine that there is an abnormal echo delay.
[0011] According to another aspect of the present application, an abnormal echo delay identification device is provided. The device includes:
[0012] A feature extraction module, configured to extract audio features from the input audio frames collected by a microphone to obtain first audio features;
[0013] A first determination module, configured to, in response to reaching a target delay, determine second audio features from candidate audio features based on the first audio features, where the candidate audio features are audio features corresponding to output audio frames for playing by a speaker, and the second audio features match the first audio features;
[0014] A second determination module, configured to determine the echo delay of the output audio frames corresponding to the second audio features;
[0015] A third determination module, configured to, in response to the echo delay being less than the target delay, determine that there is an abnormal echo delay.
[0016] According to another aspect of the present application, a terminal is provided. The terminal includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the abnormal echo delay recognition method described in the above aspect.
[0017] According to another aspect of the present application, a computer-readable storage medium is provided. At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the abnormal echo delay recognition method described in the above aspect.
[0018] According to another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the terminal reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the terminal executes the abnormal echo delay recognition method provided in the above optional implementation manner.
[0019] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0020] In the embodiments of the present application, in view of the situation of negative delay occurring during the echo cancellation process, which makes it necessary to wait for a period of time to find a matching output audio frame after receiving an input audio frame at the receiving point, a method for detecting negative delay during the echo cancellation process is proposed. After obtaining the first audio feature corresponding to the input audio frame and delaying it by the target delay, feature matching is performed based on the first audio feature, so that the echo delay can still be estimated in the case of negative delay; and the delayed feature matching makes the calculated echo delay the sum of the target delay and the propagation delay. Then, based on the relationship between the echo delay and the target delay, the positivity or negativity of the propagation delay can be determined, so that it can be timely determined whether there is an abnormal echo delay (negative delay), avoiding the echo cancellation module from continuing the echo cancellation work under the wrong echo delay, or avoiding the situation where echo cancellation cannot be performed due to the inability to calculate the echo delay, thereby improving the accuracy of echo cancellation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 shows a schematic structural diagram of an echo cancellation system in the related art;
[0023] Figure 2 shows a schematic structural diagram of an echo cancellation system shown in an exemplary embodiment of the present application;
[0024] Figure 3 shows a flowchart of an abnormal echo delay identification method provided by an exemplary embodiment of the present application;
[0025] Figure 4 shows a schematic diagram of the process of determining the echo delay shown in an exemplary embodiment of the present application;
[0026] Figure 5 shows a flowchart of an abnormal echo delay identification method provided by another exemplary embodiment of the present application;
[0027] Figure 6 shows a working schematic diagram of the delay estimation process and the abnormal delay detection process shown in an exemplary embodiment of the present application;
[0028] Figure 7 shows a schematic diagram of the extraction process of audio features shown in an exemplary embodiment of the present application;
[0029] Figure 8Shows a schematic diagram of the principle of implementing the delay function in the feature storage area shown in an exemplary embodiment of the present application;
[0030] Figure 9 Shows a schematic diagram of the process of delay feature matching shown in an exemplary embodiment of the present application;
[0031] Figure 10 Shows a schematic diagram of the feature matching process shown in an exemplary embodiment of the present application;
[0032] Figure 11 Shows a schematic diagram of the delay estimation process and the abnormal delay detection process shown in another exemplary embodiment of the present application;
[0033] Figure 12 Shows a schematic diagram of the process of audio feature extraction in the delay estimation process shown in an exemplary embodiment of the present application;
[0034] Figure 13 Shows a flowchart of an abnormal echo delay identification method provided in another exemplary embodiment of the present application;
[0035] Figure 14 Is a structural block diagram of an abnormal echo delay identification device provided in an exemplary embodiment of the present application;
[0036] Figure 15 Shows a structural block diagram of a terminal provided in an exemplary embodiment of the present application. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0038] During a voice call through an audio terminal device, the sound played through the speaker (horn), especially the sound played in the external speaker mode, is relatively loud, and the played sound is easily re-collected by the microphone of the audio terminal device, so that the sound played by the speaker will be sent to the remote end again after being collected by the microphone, and the speaker at the remote end will hear his own voice, thereby forming an echo and seriously affecting the call quality. Therefore, generally, an audio terminal device (call device) will include a software or hardware echo cancellation system, and the echo cancellation system is used to cancel the echo of the sound collected by the microphone. Figure 1The structural schematic diagram of an echo cancellation system in the related art is shown. During a voice call, the audio signal received by the terminal from the remote end needs to pass through a reference point before being sent to the speaker for playback. The audio signal collected at the reference point is generally called the reference point signal. The reference point signal is sent to the speaker for playback through the software and hardware playback logic, propagates through media such as air into the microphone, and then reaches the receiving point position through the software and hardware acquisition logic. The audio signal obtained from the receiving point is called the receiving point signal. It can be seen that when the reference point signal reaches the receiving point position, it needs to go through software and hardware playback delay (the time from the reference point to the speaker), acoustic path delay (the time from the speaker to the microphone through transmission media such as air), and software and hardware acquisition delay (the time from the microphone to the receiving point through the software and hardware acquisition logic).
[0039] As Figure 1 shown, in order to implement echo cancellation, a delay estimation module 101, a delay alignment module 102, and an echo cancellation module 103 are provided in the audio terminal device (call device). Among them, the delay estimation module 101 is used to compare the reference point signal and the receiving point signal, so as to estimate the echo delay (Tde) from the reference point signal to the receiving point signal, and transfer the Tde to the delay alignment module 102; the delay alignment module 102 is used to delay the reference point signal by Tde according to the Tde estimated by the delay estimation module 101 to obtain the delayed reference point signal, and send the delayed reference point signal to the echo cancellation module 103. The echo cancellation module 103 estimates the transfer function of the echo path based on the delayed reference point signal and the receiving point signal. Then, when a new reference point signal is received by the reference point signal, an echo copy corresponding to the reference point signal can be predicted based on the new reference point signal and the transfer function. By subtracting the echo copy from the receiving point signal, the echo signal in the receiving point signal can be eliminated.
[0040] From Figure 1 the above echo cancellation system, it can be seen that whether the delay estimation module 101 can accurately estimate the echo delay Tde will affect the accuracy of echo cancellation. During the delay estimation process, it may be affected by reasons such as the operating system's scheduling of the microphone acquisition thread and the speaker playback thread, and the stability of the application program. The echo delay from the reference point to the receiving point is not fixed. For example, when the thread for reading the receiving point signal at the receiving point is stuck, it will cause the receiving point to obtain the newly collected receiving point signal in advance, and it is impossible to find a matching audio feature from the feature storage area storing the audio features corresponding to the reference point signal, so that the echo delay Tde cannot be estimated or accurately calculated, resulting in a negative delay situation, which seriously affects the effect of echo cancellation or makes it impossible to achieve echo cancellation.
[0041] It can be seen that in the process of echo cancellation, how to effectively and timely detect whether there is abnormal echo delay in the echo cancellation system can avoid the situation that the echo cancellation system uses incorrect delay for echo cancellation or fails to perform echo cancellation, and then effective measures can be taken in time to eliminate the abnormal delay, which is the key to improving the accuracy of echo cancellation. Based on the above problems, as Figure 2 shown, it shows a schematic structural diagram of an echo cancellation system shown in an exemplary embodiment of the present application. The echo cancellation system mainly includes an abnormality detection module 201, a delay estimation module 202, a delay alignment module 203, and an echo cancellation module 204. Compared with Figure 1 , the present application adds an abnormality detection module 201 to the echo cancellation system. The abnormality detection module 201 can detect possible abnormal delays (negative delays) based on the reference point signal and the receiving point signal. When a negative delay is detected, a reset instruction can be sent to the delay estimation module so that both the delay estimation module 202 and the delay alignment module 203 can reset the algorithm and clear the cache, so that the subsequent echo cancellation process can perform echo delay estimation normally.
[0042] Please refer to Figure 3 , which shows a flowchart of an abnormal echo delay identification method provided by an exemplary embodiment of the present application. The embodiment of the present application takes the application of this method to a terminal as an example for description. The method includes:
[0043] Step 301, extract audio features from the input audio frame collected by the microphone to obtain the first audio feature.
[0044] Among them, the input audio frame is collected by the microphone. During the microphone collection process, the input audio signal (input audio frame) collected by the microphone will first be stored in the buffer corresponding to the software and hardware collection logic, and then the input audio frame is read from the buffer by calling the application program thread.
[0045] As Figure 2 shown, the reference point signal needs to pass through the software and hardware playback logic, media such as air, and software and hardware collection logic to reach the receiving point. The same audio signal is not exactly the same at the reference point and the receiving point, but should be the most similar. Therefore, in a possible implementation manner, in order to determine the audio frame most similar to the input audio frame from the reference point signal, an audio feature matching method can be used, that is, by extracting audio features from the input audio frame to obtain the first audio feature, and then based on the first audio feature and the audio features corresponding to the historical output audio frame (reference point signal) for feature matching, so as to determine the audio feature most similar to the first audio feature for subsequent echo delay estimation.
[0046] Since the first audio feature is used for audio frame matching, to improve the accuracy of audio matching, the first audio feature generally selects features unique to the audio signal. By way of illustration, the first audio feature may be Mel-frequency Cepstral Coefficients (MFCC), Fourier coefficients, Linear Prediction Coefficients (LPC), audio features based on spectral energy, etc.
[0047] Step 302: In response to reaching the target delay, based on the first audio feature, determine a second audio feature from the candidate audio features. The candidate audio features are the audio features corresponding to the output audio frames, and the output audio frames are used for speaker playback. The second audio feature matches the first audio feature.
[0048] During the echo cancellation process, when there is a lag in the process of the application thread reading the input audio frames from the buffer (the buffer corresponding to the software and hardware acquisition logic), the buffer follows the principle of first in, first out, and the data storage capacity in the buffer is limited. If the data reading thread lags, but the data writing thread still continues, when the buffer is full of data, the newly acquired input audio frames will overwrite the input audio frames historically stored in the buffer. If the lag is restored subsequently, the new input audio frames will be directly read, causing the new input audio frames to arrive at the receiving point in advance; at the same time, the lag process will also cause the candidate audio features corresponding to the historical output audio frames to stop being written to the historical feature storage area, resulting in the inability to find the output audio frames that match the current input audio frames in the historical feature storage area, leading to the inability to perform delay estimation or incorrect delay estimation, that is, the occurrence of negative delay (the audio features of the output audio frames corresponding to the current input audio frames need to be written to the historical feature storage area some time after receiving the current input audio frames); therefore, based on the characteristics of the occurrence of negative delay, in order to accurately detect the situation of negative delay, that is, in order to still calculate the echo delay when negative delay occurs, in a possible implementation manner, after the first audio feature corresponding to the input audio frame is extracted, the feature matching is not immediately performed based on this first audio feature, but after a target delay, the feature matching is performed based on this first audio feature, so as to find the matching candidate audio features.
[0049] Among them, the target delay can be set by the developer or set by the business personnel based on the actual situation. Since the target delay is to ensure that the second audio feature matching the first audio feature can still be found in the case of negative delay, the corresponding target delay needs to be greater than or equal to the maximum negative delay that the operating system can have. By way of illustration, if the maximum negative delay is 100 ms, the target delay can be 120 ms. The specific value of the target delay in the embodiments of the present application is not limited.
[0050] Optionally, the candidate audio feature is the audio feature corresponding to the output audio frame, and the output audio frame is obtained before being sent to the speaker for playback. That is, when the output audio signal passes through the reference point, it is stored in the buffer, and when a frame of input audio frame is received at the receiving point, a frame of historical output audio frame is read from the buffer, audio feature extraction is performed, and the extracted candidate audio feature is stored in the historical feature storage area for subsequent searching for the second audio feature matching the first audio feature from the historical feature storage area.
[0051] Step 303, determine the echo delay of the output audio frame corresponding to the second audio feature.
[0052] In a possible implementation manner, after the second audio feature matching the first audio feature is determined from the candidate audio features, it indicates that the input audio frame is the audio signal corresponding to the output audio frame of the second audio feature when it reaches the receiving point after passing through the echo path, and during the echo delay period, the storage position of the second audio feature in the historical feature memory moves with time. Correspondingly, the echo delay of the output audio frame corresponding to the second audio feature can be determined based on the storage position of the second audio feature in the historical feature memory.
[0053] By way of illustration, such as Figure 4As shown, it shows a schematic diagram of the process of determining the echo delay shown in an exemplary embodiment of the present application. The output audio frame Q at the reference point reaches the receiving point after a transmission delay Td to obtain the input audio frame Q'. Due to the influence of distortion and noise in the transmission channel, Q and Q' are different but similar signals. By performing audio feature extraction on the input audio frame Q' at the receiving point, the first audio feature G is obtained; assuming that the output audio frame Q is the 80th audio frame, the second audio feature obtained by performing audio feature extraction on the output audio frame Q is F(80), and F(80) is stored at the tail position (tail) of the historical feature memory. After the transmission delay Td, 31 new output audio frames are obtained at the reference point and 31 candidate audio features are calculated and placed in the historical feature storage area in sequence; therefore, when the first audio feature G corresponding to the input audio frame Q' is obtained at the receiving point, 31 candidate audio features from F(81) to F(111) are newly stored in the historical feature memory; due to the existence of the target delay, after the target delay, feature matching is performed based on the first audio feature G and the candidate audio features stored in the historical feature storage area. After the target delay, 31 candidate audio features from F(111) to F(141) are newly stored in the historical feature memory; then the echo delay at this time is the length of 61 output audio frames from F(80) to F(141), and correspondingly, the echo delay can be calculated based on the sampling interval and the number of frames of each output audio frame.
[0054] Step 304, in response to the echo delay being less than the target delay, determine that there is an abnormal echo delay.
[0055] Since in the present application, after the input audio frame is received at the receiving point, audio feature matching is performed after a target delay. When the operating system is running normally (i.e., there is no negative delay), the estimated echo delay should be the sum of the transmission delay and the target delay. The transmission delay is the delay from the reference point signal to the receiving point signal. Normally, the transmission delay is a positive integer, that is to say, the echo delay must be greater than the target delay; conversely, if the echo delay is less than the target delay, it means that the transmission delay is negative, that is, there is a negative delay situation. Therefore, in a possible implementation manner, when it is determined that the echo delay is less than the target delay, it indicates that there is a negative delay situation caused by abnormal thread calls in the system, that is, there is an abnormal echo delay. Correspondingly, the delay estimation module cannot accurately estimate the transmission delay, and it is necessary to reset the delay estimation algorithm and clear its cache to eliminate the negative delay situation.
[0056] In summary, in the embodiments of the present application, in view of the situation of negative delay occurring during echo cancellation, which causes a characteristic that after receiving an input audio frame at the receiving point, it is necessary to wait for a period of time to find a matching output audio frame, a method for detecting negative delay during echo cancellation is proposed. After obtaining the first audio feature corresponding to the input audio frame and delaying it by the target delay, feature matching is performed based on the first audio feature, so that the echo delay can still be estimated in the case of negative delay; and the delayed feature matching makes the calculated echo delay the sum of the target delay and the propagation delay. Then, based on the relationship between the echo delay and the target delay, the positivity or negativity of the propagation delay can be determined, so that it is possible to timely determine whether there is an abnormal echo delay (negative delay), avoid the echo cancellation module from continuing to perform echo cancellation work under the wrong echo delay, or avoid the situation where echo cancellation cannot be performed due to the inability to calculate the echo delay, thereby improving the accuracy of echo cancellation.
[0057] Since after the receiving point obtains the first audio feature corresponding to the input audio frame, it will not immediately perform feature matching to avoid being unable to calculate the echo delay in the case of negative delay. It will perform feature matching after reaching the target delay. During the target delay, the receiving point will still receive new input audio frames and extract audio features from them. In a possible implementation manner, a feature memory for storing the first audio feature corresponding to the input audio frame is added to enable the delayed function of feature matching through this feature memory.
[0058] In an exemplary example, as Figure 5 shown, it shows a flowchart of an abnormal echo delay identification method provided by another exemplary embodiment of the present application. The embodiments of the present application are described by taking the application of this method to a terminal as an example. The method includes:
[0059] Step 501, perform time-frequency conversion and frequency band division on the input audio frame collected by the microphone to determine M sub-bands.
[0060] It should be noted that in the embodiments of the present application, an abnormal detection module (abnormal delay detection module) is newly added to the original echo cancellation system, and both the abnormal detection module and the original delay estimation module need to extract audio features. In a possible scenario, the audio features extracted in the abnormal detection module can be the same as those extracted in the original delay estimation module; optionally, the audio features extracted in the abnormal detection module can also be different from those extracted in the original delay estimation module.
[0061] When the audio features extracted by the anomaly detection module are different from those extracted by the original delay estimation module, different feature storage areas need to be allocated for the audio features. Correspondingly, at least two feature storage areas need to be added to store the audio features, where the first feature storage area is used to store the audio features corresponding to the input audio frame, and the second feature storage area is used to store the audio features corresponding to the output audio frame.
[0062] Indicatively, Figure 6 As shown, it shows a schematic diagram of the delay estimation process and abnormal delay detection process of the present application. Figure 6 It can be seen that the feature storage area of the audio features extracted in the delay estimation process 601 is different from the feature storage area of the audio features extracted in the abnormal delay detection process 602, wherein the delay estimation process corresponds to feature storage area 1, while the abnormal delay detection process 602 corresponds to feature storage area 3 and feature storage area 4.
[0063] In the delay estimation process 601, audio features are extracted from the reference point signal (output audio frame) through the feature extraction module 1, and the extracted candidate audio features are stored in the feature storage area 1; when the receiving point signal is obtained, audio features are extracted from the receiving point signal (input audio frame) through the feature extraction module 2, and the extracted first audio feature is sent to the feature matching module 1 for feature matching, specifically, based on the first audio feature, a matching second audio feature is searched from the feature storage area 1, and then the transmission delay Tde is determined based on the delay determination strategy 1.
[0064] In the abnormal delay detection process 602, the reference point signal (output audio frame) is subjected to audio feature extraction by the feature extraction module 3, and the extracted candidate audio features are stored in the feature storage area 3, and the receiving point signal (input audio frame) is subjected to audio feature extraction by the feature extraction module 4, and the extracted first audio features are stored in the feature storage area 4, and then after the target delay, the first audio feature is read from the feature storage area 4 by the feature matching module 3, and the matching second audio feature is searched from the feature storage area 3 based on the first audio feature, and then the echo delay Tde3 is calculated by the delay determination strategy 3; at the abnormal delay detection point, it is determined whether the echo delay Tde3 is an abnormal echo delay, and if so, each module in the delay estimation process is reset.
[0065] The audio features extracted in this embodiment are frequency domain features. In a possible implementation, the input audio frame is firstly subjected to time-frequency conversion and frequency band division to obtain the required M sub-bands, and then the corresponding first audio features are extracted based on the M sub-bands.
[0066] Schematically, the value of M can be set by developers. Since the first audio feature needs to be stored in the feature storage area, to improve the calculation efficiency, M can be a quantity that is convenient for subsequent storage and calculation. Schematically, the value of M is 36. 36 sub-bands can obtain audio features of 32 binary values, and 32 bits can just be stored in a 32-bit integer variable, which is convenient for improving storage and calculation efficiency.
[0067] Optionally, the short-time Fourier transform can be used to transform the input audio frame into the frequency domain to obtain the spectral signal of the input audio frame, and then the spectral signal is divided into frequency bands. The frequency band division method can adopt linear division or Mel frequency band division according to psychoacoustic theory, so as to obtain M sub-bands.
[0068] Optionally, in a voice call scenario, the sampling rates of the reference point signal and the receiving point signal are 16 kHz or 32 kHz, etc. Taking 32 kHz as an example, according to the Nyquist sampling theorem, the effective bandwidth of the voice collected at a sampling rate of 32 kHz is 16 kHz; while the voice bandwidth required for extracting audio features does not need to be very high, because in an actual voice communication system, the frequency range that can reliably represent the voice is about 300 Hz to 3 kHz. Therefore, the actual bandwidth required is only about 3 kHz. Therefore, in order to reduce the calculation amount, in the audio feature extraction process, the audio frames with a high sampling rate at the reference point or the receiving point are downsampled to about 6 kHz (the effective audio bandwidth corresponding to a sampling rate of 6 kHz is 3 kHz), and then the subsequent audio feature extraction process is carried out.
[0069] Schematically, as Figure 7 shown, it shows a schematic diagram of the extraction process of the audio feature shown in an exemplary embodiment of the present application. The audio frame after downsampling the input audio frame is transformed into the frequency domain through a time-frequency domain conversion algorithm (such as the short-time Fourier transform) to obtain the spectral signal of the input audio frame; the spectral signal is divided into frequency bands, and the spectrum is divided into 36 sub-bands (sub-band 1 to sub-band 36).
[0070] Step 502, by performing frequency-domain energy comparison on the sub-band energies of the M sub-bands, N first frequency-domain feature scores are obtained, N is a positive integer, and M - N is a positive integer.
[0071] In a possible implementation manner, when performing time-frequency conversion and frequency band division on the input audio frame to obtain M sub-bands, the sub-band energy corresponding to each sub-band is calculated respectively, and then the frequency-domain energy comparison is performed on the sub-band energies, so as to obtain N frequency-domain feature scores.
[0072] Among them, the specific value of N can be set by developers. To improve storage and calculation efficiency, the value of N can be 32 because 32 bits can exactly be stored in a 31-bit integer variable; optionally, the value of N can also be 16, 64 or other values.
[0073] Optionally, the process of calculating the frequency-domain feature score may include the following steps, that is, step 502 may include step 502A and step 502B.
[0074] Step 502A, in response to the j-th subband energy being the maximum value among the subband energies corresponding to the (j - i)-th subband to the (j + i)-th subband, determining the first score as the first frequency-domain feature score corresponding to the j-th subband, where the j-th subband energy is the subband energy corresponding to the j-th subband, and where i is a positive integer and j - i is a positive integer, and j + i is less than or equal to M.
[0075] In a possible implementation manner, when calculating the frequency-domain feature score, comparing the current subband energy with the subband energies of the two adjacent upper and lower subbands. If the current subband energy has the maximum value, output the binary value 1, otherwise output the binary value 0; that is, for the j-th subband, if the j-th subband energy corresponding to the j-th subband is the maximum value among the subband energies corresponding to the (j - i)-th subband to the (j + i)-th subband, then determine the first score as the first frequency-domain feature score corresponding to the j-th subband. Schematically, the first score can be 1.
[0076] Among them, the value of i can be set by developers. Schematically, the value of i can be 1, corresponding to comparing the current subband energy with the subband energies of the two adjacent upper and lower subbands; if the value of i is 2, it corresponds to comparing the current subband energy with the subband energies of the two adjacent upper and lower subbands.
[0077] Step 502B, in response to the j-th subband energy being less than the maximum value among the subband energies corresponding to the (j - i)-th subband to the (j + i)-th subband, determining the second score as the first frequency-domain feature score corresponding to the j-th subband.
[0078] Conversely, for the j-th subband energy, if the j-th subband energy is less than the maximum value among the subband energies corresponding to the (j - i)-th subband to the (j + i)-th subband, then determine the second score as the first frequency-domain feature score corresponding to the j-th subband. Schematically, the second score can be 0.
[0079] Such as Figure 7As shown, after obtaining M sub-bands, the energy of each sub-band is calculated respectively to obtain 36 sub-band energies (sub-band energy 1 to sub-band energy 36); when the value of i is 2, the comparator compares the current sub-band energy with the energies of the two adjacent sub-bands above and below it; for example, comparator 1 compares the sub-band energy 3 with the energies of the two adjacent sub-bands above and below it. That is, if the sub-band energy 3 has the maximum value among the sub-band energies 1 to sub-band energy 5, comparator 1 sets the frequency domain feature score corresponding to the sub-band energy 3 to 1, that is, comparator 1 outputs G(n, 1), otherwise comparator 1 outputs G(n, 0); since the sub-band energy 4 has the maximum value among the sub-band energies 2 to sub-band energy 6, comparator 2 outputs G(n, 1); where n is the nth audio frame.
[0080] Step 503, determining the set of N first frequency domain feature scores as the first audio feature.
[0081] In a possible implementation, by performing energy comparison on the sub-band energies of M sub-bands, N first frequency domain feature scores can be obtained, and the set of N first frequency domain feature scores is the first audio feature corresponding to the input audio frame.
[0082] Exemplarily, if N is 32, the first audio feature is a set of 32-bit binary values.
[0083] It should be noted that the feature extraction process of the output audio frame can refer to the feature extraction process of the input audio frame, and this embodiment will not elaborate here.
[0084] Step 504, storing the first audio feature at the tail storage position of the first feature storage area. The first storage capacity of the first feature storage area is determined by the target delay, and the storage position of the first audio feature moves from the tail storage position to the head storage position over time.
[0085] In a possible implementation, a first feature storage area is provided for storing the first audio feature corresponding to the input audio frame, and this first feature storage area is a first-in-first-out storage area, that is, the newly added audio feature will be stored at the tail position of the first feature storage area, and for each newly added first audio feature corresponding to an input audio frame, the first feature storage area will correspondingly delete a first audio feature corresponding to an input audio frame from the head.
[0086] In order to enable the first feature storage area to achieve the function of feature matching with a delay of the target delay, in a possible implementation, the first storage capacity of the first feature storage area needs to be determined by the target delay. That is to say, when the first audio feature moves from the tail storage position to the head storage position in the first feature storage area over time for feature matching, so that the first audio feature can just perform feature matching after a delay of the target delay.
[0087] Optionally, the first storage capacity is determined by the target delay and the sampling time interval between adjacent input audio frames. Schematically, the target delay is 120 ms, and each input audio frame takes 4 ms to reach the next audio frame. This means that the first feature storage area needs to delay the first audio feature by 30 frames before performing feature matching. Correspondingly, the first feature storage area needs to store the first audio features corresponding to 31 frames of audio input frames.
[0088] In a possible implementation manner, after the first audio feature corresponding to the input audio frame is extracted, instead of directly performing feature matching based on the first audio feature, the first audio feature is stored at the tail storage position of the first feature storage area. During the target delay process, the first storage area continuously adds new first audio features, and the storage position of the first audio feature moves from the tail storage position to the head storage position as the number of audio features increases, that is, the storage position of the first audio feature moves from the tail storage position to the head storage position over time. When the first audio feature moves to the head storage position, it can be determined that the target delay is reached.
[0089] Schematically, as Figure 8 shown, it shows a schematic diagram of the principle of the feature storage area implementing the delay function shown in an exemplary embodiment of the present application. Before the first feature storage area writes the first audio feature G(130) corresponding to the 130th input audio frame, the first audio features corresponding to 31 input audio frames such as G(99) to G(129) are stored in the first feature storage area; when the first audio feature G(130) corresponding to the 130th frame is obtained and G(130) is written into the tail storage area (tail) of the first feature storage area, the first audio features stored in the first feature storage area at this time are: G(100) to G(130); during the target delay process, G(130) moves from the tail storage area to the head storage area position over time. When G(130) moves to the head storage position, it is determined that the target delay is reached. At this time, the first audio features stored in the first feature storage area are: G(130) to G(160).
[0090] Step 505, in response to the first audio feature moving to the head storage position of the first feature storage area, determine the second audio feature from the candidate audio features based on the first audio feature.
[0091] In a possible implementation manner, when the first audio feature moves to the head storage position of the first feature storage area, based on the relationship between the first storage capacity of the first feature storage area and the target delay, it is determined that the target delay is reached. Furthermore, the second audio feature can be determined from the candidate audio features based on the first audio feature.
[0092] Among them, the process of determining the second audio feature from the candidate audio features based on the first audio feature (feature matching process) may include Step 1 and Step 2.
[0093] 1. Perform feature matching on the first audio feature and the candidate audio features to obtain at least one candidate matching score. The candidate matching score is used to indicate the matching degree between the first audio feature and the candidate audio features, and the candidate matching score is negatively correlated with the matching degree.
[0094] In a possible implementation manner, after a target delay, perform feature matching on the first audio feature and each candidate audio feature stored in the current historical feature memory (the historical feature memory stores the candidate audio features corresponding to the output audio frames), obtain at least one candidate matching score, and then determine the second audio feature that matches the first audio feature from the candidate audio features according to the candidate matching score.
[0095] Among them, the number of candidate matching scores is determined by the storage capacity of the historical feature storage area. Schematically, if the historical feature storage area stores the candidate audio features corresponding to 75 output audio frames, the first audio feature is respectively matched with each candidate audio feature corresponding to 75 output audio frames, so that 75 candidate matching scores can be obtained.
[0096] Schematically, as Figure 10 shown, it shows a schematic diagram of the feature matching process shown in an exemplary embodiment of the present application. When searching for a matching second audio feature from the candidate audio features based on the first audio feature G(80) after the target delay, at this time, 75 candidate audio features such as F(67) to F(141) are stored in the historical feature memory, and correspondingly, G(80) is respectively matched with 75 candidate audio features to obtain 75 candidate matching scores: S(1) to S(75).
[0097] As can be seen from the above embodiments, each audio feature contains N frequency domain feature scores. Correspondingly, in the feature matching process, it is also necessary to perform matching operations on each frequency domain feature score respectively. In an exemplary example, Step 1 may further include Step 1 and Step 2 (that is, the process of determining the candidate matching score may further include Step 1 and Step 2).
[0098] 1. Perform a matching operation on the k-th first frequency domain feature score and the k-th second frequency domain feature score to obtain the k-th sub-matching score, where k is a positive integer less than or equal to N, and the sub-matching score is the matching degree of the k-th frequency domain feature score in the first audio feature and the second audio feature.
[0099] In order to perform feature matching, it is necessary to ensure that the candidate audio feature and the first audio feature are extracted using the same feature extraction method, and the number of frequency domain feature scores included in the candidate audio feature needs to be the same as the number of frequency domain feature scores included in the first audio feature. Schematically, the first audio feature includes N first frequency domain feature scores, and correspondingly, the candidate audio feature also includes N second frequency domain feature scores, where N is a positive integer. Schematically, when N is 32, the first audio feature includes 32 binary values, and the candidate audio feature also includes 32 binary values.
[0100] In a possible implementation manner, during the matching operation, matching operations are respectively performed on the N first frequency domain feature scores and the N second frequency domain feature scores. That is to say, a matching operation is performed on the k-th first frequency domain feature score and the k-th second audio feature to obtain the k-th sub-matching score, and then the sum of the N sub-matching scores is determined as the matching score between the first audio feature and the candidate audio feature.
[0101] Schematically, if N is 5, the first audio feature is: G1(n, 0), G2(n, 1), G3(n, 1), G4(n, 0), G5(n, 1), and the candidate audio feature is: F1(n, 1), F2(n, 0), F3(n, 0), F4(n, 1), F5(n, 0); then when performing a matching operation on the first audio feature and the candidate audio feature, a matching operation is performed on G1(n, 0) and F1(n, 1) to obtain the first sub-matching score, a matching operation is performed on G2(n, 1) and F2(n, 0) to obtain the second sub-matching score, and similarly, the third sub-matching score, the fourth sub-matching score, and the fifth sub-matching score can be obtained, and then the sum of the first sub-matching score to the fifth sub-matching score is determined as the candidate matching score.
[0102] Optionally, the matching operation can use an exclusive OR operation or other matching algorithms, and the embodiments of the present application do not limit this.
[0103] 2. Determine the sum of the N sub-matching scores as the candidate matching score.
[0104] In a possible implementation manner, the sum of the N sub-matching scores is determined as the candidate matching score. Schematically, if the matching operation uses an exclusive OR operation, the more similar the first audio feature and the candidate audio feature are, the lower the corresponding candidate matching score is, that is, the candidate matching score is negatively correlated with the matching degree.
[0105] II. Determine the second audio feature from the candidate audio features based on at least one candidate matching score.
[0106] The purpose of feature matching is to find a second audio feature that matches the first audio feature from candidate audio features. Based on the relationship between the matching degree and the candidate matching score, the higher the matching degree, the lower the candidate matching score. Correspondingly, the candidate audio feature corresponding to the minimum value in the candidate matching scores can be determined as the second audio feature.
[0107] Optionally, in other possible implementation manners, after calculating the candidate matching scores, they can be smoothed based on historical matching results to obtain feature matching scores, and then the second audio feature can be determined based on the feature matching scores. In an exemplary example, step two may further include steps 3 to 5.
[0108] 3. Smooth at least one candidate matching score to obtain at least one feature matching score.
[0109] In a possible implementation manner, each candidate matching score is smoothed to obtain at least one feature matching score.
[0110] In an exemplary example, the calculation process of the feature matching score can be expressed as:
[0111] Sm(n) = S(n) * b + Sm(n)' * (1 - b)
[0112] where Sm(n) represents the feature matching score, S(n) represents the candidate matching score, b represents the smoothing coefficient, b is a decimal between 0 and 1, the smaller b is, the higher the smoothing degree, and Sm(n)' represents the previous smoothing result.
[0113] Schematically, as Figure 10 shown, 75 candidate matching scores S(1) to S(75) are smoothed to obtain 75 feature matching scores: Sm(1) to Sm(75).
[0114] 4. Determine the minimum value in the feature matching scores as the target matching score.
[0115] Since the candidate matching score is negatively correlated with the matching degree, the feature matching score is also negatively correlated with the matching degree. That is to say, the more similar (more matching) the candidate audio feature is to the first audio feature, the smaller the feature matching score corresponding to the candidate audio feature and the first audio feature. Therefore, in a possible implementation manner, the minimum value in the feature matching scores is determined as the target matching score, and then the candidate audio feature corresponding to the target matching score is determined as the second audio feature.
[0116] 5. Determine the candidate audio feature corresponding to the target matching score as the second audio feature.
[0117] In a possible implementation, since the candidate audio feature corresponding to the target matching score is the audio feature that best matches the first audio feature, it can be directly determined as the second audio feature.
[0118] Optionally, to further improve the accuracy of feature matching, in a possible implementation, a matching score threshold is set, and only when the target matching score is less than this matching score threshold will the corresponding candidate audio feature be determined as the second audio feature.
[0119] Step 506: Obtain the target storage location of the second audio feature in the second feature storage area.
[0120] In a possible implementation, after the second audio feature is determined, since the second audio feature is also stored in the first-in, first-out second feature storage area, and the storage location of the second audio feature moves from the tail storage location to the head storage location over time, therefore, the target echo delay can be determined based on the position of the second audio feature in the second feature storage area.
[0121] Step 507: Determine the target echo delay of the output audio frame corresponding to the second audio feature based on the target storage location.
[0122] Such as Figure 9As shown, it shows a schematic diagram of the process of delayed feature matching shown in an exemplary embodiment of the present application. The output audio frame Q at the reference point reaches the receiving point after a transmission delay Td to obtain the input audio frame Q'. The feature G(100) is obtained by performing audio feature extraction on the input audio frame Q' at the receiving point and stored in the tail storage position of the first feature storage area (the position where G(130) is located in the figure); it is assumed that the output audio frame Q is the 100th audio frame. The second audio feature obtained by performing audio feature extraction on the output audio frame Q is F(100), and F(100) is stored in the tail storage position (tail) of the second feature storage area. After a transmission delay Td, 31 output audio frames are newly obtained at the reference point and 31 candidate audio features are calculated and sequentially placed in the second feature storage area; therefore, when the receiving point obtains the first audio feature G(100) corresponding to the input audio frame Q', 31 candidate audio features from F(100) to F(131) are newly stored in the second feature storage area; due to the existence of the target delay, when G(100) moves from the position where G(130) is located in the figure to the head of the first feature storage area, feature matching is performed based on the first audio feature G(100) and the candidate audio features stored in the second feature storage area; and after the target delay, 30 candidate audio features from F(131) to F(161) are newly stored in the second feature storage area; then the echo delay at this time is the length of 61 output audio frames from F(100) to F(161), and the echo delay can be calculated based on the sampling interval of each output audio frame and the number of audio frames accordingly.
[0123] In a possible implementation, based on Figure 9 the delayed feature matching process shown, the echo delay of the output audio frame corresponding to the second audio feature can be determined based on the target storage position of the second audio feature in the second feature storage area.
[0124] Schematically, as Figure 10 shown, the first audio feature G(80) matches F(80) in the second feature storage area, and F(80) corresponds to Sm(14) in the feature matching score. Then the echo delay is the length of 61 audio frames from Sm(14) to Sm(75); if each audio frame reaches the next audio frame after 4 ms, the echo delay is 61×4 ms.
[0125] Step 508, in response to the echo delay being greater than the target delay, perform echo cancellation processing on the input audio frame based on the echo delay estimated based on the echo delay.
[0126] In a possible implementation, if the echo delay is greater than the target delay, it means that no negative delay occurs, and echo cancellation processing can be performed on the input audio frame based on the echo delay obtained from the original delay estimation process.
[0127] Step 509, in response to the echo delay being less than the target delay, determine that there is an abnormal echo delay.
[0128] For the implementation manner of step 509, reference can be made to the above embodiments, and details are not described herein again in this embodiment.
[0129] Step 510, in response to the existence of an abnormal echo delay, re - estimate the echo delay.
[0130] In a possible implementation manner, if there is an abnormal echo delay, it is considered that the echo delay generated in the original delay estimation process is unreliable. Correspondingly, a reset instruction is generated to reset the delay estimation module and the delay alignment module, and their caches are cleared to eliminate the situation of negative delay, so that the delay estimation module can re - estimate the echo delay.
[0131] In this embodiment, by setting a first feature storage area for the first audio feature corresponding to the input audio frame, the function of feature delay matching for the first audio feature can be realized by storing the first audio feature in the first feature storage area. In addition, in the abnormal delay detection process, by adopting an audio feature extraction method different from the original delay estimation process, the accuracy of feature matching in the abnormal delay detection process can be improved, and then the accuracy of determining the echo delay in the abnormal delay detection process can be improved. Furthermore, the original delay estimation process can be guided whether to continue execution through a more accurate echo delay.
[0132] In another possible application scenario, in order to save computational effort and data storage space, it is set that both the abnormal detection module and the original delay estimation module adopt the same feature extraction method. Correspondingly, only one additional feature storage area is needed to store the first audio feature corresponding to the input audio frame.
[0133] Correspondingly Figure 6 , if the same feature extraction method is adopted, correspondingly, a separate feature extraction module may not be used in the abnormal delay detection process, which can reduce the computational effort generated by additional feature extraction. At the same time, the candidate audio feature corresponding to the reference point signal does not need to be stored in an additional new feature storage area, and only the feature storage area for storing the first audio feature needs to be added, which can reduce the storage space occupied by redundant audio feature storage.
[0134] On the Figure 6 basis, after deleting the feature extraction module 3, the feature extraction module 4, and the feature storage area 3, as Figure 11 shown, it shows a schematic diagram of the delay estimation process and the abnormal delay detection process shown in another exemplary embodiment of the present application. In Figure 11In this case, the candidate audio features stored in the feature storage area 1 can be used by the feature matching module 1 for feature matching to calculate the echo delay Tde generated in the delay estimation process 1101; they can also be used by the feature matching module 2 for feature matching to calculate the echo delay Tde2 generated in the abnormal delay detection process 1102. The first audio feature stored in the feature storage area 2 can be immediately used by the feature matching module 1 for feature matching to estimate the echo delay Tde when the first audio feature is obtained; it can also be used by the feature matching module 2 for feature matching to estimate the echo delay Tde2 after reaching the target delay.
[0135] Schematically, in the delay estimation process 1101, the feature extraction module 1 extracts audio features from the reference point signal (output audio frame) and stores the extracted candidate audio features in the feature storage area 1. After the received point signal is obtained, the feature extraction module 2 extracts audio features from the received point signal (input audio frame) and sends the extracted first audio feature to the feature matching module 1 for feature matching. Specifically, a matching second audio feature is found from the feature storage area 1 based on the first audio feature, and then the transmission delay Tde is determined based on the delay determination strategy 1.
[0136] In the abnormal delay detection process 1102, after the first audio feature corresponding to the input audio frame is obtained, the first audio feature is stored in the feature storage area 2. Then, after the target delay, the feature matching module 2 reads the first audio feature from the feature storage area 2, finds a second audio feature matching it from the feature storage area 1 based on this first audio feature, and then calculates the echo delay Tde2 through the delay determination strategy 2. At the abnormal delay detection, it is judged whether the echo delay Tde2 is an abnormal echo delay. If so, each module in the delay estimation process is reset.
[0137] From Figure 11 it can be seen that the same feature extraction method is adopted in both the delay estimation process and the abnormal delay detection process. Figure 12The figure shows a schematic diagram of the process of audio feature extraction in the delay estimation process shown in an exemplary embodiment of the present application. The audio frames at the reference point or the receiving point are converted to the frequency domain through downsampling and time-frequency domain conversion algorithms (such as short-time Fourier transform) to obtain the spectral signals of the audio frames; the spectral signals are divided into frequency bands, such as linear division or Mel frequency band division according to psychoacoustic theory, and the spectrum is divided into M sub-bands (M is 32), because 32 bits can just be stored in a 32-bit integer variable, which is convenient for improving the calculation efficiency. M can also be 16, 64 or other quantities convenient for storage and calculation. Calculate the sub-band energy E(m) for each sub-band, and calculate the smoothed sub-band energy 1 of each sub-band after smoothing through the smoothed sub-band energy module. The smoothed sub-band energy Ep(m) = E(m)*a + Ep(m)'*(1 - a), where a is a decimal between 0 and 1, and the smaller a is, the higher the degree of smoothing. The comparator compares the sub-band energy with the smoothed sub-band energy. If the sub-band energy is greater than the smoothed sub-band energy, it outputs the binary value 1, otherwise it outputs the binary value 0. Assume that the current frame is the nth frame. The binary output of the first sub-band is stored in the 0th bit of the nth frame of F, that is, F(n, 0). Similarly, the binary output of the second sub-band is stored in the 1st bit of the nth frame of F, that is, F(n, 1), and so on, to obtain M binary values, which are the extracted audio features.
[0138] In another possible application scenario, in order to further reduce the amount of calculation in the feature matching process, the second storage capacity of the second feature storage area can be smaller than the first storage capacity corresponding to the first feature storage area. In this setting, if an output audio frame matching the input audio frame can be found in the second feature storage area, it indicates that a negative delay situation has occurred.
[0139] On the basis of Figure 3 As shown in Figure 13 Steps 302 to 304 can be replaced by steps 1301 and 1302.
[0140] Step 1301, in response to reaching the target delay, search for candidate audio features matching the first audio feature in the second feature storage area.
[0141] Among them, the first storage capacity of the first feature storage area is determined by the target delay and the sampling time interval between adjacent input audio frames. Then, the number of frames of the input audio frames corresponding to the audio features stored in the second feature storage area is less than the number of frames of the output audio frames corresponding to the audio features stored in the first feature storage area. Schematically, if the first feature storage area can be used to store the audio features of 31 frames of audio frames, the second feature storage area can be set to be able to store the audio features of 30 frames of audio frames.
[0142] In a possible implementation, after the first audio feature is extracted and the target delay is reached, the first audio feature is used to search in the second feature storage area, that is, the first audio feature is matched with each candidate audio feature stored in the second feature storage area, and the matching scores corresponding to the first audio feature and each candidate audio feature are obtained.
[0143] It should be noted that when the second storage capacity of the second feature storage area is smaller than that of the first feature storage area, the second feature storage area cannot be shared in the abnormal delay detection process and the delay estimation process. That is to say, even if the same audio feature extraction method is used in the abnormal delay detection process and the delay estimation process, the candidate audio features to be used in the delay estimation module need to be stored in other feature storage areas.
[0144] Step 1302, in response to finding a candidate audio feature that matches the first audio feature, it is determined that there is an abnormal echo delay.
[0145] Since the abnormal delay detection process only needs to detect the case where the echo delay is less than the target delay, when the second storage capacity of the second feature storage area is smaller than the first storage capacity, normally, when performing feature matching based on the first audio feature, there is no candidate audio feature in the second storage area that matches it. On the contrary, if a matching candidate audio feature can be found in the second feature storage area, a value of the echo delay less than the target delay can be obtained. Therefore, it can be explained that a negative delay situation has occurred. Therefore, in a possible implementation, when a candidate audio feature that matches the first audio feature is found in the second storage area, it is determined that there is an abnormal echo delay.
[0146] In this embodiment, by setting the storage capacity of the second feature storage area, it is possible to determine whether there is a candidate audio feature in the second feature storage area that matches the first audio feature without calculating the echo delay. If it exists, it is determined that there is an abnormal echo delay, which can simplify the detection logic of the abnormal delay and further improve the detection efficiency of the abnormal echo delay. In addition, reducing the storage capacity of the second feature storage area can also reduce the number of feature matches and is recommended not to reduce the detection efficiency of the abnormal echo delay.
[0147] The following is an apparatus embodiment of the present application. For details not described in detail in the apparatus embodiment, reference may be made to the above method embodiment.
[0148] Figure 14 It is a structural block diagram of an abnormal echo delay recognition apparatus provided by an exemplary embodiment of the present application. The apparatus includes:
[0149] A feature extraction module 1401, configured to extract audio features from the input audio frames collected by the microphone to obtain the first audio features;
[0150] A first determination module 1402, configured to, in response to reaching a target delay, determine a second audio feature from candidate audio features based on the first audio feature, where the candidate audio features are audio features corresponding to output audio frames for speaker playback, and the second audio feature matches the first audio feature;
[0151] A second determination module 1403, configured to determine an echo delay of an output audio frame corresponding to the second audio feature;
[0152] A third determination module 1404, configured to determine that there is an abnormal echo delay in response to the echo delay being less than the target delay.
[0153] Optionally, the first determination module 1402 includes:
[0154] A storage unit, configured to store the first audio feature at a tail storage position in a first feature storage area, where a first storage capacity of the first feature storage area is determined by the target delay, and a storage position of the first audio feature moves from the tail storage position to a head storage position over time;
[0155] A first determination unit, configured to, in response to the first audio feature moving to the head storage position of the first feature storage area, determine the second audio feature from the candidate audio features based on the first audio feature.
[0156] Optionally, the candidate audio features are stored in a second feature storage area, and a storage position of the candidate audio features moves from a tail storage position to a head storage position over time;
[0157] The second determination module 1403 includes:
[0158] An acquisition unit, configured to acquire a target storage position of the second audio feature in the second feature storage area;
[0159] A second determination unit, configured to determine a target echo delay of an output audio frame corresponding to the second audio feature based on the target storage position.
[0160] Optionally, a second storage capacity of the second feature storage area is less than the first storage capacity;
[0161] The apparatus further includes:
[0162] A search module, configured to, in response to reaching the target delay, search for candidate audio features matching the first audio feature in the second feature storage area;
[0163] A fourth determination module, configured to determine that there is an abnormal echo delay in response to finding a candidate audio feature that matches the first audio feature.
[0164] Optionally, the first storage capacity is determined by the target delay and the sampling time interval between adjacent input audio frames, and the number of frames of the input audio frames corresponding to the audio features stored in the second feature storage area is less than the number of frames of the output audio frames corresponding to the audio features stored in the first feature storage area.
[0165] Optionally, the feature extraction module 1401 includes:
[0166] A third determination unit, configured to perform time-frequency conversion and frequency band division on the input audio frame collected by the microphone to determine M sub-bands;
[0167] A fourth determination unit, configured to obtain N first frequency domain feature scores by performing frequency domain energy comparison on the sub-band energies of the M sub-bands, where N is a positive integer, and M - N is a positive integer;
[0168] A fifth determination unit, configured to determine the set of N first frequency domain feature scores as the first audio feature.
[0169] Optionally, the fourth determination unit is further configured to:
[0170] In response to the energy of the j-th sub-band being the maximum value among the sub-band energies corresponding to the (j - i)-th sub-band to the (j + i)-th sub-band, determine the first score as the first frequency domain feature score corresponding to the j-th sub-band, where the energy of the j-th sub-band is the sub-band energy corresponding to the j-th sub-band, where i is a positive integer, and j - i is a positive integer, and j + i is less than or equal to M;
[0171] In response to the energy of the j-th sub-band being less than the maximum value among the sub-band energies corresponding to the (j - i)-th sub-band to the (j + i)-th sub-band, determine the second score as the first frequency domain feature score corresponding to the j-th sub-band.
[0172] Optionally, the first determination module 1402 includes:
[0173] A feature matching unit, configured to perform feature matching on the first audio feature and the candidate audio feature to obtain at least one candidate matching score, where the candidate matching score is used to indicate the matching degree between the first audio feature and the candidate audio feature, and the candidate matching score has a negative correlation with the similarity;
[0174] A sixth determination unit, configured to determine the second audio feature from the candidate audio features based on at least one of the candidate matching scores.
[0175] Optionally, the first audio feature includes N first frequency-domain feature scores, and the candidate audio feature includes N second frequency-domain feature scores, where N is a positive integer;
[0176] The feature matching unit is further configured to:
[0177] Perform a matching operation on the k-th first frequency-domain feature score and the k-th second frequency-domain feature score to obtain the k-th sub-matching score, where k is a positive integer less than or equal to N, and the sub-matching score is the matching degree of the k-th frequency-domain feature score in the first audio feature and the second audio feature;
[0178] Determine the sum of the N sub-matching scores as the candidate matching score.
[0179] Optionally, the fifth determining unit is further configured to:
[0180] Perform smoothing processing on at least one of the candidate matching scores to obtain at least one feature matching score;
[0181] Determine the minimum value among the feature matching scores as the target matching score;
[0182] Determine the candidate audio feature corresponding to the target matching score as the second audio feature.
[0183] Optionally, the apparatus further includes:
[0184] A reset module, configured to re-perform echo delay estimation in response to the existence of abnormal echo delay.
[0185] Optionally, the apparatus further includes:
[0186] An echo cancellation module, configured to perform echo cancellation processing on the input audio frame based on the echo delay obtained by echo delay estimation in response to the echo delay being greater than the target delay.
[0187] In summary, in the embodiments of the present application, in view of the situation of negative delay occurring during echo cancellation, which causes a characteristic that after receiving an input audio frame at the receiving point, it is necessary to wait for a period of time to find a matching output audio frame, a method for detecting negative delay during echo cancellation is proposed. After obtaining the first audio feature corresponding to the input audio frame and delaying it by the target delay, feature matching is performed based on the first audio feature, so that the echo delay can still be estimated in the case of negative delay; and the delayed feature matching makes the calculated echo delay the sum of the target delay and the propagation delay. Then, based on the relationship between the echo delay and the target delay, the positive or negative nature of the propagation delay can be determined, so as to timely determine whether there is an abnormal echo delay (negative delay), avoid the echo cancellation module from continuing to perform echo cancellation work under the wrong echo delay, or avoid the situation where echo cancellation cannot be performed due to the inability to calculate the echo delay, thereby improving the accuracy of echo cancellation.
[0188] Figure 15 FIG. shows a structural block diagram of a terminal 1500 provided by an exemplary embodiment of the present application. The terminal 1500 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 1500 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0189] Generally, the terminal 1500 includes: a processor 1501 and a memory 1502.
[0190] The processor 1501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1501 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0191] The memory 1502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1502 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1501 to implement the information processing method provided in the method embodiments of the present application.
[0192] In some embodiments, the terminal 1500 may further optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502, and the peripheral device interface 1503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1504, a display screen 1505, a camera module 1506, an audio circuit 1507, and a power supply 1509.
[0193] The peripheral device interface 1503 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0194] The radio frequency circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1504 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1504 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1504 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1504 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0195] The display screen 1505 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1505 is a touch display screen, the display screen 1505 also has the ability to collect touch signals on or above the surface of the display screen 1505. The touch signals can be input as control signals to the processor 1501 for processing. At this time, the display screen 1505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 1505, which is provided on the front panel of the terminal 1500; in other embodiments, there may be at least two display screens 1505, which are respectively provided on different surfaces of the terminal 1500 or are in a foldable design; in still other embodiments, the display screen 1505 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 1500. Even further, the display screen 1505 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1505 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0196] The camera module 1506 is used to capture images or videos. Optionally, the camera module 1506 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to achieve functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fused shooting functions. In some embodiments, the camera module 1506 may further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0197] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals which are input to the processor 1501 for processing, or input to the radio frequency circuit 1504 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1500. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1501 or the radio frequency circuit 1504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1507 may further include a headphone jack.
[0198] The power supply 1509 is used to supply power to each component in the terminal 1500. The power supply 1509 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1509 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0199] In some embodiments, the terminal 1500 further includes one or more sensors 1510. The one or more sensors 150 include but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, an optical sensor 1515, and a proximity sensor 1516.
[0200] The acceleration sensor 1511 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 1500. For example, the acceleration sensor 1511 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1501 can control the touch display screen 1505 to process information in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1511. The acceleration sensor 1511 can also be used for the collection of game or user's motion data.
[0201] The gyroscope sensor 1512 can detect the body direction and rotation angle of the terminal 1500. The gyroscope sensor 1512 can cooperate with the acceleration sensor 1511 to collect the 3D actions of the user on the terminal 1500. According to the data collected by the gyroscope sensor 1512, the processor 1501 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0202] The pressure sensor 1513 can be disposed on the side frame of the terminal 1500 and / or the lower layer of the touch display screen 1505. When the pressure sensor 1513 is disposed on the side frame of the terminal 1500, it can detect the holding signal of the user on the terminal 1500, and the processor 1501 performs left / right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the touch display screen 1505, the processor 1501 controls the operable controls on the UI interface according to the pressure operation of the user on the touch display screen 1505. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0203] The optical sensor 1515 is used to collect the ambient light intensity. In one embodiment, the processor 1501 can control the display brightness of the touch display screen 1505 according to the ambient light intensity collected by the optical sensor 1515. Specifically, when the ambient light intensity is high, the display brightness of the touch display screen 1505 is increased; when the ambient light intensity is low, the display brightness of the touch display screen 1505 is decreased. In another embodiment, the processor 1501 can also dynamically adjust the shooting parameters of the camera module 1506 according to the ambient light intensity collected by the optical sensor 1515.
[0204] The proximity sensor 1516, also known as the distance sensor, is usually disposed on the front panel of the terminal 1500. The proximity sensor 1516 is used to collect the distance between the user and the front of the terminal 1500. In one embodiment, when the proximity sensor 1516 detects that the distance between the user and the front of the terminal 1500 is gradually decreasing, the processor 1501 controls the touch display screen 1505 to switch from the lit screen state to the off screen state; when the proximity sensor 1516 detects that the distance between the user and the front of the terminal 1500 is gradually increasing, the processor 1501 controls the touch display screen 1505 to switch from the off screen state to the lit screen state.
[0205] Those skilled in the art can understand that Figure 15 the structure shown in does not constitute a limitation on the terminal 1500, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0206] This application also provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, a code set or an instruction set is stored, and the at least one instruction, the at least one segment of program, the code set or the instruction set is loaded and executed by a processor to implement the abnormal echo delay recognition method provided by any of the above exemplary embodiments.
[0207] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a terminal reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the terminal executes the abnormal echo delay recognition method provided in the above optional implementation manner.
[0208] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0209] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An abnormal echo delay recognition method, characterized in that, The method includes: Performing audio feature extraction on the input audio frames collected by the microphone to obtain first audio features; Storing the first audio features at the tail storage position of the first feature storage area, where the number of frames corresponding to the first storage capacity of the first feature storage area is the target delay divided by the sampling time interval between adjacent input audio frames plus 1, and the storage position of the first audio features moves from the tail storage position to the head storage position over time; In response to the first audio features moving to the head storage position of the first feature storage area, determining second audio features from candidate audio features based on the first audio features, where the candidate audio features are the audio features corresponding to output audio frames for speaker playback, and the second audio features match the first audio features; Determining the echo delay of the output audio frame corresponding to the second audio features; In response to the echo delay being less than the target delay, determining that there is an abnormal echo delay.
2. The method according to claim 1, wherein The candidate audio features are stored in the second feature storage area, and the storage position of the candidate audio features moves from the tail storage position to the head storage position over time; The determining the echo delay of the output audio frame corresponding to the second audio features includes: Obtaining the target storage position of the second audio features in the second feature storage area; Based on the target storage position, determining the target echo delay of the output audio frame corresponding to the second audio features.
3. The method according to claim 2, wherein The second storage capacity of the second feature storage area is less than the first storage capacity; After performing audio feature extraction on the input audio frames collected by the microphone to obtain first audio features, the method further includes: In response to reaching the target delay, searching in the second feature storage area for candidate audio features that match the first audio features; In response to finding candidate audio features that match the first audio features, determining that there is an abnormal echo delay.
4. The method according to claim 3, wherein The number of frames of the input audio frames corresponding to the audio features stored in the second feature storage area is less than the number of frames of the output audio frames corresponding to the audio features stored in the first feature storage area.
5. The method according to any one of claims 1 to 4, characterized in that, The performing audio feature extraction on the input audio frames collected by the microphone to obtain first audio features includes: Performing time-frequency conversion and frequency band division on the input audio frames collected by the microphone to determine M sub-bands; Obtaining N first frequency domain feature scores by performing frequency domain energy comparison on the sub-band energies of the M sub-bands, where N is a positive integer and M - N is a positive integer; Determining the set of N first frequency domain feature scores as the first audio features.
6. The method according to claim 5, characterized in that, The obtaining N first frequency domain feature scores by performing frequency domain energy comparison on the sub-band energies of the M sub-bands includes: In response to the energy of the j-th sub-band being the maximum among the sub-band energies corresponding to the (j - i)-th sub-band to the (j + i)-th sub-band, determining the first score as the first frequency domain feature score corresponding to the j-th sub-band, where the energy of the j-th sub-band is the sub-band energy corresponding to the j-th sub-band, where i is a positive integer, j - i is a positive integer, and j + i is less than or equal to M; In response to the energy of the j-th sub-band being less than the maximum value of the corresponding sub-band energies of the (j - i)-th to (j + i)-th sub-bands, the second score value is determined as the first frequency-domain feature score value corresponding to the j-th sub-band.
7. The method according to any one of claims 1 to 4, characterized in that, The determining the second audio feature from the candidate audio features based on the first audio feature includes: Performing feature matching on the first audio feature and the candidate audio features to obtain at least one candidate matching score value, where the candidate matching score value is used to indicate the degree of matching between the first audio feature and the candidate audio features, and the candidate matching score value has a negative correlation with the degree of matching; Determining the second audio feature from the candidate audio features based on at least one of the candidate matching score values.
8. The method according to claim 7, wherein The first audio feature includes N first frequency-domain feature score values, and the candidate audio features include N second frequency-domain feature score values, where N is a positive integer; The performing feature matching on the first audio feature and the candidate audio features to obtain at least one candidate matching score value includes: Performing a matching operation on the k-th first frequency-domain feature score value and the k-th second frequency-domain feature score value to obtain the k-th sub-matching score value, where k is a positive integer less than or equal to N, and the sub-matching score value is the degree of matching of the k-th frequency-domain feature score value in the first audio feature and the second audio feature; Determining the sum of the N sub-matching score values as the candidate matching score value.
9. The method according to claim 7, wherein The determining the second audio feature from the candidate audio features based on at least one of the candidate matching score values includes: Performing smoothing processing on at least one of the candidate matching score values to obtain at least one feature matching score value; Determining the minimum value among the feature matching score values as the target matching score value; Determining the candidate audio feature corresponding to the target matching score value as the second audio feature.
10. The method according to any one of claims 1 to 4, characterized in that, After determining that there is an abnormal echo delay in response to the echo delay being less than the target delay, the method further includes: In response to the existence of an abnormal echo delay, re-performing echo delay estimation.
11. According to the method described in any one of claims 1 to 4, characterized in that, After determining the echo delay of the output audio frame corresponding to the second audio feature, the method further includes: In response to the echo delay being greater than the target delay, performing echo cancellation processing on the input audio frame based on the echo delay obtained from echo delay estimation.
12. An abnormal echo delay recognition device, characterized in that, The apparatus includes: A feature extraction module, configured to perform audio feature extraction on an input audio frame collected by a microphone to obtain a first audio feature; A first determination module, configured to, in response to reaching the target delay, determine a second audio feature from candidate audio features based on the first audio feature, where the candidate audio features are audio features corresponding to an output audio frame for playing by a speaker, and the second audio feature matches the first audio feature; A second determination module, configured to determine the echo delay of the output audio frame corresponding to the second audio feature; A third determination module, configured to determine the existence of an abnormal echo delay in response to the echo delay being less than the target delay; The first determination module includes: A storage unit for storing the first audio feature at a tail storage position in a first feature storage area, where the number of frames corresponding to a first storage capacity of the first feature storage area is obtained by dividing a target delay by a sampling time interval between adjacent input audio frames and then adding 1, and the storage position of the first audio feature moves from the tail storage position to a head storage position over time; A first determination unit for, in response to the first audio feature moving to the head storage position of the first feature storage area, determining the second audio feature from the candidate audio features based on the first audio feature.
13. A terminal, characterized in that, The terminal includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the abnormal echo delay identification method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, At least one program is stored in the readable storage medium, and the at least one program is loaded and executed by a processor to implement the abnormal echo delay identification method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Echo signal processing method and apparatus thereof
CN102223456A
Negative delay time detection method and device, electronic equipment and storage medium
CN111736797A