Audio echo processing method and device, equipment, storage medium and program product

By estimating the delay of microphone and speaker signals and processing the reference signal, combined with linear and nonlinear filtering, the problem of balancing computational accuracy and efficiency in existing technologies is solved, achieving a highly efficient echo cancellation effect.

CN120833795APending Publication Date: 2025-10-24BIGO TECH PTE LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511014886.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing audio echo processing algorithms cannot balance computational accuracy and efficiency, resulting in poor echo cancellation performance.

Method used

By estimating the delay between the microphone signal and the speaker signal to be played, a matching reference signal is determined, and echo cancellation is performed based on the reference signal. This includes preliminary delay estimation and precise delay estimation under stable conditions. Combined with linear and nonlinear filtering, signal alignment and echo cancellation are achieved.

Benefits of technology

It balances computational accuracy and efficiency, improves the real-time performance of echo processing and the echo cancellation effect, and ensures audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833795A_ABST
    Figure CN120833795A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio echo processing method and device, equipment, a storage medium and a program product, and the method comprises the steps: obtaining a microphone collection signal and a loudspeaker to-be-played signal, and carrying out the first time delay estimation based on the microphone collection signal and the loudspeaker to-be-played signal, and obtaining a first time delay value; under the condition that the first time delay value meets a preset time delay stability condition, determining a first reference signal according to the first time delay value and a loudspeaker to-be-played signal, performing second time delay estimation based on the first reference signal and a microphone acquisition signal to obtain a second time delay value, and determining a second reference signal according to the second time delay value and the first reference signal, the second delay value is smaller than the first delay value; and performing echo cancellation processing on the microphone acquisition signal based on the second reference signal to obtain a first near-end voice signal. According to the scheme, the calculation precision and efficiency can be balanced, the calculation precision is improved, the calculation amount is reduced, real-time and efficient echo processing is ensured, and the echo cancellation effect is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of audio processing, and in particular to an audio echo processing method and device, equipment, a storage medium and a program product. BACKGROUND

[0002] In a voice scene such as network live microphone connection and game voice interaction, users are in a full-duplex communication state. In the full-duplex communication process, when a user terminal device is in an external playing state, the sound played by the loudspeaker will be collected by the microphone to form an acoustic echo. Since the acoustic echo will seriously affect the communication quality, an algorithm for acoustic echo cancellation needs to be used to suppress the acoustic echo to improve the communication quality.

[0003] In related technologies, an algorithm based on feature correlation or an algorithm based on adaptive filtering is used for echo delay estimation to eliminate echo interference. However, the algorithm based on feature correlation is implemented in the frequency domain, and is affected by the precision of Fourier transform, which leads to low calculation precision. Moreover, the algorithm based on adaptive filtering has high calculation precision, but has a large amount of calculation and low real-time efficiency. Therefore, the foregoing algorithms cannot balance calculation precision and efficiency, and the echo cancellation effect is poor. SUMMARY

[0004] Embodiments of the present application provide an audio echo processing method, device, equipment, a storage medium and a program product. A matching reference signal is determined by performing delay estimation on a microphone collected signal and a loudspeaker to-be-played signal, and a near-end voice signal is obtained by performing echo cancellation processing on the microphone collected signal based on the reference signal, so as to solve the problem that related technologies cannot balance precision and efficiency, and the echo cancellation effect is poor. The calculation precision and efficiency can be balanced, the calculation precision is improved while the calculation amount is reduced, the echo processing is real-time and efficient, and the echo cancellation effect is ensured.

[0005] In a first aspect, embodiments of the present application provide an audio echo processing method, which comprises: obtaining a microphone collected signal and a loudspeaker to-be-played signal, and performing first delay estimation based on the microphone collected signal and the loudspeaker to-be-played signal to obtain a first delay value; in a case where the first delay value meets a preset delay stability condition, determining a first reference signal according to the first delay value and the loudspeaker to-be-played signal, performing second delay estimation based on the first reference signal and the microphone collected signal to obtain a second delay value, determining a second reference signal according to the second delay value and the first reference signal, and the second delay value is smaller than the first delay value; performing echo cancellation processing on the microphone collected signal based on the second reference signal to obtain a first near-end voice signal.

[0006] In a second aspect, an audio echo processing apparatus is provided, comprising: a signal obtaining module configured to obtain a microphone collected signal and a speaker to be played signal; a delay estimation module configured to perform first delay estimation based on the microphone collected signal and the speaker to be played signal to obtain a first delay value; a first reference signal determination module configured to determine a first reference signal according to the first delay value and the speaker to be played signal in a case that the first delay value satisfies a preset delay stability condition; the delay estimation module is further configured to perform second delay estimation based on the first reference signal and the microphone collected signal to obtain a second delay value, the second delay value being smaller than the first delay value; the first reference signal determination module is further configured to determine a second reference signal according to the second delay value and the first reference signal; a first signal processing module configured to perform echo cancellation processing on the microphone collected signal based on the second reference signal to obtain a first near-end speech signal.

[0007] In a third aspect, an audio echo processing device is provided, comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the audio echo processing method provided by the embodiments of the present application.

[0008] In a fourth aspect, a non-volatile storage medium storing computer executable instructions is provided, the computer executable instructions, when executed by a computer processor, are configured to perform the audio echo processing method provided by the embodiments of the present application.

[0009] In a fifth aspect, a computer program product is provided, the computer program product comprising a computer program stored in a computer readable storage medium, at least one processor of a device reading and executing the computer program from the computer readable storage medium, so that the device performs the audio echo processing method provided by the embodiments of the present application.

[0010] In the embodiment of the present application, the first delay value is obtained by performing first delay estimation based on the microphone collected signal and the loudspeaker to-be-played signal, which can preliminarily perform signal delay estimation, complete signal coarse alignment, and improve the processing efficiency of subsequent signal fine alignment; in the case that the first delay value meets the preset delay stability condition, the second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone collected signal, which can avoid single estimation error, complete signal fine alignment, and ensure the accuracy of delay estimation; the first near-end speech signal is obtained by performing echo cancellation processing on the microphone collected signal based on the second reference signal, which can eliminate echo and improve audio quality. The above scheme can balance the calculation accuracy and efficiency, improve the calculation accuracy, reduce the calculation amount, ensure real-time and efficient echo processing, and guarantee the echo cancellation effect. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 2 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 3 A structural schematic diagram of a nonlinear network model provided by an embodiment of the present application; Figure 4 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 5 A structural schematic diagram of a feature alignment network provided by an embodiment of the present application; Figure 6 A process schematic diagram of an audio processing method provided by an embodiment of the present application; Figure 7 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 8 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 9 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 10 A flowchart of an audio echo processing method provided by an embodiment of the present application; Figure 11 A structural block diagram of an audio echo processing device provided by an embodiment of the present application; Figure 12A structural schematic diagram of an audio echo processing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0012] The present application will be further described below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings, rather than all the structures.

[0013] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class, and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents a "or" relationship between the front and rear associated objects.

[0014] The audio echo processing method provided by the embodiments of the present application can determine the matching reference signal by performing delay estimation on the microphone acquisition signal and the loudspeaker to-be-played signal, and can obtain the near-end speech signal by performing echo cancellation processing on the microphone acquisition signal based on the reference signal, so as to balance the calculation precision and efficiency and guarantee the echo cancellation effect. The relevant application scenarios can include network live broadcast, video call, game voice, etc. The foregoing several application scenarios are only exemplary and explanatory, and in actual application, the audio echo processing method can also be used in other scenarios, which is not limited by the embodiments of the present application.

[0015] The execution subject of each step of the audio echo processing method provided by the embodiments of the present application can be a computer device, which refers to any electronic device with data calculation, processing and storage capabilities, such as mobile phones, PC (Personal Computer), tablet computers and other terminal devices, which is not limited by the embodiments of the present application.

[0016] Figure 1 A flowchart of an audio echo processing method provided by an embodiment of the present application is shown in Figure 1 The audio echo processing method specifically includes the following steps: Step S101, obtaining a microphone acquisition signal and a loudspeaker to-be-played signal, and performing first delay estimation based on the microphone acquisition signal and the loudspeaker to-be-played signal to obtain a first delay value.

[0017] The microphone acquisition signal can be a real-time audio signal collected by the microphone, and can include near-end human voice, environmental noise, and far-end echo generated after the loudspeaker plays. The loudspeaker-to-be-played signal can be an audio signal that needs to be played through the loudspeaker, and is the main source of echo. Since there is a delay from the loudspeaker playing to the microphone recording, and the loudspeaker-to-be-played signal is obtained in advance, the microphone acquisition signal is delayed relative to the loudspeaker-to-be-played signal, and correspondingly, the loudspeaker-to-be-played signal is advanced relative to the microphone acquisition signal. Therefore, by estimating the delay between the loudspeaker-to-be-played signal and the microphone acquisition signal, a delay value is obtained, based on which the loudspeaker-to-be-played signal can be shifted to obtain a reference signal that is aligned with the microphone acquisition signal on the time axis, and subsequent acoustic echo cancellation of the microphone acquisition signal is realized. The first delay estimation can be a delay estimation algorithm based on frequency domain correlation. By performing Fourier transform or fast Fourier transform on the microphone acquisition signal and the loudspeaker-to-be-played signal, and then performing feature correlation calculation in the frequency domain, a corresponding first delay value can be obtained. The first delay value represents the delay duration of the microphone acquisition signal relative to the loudspeaker-to-be-played signal, and due to the influence of the accuracy of the frequency domain Fourier transform, it can be regarded as a preliminary estimation result with relatively low accuracy.

[0018] In step S102, if the first delay value meets the preset delay stability condition, a first reference signal is determined according to the first delay value and the loudspeaker-to-be-played signal, a second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone acquisition signal, and a second reference signal is determined according to the second delay value and the first reference signal, wherein the second delay value is less than the first delay value.

[0019] The first delay estimation is performed at a preset time interval to obtain a corresponding first delay value, for example, once every 1s, to avoid frequent execution and low processing efficiency. Since the audio signal may have delay fluctuations, the first delay value needs to be judged for delay stability judgment to avoid single estimation error and to ensure the reliability of the first delay value. After starting the first delay estimation, the first delay value obtained first can be compared with the range of the historical delay value estimated recently to determine whether the delay has changed. For example, the range of the historical delay value estimated recently is 90ms to 110ms, and if the detected first delay value is 200ms, it can be considered that the delay may have changed, and the subsequent first delay value needs to be judged whether it deviates from 90ms to 110ms. Optionally, a preset delay change condition can be set, which can be that the first delay value deviates from the range of the historical delay value is detected for a preset number of times in succession, for example, the first delay value deviates from the range of the historical delay value is detected for two times in succession, and then it is considered that the delay has changed, otherwise the historical delay calculation result is continued to be used for subsequent echo cancellation processing. Further, after judging that the delay has changed, the range of the delay value can be reset based on the latest first delay value to determine whether the delay is stable, for example, the latest detected first delay value is 200ms, and the range of the delay value can be set to 190ms to 210ms with a floating interval of 20ms. Of course, the specific range setting method is only an exemplary description, and can be adaptively adjusted according to the actual application scenario. The preset delay stability condition can be that the first delay value is detected within the set range of the delay value for a preset number of times in succession, for example, the first delay value is detected within 190ms to 210ms for three times in succession.

[0020] In one embodiment, the first reference signal is determined according to the first delay value and the speaker-to-be-played signal, specifically, the speaker-to-be-played signal is translated by the first delay value to obtain the first reference signal. In one embodiment, since there is a detection lag in the call initiation stage due to the presence of near-end components and nonlinear interference, etc., the signal alignment needs to control the reference signal within a preset lead range to ensure the subsequent echo cancellation effect. Therefore, the first reference signal is determined according to the first delay value and the speaker-to-be-played signal, specifically, a target translation position is determined according to the preset lead range, for example, the center of the preset lead range or a floating value above and below the center by a preset value, for example, the preset lead range can be 4ms to 40ms, and the corresponding target translation position can be 24ms, which can be adaptively set according to the actual application scenario, which is not limited herein. Then, the translation amount is calculated based on the first delay value and the target translation position, for example, the first delay value is 200ms, and the translation amount is ms. Thus, the first reference signal can be obtained by shifting the loudspeaker to-be-played signal by 176 ms, so that the first reference signal leads the microphone acquisition signal by between 4 ms and 40 ms. If the first delay value satisfies a preset delay stability condition and the first reference signal is determined, the second delay estimation can be continued based on the first reference signal to accurately estimate the actual delay value.

[0021] In one embodiment, the second delay estimation can be to calculate a second delay value by using an adaptive filtering-based delay estimation algorithm, for example, a time-domain LMS (Least Mean Square) algorithm, a time-domain NLMS (Normalized Least Mean Square) algorithm, etc., which are not limited herein. Thus, the second reference signal can be determined according to the second delay value and the first reference signal, and specifically, the first reference signal can be shifted by the second delay value to obtain the second reference signal. In one embodiment, due to the detection hysteresis caused by the near-end component and nonlinear interference at the beginning of the call, the signal alignment needs to control the reference signal within a preset lead range to ensure the subsequent echo cancellation effect. Therefore, the second delay value calculated by the adaptive filtering-based delay estimation algorithm can not need to be used to continue shifting the first reference signal. However, in order to be more conducive to the linear filtering processing of the subsequent echo cancellation, the second delay value calculated by the adaptive filtering-based delay estimation algorithm can be converted into a multiple of a preset frame shift time length corresponding to the linear filtering processing, and a shift amount can be calculated, so that the second reference signal obtained by shifting the first reference signal based on the shift amount can be located at a position of the multiple of the preset frame shift time length. Thus, the second reference signal can be determined according to the second delay value and the first reference signal, and specifically, a target shift position can be calculated based on the second delay value and the preset frame shift time length, and the calculation formula is as follows:

[0022] wherein, is the target shift position, is the second delay value, is a floor function. For example, the second delay value is 24 ms, and the preset frame shift time length is 5 ms, so the target shift position is . Then, the shift amount can be calculated according to the target shift position and the second delay value, for example, the second delay value is 24 ms, and the target shift position is 20 ms, so the shift amount is . Thus, the first reference signal can be shifted by 4 ms to obtain the second reference signal, so that the second reference signal can be located at a position of a multiple of the preset frame shift time length.

[0023] Step S103: performing echo cancellation processing on the microphone collected signal based on the second reference signal to obtain a first near-end speech signal.

[0024] The echo cancellation processing can be linear echo cancellation processing, nonlinear echo cancellation processing, or a combination of linear echo cancellation processing and nonlinear echo cancellation processing. The linear echo cancellation processing can be linear filtering processing based on an adaptive filtering algorithm. The nonlinear echo cancellation processing can be nonlinear filtering processing based on a neural network model. The first near-end speech signal can be near-end human voice remaining after the microphone collected signal eliminates far-end echo, environmental noise, and the like generated after the loudspeaker plays.

[0025] The first delay value is obtained by performing first delay estimation based on the microphone collected signal and the loudspeaker to-be-played signal, which can preliminarily perform signal delay estimation and complete signal coarse alignment, thereby improving the processing efficiency of subsequent signal fine alignment. In the case where the first delay value meets the preset delay stability condition, the second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone collected signal, which can avoid single estimation error while completing signal fine alignment and ensuring the accuracy of delay estimation. The first near-end speech signal is obtained by performing echo cancellation processing on the microphone collected signal based on the second reference signal, which can eliminate echo and improve audio quality. The above scheme can balance the calculation accuracy and efficiency, improve the calculation accuracy while reducing the calculation amount, ensure real-time and efficient echo processing, and guarantee the echo cancellation effect.

[0026] Figure 2 A flowchart of an audio echo processing method including a process of echo cancellation processing provided by an embodiment of the present application is shown in FIG. 2. The audio echo processing method specifically includes the following steps. Figure 2 Step S201: obtaining a microphone collected signal and a loudspeaker to-be-played signal, and performing first delay estimation based on the microphone collected signal and the loudspeaker to-be-played signal to obtain a first delay value.

[0027] Step S202: in the case where the first delay value meets a preset delay stability condition, determining a first reference signal according to the first delay value and the loudspeaker to-be-played signal, performing second delay estimation based on the first reference signal and the microphone collected signal to obtain a second delay value, and determining a second reference signal according to the second delay value and the first reference signal, wherein the second delay value is less than the first delay value.

[0028] Step S203: performing linear filtering processing on the second reference signal and the microphone collected signal to obtain a first residual signal and a third reference signal that leads the first residual signal; and inputting the first residual signal and the third reference signal into a nonlinear network model to obtain a first near-end speech signal.

[0029] ​The linear filtering process can use an adaptive filter to cancel linear echo. The first residual signal can be a processing result obtained by canceling linear echo from the microphone collected signal based on the second reference signal, and the third reference signal can be obtained by extracting a speech block with the maximum energy in the linear filtering process and shifting it to a position that is ahead of the first residual signal by a preset time length. The preset time length can be 10 ms, which is not limited in the present application. By setting the third reference signal to be ahead of the first residual signal, the non-causal problem caused by delay jitter can be prevented. In an embodiment, the nonlinear network model can include a feature extraction network, a feature alignment network, and a feature processing network, and the specific processing process is as follows: the first residual signal and the third reference signal are input into the feature extraction network to obtain a first residual feature signal and a first reference feature signal; the first residual feature signal and the first reference feature signal are input into the feature alignment network to obtain a second reference feature signal; the first residual feature signal and the second reference feature signal are spliced and input into the feature processing network to obtain a first near-end speech signal. The feature extraction network can be used to extract key feature information of the first residual signal and the third reference signal, the feature alignment network can be used to align the first reference feature signal with the first residual feature signal to achieve better residual echo cancellation effect, and the feature processing network can be used to cancel residual echo from the first residual feature signal based on the second reference feature signal to obtain the first near-end speech signal.

[0030] In an embodiment, Figure 3 A structural diagram of a nonlinear network model provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the nonlinear network model includes a linear filtering unit 101, a feature extraction network 102, a feature alignment network 103, and a feature processing network 104. Figure 3As shown, the nonlinear network module includes a feature extraction network 301, a feature alignment network 302, and a feature processing network 303. The feature extraction network 301 includes a first calculation layer 3011, a second calculation layer 3012, a first convolution layer 3013, a second convolution layer 3014, a third convolution layer 3015, and a fourth convolution layer 3016. The first calculation layer 3011 and the second calculation layer 3012 are configured to respectively perform Fourier transform on the first residual signal and the third reference signal to convert the first residual signal and the third reference signal into frequency domain signals. The first residual signal sequentially passes through the first calculation layer 3011, the first convolution layer 3013, and the second convolution layer 3014 to obtain a first residual feature signal, and the third reference signal sequentially passes through the second calculation layer 3012, the third convolution layer 3015, and the fourth convolution layer 3016 to obtain a first reference feature signal. The first residual feature signal and the first reference feature signal pass through the feature alignment network 302 to obtain a second reference feature signal. Optionally, the feature alignment network 302 can be provided with a switching flag, which can be used to control the extraction time length of a preset time domain extraction layer in the feature alignment network 302. The greater the extraction time length, the greater the time range covered by the delay estimation. The first residual feature signal and the second reference feature signal are spliced and then pass through the feature processing network 303 to obtain a first near-end speech signal.

[0031] As described above, by performing linear filtering processing on the second reference signal and the microphone acquisition signal, linear echo can be eliminated. By inputting the first residual signal and the third reference signal into the nonlinear network model, the nonlinear fitting capability of the model can be used to suppress nonlinear echo, thereby ensuring the echo cancellation effect.

[0032] Figure 4 A flowchart of an audio echo processing method provided by an embodiment of the present application is shown in FIG. 4. Figure 4 As shown in FIG. 4, the audio echo processing method specifically includes the following steps. In step S401, a microphone acquisition signal and a loudspeaker to-be-played signal are acquired, and a first delay value is obtained by performing first delay estimation based on the microphone acquisition signal and the loudspeaker to-be-played signal.

[0033] In step S402, in a case where the first delay value satisfies a preset delay change condition, the extraction time length of a preset time domain extraction layer in the feature alignment network is reduced, and the signal alignment processing is performed by the feature alignment network based on the preset time domain extraction layer with the reduced extraction time length.

[0034] The preset delay change condition can be that the first delay value deviates from the historical delay value range for a preset number of times in succession, for example, the first delay value deviating from the historical delay value range for two times in succession can be regarded as the delay changing. If the first delay value satisfies the preset delay change condition, it can be regarded that the actual delay is captured. Before the first delay value satisfies the preset delay change condition, the preset time domain extraction layer of the feature alignment network in the nonlinear network model can adopt a larger extraction time length, for example, the extraction time length is 50, which corresponds to 500 ms, to have a larger visual field range for covering a larger range of delay echoes and quickly positioning the actual delay. After the first delay value satisfies the preset delay change condition, the preset time domain extraction layer of the feature alignment network in the nonlinear network model can adopt a smaller extraction time length, for example, the extraction time length is 5, which corresponds to 50 ms, to have a smaller visual field range, reduce the calculation amount, and guarantee the fine processing effect.

[0035] Figure 5 A structural schematic diagram of a feature alignment network provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, a reference signal is input into the feature alignment network. Figure 5 Figure 3 ​The provided nonlinear network model, after the feature extraction network processing is completed, the first residual feature signal and the first reference feature signal are input to the feature alignment network for alignment processing. The feature alignment network can include a fifth convolutional layer 501, a sixth convolutional layer 502, a seventh convolutional layer 503, a multiplication and addition operation layer 504, a first time domain extraction layer 505, a second time domain extraction layer 506, an activation function layer 507, and a weighted summation layer 508. Among them, the first residual feature signal is first input to the fifth convolutional layer 501 for two-dimensional convolution processing, and the first reference feature signal is first input to the sixth convolutional layer 502 for two-dimensional convolution processing, which functions to reduce the number of channels and reduce the amount of calculation. Then, the processing result of the sixth convolutional layer 502 is input to the first time domain extraction layer 505 for local vector extraction of the matching set extraction time length. The greater the extraction time length, the greater the time range covered by the delay estimation. The first time domain extraction layer 505 and the fifth convolutional layer 501 respectively input the processing results to the multiplication and addition operation layer 504 for multiplication and addition operation to simulate the calculation of time domain correlation. The multiplication and addition operation layer 504 inputs the processing result to the seventh convolutional layer 503 for channel dimension merging. The first time domain extraction layer 505 and the seventh convolutional layer 503 respectively input the processing results to the activation function layer 507 for processing, for example, using the Softmax function, the probability density curve of the delay estimation can be obtained. The first reference feature signal is also input to the second time domain extraction layer 506 for local vector extraction of the matching set extraction time length. The second time domain extraction layer 506 and the activation function layer 507 respectively input the processing results to the weighted summation layer 508 for weighted summation to obtain the second reference feature signal. Among them, the first time domain extraction layer 505 and the second time domain extraction layer 506 can access a switching flag for controlling the size of the extraction time length. For example, before the first delay value meets the preset delay change condition, the switching flag can be set to "0", and the corresponding extraction time length can be 50, corresponding to a time range of 500ms. After the first delay value meets the preset delay change condition, the switching flag can be set to "1" to reduce the extraction time length, for example, the corresponding extraction time length is adjusted from 50 to 5, corresponding to a time range of 50ms.

[0036] In step S403, when the first delay value meets the preset delay stability condition, the first reference signal is determined according to the first delay value and the loudspeaker to be played signal, the second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone collected signal, and the second reference signal is determined according to the second delay value and the first reference signal, wherein the second delay value is less than the first delay value.

[0037] Step S404: linearly filter the second reference signal and the microphone acquisition signal to obtain a first residual signal and a third reference signal that precedes the first residual signal; input the first residual signal and the third reference signal into a nonlinear network model to obtain a first near-end speech signal, wherein the nonlinear network model includes a feature alignment network.

[0038] In one embodiment, Figure 6 A schematic diagram of a process for performing audio processing using an audio echo processing method provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the microphone signal and the speaker signal to be broadcast are input into a coarse delay estimation module 601. The coarse delay estimation module 601 processes the microphone signal and the speaker signal to be broadcast to obtain a first reference signal, which is then input into a fine delay estimation module 602. The fine delay estimation module 602 processes the first reference signal and the microphone signal to obtain a second reference signal, which is then input into a linear filtering processing module 603. The linear filtering processing module 603 processes the second reference signal and the microphone signal to obtain a first residual signal and a third reference signal that precedes the first residual signal. The first residual signal and the third reference signal are then input into a nonlinear network model 604 for processing to obtain a first near-end speech signal. The nonlinear network model includes a feature alignment network 6041. The coarse delay estimation module 601 can transmit a corresponding switching flag to the feature alignment network 6041 based on whether the first delay value meets a preset delay change condition. The switching flag can control the extraction time length of a preset time domain extraction layer in the feature alignment network 6041. For example, the initial switching flag is "0". After the first delay value meets the preset delay change condition, the coarse delay estimation module 601 can change the switching flag to "1" and pass it to the feature alignment network 6041 to reduce the extraction time length of the preset time domain extraction layer.

[0039] As mentioned above, before the first delay value meets the preset delay change condition, the preset time domain extraction layer in the feature alignment network can maintain a longer extraction time length to cover a larger range of delayed echoes and improve the efficiency of positioning the actual delay; after the first delay value meets the preset delay change condition, the extraction time length of the preset time domain extraction layer can be reduced to restore a smaller visible field of view, reduce the amount of calculation, and ensure fine processing effects.

[0040] Figure 7 A flowchart of an audio echo processing method including a process of determining a second delay value is provided in an embodiment of the present application, such as Figure 7 As shown, the audio echo processing method specifically includes the following steps: Step S701: Acquire a microphone collection signal and a speaker to-be-broadcast signal, and perform a first delay estimation based on the microphone collection signal and the speaker to-be-broadcast signal to obtain a first delay value.

[0041] Step S702: When the first delay value satisfies a preset delay stability condition, a first reference signal is determined according to the first delay value and a signal to be played by a loudspeaker.

[0042] Step S703: Perform adaptive filtering calculation based on the first reference signal and the microphone collected signal to obtain a sampling delay value; perform frequency statistics on the sampling delay value to determine a second delay value, wherein the second delay value is smaller than the first delay value.

[0043] Among them, the adaptive filtering calculation can adopt the time domain LMS algorithm, the time domain NLMS algorithm, etc., which is not limited in this application. Taking the time domain LMS algorithm as an example, its calculation formula is as follows:

[0044]

[0045]

[0046]

[0047] in, Indicates the frame index, Indicates the filter order, ranging from 4ms to 40ms, and its order is at least , The filter coefficients, The first reference signal is indexed relative to the current frame Delay The signal value of a time unit, y( ) is the filter index in the current frame The output signal value, The first reference signal is indexed by the current frame The most recent benchmark The total energy at a point in time, The index of the microphone signal collected at the current frame The signal value of Current frame index The corresponding error signal value. Filter coefficient In the continuous update, when the filter converges, the filter coefficient The index corresponding to the maximum value can be determined as a sampling delay value representing the signal delay value. By performing frequency statistics, such as histogram statistics, on the obtained sampling delay value, the numerical distribution of the sampling delay value can be determined. In order to avoid the influence of delay jitter, the sampling delay value whose frequency satisfies a preset threshold condition can be determined as the final second delay value, ensuring the reliability of the second delay estimation. It should be noted that the second delay estimation can continue to collect sampling delay values before the second delay value is determined. After detecting the sampling delay value whose frequency satisfies the preset threshold condition and determining it as the second delay value, the second delay estimation can be temporarily closed, and the existing delay estimation result can be used for subsequent echo cancellation processing to avoid the problem of missing echoes caused by frequent moving signals. The first delay estimation can continue to be executed to cope with potential small probability delay changes until the first delay value re-meets the preset delay change condition and the second delay estimation is re-enabled after meeting the preset delay stability condition.

[0048] Step S704, determining a second reference signal according to the second delay value and the first reference signal.

[0049] Step S705, performing echo cancellation processing on the microphone collected signal based on the second reference signal to obtain a first near-end speech signal.

[0050] The above, by calculating the sampling delay value based on the first reference signal and the microphone collected signal, the delay value of the sampling point can be accurately estimated, and the frequency statistics of the sampling delay value can be determined to obtain the second delay value, which can ensure the reliability of the second delay value and improve the accuracy of the delay estimation.

[0051] Figure 8 A flowchart of an audio echo processing method provided by an embodiment of the present application is shown in FIG. 8, which includes the following steps: Figure 8 Step S801, obtaining a microphone collected signal and a loudspeaker to-be-played signal.

[0052] Step S802, generating a matching template signal based on the frequency domain information corresponding to the loudspeaker to-be-played signal, and performing correlation calculation on the matching template signal and the microphone collected signal to obtain a correlation result value; in the case where the correlation result value is greater than a preset threshold, calculating a first delay value based on the generation time of the matching template signal and the current time.

[0053] ​The frequency domain information can be obtained by performing fast Fourier transform on the segmented loudspeaker-to-be-played signal, and the segment duration can be 4N milliseconds, where N is a positive integer. For example, the segment duration is 4 ms, and fast Fourier transform is performed on each 4 ms reference signal in the loudspeaker-to-be-played signal to obtain the corresponding frequency domain signal. The matching template signal is generated based on the frequency domain information corresponding to the loudspeaker-to-be-played signal. Specifically, the total power of the signal in a preset frequency range in the frequency domain signal corresponding to each 4N ms reference signal of the loudspeaker-to-be-played signal can be calculated, and the preset frequency range can be 100 Hz to 4000 Hz. Then, it is determined whether the total power of the signal of the most recent K consecutive 4N ms reference signals is greater than a preset power threshold, where K is a positive integer, K multiplied by 4N is not greater than 200, and the preset power threshold can be in the range of 45 dB to 30 dB. Optionally, K can be 32, and N can be 1, so that the most recent 128 ms of the loudspeaker-to-be-played signal is referred to. If the total power of the signal of the most recent K consecutive 4N ms reference signals is greater than the preset power threshold, the matching template signal can be generated based on the frequency domain signals corresponding to the K consecutive 4N ms reference signals. After the matching template signal is generated, the correlation between the matching template signal and the microphone acquisition signal is calculated to obtain a correlation result value. Specifically, the correlation between the frequency domain signal corresponding to each 4N ms acquisition signal in the microphone acquisition signal and the frequency domain signal corresponding to the K 4N ms reference signals in the matching template signal can be calculated, and the correlation calculation results of the K consecutive 4N ms acquisition signals in the microphone acquisition signal and the K 4N ms reference signals can be combined to form a matrix. The average of the correlation values on the diagonal line of the matrix is the correlation result value. If the correlation result value is greater than a preset threshold, it can be considered that the echo delay is detected, and the time difference between the generation time of the matching template signal and the current time can be calculated to obtain a first delay value.

[0054] In step S803, when the first delay value satisfies a preset delay stability condition, a first reference signal is determined according to the first delay value and the loudspeaker-to-be-played signal, a second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone acquisition signal, a second reference signal is determined according to the second delay value and the first reference signal, and the second delay value is less than the first delay value.

[0055] In step S804, the microphone acquisition signal is processed based on the second reference signal to obtain a first near-end speech signal.

[0056] As described above, by constructing the matching template signal and performing correlation calculation on the microphone acquisition signal, the invalid correlation calculation can be reduced, the echo delay can be quickly located, and a reliable translation result can be provided for subsequent second delay estimation.

[0057] Figure 9 A flowchart of an audio echo processing method provided by an embodiment of the present application is shown in FIG. 9, which includes the following steps: Figure 9 In step S901, a microphone collected signal and a speaker to-be-played signal are obtained, and a first delay value is obtained by performing first delay estimation based on the microphone collected signal and the speaker to-be-played signal.

[0058] In step S902, in a case where the first delay value satisfies a preset delay stability condition, a first reference signal is determined according to the first delay value and the speaker to-be-played signal, second delay estimation is performed based on the first reference signal and the microphone collected signal to obtain a second delay value, and a second reference signal is determined according to the second delay value and the first reference signal, where the second delay value is smaller than the first delay value.

[0059] In step S903, echo cancellation processing is performed on the microphone collected signal based on the second reference signal to obtain a first near-end speech signal.

[0060] In step S904, in a case where the first delay value does not satisfy the preset delay stability condition, a fourth reference signal is determined according to the first delay value and the speaker to-be-played signal.

[0061] If the first delay value does not satisfy the preset delay stability condition, the first delay value still has a certain numerical fluctuation, which is likely to affect the accuracy of the second delay estimation. Therefore, the second delay estimation can be temporarily disabled, and echo cancellation processing is performed on the microphone collected signal based on the fourth reference signal determined according to the first delay value and the speaker to-be-played signal.

[0062] In step S905, echo cancellation processing is performed on the microphone collected signal based on the fourth reference signal to obtain a second near-end speech signal.

[0063] In the case where the first delay value does not satisfy the preset delay stability condition, the second delay estimation whose estimation result is likely to be affected is disabled, so as to avoid large errors in delay estimation and guarantee the subsequent echo cancellation effect.

[0064] Figure 10 A flowchart of an audio echo processing method provided by an embodiment of the present application is shown in FIG. 9, which includes the following steps: Figure 10 In step S901, a microphone collected signal and a speaker to-be-played signal are obtained, and a first delay value is obtained by performing first delay estimation based on the microphone collected signal and the speaker to-be-played signal.

[0065] ​​In a case where the first delay value does not satisfy the preset delay change condition, the echo cancellation processing is performed on the microphone collection signal based on the loudspeaker-to-be-played signal to obtain a third near-end speech signal.

[0066] If the first delay value does not satisfy the preset delay change condition, it can be considered that the delay change is not detected, and the subsequent echo cancellation processing can be performed by using the stored historical delay calculation result. The loudspeaker-to-be-played signal is shifted by referring to the stored historical delay calculation result, and the echo cancellation processing is performed on the microphone collection signal based on the shifting result to obtain the third near-end speech signal.

[0067] In a case where the first delay value satisfies the preset delay stability condition, a first reference signal is determined according to the first delay value and the loudspeaker-to-be-played signal, a second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone collection signal, a second reference signal is determined according to the second delay value and the first reference signal, and the second delay value is smaller than the first delay value.

[0068] If the first delay value satisfies the preset delay stability condition, it means that the delay change is detected first and then the delay stability is determined, and thus it can be considered that the preset delay change condition is satisfied. Optionally, in a case where the preset delay change condition is satisfied but the preset delay stability condition is not satisfied, a fourth reference signal can be determined according to the first delay value and the loudspeaker-to-be-played signal based on the foregoing embodiments, and the echo cancellation processing is performed on the microphone collection signal based on the fourth reference signal to obtain a second near-end speech signal.

[0069] The echo cancellation processing is performed on the microphone collection signal based on the second reference signal to obtain a first near-end speech signal.

[0070] In a case where the first delay value does not satisfy the preset delay change condition, the estimation result of the first delay estimation can not be referred to until it is confirmed that the preset delay change condition is satisfied, so that the actual delay position can be positioned reliably.

[0071] Figure 11 A structural block diagram of an audio echo processing device provided by an embodiment of the present application is shown in FIG. 11. The device is configured to execute the audio echo processing method provided by the foregoing embodiments, and has the corresponding function modules and beneficial effects of the execution method. As shown in FIG. 11, the device specifically includes: Figure 11 A signal acquisition module 1101 is configured to acquire a microphone collection signal and a loudspeaker-to-be-played signal. A delay estimation module 1102 is configured to perform first delay estimation based on the microphone collection signal and the loudspeaker-to-be-played signal to obtain a first delay value. ​The first reference signal determination module 1103 is configured to determine a first reference signal according to the first delay value and the speaker-to-be-played signal in a case where the first delay value satisfies a preset delay stability condition. The delay estimation module 1102 is further configured to perform second delay estimation based on the first reference signal and the microphone-acquired signal to obtain a second delay value, the second delay value being smaller than the first delay value. The first reference signal determination module 1103 is further configured to determine a second reference signal according to the second delay value and the first reference signal. The first signal processing module 1104 is configured to perform echo cancellation processing on the microphone-acquired signal based on the second reference signal to obtain a first near-end speech signal.

[0072] The first delay value is obtained by performing first delay estimation based on the microphone-acquired signal and the speaker-to-be-played signal, which can preliminarily perform signal delay estimation and complete signal coarse alignment, so as to improve the processing efficiency of subsequent signal fine alignment. The second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone-acquired signal in a case where the first delay value satisfies a preset delay stability condition, which can avoid single estimation error and complete signal fine alignment, so as to guarantee the accuracy of delay estimation. The first near-end speech signal is obtained by performing echo cancellation processing on the microphone-acquired signal based on the second reference signal, which can eliminate echo and improve audio quality. The above scheme can balance the calculation accuracy and efficiency, improve the calculation accuracy, reduce the calculation amount, ensure real-time and efficient echo processing, and guarantee the echo cancellation effect.

[0073] In one possible embodiment, the first signal processing module 1104 is further configured to: perform linear filtering processing on the second reference signal and the microphone-acquired signal to obtain a first residual signal and a third reference signal that leads the first residual signal; input the first residual signal and the third reference signal into a nonlinear network model to obtain the first near-end speech signal.

[0074] In one possible embodiment, the nonlinear network model includes a feature alignment network; and the audio echo processing apparatus further includes: The model parameter adjustment module is configured to: in a case where the first delay value satisfies a preset delay change condition, reduce the extraction time length of a preset time domain extraction layer in the feature alignment network, so as to perform signal alignment processing based on the preset time domain extraction layer with the reduced extraction time length through the feature alignment network.

[0075] In one possible embodiment, the delay estimation module 1102 is further configured to: perform adaptive filtering calculation based on the first reference signal and the microphone-acquired signal to obtain a sampling delay value. a frequency statistic is performed on the sampling delay value to determine a second delay value.

[0076] In one possible implementation, the delay estimation module 1102 is further configured to: generate a matching template signal based on the frequency domain information corresponding to the speaker-to-be-played signal, and perform correlation calculation on the matching template signal and the microphone collected signal to obtain a correlation result value; in a case where the correlation result value is greater than a preset threshold, calculate a first delay value based on a generation time of the matching template signal and a current time.

[0077] In one possible implementation, the audio echo processing apparatus further includes: a second reference signal determination module configured to: in a case where the first delay value does not satisfy a preset delay stability condition, determine a fourth reference signal according to the first delay value and the speaker-to-be-played signal; a second signal processing module configured to: perform echo cancellation processing on the microphone collected signal based on the fourth reference signal to obtain a second near-end speech signal.

[0078] In one possible implementation, the audio echo processing apparatus further includes: a third signal processing module configured to: in a case where the first delay value does not satisfy a preset delay change condition, perform echo cancellation processing on the microphone collected signal based on the speaker-to-be-played signal to obtain a third near-end speech signal.

[0079] In one possible implementation, the nonlinear network model includes a feature extraction network, a feature alignment network, and a feature processing network. The first signal processing module 1104 is further configured to: input the first residual signal and the third reference signal into the feature extraction network to obtain a first residual feature signal and a first reference feature signal; input the first residual feature signal and the first reference feature signal into the feature alignment network to obtain a second reference feature signal; splice the first residual feature signal and the second reference feature signal and input the spliced signal into the feature processing network to obtain the first near-end speech signal.

[0080] Figure 12 A structural schematic diagram of an audio echo processing device provided by an embodiment of the present application is shown in FIG. 1, which includes a processor 1201, a memory 1202, an input device 1203, and an output device 1204. The number of processors 1201 in the device can be one or more. Figure 12 The number of processors 1201 in the device can be one or more. Figure 12The processor 1201 in the device is taken as an example; the processor 1201, the memory 1202, the input device 1203 and the output device 1204 in the device can be connected through a bus or other means, Figure 12 The memory 1202 is taken as a computer readable storage medium, which can be configured to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the audio echo processing method in the embodiments of the present application. The processor 1201 performs various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 1202, that is, implements the audio echo processing method described above. The input device 1203 can be configured to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 1204 can include a display device such as a display screen.

[0081] The embodiments of the present application also provide a non-volatile storage medium containing computer executable instructions, which are configured to perform an audio echo processing method described in the above embodiments when executed by a computer processor, and the method comprises: obtaining a microphone acquisition signal and a speaker to be played signal, performing first delay estimation based on the microphone acquisition signal and the speaker to be played signal to obtain a first delay value; in the case that the first delay value meets a preset delay stability condition, determining a first reference signal according to the first delay value and the speaker to be played signal, performing second delay estimation based on the first reference signal and the microphone acquisition signal to obtain a second delay value, and determining a second reference signal according to the second delay value and the first reference signal, wherein the second delay value is less than the first delay value; and performing echo cancellation processing on the microphone acquisition signal based on the second reference signal to obtain a first near-end speech signal.

[0082] It is worth noting that in the above embodiments of the audio echo processing device, each unit and module included is only divided according to functional logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for easy mutual distinction, and does not configure to limit the protection scope of the embodiments of the present application.

[0083] In some possible implementation manners, each aspect of the method provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is configured to make the computer device execute the steps in the method according to various exemplary embodiments of the present application described in the specification, for example, the computer device can execute the audio echo processing method described in the embodiments of the present application. The program product can be realized in the combination of one or more readable media.

Claims

1. A method of audio echo processing, characterized by, The method comprises: obtaining a microphone collection signal and a speaker to be played signal, performing first delay estimation based on the microphone collection signal and the speaker to be played signal to obtain a first delay value; in a case where the first delay value meets a preset delay stability condition, determining a first reference signal according to the first delay value and the speaker to be played signal, performing second delay estimation based on the first reference signal and the microphone collection signal to obtain a second delay value, determining a second reference signal according to the second delay value and the first reference signal, and the second delay value is smaller than the first delay value; performing echo cancellation processing on the microphone collection signal based on the second reference signal to obtain a first near-end speech signal.

2. The audio echo processing method of claim 1, wherein, The echo cancellation processing on the microphone collection signal based on the second reference signal to obtain the first near-end speech signal comprises: performing linear filtering processing on the second reference signal and the microphone collection signal to obtain a first residual signal and a third reference signal leading the first residual signal; inputting the first residual signal and the third reference signal into a nonlinear network model to obtain the first near-end speech signal.

3. The audio echo processing method of claim 2, wherein, The nonlinear network model comprises a feature alignment network. After the first delay value is obtained by performing first delay estimation based on the microphone collection signal and the speaker to be played signal, the method further comprises: in a case where the first delay value meets a preset delay change condition, reducing the extraction time length of a preset time domain extraction layer in the feature alignment network, for performing signal alignment processing based on the preset time domain extraction layer with the reduced extraction time length through the feature alignment network.

4. The audio echo processing method of claim 1, wherein, The second delay value is obtained by performing second delay estimation based on the first reference signal and the microphone collection signal, comprising: performing adaptive filtering calculation based on the first reference signal and the microphone collection signal to obtain a sample delay value; performing frequency statistics on the sample delay value to determine the second delay value.

5. The audio echo processing method of claim 1, wherein, The first delay value is obtained by performing first delay estimation based on the microphone collection signal and the speaker to be played signal, comprising: generating a matching template signal based on the frequency domain information corresponding to the speaker to be played signal, and performing correlation calculation on the matching template signal and the microphone collection signal to obtain a correlation result value; in a case where the correlation result value is greater than a preset threshold, calculating the first delay value based on the generation time of the matching template signal and the current time.

6. The audio echo processing method of claim 1, wherein, After the first delay value is obtained by performing first delay estimation based on the microphone collection signal and the speaker to be played signal, the method further comprises: in a case where the first delay value does not meet the preset delay stability condition, determining a fourth reference signal according to the first delay value and the speaker to be played signal; performing echo cancellation processing on the microphone collection signal based on the fourth reference signal to obtain a second near-end speech signal.

7. The audio echo processing method of claim 1, wherein, After the first delay value is obtained by performing first delay estimation based on the microphone collection signal and the speaker to be played signal, the method further comprises: In a case where the first delay value does not satisfy a preset delay change condition, the microphone acquisition signal is subjected to echo cancellation processing based on the loudspeaker-to-be-played signal to obtain a third near-end speech signal.

8. The audio echo processing method of claim 2, wherein, The nonlinear network model comprises a feature extraction network, a feature alignment network, and a feature processing network. The first residual signal and the third reference signal are input into a nonlinear network model to obtain a first near-end speech signal, comprising: The first residual signal and the third reference signal are input into the feature extraction network to obtain a first residual feature signal and a first reference feature signal. The first residual feature signal and the first reference feature signal are input into the feature alignment network to obtain a second reference feature signal. The first residual feature signal and the second reference feature signal are spliced and input into the feature processing network to obtain a first near-end speech signal.

9. An audio echo processing apparatus, characterized by Comprise: The signal acquisition module is configured to acquire a microphone acquisition signal and a loudspeaker-to-be-played signal. The delay estimation module is configured to perform first delay estimation based on the microphone acquisition signal and the loudspeaker-to-be-played signal to obtain a first delay value. The first reference signal determination module is configured to determine a first reference signal according to the first delay value and the loudspeaker-to-be-played signal in a case where the first delay value satisfies a preset delay stability condition. The delay estimation module is further configured to perform second delay estimation based on the first reference signal and the microphone acquisition signal to obtain a second delay value, the second delay value being smaller than the first delay value. The first reference signal determination module is further configured to determine a second reference signal according to the second delay value and the first reference signal. The first signal processing module is configured to perform echo cancellation processing on the microphone acquisition signal based on the second reference signal to obtain a first near-end speech signal.

10. An audio echo processing device, characterized by The device comprises one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the audio echo processing method of any one of claims 1-8.

11. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to perform the audio echo processing method of any one of claims 1-8 when executed by a computer processor.

12. A computer program product comprising a computer program, characterized in that, The computer program is stored in a computer readable storage medium, and at least one processor of the device reads and executes the computer program from the computer readable storage medium, so that the device executes the audio echo processing method of any one of claims 1-8.

Citation Information

Cited By

  • Delay estimation methods, echo cancellation methods and related equipment

    CN122417057A