Method, device, apparatus and computer-readable storage medium for processing echo signals

By calculating the incoherence value and coherence value in audio and video calls, determining the suppression factor to process the error signal, the problem of echo signal interfering with the voice signal of the near-end speaker is solved, and the quality of audio and video calls is improved.

CN114360565BActive Publication Date: 2025-09-05BEIJING PHOENIX AUTO INTELLIGENCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210054282.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-12-28
Filing Date
2022-01-18
Publication Date
2025-09-05
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

During an audio or video call, after the far-end signal is played through the near-end speaker, the signal collected by the near-end microphone contains the voice signal of the near-end speaker and the echo signal of the far-end signal, which interferes with the voice signal of the near-end speaker and degrades the audio or video call quality.

Method used

By acquiring the call status of the far-end signal and the near-end signal, calculating the incoherence value and the coherence value, determining the suppression factor, and processing the error signal based on the suppression factor to weaken the echo signal.

Benefits of technology

Effectively reduce the residual echo signal in the error signal and improve the quality of audio and video calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360565B_ABST
    Figure CN114360565B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device, and computer-readable storage medium for processing an echo signal, belonging to the field of signal processing technology. The method includes: obtaining a call status based on a far-end signal and a near-end signal; obtaining an estimated echo signal based on the far-end signal, removing the estimated echo signal from the near-end signal, and obtaining an error signal; calculating an incoherence value between the far-end signal and the near-end signal, and a coherence value between the near-end signal and the error signal; determining a suppression factor based on the call status, the incoherence value, and the coherence value; and processing the error signal based on the suppression factor. Since the suppression factor is determined based on the call status, the incoherence value between the far-end signal and the near-end signal, and the coherence value between the near-end signal and the error signal, the suppression factor has a high correlation with the call status, the far-end signal, the near-end signal, and the error signal. Using the suppression factor to process the error signal has a good effect on attenuating the residual echo signal in the error signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application No. 202111629809.7, filed on December 28, 2021, entitled “Method, device, apparatus and computer-readable storage medium for processing echo signals,” the entire contents of which are incorporated by reference in the embodiments of this application. Technical Field

[0002] The embodiments of the present application relate to the field of signal processing technology, and in particular to a method, apparatus, device, and computer-readable storage medium for processing an echo signal. Background Art

[0003] During an audio or video call, after the far-end signal is played through the near-end speaker, the near-end signal collected by the near-end microphone will include the near-end speaker's voice signal and an echo signal of the far-end signal. To prevent this echo signal from interfering with the near-end speaker's voice signal and degrading the audio or video call quality, a method for processing and reducing this echo signal is needed. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, device, and computer-readable storage medium for processing an echo signal, which can be used to solve problems in related technologies.

[0005] In a first aspect, an embodiment of the present application provides a method for processing an echo signal, the method comprising:

[0006] Get call status based on far-end and near-end signals;

[0007] obtaining an estimated echo signal based on the far-end signal, and removing the estimated echo signal from the near-end signal to obtain an error signal;

[0008] calculating an incoherence value between the far-end signal and the near-end signal and a coherence value between the near-end signal and the error signal;

[0009] determining a suppression factor based on the call state, the incoherence value, and the coherence value;

[0010] The error signal is processed based on the suppression factor.

[0011] In a possible implementation, determining the suppression factor based on the call status, the incoherence value, and the coherence value includes:

[0012] If it is determined based on the call state, the incoherence value, and the coherence value that the residual echo signal included in the error signal is less than or equal to a specified threshold, determining the suppression factor to be a first factor;

[0013] If it is determined based on the call state, the incoherence value, and the coherence value that the residual echo signal included in the error signal is greater than the specified threshold, determining the suppression factor to be a second factor;

[0014] The first factor is greater than the second factor.

[0015] In one possible implementation, when the incoherence value is greater than a first reference value, the coherence value is greater than a second reference value, and the call state is a near-end single-talk state, determining that a residual echo signal included in the error signal is less than or equal to the specified threshold;

[0016] Alternatively, when the incoherence value is less than a third reference value, the coherence value is less than a fourth reference value, and the call state is a double-talk state, determining that the residual echo signal included in the error signal is less than or equal to the specified threshold;

[0017] Alternatively, when the incoherence value is greater than or equal to a fifth reference value and / or the coherence value is greater than or equal to a sixth reference value, and the call state is a far-end single-talk state, it is determined that the residual echo signal included in the error signal is less than or equal to the specified threshold.

[0018] In one possible implementation, when the incoherence value is less than or equal to a first reference value and / or the coherence value is less than or equal to a second reference value, and the call state is a near-end single-talk state, determining that a residual echo signal included in the error signal is greater than the specified threshold;

[0019] Alternatively, when the incoherence value is greater than or equal to a third reference value and / or the coherence value is greater than or equal to a fourth reference value, and the call state is a double-talk state, determining that the residual echo signal included in the error signal is greater than the specified threshold;

[0020] Alternatively, when the incoherence value is less than a fifth reference value, the coherence value is less than a sixth reference value, and the call state is a far-end single-talk state, it is determined that the residual echo signal included in the error signal is greater than the specified threshold.

[0021] In a possible implementation, calculating an incoherence value between the far-end signal and the near-end signal and a coherence value between the near-end signal and the error signal includes:

[0022] performing frame processing on the far-end signal, the near-end signal, and the error signal respectively, and obtaining, based on the frame processing results, a plurality of consecutive far-end signal frames, a plurality of consecutive near-end signal frames, and a plurality of consecutive error signal frames having the same frame length, wherein a first far-end signal frame, a first near-end signal frame, and a first error signal frame have the same time domain;

[0023] For a far-end signal frame, a near-end signal frame, and an error signal frame having the same time domain, calculating an incoherence value between the far-end signal frame and the near-end signal frame;

[0024] For the far-end signal frame, near-end signal frame and error signal frame having the same time domain, a coherence value of the near-end signal frame and the error signal frame is calculated.

[0025] In a possible implementation, the time domain corresponding to the suppression factor is the time domain corresponding to the coherence value used to determine the suppression factor, and the time domain corresponding to the coherence value is the time domain of the error signal frame used to calculate the coherence value;

[0026] The processing of the error signal based on the suppression factor comprises:

[0027] Based on the suppression factor, a time-domain error signal frame corresponding to the suppression factor is processed.

[0028] In a second aspect, an embodiment of the present application provides an echo signal processing device, the device comprising:

[0029] An acquisition module, configured to acquire a call status based on a far-end signal and a near-end signal;

[0030] The acquisition module is further configured to acquire an estimated echo signal based on the far-end signal, and remove the estimated echo signal from the near-end signal to obtain an error signal;

[0031] a calculation module, configured to calculate an incoherence value between the far-end signal and the near-end signal and a coherence value between the near-end signal and the error signal;

[0032] a determination module, configured to determine a suppression factor based on the call state, the incoherence value, and the coherence value;

[0033] A processing module is configured to process the error signal based on the suppression factor.

[0034] In one possible implementation, the determination module is configured to determine that the suppression factor is a first factor if it is determined based on the call status, the incoherence value, and the coherence value that the residual echo signal included in the error signal is less than or equal to a specified threshold; and to determine that the suppression factor is a second factor if it is determined based on the call status, the incoherence value, and the coherence value that the residual echo signal included in the error signal is greater than the specified threshold; wherein the first factor is greater than the second factor.

[0035] In one possible implementation, the determination module is configured to determine that the residual echo signal included in the error signal is less than or equal to the specified threshold when the incoherence value is greater than a first reference value, the coherence value is greater than a second reference value, and the call state is a near-end single-talk state; or, when the incoherence value is less than a third reference value, the coherence value is less than a fourth reference value, and the call state is a double-talk state, determine that the residual echo signal included in the error signal is less than or equal to the specified threshold; or, when the incoherence value is greater than or equal to a fifth reference value and / or the coherence value is greater than or equal to a sixth reference value, and the call state is a far-end single-talk state, determine that the residual echo signal included in the error signal is less than or equal to the specified threshold.

[0036] In one possible implementation, the determination module is configured to determine that the residual echo signal included in the error signal is greater than the specified threshold when the incoherence value is less than or equal to a first reference value and / or the coherence value is less than or equal to a second reference value, and the call state is a near-end single-talk state; or, when the incoherence value is greater than or equal to a third reference value and / or the coherence value is greater than or equal to a fourth reference value, and the call state is a double-talk state, determine that the residual echo signal included in the error signal is greater than the specified threshold; or, when the incoherence value is less than a fifth reference value, the coherence value is less than a sixth reference value, and the call state is a far-end single-talk state, determine that the residual echo signal included in the error signal is greater than the specified threshold.

[0037] In one possible implementation, a calculation module is used to perform frame processing on the far-end signal, the near-end signal and the error signal respectively, and obtain multiple continuous far-end signal frames, multiple continuous near-end signal frames and multiple continuous error signal frames with the same frame length based on the frame processing results, wherein the first far-end signal frame, the first near-end signal frame and the first error signal frame are the same in time domain; for the far-end signal frame, the near-end signal frame and the error signal frame with the same time domain, the incoherence value of the far-end signal frame and the near-end signal frame is calculated; for the far-end signal frame, the near-end signal frame and the error signal frame with the same time domain, the coherence value of the near-end signal frame and the error signal frame is calculated.

[0038] In one possible implementation, the time domain corresponding to the suppression factor is the time domain corresponding to the coherence value used to determine the suppression factor, and the time domain corresponding to the coherence value is the time domain of the error signal frame used to calculate the coherence value; the processing module is used to process the error signal frame in the time domain corresponding to the suppression factor based on the suppression factor.

[0039] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one program code or instruction, and the at least one program code or instruction is loaded and executed by the processor to enable the electronic device to implement the echo signal processing method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0040] In a fourth aspect, a computer-readable storage medium is also provided, which stores at least one program code or instruction, and the at least one program code or instruction is loaded and executed by a processor of an electronic device to enable the electronic device to implement the echo signal processing method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0041] In a fifth aspect, a computer program or a computer program product is also provided, wherein the computer program or the computer program product stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor of an electronic device so that the electronic device implements the echo signal processing method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0042] The technical solutions provided by the embodiments of the present application bring at least the following beneficial effects:

[0043] In the technical solution provided in the embodiment of the present application, since the suppression factor is jointly determined based on the call status, the incoherence value of the far-end signal and the near-end signal, and the coherence value of the near-end signal and the error signal, the suppression factor has a high correlation with the call status, the far-end signal, the near-end signal and the error signal. The error signal is processed using the suppression factor, which has a better effect of attenuating the residual echo signal in the error signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0045] Figure 1 1 is a schematic diagram of an implementation environment of an echo signal processing method provided in an embodiment of the present application;

[0046] Figure 2 This is a flowchart of a method for processing an echo signal provided in an embodiment of the present application;

[0047] Figure 3 This is a schematic diagram of a process for processing an error signal provided by an embodiment of the present application;

[0048] Figure 4 This is a schematic diagram of a process for processing an echo signal provided by an embodiment of the present application;

[0049] Figure 5 1 is a schematic structural diagram of an echo signal processing device provided in an embodiment of the present application;

[0050] Figure 6 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0052] Figure 1 FIG. 1 is a schematic diagram of an implementation environment of an echo signal processing method provided in an embodiment of the present application, such as Figure 1 As shown, the implementation environment includes: an electronic device 101. The echo signal processing method in the embodiment of the present application can be executed by the electronic device 101. Exemplarily, the electronic device 101 can include at least one of a terminal device and a server.

[0053] The terminal device can be at least one of a smart phone, a game console, a desktop computer, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a wearable device, a Pocket PC (PPC), and a smart car computer.

[0054] The server can be a single server, a server cluster consisting of multiple servers, or any one of a cloud computing platform and a virtualization center, which are not limited in the embodiments of the present application. The server can communicate with the terminal device via a wired network or a wireless network. The server can have functions such as data processing, data storage, and data transmission and reception, which are not limited in the embodiments of the present application.

[0055] Those skilled in the art should understand that the above-mentioned electronic device 101 is only an example, and other existing or future electronic devices may also be applicable to this application, should also be included in the scope of protection of this application, and are included here by reference.

[0056] based on Figure 1In the implementation environment shown, an embodiment of the present application provides a method for processing an echo signal, which is applied to an electronic device 101. Figure 2 FIG. 1 shows a flow chart of a method for processing an echo signal provided by an embodiment of the present application. Figure 2 As shown, the method includes but is not limited to the following steps 201 to 205.

[0057] Step 201: Acquire the call status based on the far-end signal and the near-end signal.

[0058] Exemplarily, the far-end signal is a sound signal generated at the far end and transmitted to the near end, and the near-end signal is a sound signal collected at the near end. For example, the far-end signal is a sound signal obtained by collecting the voice of the far-end speaker and transmitted to the near end, and the near-end signal includes a sound signal obtained by collecting the voice of the near-end speaker and an echo signal. Exemplarily, the near-end refers to the side of the electronic device that executes the echo signal processing method, and the far-end refers to the other side that is talking to the near-end speaker. For example, A and B are having an audio or video call. When the echo signal processing method is executed by the electronic device on A's side, A's side is the near-end and B's side is the far-end; when the echo signal processing method is executed by the electronic device on B's side, B's side is the near-end and A's side is the far-end.

[0059] In some embodiments, call states include, but are not limited to, a near-end single-talk state, a far-end single-talk state, and a dual-talk state. For example, the near-end single-talk state is when only the near-end speaker is speaking, the far-end single-talk state is when only the far-end speaker is speaking, and the dual-talk state is when both the near-end speaker and the far-end speaker are speaking simultaneously. Call states may also include other states, such as a dual-end silence state, in which neither the near-end speaker nor the far-end speaker is speaking, but this embodiment of the present application is not limited thereto. For example, a double-talk detection (DTD) algorithm is used to perform the operation of acquiring the call state based on the far-end signal and the near-end signal. The embodiment of the present application is not limited to the type of DTD algorithm; it may be any DTD algorithm capable of acquiring the call state based on the far-end signal and the near-end signal. For example, the DTD algorithm may be any one of an energy-based detection algorithm, a correlation-based detection algorithm, or an echo path-based detection algorithm.

[0060] Step 202: Obtain an estimated echo signal based on the far-end signal, and remove the estimated echo signal from the near-end signal to obtain an error signal.

[0061] Exemplarily, an adaptive filtering algorithm is employed to obtain an estimated echo signal based on the far-end signal. The present embodiment does not limit the type of adaptive filtering algorithm; it may be an adaptive filtering algorithm capable of obtaining an estimated echo signal based on the far-end signal. For example, the adaptive filtering algorithm may be a Normalized Least Mean Square (NLMS) algorithm. The present embodiment does not limit the manner in which the estimated echo signal is removed from the near-end signal. For example, the difference between the near-end signal and the estimated echo signal may be obtained to remove the estimated echo signal from the near-end signal.

[0062] It should be noted that the present embodiment does not limit the execution order of step 201 and step 202. Step 201 and step 202 can be executed simultaneously, or one step can be executed first and then the other step.

[0063] Step 203 : Calculate the incoherence value between the far-end signal and the near-end signal and the coherence value between the near-end signal and the error signal.

[0064] In some embodiments, step 203 includes but is not limited to steps 2031 to 2033 .

[0065] Step 2031: perform frame processing on the far-end signal, near-end signal and error signal respectively, and obtain multiple continuous far-end signal frames, multiple continuous near-end signal frames and multiple continuous error signal frames with the same frame length based on the frame processing results, wherein the first far-end signal frame, the first near-end signal frame and the first error signal frame are the same in time domain.

[0066] Exemplarily, the far-end signal, near-end signal, and error signal are framed with the same frame length, and the frame processing results obtained by the frame processing are delayed aligned to obtain multiple continuous far-end signal frames, multiple continuous near-end signal frames, and multiple continuous error signal frames with the same frame length, and the first far-end signal frame, the first near-end signal frame, and the first error signal frame have the same time domain. It should be noted that because the multiple far-end signal frames, multiple near-end signal frames, and multiple error signal frames are all continuous, and the first far-end signal frame, the first near-end signal frame, and the first error signal frame have the same time domain, the subsequent far-end signal frames, near-end signal frames, and error signal frames with the same sequence number have the same time domain. For example, the second far-end signal frame, the second near-end signal frame, and the second error signal frame have the same time domain.

[0067] Step 2032: For the far-end signal frame, the near-end signal frame, and the error signal frame having the same time domain, calculate the incoherence value of the far-end signal frame and the near-end signal frame.

[0068] In some embodiments, the incoherence value of the far-end signal and the near-end signal includes the incoherence value of the far-end signal frame and the near-end signal frame in each time domain. For example, the calculation of the incoherence value of the far-end signal frame and the near-end signal frame in one time domain is used as an example for explanation. When calculating the incoherence value of the far-end signal frame and the near-end signal frame in other time domains, the far-end signal frame and the near-end signal frame in the time domain during the calculation process can be replaced with the far-end signal frame and the near-end signal frame in other time domains, thereby obtaining the incoherence value of the far-end signal frame and the near-end signal frame in each time domain. In this way, the incoherence value of the far-end signal and the near-end signal can be obtained.

[0069] In some embodiments, calculating the incoherence value of the far-end signal frame and the near-end signal frame includes but is not limited to step 1-1 and step 1-2.

[0070] Step 1-1: Convert the far-end signal frame and the near-end signal frame respectively to obtain a far-end signal frame represented in the frequency domain and a near-end signal frame represented in the frequency domain; calculate the auto-power spectral density of the far-end signal frame represented in the frequency domain, the auto-power spectral density of the near-end signal frame represented in the frequency domain, and the cross-power spectral density between the far-end signal frame represented in the frequency domain and the near-end signal frame represented in the frequency domain.

[0071] Exemplarily, converting the far-end signal frame and the near-end signal frame separately includes: performing windowing processing on the far-end signal frame and the near-end signal frame respectively, and performing Fourier transform on the windowed far-end signal frame and the windowed near-end signal frame. The embodiments of the present application do not limit the methods of windowing processing and Fourier transform. For example, the windowing processing is to add a Hanning window, and the Fourier transform is a fast Fourier transform (FFT).

[0072] Exemplarily, the auto-power spectral density of the far-end signal frame represented in the frequency domain and the auto-power spectral density of the near-end signal frame represented in the frequency domain are calculated according to the following formula 1 and formula 2, respectively.

[0073]

[0074] in, is the auto-power spectrum density of the kth effective point of the far-end signal frame expressed in the frequency domain, X k is the kth valid point of the far-end signal frame represented in the frequency domain, and λ1 is the first smoothing coefficient, where 0<λ1<1. Exemplarily, the far-end signal frame represented in the frequency domain includes multiple valid points, each of which is related to a Fourier transform sampling point. For example, among the multiple signal points of the far-end signal frame represented in the frequency domain, a signal point obtained based on the first half of the sampling points is considered a valid point. Exemplarily, the value of λ1 can be set based on experience or actual needs.

[0075]

[0076] in, is the auto-power spectrum density of the kth effective point of the near-end signal frame expressed in the frequency domain, D k is the kth significant point of the near-end signal frame represented in the frequency domain, and λ2 is the second smoothing coefficient, where 0 < λ2 < 1. Exemplarily, the near-end signal frame represented in the frequency domain includes multiple significant points, each of which is related to a Fourier transform sampling point. For example, among the multiple signal points of the near-end signal frame represented in the frequency domain, a signal point obtained based on the first half of the sampling points is considered a significant point. Exemplarily, the value of λ2 can be set based on experience or actual needs.

[0077] Exemplarily, the cross-power spectrum density of the far-end signal frame represented in the frequency domain and the near-end signal frame represented in the frequency domain is calculated according to the following formula 3.

[0078]

[0079] in, is the cross power spectrum density of the kth effective point of the far-end signal frame expressed in the frequency domain and the kth effective point of the near-end signal frame expressed in the frequency domain, * is conjugated, X k and D k The parameters have the same meanings as those in the above formula, and λ3 is the third smoothing coefficient, where 0 < λ3 < 1. For example, the value of λ3 can be set based on experience or actual needs, for example, the value of λ3 is between 0.85 and 0.95.

[0080] Exemplarily, step 1-1 corresponds to Figure 3 The input X shown k and D k , and related content of calculating the cross-power spectral density of a far-end signal frame represented in the frequency domain and a near-end signal frame represented in the frequency domain.

[0081] Step 1-2, based on the auto-power spectral density of the far-end signal frame represented in the frequency domain, the auto-power spectral density of the near-end signal frame represented in the frequency domain, and the cross-power spectral density between the far-end signal frame represented in the frequency domain and the near-end signal frame represented in the frequency domain, calculate the incoherence value of the far-end signal frame and the near-end signal frame.

[0082] Exemplarily, the initial coherence value of the far-end signal frame and the near-end signal frame is calculated according to the following formula 4.

[0083]

[0084] Among them, Cohlxd is the initial coherence value between the kth effective point of the far-end signal frame expressed in the frequency domain and the kth effective point of the near-end signal frame expressed in the frequency domain. The remaining parameters have the same meanings as the relevant parameters in the above formula.

[0085] For example, the cohl values ​​corresponding to each effective point are calculated. xd The coherence value of the far-end signal frame and the near-end signal frame is obtained by adding the cumulative average of . For example, the incoherence value of the far-end signal frame and the near-end signal frame is calculated according to the following formula 5.

[0086] non_c xd =1-c xd (Formula 5)

[0087] Among them, non_c xd is the incoherence value between the far-end signal frame and the near-end signal frame, c xd is the coherence value between the far-end signal frame and the near-end signal frame.

[0088] Exemplarily, the method further includes: calculating an initial incoherence value of the far-end signal frame and the near-end signal frame. For example, the initial incoherence value of the far-end signal frame and the near-end signal frame is calculated according to the following formula 6.

[0089]

[0090] Among them, non_cohl xd is the initial incoherence value of the kth effective point of the far-end signal frame expressed in the frequency domain and the kth effective point of the near-end signal frame expressed in the frequency domain. The remaining parameters have the same meanings as the relevant parameters in the above formula.

[0091] Exemplarily, steps 1-2 correspond to Figure 3 The related content of calculating the incoherence value of the far-end signal frame and the near-end signal frame is shown.

[0092] Step 2033: For the far-end signal frame, near-end signal frame, and error signal frame having the same time domain, calculate the coherence value of the near-end signal frame and the error signal frame.

[0093] In some embodiments, the coherence value between the near-end signal and the error signal includes the coherence value between the near-end signal frame and the error signal frame in each time domain. For example, the calculation of the coherence value between the near-end signal frame and the error signal frame in one time domain is used as an example for explanation. When calculating the coherence value between the near-end signal frame and the error signal frame in another time domain, the near-end signal frame and the error signal frame in the time domain can be replaced with the near-end signal frame and the error signal frame in the other time domain during the calculation process, thereby obtaining the coherence value between the near-end signal frame and the error signal frame in each time domain. Thus, the coherence value between the near-end signal and the error signal can be obtained.

[0094] In some embodiments, calculating the coherence value of the near-end signal frame and the error signal frame includes but is not limited to step 2-1 and step 2-2.

[0095] Step 2-1: Convert the near-end signal frame and the error signal frame respectively to obtain a near-end signal frame represented in the frequency domain and an error signal frame represented in the frequency domain; calculate the auto-power spectral density of the near-end signal frame represented in the frequency domain, the auto-power spectral density of the error signal frame represented in the frequency domain, and the cross-power spectral density between the near-end signal frame represented in the frequency domain and the error signal frame represented in the frequency domain.

[0096] Exemplarily, converting the near-end signal frame and the error signal frame separately includes: performing windowing processing on the near-end signal frame and the error signal frame, respectively, and performing Fourier transform on the windowed near-end signal frame and the windowed error signal frame. The windowing and Fourier transform methods are similar in principle to those in step 2032 above and are not further described here.

[0097] For the method of calculating the auto-power spectral density of the near-end signal frame represented in the frequency domain, please refer to the above formula 2. For example, the auto-power spectral density of the error signal frame represented in the frequency domain is calculated according to the following formula 7.

[0098]

[0099] in, is the autopower spectrum density of the kth effective point of the error signal frame expressed in the frequency domain, E k is the kth significant point of the error signal frame represented in the frequency domain, and λ4 is the fourth smoothing coefficient, where 0<λ4<1. Exemplarily, the error signal frame represented in the frequency domain includes multiple significant points, each of which is related to a Fourier transform sampling point. For example, among the multiple signal points of the error signal frame represented in the frequency domain, a signal point obtained based on the first half of the sampling points is used as a significant point. Exemplarily, the value of λ4 can be set based on experience or actual needs.

[0100] Exemplarily, the cross-power spectrum density between the near-end signal frame represented in the frequency domain and the error signal frame represented in the frequency domain is calculated according to the following formula 8.

[0101]

[0102] in, is the cross power spectrum density between the kth effective point of the near-end signal frame expressed in the frequency domain and the kth effective point of the error signal frame expressed in the frequency domain, D k 、E kand * have the same meanings as the relevant parameters in the above formula, λ5 is the fifth smoothing coefficient, wherein 0<λ5<1. For example, the value of λ5 can be set based on experience or actual needs, for example, the value of λ5 is a value between 0.85 and 0.95.

[0103] Exemplarily, step 2-1 corresponds to Figure 3 The input D shown k and E k , and related content of calculating the cross-power spectral density of the near-end signal frame represented in the frequency domain and the error signal frame represented in the frequency domain.

[0104] Step 2-2, based on the auto-power spectral density of the near-end signal frame represented in the frequency domain, the auto-power spectral density of the error signal frame represented in the frequency domain, and the cross-power spectral density between the near-end signal frame represented in the frequency domain and the error signal frame represented in the frequency domain, calculate the coherence value of the near-end signal frame and the error signal frame.

[0105] Exemplarily, the initial coherence value of the near-end signal frame and the error signal frame is calculated according to the following formula 9.

[0106]

[0107] Among them, Cohl de is the initial coherence value of the kth effective point of the near-end signal frame expressed in the frequency domain and the kth effective point of the error signal frame expressed in the frequency domain. The remaining parameters have the same meanings as the relevant parameters in the above formula.

[0108] For example, the cohl values ​​corresponding to each effective point are calculated. de The coherence value of the near-end signal frame and the error signal frame is obtained by adding the cumulative average of

[0109] Exemplarily, step 2-2 corresponds to Figure 3 The relevant content of calculating the coherence value of the near-end signal frame and the error signal frame is shown.

[0110] It should be noted that the embodiment of the present application does not limit the execution order of step 2032 and step 2033. Step 2032 and step 2033 can be executed simultaneously, or one step can be executed first and then the other step. For example, when one step is executed first and then the other step, the same operations in the two steps may no longer be repeated. For example, when step 2032 is executed first and then step 2033, for the near-end signal frame represented in the frequency domain and the self-power density of the near-end signal frame represented in the frequency domain obtained in step 2032, the near-end signal frame represented in the frequency domain and the self-power density of the near-end signal frame represented in the frequency domain can be obtained when step 2033 is executed, without repeating the conversion process and the calculation process.

[0111] Step 204 : Determine a suppression factor based on the call status, the incoherence value, and the coherence value.

[0112] Exemplarily, based on the call status of each time domain, the incoherence value between the far-end signal frame and the near-end signal frame, and the coherence value between the near-end signal frame and the error signal frame, the suppression factor corresponding to each time domain is determined. Since the incoherence value and the coherence value corresponding to each time domain are obtained based on the far-end signal frame, the near-end signal frame, and the error signal frame of each time domain, the time domain corresponding to the suppression factor is the time domain corresponding to the coherence value used to determine the suppression factor, and the time domain corresponding to the coherence value is the time domain of the error signal frame used to calculate the coherence value. Of course, since the incoherence value used to determine the suppression factor is the same as the time domain corresponding to the coherence value, the time domain corresponding to the suppression factor is also the time domain corresponding to the incoherence value used to determine the suppression factor.

[0113] In some embodiments, for different situations of the call status, the incoherence value, and the coherence value, the suppression factor is determined based on the call status, the incoherence value, and the coherence value, including but not limited to the following situation A and situation B. Still taking the determination of the suppression factor of one time domain as an example for explanation, when determining the suppression factors of other time domains, the call status, the incoherence value between the far-end signal frame and the near-end signal frame, and the coherence value between the near-end signal frame and the error signal frame in the determination process of the time domain can be replaced with the call status, the incoherence value between the far-end signal frame and the near-end signal frame, and the coherence value between the near-end signal frame and the error signal frame in other time domains, respectively, thereby determining the suppression factors of each time domain.

[0114] Exemplarily, for the error signal frame, the near-end signal frame, and each valid point in the error signal frame, the determined suppression factor includes the suppression factor corresponding to each valid point. The following description is given by taking the determination of the suppression factor corresponding to the kth valid point of the error signal frame represented in the frequency domain as an example. When determining the suppression factors corresponding to other valid points of the error signal frame represented in the frequency domain, the initial incoherence value of the kth valid point of the far-end signal frame represented in the frequency domain and the kth valid point of the near-end signal frame represented in the frequency domain, and the initial coherence value of the kth valid point of the near-end signal frame represented in the frequency domain and the kth valid point of the error signal frame represented in the frequency domain in the determination process can be replaced by the initial incoherence value and initial coherence value corresponding to the other valid points, respectively.

[0115] For the sake of simplicity, the initial incoherence value of the kth valid point of the far-end signal frame represented in the frequency domain and the kth valid point of the near-end signal frame represented in the frequency domain is called the first initial incoherence value, and the initial coherence value of the kth valid point of the near-end signal frame represented in the frequency domain and the kth valid point of the error signal frame represented in the frequency domain is called the first initial coherence value.

[0116] In case A, if it is determined based on the call state, the incoherence value, and the coherence value that the residual echo signal included in the error signal is less than or equal to a specified threshold, the suppression factor is determined to be the first factor.

[0117] The specified threshold can be set based on experience or actual requirements, and is not limited in the present embodiment. When the residual echo signal included in the error signal is less than or equal to the specified threshold, the residual echo signal included in the error signal is relatively small.

[0118] Exemplarily, in any of the following three situations, it is determined that the residual echo signal included in the error signal is less than or equal to a specified threshold.

[0119] In case A1, the incoherence value is greater than the first reference value, the coherence value is greater than the second reference value, and the call state is the near-end single talk state.

[0120] For example, corresponding to case A1, the first factor is 1. The first reference value and the second reference value can be determined based on experience or actual needs. For example, the first reference value is 0.9, and the second reference value is 0.98.

[0121] In case A2, the incoherence value is less than the third reference value, the coherence value is less than the fourth reference value, and the call state is a double-talk state.

[0122] For example, corresponding to case A2, the first factor is the smaller value between the first initial incoherence value and the first initial coherence value. The third reference value and the fourth reference value can be determined based on experience or actual needs. For example, the third reference value is 0.9 and the fourth reference value is 0.98.

[0123] Case A3: the incoherence value is greater than or equal to the fifth reference value and / or the coherence value is greater than or equal to the sixth reference value, and the call state is the far-end single talk state.

[0124] That is, situation A3 includes three situations: (1) the incoherence value is greater than or equal to the fifth reference value, the coherence value is greater than or equal to the sixth reference value, and the call state is the far-end single-talk state; (2) the incoherence value is less than the fifth reference value, the coherence value is greater than or equal to the sixth reference value, and the call state is the far-end single-talk state; (3) the incoherence value is greater than or equal to the fifth reference value, the coherence value is less than the sixth reference value, and the call state is the far-end single-talk state.

[0125] For example, corresponding to case A3, the first factor is the smaller value of the first initial incoherence value of the first multiple and the first initial coherence value of the first multiple. The fifth reference value, the sixth reference value, and the first multiple can be determined based on experience or actual needs. For example, the fifth reference value is 0.75, the sixth reference value is 0.78, and the first multiple is 0.9.

[0126] In case B, if it is determined based on the call state, the incoherence value, and the coherence value that the residual echo signal included in the error signal is greater than a specified threshold, the suppression factor is determined to be the second factor.

[0127] The first factor is greater than the second factor. For example, if the residual echo signal included in the error signal is greater than a specified threshold, the error signal includes a relatively large amount of residual echo signal. Therefore, the second factor needs to be smaller than the first factor to achieve a better attenuation effect on the residual echo signal.

[0128] Exemplarily, in any one of the following three situations, it is determined that the residual echo signal included in the error signal is greater than a specified threshold.

[0129] Case B1: the incoherence value is less than or equal to the first reference value and / or the coherence value is less than or equal to the second reference value, and the call state is the near-end single talk state.

[0130] That is, case B1 includes three cases: (1) the incoherence value is less than or equal to the first reference value, the coherence value is less than or equal to the second reference value, and the call state is the near-end single talk state; (2) the incoherence value is greater than the first reference value, the coherence value is less than or equal to the second reference value, and the call state is the near-end single talk state; (3) the incoherence value is less than or equal to the first reference value, the coherence value is greater than the second reference value, and the call state is the near-end single talk state. The first reference value and the second reference value are respectively the first reference value and the second reference value in the above case A1.

[0131] For example, corresponding to case B1, the second factor is the smaller value of the first initial incoherence value and the first initial coherence value of the second multiple. The second multiple is a value between 0 and 1. The second multiple can be determined based on experience or actual needs, for example, the second multiple is 0.9.

[0132] Case B2: the incoherence value is greater than or equal to the third reference value and / or the coherence value is greater than or equal to the fourth reference value, and the call state is a double-talk state.

[0133] That is, case B2 includes three cases: (1) the incoherence value is greater than or equal to the third reference value, the coherence value is greater than or equal to the fourth reference value, and the call state is a double-talk state; (2) the incoherence value is less than the third reference value, the coherence value is greater than or equal to the fourth reference value, and the call state is a double-talk state; (3) the incoherence value is greater than or equal to the third reference value, the coherence value is less than the fourth reference value, and the call state is a double-talk state. The third reference value and the fourth reference value are respectively the third reference value and the fourth reference value in case A2 above.

[0134] For example, corresponding to case B2, the second factor is the smaller value of the first initial incoherence value and the first initial coherence value of the third multiple. The third multiple is a value between 0 and 1. The third multiple can be determined based on experience or actual needs, for example, the third multiple is 0.9.

[0135] Case B3: the incoherence value is less than the fifth reference value, the coherence value is less than the sixth reference value, and the call state is the far-end single talk state.

[0136] For example, in case B3, the second factor is the smaller of the first initial incoherence value and the first initial coherence value of the fourth multiple. The fourth multiple is a value between 0 and 1, and is smaller than the first multiple in case A3. The fourth multiple can be determined based on experience or actual needs, for example, the fourth multiple is 0.1.

[0137] Exemplarily, step 204 corresponds to Figure 3 The input call status shown and the relevant content of the suppression factor corresponding to each time domain are determined based on the call status of each time domain, the incoherence value of the far-end signal frame and the near-end signal frame, and the coherence value of the near-end signal frame and the error signal frame.

[0138] Because the residual echo signal included in the error signal is indicated by the call status, the incoherence value between the far-end signal and the near-end signal, and the coherence value between the near-end signal and the error signal, the method provided in this embodiment of the application accurately determines whether the error signal includes a residual echo signal. Furthermore, because the suppression factor corresponds to the residual echo signal, subsequent processing of the error signal based on the suppression factor effectively reduces the residual echo signal in the error signal.

[0139] Step 205 : Process the error signal based on the suppression factor.

[0140] In some embodiments, processing the error signal based on the suppression factor includes: processing the error signal frame in the time domain corresponding to the suppression factor based on the suppression factor. Exemplarily, processing the error signal frame in the time domain corresponding to the suppression factor based on the suppression factor includes: multiplying the suppression factor with the frequency spectrum of the error signal frame in the time domain corresponding to the suppression factor. In one possible implementation, multiplying the suppression factor with the frequency spectrum of the error signal frame in the time domain corresponding to the suppression factor includes: for each valid point in the error signal frame, multiplying the suppression factor corresponding to the valid point by the frequency of the valid point. For example, for the kth valid point in the error signal frame, multiplying the suppression factor corresponding to the kth valid point by the frequency of the kth valid point.

[0141] In one possible implementation, before multiplying the suppression factor with the spectrum of the error signal frame in the time domain corresponding to the suppression factor, the method also includes: obtaining the energy of the near-end signal frame and the energy of the error signal frame in each time domain; and determining the spectrum of the error signal frame to be multiplied with the suppression factor corresponding to each time domain based on the energy of the near-end signal frame and the energy of the error signal frame in each time domain.

[0142] Exemplarily, obtaining the energy of a near-end signal frame and the energy of an error signal frame in one time domain is used as an example for description. When obtaining the energy of a near-end signal frame and the energy of an error signal frame in another time domain, the near-end signal frame and the error signal frame in the time domain during the acquisition process can be replaced with the near-end signal frame and the error signal frame in the other time domain, respectively. In this way, the energy of the near-end signal frame and the energy of the error signal frame in each time domain can be obtained.

[0143] In some embodiments, obtaining the energy of the near-end signal frame and the energy of the error signal frame includes: calculating the power spectral density of the near-end signal frame and the power spectral density of the error signal frame; calculating the energy of the near-end signal frame based on the power spectral density of the near-end signal frame; and calculating the energy of the error signal frame based on the power spectral density of the error signal frame.

[0144] Exemplarily, the power spectrum density of the near-end signal frame and the power spectrum density of the error signal frame are calculated according to the following formula 10 and formula 11, respectively.

[0145]

[0146] in, is the power spectrum density of the kth effective point of the near-end signal frame expressed in the frequency domain, D k is the kth effective point of the near-end signal frame represented in the frequency domain, λ6 is the sixth smoothing coefficient, wherein 0<λ6<1. For example, the value of λ6 can be set based on experience or actual needs, for example, the value of λ6 is a value between 0.85 and 0.95.

[0147]

[0148] in, is the power spectrum density of the kth effective point of the error signal frame expressed in the frequency domain, E k is the kth effective point of the error signal frame represented in the frequency domain, and λ7 is the seventh smoothing coefficient, wherein 0<λ7<1. For example, the value of λ7 can be set based on experience or actual needs, for example, the value of λ7 is a value between 0.85 and 0.95.

[0149] Exemplarily, the above content corresponds to Figure 3 The related contents of calculating the power spectrum density of the near-end signal frame represented in the frequency domain and the power spectrum density of the error signal frame represented in the frequency domain are shown.

[0150] For example, the energy of the near-end signal frame is the energy of each valid point corresponding to The energy of the error signal frame is the cumulative sum of the corresponding The cumulative sum of . Exemplarily, this step corresponds to Figure 3 The related contents of calculating the energy of the near-end signal frame and the energy of the error signal frame are shown.

[0151] According to the content of the above step 203, the near-end signal and the error signal are converted respectively to obtain a near-end signal frame represented in the frequency domain and an error signal frame represented in the frequency domain, thereby obtaining the spectrum of the near-end signal frame and the spectrum of the error signal frame.

[0152] Exemplarily, the spectrum of the error signal frame for multiplication with the suppression factor is determined according to the energy of the near-end signal frame and the energy of the error signal frame, including but not limited to the following Case 1 and Case 2.

[0153] Case 1: The energy of the error signal frame is greater than the energy of the near-end signal frame.

[0154] Exemplarily, this situation 1 indicates that the reliability of the error signal frame is low.

[0155] In one possible implementation, if the energy of the error signal frame is greater than the energy of the near-end signal frame and less than or equal to the fifth multiple of the energy of the near-end signal frame, the error signal frame can still be retained, and the spectrum of the error signal frame is replaced with the spectrum of the near-end signal frame. The spectrum of the replaced error signal frame is used as the spectrum of the error signal frame for multiplication by the suppression factor. The fifth multiple can be determined based on experience or actual needs, for example, the fifth multiple is 19.95.

[0156] In another possible implementation, if the energy of the error signal frame is greater than the fifth multiple of the energy of the near-end signal frame, it indicates that the error signal frame cannot be retained, and the spectrum of the error signal frame is set to zero. The spectrum of the error signal frame after being set to zero is used as the spectrum of the error signal frame for multiplication by the suppression factor. In other words, the error signal frame is eliminated.

[0157] Case 2: The energy of the error signal frame is less than or equal to the energy of the near-end signal frame.

[0158] Exemplarily, case 2 indicates that the reliability of the error signal frame is high. Exemplarily, when the energy of the error signal frame is less than or equal to the energy of the near-end signal frame, the spectrum of the error signal frame is used as the spectrum of the error signal frame for multiplying the suppression factor.

[0159] Exemplarily, step 205 corresponds to Figure 3 The application of the suppression factor to the relevant content of the error signal frame is shown.

[0160] In the method provided in the embodiment of the present application, since the suppression factor is jointly determined based on the call status, the incoherence value of the far-end signal and the near-end signal, and the coherence value of the near-end signal and the error signal, the suppression factor has a high correlation with the call status, the far-end signal, the near-end signal and the error signal. The error signal is processed using the suppression factor, which has a better effect of attenuating the residual echo signal in the error signal.

[0161] The above describes the echo signal processing method provided in the embodiment of the present application from the perspective of method steps. Next, the echo signal processing method provided in the embodiment of the present application will be described in conjunction with a scenario. Figure 4 The figure shows a schematic diagram of a process for processing an echo signal provided by an embodiment of the present application.

[0162] For example, Figure 4The illustrated scenario includes a speaker module, a microphone module, a two-end speech detection module, an adaptive filtering module, and a nonlinear processing module. The speaker module is used to play the far-end signal, the microphone module is used to collect the near-end signal, the two-end detection module is used to determine the call status, the adaptive filtering module is used to obtain an estimated echo signal, and the nonlinear processing module is used to perform nonlinear processing, which includes steps 203 to 205 described above. For example, the far-end signal is played through the speaker module, and an echo signal is generated based on the far-end signal. The path for generating the echo signal is called an echo channel. The microphone module collects the near-end speaker's voice signal and the echo signal. The far-end signal and the near-end signal are input to the two-end speech detection module for two-end speech detection to obtain the call status. This call status can be input to the adaptive filtering module along with the far-end signal to obtain an estimated echo signal. Whether to input the call status to the adaptive filtering module can be determined based on the adaptive filtering algorithm applied to the adaptive filtering module. The estimated echo signal is removed from the near-end signal to obtain an error signal. The call status and the error signal are input to the nonlinear processing module to perform steps 203 to 205 described above to obtain an output signal.

[0163] Figure 5 FIG. 1 is a schematic diagram of the structure of an echo signal processing device provided in an embodiment of the present application, as shown in FIG. Figure 5 As shown, the device includes:

[0164] An acquisition module 501 is configured to acquire a call status based on a far-end signal and a near-end signal;

[0165] The acquisition module 501 is further configured to acquire an estimated echo signal based on the far-end signal, and remove the estimated echo signal from the near-end signal to obtain an error signal;

[0166] A calculation module 502 is configured to calculate an incoherence value between the far-end signal and the near-end signal and a coherence value between the near-end signal and the error signal;

[0167] a determination module 503 for determining a suppression factor based on the call state, the incoherence value, and the coherence value;

[0168] The processing module 504 is configured to process the error signal based on the suppression factor.

[0169] In one possible implementation, the determination module 503 is configured to determine that the suppression factor is a first factor if the residual echo signal included in the error signal is determined to be less than or equal to a specified threshold based on the call status, the incoherence value, and the coherence value; and to determine that the suppression factor is a second factor if the residual echo signal included in the error signal is determined to be greater than the specified threshold based on the call status, the incoherence value, and the coherence value; wherein the first factor is greater than the second factor.

[0170] In one possible implementation, the determination module 503 is configured to determine that the residual echo signal included in the error signal is less than or equal to a specified threshold when the incoherence value is greater than a first reference value, the coherence value is greater than a second reference value, and the call state is a near-end single-talk state; or, when the incoherence value is less than a third reference value, the coherence value is less than a fourth reference value, and the call state is a double-talk state, determine that the residual echo signal included in the error signal is less than or equal to the specified threshold; or, when the incoherence value is greater than or equal to a fifth reference value and / or the coherence value is greater than or equal to a sixth reference value, and the call state is a far-end single-talk state, determine that the residual echo signal included in the error signal is less than or equal to the specified threshold.

[0171] In one possible implementation, the determination module 503 is configured to determine that the residual echo signal included in the error signal is greater than a specified threshold when the incoherence value is less than or equal to a first reference value and / or the coherence value is less than or equal to a second reference value, and the call state is a near-end single-talk state; or, when the incoherence value is greater than or equal to a third reference value and / or the coherence value is greater than or equal to a fourth reference value, and the call state is a double-talk state, determine that the residual echo signal included in the error signal is greater than the specified threshold; or, when the incoherence value is less than a fifth reference value, the coherence value is less than a sixth reference value, and the call state is a far-end single-talk state, determine that the residual echo signal included in the error signal is greater than the specified threshold.

[0172] In one possible implementation, the calculation module 502 is used to perform frame processing on the far-end signal, the near-end signal and the error signal respectively, and obtain multiple continuous far-end signal frames, multiple continuous near-end signal frames and multiple continuous error signal frames with the same frame length based on the frame processing results, wherein the first far-end signal frame, the first near-end signal frame and the first error signal frame are the same in time domain; for the far-end signal frame, the near-end signal frame and the error signal frame with the same time domain, calculate the incoherence value of the far-end signal frame and the near-end signal frame; for the far-end signal frame, the near-end signal frame and the error signal frame with the same time domain, calculate the coherence value of the near-end signal frame and the error signal frame.

[0173] In one possible implementation, the time domain corresponding to the suppression factor is the time domain corresponding to the coherence value used to determine the suppression factor, and the time domain corresponding to the coherence value is the time domain of the error signal frame used to calculate the coherence value; the processing module 504 is used to process the error signal frame in the time domain corresponding to the suppression factor based on the suppression factor.

[0174] In the above-mentioned device, since the suppression factor is determined based on the call status, the incoherence value of the far-end signal and the near-end signal, and the coherence value of the near-end signal and the error signal, the suppression factor has a high correlation with the call status, the far-end signal, the near-end signal and the error signal. The error signal is processed using the suppression factor, and the attenuation effect is better.

[0175] It should be understood that the above Figure 5 The provided device is illustrated only by the division of the above-mentioned functional modules when implementing its functions. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the corresponding method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0176] Figure 6 The following is a block diagram of an electronic device 600 according to an exemplary embodiment of the present application. The electronic device 600 may be a portable mobile terminal, such as a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. The electronic device 600 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other similar names.

[0177] Typically, the electronic device 600 includes a processor 601 and a memory 602 .

[0178] The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0179] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one instruction, which is executed by the processor 601 to implement the echo signal processing method provided in the method embodiment of the present application.

[0180] In some embodiments, electronic device 600 may optionally include a peripheral device interface 603 and at least one peripheral device. Processor 601, memory 602, and peripheral device interface 603 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 603 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, a positioning assembly 608, and a power supply 609.

[0181] The peripheral device interface 603 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0182] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 604 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or Wi-Fi (Wireless Fidelity) networks. In some embodiments, the radio frequency circuit 604 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.

[0183] Display screen 605 is used to display a user interface (UI). This UI can include graphics, text, icons, videos, or any combination thereof. When display screen 605 is a touch screen display, it can also capture touch signals on or above the surface of display screen 605. These touch signals can be input as control signals to processor 601 for processing. Display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 605, located on the front panel of electronic device 600. In other embodiments, there can be at least two display screens 605, located on different surfaces of electronic device 600 or in a foldable design. In other embodiments, display screen 605 can be a flexible display, located on a curved or foldable surface of electronic device 600. Display screen 605 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 605 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0184] The camera assembly 606 is used to capture images or videos. Optionally, the camera assembly 606 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0185] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals to be input into the processor 601 for processing, or input into the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, which are respectively arranged in different parts of the electronic device 600. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signals into sound waves audible to humans, but also convert the electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 607 may also include a headphone jack.

[0186] The positioning component 608 is used to locate the current geographic location of the electronic device 600 to implement navigation or LBS (Location Based Service). The positioning component 608 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.

[0187] Power supply 609 is used to power the various components in electronic device 600. Power supply 609 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 609 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0188] In some embodiments, the electronic device 600 further includes one or more sensors 610 , including but not limited to: an acceleration sensor 611 , a gyroscope sensor 612 , a pressure sensor 613 , a fingerprint sensor 614 , an optical sensor 615 , and a proximity sensor 616 .

[0189] The accelerometer 611 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the electronic device 600. For example, the accelerometer 611 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 601 can control the display screen 605 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 611. The accelerometer 611 can also be used to collect game or user motion data.

[0190] The gyroscope sensor 612 can detect the orientation and rotation angle of the electronic device 600. It can work in conjunction with the accelerometer 611 to collect the user's 3D movements of the electronic device 600. Based on the data collected by the gyroscope sensor 612, the processor 601 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0191] The pressure sensor 613 can be set on the side frame of the electronic device 600 and / or the lower layer of the display screen 605. When the pressure sensor 613 is set on the side frame of the electronic device 600, it can detect the user's grip signal of the electronic device 600, and the processor 601 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 613. When the pressure sensor 613 is set on the lower layer of the display screen 605, the processor 601 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 605. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0192] The fingerprint sensor 614 is used to collect the user's fingerprint. The processor 601 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 614, or the fingerprint sensor 614 identifies the user's identity based on the collected fingerprint. When the user's identity is recognized as a trusted identity, the processor 601 authorizes the user to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 614 can be set on the front, back, or side of the electronic device 600. When a physical button or manufacturer logo is set on the electronic device 600, the fingerprint sensor 614 can be integrated with the physical button or manufacturer logo.

[0193] Optical sensor 615 is used to detect ambient light intensity. In one embodiment, processor 601 can control the display brightness of display screen 605 based on the ambient light intensity detected by optical sensor 615. Specifically, when the ambient light intensity is high, the display brightness of display screen 605 is increased; when the ambient light intensity is low, the display brightness of display screen 605 is decreased. In another embodiment, processor 601 can also dynamically adjust the shooting parameters of camera assembly 606 based on the ambient light intensity detected by optical sensor 615.

[0194] Proximity sensor 616, also known as a distance sensor, is typically located on the front panel of electronic device 600. Proximity sensor 616 is used to detect the distance between the user and the front of electronic device 600. In one embodiment, when proximity sensor 616 detects that the distance between the user and the front of electronic device 600 is gradually decreasing, processor 601 controls display screen 605 to switch from the screen-on state to the screen-off state. When proximity sensor 616 detects that the distance between the user and the front of electronic device 600 is gradually increasing, processor 601 controls display screen 605 to switch from the screen-off state to the screen-on state.

[0195] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation on the electronic device 600, and the electronic device 600 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0196] In an exemplary embodiment, a computer-readable storage medium is further provided. The storage medium stores at least one program code. The at least one program code is loaded and executed by a processor of an electronic device to enable the electronic device to implement any of the above-mentioned echo signal processing methods.

[0197] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0198] In an exemplary embodiment, a computer program or a computer program product is further provided. The computer program or the computer program product stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor of an electronic device to enable the electronic device to implement any of the above-mentioned echo signal processing methods.

[0199] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0200] The terms "first," "second," "third," and "fourth," etc. in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0201] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for processing an echo signal, characterized in that: The method comprises: Get call status based on far-end and near-end signals; obtaining an estimated echo signal based on the far-end signal, and removing the estimated echo signal from the near-end signal to obtain an error signal; calculating an incoherence value between the far-end signal and the near-end signal and a coherence value between the near-end signal and the error signal; Under different call states, the incoherence value and the coherence value are compared with respective reference values ​​to determine whether a residual echo signal included in the error signal is less than or equal to a specified threshold; if the residual echo signal included in the error signal is less than or equal to the specified threshold, determining the suppression factor to be a first factor; if the residual echo signal included in the error signal is greater than the specified threshold, determining the suppression factor to be a second factor; wherein the first factor is greater than the second factor; The error signal is processed based on the suppression factor.

2. The method according to claim 1, characterized in that When the incoherence value is greater than a first reference value, the coherence value is greater than a second reference value, and the call state is a near-end single-talk state, determining that the residual echo signal included in the error signal is less than or equal to the specified threshold; Alternatively, when the incoherence value is less than a third reference value, the coherence value is less than a fourth reference value, and the call state is a double-talk state, determining that the residual echo signal included in the error signal is less than or equal to the specified threshold; Alternatively, when the incoherence value is greater than or equal to a fifth reference value and / or the coherence value is greater than or equal to a sixth reference value, and the call state is a far-end single-talk state, it is determined that the residual echo signal included in the error signal is less than or equal to the specified threshold.

3. The method according to claim 1, characterized in that When the incoherence value is less than or equal to a first reference value and / or the coherence value is less than or equal to a second reference value, and the call state is a near-end single-talk state, determining that the residual echo signal included in the error signal is greater than the specified threshold; Alternatively, when the incoherence value is greater than or equal to a third reference value and / or the coherence value is greater than or equal to a fourth reference value, and the call state is a double-talk state, determining that the residual echo signal included in the error signal is greater than the specified threshold; Alternatively, when the incoherence value is less than a fifth reference value, the coherence value is less than a sixth reference value, and the call state is a far-end single-talk state, it is determined that the residual echo signal included in the error signal is greater than the specified threshold.

4. The method according to any one of claims 1 to 3, characterized in that: The calculating the incoherence value between the far-end signal and the near-end signal and the coherence value between the near-end signal and the error signal includes: performing frame processing on the far-end signal, the near-end signal, and the error signal respectively, and obtaining, based on the frame processing results, a plurality of consecutive far-end signal frames, a plurality of consecutive near-end signal frames, and a plurality of consecutive error signal frames having the same frame length, wherein a first far-end signal frame, a first near-end signal frame, and a first error signal frame have the same time domain; For a far-end signal frame, a near-end signal frame, and an error signal frame having the same time domain, calculating an incoherence value between the far-end signal frame and the near-end signal frame; For the far-end signal frame, near-end signal frame and error signal frame having the same time domain, a coherence value of the near-end signal frame and the error signal frame is calculated.

5. The method according to claim 4, characterized in that The time domain corresponding to the suppression factor is the time domain corresponding to the coherence value used to determine the suppression factor, and the time domain corresponding to the coherence value is the time domain of the error signal frame used to calculate the coherence value; The processing of the error signal based on the suppression factor comprises: Based on the suppression factor, a time-domain error signal frame corresponding to the suppression factor is processed.

6. An echo signal processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a call status based on a far-end signal and a near-end signal; The acquisition module is further configured to acquire an estimated echo signal based on the far-end signal, and remove the estimated echo signal from the near-end signal to obtain an error signal; a calculation module, configured to calculate an incoherence value between the far-end signal and the near-end signal and a coherence value between the near-end signal and the error signal; a determination module, configured to compare the incoherence value and the coherence value with reference values ​​respectively under different call states to determine whether a residual echo signal included in the error signal is less than or equal to a specified threshold, and if so, determine a suppression factor as a first factor; and if so, determine a suppression factor as a second factor; wherein the first factor is greater than the second factor; A processing module is configured to process the error signal based on the suppression factor.

7. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one program code or instruction, and the at least one program code or instruction is loaded and executed by the processor, so that the electronic device implements the echo signal processing method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one program code, and the at least one program code is loaded and executed by a processor of an electronic device, so that the electronic device implements the echo signal processing method according to any one of claims 1 to 5.

9. A computer program product, characterized in that The computer program product stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor of an electronic device, so that the electronic device implements the echo signal processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Call signal processing method and device, electronic equipment and storage medium

    CN110971769A