Echo control method, apparatus, device, medium, and computer program product

By predicting the energy proportion of the simulated echo signal and adjusting the power spectrum during voice calls, the problem of poor echo cancellation in traditional echo control methods is solved, and higher quality voice calls are achieved.

CN116805492BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-03-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Traditional echo control methods have poor echo cancellation effects, resulting in low call quality.

Method used

By acquiring the current voice signal from the other end during a voice call, echo prediction processing is performed to obtain a simulated echo signal. The energy proportion of this signal at the local end is predicted, and the power spectrum is adjusted according to the estimated signal-to-echo ratio to reduce the energy of the voice signal to be played, thereby eliminating the real echo signal.

Benefits of technology

It effectively avoids the occurrence of large echo signals, improves the voice quality of calls, reduces echo residue, and enhances the clarity of voice calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805492B_ABST
    Figure CN116805492B_ABST
Patent Text Reader

Abstract

The application relates to an echo control method, device, equipment, medium and computer program product, and belongs to the technical field of voice processing. The method comprises the following steps: acquiring a current opposite-end voice signal in a voice call; performing echo prediction processing on the current opposite-end voice signal to obtain a simulated echo signal; predicting an estimated signal-to-echo ratio of the simulated echo signal; the estimated signal-to-echo ratio is used for representing the energy proportion of the simulated echo signal in the current acquisition signal at the local end; performing power spectrum adjustment processing on the current opposite-end voice signal according to the estimated signal-to-echo ratio to obtain an adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than that of the current opposite-end voice signal; and performing echo cancellation processing on a real echo signal generated based on the current voice signal to be played. The method can improve the voice quality of the call.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to an echo control method, apparatus, device, medium, and computer program product. Background Technology

[0002] With the development of computer technology, people have increasingly higher requirements for the quality of voice calls. Echo is a significant factor affecting voice call quality. Echo occurs when the sound signal emitted from the speaker is picked up by the microphone and transmitted back to the other end of the call. Essentially, an echo is the sound a user hears of themselves speaking, but this sound travels back to the other end within a short period. The presence of echo reduces voice quality and severely interferes with the user's communication. Therefore, echo cancellation is necessary to improve voice call quality.

[0003] Traditional techniques primarily eliminate echoes by directly suppressing the generated echo signal. However, traditional echo control methods are relatively ineffective at eliminating echoes, resulting in lower voice quality during calls. Summary of the Invention

[0004] Therefore, it is necessary to provide an echo control method, device, equipment, medium, and computer program product that can improve the quality of voice calls, addressing the aforementioned technical problems.

[0005] In a first aspect, this application provides an echo control method, the method comprising:

[0006] Get the current voice signal from the other end in a voice call;

[0007] Echo prediction processing is performed on the current peer voice signal to obtain a simulated echo signal;

[0008] The estimated signal-to-return ratio of the simulated echo signal is predicted; the estimated signal-to-return ratio is used to characterize the energy proportion of the simulated echo signal in the currently acquired signal at this end.

[0009] Based on the estimated signal-to-return ratio, the current peer voice signal is subjected to power spectrum adjustment processing to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal.

[0010] Echo cancellation processing is performed on the actual echo signal generated based on the currently playing audio signal.

[0011] Secondly, this application provides an echo control device, the device comprising:

[0012] The acquisition module is used to acquire the current voice signal from the other end during a voice call;

[0013] The prediction module is used to perform echo prediction processing on the current peer voice signal to obtain a simulated echo signal; predict the estimated signal-to-echo ratio of the simulated echo signal; the estimated signal-to-echo ratio is used to characterize the energy proportion of the simulated echo signal in the current acquired signal at the local end;

[0014] The adjustment module is used to perform power spectrum adjustment processing on the current peer voice signal according to the estimated signal-to-return ratio to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal.

[0015] The echo cancellation module is used to perform echo cancellation processing on the real echo signal generated based on the current audio signal to be played.

[0016] In one embodiment, the prediction module is further configured to acquire the current acquisition signal at the local end; determine the power spectrum of the current acquisition signal; determine the power spectrum of the analog echo signal; and determine the estimated signal-to-return ratio of the analog echo signal based on the power spectrum of the current acquisition signal and the power spectrum of the analog echo signal.

[0017] In one embodiment, the prediction module is further configured to perform frame segmentation processing on the current acquisition signal to obtain multiple frames of acquisition signals; for each frame of acquisition signal in the multiple frames of acquisition signals, determine the initial power value of the acquisition signal; obtain the prior power value of the previously acquired signal; the prior acquired signal is the previous frame of acquisition signal in the multiple frames of acquisition signals; and determine the power spectrum of the current acquisition signal based on the initial power value of each frame of acquisition signal and the prior power value of the corresponding previously acquired signal.

[0018] In one embodiment, the prediction module is further configured to, for each frame of the multi-frame acquisition signal, weight the prior power value corresponding to the previously acquired signal by a preset first weighting coefficient to obtain a first prior power value; select the maximum value from the first prior power value and the initial power value of the acquisition signal as the smoothed power value of the acquisition signal; perform smoothing processing on the smoothed power value of the acquisition signal and the prior power value corresponding to the previously acquired signal according to a preset first acquisition signal smoothing coefficient and a preset second acquisition signal smoothing coefficient to obtain the current power value of the acquisition signal; and determine the power spectrum of the current acquisition signal based on the current power value of each frame of the acquisition signal.

[0019] In one embodiment, the prediction module is further configured to perform frame segmentation processing on the simulated echo signal to obtain multiple frames of echo signals; for each frame of echo signal in the multiple frames of echo signals, determine the initial power value of the echo signal; obtain the prior power value of the preceding echo signal; the preceding echo signal is the previous frame echo signal of the echo signal in the multiple frames of echo signals; and determine the power spectrum of the simulated echo signal based on the initial power value of each frame of echo signal and the prior power value of the corresponding preceding echo signal.

[0020] In one embodiment, the prediction module is further configured to, for each frame of the multi-frame echo signal, weight the prior power value corresponding to the prior echo signal of the echo signal using a preset second weighting coefficient to obtain a second prior power value; select the maximum value between the second prior power value and the initial power value of the echo signal as the smoothed power value of the echo signal; perform smoothing processing on the smoothed power value of the echo signal and the prior power value corresponding to the prior echo signal according to a preset first echo signal smoothing coefficient and a preset second echo signal smoothing coefficient to obtain the current power value of the echo signal; and determine the power spectrum of the current echo signal based on the current power value of each frame of the echo signal.

[0021] In one embodiment, the estimated signal-to-return ratio (SRR) of the simulated echo signal includes the estimated SRR of each frequency point in the power spectrum of the simulated echo signal; the prediction module is further configured to determine a candidate power value for each frequency point in the power spectrum based on the difference between the power value of the acquired signal and the power value of the echo signal; the power value of the acquired signal is the power value in the power spectrum of the currently acquired signal corresponding to the frequency point; the power value of the echo signal is the power value in the power spectrum of the simulated echo signal corresponding to the frequency point; the maximum value is selected from the candidate power value and the preset minimum power value as the target power value; and the estimated SRR of the frequency point is determined based on the ratio of the target power value to the power value of the echo signal.

[0022] In one embodiment, the adjustment module is further configured to determine a power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and a preset signal-to-return ratio threshold; and to perform power spectrum adjustment processing on the current peer voice signal based on the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played.

[0023] In one embodiment, the adjustment module is further configured to determine a first candidate power spectrum adjustment coefficient based on the ratio of the estimated signal-to-return ratio to the preset signal-to-return ratio threshold and a preset minimum power spectrum adjustment coefficient; select the minimum value from the first candidate power spectrum adjustment coefficient and the second candidate power spectrum adjustment coefficient as the power spectrum adjustment coefficient for the current peer voice signal; the second candidate power spectrum adjustment coefficient is a coefficient that keeps the power spectrum of the current peer voice signal unchanged.

[0024] In one embodiment, the apparatus further includes:

[0025] A determination module is used to determine the signal correlation between the currently acquired signal and the prior voice signal to be played; the prior voice signal to be played is the voice signal to be played that precedes the current voice signal to be played and is time-aligned with the current acquired signal; the echo state of the voice call is determined based on the signal correlation.

[0026] The adjustment module is also used to determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state.

[0027] In one embodiment, the preset signal-to-return ratio (SRR) threshold includes a first preset SRR threshold set for a one-way talk state and a second preset SRR threshold set for a two-way talk state; the first preset SRR threshold is less than the second preset SRR threshold; the adjustment module is further configured to, if the echo state of the voice call is a one-way talk state, determine a power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the first preset SRR threshold; if the echo state of the voice call is a two-way talk state, determine a power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the second preset SRR threshold.

[0028] In one embodiment, the adjustment module is further configured to perform frequency domain transformation on the current peer voice signal to obtain an initial power spectrum corresponding to the current peer voice signal; the initial power spectrum includes the initial amplitude of each frequency point corresponding to the current peer voice signal; for each frequency point, the initial amplitude of the frequency point is weighted and calculated with the power spectrum adjustment coefficient corresponding to the frequency point to obtain the target amplitude of the frequency point; based on the target amplitude of each frequency point, the current peer voice signal is subjected to power spectrum adjustment processing; and the adjusted power spectrum is transformed in the time domain to obtain the adjusted current voice signal to be played.

[0029] In one embodiment, the cancellation module is further configured to perform echo simulation on the current audio signal to be played to obtain a target simulated echo signal; perform linear echo cancellation processing on the real echo signal generated after the current audio signal to be played is played based on the target simulated echo signal to obtain a residual signal including a nonlinear echo signal; and perform nonlinear echo cancellation processing on the nonlinear echo signal in the residual signal.

[0030] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the various method embodiments of this application.

[0031] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the various method embodiments of this application.

[0032] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the various method embodiments of this application.

[0033] The aforementioned echo control method, apparatus, device, medium, and computer program product acquire the current peer voice signal during a voice call, perform echo prediction processing on the current peer voice signal, and obtain a simulated echo signal. By predicting the estimated signal-to-return ratio (SRR) of the simulated echo signal, where the estimated SRR characterizes the energy proportion of the simulated echo signal in the currently acquired signal at the local end, and performing power spectrum adjustment processing on the current peer voice signal based on the estimated SRR, an adjusted current voice signal to be played can be obtained, wherein the energy of the current voice signal to be played is lower than the energy of the current peer voice signal. Then, echo cancellation processing is performed on the actual echo signal generated based on the current voice signal to be played. Compared to traditional methods that directly cancel existing echo signals, this application predicts a simulated echo signal based on the current peer's speech signal and estimates the signal-to-return ratio (SRR), which characterizes the energy proportion of the simulated echo signal in the current acquired signal at the local end. Based on the estimated SRR of the simulated echo signal, the power spectrum of the current peer's speech signal is first adjusted to avoid the occurrence of large echo signals. Then, the smaller echo signals generated by the adjusted current speech signal to be played are echo-cancelled. The echo cancellation effect is better, thereby improving the voice quality of the call. Attached Figure Description

[0034] Figure 1 This is a diagram illustrating the application environment of the echo control method in one embodiment.

[0035] Figure 2This is a flowchart illustrating an echo control method in one embodiment;

[0036] Figure 3 This is a block diagram illustrating the principle of traditional echo control methods.

[0037] Figure 4 This is a flowchart illustrating the process of predicting the signal-to-return ratio for a simulated echo signal in one embodiment.

[0038] Figure 5 This is a schematic diagram of the process for power spectrum adjustment of the current peer voice signal in one embodiment;

[0039] Figure 6 This is a block diagram illustrating the principle of the echo control method of this application in one embodiment;

[0040] Figure 7 This is a flowchart illustrating the echo control method in another embodiment;

[0041] Figure 8 This is a structural block diagram of an echo control device in one embodiment;

[0042] Figure 9 This is a structural block diagram of the echo control device in another embodiment;

[0043] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0045] The echo control method provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on other servers. Terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal 102 and server 104 can be directly or indirectly connected via wired or wireless communication; this application does not impose any restrictions on this connection.

[0046] Terminal 102 can obtain the current peer voice signal from server 104 during a voice call and perform echo prediction processing on the current peer voice signal to obtain a simulated echo signal. Terminal 102 can predict the estimated signal-to-return ratio (SRR) of the simulated echo signal, where the estimated SRR characterizes the energy proportion of the simulated echo signal in the currently acquired signal at this end. Terminal 102 can perform power spectrum adjustment processing on the current peer voice signal based on the estimated SRR to obtain an adjusted current voice signal to be played, where the energy of the current voice signal to be played is lower than the energy of the current peer voice signal. Terminal 102 can perform echo cancellation processing on the actual echo signal generated based on the current voice signal to be played.

[0047] In one embodiment, such as Figure 2 As shown, an echo control method is provided. This embodiment applies this method to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps:

[0048] Step 202: Obtain the current voice signal from the other end in the voice call.

[0049] Here, the current peer voice signal is the voice signal sent from the peer to this end at the current moment during a voice call. It can be understood that the terminal executing the echo control method of this application for echo cancellation is this end, and the other terminal conducting a voice call with this end is the peer end.

[0050] Specifically, this end can establish a communication connection with the other end through a server. During a voice call, the other end can send voice signals to this end, i.e., the current voice signal from the other end. This end can receive the current voice signal from the other end during the voice call.

[0051] Step 204: Perform echo prediction processing on the current peer voice signal to obtain a simulated echo signal.

[0052] Among them, the simulated echo signal is an echo signal simulated based on the current peer voice signal. It can be understood that the simulated echo signal is an echo signal simulated by this end based on the current peer voice signal before the current peer voice signal is played by the speaker of this end. The simulated echo signal is not the real echo signal actually generated after the current peer voice signal is played by this end.

[0053] In one embodiment, the terminal can input the current peer voice signal into a pre-set analog adaptive filter to filter the current peer voice signal and obtain an analog echo signal.

[0054] Step 206: Predict the estimated signal-to-return ratio of the analog echo signal; the estimated signal-to-return ratio is used to characterize the energy proportion of the analog echo signal in the currently acquired signal at this end.

[0055] Here, the currently acquired signal is the signal captured by the microphone at the local end at the current moment. The estimated signal-to-return ratio (SNR) is the ratio of the energy of the currently acquired signal to the energy of the analog echo signal. The energy proportion of the analog echo signal in the currently acquired signal refers to the proportion of the analog echo signal's energy within the total energy of the currently acquired signal. It can be understood that the larger the analog echo signal, the smaller the estimated SNR, indicating a larger energy proportion of the analog echo signal in the currently acquired signal.

[0056] In one embodiment, the terminal can acquire the currently acquired signal and determine the energy spectrum of the currently acquired signal. The terminal can determine the energy spectrum of the analog echo signal and, based on the energy spectrum of the currently acquired signal and the energy spectrum of the analog echo signal, determine the estimated signal-to-echo ratio of the analog echo signal.

[0057] In one embodiment, the terminal can perform frame segmentation on the currently acquired signal to obtain multiple frames of acquired signals. For each frame of acquired signals, the terminal can determine the initial energy value of the acquired signal. The terminal can obtain the prior energy value of previously acquired signals and determine the energy spectrum of the current acquired signal based on the initial energy value of each frame and the prior energy value of the corresponding previously acquired signals. Here, the previously acquired signal is the frame preceding the acquired signal in the multi-frame acquisition signal sequence.

[0058] In one embodiment, the terminal can perform frame-segmentation processing on the analog echo signal to obtain multiple frames of echo signals. For each frame of echo signals, the terminal can determine the initial energy value of the echo signal and obtain the prior energy value of the preceding echo signal. The terminal can determine the energy spectrum of the analog echo signal based on the initial energy value of each frame of echo signal and the corresponding prior energy value of the preceding echo signal. The preceding echo signal is the previous frame of echo signal in the multiple frames of echo signals.

[0059] Step 208: Based on the estimated signal-to-response ratio, perform power spectrum adjustment processing on the current peer voice signal to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal.

[0060] The power spectrum adjustment process involves adjusting the power values ​​in the power spectrum of the current peer's audio signal. The audio signal to be played is time-aligned with the current peer's audio signal. The power value of a signal can be understood as a representation of its energy; signals can carry energy, and the power value is a measure of that energy, indicating the energy absorbed or emitted by the signal per unit time. The higher the power value, the more energy the signal emits per unit time.

[0061] Specifically, the terminal can adjust the power value in the power spectrum of the current peer voice signal based on the estimated signal-to-return ratio of the analog echo signal to obtain the adjusted current voice signal to be played. In other words, the terminal can reduce the power value in the power spectrum of the current peer voice signal based on the estimated signal-to-return ratio of the analog echo signal, so that the energy of the adjusted current voice signal to be played is lower than the energy of the current peer voice signal, thus avoiding the generation of a large echo signal later.

[0062] In one embodiment, the terminal can determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio. Then, the terminal can perform power spectrum adjustment processing on the current peer voice signal based on the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played. The power spectrum adjustment coefficient is a coefficient used to adjust the power values ​​in the power spectrum of the current peer voice signal.

[0063] Step 210: Perform echo cancellation processing on the real echo signal generated based on the current audio signal to be played.

[0064] The true echo signal is the actual echo signal generated after the voice signal to be played is played by the speaker on this end.

[0065] Specifically, when the audio signal to be played is played by the local speaker, a real echo signal is generated. The terminal can then perform echo cancellation processing on the real echo signal generated based on the audio signal to be played.

[0066] In one embodiment, the terminal can perform linear echo cancellation processing on the real echo signal generated after the current audio signal to be played is played, so as to eliminate the linear echo component in the real echo signal.

[0067] In one embodiment, the terminal can perform nonlinear echo cancellation processing on the real echo signal generated after the current voice signal to be played is played, so as to eliminate the nonlinear echo component in the real echo signal.

[0068] In the aforementioned echo control method, by acquiring the current peer voice signal in a voice call and performing echo prediction processing on it, a simulated echo signal can be obtained. The estimated signal-to-return ratio (SRR) of the simulated echo signal is predicted, where the estimated SRR characterizes the energy proportion of the simulated echo signal in the currently acquired signal at the local end. Based on the estimated SRR, power spectrum adjustment processing is performed on the current peer voice signal to obtain the adjusted current voice signal to be played, where the energy of the current voice signal to be played is lower than that of the current peer voice signal. Finally, echo cancellation processing is performed on the actual echo signal generated based on the current voice signal to be played. Compared to traditional methods that directly cancel existing echo signals, this application predicts a simulated echo signal based on the current peer's speech signal and estimates the signal-to-return ratio (SRR), which characterizes the energy proportion of the simulated echo signal in the current acquired signal at the local end. Based on the estimated SRR of the simulated echo signal, the power spectrum of the current peer's speech signal is first adjusted to avoid the occurrence of large echo signals. Then, the smaller echo signals generated by the adjusted current speech signal to be played are echo-cancelled. The echo cancellation effect is better, thereby improving the voice quality of the call.

[0069] In traditional echo control methods, such as Figure 3 As shown, the terminal can receive the current voice signal from the other end and acquire the currently collected signal through the microphone. For example... Figure 3As shown, both the local and remote ends are speaking. The currently acquired signal can include the echo signal generated after the remote end's voice signal is played by the local end's speaker, and the current local voice signal generated by the local end's speech. The terminal can input the current acquired signal and the current remote end's voice signal to the echo delay detection unit. The echo delay detection unit determines the remote end's voice signal that is time-aligned with the current acquired signal, and inputs the aligned remote end's voice signal to an adaptive filter for filtering to obtain a simulated echo signal. The simulated echo signal is used to perform linear echo cancellation on the current acquired signal to obtain a residual signal. The terminal can input the residual signal to a nonlinear cancellation unit for nonlinear echo cancellation to obtain the transmission signal sent to the remote end. Traditional echo control methods directly cancel the generated echo signal to achieve echo cancellation. However, traditional echo control methods have poor echo cancellation effects, especially for large echo signals, easily resulting in echo residue and low voice quality.

[0070] The echo control method of this application does not directly eliminate the already generated echo signal. Instead, it first predicts a simulated echo signal based on the current peer voice signal and predicts the estimated signal-to-echo ratio, which can characterize the energy proportion of the simulated echo signal in the current acquired signal at the local end. Then, based on the estimated signal-to-echo ratio, the power spectrum of the current peer voice signal is adjusted to avoid the occurrence of large echo signals. Finally, the smaller echo signals generated by the adjusted current voice signal to be played are echo-cancelled. The echo cancellation effect is good, avoiding the problem of echo residue, thereby improving the voice quality of the call.

[0071] In one embodiment, such as Figure 4 As shown, the steps for predicting the estimated signal-to-return ratio of an analog echo signal include:

[0072] Step 402: Obtain the current acquisition signal from this end.

[0073] Specifically, the terminal may include a microphone responsible for acquiring signals, and the terminal can acquire the currently acquired signals through the microphone.

[0074] Step 404: Determine the power spectrum of the currently acquired signal.

[0075] It can be understood that the power spectrum of the currently acquired signal is one way of expressing the energy spectrum of the aforementioned currently acquired signal.

[0076] In one embodiment, the terminal can perform frame segmentation on the currently acquired signal to obtain multiple frames of acquired signals. For each frame of the multi-frame acquired signal, the terminal can determine the power value of that acquired signal. The terminal can determine the power spectrum of the currently acquired signal based on the power values ​​of each frame of the acquired signal. It can be understood that the power spectrum of the currently acquired signal includes the power values ​​of the acquired signal corresponding to each frequency point, and one frequency point can correspond to multiple frames of acquired signals.

[0077] For example, the terminal can perform frame segmentation on the currently acquired signal, taking each 20 milliseconds of the currently acquired signal as a frame, thus obtaining multiple frames of acquired signal.

[0078] In one embodiment, for each frame of a multi-frame acquisition signal, the terminal can perform frequency domain transformation on the acquisition signal. For example, the terminal can perform a Fast Fourier Transform on the acquisition signal to obtain the amplitude of the acquisition signal. Then, the terminal can determine the power value of the acquisition signal based on the amplitude of the acquisition signal.

[0079] Step 406: Determine the power spectrum of the analog echo signal.

[0080] It can be understood that the power spectrum of the analog echo signal is one way of expressing the energy spectrum of the analog echo signal.

[0081] In one embodiment, the terminal can perform frame segmentation processing on the analog echo signal to obtain multiple frames of echo signals. For each frame of the multi-frame echo signal, the power value of the echo signal is determined. The terminal can determine the power spectrum of the analog echo signal based on the power value of each frame of the echo signal. It can be understood that the power spectrum of the analog echo signal includes the power values ​​of the echo signals corresponding to each frequency point, and one frequency point can correspond to multiple frames of echo signals.

[0082] For example, the terminal can perform frame processing on the analog echo signal, taking each 20 milliseconds of the analog echo signal as one frame of echo signal, thus obtaining multiple frames of echo signal.

[0083] In one embodiment, for each frame of the multi-frame echo signal, the terminal can perform frequency domain transformation on the echo signal. For example, the terminal can perform a Fast Fourier Transform on the echo signal to obtain the amplitude of the echo signal. Then, the terminal can determine the power value of the echo signal based on the amplitude of the echo signal.

[0084] Step 408: Determine the estimated signal-to-return ratio of the analog echo signal based on the power spectrum of the currently acquired signal and the power spectrum of the analog echo signal.

[0085] In one embodiment, the estimated signal-to-return ratio (SRR) of the simulated echo signal includes the estimated SRR at each frequency point in the power spectrum of the simulated echo signal. For each frequency point in the power spectrum, the terminal can determine the estimated SRR at that frequency point based on the ratio of the acquired signal power value to the echo signal power value.

[0086] In one embodiment, for each frequency point in the power spectrum, the terminal directly uses the ratio of the acquired signal power value to the echo signal power value as the estimated signal-to-echo ratio for that frequency point.

[0087] In the above embodiments, by using the power spectrum of the currently acquired signal and the power spectrum of the analog echo signal to determine the estimated signal-to-echo ratio of the analog echo signal, the accuracy of the estimated signal-to-echo ratio can be improved.

[0088] In one embodiment, determining the power spectrum of the currently acquired signal includes: performing frame segmentation on the currently acquired signal to obtain multiple frames of acquired signals; determining the initial power value of the acquired signal for each frame of the multiple frames of acquired signals; obtaining the prior power value of the previously acquired signal; the previously acquired signal is the previous frame of the acquired signal in the multiple frames of acquired signals; and determining the power spectrum of the currently acquired signal based on the initial power value of each frame of acquired signal and the prior power value of the corresponding previously acquired signal.

[0089] The initial power value of the acquired signal is determined based on the amplitude of the acquired signal. This can be understood as the amplitude of the acquired signal being obtained through frequency domain transformation. The earlier acquired signal's power value is the power value of the previously acquired signal.

[0090] Specifically, the terminal can perform frame segmentation on the currently acquired signal to obtain multiple frames of acquired signals. For each frame of the multi-frame acquired signal, the terminal can determine the initial power value of the acquired signal and obtain the prior power value of the previously acquired signal. The terminal can determine the power spectrum of the currently acquired signal based on the initial power value of each frame and the prior power value of the corresponding previously acquired signal. It can be understood that the power spectrum of the currently acquired signal includes the current power value of each frequency point in the power spectrum of the currently acquired signal, and one frequency point can correspond to multiple frames of acquired signals. It can be understood that the power spectrum of the currently acquired signal includes the current power value of each frame of the multi-frame acquired signal. For each frame of the multi-frame acquired signal, the terminal can determine the initial power value of the acquired signal and obtain the prior power value of the previously acquired signal. Furthermore, the terminal can determine the current power value of the acquired signal based on the initial power value of the acquired signal and the prior power value of the corresponding previously acquired signal.

[0091] In one embodiment, for each frame of a multi-frame acquisition signal, the terminal can perform a weighted summation of the initial power value of the acquisition signal and the previous power value of the corresponding previously acquired signal to obtain the current power value of the acquisition signal.

[0092] In the above embodiments, by determining the power spectrum of the current acquisition signal based on the initial power value of each frame of the acquired signal and the previous power value of the corresponding previously acquired signal, it is possible to avoid drastic changes in the power values ​​of adjacent frequency points in the power spectrum of the current acquisition signal, making the power spectrum of the current acquisition signal smoother and further improving the accuracy of the predicted signal-to-return ratio.

[0093] In one embodiment, determining the power spectrum of the current acquired signal based on the initial power value of each frame of the acquired signal and the prior power value of the corresponding previously acquired signal includes: for each frame of the acquired signal in the multi-frame acquired signal, weighting the prior power value corresponding to the previously acquired signal of the acquired signal by a preset first weighting coefficient to obtain a first prior power value; selecting the maximum value from the first prior power value and the initial power value of the acquired signal as the smoothed power value of the acquired signal; smoothing the smoothed power value of the acquired signal and the prior power value corresponding to the previously acquired signal of the acquired signal according to a preset first acquisition signal smoothing coefficient and a preset second acquisition signal smoothing coefficient to obtain the current power value of the acquired signal; and determining the power spectrum of the current acquired signal based on the current power value of each frame of the acquired signal.

[0094] Here, the first weighting coefficient is used to weight the power values ​​corresponding to previously acquired signals of the acquired signal. The first prior power value is the power value obtained by weighting the power values ​​corresponding to previously acquired signals of the acquired signal using the first weighting coefficient. The smoothed power value of the acquired signal is the maximum power value selected from the first prior power value and the initial power value of the acquired signal. The first and second smoothing coefficients of the acquired signal are coefficients used to smooth the smoothed power value of the acquired signal and the power values ​​corresponding to previously acquired signals of the acquired signal.

[0095] Specifically, for each frame of the multi-frame acquisition signal, the terminal can weight the prior power value corresponding to the previously acquired signal using a preset first weighting coefficient to obtain a first prior power value. The terminal can compare the first prior power value with the initial power value of the acquisition signal to select the maximum value as the smoothed power value of the acquisition signal. The terminal can then smooth the smoothed power value of the acquisition signal and the prior power value corresponding to the previously acquired signal using preset first and second acquisition signal smoothing coefficients to obtain the current power value of the acquisition signal. Furthermore, the terminal can determine the power spectrum of the current acquisition signal based on the current power value of each frame. It can be understood that the power spectrum of the current acquisition signal includes the current power values ​​of each frame of the multi-frame acquisition signal.

[0096] In one embodiment, the current power value of each frame of acquired signal can be calculated using the following formula:

[0097] Ed(i,j)=a1*Ed(i-1,j)+(1-a1)*max(b1*Ed(i-1,j),Ed0(i,j))

[0098] Where Ed(i,j) represents the current power value corresponding to the j-th frequency point of the i-th frame acquisition signal. b1 represents the preset first weighting coefficient, and Ed(i-1,j) represents the prior power value corresponding to the j-th frequency point of the prior acquisition signal of the i-th frame acquisition signal. It can be understood that b1*Ed(i-1,j) is the first prior power value. Ed0(i,j) represents the initial power value corresponding to the j-th frequency point of the i-th frame acquisition signal. It can be understood that max(b1*Ed(i-1,j), Ed0(i,j)) is the smoothed power value of the acquisition signal. a1 represents the first acquisition signal smoothing coefficient. (1-a1) represents the second acquisition signal smoothing coefficient. It should be noted that the values ​​of a1 and b1 are in the range of (0,1).

[0099] In the above embodiments, a first prior power value can be obtained by weighting the prior power value corresponding to the previously acquired signal using a preset first weighting coefficient. The smoothed power value of the acquired signal can be obtained by selecting the maximum value from the first prior power value and the initial power value of the acquired signal. By smoothing the smoothed power value of the acquired signal and the prior power value corresponding to the previously acquired signal using preset first and second acquisition signal smoothing coefficients, a more stable current power value of the acquired signal can be obtained. Furthermore, based on the relatively stable current power value of each frame of the acquired signal, the smoothness of the power spectrum of the current acquired signal can be further improved, thereby further improving the accuracy of the estimated signal-to-return ratio.

[0100] In one embodiment, determining the power spectrum of the analog echo signal includes: performing frame segmentation on the analog echo signal to obtain multiple frames of echo signals; determining the initial power value of the echo signal for each frame of the echo signal in the multiple frames of echo signals; obtaining the prior power value of the preceding echo signal; the preceding echo signal is the echo signal of the previous frame in the multiple frames of echo signals; and determining the power spectrum of the analog echo signal based on the initial power value of each frame of echo signal and the prior power value of the corresponding preceding echo signal.

[0101] The initial power value of the echo signal is determined based on the amplitude of the echo signal. This can be understood as the amplitude of the echo signal being obtained through frequency domain transformation. The earlier power value of the previous echo signal is the power value of the earlier echo signal.

[0102] Specifically, the terminal can perform frame segmentation on the current echo signal to obtain multiple frames of echo signals. For each frame of the echo signal, the terminal can determine the initial power value of the echo signal and obtain the prior power value of the preceding echo signal. The terminal can determine the power spectrum of the current echo signal based on the initial power value of each frame and the corresponding prior power value of the preceding echo signal. It can be understood that the power spectrum of the current echo signal includes the current power value of each frequency point in the power spectrum, and one frequency point can correspond to multiple frames of echo signals. It can be understood that the power spectrum of the current echo signal includes the current power value of each frame of the echo signal. For each frame of the echo signal, the terminal can determine the initial power value of that echo signal and obtain the prior power value of the preceding echo signal. Furthermore, the terminal can determine the current power value of that echo signal based on the initial power value of that echo signal and the corresponding prior power value of the preceding echo signal.

[0103] In one embodiment, for each frame of echo signal in a multi-frame echo signal, the terminal can perform a weighted summation of the initial power value of the echo signal and the prior power value of the corresponding prior echo signal to obtain the current power value of the echo signal.

[0104] In the above embodiments, by determining the power spectrum of the simulated echo signal based on the initial power value of each frame of echo signal and the prior power value of the corresponding prior echo signal, it is possible to avoid drastic changes in the power values ​​of adjacent frequency points in the power spectrum of the simulated echo signal, making the power spectrum of the simulated echo signal smoother and further improving the accuracy of the predicted signal-to-echo ratio.

[0105] In one embodiment, determining the power spectrum of the simulated echo signal based on the initial power value of each frame of the echo signal and the prior power value of the corresponding prior echo signal includes: for each frame of the echo signal in the multi-frame echo signal, weighting the prior power value corresponding to the prior echo signal of the echo signal by a preset second weighting coefficient to obtain a second prior power value; selecting the maximum value from the second prior power value and the initial power value of the echo signal as the smoothed power value of the echo signal; smoothing the smoothed power value of the echo signal and the prior power value corresponding to the prior echo signal of the echo signal according to a preset first echo signal smoothing coefficient and a preset second echo signal smoothing coefficient to obtain the current power value of the echo signal; and determining the power spectrum of the current echo signal based on the current power value of each frame of the echo signal.

[0106] The second weighting coefficient is used to weight the prior power values ​​corresponding to the preceding echo signals of the echo signal. The second prior power value is the power value obtained by weighting the prior power values ​​corresponding to the preceding echo signals of the echo signal using the second weighting coefficient. The smoothed echo signal power value is the maximum power value selected from the second prior power value and the initial power value of the echo signal. The first and second echo signal smoothing coefficients are coefficients used to smooth the smoothed echo signal power value and the prior power values ​​corresponding to the preceding echo signals of the echo signal.

[0107] Specifically, for each echo signal in a multi-frame echo signal, the terminal can weight the prior power value corresponding to the preceding echo signal using a preset second weighting coefficient to obtain a second prior power value. The terminal can compare the second prior power value with the initial power value of the echo signal, selecting the maximum value as the smoothed echo signal power value. The terminal can smooth the smoothed echo signal power value and the preceding power value corresponding to the preceding echo signal using preset first and second echo signal smoothing coefficients to obtain the current power value of the echo signal. Furthermore, the terminal can determine the power spectrum of the current echo signal based on the current power value of each echo signal frame. It can be understood that the power spectrum of the current echo signal includes the current power values ​​of each echo signal frame in the multi-frame echo signal.

[0108] In one embodiment, the current power value of each frame of the echo signal can be calculated using the following formula:

[0109] Ey(i,j)=a2*Ey(i-1,j)+(1-a2)*max(b2*Ey(i-1,j),Ey0(i,j))

[0110] Where Ey(i,j) represents the current power value corresponding to the j-th frequency point of the i-th frame echo signal. B2 represents the preset second weighting coefficient, and Ey(i-1,j) represents the prior power value corresponding to the j-th frequency point of the prior echo signal of the i-th frame echo signal. It can be understood that b2*Ey(i-1,j) is the second prior power value. Ey0(i,j) represents the initial power value corresponding to the j-th frequency point of the i-th frame echo signal. It can be understood that max(b2*Ey(i-1,j), Ey0(i,j)) is the smoothed power value of the echo signal. a2 represents the first echo signal smoothing coefficient. (1-a2) represents the second echo signal smoothing coefficient. It should be noted that the values ​​of a2 and b2 are in the range of (0,1).

[0111] In the above embodiments, a second prior power value can be obtained by weighting the prior power value corresponding to the prior echo signal of the echo signal using a preset second weighting coefficient. The smoothed power value of the echo signal can be obtained by selecting the maximum value from the second prior power value and the initial power value of the echo signal. By smoothing the smoothed power value of the echo signal and the prior power value corresponding to the prior echo signal according to preset first and second echo signal smoothing coefficients, a more stable current power value of the echo signal is obtained. Furthermore, based on the relatively stable current power value of each frame of the echo signal, the smoothness of the power spectrum of the current echo signal can be further improved, thereby further improving the accuracy of the estimated signal-to-return ratio.

[0112] In one embodiment, the estimated signal-to-return ratio (SPR) of the simulated echo signal includes the estimated SPR at each frequency point in the power spectrum of the simulated echo signal. Determining the estimated SPR of the simulated echo signal based on the power spectrum of the currently acquired signal and the power spectrum of the simulated echo signal includes: for each frequency point in the power spectrum, determining a candidate power value based on the difference between the power value of the acquired signal and the power value of the echo signal; the power value of the acquired signal is the power value corresponding to the frequency point in the power spectrum of the currently acquired signal; the power value of the echo signal is the power value corresponding to the frequency point in the power spectrum of the simulated echo signal; selecting the maximum value from the candidate power value and a preset minimum power value as the target power value; and determining the estimated SPR at the frequency point based on the ratio of the target power value to the echo signal power value.

[0113] The candidate power value is determined based on the difference between the power value of the acquired signal and the power value of the echo signal. The target power value is the largest power value selected from the candidate power values ​​and the preset minimum power value.

[0114] Specifically, the power spectrum of the currently acquired signal includes the power values ​​corresponding to each frequency point in the power spectrum. The power spectrum of the simulated echo signal includes the power values ​​corresponding to each frequency point in the power spectrum. The estimated signal-to-return ratio (SRR) of the simulated echo signal includes the estimated SRR for each frequency point in the power spectrum. For each frequency point in the power spectrum, the terminal can determine a candidate power value for that frequency point based on the difference between the power value of the acquired signal and the power value of the echo signal. The terminal can select the maximum value from the candidate power value and the preset minimum power value as the target power value for that frequency point. The terminal can determine the estimated SRR for that frequency point based on the ratio of the target power value to the echo signal power value.

[0115] In one embodiment, the estimated signal-to-return ratio at each frequency point in the power spectrum of the simulated echo signal can be calculated using the following formula:

[0116]

[0117] Where ser(i,j) represents the estimated signal-to-return ratio corresponding to the j-th frequency point of the i-th frame echo signal in the analog echo signal. Emin(j) represents the preset minimum power value corresponding to the j-th frequency point.

[0118] In the above embodiments, for each frequency point in the power spectrum, a candidate power value can be determined based on the difference between the acquired signal power value and the echo signal power value. The target power value is obtained by selecting the maximum value from the candidate power value and the preset minimum power value. Furthermore, the estimated signal-to-return ratio of the frequency point can be determined based on the ratio of the target power value to the echo signal power value, thereby further improving the accuracy of the estimated signal-to-return ratio.

[0119] In one embodiment, such as Figure 5 As shown, the steps for performing power spectrum adjustment on the current peer-end voice signal based on the estimated signal-to-return ratio to obtain the adjusted current voice signal to be played include:

[0120] Step 502: Determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold.

[0121] In one embodiment, the terminal can compare the estimated signal-to-return ratio (SRR) with a preset SRR threshold to obtain a comparison result. Then, based on the comparison result, the terminal can determine the power spectrum adjustment coefficient for the current peer's voice signal.

[0122] In one embodiment, if the comparison result shows that the estimated signal-to-return ratio is greater than or equal to a preset signal-to-return ratio threshold, then the power spectrum adjustment coefficient for the current peer voice signal is 1. If the comparison result shows that the estimated signal-to-return ratio is less than the preset signal-to-return ratio threshold, then the ratio of the estimated signal-to-return ratio to the preset signal-to-return ratio threshold is directly used as the power spectrum adjustment coefficient for the current peer voice signal.

[0123] Step 504: Perform power spectrum adjustment processing on the current peer voice signal according to the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played.

[0124] Specifically, the terminal can perform power spectrum adjustment processing on the power spectrum of the current peer voice signal according to the power spectrum adjustment coefficient, and determine the voice signal to be played based on the adjusted power spectrum.

[0125] In the above embodiments, by estimating the signal-to-return ratio and setting a preset signal-to-return ratio threshold, the power spectrum adjustment coefficient for the current peer voice signal can be determined. Then, the power spectrum adjustment processing of the current peer voice signal is performed according to the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played, which improves the accuracy of the current voice signal to be played and can effectively avoid the occurrence of large echo signals, thereby further improving the echo cancellation effect.

[0126] In one embodiment, determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and a preset signal-to-return ratio threshold includes: determining a first candidate power spectrum adjustment coefficient based on the ratio of the estimated signal-to-return ratio to the preset signal-to-return ratio threshold and a preset minimum power spectrum adjustment coefficient; selecting the minimum value from the first candidate power spectrum adjustment coefficient and the second candidate power spectrum adjustment coefficient as the power spectrum adjustment coefficient for the current peer voice signal; the second candidate power spectrum adjustment coefficient is a coefficient that keeps the power spectrum of the current peer voice signal unchanged.

[0127] The first candidate power spectrum adjustment coefficient is a coefficient determined based on the ratio of the estimated signal-to-response ratio to the preset signal-to-response ratio threshold and the preset minimum power spectrum adjustment coefficient, and is used to calculate the power spectrum adjustment coefficient.

[0128] Specifically, the terminal can determine the ratio of the estimated signal-to-return ratio (SRR) to a preset SRR threshold, and calculate a first candidate power spectrum adjustment coefficient based on this ratio and a preset minimum power spectrum adjustment coefficient. The terminal can then select the minimum value from the first and second candidate power spectrum adjustment coefficients as the power spectrum adjustment coefficient for the current peer's voice signal. The second candidate power spectrum adjustment coefficient is a coefficient that keeps the power spectrum of the current peer's voice signal unchanged. It can be understood that the second candidate power spectrum adjustment coefficient is 1.

[0129] In one embodiment, the power spectrum adjustment coefficient for the current peer voice signal can be calculated using the following formula:

[0130]

[0131] Where g(i,j) represents the power spectrum adjustment coefficient corresponding to the j-th frequency point of the i-th frame of the current peer's speech signal. ser This indicates the preset signal-to-response ratio threshold. Gmin represents the preset minimum power spectrum adjustment coefficient. 1 represents the second candidate power spectrum adjustment coefficient. This can be understood as... This is the adjustment coefficient for the first candidate power spectrum.

[0132] In the above embodiments, by estimating the ratio of the signal-to-return ratio to a preset signal-to-return ratio threshold and using a preset minimum power spectrum adjustment coefficient, a first candidate power spectrum adjustment coefficient can be determined. Then, by selecting the minimum value from the first and second candidate power spectrum adjustment coefficients, the power spectrum adjustment coefficient for the current peer voice signal can be obtained, improving the accuracy of the power spectrum adjustment coefficient and further enhancing the echo cancellation effect. Simultaneously, by setting a preset minimum power spectrum adjustment coefficient, excessive attenuation of the current peer voice signal can be avoided, preventing distortion of the played sound.

[0133] In one embodiment, the method further includes: determining the signal correlation between the currently acquired signal and the prior voice signal to be played; the prior voice signal to be played is a voice signal to be played that precedes the current voice signal to be played and is time-aligned with the currently acquired signal; determining the echo state of the voice call based on the signal correlation; and determining a power spectrum adjustment coefficient for the current peer voice signal according to an estimated signal-to-return ratio and a preset signal-to-return ratio threshold, including: determining the power spectrum adjustment coefficient for the current peer voice signal according to the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state.

[0134] Signal correlation is used to characterize the degree of correlation between the currently acquired signal and the previously played audio signal. The echo state of a voice call includes one-way and two-way states. One-way state refers to a voice call where only the other end is speaking; in this state, the currently acquired signal only includes the echo signal generated after the previously played audio signal is played. Two-way state refers to a voice call where both the other end and the local end are speaking simultaneously; in this state, the currently acquired signal includes the echo signal generated after the previously played audio signal is played, as well as the current local audio signal generated when the local end is speaking.

[0135] Specifically, the terminal can determine the signal correlation between the currently acquired signal and the previously played voice signal, and determine the echo state of the current voice call based on the signal correlation. It can be understood that different echo states correspond to different preset signal-to-return ratio (SRR) thresholds. The terminal can determine the power spectrum adjustment coefficient for the current peer's voice signal based on the estimated SRR and the preset SRR threshold corresponding to the echo state of the current voice call.

[0136] In the above embodiments, by determining the signal correlation between the currently acquired signal and the previously played voice signal, and determining the echo state of the voice call based on the signal correlation, and then determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state, different echo cancellation strengths can be adopted for different echo states, thereby ensuring the voice call quality while ensuring the echo cancellation effect.

[0137] In one embodiment, the preset signal-to-return ratio (SRR) threshold includes a first preset SRR threshold set for a one-way talk state and a second preset SRR threshold set for a two-way talk state; the first preset SRR threshold is less than the second preset SRR threshold; determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the preset SRR threshold includes: if the echo state of the voice call is a one-way talk state, then determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the first preset SRR threshold; if the echo state of the voice call is a two-way talk state, then determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the second preset SRR threshold.

[0138] The first preset signal-to-response ratio threshold is a preset signal-to-response ratio threshold set for single-talk mode. The second preset signal-to-response ratio threshold is a preset signal-to-response ratio threshold set for dual-talk mode.

[0139] Specifically, if the echo state of the current voice call is one-way, the terminal can determine the power spectrum adjustment coefficient for the current peer's voice signal based on the estimated signal-to-return ratio and a first preset signal-to-return ratio threshold set for one-way. If the echo state of the current voice call is two-way, the terminal can determine the power spectrum adjustment coefficient for the current peer's voice signal based on the estimated signal-to-return ratio and a second preset signal-to-return ratio threshold set for two-way.

[0140] In the above embodiments, if the echo state of the voice call is a one-way state, the power spectrum adjustment coefficient for the current peer voice signal is determined by estimating the signal-to-return ratio (SRR) and using a small first preset SRR threshold. This ensures the echo cancellation effect in the one-way state while preventing severe distortion of the current peer voice signal, further guaranteeing the voice call quality. If the echo state of the voice call is a two-way state, the power spectrum adjustment coefficient for the current peer voice signal is determined by estimating the SRR and using a second preset SRR threshold, ensuring the echo cancellation effect in the two-way state.

[0141] In one embodiment, power spectrum adjustment processing is performed on the current peer-end voice signal according to the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played. This includes: performing frequency domain transformation on the current peer-end voice signal to obtain the initial power spectrum corresponding to the current peer-end voice signal; the initial power spectrum includes the initial amplitude of each frequency point corresponding to the current peer-end voice signal; for each frequency point, the initial amplitude of the frequency point is weighted and calculated with the power spectrum adjustment coefficient corresponding to the frequency point to obtain the target amplitude of the frequency point; based on the target amplitude of each frequency point, the power spectrum adjustment processing is performed on the current peer-end voice signal; and the adjusted power spectrum is transformed in the time domain to obtain the adjusted current voice signal to be played.

[0142] Here, the current peer voice signal and the voice signal to be played are time-domain signals. The initial power spectrum corresponding to the current peer voice signal is the power spectrum obtained after frequency-domain transformation of the current peer voice signal. The initial amplitude is the amplitude corresponding to each frequency point in the initial power spectrum of the current peer voice signal. The target amplitude is the amplitude obtained after adjusting the initial power spectrum using power spectrum adjustment coefficients.

[0143] Specifically, the terminal can perform frequency domain conversion on the current peer-end voice signal, which belongs to the time domain signal, to obtain the initial power spectrum corresponding to the current peer-end voice signal. The initial power spectrum includes the initial amplitude of each frequency point corresponding to the current peer-end voice signal. For each frequency point, the terminal can perform a weighted calculation with the initial amplitude of the frequency point and the corresponding power spectrum adjustment coefficient to obtain the target amplitude of the frequency point. Based on the target amplitude of each frequency point, the terminal can determine the adjusted power value corresponding to each frequency point to achieve power spectrum adjustment processing of the current peer-end voice signal. Furthermore, the terminal can perform time domain conversion on the adjusted power spectrum to obtain the adjusted, time-domain signal of the current voice signal to be played.

[0144] In the above embodiments, by performing frequency domain transformation on the current peer voice signal, the initial power spectrum corresponding to the current peer voice signal can be obtained. For each frequency point, the target amplitude of the frequency point can be obtained by weighting the initial amplitude of the frequency point with the power spectrum adjustment coefficient corresponding to the frequency point. Then, based on the target amplitude of each frequency point, the power spectrum of the current peer voice signal can be adjusted, and the adjusted power spectrum can be transformed in the time domain to obtain the adjusted current voice signal to be played, further improving the accuracy of the current voice signal to be played.

[0145] In one embodiment, echo cancellation processing is performed on the real echo signal generated based on the current speech signal to be played, including: performing echo simulation on the current speech signal to be played to obtain a target simulated echo signal; performing linear echo cancellation processing on the real echo signal generated after the current speech signal to be played is played according to the target simulated echo signal to obtain a residual signal including a nonlinear echo signal; and performing nonlinear echo cancellation processing on the nonlinear echo signal in the residual signal.

[0146] The target simulated echo signal is the echo signal obtained by simulating the echo of the current audio signal to be played. Linear echo cancellation processing refers to the processing method of eliminating the linear echo signal portion in the target simulated echo signal. Nonlinear echo cancellation processing refers to the processing method of eliminating the nonlinear echo signal portion in the residual signal. The residual signal includes at least the nonlinear echo signal in the target simulated echo signal. It can be understood that if the current voice call is in a one-way echo state, the residual signal includes the nonlinear echo signal in the target simulated echo signal. If the current voice call is in a two-way echo state, the residual signal includes the nonlinear echo signal in the target simulated echo signal, as well as the current local audio signal generated when the speaker is speaking.

[0147] Specifically, the terminal can input the currently playing audio signal to a pre-set adaptive filter to filter the audio signal and obtain a target simulated echo signal. Based on the target simulated echo signal, the terminal can perform linear echo cancellation processing on the actual echo signal generated after the currently playing audio signal is played, obtaining a residual signal including nonlinear echo signals. The terminal can then input the residual signal to a pre-set nonlinear cancellation unit to perform nonlinear echo cancellation processing on the nonlinear echo signals in the residual signal.

[0148] In one embodiment, the terminal can determine the difference between the current acquired signal and the target analog echo signal, and based on the difference between the current acquired signal and the target analog echo signal, determine a residual signal including the nonlinear echo signal.

[0149] In one embodiment, the terminal may use the difference between the currently acquired signal and the target analog echo signal as a residual signal that includes the nonlinear echo signal.

[0150] In the above embodiments, by using the target simulated echo signal, linear echo cancellation processing can be performed on the real echo signal generated after the current voice signal to be played is played, resulting in a residual signal including nonlinear echo signals. Then, nonlinear echo cancellation processing is performed on the nonlinear echo signals in the residual signal, which can further improve the echo cancellation effect and thus further improve the voice quality of the call.

[0151] In one embodiment, such as Figure 6 As shown, in the echo control method of this application, the terminal can receive the current peer voice signal and acquire the current acquisition signal through a microphone. The terminal can input the current peer voice signal to an analog adaptive filter for filtering to obtain an analog echo signal. The terminal can input the current acquisition signal and the analog echo signal to a signal-to-return ratio (SRR) estimation unit to predict the estimated SRR of the analog echo signal. The terminal can input the current acquisition signal to an echo delay detection unit to determine the prior voice signal to be played that is time-aligned with the current acquisition signal. The aligned prior voice signal to be played and the current acquisition signal are input to a single / dual-talk detection unit to determine the echo state of the current voice call. The power spectrum adjustment coefficient estimation unit determines the power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the preset SRR threshold corresponding to the echo state, and performs power spectrum adjustment processing on the current peer voice signal using the power spectrum adjustment coefficient to obtain the current voice signal to be played. Furthermore, the terminal can perform echo cancellation processing on the real echo signal generated after the current audio signal to be played is played. Specifically, the terminal can input a subsequently acquired signal, which includes at least the real echo signal, to the echo delay detection unit, so that the echo delay detection unit can determine the current audio signal to be played that is time-aligned with the subsequently acquired signal. It can be understood that the subsequently acquired signal is the acquired signal that follows the current acquired signal and is time-aligned with the current audio signal to be played. The terminal can input the current audio signal to be played to an adaptive filter for filtering to obtain the target simulated echo signal. Based on the target simulated echo signal, the terminal can perform linear echo cancellation processing on the real echo signal generated after the current audio signal to be played is played, obtaining a residual signal including nonlinear echo signals. Furthermore, the terminal can input the residual signal to a nonlinear cancellation unit for nonlinear echo cancellation processing of the nonlinear echo signals in the residual signal, obtaining the transmission signal to be sent to the other end.

[0152] Compared to Figure 3 Traditional echo control methods, such as Figure 6 As shown, the echo control method of this application does not directly eliminate the already generated echo signal. Instead, it first predicts a simulated echo signal based on the current peer voice signal and predicts the estimated signal-to-echo ratio, which can characterize the energy proportion of the simulated echo signal in the current acquired signal at the local end. Then, based on the estimated signal-to-echo ratio, the power spectrum of the current peer voice signal is adjusted to avoid the occurrence of large echo signals. Finally, the smaller echo signals generated by the adjusted current voice signal to be played are echo-cancelled. The echo cancellation effect is good, avoiding the problem of echo residue, thereby improving the voice quality of the call.

[0153] like Figure 7 As shown, in one embodiment, an echo control method is provided. This embodiment applies this method to... Figure 1 Taking terminal 102 as an example, the method specifically includes the following steps:

[0154] Step 702: Obtain the current peer voice signal in the voice call, and perform echo prediction processing on the current peer voice signal to obtain a simulated echo signal.

[0155] Step 704: Obtain the current acquisition signal from the local end, and perform frame segmentation processing on the current acquisition signal to obtain multiple frames of acquisition signal.

[0156] Step 706: For each frame of the multi-frame acquisition signal, determine the initial power value of the acquisition signal.

[0157] Step 708: Obtain the prior power value of the previously acquired signal; the previously acquired signal is the previous frame acquisition signal in the multi-frame acquisition signal.

[0158] Step 710: Determine the power spectrum of the current acquisition signal based on the initial power value of each frame of acquired signal and the prior power value of the corresponding previously acquired signal.

[0159] Step 712: Perform frame segmentation on the analog echo signal to obtain multi-frame echo signals.

[0160] Step 714: For each frame of echo signal in the multi-frame echo signal, determine the initial power value of the echo signal.

[0161] Step 716: Obtain the prior power value of the prior echo signal; the prior echo signal is the previous frame echo signal in the multi-frame echo signal.

[0162] Step 718: Determine the power spectrum of the analog echo signal based on the initial power value of each frame of echo signal and the prior power value of the corresponding prior echo signal.

[0163] Step 720: Determine the estimated signal-to-return ratio of the analog echo signal based on the power spectrum of the currently acquired signal and the power spectrum of the analog echo signal; the estimated signal-to-return ratio is used to characterize the energy proportion of the analog echo signal in the currently acquired signal at this end.

[0164] Step 722: Determine the signal correlation between the currently acquired signal and the previously acquired audio signal to be played; the previously acquired audio signal to be played is the audio signal to be played that precedes the currently acquired audio signal and is time-aligned with the currently acquired signal.

[0165] Step 724: Determine the echo state of the voice call based on signal correlation.

[0166] Step 726: Determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state.

[0167] Step 728: Perform power spectrum adjustment processing on the current peer voice signal according to the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal.

[0168] Step 730: Perform echo simulation on the current audio signal to be played to obtain the target simulated echo signal.

[0169] Step 732: Based on the target simulated echo signal, perform linear echo cancellation processing on the real echo signal generated after the current audio signal to be played is played, to obtain a residual signal including the nonlinear echo signal.

[0170] Step 734: Perform nonlinear echo cancellation processing on the nonlinear echo signal in the residual signal.

[0171] This application also provides an application scenario in which the echo control method described above is applied. Specifically, the echo control method can be applied to echo cancellation scenarios during instant messaging. It can be understood that the echo control method is applied to an instant messaging application running on a terminal. The terminal can acquire the current real-time voice signal from the other end in an instant voice call and perform echo prediction processing on the current real-time voice signal from the other end to obtain a simulated echo signal. The terminal acquires the current acquisition signal and performs frame segmentation processing on the current acquisition signal to obtain multiple frames of acquisition signals. For each frame of the multi-frame acquisition signal, the initial power value of the acquisition signal is determined. The prior power value of the previously acquired signal is acquired; the prior acquired signal is the previous frame of the acquisition signal in the multi-frame acquisition signal. Based on the initial power value of each frame of acquisition signal and the prior power value of the corresponding previously acquired signal, the power spectrum of the current acquisition signal is determined. The simulated echo signal is processed by framing to obtain multiple frames of echo signals. For each frame of the multi-frame echo signal, the initial power value of the echo signal is determined. Obtain the prior power value of the preceding echo signal; the preceding echo signal is the echo signal of the previous frame in a multi-frame echo signal. Determine the power spectrum of the analog echo signal based on the initial power value of each frame echo signal and the prior power value of the corresponding preceding echo signal.

[0172] The terminal can determine the estimated signal-to-return ratio (SRR) of the analog echo signal based on the power spectrum of the currently acquired signal and the power spectrum of the analog echo signal. The estimated SRR characterizes the energy proportion of the analog echo signal in the currently acquired signal at this end. It then determines the signal correlation between the currently acquired signal and the preceding real-time voice signal to be played. The preceding real-time voice signal to be played is the real-time voice signal that precedes the current real-time voice signal to be played and is time-aligned with the current acquired signal. Based on the signal correlation, the echo state of the real-time voice call is determined. The terminal can determine the power spectrum adjustment coefficient for the current peer's real-time voice signal based on the estimated SRR and the preset SRR threshold corresponding to the echo state. The power spectrum of the current peer's real-time voice signal is adjusted according to the power spectrum adjustment coefficient to obtain the adjusted current real-time voice signal to be played; the energy of the current real-time voice signal to be played is lower than the energy of the current peer's real-time voice signal.

[0173] The terminal can simulate the echo of the currently playing real-time audio signal to obtain a target simulated echo signal. Based on the target simulated echo signal, linear echo cancellation processing is performed on the real echo signal generated after the currently playing real-time audio signal is played, resulting in a residual signal that includes nonlinear echo signals. Nonlinear echo cancellation processing is then performed on the nonlinear echo signals in the residual signal.

[0174] This application also provides another application scenario where the echo control method described above is applied. Specifically, the echo control method can be applied to video call scenarios, live video streaming scenarios, and video conferencing scenarios. For example, in a video call scenario, the terminal performs echo prediction processing on the current peer's voice signal to obtain a simulated echo signal, predicts the estimated signal-to-return ratio (SPR) of the simulated echo signal, and then adjusts the power spectrum of the current peer's voice signal based on the estimated SPR to obtain the current voice signal to be played in the video call scenario, thereby avoiding excessive echo signals during the video call. Furthermore, the terminal performs echo cancellation processing on the smaller real echo signal generated based on the current voice signal to be played, improving the echo cancellation effect and thus improving the call voice quality in the video call scenario.

[0175] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially, these steps are not necessarily executed in that order. Unless otherwise expressly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the above embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0176] In one embodiment, such as Figure 8 As shown, an echo control device 800 is provided. This device can be a software module, a hardware module, or a combination of both, integrated into a computer device. Specifically, the device includes:

[0177] The acquisition module 802 is used to acquire the current voice signal of the other end in a voice call.

[0178] The prediction module 804 is used to perform echo prediction processing on the current peer voice signal to obtain a simulated echo signal; predict the estimated signal-to-echo ratio of the simulated echo signal; the estimated signal-to-echo ratio is used to characterize the energy proportion of the simulated echo signal in the current acquired signal at the local end.

[0179] The adjustment module 806 is used to perform power spectrum adjustment processing on the current peer voice signal according to the estimated signal-to-return ratio to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal.

[0180] The echo cancellation module 808 is used to perform echo cancellation processing on the real echo signal generated based on the current audio signal to be played.

[0181] In one embodiment, the prediction module 804 is further configured to acquire the current acquisition signal at the local end; determine the power spectrum of the current acquisition signal; determine the power spectrum of the analog echo signal; and determine the estimated signal-to-return ratio of the analog echo signal based on the power spectrum of the current acquisition signal and the power spectrum of the analog echo signal.

[0182] In one embodiment, the prediction module 804 is further configured to perform frame segmentation processing on the current acquired signal to obtain multiple frames of acquired signals; for each frame of acquired signals in the multiple frames of acquired signals, determine the initial power value of the acquired signal; obtain the prior power value of the previously acquired signal; the previously acquired signal is the previous frame of the acquired signal in the multiple frames of acquired signals; and determine the power spectrum of the current acquired signal based on the initial power value of each frame of acquired signal and the prior power value of the corresponding previously acquired signal.

[0183] In one embodiment, the prediction module 804 is further configured to, for each frame of the multi-frame acquisition signal, weight the prior power value corresponding to the prior acquisition signal of the acquisition signal by a preset first weighting coefficient to obtain a first prior power value; select the maximum value from the first prior power value and the initial power value of the acquisition signal as the smoothed power value of the acquisition signal; perform smoothing processing on the smoothed power value of the acquisition signal and the prior power value corresponding to the prior acquisition signal of the acquisition signal according to a preset first acquisition signal smoothing coefficient and a preset second acquisition signal smoothing coefficient to obtain the current power value of the acquisition signal; and determine the power spectrum of the current acquisition signal according to the current power value of each frame of the acquisition signal.

[0184] In one embodiment, the prediction module 804 is further configured to perform frame-segmentation processing on the analog echo signal to obtain multiple frames of echo signals; determine the initial power value of the echo signal for each frame of echo signal in the multiple frames of echo signals; obtain the prior power value of the preceding echo signal; the preceding echo signal is the previous frame echo signal in the multiple frames of echo signals; and determine the power spectrum of the analog echo signal based on the initial power value of each frame of echo signal and the prior power value of the corresponding preceding echo signal.

[0185] In one embodiment, the prediction module 804 is further configured to, for each frame of echo signal in the multi-frame echo signal, weight the prior power value corresponding to the prior echo signal of the echo signal by a preset second weighting coefficient to obtain a second prior power value; select the maximum value from the second prior power value and the initial power value of the echo signal as the smoothed power value of the echo signal; perform smoothing processing on the smoothed power value of the echo signal and the prior power value corresponding to the prior echo signal according to a preset first echo signal smoothing coefficient and a preset second echo signal smoothing coefficient to obtain the current power value of the echo signal; and determine the power spectrum of the current echo signal according to the current power value of each frame of echo signal.

[0186] In one embodiment, the estimated signal-to-return ratio (SRR) of the simulated echo signal includes the estimated SRR of each frequency point in the power spectrum of the simulated echo signal. The prediction module 804 is further configured to determine a candidate power value for each frequency point in the power spectrum based on the difference between the power value of the acquired signal and the power value of the echo signal. The power value of the acquired signal is the power value corresponding to the frequency point in the power spectrum of the currently acquired signal. The power value of the echo signal is the power value corresponding to the frequency point in the power spectrum of the simulated echo signal. The maximum value is selected from the candidate power value and the preset minimum power value as the target power value. The estimated SRR of the frequency point is determined based on the ratio of the target power value to the power value of the echo signal.

[0187] In one embodiment, the adjustment module 806 is further configured to determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold; and to perform power spectrum adjustment processing on the current peer voice signal based on the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played.

[0188] In one embodiment, the adjustment module 806 is further configured to determine a first candidate power spectrum adjustment coefficient based on the ratio of the estimated signal-to-return ratio to a preset signal-to-return ratio threshold and a preset minimum power spectrum adjustment coefficient; select the minimum value from the first candidate power spectrum adjustment coefficient and the second candidate power spectrum adjustment coefficient as the power spectrum adjustment coefficient for the current peer voice signal; the second candidate power spectrum adjustment coefficient is a coefficient that keeps the power spectrum of the current peer voice signal unchanged.

[0189] In one embodiment, such as Figure 9 As shown, the echo control device 800 also includes:

[0190] The determination module 810 is used to determine the signal correlation between the currently acquired signal and the prior voice signal to be played; the prior voice signal to be played is the voice signal to be played that precedes the current voice signal to be played and is time-aligned with the current acquired signal; the echo state of the voice call is determined based on the signal correlation.

[0191] The adjustment module 806 is also used to determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state.

[0192] In one embodiment, the preset signal-to-return ratio (SRR) threshold includes a first preset SRR threshold set for a one-way talk state and a second preset SRR threshold set for a two-way talk state; the first preset SRR threshold is less than the second preset SRR threshold; the adjustment module 806 is further configured to, if the echo state of the voice call is a one-way talk state, determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the first preset SRR threshold; if the echo state of the voice call is a two-way talk state, determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated SRR and the second preset SRR threshold.

[0193] In one embodiment, the adjustment module 806 is further configured to perform frequency domain conversion on the current peer voice signal to obtain the initial power spectrum corresponding to the current peer voice signal; the initial power spectrum includes the initial amplitude of each frequency point corresponding to the current peer voice signal; for each frequency point, the initial amplitude of the frequency point is weighted and calculated with the power spectrum adjustment coefficient corresponding to the frequency point to obtain the target amplitude of the frequency point; based on the target amplitude of each frequency point, the current peer voice signal is subjected to power spectrum adjustment processing; and the adjusted power spectrum is converted in the time domain to obtain the adjusted current voice signal to be played.

[0194] In one embodiment, the cancellation module 808 is further configured to perform echo simulation on the current audio signal to be played to obtain a target simulated echo signal; perform linear echo cancellation processing on the real echo signal generated after the current audio signal to be played is played based on the target simulated echo signal to obtain a residual signal including a nonlinear echo signal; and perform nonlinear echo cancellation processing on the nonlinear echo signal in the residual signal.

[0195] The aforementioned echo control device acquires the current peer voice signal during a voice call, performs echo prediction processing on the current peer voice signal to obtain a simulated echo signal. By predicting the estimated signal-to-return ratio (SRR) of the simulated echo signal, which characterizes the energy proportion of the simulated echo signal in the currently acquired signal at the local end, and performing power spectrum adjustment processing on the current peer voice signal based on the estimated SRR, an adjusted current voice signal to be played can be obtained. The energy of the current voice signal to be played is lower than that of the current peer voice signal. Then, echo cancellation processing is performed on the actual echo signal generated based on the current voice signal to be played. Compared to traditional methods that directly cancel existing echo signals, this application predicts a simulated echo signal based on the current peer's speech signal and estimates the signal-to-return ratio (SRR), which characterizes the energy proportion of the simulated echo signal in the current acquired signal at the local end. Based on the estimated SRR of the simulated echo signal, the power spectrum of the current peer's speech signal is first adjusted to avoid the occurrence of large echo signals. Then, the smaller echo signals generated by the adjusted current speech signal to be played are echo-cancelled. The echo cancellation effect is better, thereby improving the voice quality of the call.

[0196] Each module in the aforementioned echo control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0197] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an echo control method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0198] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0199] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0200] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0201] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0202] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0203] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0204] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0205] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An echo control method, characterized in that, The method includes: Get the current voice signal from the other end in a voice call; Echo prediction processing is performed on the current peer voice signal to obtain a simulated echo signal; The estimated signal-to-return ratio of the simulated echo signal is predicted; the estimated signal-to-return ratio is used to characterize the energy proportion of the simulated echo signal in the currently acquired signal at this end. Based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold, determine the power spectrum adjustment coefficient for the current peer voice signal; The power spectrum adjustment process is performed on the current peer voice signal according to the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal. Echo cancellation processing is performed on the actual echo signal generated based on the currently playing audio signal.

2. The method according to claim 1, characterized in that, The prediction of the signal-to-return ratio of the simulated echo signal includes: Obtain the current acquisition signal from this terminal; Determine the power spectrum of the currently acquired signal; Determine the power spectrum of the simulated echo signal; The estimated signal-to-return ratio of the simulated echo signal is determined based on the power spectrum of the currently acquired signal and the power spectrum of the simulated echo signal.

3. The method according to claim 2, characterized in that, Determining the power spectrum of the currently acquired signal includes: The currently acquired signal is processed into frames to obtain multiple frames of acquired signal; For each frame of the multi-frame acquisition signal, determine the initial power value of the acquisition signal; Obtain the prior power value of the previously acquired signal; the previously acquired signal is the previous frame acquisition signal among the multiple frames of acquisition signals. The power spectrum of the currently acquired signal is determined based on the initial power value of the acquired signal in each frame and the prior power value of the corresponding previously acquired signal.

4. The method according to claim 3, characterized in that, Determining the power spectrum of the currently acquired signal based on the initial power value of the acquired signal in each frame and the prior power value of the corresponding previously acquired signal includes: For each frame of the multi-frame acquisition signal, the prior power value corresponding to the prior acquisition signal is weighted by a preset first weighting coefficient to obtain a first prior power value; The maximum value between the first prior power value and the initial power value of the acquired signal is selected as the smoothed power value of the acquired signal; Based on the preset first acquisition signal smoothing coefficient and the preset second acquisition signal smoothing coefficient, the smoothed power value of the acquisition signal and the prior power value corresponding to the previous acquisition signal of the acquisition signal are smoothed to obtain the current power value of the acquisition signal. The power spectrum of the current acquired signal is determined based on the current power value of the acquired signal in each frame.

5. The method according to claim 2, characterized in that, Determining the power spectrum of the analog echo signal includes: The simulated echo signal is processed into frames to obtain multiple frames of echo signal; For each frame of the multi-frame echo signal, determine the initial power value of the echo signal; Obtain the prior power value of the prior echo signal; the prior echo signal is the echo signal of the previous frame in the multi-frame echo signal; The power spectrum of the simulated echo signal is determined based on the initial power value of the echo signal in each frame and the prior power value of the corresponding prior echo signal.

6. The method according to claim 5, characterized in that, Determining the power spectrum of the analog echo signal based on the initial power value of the echo signal in each frame and the prior power value of the corresponding prior echo signal includes: For each frame of the multi-frame echo signal, the prior power value corresponding to the prior echo signal is weighted by a preset second weighting coefficient to obtain a second prior power value. The maximum value between the second prior power value and the initial power value of the echo signal is selected as the smoothed power value of the echo signal. Based on a preset first echo signal smoothing coefficient and a preset second echo signal smoothing coefficient, the smoothed power value of the echo signal and the prior power value corresponding to the prior echo signal of the echo signal are smoothed to obtain the current power value of the echo signal. The power spectrum of the simulated echo signal is determined based on the current power value of the echo signal in each frame.

7. The method according to claim 2, characterized in that, The estimated signal-to-return ratio (SPR) of the simulated echo signal includes the estimated SPR at each frequency point in the power spectrum of the simulated echo signal; determining the estimated SPR of the simulated echo signal based on the power spectrum of the currently acquired signal and the power spectrum of the simulated echo signal includes: For each frequency point in the power spectrum, a candidate power value is determined based on the difference between the power value of the acquired signal and the power value of the echo signal; the power value of the acquired signal is the power value in the power spectrum of the currently acquired signal corresponding to the frequency point; the power value of the echo signal is the power value in the power spectrum of the simulated echo signal corresponding to the frequency point. The maximum value is selected from the candidate power values ​​and the preset minimum power value as the target power value; The estimated signal-to-return ratio at the frequency point is determined based on the ratio of the target power value to the echo signal power value.

8. The method according to claim 1, characterized in that, The step of determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold includes: The first candidate power spectrum adjustment coefficient is determined based on the ratio of the estimated signal-to-response ratio to the preset signal-to-response ratio threshold and the preset minimum power spectrum adjustment coefficient. The minimum value is selected from the first candidate power spectrum adjustment coefficient and the second candidate power spectrum adjustment coefficient as the power spectrum adjustment coefficient for the current peer voice signal; the second candidate power spectrum adjustment coefficient is a coefficient that keeps the power spectrum of the current peer voice signal unchanged.

9. The method according to claim 1, characterized in that, The method further includes: Determine the signal correlation between the currently acquired signal and the prior audio signal to be played; the prior audio signal to be played is the audio signal to be played that precedes the current audio signal to be played and is time-aligned with the current acquired signal. The echo state of the voice call is determined based on the signal correlation. The step of determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold includes: Based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state, the power spectrum adjustment coefficient for the current peer voice signal is determined.

10. The method according to claim 1, characterized in that, The preset signal-to-response ratio threshold includes a first preset signal-to-response ratio threshold set for single-talk mode and a second preset signal-to-response ratio threshold set for dual-talk mode; The first preset signal-to-response ratio threshold is less than the second preset signal-to-response ratio threshold; The step of determining the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold includes: If the echo state of the voice call is a one-way state, then the power spectrum adjustment coefficient for the current peer voice signal is determined based on the estimated signal-to-return ratio and the first preset signal-to-return ratio threshold. If the echo state of the voice call is two-way talk, then the power spectrum adjustment coefficient for the current peer voice signal is determined based on the estimated signal-to-return ratio and the second preset signal-to-return ratio threshold.

11. The method according to claim 1, characterized in that, The step of performing power spectrum adjustment processing on the current peer-end voice signal according to the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played includes: The current peer voice signal is frequency domain transformed to obtain the initial power spectrum corresponding to the current peer voice signal; the initial power spectrum includes the initial amplitude of each frequency point corresponding to the current peer voice signal. For each frequency point, the initial amplitude of the frequency point is weighted and calculated with the power spectrum adjustment coefficient corresponding to the frequency point to obtain the target amplitude of the frequency point; Based on the target amplitude at each frequency point, the current peer voice signal is subjected to power spectrum adjustment processing. The adjusted power spectrum is then transformed in the time domain to obtain the adjusted current audio signal to be played.

12. The method according to any one of claims 1 to 11, characterized in that, The echo cancellation process for the real echo signal generated based on the current audio signal to be played includes: Echo simulation is performed on the current audio signal to be played to obtain the target simulated echo signal; Based on the target simulated echo signal, the real echo signal generated after the current audio signal to be played is subjected to linear echo cancellation processing to obtain a residual signal including nonlinear echo signals. The nonlinear echo signal in the residual signal is subjected to nonlinear echo cancellation processing.

13. An echo control device, characterized in that, The device includes: The acquisition module is used to acquire the current voice signal from the other end during a voice call; The prediction module is used to perform echo prediction processing on the current peer voice signal to obtain a simulated echo signal; predict the estimated signal-to-echo ratio of the simulated echo signal; the estimated signal-to-echo ratio is used to characterize the energy proportion of the simulated echo signal in the current acquired signal at the local end; The adjustment module is used to determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold; perform power spectrum adjustment processing on the current peer voice signal based on the power spectrum adjustment coefficient to obtain the adjusted current voice signal to be played; the energy of the current voice signal to be played is lower than the energy of the current peer voice signal; The echo cancellation module is used to perform echo cancellation processing on the real echo signal generated based on the current audio signal to be played.

14. The echo control device according to claim 13, characterized in that, The prediction module is also used to acquire the current acquisition signal at the local end; determine the power spectrum of the current acquisition signal; determine the power spectrum of the analog echo signal; and determine the estimated signal-to-return ratio of the analog echo signal based on the power spectrum of the current acquisition signal and the power spectrum of the analog echo signal.

15. The echo control device according to claim 14, characterized in that, The prediction module is further configured to perform frame segmentation processing on the current acquisition signal to obtain multiple frames of acquisition signals; for each frame of acquisition signal in the multiple frames of acquisition signals, determine the initial power value of the acquisition signal; obtain the prior power value of the previously acquired signal; the previously acquired signal is the previous frame of acquisition signal in the multiple frames of acquisition signals; and determine the power spectrum of the current acquisition signal based on the initial power value of each frame of acquisition signal and the prior power value of the corresponding previously acquired signal.

16. The echo control device according to claim 15, characterized in that, The prediction module is further configured to, for each frame of the multi-frame acquisition signal, weight the prior power value corresponding to the earlier acquisition signal of the acquisition signal using a preset first weighting coefficient to obtain a first prior power value; select the maximum value from the first prior power value and the initial power value of the acquisition signal as the smoothed power value of the acquisition signal; perform smoothing processing on the smoothed power value of the acquisition signal and the prior power value corresponding to the earlier acquisition signal of the acquisition signal according to a preset first acquisition signal smoothing coefficient and a preset second acquisition signal smoothing coefficient to obtain the current power value of the acquisition signal; and determine the power spectrum of the current acquisition signal based on the current power value of each frame of the acquisition signal.

17. The echo control device according to claim 14, characterized in that, The prediction module is further configured to perform frame-segmentation processing on the simulated echo signal to obtain multiple frames of echo signals; for each frame of echo signal in the multiple frames of echo signals, determine the initial power value of the echo signal; obtain the prior power value of the preceding echo signal; the preceding echo signal is the echo signal of the previous frame in the multiple frames of echo signals; and determine the power spectrum of the simulated echo signal based on the initial power value of each frame of echo signal and the prior power value of the corresponding preceding echo signal.

18. The echo control device according to claim 17, characterized in that, The prediction module is further configured to, for each frame of the multi-frame echo signal, weight the prior power value corresponding to the prior echo signal of the echo signal using a preset second weighting coefficient to obtain a second prior power value; select the maximum value between the second prior power value and the initial power value of the echo signal as the smoothed power value of the echo signal; perform smoothing processing on the smoothed power value of the echo signal and the prior power value corresponding to the prior echo signal according to a preset first echo signal smoothing coefficient and a preset second echo signal smoothing coefficient to obtain the current power value of the echo signal; and determine the power spectrum of the simulated echo signal based on the current power value of each frame of the echo signal.

19. The echo control device according to claim 14, characterized in that, The estimated signal-to-return ratio (SRR) of the simulated echo signal includes the estimated SRR for each frequency point in the power spectrum of the simulated echo signal. The prediction module is further configured to determine a candidate power value for each frequency point in the power spectrum based on the difference between the power value of the acquired signal and the power value of the echo signal. The power value of the acquired signal is the power value corresponding to the frequency point in the power spectrum of the currently acquired signal. The power value of the echo signal is the power value corresponding to the frequency point in the power spectrum of the simulated echo signal. The maximum value is selected from the candidate power value and the preset minimum power value as the target power value. The estimated SRR of the frequency point is determined based on the ratio of the target power value to the power value of the echo signal.

20. The echo control device according to claim 13, characterized in that, The adjustment module is further configured to determine a first candidate power spectrum adjustment coefficient based on the ratio of the estimated signal-to-response ratio to the preset signal-to-response ratio threshold and a preset minimum power spectrum adjustment coefficient; The minimum value is selected from the first candidate power spectrum adjustment coefficient and the second candidate power spectrum adjustment coefficient as the power spectrum adjustment coefficient for the current peer voice signal; the second candidate power spectrum adjustment coefficient is a coefficient that keeps the power spectrum of the current peer voice signal unchanged.

21. The echo control device according to claim 13, characterized in that, The device further includes a determining module, which is used to determine the signal correlation between the currently acquired signal and the prior voice signal to be played; the prior voice signal to be played is a voice signal to be played that precedes the current voice signal to be played and is time-aligned with the current acquired signal; and the echo state of the voice call is determined based on the signal correlation. The adjustment module is also used to determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the preset signal-to-return ratio threshold corresponding to the echo state.

22. The echo control device according to claim 13, characterized in that, The preset signal-to-response ratio threshold includes a first preset signal-to-response ratio threshold set for single-talk mode and a second preset signal-to-response ratio threshold set for two-talk mode; the first preset signal-to-response ratio threshold is less than the second preset signal-to-response ratio threshold; The adjustment module is further configured to, if the echo state of the voice call is a single-talk state, determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the first preset signal-to-return ratio threshold; if the echo state of the voice call is a two-talk state, determine the power spectrum adjustment coefficient for the current peer voice signal based on the estimated signal-to-return ratio and the second preset signal-to-return ratio threshold.

23. The echo control device according to claim 13, characterized in that, The adjustment module is also used to perform frequency domain conversion on the current peer voice signal to obtain the initial power spectrum corresponding to the current peer voice signal; the initial power spectrum includes the initial amplitude of each frequency point corresponding to the current peer voice signal. For each frequency point, the initial amplitude of the frequency point is weighted and calculated with the power spectrum adjustment coefficient corresponding to the frequency point to obtain the target amplitude of the frequency point; based on the target amplitude of each frequency point, the current peer voice signal is subjected to power spectrum adjustment processing. The adjusted power spectrum is then transformed in the time domain to obtain the adjusted current audio signal to be played.

24. The echo control device according to any one of claims 13 to 23, characterized in that, The elimination module is also used to perform echo simulation on the current audio signal to be played to obtain a target simulated echo signal; based on the target simulated echo signal, to perform linear echo cancellation processing on the real echo signal generated after the current audio signal to be played is played to obtain a residual signal including nonlinear echo signals. The nonlinear echo signal in the residual signal is subjected to nonlinear echo cancellation processing.

25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.

26. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.