Voice signal processing method, processing device, terminal and storage medium thereof

By acquiring the voice signal and its echo signal for echo processing, and suppressing and adjusting according to the correlation, the nonlinear distortion problem of speakers is solved and the quality of voice communication is improved.

CN111145771BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010139686.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-03
Publication Date
2025-06-27
Estimated Expiration
2040-03-03

AI Technical Summary

Technical Problem

The prior art cannot effectively deal with the nonlinear distortion problem of speakers, resulting in a degradation of voice communication quality.

Method used

By acquiring the first voice signal and its echo signal, echo processing is performed to obtain the residual signal, and then the first voice signal is suppressed and adjusted according to the correlation between the first voice signal and the residual signal to suppress the nonlinear distortion signal emitted by the speaker.

Benefits of technology

It effectively suppresses the nonlinear distortion signal emitted by the speaker and improves the quality of voice communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111145771B_ABST
    Figure CN111145771B_ABST
Patent Text Reader

Abstract

The present application discloses a voice signal processing method, a processing device, a terminal, and a storage medium thereof. The voice signal processing method includes: obtaining a first voice signal, and obtaining a second voice signal that is the echo generated when the first voice signal is played, then performing echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal, and then performing suppression adjustment on the first voice signal according to the correlation between the first voice signal and the residual signal. According to the solution provided by the embodiments of the present application, after obtaining the residual signal corresponding to the first voice signal by performing echo processing on the second voice signal, the severity of the non-linear distortion signal in the echo can be judged according to the correlation between the residual signal and the first voice signal, that is, the first voice signal can be subjected to suppression adjustment according to the correlation between the first voice signal and the residual signal, which can suppress the non-linear distortion signal emitted by the speaker and improve the quality of voice communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of voice signal processing, and in particular, to a voice signal processing method, a processing device, a terminal, and a computer-readable storage medium thereof. Background Art

[0002] With the continuous development of voice signal processing technology, users' requirements for voice quality are also getting higher and higher. If there is echo in the voice, it will seriously affect the voice quality. The principle of echo generation: The voice signal is played through the speaker and undergoes multiple reflections in a closed or semi-closed environment, resulting in signal distortion. Finally, it is collected by the microphone together with the local voice to form an echo.

[0003] In order to eliminate the influence of echo on voice quality, traditional echo cancellation methods mainly directly perform echo cancellation on the voice signal collected by the microphone through an echo cancellation system. However, using the above echo cancellation method, the problem of non-linear distortion of the speaker cannot be solved. The non-linear distortion of the speaker is manifested as many additional distortion signals that are not the original voice signals in the sound output by the speaker. The distortion signals are non-linear distortion signals caused by the non-linear distortion of the speaker. The existing echo cancellation system cannot effectively process the non-linear distortion signals, thus affecting the quality of voice communication. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.

[0005] The present application provides a voice signal processing method, a processing device, a terminal, and a computer-readable storage medium thereof, which can suppress the non-linear distortion emitted by the speaker and improve the quality of voice communication.

[0006] According to the first aspect of the present application, a voice signal processing method is provided, including:

[0007] Obtain a first voice signal;

[0008] Obtain a second voice signal, where the second voice signal is a voice acquisition signal of the echo generated when the first voice signal is played;

[0009] Perform echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal;

[0010] Suppress and adjust the first voice signal according to the correlation between the first voice signal and the residual signal.

[0011] According to the second aspect of the present application, a voice signal processing device is provided, including:

[0012] A voice input end, configured to obtain a first voice signal;

[0013] A loudspeaker, connected to the voice input end, configured to play the first voice signal;

[0014] A microphone, configured to obtain a second voice signal, where the second voice signal is a voice acquisition signal of an echo generated when the first voice signal is played;

[0015] An echo cancellation module, configured to perform echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal;

[0016] A gain adjustment module, disposed between the voice input end and the loudspeaker, configured to perform suppression adjustment on the first voice signal according to the correlation between the first voice signal and the residual signal.

[0017] According to a third aspect of the present application, there is provided a voice signal processing device, including:

[0018] A voice acquisition unit, configured to obtain a first voice signal;

[0019] A voice acquisition unit, configured to obtain a second voice signal, where the second voice signal is a voice acquisition signal of an echo generated when the first voice signal is played;

[0020] An echo processing unit, configured to perform echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal;

[0021] A suppression adjustment unit, configured to perform suppression adjustment on the first voice signal according to the correlation between the first voice signal and the residual signal.

[0022] According to a fourth aspect of the present application, there is provided a voice signal processing device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the voice signal processing method described in the first aspect above is implemented.

[0023] According to a fifth aspect of the present application, there is provided a terminal, including the voice signal processing device described in the second aspect above, or the voice signal processing device described in the third aspect above, or the voice signal processing device described in the fourth aspect above.

[0024] According to a sixth aspect of the present application, there is provided a computer-readable storage medium, storing computer-executable instructions, where the computer-executable instructions are used to execute the voice signal processing method described in the first aspect above.

[0025] After obtaining the residual signal corresponding to the first speech signal by performing echo processing on the second speech signal, the technical solution provided by this application can determine the severity of the non-linear distortion signal in the echo according to the correlation between the residual signal and the first speech signal. That is, the first speech signal can be suppressed and adjusted according to the correlation between the first speech signal and the residual signal, which can suppress the non-linear distortion signal emitted by the speaker and improve the quality of voice communication.

[0026] Other features and advantages of this application will be described in the following specification. And, partly, they will become obvious from the specification, or be understood by implementing this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Brief Description of the Drawings

[0027] The drawings are used to provide a further understanding of the technical solution of this application, and constitute a part of the specification. Together with the embodiments of this application, they are used to explain the technical solution of this application, and do not constitute a limitation to the technical solution of this application.

[0028] Figure 1 is the system schematic diagram of the prior art echo cancellation system;

[0029] Figure 2 is the system principle block diagram of an application scenario of the speech signal processing method and the speech signal processing device according to the embodiments of this application;

[0030] Figure 3 is the speech signal processing flow chart of the speech signal processing method according to the embodiments of this application;

[0031] Figure 4 is the method flow chart of the speech signal processing method according to the embodiments of this application;

[0032] Figure 5 is Figure 4 the specific method flow chart of step 404 in

[0033] Figure 6 is the system principle block diagram of the speech signal processing device according to the embodiments of this application;

[0034] Figure 7 is the system principle block diagram of the suppression and adjustment unit according to the embodiments of this application;

[0035] Figure 8 is the system principle block diagram of the echo processing unit according to the embodiments of this application;

[0036] Figure 9 is the system schematic diagram of the speech signal processing device according to the embodiments of this application;

[0037] Figure 10 is a system principle block diagram of a voice signal processing device according to an embodiment of the present application;

[0038] Figure 11 is a system principle block diagram of a terminal according to an embodiment of the present application. Detailed implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.

[0040] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0041] Among the many factors affecting the quality of call voice, the echo problem is a relatively prominent one. Refer to Figure 1 as shown Figure 1 is a system principle diagram of an existing echo cancellation system, and the echo cancellation system is applied to the scenario of voice communication or video communication of a local terminal. Among them, the user of the terminal is the near end, and the communication object of the terminal user is the far end. The terminal includes a speaker and a microphone. The speaker is used to play the first voice signal, and the echo is the sound emitted from the speaker and then collected by the microphone of the terminal. If the echo is not processed, the microphone will send the near-end voice signal of the near-end terminal user and the echo signal to the far end together. In this way, the user at the far end will receive the echo signal, for example, hear his own voice, and this voice is transmitted back from the other terminal within a short time, thus greatly affecting the quality of the voice call.

[0042] In the prior art, the echo problem is solved by an echo cancellation module (AEC, Acoustic Echo Cancellation), and the signal collected by the microphone is echo processed. The echo processing steps of the echo cancellation module include: the terminal synchronizes the collected voice signal obtained by the microphone with the first voice signal, adaptively filters the first voice signal and performs inversion processing, and then linearly superimposes the synchronized collected signal and the first voice signal after inversion processing, thereby eliminating a part of the linear echo signal in the collected voice signal to obtain a residual signal, and the nonlinear echo signal remaining in the residual signal is further suppressed by a nonlinear suppression (NLP, Non-Linear Processor) module. The echo processing effect of the above-mentioned echo cancellation module depends on the AEC algorithm of the terminal, such as the mobile phone hardware, but the AEC algorithm performance of some mobile phone hardware does not meet the ideal index, and some residual echoes that are not suppressed cleanly appear, or the hardware AEC unit of the mobile phone cannot be used in some application scenarios.

[0043] One of the reasons why the echo problem has been difficult to solve is the nonlinear distortion in the echo signal. The nonlinear distortion problem mainly comes from the speaker, which is manifested as many additional non-original signal components in the sound output by the speaker. These components are caused by the nonlinear distortion of the speaker. The nonlinear distortion of the speaker mainly includes harmonic distortion, modulation distortion, transient distortion and subharmonic distortion. Due to the existence of nonlinear distortion of the speaker, the echo signal collected by the microphone also has nonlinear components.

[0044] Compared with linear echo, the characteristics of nonlinear echo are unstable. However, it is difficult for echo cancellation algorithms to accurately predict the size of nonlinear echo and completely and effectively suppress it. Therefore, a common problem of NLP algorithms in existing solutions is that while suppressing the nonlinear components of the echo, the normal near-end voice signal may also be suppressed. When the nonlinearity of the echo is more serious, there are two main strategies for existing echo cancellation solutions: 1. Ensure the fidelity of the near-end normal voice as much as possible without damaging it, while allowing the residual echo to exist. The result of this strategy will make the user feel uncomfortable during the call; 2. Increase the NLP suppression strength to suppress the residual echo, but at the same time it will also damage the near-end normal voice, resulting in clipping and intermittent near-end voice. Obviously, the existing echo cancellation method cannot solve the problem of severe nonlinear echo.

[0045] To this end, the embodiment of the present application provides a speech signal processing method, which can suppress the nonlinear distortion signal in the echo emitted by the speaker and improve the quality of voice communication. The speech signal processing method can be applied to Figure 2The application environment shown, which includes the terminal 210, the server 220, and the terminal 230. The terminal 210, the server 220, and the terminal 230 are connected through a network. The terminal 210 and the terminal 230 can be desktop terminals or mobile terminals. The mobile terminal can specifically be at least one of a mobile phone, a tablet computer, and a laptop computer. When the terminal 210 is the local end (proximal end), the terminal 230 is the remote end. The server 220 can be an independent server providing voice call support for the terminal 210 and the terminal 230, or a cluster server composed of multiple servers. The voice signal processing method in this embodiment can also be applied to application scenarios with only terminals or composed of only terminals and servers, such as a human-robot chat program or an artificial intelligence program with the remote end set in a terminal or a server.

[0046] When the voice signal processing method of the embodiment of the present application is applied to Figure 2 the shown terminal 210, with the terminal 210 as the local proximal end and the terminal 230 as the remote end communicating with the terminal 210, referring to Figure 3 as shown, the terminal 210 includes a microphone 310, a speaker 320, a voice input end 330, and a voice output end 340. Among them, the voice input end 330 is used to obtain Figure 2 the first voice signal sent by the terminal 230 therein and play it through the speaker. The voice output end 340 is used to send the third voice signal, or the local voice signal, to the terminal 230.

[0047] Referring to Figure 4 as shown, the voice signal processing method of the embodiment of the present application specifically includes step 401, step 402, step 403, and step 404.

[0048] Step 401, obtain the first voice signal.

[0049] Among them, the voice signal processing method of the embodiment of the present application can be applied to the application scenario of voice communication, and can also be applied to the human-computer interaction scenario with speaker playback, such as intelligent devices with voice calls, such as intelligent robots, smart speakers, and smart watches. The first voice signal can include, but is not limited to, user voices (including call voices), music, other background sounds, synthetic sounds, and prompt sounds and other audio signals.

[0050] In one embodiment, if the voice signal processing method is applied to a scenario such as voice communication with the above-mentioned terminal 230, the first voice signal is a voice signal obtained by the terminal 230 collecting ambient voice, such as a voice signal formed by collecting sound through the microphone 310 of the terminal 230. The voice signal can be a frequency-domain or time-domain signal, and the frequency-domain signal can be obtained by performing a Fourier transform on the time-domain signal.

[0051] In another embodiment, if the voice signal processing method is applied to a human-computer interaction scenario, the first voice signal is a voice signal obtained by a machine or an artificial intelligence device through voice synthesis, including but not limited to converting voice text into a synthesized voice signal.

[0052] Step 402: Obtain a second voice signal, where the second voice signal is a voice acquisition signal of the echo generated when the first voice signal is played.

[0053] In one embodiment, the first voice signal is played through the speaker 320 so that the user of the terminal 210 can hear the voice signal of the terminal 230 or the machine. The echo signal generated by the speaker 320 includes a linear echo signal and a non-linear echo signal. The linear echo signal may be the linear echo generated by the voice signal played by the speaker 320 due to reflection in the environment, etc. The non-linear echo signal may be the non-linear echo signal generated during the play due to the non-linear characteristics of the speaker 320. The second voice signal includes both the linear echo signal and the non-linear echo signal emitted by the speaker 320, and also includes the proximal voice signal input by the user on the terminal 210 side (i.e., the proximal terminal user) through the microphone 310.

[0054] In one embodiment, the first voice signal is played through the speaker 320 after gain adjustment. The gain adjustment may be to adjust the playback gain of the first voice signal. The first voice signal after gain adjustment is the playback signal for driving the speaker 320 to play. When the gain adjustment module does not make adjustments, the playback signal is the same as the first voice signal. When the gain adjustment module makes adjustments, the second voice signal obtained in step 402 is the echo generated when the adjusted first voice signal is played.

[0055] Step 403: Perform echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal.

[0056] In one embodiment, the terminal performs adaptive filtering processing on the obtained first voice signal or the first voice signal after gain adjustment to obtain a linear echo signal, and performs an inverting process on the linear echo signal. In addition, the terminal aligns the second voice signal collected by the microphone 310 with the first voice signal or the first voice signal after gain adjustment, and linearly superimposes the aligned second voice signal and the linearly inverted linear echo signal, so as to eliminate at least a part of the echo in the second voice signal, and further obtain a residual signal. The residual signal includes a proximal voice signal and a non-linear echo signal, and the non-linear echo signal is the non-linear residue in the above echo signal.

[0057] In one embodiment, the adaptive filtering algorithm may adopt the Least Mean Square (LMS) algorithm, the Normalized Least Mean Square (NLMS) algorithm, the Average Precision (AP) algorithm, or the Recursive Least Square (RLS) algorithm. The above adaptive filtering algorithms are prior arts and will not be elaborated herein.

[0058] Step 404: Perform suppression adjustment on the first voice signal according to the correlation between the first voice signal and the residual signal.

[0059] In one embodiment, the residual signal includes the non-linear echo signal generated by the speaker 320. Since the residual signal is obtained by performing echo processing based on the second voice signal, and the second voice signal includes the echo generated when the first voice signal is played, there is a correlation between the non-linear echo signal in the residual signal and the first voice signal. Therefore, by judging the correlation between the residual signal and the first voice signal, the presence of the non-linear echo signal in the residual signal can be detected. The higher the correlation, the more non-linear echo components in the residual signal; the lower the correlation, the fewer non-linear echo components in the residual signal.

[0060] Since the non-linear distortion problem of the speaker 320 is more prominent when the signal amplitude is large, performing suppression adjustment on the first voice signal according to the correlation between the residual signal and the first voice signal can suppress the non-linear echo signal emitted by the speaker 320 by adjusting the gain of the first voice signal to an appropriate value. For example, when the correlation between the residual signal and the first voice signal is low, it indicates that there are fewer non-linear echo components in the residual signal. To improve the quality of voice communication, the first voice signal may not be adjusted or slightly adjusted. When the correlation between the residual signal and the first voice signal is high, to suppress the non-linear echo signal emitted by the speaker 320, a relatively large adjustment is made to the first voice signal to reduce the non-linear echo components in the residual signal.

[0061] In this embodiment, the gain adjustment of the first voice signal includes suppressing the gain of the time-domain signal of the first voice signal of the far-end voice signal, and also includes suppressing the gain of the frequency-domain signal of the first voice signal. For example, the correlation between the power value of the frequency-domain signal of the first voice signal and the power value of the corresponding frequency point of the residual signal in the frequency domain is analyzed, and the power gain of each frequency point corresponding to the first voice signal is individually suppressed according to the correlation of the power value of the corresponding frequency point. The adjusted first voice signal is a playback signal used to drive the speaker 320 to play, and the playback signal is adjusted accordingly according to the component of the non-linear echo in the residual signal obtained at the previous moment, that is, the part of the first voice signal that is likely to cause the speaker 320 to emit non-linear echo signals is suppressed and adjusted, so that the non-linear echo signals emitted by the current speaker 320 can be suppressed. Since the acquisition of the residual signal, the correlation analysis between the residual signal and the first voice signal, and the gain adjustment of the first voice signal are continuously online, an overall feedback closed-loop control is formed, which can suppress the non-linear distortion signals emitted by the speaker 320 in real time, so that the non-linear echo part of the second voice signal collected by the microphone 310 is smaller. Through echo cancellation processing and non-linear suppression processing, the non-linear echo part in the second voice signal can be effectively eliminated, improving the quality of voice communication.

[0062] An embodiment of the present application proposes a method for processing voice signals. By acquiring a first voice signal and acquiring a second voice signal of the echo generated when the first voice signal is played, then performing echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal, and then suppressing and adjusting the first voice signal according to the correlation between the first voice signal and the residual signal. According to the solution provided by the embodiment of the present application, after obtaining the residual signal corresponding to the first voice signal by performing echo processing on the second voice signal, the severity of the non-linear distortion signal in the echo can be judged according to the correlation between the residual signal and the first voice signal, that is, the first voice signal can be suppressed and adjusted according to the correlation between the first voice signal and the residual signal, which can suppress the non-linear distortion signals emitted by the speaker and improve the quality of voice communication.

[0063] In one embodiment, referring to Figure 5 As shown, step 404 further includes step 501 and step 502.

[0064] Step 501, calculate the correlation between the first voice signal and the residual signal.

[0065] In one embodiment, the terminal calculates the correlation between the acquired first voice signal and the residual signal, for example, the correlation calculation can be performed in the frequency domain or the time domain. Through the correlation calculation, a correlation value representing the correlation degree between the first voice signal and the residual signal can be obtained. According to the correlation value, the playback gain of the first voice signal is suppressed and adjusted, which can provide a basis for conditional judgment for the suppression and adjustment in the following steps.

[0066] Step 502: Obtain the corresponding gain adjustment parameter according to the correlation, and use the gain adjustment parameter to suppress and adjust the first voice signal.

[0067] In one embodiment, a correspondence between the correlation and the gain adjustment parameter is established in advance. In this way, when the correlation between the first voice signal and the residual signal is calculated, the specific gain adjustment parameter can be obtained. The gain adjustment parameter corresponds to the non-linear echo signal part in the residual signal. The gain adjustment parameter can suppress the part of the first voice signal that is likely to generate non-linear echo signals in the speaker. The gain adjustment parameter in this embodiment is a value in the range of 0 to 1. Therefore, the gain adjustment process in this embodiment can suppress the amplitude of the corresponding signal in the first voice signal that is likely to generate non-linear echo in the speaker. When suppressing and adjusting the current first voice signal, the suppression adjustment is achieved by multiplying the first voice signal by the gain adjustment parameter.

[0068] The gain adjustment parameter can adjust the overall gain of the first voice signal in the time domain or can adjust the gain corresponding to each frequency point of the first voice signal in the frequency domain. Since the non-linear distortion problem of the speaker may only occur for a certain frequency point or be more obvious for a certain frequency point, by suppressing the gain of some frequency points of the first voice signal, the remaining echo volume can be reduced and is not easily noticed, which can reduce the impact on the voice call due to suppressing and adjusting the first voice signal. Based on this, it is necessary to calculate the correlation between the residual signal and the first voice signal at each frequency point respectively. The gain adjustment parameter is the gain adjustment coefficient corresponding to each frequency point of the first voice signal in the frequency domain.

[0069] In one embodiment, the above correspondence between the correlation and the gain adjustment can be linear, for example, in the relationship of a linear function, or can be non-linear, for example, in the relationship of a piecewise function.

[0070] Adopting the relationship of a linear function can better adjust the gain of the first voice signal for the non-linear echo component of the acquired residual signal. However, considering that the non-linear distortion characteristics of the speakers of different terminals are different, it is relatively complex to design a linear function suitable for various types and models of terminals. Therefore, in one embodiment, by setting a piecewise function related to the correlation, the applicability is better.

[0071] As an example, step 502 specifically includes:

[0072] When the relevance is less than the first preset threshold, set the gain adjustment parameter to the first gain value, and use the first gain value to perform suppression adjustment on the first voice signal;

[0073] When the relevance is greater than or equal to the first preset threshold and less than or equal to the second preset threshold, set the gain adjustment parameter to the second gain value, and use the second gain value to perform suppression adjustment on the first voice signal;

[0074] When the relevance is greater than the second preset threshold, set the gain adjustment parameter to the third gain value, and use the third gain value to perform suppression adjustment on the first voice signal, where the second gain value is greater than the third gain value and less than the first gain value, and the first gain value is less than or equal to 1.

[0075] In one embodiment, the gain adjustment parameter includes three levels, where the first gain value > the second gain value > the third gain value. The larger the gain value, the smaller the degree of suppression adjustment on the first voice signal, and the smaller the gain value, the higher the degree of suppression adjustment on the first voice signal. The processing gain value is generally set to 1, that is, no processing is performed. The second gain value and the third gain value are formulated according to the actual situation. In order to avoid affecting the voice call quality, the third gain value generally needs to be greater than 0.5. In addition, the number of levels of the gain adjustment parameter can be increased or decreased according to needs.

[0076] In this embodiment, when the gain adjustment parameter is the gain adjustment coefficient corresponding to each frequency point of the first voice signal in the frequency domain, it is necessary to calculate the relevance between the first voice signal and the residual signal at each frequency point respectively, then determine the gain adjustment parameter corresponding to each frequency point, and further adjust the gain of each frequency point of the first voice signal, which can reduce the volume of the residual echo and make it less noticeable, and can reduce the impact on the voice call due to suppressing the first voice signal.

[0077] Refer to Figure 3 As shown, the first voice signal is x(n), the second voice signal is d(n), and w(n) is the weight of the adaptive filter. First, align the second voice signal d(n) with the first voice signal x(n), then perform adaptive filtering processing on the first voice signal x(n) to obtain the linear echo signal w(n)x(n). Then, after taking the inverse of the linear echo signal w(n)x(n) and linearly adding it to the second voice signal d(n), the residual signal is obtained as e(n) = d(n) - w(n)x(n), where n is the frame number.

[0078] To compare the correlation between the residual signal e(n) and the first speech signal x(n) at each frequency point, the residual signal e(n) is transformed into the frequency-domain signal E(k) through the fast Fourier transform, and the power value |E(k)| of the residual signal e(n) at the k-th frequency point in the frequency domain is calculated. Here, k is the frequency-point serial number, representing the N-th frequency point, where N is a positive integer;

[0079] Then, the power value |X(k)| of the frequency-domain signal X(k) obtained after the first speech signal x(n) undergoes the fast Fourier transform is acquired;

[0080] Next, cross-correlation calculation is performed to obtain the power correlation coefficient between the residual signal e(n) and the first speech signal x(n) at each frequency point as:

[0081] The power correlation coefficient corr(k) between the residual signal e(n) and the first speech signal x(n) at each frequency point reflects the degree of non-linear echo signal at each frequency point in the residual signal e(n). When the value of the power correlation coefficient corr(k) is relatively high, it indicates a relatively high degree of non-linear echo at the current frequency point. When the value of the power correlation coefficient corr(k) is relatively low, it indicates a relatively low degree of non-linear echo at the current frequency point. Therefore, the gain adjustment coefficient g(k) corresponding to each frequency point in the frequency domain of the first speech signal can be set according to the power correlation coefficient corr(k).

[0082] In one embodiment, the gain adjustment coefficient g(k) is divided into three levels, including Gmax, Gnor, and Gmin, where Gmax is the maximum gain (usually 1), Gnor is the second-level gain (a value less than 1), and Gmin is the third-level gain, whose value is less than Gnor.

[0083] Configuring the gain coefficient g(k) is to multiply the frequency-domain signal X(k) obtained after the first speech signal x(n) undergoes the fast Fourier transform by the gain coefficient corresponding to each frequency point: |X(k)| * g(k), and keep the phase unchanged, and then obtain the played signal x′(n) after frequency-domain gain adjustment through the inverse fast Fourier transform, and finally send the played signal x′(n) to the speaker 320 for playback.

[0084] In one embodiment, the gain coefficient g(k) is a piecewise function, and the formula of this piecewise function is:

[0085]

[0086]

[0087] In the above formula, m is the corrected power correlation coefficient. When m < a1, it indicates that the non-linear echo degree of the current frequency point is relatively low, and g(k) takes the maximum value Gmax. When a1 ≤ m ≤ a2, it indicates that the non-linear echo degree of the current frequency point is average, and g(k) takes the intermediate value Gnor. When m > a2, it indicates that the non-linear echo degree of the current frequency point is relatively high, and g(k) takes the minimum value Gmin. The value of m is related to the power values of each frequency point of the residual signal e(n) and the first speech signal x(n). When |E(k)| > |X(k)|, m normally takes the value of corr(k). When |E(k)| ≤ |X(k)|, it indicates that the non-linear echo degree in the residual signal is relatively low. In order to avoid affecting normal voice communication, multiply corr(k) by Since Therefore, the power correlation coefficient m can be correspondingly reduced.

[0088] Of course, the above piecewise function is only one implementation manner of the gain adjustment coefficient. Those skilled in the art can also calculate the gain adjustment coefficient corresponding to the power correlation coefficient through a linear function according to needs.

[0089] The voice signal processing method in any of the above embodiments can be carried by the system program of the terminal, or can be carried by an application program that needs to use a microphone and a speaker in the terminal. The embodiments of the present application do not make specific limitations. When the voice signal processing method in any of the above embodiments is carried by the system program of the terminal, the user only needs to upgrade and update the system program of the terminal to use the voice signal processing method in any of the above embodiments, without the need to perform factory debugging on the terminal, which is convenient for the user to update and use. When the voice signal processing method in any of the above embodiments is carried by an application program that needs to use a microphone and a speaker in the terminal, only by installing the application program can the voice signal processing method in any of the above embodiments be used. Similarly, there is no need to perform factory debugging on the terminal, which is convenient for the user to update and use. The application program can be an independent voice signal processing software or plug-in, or can be an algorithm program built into the voice call software, such as application programs for WeChat audio and video calls, live broadcasts, broadcasts, etc. The embodiments of the present application do not make specific limitations.

[0090] As an example, it is used to illustrate the voice signal processing method provided in the above embodiments of the present application, which is applied to a scenario such as Figure 2 shown, where both the terminal 210 and the terminal 230 are installed with voice communication software capable of voice communication. The voice communication software can be the audio and video call function in chat and social software, or the live broadcast and broadcast functions.

[0091] First, a voice communication connection is established between the terminal 210 and the terminal 230 through the server. The microphone 310 of the terminal 230 collects the first voice signal from the remote end and sends the collected first voice signal to the local terminal 210 through the server. Refer to Figure 3As shown, at the moment when the voice communication starts, there is no echo signal yet. Therefore, there is no non-linear echo signal in the residual signal. The terminal 210 does not perform suppression adjustment on the first voice signal to obtain the playback signal. The voice signal output by the speaker 320 according to the playback signal generates an echo signal due to reflection in the environment and other reasons, and this echo signal is collected by the microphone 310 of the terminal 210. Since after the voice communication starts, the microphone 310 also collects the proximal voice signal of the local user at the same time, for example, the voice signal when the user speaks to the microphone 310. Therefore, the microphone 310 collects the second voice signal of the echo generated when the first voice signal is played. The second voice signal includes the linear echo signal and the non-linear echo signal emitted by the speaker 320. In order to remove the linear echo signal, on the one hand, the terminal 210 obtains the adjusted playback signal and generates a linear echo signal as the reference for echo cancellation through adaptive filtering. On the other hand, the terminal 210 performs time delay synchronization processing on the second voice signal according to the playback signal, aiming at the phase synchronization of the subsequent linear superposition. By linearly superposing the second voice signal after time delay processing and the linearly filtered and inverted linear echo signal, the linear echo information in the second voice signal is eliminated, and a residual signal is generated. After the terminal 210 performs non-linear suppression on the residual signal, it outputs the voice signal after echo processing to the terminal 230 through the server. At this time, the non-linear suppression processing cannot completely or preferably eliminate the non-linear echo signal in the residual signal. Therefore, at the voice input end 330, the first voice signal at the next moment is suppressed and adjusted according to the correlation between the residual signal and the first voice signal to suppress the non-linear echo signal emitted by the speaker. Specifically, by analyzing the correlation between the residual signal and the first voice signal at each frequency point, it can be known at which frequency points the non-linear echo signal with a higher degree of non-linear distortion appears, so as to suppress and adjust the gain adjustment coefficient of the first voice signal at each frequency point at the next moment. When the suppressed and adjusted playback signal drives the speaker 320, it can suppress the speaker 320 from emitting non-linear echo signals at the corresponding frequency points, thereby reducing the degree of non-linear echo signals in the second voice signal and effectively eliminating the echo signal of the second voice signal when performing echo processing on the second voice signal. It should be noted that the suppression adjustment of the first voice signal, the speaker 320 playing the first voice signal to generating a residual signal by performing echo processing on the second voice signal, and until the first voice signal is suppressed and adjusted according to the correlation between the residual signal and the first voice signal constitute a feedback closed-loop control. The next moment in the above embodiment is determined according to the frame sequence of the transmitted voice signal. Substantially, the entire processing process takes a short time and can be regarded as real-time processing.

[0092] As another example, the method for processing a voice signal provided in the above embodiments of the present application is applied to a scenario with only a terminal and a server. A voice response program with voice recognition and text-to-speech functions is set in the server, such as a voice assistant or an artificial intelligence response program. The terminal can be an intelligent device such as a smart phone or a smart speaker, and the terminal interacts with the server to implement an intelligent voice conversation.

[0093] The method for processing a voice signal according to the embodiments of the present application is applied to a terminal. The terminal establishes a voice communication connection with the server. Referring to Figure 3 as shown, the first voice signal can be that the voice response program in the server converts text information into a voice signal. Referring to Figure 3 as shown, at the moment when the voice communication starts, there is no echo signal yet, so there is no non-linear echo signal in the residual signal. The terminal does not perform suppression adjustment on the first voice signal to obtain a playback signal. The voice signal output by the speaker 320 according to the playback signal generates an echo signal due to reflection in the environment, and this echo signal is collected by the microphone 310 of the terminal. Since after the voice communication starts, the microphone 310 also collects the proximal voice signal of the local user, such as the voice signal when the user speaks to the microphone 310, the microphone 310 collects the second voice signal of the echo generated when the first voice signal is played. The second voice signal includes a linear echo signal and a non-linear echo signal emitted by the speaker 320. To remove the linear echo signal, on the one hand, the terminal obtains an adjusted playback signal and generates a linear echo signal as an echo cancellation reference through adaptive filtering. On the other hand, the terminal performs time delay synchronization processing on the second voice signal according to the playback signal, aiming at the phase synchronization of subsequent linear superposition. By linearly superposing the second voice signal after time delay processing and the linearly filtered and inverted linear echo signal, the linear echo information in the second voice signal is eliminated, and a residual signal is generated. After the terminal performs non-linear suppression on the residual signal, it outputs the voice signal after echo processing to the server. At this time, the non-linear suppression processing cannot completely or preferably eliminate the non-linear echo signal in the residual signal, which will affect the voice recognition effect of the voice response program in the server.

[0094] In this regard, at the voice input end 330, the first voice signal at the next moment is suppressed and adjusted according to the correlation between the residual signal and the first voice signal, so as to suppress the non-linear echo signal emitted by the speaker. Specifically, by analyzing the correlation between the residual signal and the first voice signal at each frequency point, it is known at which frequency points the non-linear echo signal with a higher degree of non-linear distortion appears, so as to suppress and adjust the gain adjustment coefficient of the first voice signal at each frequency point at the next moment. When the suppressed and adjusted playback signal drives the speaker 320, the speaker 320 can be suppressed from emitting non-linear echo signals at the corresponding frequency points, thereby reducing the degree of non-linear echo signals in the second voice signal. When performing echo processing on the second voice signal, the echo signal of the second voice signal can be effectively eliminated, and the voice recognition effect of the voice response program in the server can be improved.

[0095] As Figure 6 shown, another embodiment of the present application further provides a voice signal processing device, which includes:

[0096] A voice acquisition unit 510, configured to acquire a first voice signal;

[0097] A voice collection unit 520, configured to acquire a second voice signal, where the second voice signal is a voice collection signal of the echo generated when the first voice signal is played;

[0098] An echo processing unit 530, configured to perform echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal;

[0099] A suppression adjustment unit 540, configured to suppress and adjust the first voice signal according to the correlation between the first voice signal and the residual signal.

[0100] In one embodiment, the first voice signal acquired by the voice acquisition unit 510 includes, but is not limited to, other user voices (including call voices), music, other background sounds, synthetic sounds, and prompt sounds and other audio signals.

[0101] In one embodiment, when an echo signal is generated during the playback of the first voice signal, the echo signal may be received by the voice signal processing device, thereby affecting the normal voice signal received by the voice signal processing device. To avoid the influence of the echo signal generated during the playback of the first voice signal, the voice acquisition unit 520 acquires the echo signal to form a second voice signal. Then, the voice acquisition unit 520 sends the second voice signal to the echo processing unit 530, so that the echo processing unit 530 performs echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal. Then, the echo processing unit 530 sends the residual signal to the suppression adjustment unit 540. At this time, the suppression adjustment unit 540 can judge the severity of the non-linear distortion signal in the echo signal according to the correlation between the first voice signal and the residual signal, and can perform suppression adjustment on the first voice signal according to the severity, so as to suppress the non-linear distortion signal in the echo signal and improve the quality of voice communication.

[0102] As Figure 7 shown, in one embodiment, the suppression adjustment unit 540 includes:

[0103] A correlation calculation unit 610 for calculating the correlation between the first voice signal and the residual signal;

[0104] A gain adjustment unit 620 for obtaining a corresponding gain adjustment parameter according to the correlation and using the gain adjustment parameter to perform suppression adjustment on the first voice signal.

[0105] In one embodiment, the correlation calculation unit 610 calculates the correlation between the acquired first voice signal and the residual signal. For example, it can be a correlation calculation in the frequency domain or the time domain. Through the correlation calculation, a correlation value representing the correlation degree between the first voice signal and the residual signal can be obtained. According to the correlation value, the playback gain of the first voice signal is suppressed and adjusted, which can provide a condition judgment basis for the subsequent suppression adjustment steps.

[0106] In one embodiment, a corresponding relationship between the correlation and the gain adjustment parameter can be established in advance. In this way, when the correlation between the first voice signal and the residual signal is calculated, the specific gain adjustment parameter can be obtained. The gain adjustment parameter can suppress the part of the first voice signal that is likely to generate a non-linear echo signal in the speaker. The gain adjustment parameter in this embodiment is a value in the range of 0 to 1. Therefore, the gain adjustment can suppress the amplitude of the part of the first voice signal that is likely to generate a non-linear echo in the speaker. When performing suppression adjustment on the current first voice signal, the first voice signal is multiplied by the gain adjustment parameter to achieve the above-mentioned suppression adjustment.

[0107] The gain adjustment parameter can be used to adjust the overall gain of the first voice signal in the time domain or to adjust the gains corresponding to each frequency point of the first voice signal in the frequency domain. Since the non-linear distortion problem of the speaker may only occur for a certain frequency point or be more obvious for a certain frequency point, by suppressing the gains of some frequency points of the first voice signal, the volume of the residual echo can be reduced and is not easily noticeable, which can reduce the impact on the voice call due to suppressing and adjusting the first voice signal. Based on this, it is necessary to calculate the correlation between the residual signal and the first voice signal at each frequency point respectively, and the gain adjustment parameter is the gain adjustment coefficient corresponding to each frequency point of the first voice signal in the frequency domain.

[0108] In one embodiment, the above-mentioned correlation and gain adjustment correspondence can be linear, such as a linear function relationship, or non-linear, such as a piecewise function relationship.

[0109] Adopting a linear function relationship can better adjust the gain of the first voice signal for the non-linear echo components of the obtained residual signal. However, considering that the non-linear distortion characteristics of the speakers of different terminals are different, it is relatively complex to design a linear function suitable for various types and models of terminals. Therefore, in the embodiment, by setting a piecewise function related to the correlation, the applicability is better.

[0110] In one embodiment, the piecewise threshold can be set according to the specific value of the correlation, so that different degrees of suppression adjustment can be performed on the first voice signal according to different values of the correlation. For example:

[0111] When the correlation is less than the first preset threshold, set the gain adjustment parameter to the first gain value, and use the first gain value to perform suppression adjustment on the first voice signal;

[0112] When the correlation is greater than or equal to the first preset threshold and less than or equal to the second preset threshold, set the gain adjustment parameter to the second gain value, and use the second gain value to perform suppression adjustment on the first voice signal;

[0113] When the correlation is greater than the second preset threshold, set the gain adjustment parameter to the third gain value, and use the third gain value to perform suppression adjustment on the first voice signal, where the second gain value is greater than the third gain value and less than the first gain value, and the first gain value is less than or equal to 1.

[0114] In this embodiment, the gain adjustment parameter includes three levels, where the first gain value > the second gain value > the third gain value. The larger the gain value, the smaller the degree of suppression adjustment for the first voice signal, and the smaller the gain value, the higher the degree of suppression adjustment for the first voice signal. Generally, the first gain value is set to 1, that is, no processing is performed. The second gain value and the third gain value are formulated according to the actual situation. To avoid affecting the voice call quality, the third gain value generally needs to be greater than 0.5. In addition, the number of levels of the gain adjustment parameter can also be increased or decreased as needed.

[0115] In one embodiment, when the gain adjustment parameter is the gain adjustment coefficient corresponding to each frequency point of the first voice signal in the frequency domain, it is necessary to calculate the correlation degree between the first voice signal and the residual signal at each frequency point respectively, then determine the gain adjustment parameter corresponding to each frequency point, and further adjust the gain of each frequency point of the first voice signal, which can reduce the residual echo volume and make it less noticeable, and can reduce the impact on the voice call due to suppressing the first voice signal.

[0116] As Figure 8 shown, in one embodiment, the echo processing unit 530 includes:

[0117] A delay adjustment unit 710, configured to align the second voice signal with the playback signal, where the playback signal is obtained according to the first voice signal and is used to drive the speaker;

[0118] An adaptive filtering unit 720, configured to perform adaptive filtering processing on the playback signal to obtain a linear echo signal;

[0119] A linear superposition unit 730, configured to perform an inverting process on the linear echo signal and linearly superpose it with the aligned second voice signal to obtain a residual signal corresponding to the first voice signal.

[0120] In one embodiment, the adaptive filtering unit 720 performs adaptive filtering processing on the acquired first voice signal to obtain a linear echo signal. Then, the adaptive filtering unit 70 transmits the linear echo signal to the linear superposition unit 730 for inverting processing. In addition, the delay adjustment unit 710 aligns the collected second voice signal with the playback signal and transmits the aligned second voice signal to the linear superposition unit 730. At this time, the linear superposition unit 730 linearly superposes the aligned second voice signal with the linearly echo signal after inverting processing, so as to eliminate at least a part of the echo in the second voice signal, and further obtain a residual signal, thereby providing a condition judgment basis for the suppression adjustment processing of the suppression adjustment unit 540.

[0121] Referring to Figure 9As shown in the figure, an embodiment of the present application provides a voice signal processing device, which can suppress the non-linear distortion signal in the echo emitted by the speaker and improve the quality of voice communication. The voice signal processing device is applied to an application environment as shown in Figure 2 As shown, the application environment includes a terminal 210, a server 220, and a terminal 230. The terminal 210, the server 220, and the terminal 230 are connected through a network. The terminal 210 and the terminal 230 can be desktop terminals or mobile terminals. The mobile terminal can specifically be at least one of a mobile phone, a tablet computer, and a laptop computer. When the terminal 210 is the local end (proximal end), the terminal 230 is the remote end. The server 220 can be an independent server providing voice call support for the terminal 210 and the terminal 230 or a cluster server composed of multiple servers.

[0122] In addition, as another implementation manner of the present application, the voice signal processing device can also be applied to an application environment consisting of only terminals or only terminals and servers. For example, a human-robot chat program or an artificial intelligence program with the remote end set in a terminal or a server.

[0123] Referring to Figure 9 As shown in the figure, the voice signal processing device can be applied to an application environment as shown in Figure 2 As shown, the application environment includes a terminal 210, a server 220, and a terminal 230. The voice signal processing device is provided on the terminal 210. The voice input end 330 is used to receive the first voice signal from the remote end, and the voice output end 340 is used to send the voice signal on the proximal user side to the remote end. In addition, the voice signal processing device in this embodiment can also be applied to an application scenario consisting of only terminals or only terminals and servers. The following embodiments only use the Figure 2 application environment shown in the figure as an example for illustration, and is not limited to being applied in the Figure 2 application environment shown in the figure.

[0124] Referring to Figure 9 As shown in the figure, the voice signal processing device includes a voice input end 330, a voice output end 340, a speaker 320, and a microphone 310. The voice input end 330 and the voice output end 340 can be unified into a voice signal transmission port with uplink and downlink transceiver functions, including but not limited to a wired communication module and a wireless communication module. The wireless communication module can be a Wi-Fi communication module or a mobile network communication module. The mobile network communication module can be a mobile network communication module supporting 3G, 4G, 5G, or a future higher standard. The first voice signal can include but not limited to user voice (including call voice), music, other background sounds, synthesized sound, and prompt sound, etc.

[0125] The voice signal processing device further includes a gain adjustment module 810. The gain adjustment module 810 is disposed between the voice input end 330 and the speaker 320 and is used to adjust the gain of the first voice signal obtained by the voice input end 330. After being adjusted by the gain adjustment module 810, the first voice signal becomes a playback signal for driving the speaker 320. Among them, when the gain adjustment module 810 does not adjust the first voice signal, the playback signal is the same as the first voice signal. The voice output by the speaker 320 generates an echo signal due to reflection and other reasons in the environment, and the echo signal is acquired by the microphone 310. Since there are many additional non-original signal components in the sound output by the speaker 320, these signal components are non-linear distortion signals. Therefore, the echo signal output by the speaker 320 includes a linear echo signal and a non-linear echo signal.

[0126] The microphone 310 acquires the second voice signal of the echo generated when the first voice signal or the first voice signal after gain adjustment is played. At this time, the second voice signal includes a linear echo signal, a non-linear echo signal, and the voice signal on the proximal user side. Among the second voice signals, both the linear echo signal and the non-linear echo signal are signals that are not desired to appear. Therefore, an echo cancellation module 820 is also disposed between the voice output end 340 and the microphone 310. The echo cancellation module 820 can better cancel the linear echo signal, but it is difficult to completely cancel the non-linear echo signal.

[0127] Since in the echo cancellation module 820, a residual signal is generated when performing echo processing on the second voice signal, and the residual signal includes the voice signal on the proximal user side and the non-linear echo signal. Therefore, by connecting the gain adjustment module 810 to the echo cancellation module 820, the gain adjustment module 810 can obtain the residual signal output by the echo cancellation module 820. Then, the gain adjustment module 810 performs suppression adjustment on the acquired first voice signal according to the correlation between the residual signal and the first voice signal. The adjusted first voice signal is the playback signal for driving the speaker 320 to play, and the playback signal is adjusted accordingly according to the non-linear echo component in the residual signal acquired at the previous moment, that is, the part of the first voice signal that is likely to cause the speaker 320 to emit a non-linear echo signal is suppressed and adjusted. In this way, the non-linear echo signal emitted by the current speaker 320 can be suppressed. Since the acquisition of the residual signal, the correlation analysis between the residual signal and the first voice signal, and the gain adjustment of the first voice signal are continuously online, an overall feedback closed-loop control is formed, which can suppress the non-linear distortion signal emitted by the speaker 320 in real time, so that the non-linear echo part of the second voice signal collected by the microphone 310 is smaller. Through echo cancellation processing and non-linear suppression processing, the non-linear echo part in the second voice signal can be effectively eliminated, improving the quality of voice communication.

[0128] In one embodiment, the echo cancellation module 820 includes a delay adjustment module 830, an adaptive filtering module 840, a linear superposition module 850, and a non-linear suppression module 860. The microphone 310, the delay adjustment module 830, the linear superposition module 850, the non-linear suppression module 860, and the voice output end 340 are connected in sequence. The output end of the gain adjustment module 810 is connected to the linear superposition module 850 through the adaptive filtering module 840. The adaptive filtering module 840 performs adaptive filtering on the playback signal output by the gain adjustment module 810 and then performs an inverting process to form a linear echo signal as the echo cancellation reference signal. The delay adjustment module 830 is used to align and adjust the delay of the second voice signal collected by the microphone 310 with the acquired first voice signal. The linear superposition module 850 linearly superimposes the delay-adjusted second voice signal and the inverted second voice signal after adaptive filtering, so that the linear echo signal in the second voice signal can be superimposed and cancelled, and then a residual signal is generated. The residual signal passes through the non-linear suppression module 860 to suppress the non-linear residual part signal and then outputs the voice signal to the remote end through the voice output end 340. In addition, the residual signal is also transmitted to the gain adjustment module 810. The gain adjustment module 810 performs suppression adjustment on the first voice signal according to the correlation between the residual signal and the first voice signal. Specifically, it suppresses and adjusts the playback gain of the playback signal. When the adjusted playback signal drives the speaker 320 to play and output, it can suppress the non-linear echo signal emitted by the speaker 320, so that the degree of the non-linear echo signal of the second voice signal collected by the microphone 310 is smaller, improving the echo cancellation effect of the echo cancellation module 820 and improving the quality of voice communication.

[0129] In one embodiment, the delay adjustment module 830 is a delay device, which can be an analog delay or a digital delay. The adaptive filtering module 840 is an adaptive filter, the linear superposition module 850 is an adder, and the non-linear suppression module 860 is a signal suppressor.

[0130] The gain adjustment module 810 can suppress and adjust the gain of the time-domain signal of the first voice signal, or suppress and adjust the gain of the frequency-domain signal of the first voice signal. The gain adjustment module 810 pre-establishes the correspondence between the correlation degree and the gain adjustment parameter. In this way, after calculating the correlation degree between the first voice signal and the residual signal, the specific gain adjustment parameter can be obtained. The gain adjustment parameter corresponds to the non-linear echo signal part in the residual signal, and the gain adjustment parameter can suppress the part of the first voice signal that is likely to generate non-linear echo signals in the speaker 320. The gain adjustment parameter in this embodiment is a value within the range of 0 to 1. Therefore, the gain adjustment can suppress the amplitude of the part of the first voice signal that is likely to generate non-linear echo in the speaker 320. When performing suppression adjustment on the current first voice signal, the first voice signal is multiplied by the gain adjustment parameter to achieve the suppression adjustment.

[0131] Among them, the correlation between the residual signal and the first voice signal can be processed by the gain adjustment module 810, or an additional processor can be set to process it, and the processing result of the correlation degree is fed back to the gain adjustment module 810 to perform corresponding adjustment on the first voice signal.

[0132] The gain adjustment parameter can be used to adjust the overall gain of the first voice signal in the time domain, or to adjust the gain corresponding to each frequency point of the first voice signal in the frequency domain. Since the non-linear distortion problem of the speaker 320 may only occur for a certain frequency point or be more obvious for a certain frequency point, by suppressing the gain of some frequency points of the first voice signal, the remaining echo volume can be reduced and is not easily noticeable, which can reduce the impact on the voice call due to suppressing the first voice signal. Based on this, it is necessary to calculate the correlation degree between the residual signal and the first voice signal at each frequency point respectively, and the gain adjustment parameter is the gain adjustment coefficient corresponding to each frequency point of the first voice signal in the frequency domain.

[0133] In one embodiment, the above correspondence between the correlation degree and the gain adjustment can be linear, such as a linear function relationship, or non-linear, such as a piecewise function relationship.

[0134] Adopting a linear function relationship can better adjust the gain of the first voice signal for the non-linear echo components of the obtained residual signal. However, considering that the non-linear distortion characteristics of the speaker 320 for different terminals are different, it is relatively complex to design a linear function suitable for various types and models of terminals. Therefore, in one embodiment, the gain adjustment module 810 can better adapt to different terminal types by setting a piecewise function corresponding to the correlation degree.

[0135] In one embodiment, when the correlation degree is less than the first preset threshold, the gain adjustment parameter is set to the first gain value, and the first voice signal is suppressed and adjusted by using the first gain value;

[0136] When the correlation degree is greater than or equal to the first preset threshold and less than or equal to the second preset threshold, the gain adjustment parameter is set to the second gain value, and the first voice signal is suppressed and adjusted by using the second gain value;

[0137] When the correlation degree is greater than the second preset threshold, the gain adjustment parameter is set to the third gain value, and the first voice signal is suppressed and adjusted by using the third gain value, where the second gain value is greater than the third gain value and less than the first gain value, and the first gain value is less than or equal to 1.

[0138] In this embodiment, the gain adjustment parameter includes three levels, where the first gain value > the second gain value > the third gain value. The larger the gain value, the smaller the degree of suppression adjustment for the first voice signal, and the smaller the gain value, the higher the degree of suppression adjustment for the first voice signal. Generally, the first gain value is set to 1, that is, no processing is performed. The second gain value and the third gain value are formulated according to the actual situation. To avoid affecting the voice call quality, the third gain value generally needs to be greater than 0.5. In addition, the number of levels of the gain adjustment parameter can be increased or decreased according to needs.

[0139] In this embodiment, when the gain adjustment parameter is the gain adjustment coefficient corresponding to each frequency point of the first voice signal in the frequency domain, it is necessary to calculate the correlation degree between the first voice signal and the residual signal at each frequency point respectively, then determine the gain adjustment parameter corresponding to each frequency point, and further adjust the gain of each frequency point of the first voice signal, which can reduce the residual echo volume and make it less noticeable, and can reduce the impact on the voice call due to suppressing and adjusting the first voice signal.

[0140] Refer to Figure 10 As shown, the embodiment of the present application provides a voice signal processing device. The voice signal processing device can be any type of intelligent terminal, such as a smart phone, a tablet computer, a laptop computer. Specifically, the voice signal processing device includes: a memory 910, a processor 920, and a computer program stored on the memory and executable on the processor. The processor 920 and the memory 910 can be connected through a bus or other means. Figure 10 Taking the connection through the bus as an example.

[0141] The non-transitory software program and instructions required to implement the voice signal processing method in the above embodiment are stored in the memory 910. When executed by the processor 920, the voice signal processing method in the above embodiment is executed. For example, the method steps 401 to 404 described above are executed. Figure 4 in Figure 5The method steps 501 to 502 in

[0142] Referring to Figure 11 As shown, an embodiment of the present application provides a terminal, which includes a voice signal processing device such as Figure 6 or a voice signal processing device such as Figure 9 shown or a voice signal processing device such as Figure 10 shown. The terminal can be a desktop terminal or a mobile terminal, and the mobile terminal can specifically be at least one of a mobile phone, a tablet computer, and a laptop computer.

[0143] An embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are executed by a processor or a controller. For example, they are executed by Figure 9 a processor 920 in Figure 4 to enable the above-mentioned processor 920 to execute the voice signal processing method in the above embodiment. For example, to execute the method steps 401 to 404 in Figure 5 and the method steps 501 to 502 in

[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] Those of ordinary skill in the art will appreciate that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery media.

[0146] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A method for processing a voice signal, comprising: Obtaining a first voice signal; Obtaining a second voice signal, where the second voice signal is a voice acquisition signal of an echo generated when the first voice signal is played; Performing echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal; Calculating a first power value of a corresponding frequency point of the first voice signal in the frequency domain; Calculating a second power value of a corresponding frequency point of the residual signal in the frequency domain; Calculating a correlation degree between the first voice signal and the residual signal according to the first power value and the second power value, where the correlation degree is a power correlation coefficient of corresponding frequency points of the first voice signal and the residual signal in the frequency domain; Obtaining a corresponding gain adjustment parameter according to the correlation degree, and using the gain adjustment parameter to perform suppression adjustment on the first voice signal.

2. The method according to claim 1, characterized in that The calculation formula of the power correlation coefficient includes: Among them, E(k) is the frequency-domain signal corresponding to the k-th frequency point after Fourier transform of the residual signal, X(k) is the frequency-domain signal corresponding to the k-th frequency point after Fourier transform of the first speech signal, k represents the N-th frequency point, where N is a positive integer.

3. The method according to claim 1, characterized in that, The obtaining a corresponding gain adjustment parameter according to the correlation degree, and using the gain adjustment parameter to perform suppression adjustment on the first voice signal includes: Obtaining a corresponding gain adjustment parameter according to the power correlation coefficient, where the gain adjustment parameter is a gain adjustment coefficient of a corresponding frequency point in the frequency domain; Performing suppression adjustment on the power value of the corresponding frequency point of the first voice signal in the frequency domain according to the gain adjustment coefficient.

4. The method according to claim 1, wherein The obtaining a corresponding gain adjustment parameter according to the correlation degree, and using the gain adjustment parameter to perform suppression adjustment on the first voice signal includes: When the correlation degree is less than a first preset threshold, setting the gain adjustment parameter to a first gain value, and using the first gain value to perform suppression adjustment on the first voice signal; When the correlation degree is greater than or equal to the first preset threshold and less than or equal to a second preset threshold, setting the gain adjustment parameter to a second gain value, and using the second gain value to perform suppression adjustment on the first voice signal; When the correlation degree is greater than the second preset threshold, setting the gain adjustment parameter to a third gain value, and using the third gain value to perform suppression adjustment on the first voice signal, where the second gain value is greater than the third gain value and less than the first gain value, and the first gain value is less than or equal to 1.

5. The method according to any one of claims 1 to 4, characterized in that, The echo generated when the first voice signal is played is generated by driving a speaker with a playback signal obtained according to the first voice signal. The performing echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal includes: Aligning the second voice signal with the playback signal; Performing adaptive filtering processing on the playback signal to obtain a linear echo signal; Performing an inverting process on the linear echo signal and linearly adding it to the aligned second voice signal to obtain a residual signal corresponding to the first voice signal.

6. A voice signal processing apparatus, comprising: A voice acquisition unit for obtaining a first voice signal; A voice collection unit for obtaining a second voice signal, where the second voice signal is a voice collection signal of an echo generated when the first voice signal is played; An echo processing unit for performing echo processing on the second voice signal to obtain a residual signal corresponding to the first voice signal; A suppression adjustment unit for calculating a first power value of a corresponding frequency point of the first voice signal in the frequency domain; Calculating a second power value of a corresponding frequency point of the residual signal in the frequency domain; calculating a correlation degree between the first voice signal and the residual signal according to the first power value and the second power value, where the correlation degree is a power correlation coefficient of corresponding frequency points of the first voice signal and the residual signal in the frequency domain; obtaining a corresponding gain adjustment parameter according to the correlation degree, and using the gain adjustment parameter to perform suppression adjustment on the first voice signal.

7. A voice signal processing device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the voice signal processing method according to any one of claims 1 to 5 is implemented.

8. A terminal, characterized in that, Including the voice signal processing device according to claim 6 or the voice signal processing device according to claim 7.

9. A computer-readable storage medium storing computer-executable instructions for executing the voice signal processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Echo cancellation method and device, computer readable storage medium and computer equipment

    CN110177317A