Audio data processing method and apparatus, call method, audio processing chip, electronic device, and computer-readable storage medium
By performing linear filtering and weighting filtering on the audio data, the weighting factors related to the call status are determined, and the problem of poor echo residue elimination in the prior art is solved, and a more efficient echo suppression effect is achieved, and the audio call quality is improved.
Patent Information
- Application Number
- CN202011073889.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-09
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-10-09
AI Technical Summary
In the prior art, the echo residual elimination effect is poor, and it is difficult to meet the audio call quality requirements in variable and complex scenarios.
Linear echo data is determined by linear filtering of the audio data sent by the first calling party and the audio data collected by the second calling party. Then, based on the linear echo data and the second audio data, the weighting factor related to the call status is determined and a weighted filtering process is performed to suppress the echo.
Improves the echo residual suppression effect, improves call quality, and is suitable for varied and complex audio call scenarios.
Smart Images

Figure CN114333867B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio data processing, and in particular to an audio data processing method and device, a call method, an audio processing chip, an electronic device, and a computer-readable storage medium. Background Art
[0002] As the application scenarios of audio call technology become more and more extensive, people's requirements for call quality are also getting higher and higher. In a normal call process, after one party of the call speaks a voice, it is collected by the call device of the party and transmitted to the other party of the call and played by the voice playback device of the other party's call device, so that the other party of the call can listen. In this process, when the voice audio of one party of the call is played by the voice playback device of the call device of the other party of the call, an echo will be generated in the space where the other party is located, that is, the played voice audio is reflected by the surfaces of various walls or objects in the space, and then when the other party responds to the voice of one party of the call and makes a voice response, it is collected by the voice collection device of the call device of the other party, and then it is regarded as the voice of the other party of the call and transmitted back to the second party of the call. Therefore, one party of the call will receive the audio sound of his own voice transmitted to the other party of the call and then transmitted back again while speaking, that is, a call echo is generated, and such a call echo seriously affects the call experience of the caller.
[0003] In the prior art, the residual energy of the echo is usually estimated based on the delay estimation result of the audio data and the output of the linear filter, so as to adjust the spectral gain of the signal after the linear echo processing. However, the existing echo processing scheme only considers the residual energy of the echo in the audio data during the call. However, with the increasing diversification of the application scenarios of audio call technology, the echo will show different characteristics in different scenarios and environments. Therefore, using a unified residual energy as a benchmark to suppress the echo is difficult to meet people's requirements for audio call quality in complex and changing scenarios. Summary of the invention
[0004] The embodiments of the present application provide an audio data processing method and device, a call method, an audio processing chip, an electronic device, and a computer-readable storage medium to solve the defect of poor echo residual elimination effect in the prior art.
[0005] To achieve the above object, the present application provides an audio data processing method, including:
[0006] Performing linear filtering on first audio data sent by a first caller and second audio data collected by a second caller to obtain linear echo data, wherein the first caller and the second caller are in the same call activity;
[0007] Determine linear output data according to the second audio data collected by the second caller and the linear echo data;
[0008] Determining first state data and second state data for identifying a state of an audio call conducted between the first calling party and the second calling party based on the first audio data and the second audio data, wherein the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-frequency band, and the first state data is an average value of ratios of the linear echo data to the second audio data in each sub-frequency band;
[0009] A weight factor related to the call state is determined according to the first state data and the second state data, so as to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first call party.
[0010] The present application also provides an audio data processing method, including:
[0011] Performing linear filtering on first audio data sent by a first caller and second audio data collected by a second caller to obtain linear echo data, wherein the first caller and the second caller are in the same call activity;
[0012] Determine linear output data according to the second audio data collected by the second caller and the linear echo data;
[0013] Determining first state data and second state data for identifying a state of an audio call conducted between the first calling party and the second calling party based on the first audio data and the second audio data, wherein the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-frequency band, and the first state data is an average value of ratios of the linear echo data to the second audio data in each sub-frequency band;
[0014] According to the first state data and the second state data, selecting a signal reduction amplitude value corresponding to the first state data and the second state data;
[0015] The linear output audio data is subjected to a signal amplitude reduction operation according to the signal reduction amplitude value to obtain audio data sent to the first talking party.
[0016] The embodiment of the present application also provides a calling method, including:
[0017] receiving first audio data;
[0018] Playing the first audio data;
[0019] Performing audio collection processing to generate second audio data, wherein the second audio data at least includes audio data collected when playing the first audio data;
[0020] Performing linear filtering on the second audio data to obtain linear echo data;
[0021] determining linear output data according to the second audio data and the linear echo data;
[0022] Determining first status data and second status data for identifying an audio call status according to the first audio data and the second audio data, wherein the first status data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-frequency band, and the first status data is an average value of ratios of the linear echo data to the second audio data in each sub-frequency band;
[0023] Determine a weight factor related to the call state according to the first state data and the second state data, so as to perform weighted filtering processing on the linear output data to obtain third audio data;
[0024] The third audio data is output to a party in the call.
[0025] The present application also provides an audio processing chip, including:
[0026] An audio receiving module, used for receiving first audio data;
[0027] Audio output, used for playing the first audio data;
[0028] A sound pickup module, configured to perform audio collection processing to generate second audio data, wherein the second audio data at least includes audio data collected by the sound pickup module when playing the first audio data;
[0029] A filtering module, used for performing linear filtering on the second audio data to obtain linear echo data;
[0030] a processing module, configured to determine linear output data according to the second audio data and the linear echo data, determine first state data and second state data for identifying an audio call state according to the first audio data and the second audio data, wherein the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-frequency band, and the first state data is an average value of ratios of the linear echo data to the second audio data in each sub-frequency band; and determine a weight factor associated with the call state according to the first state data and the second state data,
[0031] The filtering module is used to perform weighted filtering on the linear output data to obtain third audio data, and
[0032] The audio output module is used to output the third audio data to a calling party.
[0033] The present application also provides an audio data processing device, including:
[0034] A filtering module, configured to perform linear filtering on first audio data sent by a first call party and second audio data collected by a second call party to obtain linear echo data, wherein the first call party and the second call party are in the same call activity;
[0035] A linear output module, used to determine linear output data according to the second audio data collected by the second call party and the linear echo data;
[0036] a state determination module, configured to determine first state data and second state data for identifying a state of an audio call conducted between the first calling party and the second calling party based on the first audio data and the second audio data, wherein the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-frequency band, and the first state data is an average value of ratios of the linear echo data to the second audio data in each sub-frequency band;
[0037] The suppression module is used to determine a weight factor related to the call state according to the first state data and the second state data, so as to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first call party.
[0038] The present application also provides an electronic device, including:
[0039] Memory, used to store programs;
[0040] The processor is used to run the program stored in the memory, and the audio data processing method provided in the embodiment of the present application is executed when the program is run.
[0041] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program executable by a processor is stored, wherein the program, when executed by the processor, implements the audio data processing method provided in the embodiment of the present application.
[0042] The audio data processing method and device, call method, audio processing chip, electronic device and computer-readable storage medium provided in the embodiments of the present application can determine the first state data and the second state data used to identify the current audio call state according to the audio data sent by the first call party and the collected data of the second call party; further, the microphone collected data after linear filtering can be subjected to targeted suppression processing according to the weighted coefficient or corresponding suppression scheme related to the current audio call state determined by the first state data and the second state data. Thus, weighted filtering can be performed based on the current call state or corresponding suppression scheme can be adopted for processing, so that echo residual suppression processing can be performed considering the component characteristics of echo residual under different call states, which can improve the echo residual suppression effect and effectively improve the call quality.
[0043] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0045] Figure 1 A schematic diagram of an application scenario of the audio data processing method provided in an embodiment of the present application;
[0046] Figure 2 A flowchart of an embodiment of the audio data processing method provided by the present application;
[0047] Figure 3 A flowchart of another embodiment of the audio data processing method provided by the present application;
[0048] Figure 4 A schematic diagram of the structure of an embodiment of an audio data processing device provided by the present application;
[0049] Figure 5 A schematic diagram of the structure of an electronic device embodiment provided in this application. DETAILED DESCRIPTION
[0050] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0051] Embodiment 1
[0052] The solution provided in the embodiments of the present application can be applied to any communication system with audio processing capabilities, such as a communication device equipped with an audio processing module, etc. Figure 1 A schematic diagram of an application scenario of the audio data processing method provided in an embodiment of the present application, Figure 1 The scenario shown is only one example of a scenario to which the technical solution of the present application can be applied.
[0053] With the development of audio technology, the application scenarios of audio call technology are becoming more and more extensive, not only in daily audio calls, but also in the business field. Especially with the recent rise of remote work, more and more users use video conferencing or audio conferencing to communicate. Therefore, users' requirements for call quality are getting higher and higher.
[0054] In a normal two-end call, when one end of the call, for example, the other end opposite to the local end, also called the far end, speaks, the voice is collected by the other party or the far-end call device and then transmitted to the local end of the call, for example, the near end, and played by the audio playback device of the near-end call device, so that the near-end can hear the voice audio emitted by the far-end.
[0055] In this process, when the voice audio emitted by the far end is played by the voice playback device of the call equipment at the near end, an echo will be generated in the space where the near end is located, that is, the played voice audio is propagated in the space where the near end is located, and then collected by the voice collection device of the call equipment at the near end, and thus is treated as the voice audio of the near end and transmitted back to the far end. In this case, the far end is usually still speaking at this time, that is, continuously emitting voice, while the near end is actually in a listening state and not speaking. However, since the echo generated by the audio emitted by the far end at the near end is transmitted back to the call equipment at the far end, the far end's own voice that is transmitted back will be played out while the far end is speaking, so that the far end can also hear its own voice when speaking, that is, the echo of the previously spoken words, that is, a call echo is generated at the far end, and such a call echo seriously affects the call experience of the caller.
[0056] To this end, the prior art has proposed estimating the energy of the echo residue based on the delay estimation result of the transmitted audio data and the output of the linear filter, so as to adjust the spectrum gain of the signal after the linear echo processing, so as to suppress the echo residue in the call audio. Figure 1 As shown in , the downlink data from the far end, that is, the audio data sent from the far end, is played at the near end by a playing device such as a speaker in the near end call device, and can be transmitted along the lines such as Figure 1 The echo path shown in is collected by an audio collection device such as a microphone of a near-end call device, so that the collected data can be delayed together with the far-end downlink data through a delay estimation module to align with the signal including the echo of the far-end audio data collected by the audio collection device such as a microphone in the near-end call device, so as to obtain the delay estimation result of the far-end downlink data at the near-end, and then the delay-adjusted signal and the signal collected by the audio collection device are input into a linear filter for linear echo estimation processing, so as to finally estimate the residual energy of the echo in the near-end uplink data based on the delay estimation result and the linear echo estimation result, so as to suppress the echo in the uplink data.
[0057] However, in actual use, there are many changes in the situations of the two parties in a call. For example, one party may use headphones, the other party may use a speaker, or both parties may be talking, etc. In these cases, the audio signals collected are also very different. Figure 1 In the scenario shown in , when headphones are used at the near end, the echo residual collected by the microphone is very small. Therefore, in the technical solution of the prior art, the echo energy is still estimated through delay estimation and linear filtering estimation, which may mistakenly identify the voice emitted by the near end as an echo. In this case, the near-end voice may be mistakenly suppressed, which in turn affects the far-end listening to the near-end voice audio.
[0058] Therefore, with the increasing diversity of application scenarios of audio call technology, the echo components will present different characteristics in different scenarios and environments. Therefore, using a unified residual energy as a benchmark to suppress echoes is difficult to meet people's requirements for audio call quality in changing and complex scenarios.
[0059] For this reason, Figure 1 As shown in FIG. 1 , a scenario of eliminating the call echo heard by the remote end, that is, the other party at the local end of the call. In the prior art, the first party at the remote end can transmit the voice data sent as downlink data to the call device at the near end in a wired or wireless manner, for example, Figure 1In the process, the downlink data x(t) is transmitted to the speaker and delay estimation module in the call device of the second call party at the near end, so that while the far-end voice audio is played by the speaker of the call device at the near end, the delay estimation module performs delay alignment processing on the downlink data x(t) and the audio data d(t) collected by the microphone of the call device at the near end to obtain a delay estimation result x(t'). Next, the delay estimation result x(t') can be further input into the existing linear echo cancellation filter together with the audio data d(t) collected by the microphone to perform linear echo estimation calculation, so as to obtain a linear echo estimation result y(t) and a corresponding linear output.
[0060] Different from the prior art, in the present application, a processing module may be provided in a call device as the second call party at the near end or in an audio processing chip in the call device, so that after obtaining the delay estimation result x(t') and the linear echo estimation result y(t) as in the prior art solution, the current call state of the second call party may be further determined in the processing module based on the first audio data x(t) as downlink data, the linear echo data y(t) and the second audio data d(t) collected by the microphone. For example, in an embodiment of the present application, in the processing module, the first state data may be determined based on the first audio data x(t) transmitted by the first call party and the audio data d(t) collected by the microphone of the near-end call device, and the second state data may be determined based on the second audio data d(t) transmitted by the second call party and the linear echo data y(t).
[0061] For example, in the embodiment of the present application, Figure 1In the scheme shown in , the first state data may be the average value of the correlation coefficients of the first audio data and the second audio data in each sub-band, and the second state data may be the average value of the ratios of the linear echo data and the second audio data in each sub-band. Therefore, according to the embodiment of the present application, compared with the prior art, the current call state can be further determined according to the original output of the prior art, and the weight factor with the call state can be determined according to the call state, and the weight factor is introduced into the echo suppression processing in the prior art to consider the call state of the second call party, that is, the echo residual component that may exist in the linear output result can be more clearly determined according to the different call states of the second call party, so that the echo residual component can be further suppressed for the linear output result of the prior art scheme. Alternatively, in the embodiment of the present application, a mapping table of the call state and the gain adjustment scheme of the linear output can also be established according to historical experience data, so that when the first state data and the second state data are determined, the pre-established mapping table can be directly queried according to the determined first state data and the second state data to select the corresponding adjustment scheme or gain adjustment factor to directly perform echo residual processing on the linear output result.
[0062] Therefore, according to the present invention, the first state data and the second state data for identifying the current audio call state can be determined according to the audio data sent by the first caller and the collected data of the second caller; further, the microphone collected data after linear filtering can be subjected to targeted suppression processing according to the weighted coefficients or corresponding suppression schemes related to the current audio call state determined by the first state data and the second state data. Thus, weighted filtering can be performed based on the current call state or corresponding suppression schemes can be adopted for processing, so that echo residual suppression processing can be performed considering the component characteristics of echo residual under different call states, which can improve the echo residual suppression effect and effectively improve call quality.
[0063] The above embodiments are illustrations of the technical principles and exemplary application frameworks of the embodiments of the present application. The specific technical solutions of the embodiments of the present application are further described in detail below through multiple embodiments.
[0064] Embodiment 2
[0065] Figure 2 This is a flowchart of an embodiment of the audio data processing method provided by the present application. The execution subject of the method can be various IoT terminals or devices with audio processing capabilities, or it can be a device or chip integrated on these devices. Figure 2 As shown, the audio data processing method includes the following steps:
[0066] S201, performing linear filtering processing on first audio data sent by a first calling party and second audio data collected by a second calling party to obtain linear echo data.
[0067] In the embodiment of the present application, when a first call direction of a remote party at a remote end relative to the local place sends a voice as the first audio data to a second call party at a near end in the same call activity, such as Figure 1 As shown in , when the local end, i.e., the second caller, receives the first audio data, the first audio data can be played, for example, through the speaker of the near-end call device, and the second audio data can be obtained by collecting it through a microphone, so that the first audio data and the second audio data can be linearly filtered. For example, the first audio data received by the first caller and the second audio data collected by the local call device can be input into a linear AEC (Acoustic Echo Cancellation) filter module to obtain linear echo data. That is, the linear echo data is separated by the processing of a linear filter.
[0068] S202: Determine linear output data according to second audio data and linear echo data collected by the second calling party.
[0069] After the linear echo data is obtained in step S201, since the linear echo data can reflect the echo component related to the first audio data sent by the first call party at the far end in the second audio data collected by the communication device at the local place, the linear output data can be further determined based on the second audio data collected by the second call party at the near end as the local place and the linear echo data obtained in S201.
[0070] S203: Determine first status data and second status data for identifying a status of an audio call between a first calling party and a second calling party according to the first audio data and the second audio data.
[0071] In the embodiment of the present application, since the first audio data is the voice data sent by the first party at the far end, and the second audio data is the data collected by the communication device of the second party at the local end, such as a microphone, under normal circumstances, the local end as the near end may have three call states: using headphones to talk, using a speaker to listen to the call of the far end and the near end is not talking, and using a speaker to listen to the call of the far end and the near end is talking.
[0072] For example, in the above three states, there are differences in the situations in which the second audio data collected by the near-end communication device can include the echo data of the first audio data. For example, when the near-end uses headphones to listen to the call at the far-end or uses headphones to speak, the second audio data collected by the microphone of the near-end communication device hardly includes the first audio data sent by the far-end. When the near-end uses a loudspeaker to play the call at the far-end and the near-end does not speak, since the near-end uses the loudspeaker to play the first audio data at the far-end, the first audio data propagated in the space where the near-end is located will be collected by the near-end communication device and therefore included in the second audio data collected by the near-end communication device. In particular, in this state, since the near-end does not speak, the second audio data collected by the near-end communication device is almost all the first audio data. When the near-end uses a loudspeaker to play the first audio data of the far-end and the near-end is speaking at the same time, what is propagated in the space where the near-end is located is not only the first audio data of the far-end, but also the voice data emitted by the near-end in the space. Therefore, the second audio data that can be collected by the near-end communication device can include both the components of the first audio data of the far-end, that is, the echo data, and the words that the near-end is saying.
[0073] Therefore, in the embodiment of the present application, the call state of the near end can be determined based on the first audio data and the second audio data in step S203, that is, the first state data and the second state data identifying the call state of the near end are determined. For example, in the embodiment of the present application, the first state data can be obtained by band averaging the correlation coefficients of the first audio signal and the second audio signal in each self-frequency band. For example, in the embodiment of the present application, the second state data can be obtained by band averaging the ratio of the linear echo data and the second audio data separated by the linear filter processing in each self-frequency band. That is, the first state data is the average value of the correlation coefficients of the first audio data and the second audio data in each sub-frequency band, and the second state data is the average value of the ratio of the linear echo data to the second audio data in each sub-frequency band. Therefore, the first state data and the second state data determined in this way can well distinguish the above three call states of the second call party, so that weighted filtering can be performed based on the call state in subsequent processing.
[0074] S204, determining a weight factor related to the call state according to the first state data and the second state data, so as to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first call party.
[0075] Therefore, the weighting factor related to the current call status of the second call party can be determined based on the first status data and the second status data determined in step S203, and the outputs of steps S201 and S202 can be directly used in step S204 to further perform weighted filtering processing based on the current call status, that is, the weight factor of the call status.
[0076] In other words, the embodiment of the present application can introduce the weight factor into the echo suppression processing in the prior art to take into account the call status of the second caller, that is, according to the different call status of the second caller, the echo residual component that may exist in the linear output result can be more clearly determined, so that the echo residual component of the linear output result of the prior art solution can be further suppressed. Or, in the embodiment of the present application, a mapping table of the call status and the gain adjustment scheme of the linear output can be established according to historical experience data, so that when the first state data and the second state data are determined, the pre-established mapping table can be directly queried according to the determined first state data and the second state data to select the corresponding adjustment scheme or gain adjustment factor to directly perform echo residual processing on the linear output result.
[0077] Therefore, according to the present invention, the first state data and the second state data for identifying the current audio call state can be determined according to the audio data sent by the first caller and the collected data of the second caller; further, the microphone collected data after linear filtering can be subjected to targeted suppression processing according to the weighted coefficients or corresponding suppression schemes related to the current audio call state determined by the first state data and the second state data. Thus, weighted filtering can be performed based on the current call state or corresponding suppression schemes can be adopted for processing, so that echo residual suppression processing can be performed considering the component characteristics of echo residual under different call states, which can improve the echo residual suppression effect and effectively improve call quality.
[0078] Embodiment 3
[0079] Figure 3 This is a flowchart of another embodiment of the audio data processing method provided by the present application. The execution subject of the method can be various communication terminals or devices with audio processing capabilities, or a device or chip integrated on these devices. Figure 3 As shown, the audio data processing method includes the following steps:
[0080] S301: Perform audio activity detection on first audio data to determine whether the first audio data contains voice audio.
[0081] In an embodiment of the present application, since each module for audio processing in the communication terminal consumes a large amount of power, and the two parties usually do not talk all the time during a call, the near end as the local party can perform audio activity detection (Voice Activity Detection, VAD) on the first audio data received from the far end when making a call, so as to start the relevant audio processing module only when it is determined that the first audio data contains voice data, and at other times the relevant audio processing modules are all in standby or dormant state, and until the VAD processing in this step determines that the first audio data contains voice data, that is, it is determined that the first communication party at the far end is talking, a wake-up signal is sent to the relevant audio processing module or a notification signal is sent to the central processor or controller of the communication device, etc., so that the communication device at the far end can wake up the relevant audio processing module to process the first audio data containing the voice of the far end.
[0082] S302: Perform delay alignment processing on the first audio data according to the second audio data.
[0083] In the case where it is determined in step S301 that the received first audio data contains the audio data of the far end, the received first audio data can be delayed aligned, for example, through a delay estimation module, according to the second audio data collected by the near-end communication device in step S302, so as to adjust the delay of the received first audio. In general, since the near-end as the local plays the first audio data through a playing device such as a microphone of the communication device and collects the audio data propagated in the space through a collection module such as a speaker, that is, when obtaining the second audio data, the second audio data needs to be propagated in the space where the near-end is located before it can be collected by the near-end communication device. Therefore, even if the near-end does not speak, the far-end speech component contained in the second audio data collected by the near-end communication device will have a time difference with the first audio data, that is, it takes a period of time for the first audio data to be played by the speaker of the near-end communication device and be received by the microphone after propagating in the space, so it is not aligned with the first audio data on the timeline. Therefore, in step S302, the first audio data may be used to align the second audio data collected by the near-end communication device on the timeline, thereby facilitating the accuracy and efficiency of subsequent echo recognition and comparison processing.
[0084] S303: Perform linear filtering on the first audio data sent by the first call party and the second audio data collected by the second call party to obtain linear echo data.
[0085] In the embodiment of the present application, when a far-end party at a far-end opposite to the local party sends a voice to the near-end party as the first audio data, such as Figure 1 As shown in , when the local end, i.e., the near end, receives the first audio data, it can play the first audio data, for example, through the speaker of the near-end call device, and obtain the second audio data by collecting it through a microphone, so that the first audio data and the second audio data can be linearly filtered. That is, the first audio data is the voice data sent by the first call party to the second call party during the call activity, and the second audio data is the audio data collected by the second call party when playing the first audio data. For example, the first audio data received by the first call party and the second audio data collected by the local call device can be input into, for example, a linear AEC (Acoustic Echo Cancellation) filter module to obtain linear echo data. That is, the linear echo data is separated by the processing of a linear filter.
[0086] S304: Determine linear output data according to the second audio data collected by the second calling party and the linear echo data.
[0087] After the linear echo data is obtained in step S301, since the linear echo data can reflect the echo component related to the first audio data sent by the first call party at the far end in the second audio data collected by the communication device at the local place, the linear output data can be further determined based on the second audio data collected by the second call party at the near end as the local place and the linear echo data obtained in S301.
[0088] S305: Determine the correlation coefficient between the first audio data and the second audio data in each sub-frequency band, and determine an average value of each correlation coefficient as the first state data.
[0089] S306: Determine the ratio of the linear echo data to the second audio data in each sub-frequency band, and determine an average value of each ratio as the second state data.
[0090] In the embodiment of the present application, since the first audio data is the voice data sent by the first party at the far end, and the second audio data is the data collected by the communication device of the second party at the local end, such as a microphone, under normal circumstances, the local end as the near end may have three call states: using headphones to talk, using a speaker to listen to the call of the far end and the near end is not talking, and using a speaker to listen to the call of the far end and the near end is talking.
[0091] For example, in the above three states, there are differences in the first audio data, that is, the echo data, included in the second audio data collected by the near-end communication device. For example, when the near-end uses headphones to listen to the call at the far-end or uses headphones to speak, the second audio data collected by the microphone of the near-end communication device almost does not contain the first audio data sent by the far-end. When the near-end uses a loudspeaker to play the call at the far-end and the near-end does not speak, since the near-end uses the loudspeaker to play the first audio data at the far-end, the first audio data propagated in the space where the near-end is located will be collected by the near-end communication device and therefore included in the second audio data collected by the near-end communication device. In particular, in this state, since the near-end does not speak, the second audio data collected by the near-end communication device is almost all the first audio data. When the near-end uses a loudspeaker to play the first audio data of the far-end and the near-end is speaking at the same time, what is propagated in the space where the near-end is located is not only the first audio data of the far-end, but also the voice data emitted by the near-end in the space. Therefore, the second audio data that can be collected by the near-end communication device can include both the components of the first audio data of the far-end, that is, the echo data, and the words that the near-end is saying.
[0092] Therefore, in the embodiment of the present application, the call status of the near end can be determined based on the first audio data and the second audio data in steps S303 and S304, that is, in step S303, the first status data Coh can be obtained by band averaging the correlation coefficients of the first audio signal and the second audio signal in each sub-band. XD , and the second state data YDR can be obtained by band averaging the ratio of the linear echo data and the second audio data separated by the linear filter processing in each sub-band in step S304. Therefore, in this case, the first state data obtained in step S303 and the second state data obtained in step S304 can be used to identify the state of the audio call between the first call party and the second call party. Therefore, the first state data and the second state data determined in this way can well distinguish the above three call states of the second call party, so that weighted filtering can be performed based on the call state in subsequent processing.
[0093] S307: Determine a trade-off factor for controlling speech distortion according to the first state data and the second state data.
[0094] In the embodiment of the present application, the first state data obtained in step S303 and the second state data obtained in step S304 are used to determine the trade-off factor. In the embodiment of the present application, the following formula can be used to determine the trade-off factor.
[0095]
[0096] Among them, Φ YY and ΦEE are covariance matrices calculated using different frame signals. Therefore, in the embodiment of the present application, a weighting factor that can be used to determine the echo residual component can be obtained through the first state data and the second state data.
[0097] S308, determining inter-frame Wiener filter coefficients according to the weighing factor, the linear echo data and the linear output data.
[0098] After the weighing factor is obtained in step S307, the inter-frame Wiener filter coefficients can be further determined in step S308 according to the weighing factor in step S307 and the linear echo data and linear output data obtained in steps S303 and S304. For example, the inter-frame Wiener filter coefficients can be determined using the following formula.
[0099]
[0100] Where t is time, f is the frequency of the sound, e 1 Represents a calculation parameter.
[0101] S309: filter the linear output data according to the Wiener filter coefficient to obtain first output audio data.
[0102] In step S309, the Wiener filter coefficient determined in step S308 can be used to filter the linear output data obtained in step S304 using a Wiener filter, so as to obtain the first output audio data. The embodiment of the present application can introduce the trade-off factor in the echo suppression process in the prior art to consider the call status of the second call party, that is, the echo residual component that may exist in the linear output result can be more clearly determined according to the different call status of the second call party, so that the echo residual component can be further suppressed for the linear output result of the prior art solution.
[0103] Therefore, in the embodiment of the present application, the near-end call status can be considered to introduce the judgment result of the current near-end call status as a parameter when determining the filter coefficient, thereby improving the effect of the echo suppression processing.
[0104] S310: Determine a frequency band average gain of first output audio data.
[0105] S311, selecting a signal reduction amplitude value corresponding to the frequency band average gain.
[0106] S312: performing an operation of reducing the signal amplitude of the first output audio data according to the signal reduction amplitude value.
[0107] After the first output audio data is obtained in step S309, the frequency band average gain of the first output audio data obtained in step S309 can be further determined in step S310, and the corresponding signal reduction amplitude value can be selected according to the frequency band average gain obtained in step S309 in step S311, so that the signal amplitude of the first output audio data obtained in step S309 can be reduced by using the selected signal reduction amplitude value in step S312, thereby further suppressing the echo residual contained in the first output audio data.
[0108] In addition, in an embodiment of the present application, a mapping table of call status and gain adjustment scheme of linear output can also be established based on historical experience data, so that when the first state data and the second state data are determined, the pre-established mapping table can be directly queried based on the determined first state data and the second state data to select the corresponding adjustment scheme or gain adjustment factor to directly perform echo residual processing on the linear output result.
[0109] For example, in an embodiment of the present application, instead of determining a trade-off factor, a signal reduction amplitude value corresponding to the first state data and the second state data can be selected, and the signal amplitude can be reduced for the linear output audio data according to the signal reduction amplitude value, thereby being able to quickly determine an incremental adjustment scheme for removing echo residuals based on historical experience data.
[0110] Therefore, according to the present invention, the first state data and the second state data for identifying the current audio call state can be determined according to the audio data sent by the first caller and the collected data of the second caller; further, the microphone collected data after linear filtering can be subjected to targeted suppression processing according to the weighted coefficients or corresponding suppression schemes related to the current audio call state determined by the first state data and the second state data. Thus, weighted filtering can be performed based on the current call state or corresponding suppression schemes can be adopted for processing, so that echo residual suppression processing can be performed considering the component characteristics of echo residual under different call states, which can improve the echo residual suppression effect and effectively improve call quality.
[0111] Embodiment 4
[0112] Figure 4 The structure diagram of the audio data processing device embodiment provided in the present application can be used to perform the following steps: Figure 2 and Figure 3 The method steps shown. Figure 4 As shown, the audio data processing device may include: a filtering module 41 , a linear output module 42 , a state determination module 43 and a suppression module 44 .
[0113] The filtering module 41 may be used to perform linear filtering on the first audio data sent by the first call party and the second audio data collected by the second call party to obtain linear echo data.
[0114] In the embodiment of the present application, when a first call direction of a remote party at a remote end relative to the local place sends a voice as the first audio data to a second call party at a near end in the same call activity, such as Figure 1 As shown in , when the local end, i.e., the second caller, receives the first audio data, the first audio data can be played, for example, through the speaker of the near-end call device, and the second audio data can be obtained by collecting it through a microphone, so that the first audio data and the second audio data can be linearly filtered. For example, the first audio data received by the first caller and the second audio data collected by the local call device can be input into a linear AEC (Acoustic Echo Cancellation) filter module to obtain linear echo data. That is, the linear echo data is separated by the processing of a linear filter.
[0115] The linear output module 42 is used to determine linear output data according to the second audio data and the linear echo data collected by the second call party.
[0116] After the filtering module 41 obtains the linear echo data, since the linear echo data can reflect the echo component related to the first audio data sent by the first call party at the far end in the second audio data collected by the communication device at the local place, the linear output module 42 can be further used to determine the linear output data based on the second audio data collected by the second call party at the near end as the local place and the linear echo data obtained by the filtering module 41.
[0117] The state determination module 43 is used to determine first state data and second state data for identifying the state of the audio call between the first calling party and the second calling party according to the first audio data and the second audio data.
[0118] In the embodiment of the present application, since the first audio data is the voice data sent by the first party at the far end, and the second audio data is the data collected by the communication device of the second party at the local end, such as a microphone, under normal circumstances, the local end as the near end may have three call states: using headphones to talk, using a speaker to listen to the call of the far end and the near end is not talking, and using a speaker to listen to the call of the far end and the near end is talking.
[0119] Therefore, in the embodiment of the present application, the state determination module 43 can determine the call state of the near end based on the first audio data and the second audio data, that is, determine the first state data and the second state data that identify the call state of the near end. For example, in the embodiment of the present application, the first state data can be obtained by the first determination unit 431 of the state determination module 43 by band averaging the correlation coefficients of the first audio signal and the second audio signal in each self-frequency band. For example, in the embodiment of the present application, the second state data can be obtained by the second determination unit 432 of the state determination module 43 by band averaging the ratio of the linear echo data and the second audio data separated by the linear filter processing in each self-frequency band.
[0120] The suppression module 44 may be configured to determine a weight factor associated with the call state according to the first state data and the second state data, so as to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first call party.
[0121] Therefore, the weighting factor related to the current call state of the second call party can be determined according to the first state data and the second state data determined by the state determination module 43, and the suppression module 44 can directly use the output of the filtering module 41 and the linear output module 42 to further perform weighted filtering processing according to the current call state, that is, the weight factor of the call state.
[0122] In other words, the embodiment of the present application can introduce the weight factor into the echo suppression processing in the prior art to take into account the call status of the second caller, that is, according to the different call status of the second caller, the echo residual component that may exist in the linear output result can be more clearly determined, so that the echo residual component of the linear output result of the prior art solution can be further suppressed. Or, in the embodiment of the present application, a mapping table of the call status and the gain adjustment scheme of the linear output can be established according to historical experience data, so that when the first state data and the second state data are determined, the pre-established mapping table can be directly queried according to the determined first state data and the second state data to select the corresponding adjustment scheme or gain adjustment factor to directly perform echo residual processing on the linear output result.
[0123] For example, in the embodiment of the present application, the suppression module 44 may include: a third determination unit 441 , a fourth determination unit 442 , and a filtering unit 443 .
[0124] The third determining unit 441 may be configured to determine a trade-off factor for controlling speech distortion according to the first state data and the second state data.
[0125] In the embodiment of the present application, the trade-off factor is determined by using the first state data obtained by the first determination unit 431 and the second state data obtained by the second determination unit 432. In the embodiment of the present application, the trade-off factor μ can be determined using the following formula.
[0126]
[0127] Among them, Φ YY and Φ EE is a covariance matrix determined using different frame signals. Therefore, in the embodiment of the present application, such a weighting factor that can be used to determine the echo residual component can be obtained through the first state data and the second state data.
[0128] The fourth determination unit 442 may be configured to determine an inter-frame Wiener filter coefficient according to the trade-off factor, the linear echo data, and the linear output data.
[0129] After the third determination unit 441 obtains the trade-off factor, the fourth determination unit 442 may further determine the inter-frame Wiener filter coefficient according to the trade-off factor determined by the third determination unit 441, the linear echo data obtained by the filter module 41, and the linear output data obtained by the linear output module 42. For example, the inter-frame Wiener filter coefficient w(t,f) may be determined using the following formula.
[0130]
[0131] Where t is time, f is the frequency of the sound, e 1 Represents a calculation parameter.
[0132] The filtering unit 443 can be used to filter the linear output data according to the Wiener filter coefficient to obtain the first output audio data.
[0133] In the embodiment of the present application, the filtering unit 443 can use the Wiener filter coefficient determined by the third determining unit 441 to filter the linear output data obtained by the linear output module 42 using the Wiener filter, so as to obtain the first output audio data. Therefore, in the embodiment of the present application, the trade-off factor can be introduced into the echo suppression process in the prior art to consider the call state of the second call party, that is, the echo residual component that may exist in the linear output result can be more clearly determined according to the different call states of the second call party, so that the echo residual component can be further suppressed for the linear output result of the prior art solution.
[0134] In addition, the audio processing device of the embodiment of the present application may further include: a gain determination module 47 , a selection module 48 and a signal amplitude adjustment module 49 .
[0135] The gain determination module 47 may be used to determine a frequency band average gain of the first output audio data;
[0136] The selection module 48 may be used to select a signal reduction amplitude value corresponding to the average gain of the frequency band;
[0137] The signal amplitude adjustment module 49 may be configured to reduce the signal amplitude of the first output audio data according to the signal amplitude reduction value.
[0138] After the filtering unit 443 obtains the first output audio data, the gain determination module 47 can be further used to determine the frequency band average gain of the first output audio data obtained by the filtering unit 443, and the selection module 48 can be used to select the corresponding signal reduction amplitude value according to the frequency band average gain obtained by the gain determination module 47, so as to use the signal amplitude adjustment module 49 to use the selected signal reduction amplitude value to reduce the signal amplitude of the first output audio data obtained by the filtering unit 443, thereby achieving further suppression of the echo residual contained in the first output audio data.
[0139] In addition, the suppression module 44 may further include: a selection unit 444 and a signal amplitude adjustment unit 445 .
[0140] The selection unit 444 may be configured to select a signal reduction amplitude value corresponding to the first state data and the second state data.
[0141] The signal amplitude adjustment unit 445 may be configured to reduce the signal amplitude of the linear output audio data according to the signal reduction amplitude value.
[0142] In addition, after the first state data and the second state data have been obtained in the embodiment of the present application, the selection unit 444 can also be used to directly query the signal reduction amplitude value corresponding to the state data from, for example, a pre-set table according to the current call state of the near end identified by the first state data and the second state data, and the signal amplitude adjustment unit 445 can use the signal reduction amplitude value selected by the step selection unit 444 to reduce the signal amplitude of the linear output audio obtained by the linear output module 42.
[0143] The voice detection module 45 may be configured to perform audio activity detection on the first audio data to determine whether the first audio data contains voice audio.
[0144] In the embodiment of the present application, since the various modules for audio processing in the communication terminal consume a large amount of power, and the two parties usually do not talk all the time during a call, the near end as the local party can use the voice detection module 45 to perform audio activity detection (Voice Activity Detection, VAD) on the first audio data received from the far end when making a call, so as to start the relevant audio processing module only when it is determined that the first audio data contains voice data, and make the relevant audio processing modules in standby or dormant state at other times, and until the VAD processing in this step determines that the first audio data contains voice data, that is, it is determined that the first communication party at the far end is talking, a wake-up signal is sent to the relevant audio processing module or a notification signal is sent to the central processor or controller of the communication device, so that the communication device at the far end can wake up the relevant audio processing module to process the first audio data containing the voice of the far end.
[0145] The delay alignment module 46 may be configured to perform delay alignment processing on the first audio data according to the second audio data.
[0146] When the voice detection module 45 determines that the received first audio data contains the far-end audio data, the delay alignment module 46 can be used to perform delay alignment on the received first audio data according to the second audio data collected by the near-end communication device, for example, through the delay estimation module, so as to adjust the delay of the received first audio. In general, since the near-end as the local plays the first audio data through the playing device of the communication device, such as a microphone, and collects the audio data propagated in the space through the collection module, such as a speaker, that is, when obtaining the second audio data, the second audio data needs to be propagated in the space where the near-end is located before it can be collected by the near-end communication device. Therefore, even if the near-end does not speak, the far-end speech component contained in the second audio data collected by the near-end communication device will have a time difference with the first audio data, that is, it takes a period of time for the first audio data to be played by the speaker of the near-end communication device and be received by the microphone after propagating in the space, so it is not aligned with the first audio data on the timeline. Therefore, the first audio data can be used, for example, by the delay alignment module 46 of the near-end communication device to align the second audio data collected by the near-end communication device on the timeline, thereby facilitating the accuracy and efficiency of subsequent echo recognition and comparison processing.
[0147] Therefore, according to the audio data processing device of the embodiment of the present application, the first state data and the second state data used to identify the current audio call state can be determined according to the audio data sent by the first call party and the collected data of the second call party; further, the microphone collected data after linear filtering can be subjected to targeted suppression processing according to the weighted coefficient or corresponding suppression scheme related to the current audio call state determined by the first state data and the second state data. Thus, weighted filtering can be performed based on the current call state or corresponding suppression schemes can be adopted for processing, so that echo residual suppression processing can be performed considering the component characteristics of echo residual under different call states, which can improve the echo residual suppression effect and effectively improve call quality.
[0148] Embodiment 5
[0149] The internal functions and structure of the audio data processing device are described above. The device can be implemented as an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device embodiment provided by this application. Figure 5 As shown, the electronic device includes a memory 51 and a processor 52 .
[0150] The memory 51 is used to store programs. In addition to the above programs, the memory 51 can also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0151] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0152] The processor 52 is not limited to a central processing unit (CPU), but may also be a processing chip such as a graphics processing unit (GPU), a field programmable gate array (FPGA), an embedded neural network processor (NPU) or an artificial intelligence (AI) chip. The processor 52 is coupled to the memory 51 and executes a program stored in the memory 51. When the program is executed, the audio data processing method of the above-mentioned embodiments 2 and 3 is executed.
[0153] Further, if Figure 5 As shown, the electronic device may also include: a communication component 53, a power component 54, an audio component 55, a display 56 and other components. Figure 5Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 Components shown.
[0154] The communication component 53 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 53 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 53 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0155] The power supply component 54 provides power to various components of the electronic device. The power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0156] The audio component 55 is configured to output and / or input audio signals. For example, the audio component 55 includes a microphone (MIC), and when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 51 or sent via the communication component 53. In some embodiments, the audio component 55 also includes a speaker for outputting audio signals.
[0157] The display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0158] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio data processing method, comprising: performing linear filtering processing on first audio data sent by a first calling party and second audio data collected by a second calling party to obtain linear echo data, wherein the first calling party and the second calling party are in the same call activity; determining linear output data according to the second audio data collected by the second calling party and the linear echo data; determining first state data and second state data for identifying an audio call state between the first calling party and the second calling party according to the first audio data and the second audio data, wherein the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-band, and the second state data is an average value of ratios of the linear echo data and the second audio data in each sub-band; determining a weight factor related to the call state according to the first state data and the second state data to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first calling party.
2. The audio data processing method according to claim 1, wherein, the first audio data is voice data sent by the first calling party to the second calling party during the call activity, and the second audio data is audio data collected by the second calling party when playing the first audio data.
3. The audio data processing method according to claim 1 or 2, wherein, the determining a weight factor related to the call state according to the first state data and the second state data to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first calling party includes: determining a weight factor for controlling voice distortion according to the first state data and the second state data; determining an inter-frame Wiener filter coefficient according to the weight factor, the linear echo data and the linear output data; performing filtering processing on the linear output data according to the Wiener filter coefficient to obtain first output audio data as the third audio data.
4. The audio data processing method according to claim 3, wherein, the method further includes: determining a band average gain of the first output audio data; selecting a signal reduction amplitude value corresponding to the band average gain; performing an operation of reducing the signal amplitude on the first output audio data according to the signal reduction amplitude value.
5. The audio data processing method according to claim 1, wherein, before the determining a weight factor related to the call state according to the first state data and the second state data to perform weighted filtering processing on the linear output data to obtain third audio data sent to the first calling party, the method further includes: performing audio activity detection on the first audio data to determine whether the first audio data contains voice audio.
6. The audio data processing method according to claim 1, wherein, Before determining, according to the first audio data and the second audio data, first state data and second state data for identifying an audio call state between the first calling party and the second calling party, the method further includes: Performing delay alignment processing on the first audio data according to the second audio data.
7. An audio data processing method, including: Performing linear filtering processing on first audio data sent by a first calling party and second audio data collected by a second calling party to obtain linear echo data, where the first calling party and the second calling party are in the same call activity; Determining linear output data according to the second audio data collected by the second calling party and the linear echo data; Determining, according to the first audio data and the second audio data, first state data and second state data for identifying an audio call state between the first calling party and the second calling party, where the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-band, and the first state data is an average value of ratios of the linear echo data and the second audio data in each sub-band; Selecting a signal reduction amplitude value corresponding to the first state data and the second state data according to the first state data and the second state data; Performing an operation of reducing the signal amplitude on the linear output audio data according to the signal reduction amplitude value to obtain audio data sent to the first calling party.
8. A call method, including: Receiving first audio data; Playing the first audio data; Performing audio acquisition processing to generate second audio data, where the second audio data at least includes audio data collected when playing the first audio data; Performing linear filtering processing on the second audio data to obtain linear echo data; Determining linear output data according to the second audio data and the linear echo data; Determining, according to the first audio data and the second audio data, first state data and second state data for identifying an audio call state, where the first state data is an average value of correlation coefficients of the first audio data and the second audio data in each sub-band, and the first state data is an average value of ratios of the linear echo data and the second audio data in each sub-band; Determining a weight factor related to the call state according to the first state data and the second state data to perform weighted filtering processing on the linear output data to obtain third audio data; Outputting the third audio data to a calling party in a call.
9. An audio processing chip, including: An audio receiving module for receiving first audio data; An audio output module for playing the first audio data; A sound pickup module for performing audio acquisition processing to generate second audio data, where the second audio data at least includes audio data collected by the sound pickup module when playing the first audio data; A filtering module, configured to perform linear filtering on the second audio data to obtain linear echo data; A processing module, configured to determine linear output data according to the second audio data and the linear echo data, and determine first state data and second state data for identifying the audio call state according to the first audio data and the second audio data, where the first state data is the average value of the correlation coefficients of the first audio data and the second audio data in each sub-band, and the first state data is the average value of the ratios of the linear echo data and the second audio data in each sub-band; and determine a weight factor related to the call state according to the first state data and the second state data, wherein, the filtering module is configured to perform weighted filtering on the linear output data to obtain third audio data, and the audio output module is configured to output the third audio data to the calling party of the call.
10. An audio data processing device, comprising: A filtering module, configured to perform linear filtering on first audio data sent by a first calling party and second audio data collected by a second calling party to obtain linear echo data, where the first calling party and the second calling party are in the same call activity; A linear output module, configured to determine linear output data according to the second audio data collected by the second calling party and the linear echo data; A state determination module, configured to determine first state data and second state data for identifying the audio call state between the first calling party and the second calling party according to the first audio data and the second audio data, where the first state data is the average value of the correlation coefficients of the first audio data and the second audio data in each sub-band, and the first state data is the average value of the ratios of the linear echo data and the second audio data in each sub-band; A suppression module, configured to determine a weight factor related to the call state according to the first state data and the second state data, so as to perform weighted filtering on the linear output data to obtain third audio data sent to the first calling party.
11. An electronic device, comprising: A memory, configured to store a program; A processor, configured to run the program stored in the memory, and when the program runs, execute the method according to any one of claims 1 to 8.
12. A computer-readable storage medium, on which a computer program executable by a processor is stored, wherein, when the program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Call signal processing method and device, electronic equipment and storage medium
CN110971769A
Filtering method and device and electronic equipment
CN111524498A