Audio processing method, system and storage medium

By intercepting audio in the first electronic device and transmitting it to the second electronic device for playback and compensation calculations, the echo cancellation problem under the audio player of the non-electronic device itself is solved, thereby improving the quality of voice communication and user experience.

CN115802243BActive Publication Date: 2025-12-19LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211526342.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-01
Publication Date
2025-12-19
Estimated Expiration
2042-12-01

AI Technical Summary

Technical Problem

In real-time audio calls or live streaming scenarios, when using an audio player that is not the electronic device's own, the correct reference audio cannot be obtained, resulting in echo cancellation failure and affecting call quality and user experience.

Method used

By intercepting audio in the first electronic device and transmitting it to the audio player of the second electronic device for playback, and combining audio delay and EQ compensation calculations, a reference audio is obtained for echo cancellation.

Benefits of technology

It achieves high-quality echo cancellation in scenarios where the audio player is not on the electronic device itself, thus improving the quality of voice communication and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115802243B_ABST
    Figure CN115802243B_ABST
Patent Text Reader

Abstract

The application provides an audio processing method, system and storage medium. When a first audio collector of a first electronic device is in a collection state and a second audio player of a second electronic device is in a playing state, after a first audio to be played by the first electronic device is obtained, the first audio can be sent to the second audio player of the second electronic device for playing. The first electronic device acquires a second audio collected by the first audio collector, which can include audio played by the second audio player. Then, compensation operation is performed on the first audio and the second audio to obtain reference audio for the second audio player. Thus, the reference audio is used to perform noise reduction operation on the second audio, and the obtained third audio is output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application mainly relates to the technical field of signal processing, and more particularly to an audio processing method, system and storage medium. BACKGROUND

[0002] In real-time audio call or live broadcast scenarios, electronic devices can use techniques such as AEC (Acoustic Echo Cancelling), AGC (Automatic Gain Control), and / or ANS (Active Noise Control) to improve call or live broadcast quality and user experience.

[0003] Among them, the echo cancellation technology is realized by using the echo cancellation method, which requires the use of the loudspeaker of the electronic device to play the received audio signal, so as to realize the echo filtering of the audio signal collected by the microphone, which limits the application scenarios and the types of electronic devices used. SUMMARY

[0004] Therefore, the present application provides an audio processing method, which comprises:

[0005] When the first audio collector of the first electronic device is in a collection state and the second audio player of the second electronic device connected to the first electronic device is in a playing state, obtaining a first audio to be played by the first electronic device;

[0006] Obtaining a second audio collected by the first audio collector, wherein the second audio comprises an audio played by the second audio player after the second electronic device obtains the first audio from the first electronic device;

[0007] Compensating the first audio and the second audio to obtain a reference audio;

[0008] Using the reference audio to perform a noise reduction operation on the second audio, and outputting a third audio obtained.

[0009] Optionally, the obtaining of the first audio to be played by the first electronic device comprises:

[0010] Detecting a first audio transmitted to the first audio player of the first electronic device, and intercepting the first audio to prevent the first audio player from playing the first audio;

[0011] The method further comprises:

[0012] transmit the intercepted first audio to the second electronic device, and play the first audio by a second audio player of the second electronic device.

[0013] Optionally, the obtaining of the first audio to be played by the first electronic device comprises:

[0014] receiving the first audio, and transmitting the first audio to an audio separation device, so that the audio separation device obtains two identical first audios;

[0015] receiving one of the first audios fed back by the audio separation device;

[0016] The audio separation device also transmits the other first audio to a second electronic device connected to the first electronic device.

[0017] The application further provides an audio processing method, which comprises:

[0018] receiving first audio sent by a first electronic device when a first audio player of the first electronic device is in a playing state and a second audio collector of a second electronic device is in a collecting state, wherein the first audio is audio intercepted by the first electronic device and sent to a virtual audio player of the second electronic device, and the intercepted first audio can be transmitted to the first audio player for playing;

[0019] obtaining second audio collected by the second audio collector of the second electronic device, wherein the second audio comprises audio played by the first audio player;

[0020] performing compensation operation on the first audio and the second audio to obtain reference audio;

[0021] performing noise reduction operation on the second audio by using the reference audio, and outputting obtained third audio.

[0022] Optionally, the compensation operation on the first audio and the second audio to obtain reference audio comprises:

[0023] performing time delay compensation on the first audio by using pre-stored audio time delay parameter to obtain fourth audio;

[0024] performing EQ compensation operation on the second audio and the fourth audio to obtain EQ compensation parameter for the second audio player;

[0025] performing EQ compensation on the fourth audio by using the EQ compensation parameter to obtain reference audio for the second audio player.

[0026] Optionally, the EQ compensation operation on the second audio and the fourth audio obtains the EQ compensation parameter for the second audio player, including:

[0027] The frequency domain analysis is respectively performed on the second audio and the fourth audio to obtain corresponding second audio spectrum and fourth audio spectrum;

[0028] The comparative analysis is performed on the second audio spectrum and the fourth audio spectrum to obtain the frequency band compensation parameter of the fourth audio;

[0029] The time domain conversion processing is performed on the frequency band compensation parameter to obtain the EQ compensation parameter for the second audio player.

[0030] Optionally, the comparative analysis on the second audio spectrum and the fourth audio spectrum to obtain the frequency band compensation parameter of the fourth audio includes:

[0031] The normalization processing is respectively performed on the second audio spectrum and the fourth audio spectrum to obtain the amplitude-normalized second audio spectrum and fourth audio spectrum;

[0032] The difference between the signal amplitudes of the amplitude-normalized second audio spectrum and fourth audio spectrum under the same frequency band is obtained;

[0033] The difference under the different frequency bands is restored to obtain the frequency band compensation parameter of the fourth audio; wherein, the restoration processing is opposite to the normalization processing.

[0034] Optionally, the audio delay parameter obtaining process includes:

[0035] The difference value operation is performed on the first audio before and after the interception to obtain an interception delay parameter; and / or,

[0036] The difference value operation is performed on the first audio after the interception and the first audio received by the second electronic device to obtain a transmission delay parameter;

[0037] The audio delay parameter is obtained by using the obtained interception delay parameter and / or transmission delay parameter.

[0038] The application further proposes an audio processing system, the system including a first electronic device and a second electronic device in communication connection, wherein:

[0039] When the first audio collector of the first electronic device is in the collection state and the second audio player of the second electronic device is in the playing state, the first electronic device obtains the first audio to be played, and sends the first audio to the second audio player of the second electronic device for playing;

[0040] The first electronic device acquires second audio collected by the first audio collector, performs compensation operation on the first audio and the second audio, obtains reference audio, performs noise reduction operation on the second audio by using the reference audio, and outputs obtained third audio; wherein the second audio includes audio played by the second audio player.

[0041] Alternatively,

[0042] In a case where the first audio player of the first electronic device is in a playing state and the second audio collector of the second electronic device is in a collecting state, the first electronic device intercepts audio sent to the virtual audio player of the second electronic device, and transmits the intercepted first audio to the first audio player for playing.

[0043] The second electronic device receives the first audio sent by the first electronic device, acquires second audio collected by the second audio collector, performs compensation operation on the first audio and the second audio, obtains reference audio, performs noise reduction operation on the second audio by using the reference audio, and outputs obtained third audio; wherein the second audio includes audio played by the first audio player. The application also proposes a computer readable storage medium having a computer program stored thereon, wherein the computer program is loaded and executed by a processor to implement the audio processing method as described above.

[0044] Therefore, the application proposes an audio processing method, system and storage medium. In a case where the first audio collector of the first electronic device is in a collecting state and the second audio player of the second electronic device is in a playing state, after obtaining first audio to be played by the first electronic device, the first audio can be sent to the second audio player of the second electronic device for playing. The first electronic device acquires second audio collected by the first audio collector, which can include audio played by the second audio player. Then, compensation operation is performed on the first audio and the second audio to obtain reference audio for the second audio player, so that noise reduction operation is performed on the second audio by using the reference audio, and obtained third audio is output. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on the provided drawings.

[0046] Figure 1Structural schematic diagram of an optional example of an audio processing system proposed in the present application;

[0047] Figure 2 Structural schematic diagram of another optional example of an audio processing system proposed in the present application;

[0048] Figure 3a Structural schematic diagram of an optional voice communication scenario suitable for an audio processing method proposed in the present application;

[0049] Figure 3b Structural schematic diagram of another optional voice communication scenario suitable for an audio processing method proposed in the present application;

[0050] Figure 4 Structural schematic diagram of another optional example of an audio processing system proposed in the present application;

[0051] Figure 5 Structural schematic diagram of another optional voice communication scenario suitable for an audio processing method proposed in the present application;

[0052] Figure 6 Structural schematic diagram of another optional example of an audio processing method proposed in the present application;

[0053] Figure 7 Structural schematic diagram of another optional example of an audio processing method proposed in the present application;

[0054] Figure 8a Structural schematic diagram of another optional voice communication scenario suitable for an audio processing method proposed in the present application;

[0055] Figure 8b Structural schematic diagram of another optional voice communication scenario suitable for an audio processing method proposed in the present application;

[0056] Figure 9 Structural schematic diagram of another optional example of an audio processing method proposed in the present application;

[0057] Figure 10 Structural schematic diagram of another optional example of an audio processing method proposed in the present application;

[0058] Figure 11 Structural schematic diagram of another optional example of an audio processing method proposed in the present application;

[0059] Figure 12 Structural schematic diagram of another optional voice communication scenario suitable for an audio processing method proposed in the present application;

[0060] Figure 13 Structural schematic diagram of another optional example of an audio processing method proposed in the present application;

[0061] Figure 14 Structure diagram of an optional example of the audio processing device proposed in the present application;

[0062] Figure 15 Structure diagram of another optional example of the audio processing device proposed in the present application. DETAILED DESCRIPTION

[0063] For the description in the background section, in the case of using a non-electronic device's audio player to play the audio received by the electronic device in the speech communication scenario such as video conference, live broadcast, etc., the correct reference audio cannot be obtained in the echo cancellation process, resulting in the failure to eliminate the echo audio, so that the remote user in communication with the electronic device can obviously hear his own echo, and the voice of the near-end user will also be eliminated as noise, resulting in the remote user's inability to hear the speaking content of the near-end user, which limits the use scenario and customer experience of the electronic device.

[0064] To solve the above problems, the present application proposes to directly collect the audio played by another electronic device, and intercept the audio sent by the remote end through software or hardware, and then perform compensation operation on the two audio to obtain the reference audio, which is used to realize echo cancellation on the collected audio, output the obtained high-quality audio, solve the echo interference problem, and improve the speech communication quality.

[0065] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0066] Reference Figure 1 Structure diagram of an optional example of the audio processing system proposed in the present application, which can include a first electronic device 100 and a second electronic device 200, wherein:

[0067] In different speech communication scenarios, the device types and their constituent structures of the first electronic device 100 and the second electronic device 200 can be different, including but not limited to the structure of the e-book device and its connection mode described in the following embodiments.

[0068] In the speech communication scenario using the audio collection function of the first electronic device 100 and the audio playing function of the second electronic device 200 connected to the first electronic device 100, such as Figure 2As shown, the first electronic device 100 can include a first communication interface 110, a first audio collector 120, a first memory 130 and a first processor 140; the second electronic device 200 can include at least a second communication interface 210 and a second audio player 220, wherein:

[0069] The first communication interface 110 can be connected with the second communication interface 210 to realize data transmission between the first electronic device 100 and the second electronic device 200, and the interface type of the communication interface is not limited in the present application, which can be determined according to the device type of the electronic device.

[0070] For example, in an optional voice communication scenario as shown in Figure 3a , the first electronic device 100 is a HEC all-in-one machine, an AIO all-in-one machine or an NB transmission device, and the second electronic device 200 is a display or a TV, and the two electronic devices can be connected through an HDMI (High Definition Multimedia Interface) data line. Therefore, the first communication interface 110 and the second communication interface 210 can both include an HDMI interface, but are not limited thereto.

[0071] In another optional voice communication scenario as shown in Figure 3b , the first electronic device 100 can be a TDT / NB device, and the second electronic device 200 can be an external playback device, such as at least one loudspeaker as shown in Figure 3b , and the two electronic devices can be connected through a USB data line or a headphone line. At this time, the first communication interface 110 and the second communication interface 210 can both include a corresponding USB interface or a headphone port, such as a 3.5mm interface.

[0072] It should be understood that if the first electronic device 100 and the second electronic device 200 are connected in other ways, the two electronic devices can be configured with corresponding types of communication interfaces, and in order to meet the connection of the same electronic device with different types of electronic devices and meet various voice communication scenarios, the first communication interface 110 and the second communication interface 210 can both include multiple different types of communication interfaces, including but not limited to the types of interfaces listed above. In addition, the first communication interface 110 can also include an interface capable of accessing a wireless communication network, such as a WIFI communication interface, a GPRS communication interface, etc., to realize voice communication between the first electronic device 100 and a third electronic device, and the voice communication process will not be described in detail.

[0073] The first audio collector 120 in the first electronic device 100 can be, for example, Figure 3a , and Figure 3bThe microphone shown; the second audio player in the second electronic device 200 can be, for example Figure 3a and Figure 3b The loudspeaker / speaker, etc. shown, and the number, structure and deployment position of the loudspeaker and microphone in each electronic device are not limited in the present application, and can be determined as appropriate.

[0074] The first memory 130 can be used to store a first program for implementing the audio processing method described in the following method embodiments performed by the first electronic device 100; the first processor 140 can load and execute the first program stored in the first memory 130 to implement each step of the audio processing method described in the following method embodiments, and the specific implementation process can be referred to the description of the corresponding part of the following embodiments, which will not be described in detail herein.

[0075] The first memory 130 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device or other volatile solid-state storage device. The first processor 140 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready-to-program gate array (FPGA) or other programmable logic device, etc.

[0076] It should be understood that Figure 2 The structure of the first electronic device and the second electronic device shown does not constitute a limitation on the corresponding electronic devices in the embodiments of the present application, and in actual applications, the first electronic device and the second electronic device can include more or fewer components than those shown, or combine some components, such as a display, an antenna, a power module, and an audio separation circuit for intercepting the first audio sent to the first audio player in the first electronic device, etc., which will not be enumerated one by one herein. Figure 2

[0077] In some other embodiments of the present application, voice communication can also be performed using the audio playing function of the first electronic device 100 and the audio collecting function of the second electronic device 200, at which time the first electronic device 100 can be connected to a third electronic device through wireless communication or wired communication to receive the first audio sent by the third electronic device. Based on this, for example Figure 4 ​As shown, the first electronic device 100 can further include a first audio player 150, and of course, if it includes the first audio collector 120, the first electronic device 100 can be prohibited from working, i.e., the voice collection function of the first electronic device 100 is disabled in the current voice communication scenario. Alternatively, the selected first electronic device 100 can not have the first audio collector 120 and the like.

[0078] Correspondingly, the second electronic device 200 having the audio collection function can include a second communication interface 210, a second audio collector 230, a second memory 240, and a second processor 250. Of course, it can also include a second audio player 220, but in the voice communication scenario of the present embodiment, the second audio player 220 is turned off, and the audio playing function of the second electronic device 200 is disabled. As for the types of the second audio collector 230, the second memory 240, and the second processor 250 in the present embodiment, reference can be made to the related description of the same types of devices in the first electronic device 100 above, and the present embodiment will not be described in detail.

[0079] It should be noted that in the voice communication scenario as shown in Figure 5 As shown, in the voice communication scenario, after the first electronic device 100 receives the first audio, it can send the first audio to the first audio player 150 for playing, and at the same time, it can also send the first audio to the second electronic device 200, so that the second electronic device 200 performs the audio processing method described below from the method embodiment executed by the second electronic device side, compensates the first audio and the second audio collected by the second audio collector 230, uses the reference audio obtained for the first audio player 150 to perform echo cancellation on the collected second audio, and sends the obtained third audio to the first electronic device 100. After that, the first electronic device 100 can send the third audio to the third electronic device through the wireless / wired communication network, thereby ensuring the call quality in the current voice communication scenario.

[0080] As for the implementation process of the audio processing method executed by the first electronic device and the second electronic device connected thereto in different voice communication scenarios described above, reference can be made to the description of the corresponding part of the embodiment below, and the present embodiment will not be described in detail.

[0081] It should be understood that the structure of the audio processing system shown in the drawings does not constitute a limitation on the audio processing system in the present embodiment, and in actual applications, the audio processing system can include more or fewer devices than those shown in the drawings, such as a communication server supporting remote voice communication of the first electronic device, a third electronic device for voice communication with the first electronic device, a database for storing communication content, and the like, which will not be enumerated one by one.

[0082] In combination with the audio processing system and the voice communication scenarios described above, the implementation process of the audio processing method will be described from the first electronic device or the second electronic device side for different types of voice communication scenarios, but it is not limited to the implementation method described in the following embodiments, and can be adjusted as needed.

[0083] Referring to Figure 6 An optional example of the audio processing method proposed in this application is shown in the flowchart. The method can be described from the first electronic device side, which is installed with a filtering driver or an audio separation circuit with audio interception function, etc. Referring to the voice communication scenarios shown in Figure 3a or Figure 3b When the audio acquisition function of the first electronic device and the audio playback function of the second electronic device connected to the first electronic device meet the voice communication requirements in the corresponding voice communication scenario, the first electronic device performs the audio processing method as shown in Figure 6 The audio processing method performed by the first electronic device can include:

[0084] In step S61, when the first audio collector of the first electronic device is in the acquisition state, and the second audio player of the second electronic device connected to the first electronic device is in the playback state, the first audio to be played by the first electronic device is obtained.

[0085] In practical application, if the first electronic device determines that its first audio collector is in the acquisition state, and the second audio player of the second electronic device connected to the first electronic device is in the playback state, it can be explained that the current voice communication scenario can be as shown in Figure 3a or Figure 3b The first electronic device can use its own microphone and the speaker of the second electronic device for voice communication. The local audio is collected by the microphone of the first electronic device, and the audio sent by the third electronic device (i.e. any terminal for voice communication with the first electronic device) to the first electronic device is played by the speaker of the second electronic device.

[0086] In the above method, when the first electronic device is in the voice communication scenario as shown in Figure 3a or Figure 3b Echo cancellation can be used for noise reduction in this type of voice communication scenario. After the first electronic device receives the first audio sent by the third electronic device, it can intercept the first audio before sending it to the first audio player of the first electronic device, i.e. obtaining the first audio to be played by the first electronic device. The first audio interception process can be realized by software / hardware such as filtering driver or audio separation circuit in the first electronic device, which is not limited in this application.

[0087] Step S62, obtaining the second audio collected by the first audio collector;

[0088] In the embodiments of the present application, in order to obtain the reference audio for the second audio player of the second electronic device, so as to realize the echo cancellation in the voice communication scenario described above, the first audio sent to the first audio player and intercepted will also be transmitted to the second electronic device and played by the second audio player of the second electronic device. In this way, the second audio played by the second audio player can be collected by the first audio collector of the first electronic device.

[0089] It can be seen that the second audio obtained in step S62 can include the audio played by the second audio player after the first audio from the first electronic device is obtained by the second electronic device. Of course, if the user of the first electronic device speaks during this process, the second audio can also include the audio of the user, etc. The content of the second audio can be determined according to the call situation in the voice communication scenario and the environment in which it is located.

[0090] Step S63, compensating the first audio and the second audio to obtain the reference audio;

[0091] Step S64, using the reference audio to perform noise reduction operation on the second audio, and outputting the obtained third audio.

[0092] In combination with the related description above of the implementation process of how the first electronic device obtains the first audio and the second audio, considering the time delay problem caused by the interception process of the first audio and the transmission process to the second electronic device, and the difference in different frequency bands between the audio played by the second audio player and the first audio directly received by the first electronic device, the first audio and the second audio can be analyzed to determine the compensation value in multiple aspects. After the intercepted first audio is compensated, the reference audio for the second audio player is obtained. Then, the echo cancellation on the second audio collected by the first audio collector can be realized according to this, such as subtracting the reference audio from the second audio to obtain the third audio.

[0093] In order to improve the signal quality of the third audio sent by the first electronic device to the third electronic device, in addition to the above-mentioned echo cancellation and noise reduction processing method, noise reduction techniques such as AGC (Automatic Gain Control) and / or ANS (Active Noise Control) can also be used to reduce the noise of the collected second audio, improve the call quality, and the implementation process will not be described in detail.

[0094] It should be understood that after the first electronic device obtains the third audio in the manner described above, according to actual needs in the voice communication scenario, the first electronic device can send the third audio to the third electronic device through the first communication module. If the voice communication scenario of the third electronic device is similar to the voice communication scenario of the first electronic device, the third electronic device can also use the audio processing method described above to perform echo cancellation on the received third audio, play the obtained fourth audio, and the like, and the implementation process is not described herein. Of course, after the first electronic device obtains the third audio, the first electronic device can also send the third audio to the first memory of the first electronic device for storage, and the output mode of the third audio is not limited in the present application.

[0095] In summary, in the embodiments of the present application, in the voice communication scenario in which the first audio collector of the first electronic device is in the collection state and the second audio player of the second electronic device connected to the first electronic device is in the playing state, after the first electronic device obtains the first audio to be played, the first electronic device sends the first audio to the second audio player for playing, and the first audio collector collects audio to obtain second audio containing the audio played by the second audio player. By compensating the first audio and the second audio, the reference audio for the second audio player can be obtained to realize echo cancellation of the second audio collected by the first audio collector, and high-quality third audio can be output. It can be seen that the audio processing method proposed in the present application can be applied to an online voice communication scenario in which a non-electronic device audio player is used to play audio received by an electronic device, meets the echo cancellation needs in the scenario, and improves the voice communication quality and customer experience.

[0096] Reference Figure 7 For another optional example of the audio processing method proposed in the present application, the present embodiment can describe an optional refinement of the audio processing method described above. Referring to Figure 2 or the voice communication scenario shown in FIG. 3, that is, the first audio collector of the first electronic device is in the collection state, and the second audio player of the second electronic device connected to the first electronic device is in the playing state, as shown in Figure 7 The audio processing method executed by the first electronic device can include the following steps:

[0097] Step S71, detecting the first audio transmitted to the first audio player of the first electronic device, and intercepting the first audio to prevent the first audio player from playing the first audio;

[0098] In the embodiments of the present application, in combination with Figure 3a and Figure 3bIn the illustrated voice communication scenario, after the first electronic device obtains the first audio through the wired communication network or the wireless communication network, if the first electronic device has a first audio player, the first electronic device usually sends the first audio to the first audio player for playing. However, in the voice communication scenario described in this embodiment, the first audio needs to be played by using the second audio player of the second electronic device. Therefore, after receiving the first audio, the first electronic device intercepts the first audio sent to the first audio player, and prevents the first audio player from playing the first audio. The method for intercepting the first audio is not limited in this application.

[0099] Optionally, referring to Figure 8a In the illustrated flowchart, the first electronic device can start an audio filtering driver to intercept the first audio to be played, such as the audio received by a network device such as a VoIP (Voice over Internet Protocol) gateway from the cloud, and prevent the first audio from being transmitted to the loudspeaker of the first electronic device for playing. In some other embodiments, as illustrated in Figure 8b In the illustrated flowchart, the first electronic device can also deploy an HDMI audio separation module to intercept the first audio to be played received by a network device in a hardware manner, and prevent the loudspeaker of the first electronic device from playing the first audio. The structure of the hardware circuit of the HDMI audio separation module is not limited in this application.

[0100] In some other embodiments of the present application, as illustrated in Figure 8b In the illustrated flowchart, when the HDMI audio separation module is used as a separate audio separation device, the audio separation device can be connected to the HDMI interfaces, i.e., multimedia communication interfaces, of the first electronic device and the second electronic device respectively. After the first electronic device receives the first audio sent from the cloud, the first electronic device transmits the first audio to the audio separation device through the HDMI interface, and the audio separation device obtains two identical first audios and transmits the two first audios to the first electronic device and the second electronic device respectively. The first electronic device receives one of the first audios fed back by the audio separation device, and the second electronic device receives the other first audio sent by the audio separation device, and then continues to process the first audio according to the method described below.

[0101] It should be noted that the method for intercepting the first audio described in step S71 includes but is not limited to the implementation manners described above.

[0102] In step S72, the intercepted first audio is transmitted to the second electronic device, and the second audio player of the second electronic device plays the first audio.

[0103] In step S73, the second audio collected by the first audio collector is obtained. The second audio includes the audio played by the second audio player.

[0104] In combination with the above description of the voice communication scenario applicable to the present embodiment, the second audio player of the second electronic device is used to play the first audio, while the first audio collector of the first electronic device will perform audio collection to obtain the second audio, and the implementation process can refer to the description of the corresponding part in the above embodiment.

[0105] In step S74, the first audio is compensated for delay using the pre-stored audio delay parameter to obtain the fourth audio.

[0106] In actual applications, the interception of the first audio received by the first electronic device by the above method will cause a certain delay, and the transmission process of the intercepted first audio to the second audio player of the second electronic device for playing will also have a certain delay, and the delay caused by different transmission methods is different. To this end, in order to ensure the reliability of the subsequent echo cancellation, the intercepted first audio can be compensated for delay, and the obtained fourth audio can be used in the subsequent echo cancellation algorithm.

[0107] Based on this, the present application can consider the above-mentioned multi-aspect delay problem for delay testing in the voice communication scenario described above, and store the audio delay parameter after obtaining it. During the delay testing process, the difference between the first audio before and after interception can be calculated according to the audio interception method described above to obtain the interception delay parameter; and / or, the difference between the intercepted first audio and the first audio received by the second electronic device can be calculated to obtain the transmission delay parameter. Then, the obtained interception delay parameter and / or transmission delay parameter can be used to obtain the audio delay parameter, which can be determined according to the delay compensation accuracy and is not limited to the audio delay parameter acquisition method described in the present embodiment.

[0108] In step S75, the second audio and the fourth audio are subjected to EQ compensation operation to obtain the EQ compensation parameter for the second audio player.

[0109] In step S76, the fourth audio is compensated for EQ using the EQ compensation parameter to obtain the reference audio for the second audio player.

[0110] Since the first audio intercepted by the first electronic device cannot know the post-processing EQ compensation parameter of the second audio player, the present application proposes to analyze the frequency domain signal difference between the second audio and the fourth audio, determine the compensation value of each frequency band, convert it to the time domain, and obtain the EQ compensation parameter for the second audio player, so as to realize the EQ compensation of the fourth audio after delay compensation, and obtain the reference audio of the second audio player in the voice communication scenario of the present embodiment, i.e. the reference signal for echo cancellation. The implementation process of the EQ compensation operation is not described in detail.

[0111] Step S77, using the reference audio, performing noise reduction operation on the second audio, and outputting the obtained third audio.

[0112] In yet some embodiments of the present disclosure, as shown in the flow diagram of another optional example of the audio processing method of the present disclosure, Figure 9 As shown in the flow diagram of another optional example of the audio processing method of the present disclosure, Figure 9 As shown in the flow diagram of another optional example of the audio processing method of the present disclosure,

[0113] Step S91, using the pre-stored audio delay parameter, performing delay compensation on the intercepted first audio to obtain fourth audio;

[0114] For the implementation process of step S91, refer to the description of the corresponding part of the above embodiments, and the present embodiment will not be repeated.

[0115] Step S92, performing frequency domain analysis on the fourth audio and the second audio collected by the first audio collector to obtain the corresponding second audio spectrum and fourth audio spectrum;

[0116] Step S93, performing normalization processing on the second audio spectrum and the fourth audio spectrum respectively to obtain the amplitude-normalized second audio spectrum and fourth audio spectrum;

[0117] Refer to Figure 10 Another optional flow diagram of the audio processing method applicable to the voice communication scenario Figure 3a Since the above audios are time domain signals, they can be converted to frequency domain first, and then the compensation values in each frequency band are analyzed. In order to facilitate subsequent operations, the amplitudes of the second audio spectrum and the fourth audio spectrum can be normalized according to the pre-set amplitude range to obtain the amplitude-normalized audio spectrum. The conversion processing method between the time domain and the frequency domain of the audio, and the amplitude normalization processing method will not be described in detail.

[0118] Then, the present disclosure can compare and analyze the second audio spectrum and the fourth audio spectrum to obtain the frequency band compensation parameter of the fourth audio. The implementation process can refer to but not limited to the implementation method described in the following steps S94 and S95, and the present disclosure will be described by taking this as an example.

[0119] Step S94, obtaining the difference between the signal amplitudes of the amplitude-normalized second audio spectrum and the fourth audio spectrum in the same frequency band;

[0120] Step S95, the difference value under the different frequency bands obtained is restored to obtain the frequency band compensation parameter of the fourth audio;

[0121] Step S96, the time domain conversion processing is performed on the frequency band compensation parameter to obtain the EQ compensation parameter for the second audio player;

[0122] After the above amplitude normalization processing, the signal amplitude of the second audio and the fourth audio is converted to the preset range, and then the signal amplitude corresponding to the same frequency band of the two is calculated to obtain the compensation value of each frequency band of the reference audio. After the amplitude restoration processing opposite to the above amplitude normalization processing, and after conversion to the time domain, the EQ compensation parameter for the second audio player can be obtained. The calculation process is not described in detail. The method for obtaining the EQ compensation parameter includes but is not limited to the method described in the embodiment.

[0123] Step S97, the EQ compensation parameter is used to perform EQ compensation on the fourth audio to obtain the reference audio for the second audio player.

[0124] It can be seen that in the voice communication scenario using the first electronic device and the third electronic device, the audio is played using the first electronic device, the audio from the third electronic device is played using the second audio player of the second electronic device, and the EQ compensation parameter of the second audio player is obtained by the above method before the first electronic device feeds back the collected second audio. The reference audio for the echo cancellation of the second audio player is obtained, so that the automatic echo canceller can use the reference audio to perform echo cancellation on the collected second audio, solve the interference of the echo signal, and improve the call quality.

[0125] Optionally, as shown in Figure 10 After the first electronic device receives the first audio, in order to ensure the reliability of the subsequent compensation operation, the first audio can be subjected to automatic gain control AGC, equalization processing, amplification processing, etc. The third audio after echo cancellation can also be subjected to noise suppression processing by using other noise reduction technologies, and then transmitted to the third electronic device after automatic gain control AGC, to further improve the voice communication quality. The implementation process is not described in detail.

[0126] Referring to Figure 11 , the flowchart of another optional example of the audio processing method proposed in the present application can be applied to Figure 5The illustrated voice communication scenario is that the second electronic device has a separate DSP and the like processor, and a microphone and a loudspeaker and the like hardware, supports data processing, audio acquisition and audio playing functions, and the first electronic device has a loudspeaker and supports audio playing function, uses the audio playing function of the first electronic device and the audio acquisition function and data processing function of the second electronic device to meet the voice communication requirement in the semantic communication scenario, such as Figure 11 As shown, the audio processing method executed by the second electronic device can include:

[0127] Step S111, when the first audio player of the first electronic device connected by the second electronic device is in a playing state and the second audio collector of the second electronic device is in an acquisition state, receiving the first audio sent by the first electronic device;

[0128] In combination with Figure 5 As shown in the voice communication scenario, the second electronic device (such as a USB portable audio device) can report a virtual loudspeaker to the first electronic device, and the audio filtering driver of the first electronic device can transmit the intercepted audio to the virtual loudspeaker in real time, therefore, the first audio received by the second electronic device can be the audio sent by the first electronic device to the virtual audio player of the second electronic device, and the interception process of the first audio can refer to the description of the corresponding part of the embodiment described above on the first electronic device side, and this embodiment will not be described in detail. In addition, the first electronic device will also transmit the intercepted first audio to the first audio player for playing, so that the user of the first electronic device can know the voice communication content of the communication opposite end.

[0129] In the embodiment of the present application, in combination with the related description of the system embodiment and the scenario embodiment above, the first electronic device can send the intercepted first audio to the second electronic device through the USB data line. Therefore, the second processor (such as DSP) of the second electronic device can receive the first audio received by the USB interface connected to the first electronic device.

[0130] Step S112, acquiring the second audio collected by the second audio collector of the second electronic device; the second audio includes the audio played by the first audio player;

[0131] As described above, the present application uses the audio acquisition function of the second electronic device for voice communication, and can acquire the second audio collected by the second audio collector in real time. In this process, if the first audio player plays the intercepted first audio, the second audio can include the interference audio played by the first audio player.

[0132] Step S113, compensating the first audio and the second audio to obtain a reference audio;

[0133] Step S114, using the reference audio, performing noise reduction operation on the second audio, and outputting the obtained third audio.

[0134] In some embodiments, referring to Figure 12 the flowchart shown, the second electronic device has echo delay compensation and adaptive EQ compensation functions, therefore, regarding the implementation process of step S113, please refer to Figure 13 and the related description of the above detailed method embodiments for the step of "performing compensation operation on the first audio and the second audio to obtain the reference audio", the present embodiments will not be repeated here.

[0135] According to the above description of various embodiments, for different voice communication scenarios, the corresponding compensation algorithm can be configured in the first electronic device or the second electronic device, and the software or hardware circuit for implementing the audio interception function can be configured to obtain the reference audio for echo cancellation of the audio player of the other electronic device, to realize echo cancellation of the audio collected by the present electronic device, and to improve the call quality and user experience.

[0136] Referring to Figure 14 , the structure diagram of an optional example of the audio processing device proposed in the present application, which can be described from the first electronic device side, as shown in Figure 14 , the device can include:

[0137] The first audio obtaining module 141 is configured to obtain the first audio to be played by the first electronic device when the first audio collector of the first electronic device is in a collecting state and the second audio player of the second electronic device connected to the first electronic device is in a playing state.

[0138] The second audio obtaining module 142 is configured to obtain the second audio collected by the first audio collector, wherein the second audio includes the audio played by the second audio player after the second electronic device obtains the first audio from the first electronic device.

[0139] The reference audio obtaining module 143 is configured to perform compensation operation on the first audio and the second audio to obtain the reference audio.

[0140] The noise reduction module 144 is configured to use the reference audio to perform noise reduction operation on the second audio, and output the obtained third audio.

[0141] Optionally, the first audio obtaining module 141 can include:

[0142] The first audio interception unit is configured to detect the first audio transmitted to the first audio player of the first electronic device, intercept the first audio, and prevent the first audio player from playing the first audio.

[0143] Based on this, the above-mentioned device can further include:

[0144] The first audio output module is configured to transmit the intercepted first audio to the second electronic device, and the first audio is played by a second audio player of the second electronic device.

[0145] Optionally, in the case where the first electronic device is connected to the audio separation device and the second electronic device respectively, the first audio obtaining module 141 can also include:

[0146] The first audio output unit is configured to receive the first audio and transmit the first audio to the audio separation device, so that the audio separation device obtains two identical first audios;

[0147] The first audio receiving unit is configured to receive one of the first audios fed back by the audio separation device;

[0148] The audio separation device also transmits the other first audio to the second electronic device connected to the first electronic device.

[0149] Referring to Figure 15 For another optional example of the audio processing device proposed in the present application, the device can be described from the second electronic device side. As shown in Figure 15 The device can include:

[0150] The first audio receiving module 151 is configured to receive the first audio sent by the first electronic device when a first audio player of the first electronic device connected to the second electronic device is in a playing state and a second audio collector of the second electronic device is in a collecting state;

[0151] The first audio is the audio intercepted by the first electronic device and sent to the virtual audio player of the second electronic device, and the intercepted first audio can be transmitted to the first audio player for playing;

[0152] The second audio obtaining module 152 is configured to obtain the second audio collected by the second audio collector of the second electronic device; the second audio includes the audio played by the first audio player;

[0153] The reference audio obtaining module 153 is configured to perform compensation operation on the first audio and the second audio to obtain a reference audio;

[0154] The noise reduction module 154 is configured to perform noise reduction operation on the second audio by using the reference audio, and output the obtained third audio.

[0155] In some embodiments, the reference audio obtaining module 143 and the reference audio obtaining module 153 in the above embodiments can both include:

[0156] a delay compensation unit, configured to perform delay compensation on the first audio by using a pre-stored audio delay parameter, to obtain fourth audio;

[0157] an EQ compensation parameter obtaining unit, configured to perform EQ compensation operation on the second audio and the fourth audio, to obtain an EQ compensation parameter for the second audio player;

[0158] a reference audio obtaining unit, configured to perform EQ compensation on the fourth audio by using the EQ compensation parameter, to obtain reference audio for the second audio player.

[0159] Optionally, the EQ compensation parameter obtaining unit can include:

[0160] a frequency domain analysis unit, configured to perform frequency domain analysis on the second audio and the fourth audio respectively, to obtain corresponding second audio spectrum and fourth audio spectrum;

[0161] a comparison analysis unit, configured to perform comparison analysis on the second audio spectrum and the fourth audio spectrum, to obtain a frequency band compensation parameter of the fourth audio;

[0162] a time domain conversion unit, configured to perform time domain conversion processing on the frequency band compensation parameter, to obtain the EQ compensation parameter for the second audio player.

[0163] Optionally, the comparison analysis unit can include:

[0164] a normalization processing unit, configured to perform normalization processing on the second audio spectrum and the fourth audio spectrum respectively, to obtain amplitude-normalized second audio spectrum and fourth audio spectrum;

[0165] a difference value obtaining unit, configured to obtain a difference value between signal amplitudes of the amplitude-normalized second audio spectrum and fourth audio spectrum in the same frequency band;

[0166] a restoration processing unit, configured to perform restoration processing on the difference values in different frequency bands obtained, to obtain the frequency band compensation parameter of the fourth audio; wherein the restoration processing is opposite to the normalization processing.

[0167] In order to obtain the above-mentioned audio delay parameter, the delay compensation unit can include:

[0168] an intercepted delay parameter obtaining unit, configured to perform difference operation on the first audio before and after interception, to obtain an intercepted delay parameter; and / or,

[0169] The transmission delay parameter obtaining unit is configured to obtain a transmission delay parameter by performing a difference operation on the intercepted first audio and the first audio received by the second electronic device.

[0170] The audio delay parameter obtaining unit is configured to obtain an audio delay parameter by using the obtained interception delay parameter and / or the transmission delay parameter.

[0171] It should be noted that, as to various modules, units, etc. in the above-mentioned device embodiments, they can be stored in the memory of the corresponding electronic device as program modules, and the processor of the electronic device executes the above-mentioned program modules stored in the memory to realize the corresponding functions. As to the functions realized by each program module and its combination, and the technical effects achieved, reference can be made to the descriptions of the corresponding parts of the above-mentioned method embodiments, and the present embodiment will not be described herein.

[0172] The present application also provides a computer readable storage medium, on which computer readable instructions can be stored, which can be called and loaded by a processor to realize each step of the above-mentioned audio processing method.

[0173] Finally, it should be noted that, as to each embodiment described above, unless the context clearly indicates otherwise, “one”, “a”, “an” and / or “the” do not necessarily specify a single number, but can also include a plurality. Generally speaking, the terms “comprise” and “include” only indicate that the steps and elements explicitly identified are included, and these steps and elements do not constitute an exclusive list, and the method or device can also include other steps or elements. The element defined by the statement “comprises one” does not exclude the presence of another identical element in the process, method, product or device comprising the element.

[0174] In the description of the embodiments of the present application, unless otherwise specified, “ / ” represents or, for example, A / B can represent A or B; “and / or” in the present text only describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, “multiple” means two or more than two.

[0175] The terms such as “first”, “second” etc. involved in the present application are only for the purpose of description, used to distinguish one operation, unit or module from another operation, unit or module, and do not necessarily require or imply any such actual relationship or sequence between these units, operations or modules. And it cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features, therefore, the features with “first”, “second” can explicitly or implicitly include one or more of the features.

[0176] In addition, various embodiments described in this specification are presented by way of example, and the implementation of any embodiment described herein is not limited to the specific embodiments described. The various embodiments can be implemented in any combination of hardware and / or software. Each of the embodiments can be implemented alone or in combination with one another. Various embodiments can be implemented using any hardware device or group of hardware devices.

[0177] The foregoing description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio processing method, the method comprising: When the first audio collector of the first electronic device is in the acquisition state and the second audio player of the second electronic device connected to the first electronic device is in the playback state, the first audio to be played by the first electronic device is obtained. Acquire the second audio collected by the first audio collector, wherein the second audio includes the audio played by the second audio player after the second electronic device obtains the first audio from the first electronic device; Compensation calculations are performed on the first audio and the second audio to obtain a reference audio. Using the reference audio, noise reduction is performed on the second audio to output the resulting third audio; The step of obtaining the first audio to be played by the first electronic device includes: The system detects first audio being transmitted to the first audio player of the first electronic device, and intercepts the first audio to prevent the first audio player from playing the first audio. The intercepted first audio is transmitted to the second electronic device, where the second audio player plays the first audio.

2. The method according to claim 1, wherein obtaining the first audio to be played by the first electronic device comprises: Upon receiving the first audio, the first audio is transmitted to the audio separation device so that the audio separation device obtains two identical first audio streams; Receive one of the first audio signals fed back by the audio separation device; The audio separation device also transmits another audio signal from the first audio source to a second electronic device connected to the first electronic device.

3. An audio processing method, the method comprising: When the first audio player of the first electronic device connected to the second electronic device is in playback mode and the second audio collector of the second electronic device is in acquisition mode, the first audio is received by the first electronic device; wherein, the first audio is the audio intercepted by the first electronic device and sent to the virtual audio player of the second electronic device, and the intercepted first audio can be transmitted to the first audio player for playback; Acquire a second audio signal collected by a second audio acquisition device of a second electronic device; the second audio signal includes the audio played by the first audio player; Compensation calculations are performed on the first audio and the second audio to obtain a reference audio. Using the reference audio, noise reduction is performed on the second audio, and the resulting third audio is output.

4. The method according to any one of claims 1-3, wherein performing compensation operations on the first audio and the second audio to obtain a reference audio comprises: Using pre-stored audio delay parameters, delay compensation is applied to the first audio to obtain the fourth audio. Perform EQ compensation calculations on the second audio and the fourth audio to obtain EQ compensation parameters for the second audio player; Using the EQ compensation parameters, the fourth audio is EQ compensated to obtain a reference audio for the second audio player.

5. The method according to claim 4, wherein performing EQ compensation calculation on the second audio and the fourth audio to obtain EQ compensation parameters for the second audio player includes: Frequency domain analysis was performed on the second audio and the fourth audio respectively to obtain the corresponding second audio spectrum and fourth audio spectrum; By comparing and analyzing the spectrum of the second audio and the spectrum of the fourth audio, the frequency band compensation parameters of the fourth audio are obtained; The frequency band compensation parameters are subjected to time-domain transformation to obtain the EQ compensation parameters for the second audio player.

6. The method according to claim 5, wherein the step of comparing and analyzing the second audio spectrum and the fourth audio spectrum to obtain the frequency band compensation parameters of the fourth audio includes: The second audio spectrum and the fourth audio spectrum are normalized respectively to obtain the second audio spectrum and the fourth audio spectrum after amplitude normalization; Obtain the difference between the signal amplitudes of the second audio spectrum and the fourth audio spectrum after amplitude normalization in the same frequency band; The differences obtained under different frequency bands are restored to obtain the frequency band compensation parameters of the fourth audio; wherein the restoration process is the reverse of the normalization process.

7. The method according to claim 4, wherein the audio delay parameter acquisition process includes: The interception delay parameter is obtained by performing a difference calculation on the first audio before and after the interception. And / or, The transmission delay parameter is obtained by performing a difference calculation between the intercepted first audio and the first audio received by the second electronic device. The audio delay parameter is obtained by using the interception delay parameter and / or the transmission delay parameter.

8. An audio processing system, the system comprising a first electronic device and a second electronic device communicatively connected, wherein: When the first audio collector of the first electronic device is in the acquisition state and the second audio player of the second electronic device is in the playback state, the first electronic device obtains the first audio to be played and sends the first audio to the second audio player of the second electronic device for playback; wherein, the first electronic device obtaining the first audio to be played includes: detecting the first audio being transmitted to the first audio player of the first electronic device, intercepting the first audio to prevent the first audio player from playing the first audio; and transmitting the intercepted first audio to the second electronic device, whereby the second audio player of the second electronic device plays the first audio; The first electronic device acquires the second audio collected by the first audio collector, performs compensation operations on the first audio and the second audio to obtain a reference audio, uses the reference audio to perform noise reduction on the second audio, and outputs the obtained third audio; wherein, the second audio includes the audio played by the second audio player; or, When the first audio player of the first electronic device is in playback mode and the second audio collector of the second electronic device is in acquisition mode, the first electronic device intercepts the audio sent to the virtual audio player of the second electronic device and transmits the intercepted first audio to the first audio player for playback. The second electronic device receives the first audio sent by the first electronic device, acquires the second audio collected by the second audio collector, performs compensation operations on the first audio and the second audio to obtain a reference audio, uses the reference audio to perform noise reduction on the second audio, and outputs the obtained third audio; wherein, the second audio includes the audio played by the first audio player.

9. A computer-readable storage medium having a computer program stored thereon, the computer program being loaded and executed by a processor to implement the audio processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and apparatus for earphone sound effect compensation and an earphone

    US20170171657A1

  • Electronic device for performing wireless communication, and wireless communication method

    WO2021107589A1