Audio playing method, device and conference system

CN122802625APending Publication Date: 2026-09-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329485.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,目前的网络会议会受限于播放源,一旦播放源是单声道格式,即便播放设备具备立体声播放的硬件条件,也无法实现出立体声或者更高阶的播放效果

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802625A_ABST
    Figure CN122802625A_ABST
Patent Text Reader

Abstract

An audio playing method, an audio processing method, an apparatus, a device and a conference system. The audio playing method comprises receiving audio carrying sound source direction information, determining the direction of a speaker of the audio relative to a terminal, and processing the single-channel audio into at least two channels of audio based on the direction to achieve a stereo or higher-order playing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, and in particular to an audio playback method, an audio processing method, an apparatus, a device, and a conference system. Background Technology

[0002] With the development of network technology, web conferencing is becoming increasingly accepted and is gradually becoming a mainstream communication method. However, current web conferencing is limited by the playback source. If the playback source is in mono format, even if the playback device has the hardware for stereo playback, it cannot achieve stereo or higher-order playback effects. Therefore, how to achieve stereo or higher-order playback effects based on mono audio is an important research direction. Summary of the Invention

[0003] This application provides an audio playback method, an audio processing method, an apparatus, a device, and a conference system.

[0004] Firstly, an audio playback method is provided, in which the playback end receives audio carrying sound source location information to determine the speaker's location relative to the terminal, and then processes the mono audio into at least two audio streams based on the location to achieve stereo or higher-order playback effects.

[0005] Optionally, the sound source location information is carried in the target audio in the form of an audio watermark, thus providing a processing method with lower processing difficulty and avoiding excessive processing pressure on the terminal. Furthermore, since the audio watermark has minimal impact on the audio itself, it will not affect the playback effect for playback terminals that do not possess the aforementioned mono audio processing capabilities, greatly expanding the applicability of this audio playback method.

[0006] Optionally, when the playback device acquires at least two audio channels, it can adjust at least one of the delay and amplitude of the mono audio based on the sound source location information to obtain the left channel audio and the right channel audio.

[0007] The aforementioned adjustments can refer to adjusting the delay of the mono audio to obtain an audio that is delayed relative to the existing mono audio. This adjusted audio is used as the right channel audio. The amplitude of the mono audio is then adjusted to obtain a left channel audio that is distinct from the right channel audio. These are then played through different channels to achieve stereo or higher-order playback effects. Through these adjustments, the two audio streams can be made to differ in playback time, intensity, or volume, creating a stereo effect. Optionally, the delayed audio can be used as the left channel audio, and the amplitude-adjusted audio as the right channel audio; this is not a limitation.

[0008] The adjustments described above are not limited to the methods mentioned. Other methods can also be used. For example, based on the sound source's location information, the amplitude of the mono audio can be adjusted to obtain the left channel audio; the delay of the mono audio can be adjusted based on the sound source's location information to obtain the right channel audio. Another example is adjusting both the amplitude and delay of the mono audio based on the sound source's location information to obtain the left channel audio; the same applies to the right channel audio. In this case, the left and right channel audios differ in both amplitude and delay. The choice of adjustment method can be based on the terminal settings, as long as it results in differences in playback time, intensity, or volume between the two audio streams.

[0009] Optionally, the microphone and playback end are conference terminals. Users can access network conferences through the conference terminals. The microphone adds sound source location information to the audio and transmits the audio over the network. The playback end then analyzes and processes the received audio to obtain left and right channel audio for playback, thereby producing stereo or even better sound effects.

[0010] Secondly, an audio processing method is provided in which the microphone embeds the speaker's location information into the acquired audio and sends it to other terminals for playback. This ensures that even if the audio acquired by the microphone is mono, the sound source location information can still provide a reference for other terminals to process the audio and obtain multi-channel audio for stereo playback.

[0011] Optionally, the sound source location information is embedded into the audio in the form of an audio watermark, thus providing a processing method with lower processing difficulty and avoiding excessive processing pressure on the terminal.

[0012] Optionally, the terminal is equipped with an array microphone (MIC) for audio acquisition. The array microphone can serve as the hardware basis for determining the location information of a sound source. For example, the terminal can locate the sound source based on the array microphone to obtain the sound source location information. Since the array microphone is a common audio acquisition component, this positioning method does not require high hardware capabilities from the terminal, making the method provided in this application highly practical.

[0013] Optionally, audio acquisition can be initiated based on human voice to avoid consuming terminal computing resources. In other embodiments, human voice can be extracted from the acquired raw audio to avoid interference from other sounds, which greatly improves audio quality.

[0014] Optionally, the above method can be applied to two types of conferencing systems: one type encodes audio on the client side for data transmission, and the other type encodes audio in the cloud for data transmission. The steps performed by the terminal differ slightly between these two types of systems. For example, in a conferencing system that encodes audio on the client side, the microphone and playback terminals can be conferencing terminals themselves. The microphone encodes the target audio and sends the encoded target audio to the playback terminal. In a conferencing system that encodes audio in the cloud, the microphone and playback terminals can conduct the meeting through the cloud. That is, the microphone sends the target audio to the target service, which encodes the target audio and sends it to the playback terminal.

[0015] Thirdly, an audio playback device is provided, the device including at least one functional module for implementing the method provided in the first aspect or any alternative method of the first aspect.

[0016] Fourthly, an audio processing apparatus is provided, the apparatus including at least one functional module for implementing the method provided in the second aspect or any alternative method of the second aspect.

[0017] Fifthly, a terminal is provided, the terminal including a processor and a memory, the processor being connected to the memory, and an audio device being used to implement the method provided by the first aspect, the second aspect, or any alternative approach of the first and second aspects.

[0018] In a sixth aspect, a computer-readable storage medium is provided for storing at least one piece of program code for implementing the method provided by any of the above aspects or any alternative method of any of the above aspects.

[0019] In a seventh aspect, a computer program product is provided, which is used to implement the method provided in any of the above aspects or any alternative manner of any of the above aspects.

[0020] Eighthly, a conferencing system is provided, comprising a first terminal and a second terminal, the second terminal being used to acquire audio from a local environment; embedding sound source location information of the audio in the local environment into the audio to obtain target audio, the sound source location information indicating the location of the audio source relative to the first terminal; sending the target audio to the first terminal; the first terminal being used to receive the target audio from the second terminal, the target audio carrying sound source location information indicating the location of the audio source relative to the second terminal; processing the target audio to obtain mono audio and the sound source location information; processing the mono audio based on the sound source location information to obtain at least two audio streams; and playing the at least two audio streams. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application;

[0022] Figure 2 A schematic diagram of a principle provided for an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of the architecture of a conference system provided in an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of the architecture of a conference system provided in an embodiment of this application;

[0025] Figure 5 This is an exemplary flowchart of an audio playback method provided in this application embodiment;

[0026] Figure 6 This is an exemplary flowchart of an audio playback method provided in this application embodiment;

[0027] Figure 7 This is a schematic diagram of the structure of an audio playback device provided in an embodiment of this application;

[0028] Figure 8 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application;

[0029] Figure 9 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that all information (including but not limited to interfaces), data (including but not limited to files, directories, etc.) and signals involved in this application are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the interfaces involved in this application are all obtained under fully authorized conditions.

[0031] To facilitate understanding, the key terms and concepts involved in this application will be explained below.

[0032] An array microphone is a system consisting of multiple microphones arranged in a certain spatial distribution and topology. Common array microphones include linear arrays, planar arrays, and stereo arrays. The array configuration is based on the requirements for sound acquisition.

[0033] Sound source localization refers to using information such as the time difference of arrival (TDOA) or phase difference when multiple microphones receive the same sound signal, and then using algorithms to calculate the azimuth and elevation angles of the sound source relative to the microphone array.

[0034] Sound source location information indicates the location of the sound source relative to a certain terminal, that is, direction and position, such as relative direction and distance between the two.

[0035] With the development of internet technology, online conferencing has become an effective communication method. Currently, online conferencing typically requires participants to access the meeting through a terminal. Taking a typical online conferencing example, the microphone terminal, or the speaker's terminal, can acquire two microphone audio streams (left and right channels) using a microphone array or dedicated stereo acquisition equipment. These two microphone audio streams are then processed with techniques such as echo cancellation, noise reduction, and automatic gain control (AGC) audio pre-processing. Based on the advanced audio coding-low delay (AAC-LD) dual-channel protocol, the two processed audio streams are encoded and transmitted over the network to the playback terminal, or the listener's terminal. Upon receiving the bitstream, the playback terminal decodes the two pulse code modulation (PCM) format data based on the AAC-LD dual-channel protocol. The decoded data is then post-processed, and the post-processed audio streams are output to the left and right speakers for playback. During this process, the left and right channels are processed independently at both the pickup and playback ends to achieve a stereo playback effect.

[0036] However, the above process requires the audio pickup end to have dual-channel acquisition capabilities. Currently, due to differences in the capabilities of existing network protocols, web conferencing generally does not support multi-channel audio. Furthermore, the audio pickup end needs to use beamforming or dedicated hardware to pick up stereo sound, and the stereo pickup effect cannot be further enhanced by software, resulting in a generally poor stereo experience.

[0037] Therefore, this application provides an audio playback method that can improve the audio playback effect. By embedding the speaker's sound source location information from the pickup end into the acquired mono audio in real time, the playback end parses the sound source location information from the received audio and uses the sound source location information to process the mono audio into at least two audio channels for playback, thereby using mono speech to achieve stereo or higher quality playback effect.

[0038] The technical solution of this application will be described in detail below from multiple perspectives, including application scenarios, systems, methods and processes, software devices, and hardware devices.

[0039] Figure 1 This is a schematic diagram illustrating an application scenario provided in an embodiment of this application. For example... Figure 1As shown, this application scenario includes a first terminal 101, a second terminal 102, and a control unit / cloud 103. The first terminal 101 has audio playback functionality, while the second terminal 102 has audio acquisition and processing functions; for example, the second terminal 102 can acquire mono audio. Taking the first terminal 101 playing audio from the second terminal 102 as an example, the second terminal 102 is used to acquire audio from the environment where the terminal is located, embed the sound source location information of the local environment into the audio to obtain the target audio, and send the target audio to the first terminal. The sound source location information indicates the location of the sound source relative to the second terminal. After receiving the target audio from the second terminal 102, the first terminal 101 processes the target audio to obtain mono audio and sound source location information. Based on the sound source location information, it further processes the mono audio to obtain at least two audio channels, including left and right channel audio, and plays the processed at least two audio channels.

[0040] Optionally, the second terminal 102 collects audio from the local environment via an array of microphones (MICs), and determines the sound source location information based on the audio collected by the array of microphones. The sound source location information can be represented in the form of angles and distances. The sound source localization method based on the array of microphones includes a time delay estimation method and a beamforming method. The time delay estimation method includes calculating the sound source location by using the time difference between multiple microphones receiving the same sound source signal in the microphone array. Assuming the distance between the microphones is known, based on the speed of sound propagation, the distance difference from the sound source to each microphone can be calculated by measuring the time difference of sound received by different microphones, thus determining the sound source location. For example, an array of microphones (A, B, and C) is arranged horizontally on a plane. When the sound source emits sound, microphone A receives the sound 0.1 seconds earlier than microphone B. Given that the speed of sound in air is 340 meters per second, the distance difference from the sound source to A and B can be calculated to be 34 meters. Combining this with the time information of sound received by microphone C, the specific location of the sound source can be calculated using the principle of triangulation. The beamforming method includes: weighted summation of signals received by the array microphones to form a directional beam, which forms a main lobe in the direction of the sound source to obtain maximum output power, thereby determining the direction of the sound source. In this embodiment, the method used for sound source localization is not limited and can be determined based on the array microphone settings.

[0041] Optionally, the audio collected by the second terminal 102 includes ambient audio generated in the environment where the terminal is located. This ambient audio includes, but is not limited to, ambient noise audio and / or human voice audio. Ambient noise audio includes, for example, low-frequency noise generated by the operation of equipment such as air conditioners and fans in the local environment, and / or noise emitted by any sound-emitting device deployed in the local environment. Human voice audio refers to the voice of a speaker in the local environment. Therefore, optionally, the second terminal 102 can detect the presence of human voices in the local environment through human voice detection. Only when human voices are detected will the human voice acquisition function be activated to collect human voice audio from the local environment. This method saves energy; if no human voice is detected, the acquisition function is not executed, thus relatively reducing the energy consumption of the terminal side for acquisition. Optionally, the human voice audio can also be extracted from the original audio in the local environment. This method ensures the integrity of the audio acquisition and prevents information loss.

[0042] The second terminal described above is a device that integrates a microphone (including an array microphone) and conferencing functions. It can be understood as allowing users to access and participate in a meeting through a single terminal. In other scenarios, the second terminal and the microphone are two independently deployed devices. The microphone can collect audio from the environment in real-time or periodically, while the second terminal can serve as the implementation point for some conferencing functions. If the second terminal 102 integrates a microphone, the location information of the audio source can be directly obtained to reflect the positional relationship between the speaker and the second terminal 102. If the microphone and the second terminal 102 are separate, the location information of the audio source can be determined based on the positional relationship between the microphone and the second terminal 102, as well as the positional relationship between the speaker and the microphone.

[0043] Combination Figure 2 The principle explained is that after the second terminal 102 determines the sound source location information, an audio watermark can be generated based on the sound source location information, and the audio watermark can be embedded into the audio to obtain the target audio. The process of embedding the audio watermark into the audio can employ any of the following methods: time domain method, frequency domain method, or transform domain method.

[0044] The embedding process based on the time-domain method includes: generating a target time-domain signal based on the sound source location information, and superimposing the target time-domain signal and the audio time-domain signal to obtain the target audio time-domain signal.

[0045] The embedding process based on frequency domain methods includes: generating a target frequency domain signal based on the sound source's location information; converting the acquired audio from the time domain to the frequency domain, such as using methods like Discrete Fourier Transform (DFT) or Discrete Cosine Transform (DCT); and then selecting appropriate frequency components in the frequency domain to embed the target frequency domain signal. For example, adjusting the amplitude or phase of certain frequency components, and then converting the modified frequency domain signal back to the time domain to obtain the target audio with the embedded audio watermark.

[0046] The embedding process based on frequency domain methods includes: decomposing the audio into different sub-bands or scales using wavelet transform, and then embedding audio watermarks generated based on sound source location information into these sub-bands or scales. Because wavelet transform can better reflect the local characteristics of audio signals, it can achieve more effective watermark embedding and extraction in some cases.

[0047] This embedding method allows sound source location information to be transmitted to a remote terminal using the same transmission method without modifying the transmission protocol. The remote terminal only needs to decode the sound source location information after receiving the target audio. Furthermore, since the audio watermark has minimal impact on the audio itself, it will not affect the playback quality for playback devices that lack the aforementioned mono audio processing capabilities, greatly expanding the applicability of this audio playback method.

[0048] It should be noted that the embedding position of the aforementioned audio watermark can be located at certain specific locations, such as positions where the energy exceeds an energy threshold. This allows the remote terminal to determine the embedding position based on the energy after receiving the target audio, thereby extracting the audio watermark. The embedding position can also exist periodically, for example, embedding the audio watermark once every target duration. This avoids information loss due to a single detection failure, improving the stability of the function.

[0049] On the first terminal 101 side, after receiving the target audio, the first terminal separates the audio watermark from the target audio to obtain mono audio, and then analyzes the audio watermark to obtain the sound source location information. When separating the audio watermark, separation can be performed based on the aforementioned embedding position to obtain a complete audio watermark.

[0050] After obtaining the sound source location information, the first terminal 101 can reconstruct a stereo effect based on the sound source location information. For example, the first terminal 101 can adjust the mono audio to obtain an audio that differs from the original mono audio. The original mono audio and the adjusted mono audio are then used as the left and right channel audio, respectively, and played through the left and right channels of the first terminal 101 to achieve a stereo effect. In some embodiments, the above adjustment may include at least one of delaying the audio and adjusting the amplitude of the audio. For example, based on the sound source location information, the delay of the mono audio is adjusted to obtain the left channel audio; based on the sound source location information, the amplitude of the mono audio is adjusted to obtain the right channel audio. As another example, based on the sound source location information, the amplitude of the mono audio is adjusted to obtain the left channel audio; based on the sound source location information, the delay of the mono audio is adjusted to obtain the right channel audio. For example, based on the sound source location information, the amplitude and delay of the mono audio are adjusted in the first way to obtain the left channel audio; based on the sound source location information, the amplitude and delay of the mono audio are adjusted in the second way to obtain the right channel audio.

[0051] Because sound reaches each ear at different times, the time difference between audio recordings allows the human ear to perceive the location of the sound source. Therefore, by generating two audio streams with a time difference from a mono audio source, left and right channel audio can be obtained, thus simulating real human hearing. Furthermore, the amplitude difference between the left and right channels can create a stereo effect. Therefore, by adjusting the amplitude as described above, that is, by generating two audio streams with an amplitude difference from a mono audio source, left and right channel audio can be obtained, thus simulating real human hearing.

[0052] This application uses a conference scenario as an example for illustration. This scenario can be an audio / video conferencing system or a cloud-based conferencing system. Accordingly, the first and second terminals mentioned above are conference terminals. The system may include a conference management terminal, typically used by conference administrators. In some conference scenarios, the conference terminal may integrate a microphone, which has audio acquisition capabilities. The conference terminal can acquire audio through the microphone, embed an audio watermark indicating the speaker's location into the audio, and then send it to a remote conference terminal (conference service platform or conference management platform) through the conference management platform. Additionally, the conference terminal performs preprocessing such as noise reduction on the raw audio acquired by the microphone. In other conference scenarios, a microphone can be deployed separately. This microphone is a device with audio acquisition capabilities deployed within the conference scenario. The microphone can send the acquired audio to the conference terminal, which then embeds an audio watermark indicating the speaker's location into the audio before sending it to a remote conference terminal (conference service platform or conference management platform) through the conference management platform. In addition, the conference terminal also performs preprocessing such as noise reduction on the raw audio collected by the audio pickup device.

[0053] Optionally, the conference terminal can be a dedicated physical device or a software program with conferencing capabilities. This software program can run on various computing devices, such as mobile phones, tablets, and computers. In this case, the computing device running the software program can also be considered a conference terminal. The conference terminal joins the video conference through a conference service platform. Specifically, the conference terminal can acquire conference data (including audio data) from the conference service platform and send the locally acquired conference data to the platform, which then forwards it to other participating conference terminals. Conference terminals can connect wirelessly, allowing participants to join the video conference smoothly regardless of geographical location. In some cases, a participant may consist of only one user, such as joining the conference through conferencing software running on a personal mobile phone. In other cases, a participant may include multiple users, such as in a conference room scenario where multiple users in the conference room join the conference through a single conference terminal within the conference room. The conference service platform can be a multipoint control unit (MCU).

[0054] The following is an example of a conference system according to an embodiment of this application.

[0055] For example, Figure 3 This is a schematic diagram of the structure of a conference system provided in an embodiment of this application. For example... Figure 3As shown, the conference system includes a conference terminal and a conference service platform (MCU). Optionally, the conference system also includes a conference management terminal. The conference terminal is communicatively connected to the conference service platform. The conference service platform is communicatively connected to the conference management terminal.

[0056] The aforementioned conference service platform includes an audio / video transceiver module and a business message module. The audio / video transceiver module is used to send and receive network-encoded audio / video data with other devices (such as conference terminals). Optionally, the conference service platform also includes an audio encoding / decoding module, which encodes received audio / video data into its original format; encodes the original format audio / video data into its network format; and performs noise reduction processing, such as removing playback noise, on the original format audio / video data before encoding it into its network format. Optionally, the conference service platform also includes an audio watermarking module, which embeds an audio watermark into the original format audio data from the conference terminal and also parses the audio watermark from the original format audio data from the conference terminal. That is, the embedding and parsing of the audio watermark can also be performed by the conference service platform to reduce the load on the conference terminal. The conference terminal sends the sound source location information to the conference service platform via business messages, and the conference service platform performs the embedding process. This embodiment of the application does not limit this.

[0057] The conference terminal includes an audio / video transceiver module, an audio / video encoding / decoding module, an audio / video acquisition module, an audio / video playback module, and an audio watermarking module. The audio / video transceiver module is used to send and receive network-encoded audio / video data with other devices (such as a conference service platform). The audio / video encoding / decoding module encodes received audio / video data into its original format; encodes the original format audio / video data into its network format; and performs noise reduction processing, such as removing playback noise, on the original format audio / video data before encoding it into its network format. The audio / video acquisition module acquires audio / video data in its original format. The audio / video playback module plays the decoded original format audio / video data. The audio watermarking module embeds an audio watermark into the original format audio data from the audio / video encoding / decoding module and also extracts the audio watermark from the original format audio data from the audio / video encoding / decoding module. The original format is, for example, WAV format (a standard digital audio file format), and the network format is, for example, Real-Time Transport Protocol (RTP) format.

[0058] The conference management terminal is used to manage conferences, such as pausing a conference or remotely controlling a conference terminal in response to user actions. Optionally, the conference management terminal also includes an audio watermarking processing module. This module embeds audio watermarks into the raw audio data from the audio / video codec module and also extracts the audio watermarks from the raw audio data.

[0059] and Figure 4 This is a schematic diagram of another conference system provided in an embodiment of this application, such as... Figure 4 As shown, this is a cloud conferencing system, which may also include a conferencing service platform. The functions of the conferencing service platform in the two types of conferencing systems may be the same or different. For example, for cloud conferencing, its conferencing service platform can provide conferencing services to multiple terminals through cloud services. However, a cloud conferencing system does not include conferencing terminals; instead, it runs on the terminals in the form of a cloud conferencing application and provides conferencing services through cloud resources.

[0060] Accordingly, based on the aforementioned conference system, an exemplary embodiment provided by this application is introduced. Users bring smart terminals, video conferencing terminals, and soft terminals into the same conference via an in-place conference. Since the soft terminal does not support the AAC-LD dual-channel protocol, the conference is a mono conference. See also... Figure 5 The audio playback method includes:

[0061] 501. After the smart terminal picks up the speaker's voice signal through the array microphone, it performs audio preprocessing steps such as beamforming, echo cancellation, noise reduction, dereverberation, and automatic gain control to obtain the audio. Based on the audio, the smart terminal determines the speaker's sound source location information relative to the smart terminal. Here, the smart terminal is an example of the aforementioned second terminal, and the audio it can acquire is mono audio.

[0062] The audio preprocessing performed by the smart terminal is only one example of the embodiments of this application. Other audio preprocessing can also be performed, and the embodiments of this application do not limit this.

[0063] 502. Based on the sound source location information, the smart terminal embeds the sound source location information into the audio in real time in the form of an audio watermark.

[0064] 503. The intelligent terminal encodes the audio containing the location information of the sound source and sends it to the remote terminal in the form of a bitstream via the network.

[0065] 504. The remote terminal decodes the received bitstream to obtain the audio in its original format, for example, the original audio in the format of pulse code modulation (PCM) raw data.

[0066] 505. The remote terminal extracts the sound source location information from the original audio format.

[0067] 506. The remote terminal generates two left and right channel audios from the mono sound based on the sound source location information, so as to form data that can realize stereo sound.

[0068] 507. The remote terminal plays the corresponding audio through the left and right channels respectively. For example, a stereo effect can be achieved through left and right headphones and two external speakers.

[0069] In the above embodiments, the audio encoding is performed at the pickup end, rather than by the conference service platform in the conference system. Similarly, the audio decoding is also performed at the playback end, and does not require the conference service platform in the conference system to perform it. This ensures that the conference service platform does not participate in the encoding and decoding process and does not increase the processing pressure on the conference service platform.

[0070] In the above embodiments, the smart terminal is also the second terminal involved in the embodiments of this application, and the remote terminal is also the first terminal involved in the embodiments of this application. For a terminal, it can have both audio acquisition capability and audio playback capability. That is, as long as the sound pickup end can obtain the sound source location information and the playback end can have at least two-channel playback capability, a stereo effect can be achieved on the playback end, thereby expanding the mono conference into a conference with stereo effect and greatly enhancing the realism of the conference.

[0071] Accordingly, based on the aforementioned conference system, an exemplary embodiment provided by this application is introduced. When a user pulls multiple terminals, such as smart terminals, video conferencing terminals, and soft terminals, into the same cloud conference, the conference is a mono conference, and all of the aforementioned terminals have the cloud conferencing application installed. See also... Figure 6 The audio playback method includes:

[0072] 601. After the smart terminal picks up the speaker's voice signal through the array MIC, it obtains the audio through audio preprocessing processes such as beamforming, echo cancellation, noise reduction, dereverberation, and automatic gain control. Based on the audio, the smart terminal determines the speaker's sound source location information relative to the smart terminal.

[0073] 602. Based on the sound source location information, the smart terminal embeds the sound source location information into the audio in real time in the form of an audio watermark.

[0074] 603. The smart terminal transmits audio with embedded sound source location information to the cloud conferencing application via the audio uplink path as microphone input.

[0075] 604. The cloud conferencing application's conferencing service platform encodes the received audio and sends it as a stream over the network to the remote terminal.

[0076] 605. The cloud conferencing application on the remote terminal decodes the received bitstream to obtain the original audio format, which is then sent to the terminal's underlying application for playback via the audio downlink channel. For example, the original audio format is raw PCM data.

[0077] 606. The terminal-level application on the remote terminal parses the sound source location information from the original audio format.

[0078] 607. The remote terminal generates two left and right channel audios from a mono sound based on the sound source location information, so as to form data that can realize stereo sound.

[0079] 608. The remote terminal plays the corresponding audio through the left and right channels respectively. For example, a stereo effect can be achieved through left and right headphones and two external speakers.

[0080] In the above embodiments, the audio encoding is performed in the cloud, rather than by the terminal itself. Similarly, the audio decoding is performed in the cloud conferencing application on the playback end, and the terminal does not need to have hardware encoding and decoding capabilities, thereby reducing the hardware requirements of the conferencing terminal.

[0081] In the above embodiments, the smart terminal is also the second terminal involved in the embodiments of this application, and the remote terminal is also the first terminal involved in the embodiments of this application. For a terminal, it can have both audio acquisition capability and audio playback capability. That is, as long as the sound pickup end can obtain the sound source location information and the playback end can have at least two-channel playback capability, a stereo effect can be achieved on the playback end, thereby expanding the mono conference into a conference with stereo effect and greatly enhancing the realism of the conference.

[0082] In the above embodiments, the addition of audio watermarks is taken as an example of being performed on the terminal side. Alternatively, after the second terminal obtains the sound source location information, it can submit the sound source location information and the collected audio together to the cloud for further processing. The cloud can embed the audio watermark, and then perform corresponding parsing and other processing at the remote terminal. This can also reduce the processing pressure on the sound pickup side to a certain extent.

[0083] This application also provides an audio playback device. For example... Figure 7 As shown, Figure 7This is a schematic diagram of the structure of an audio playback device provided in an embodiment of this application. Figure 7 As shown, the device includes the following structure.

[0084] The receiving module 701 is used to receive target audio from the second terminal, wherein the target audio carries sound source location information, and the sound source location information indicates the location of the sound source of the audio relative to the second terminal.

[0085] Information processing module 702 is used to process the target audio to obtain mono audio and the sound source location information;

[0086] The audio processing module 703 is used to process the mono audio based on the sound source location information to obtain at least two audio channels;

[0087] The playback module 704 is used to play the at least two audio channels.

[0088] In some embodiments, the information processing module 702 is configured to separate an audio watermark from the target audio to obtain the mono audio, wherein the audio watermark indicates the sound source location information; and to parse the audio watermark to obtain the sound source location information.

[0089] In some embodiments, the audio processing module 703 is configured to adjust the delay of the mono audio based on the sound source orientation information to obtain the left channel audio; and to adjust the amplitude of the mono audio based on the sound source orientation information to obtain the right channel audio.

[0090] In some embodiments, the audio processing module 703 is configured to adjust the amplitude of the mono audio based on the sound source orientation information to obtain the left channel audio; and to adjust the delay of the mono audio based on the sound source orientation information to obtain the right channel audio.

[0091] In some embodiments, the audio processing module 703 is configured to perform a first adjustment on the amplitude and delay of the mono audio based on the sound source location information to obtain left channel audio; and to perform a second adjustment on the amplitude and delay of the mono audio based on the sound source location information to obtain right channel audio.

[0092] In some embodiments, the first terminal and the second terminal are conference terminals.

[0093] It should be understood that the device provided in the above embodiments is only illustrated by the division of the above functional units when playing audio. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio playback device and the audio playback method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0094] This application also provides an audio processing apparatus. For example... Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application. Figure 8 As shown, the device includes...

[0095] Acquisition module 801 is used to acquire audio from the local environment;

[0096] The embedding module 802 is used to embed the sound source location information of the audio in the local environment into the audio to obtain the target audio, wherein the sound source location information indicates the location of the sound source of the audio relative to the first terminal;

[0097] The sending module 803 is used to send the target audio to the first terminal.

[0098] In some embodiments, the embedding module 802 is used to generate an audio watermark based on the sound source location information of the audio in the environment, and embed the audio watermark into the audio to obtain the target audio.

[0099] In some embodiments, the acquisition module 801 is used to acquire audio in the local environment through the array MIC configured in the second terminal;

[0100] The device further includes:

[0101] The sound source localization module is used to determine the sound source location information of the audio based on the audio collected by the array MIC.

[0102] In some embodiments, the acquisition module 801 is configured to acquire human voice audio in the local environment in response to detecting the presence of human voice in the local environment; or, acquire raw audio in the local environment and extract human voice audio from the raw audio.

[0103] In some embodiments, the sending module 803 is configured to encode the target audio and send the encoded target audio to the first terminal; or, send the target audio to a target service, whereby the target service encodes the target audio and sends it to the first terminal.

[0104] It should be understood that the device provided in the above embodiments is only illustrated by the division of the above functional units when performing audio processing. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device and the audio processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0105] Figure 9 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. See also... Figure 9 Terminal 900 may include a processor 901, a memory 902, and an audio circuit 903. It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on terminal 900. In other embodiments of this application, terminal 900 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0106] Processor 901 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. Processor 901 may also include memory for storing instructions and data. In some embodiments, the memory in processor 901 is a cache memory.

[0107] Memory 902 can be used to store computer executable program code, which includes instructions. Memory 902 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, and universal flash storage (UFS). Processor 901 executes various functional applications and data processing of terminal 900 by running instructions stored in memory 902 and / or instructions stored in memory disposed in the processor.

[0108] The audio circuit 903 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 901 for processing, or to transmit them for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from or received by the processor 901 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 903 may also include a headphone jack.

[0109] It is understood that the above structure is only one hardware possibility for implementing the terminal provided in the embodiments of this application, and the embodiments of this application do not specifically limit it.

[0110] This application provides a computer-readable storage medium for storing at least one piece of program code, which is used to implement the above-described audio playback method or audio processing method.

[0111] This application provides a computer program product for implementing the above-described audio playback method or audio processing method.

[0112] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have substantially the same function and purpose. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another. For example, without departing from the various examples described, a first key can be referred to as a second key, and similarly, a second key can be referred to as a first key. Both the first key and the second key can be keys, and in some cases, they can be separate and different keys.

[0113] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple keys means two or more keys.

[0114] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0115] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of program structure information. This program structure information includes one or more program instructions. When these program instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of this application are generated, in whole or in part.

[0116] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0117] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An audio playback method, characterized in that, Applied to a first terminal, the method includes: Receive target audio from a second terminal, the target audio carrying sound source location information, the sound source location information indicating the location of the audio source relative to the second terminal; The target audio is processed to obtain mono audio and the sound source location information; Based on the sound source location information, the mono audio is processed to obtain at least two audio channels; Play at least two audio streams.

2. The method according to claim 1, characterized in that, The process of processing the target audio to obtain mono audio and the sound source location information includes: The audio watermark is extracted from the target audio to obtain the mono audio, and the audio watermark indicates the location information of the sound source; The audio watermark is analyzed to obtain the sound source location information.

3. The method according to claim 1, characterized in that, The process of processing the mono audio based on the sound source location information to obtain at least two audio streams includes: Based on the sound source location information, the delay of the mono audio is adjusted to obtain the left channel audio; Based on the sound source location information, the amplitude of the mono audio is adjusted to obtain the right channel audio.

4. The method according to claim 1, characterized in that, The process of processing the mono audio based on the sound source location information to obtain at least two audio streams includes: Based on the sound source location information, the amplitude of the mono audio is adjusted to obtain the left channel audio. Based on the sound source location information, the delay of the mono audio is adjusted to obtain the right channel audio.

5. The method according to claim 1, characterized in that, The process of processing the mono audio based on the sound source location information to obtain at least two audio streams includes: Based on the sound source location information, the amplitude and delay of the mono audio are adjusted first to obtain the left channel audio; Based on the sound source location information, the amplitude and delay of the mono audio are adjusted a second time to obtain the right channel audio.

6. The method according to claim 1, characterized in that, The first terminal and the second terminal are conference terminals.

7. An audio processing method, characterized in that, Applied to a second terminal, the method includes: Capture audio from the local environment; The source location information of the audio in the local environment is embedded into the audio to obtain the target audio, wherein the source location information indicates the location of the audio source relative to the first terminal; The target audio is sent to the first terminal.

8. The method according to claim 7, characterized in that, The step of embedding the sound source location information of the audio in the environment into the audio to obtain the target audio includes: Based on the sound source location information of the audio in the environment, an audio watermark is generated, and the audio watermark is embedded in the audio to obtain the target audio.

9. The method according to claim 7, characterized in that, The audio collected from the local environment includes: The second terminal uses an array microphone to collect audio from the local environment. The method further includes: Based on the audio collected by the array MIC, the sound source location information of the audio is determined.

10. The method according to claim 7, characterized in that, The audio collected from the local environment includes: In response to detecting the presence of human voices in the local environment, the audio of the human voices in the local environment is acquired; or, Collect raw audio from the local environment and extract human voice audio from the raw audio.

11. The method according to claim 7, characterized in that, Sending the target audio to the first terminal includes: The target audio is encoded, and the encoded target audio is sent to the first terminal; or, The target audio is sent to the target service, which encodes the target audio and sends it to the first terminal.

12. An audio playback device, characterized in that, The device includes: A receiving module is used to receive target audio from a second terminal, wherein the target audio carries sound source location information, and the sound source location information indicates the location of the sound source of the audio relative to the second terminal. The information processing module is used to process the target audio to obtain mono audio and the sound source location information; An audio processing module is used to process the mono audio based on the sound source location information to obtain at least two audio channels; A playback module is used to play the at least two audio streams.

13. An audio processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire audio from the local environment; An embedding module is used to embed the sound source location information of the audio in the local environment into the audio to obtain target audio, wherein the sound source location information indicates the location of the audio source relative to the first terminal; The sending module is used to send the target audio to the first terminal.

14. A terminal, characterized in that, The terminal includes: Processor and memory; The memory is used to store computer programs, the computer programs including program instructions; The processor is configured to invoke the computer program to implement the method as described in any one of claims 1 to 11.

15. A conference system, characterized in that, The conference system includes a first terminal and a second terminal. The second terminal is used to collect audio in the local environment; the sound source location information of the audio in the local environment is embedded into the audio to obtain the target audio, wherein the sound source location information indicates the location of the audio source relative to the first terminal; The target audio is sent to the first terminal; The first terminal is used to receive target audio from the second terminal. The target audio carries sound source location information, which indicates the location of the sound source of the audio relative to the second terminal. The target audio is processed to obtain a mono audio and the sound source location information; based on the sound source location information, the mono audio is processed to obtain at least two audio channels; the at least two audio channels are played.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which is used to implement the audio playback method of any one of claims 1 to 6 or the audio processing method of any one of claims 7 to 11.

17. A computer program product, characterized in that, The computer program product is used to implement the audio playback method of any one of claims 1 to 6 or the audio processing method of any one of claims 7 to 11.