Audio processing method and apparatus, and electronic device and computer program product

By performing spatial audio processing on the audio data in multi-person live streaming, and processing and mixing the audio according to the speaker's position on the live stream screen, the problem of mixed sound in multi-person live streaming is solved, and the audio experience and recognizability are improved.

WO2026156841A1PCT designated stage Publication Date: 2026-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-01-26
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

In multi-person live streaming, the voices of various hosts are mixed together without any distinction of importance or direction, resulting in a poor audio experience for users due to the lack of stereo effect.

Method used

Spatial audio processing is performed on the audio data of different speakers to make their positions on the terminal correspond to their positions in the live broadcast. A stereo audio stream is generated using spatial audio processing algorithms such as HRTF, and then mixed.

Benefits of technology

It enhances the audience's sense of presence and the speaker's recognizability, improving the audio experience during real-time conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075318_30072026_PF_FP_ABST
    Figure CN2025075318_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to and provides an audio processing method and apparatus, and an electronic device and a computer program product. The method comprises: receiving a plurality of audio streams for a real-time session, wherein each audio stream is associated with a corresponding speaker among a plurality of speakers. The method further comprises: performing spatial audio processing on the plurality of audio streams on the basis of a layout of the plurality of speakers presented on a terminal. The method further comprises: mixing the plurality of audio streams, which have undergone spatial audio processing, wherein the mixed audio streams are transmitted to the terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing methods, apparatuses, electronic devices and computer program products Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to an audio processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] Live streaming is a media format that uses internet technology to transmit content in real time through video, audio, and other means. Hosts use specialized live streaming platforms or software to showcase their talents and lives, share professional knowledge, and other information to a wide audience online. Viewers can log in to the live streaming platform via mobile phones, computers, and other devices to watch the live stream and interact with the host.

[0003] Multi-person live streaming is an advanced form of online live streaming. It allows multiple hosts to interact online simultaneously. In multi-person live streaming, participants can communicate with each other, collaborate on performances, and so on. In multi-person live streaming scenarios, the audio data of the hosts is usually mixed before being sent to the audience for playback. Summary of the Invention

[0004] Embodiments of this disclosure provide an audio processing scheme for real-time conversations involving multiple speakers.

[0005] According to a first aspect of this disclosure, an audio processing method is provided. The method includes: receiving a plurality of audio streams for a real-time session, each audio stream being associated with a corresponding speaker among a plurality of speakers; performing spatial audio processing on the plurality of audio streams based on a layout of the plurality of speakers presented on a terminal; and mixing the spatially audio-processed plurality of audio streams, wherein the mixed audio streams are transmitted to the terminal.

[0006] According to a second aspect of this disclosure, an audio processing apparatus is provided. The apparatus includes: an audio stream receiving unit configured to receive a plurality of audio streams for a real-time session, each audio stream being associated with a corresponding speaker among a plurality of speakers; a spatial audio processing unit configured to perform spatial audio processing on the plurality of audio streams based on a layout of the plurality of speakers presented on a terminal; and a mixing unit configured to mix the spatially audio-processed plurality of audio streams, wherein the mixed audio streams are transmitted to the terminal.

[0007] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform an audio processing method, the method comprising: receiving a plurality of audio streams for a real-time session, each audio stream being associated with a corresponding speaker among a plurality of speakers; performing spatial audio processing on the plurality of audio streams based on a layout of the plurality of speakers presented on a terminal; and mixing the spatially audio-processed plurality of audio streams, wherein the mixed audio streams are transmitted to the terminal.

[0008] According to a fourth aspect of this disclosure, a non-transient computer storage medium is provided, including machine-executable instructions that, when executed by a device, cause the device to perform the method as described in the first aspect of this disclosure.

[0009] According to a fifth aspect of this disclosure, a computer program product is provided, including machine-executable instructions that, when executed by a device, cause the device to perform the method as described in the first aspect of this disclosure.

[0010] It should be understood that the summary section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other objects, features, and advantages of embodiments of the present disclosure will become more readily understood from the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the present disclosure will be described by way of example and non-limitation, wherein:

[0012] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0013] Figure 2 shows a schematic flowchart of an audio processing method according to an embodiment of the present disclosure;

[0014] Figure 3 shows a schematic block diagram of a server for implementing spatial audio processing according to an embodiment of the present disclosure;

[0015] Figure 4 shows a schematic diagram of the spatial audio coordinate system and speaker layout according to an embodiment of the present disclosure;

[0016] Figures 5A-5C show schematic diagrams of layouts based on the number of speakers in a live broadcast room according to embodiments of the present disclosure;

[0017] Figure 6 shows a schematic flowchart of the dynamically updated spatial audio processing procedure according to an embodiment of the present disclosure.

[0018] Figure 7 shows a block diagram of a data protection device according to an embodiment of the present disclosure; and

[0019] Figure 8 shows a block diagram of an electronic device according to an embodiment of the present disclosure.

[0020] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation

[0021] This disclosure will now be discussed with reference to several example implementations. It should be understood that these implementations are discussed only to enable those skilled in the art to better understand and thus implement this disclosure, and not to imply any limitation on the scope of this disclosure.

[0022] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0027] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.

[0028] In the current live streaming scenario, whether using client-side or server-side merging, the standard procedure involves first performing a simple mixing of the audio data from each streamer before transmitting it to the viewer. This method has the following problems: the voices of each streamer are mixed together without distinction of primary or secondary voices or direction, lacking stereo and spatial effects, resulting in a poor audio experience for the user.

[0029] In view of this, embodiments of the present disclosure provide an audio processing solution that performs spatial audio processing on the audio data of speakers (e.g., broadcasters or other participants in the conversation) in real-time conversation scenarios with multiple speakers, so that the direction of the speaker's voice heard by the audience corresponds to the position of the speaker on the terminal, thereby enhancing the audience's sense of presence and the speaker's recognizability, and thus improving the audio experience during real-time conversations.

[0030] In general, to make the sound heard by viewers of a live stream more three-dimensional and directional, thereby enhancing the user experience and increasing their willingness to watch, the viewing perspective can be considered. For example, if the viewer is directly in front of the live stream playback device (such as a mobile phone, tablet, or other terminal), different spatial audio processing can be applied to the audio data of different speakers based on their position in the live stream. For instance, for a speaker on the left side of the live stream, spatial audio processing can be used to make their audio data appear to come from the left front; for a speaker on the right side, spatial audio processing can be used to make their audio data appear to come from the right front. After processing, the spatially processed audio data is mixed and sent to the viewer's (user's) device.

[0031] The embodiments of this disclosure will be described in detail below with reference to Figures 1 to 8. It should be noted that the embodiments of this disclosure are mainly described below using a scenario of multiple people interacting in a live broadcast room as an example. However, the embodiments of this disclosure are also applicable to other real-time conversation scenarios, such as video conferences with multiple participants, and this disclosure does not impose any limitations in this regard.

[0032] Figure 1 illustrates a schematic diagram of an example environment 100 that can be implemented according to embodiments of the present disclosure. Environment 100 can be a multi-person live streaming scenario where multiple streamers interact or perform in the same live streaming room, and users can simultaneously watch the live streams of multiple streamers on their devices. Each streamer can occupy a portion of the screen on the user's device, forming a corresponding streamer layout. For example, if a user links or jumps from a streamer's live streaming room to the current multi-person live streaming scenario, the user can set that streamer in their preferred position, such as the left side of the screen. The positions of other streamers can be determined according to system configuration or user settings. In some implementations, the user can manually change the streamer layout.

[0033] Referring to Figure 1, audio streams from multiple speakers (e.g., broadcasters) 101 can be sent to server 110. In some embodiments, server 110 may include a real-time communication (RTC) server. As shown, server 110 includes a spatial audio processing component 120 for performing spatial audio processing on the received audio streams. In some embodiments, the spatial audio processing component 120 can perform spatial audio processing on the audio stream of speaker 101 based on the speaker's position information in the live broadcast frame, the direction the speaker is speaking, and the direction the audience is listening to the sound. In some embodiments, it can be assumed that the speaker is always facing forward, and the audience is always facing forward to listen to the sound. The spatial audio processing component 120 can convert the position information into spatial audio parameters and input them into an audio localization algorithm (e.g., Head-Related Transform Function (HRTF) or other algorithms) to generate a spatially processed audio stream. The speaker's position information may depend on their layout on the live broadcast device. For example, broadcaster A may be located on the left side of the live broadcast frame on one device, and on the right side or in another position on another frame, requiring different processing of broadcaster A's audio stream accordingly.

[0034] Server 110 also includes a mixer 130. The mixer 130 receives multiple audio streams that have undergone spatial audio processing and combines them to obtain a mixed audio stream. These audio streams have spatial directional characteristics, thus the mixed audio stream has a stereo effect. The mixer 130 can combine audio streams to form a mixed audio stream based on the layout of speakers in the live broadcast. For example, in a scenario with two broadcasters, the live broadcast might show broadcaster A on the left and broadcaster B on the right, or vice versa. The mixer 130 can merge the spatially processed audio stream of broadcaster A on the left and broadcaster B on the right, or vice versa. In other words, the mixer 130 mixes the processed audio streams to ensure that the mixed audio stream matches the live broadcast.

[0035] Server 110 can send the mixed audio stream to network 140, which then distributes the mixed audio stream to terminal 150. Network 140 can be, for example, a content delivery network (CDN). Thus, terminal 150 can receive an audio stream that matches the speaker layout in its live stream.

[0036] It should be noted that the environment 100 shown in Figure 1 is merely exemplary, and the embodiments of this disclosure can also be implemented in different environments, such as video conferencing. Environment 100 may also include more or fewer components; for example, server 110 may also include audio stream encoding / decoding components, etc.

[0037] Figure 2 shows a schematic flowchart of an audio processing method 200 according to an embodiment of the present disclosure. In some embodiments, method 200 may be implemented by, for example, the server 110 shown in Figure 1. It should be understood that method 200 may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0038] In box 210, method 200 includes receiving multiple audio streams for a real-time session, each audio stream being associated with a corresponding speaker among multiple speakers. For example, in a multi-person live streaming scenario, server 110 may receive audio streams from multiple hosts, each audio stream including audio data of the corresponding host.

[0039] In box 220, method 200 includes performing spatial audio processing on multiple audio streams based on the layout of multiple speakers presented on the terminal. The position of each speaker in the live stream can be used to determine how to process that speaker's audio stream. In some embodiments, server 110 can determine the spatial audio parameters of each speaker based on the layout of the multiple speakers, wherein the layout can be determined based on the number of speakers currently participating in a multi-person live stream. Then, server 110 performs spatial audio processing on the associated audio streams based on the spatial audio parameters. For example, if the audio localization algorithm requires input of the audience coordinates, sound source coordinates, the audience's listening direction, the sound source's emission direction, etc., in a spatial coordinate system, server 110 needs to determine the above information and provide it to the audio localization algorithm. This is described in detail below with reference to Figures 4 and 5A-5C.

[0040] In box 230, method 200 includes: mixing multiple audio streams that have undergone spatial audio processing, wherein the mixed audio streams are transmitted to a terminal.

[0041] In some embodiments, server 110 can generate multiple spatially processed audio streams for an audio stream, each processed audio stream corresponding to a different position of a speaker in the live broadcast layout. Server 110 can select one of the multiple processed audio streams for mixing based on the layout. The mixed audio stream can then be distributed to the terminal 150 that applies the layout.

[0042] Figure 3 shows a schematic block diagram of a server 310 for implementing spatial audio processing according to an embodiment of the present disclosure. Server 310 may be an exemplary implementation of server 110 of Figure 1, such as an RTC merging server.

[0043] As shown in the figure, server 310 includes an audio decoding component 302, which is used to perform decoding operations on the received audio stream to obtain decoded audio data, such as audio data in Pulse Code Adjustment (PCM) format.

[0044] Server 310 also includes a spatial audio processing component 320, which includes an audio localization algorithm 322 and stores layout and spatial audio parameters 325. The audio localization algorithm 322 can be, for example, an HRTF algorithm, a Vector Basis Amplitude Shift (VBAP) algorithm, a Sound Field Synthesis (SFS) algorithm, etc. For each decoded audio stream, the audio localization algorithm 320 processes the audio stream based on the corresponding spatial audio parameters.

[0045] The processed audio stream is provided to mixer 330. Mixer 330 combines the spatially information-containing audio streams based on the layout to obtain a mixed audio stream, which is then provided to the high-level audio codec AAC encoder 340. Specifically, the speaker's audio stream can generate multiple processed audio streams after spatial audio processing, each corresponding to a different position of the speaker in the layout. Mixer 330 can select one of the multiple processed audio streams for mixing based on the layout.

[0046] The audio stream encoded by the AAC encoder 340 is sent to the content delivery network (CDN) for distribution to the terminal.

[0047] For ease of explanation, the HRTF algorithm will be used as an example to illustrate the details of determining spatial audio parameters. To determine the spatial audio parameters of multiple speakers, server 310 can determine the virtual spatial position of the speaker relative to the terminal user (i.e., the audience) based on the speaker's position in the layout. The virtual spatial position can include coordinates in the spatial audio coordinate system.

[0048] Figure 4 illustrates a schematic diagram of the spatial audio coordinate system and speaker layout according to an embodiment of the present disclosure. As shown in Figure 4, the user views the screen from the front, and several possible positions of the speaker are arranged on the screen: position 1 corresponds to the user's left side, position 2 corresponds to the user's right side, position 3 corresponds to the user's upper left, position 4 corresponds to the user's lower left, position 5 corresponds to the user's upper right, and position 6 corresponds to the user's lower right. It is understood that the possible positions on the screen may differ from this, and more or fewer possible positions may be set. In some embodiments, the speaker layout is determined from these possible positions based on the number of speakers in the current live broadcast room.

[0049] Figures 5A-5C illustrate schematic diagrams of layouts based on the number of speakers in a live streaming room according to embodiments of the present disclosure. Figure 5A shows an exemplary layout for a two-speaker scenario, where speaker A is located on the left (i.e., position 1) and speaker B is located on the right (i.e., position 2). Figure 5B shows an exemplary layout for a three-speaker scenario, where speaker A is located on the left (i.e., position 1), speaker B is located in the upper right (i.e., position 5), and speaker C is located in the lower right (i.e., position 6). Figure 5C shows an exemplary layout for a four-speaker scenario, where speaker A is located in the upper left (i.e., position 3), speaker B is located in the lower left (i.e., position 4), speaker C is located in the upper right (i.e., position 5), and speaker D is located in the lower right (i.e., position 6).

[0050] In some embodiments, if the speaker is the main speaker on the terminal, the speaker's position can be determined as the left, upper left, or lower left (i.e., position 1, 3, or 4). The main speaker refers to the user who enters the multi-person live chat scene from the speaker's live stream room.

[0051] The following describes an exemplary method for calculating the spatial coordinates of positions 1 to 6. In the spatial coordinate system of Figure 4, the horizontal direction within the screen is the X-axis, the vertical direction is the Y-axis, and the direction perpendicular to the screen is the Z-axis (the user's direction is positive). Assuming the user's coordinates are (0, 0, L), where L is greater than 0, then the coordinates of position 1 can be set to (0, -A, 0), and the coordinates of position 2 can be set to (0, A, 0), where A is greater than 0. Furthermore, it can be assumed that the distances from positions 3, 4, 5, and 6 to the X-axis are equal to A.

[0052] It can be seen that L and A together determine the size of the angle θ (theta) between the speaker's voice from different directions (relative to the Z-axis) heard by the user, specifically theta = atan(A / L), where atan is the inverse function of the tangent function tan. In some embodiments, the corresponding values ​​of A and L can be derived based on the target angle. For example, assuming the target angle theta is 20 degrees, then A / L = tan 20° ≈ 0.364. Therefore, as long as the ratio of A to L remains constant, the direction of the sound remains unchanged. Let L = 1, and according to the above settings, with the angle between the speaker and the user being theta, the spatial coordinates from position 1 to position 6 can be determined as follows:

[0053] In some embodiments, to determine the virtual spatial position of a speaker relative to the user of the terminal, the server 310 can determine the virtual spatial position of the speaker relative to the user based on the speaker's position in the layout and the user's virtual listening angle (i.e., theta). Specifically, the effect of spatial audio can be adjusted by setting different angles theta. If a strong spatial audio angle is required, a larger theta can be set; if only a slight distinction in spatial position is desired, a smaller theta can be set. In some embodiments, to reduce computational complexity, vertical spatial information can be disregarded, and only horizontal spatial information can be considered. That is, the same spatial audio parameters about the X-axis can be applied to the speakers; for example, all left-hand speakers can apply the parameters of position 1, and all right-hand speakers can apply the parameters of position 2.

[0054] In some embodiments, determining the speaker's spatial audio parameters may include determining the user's orientation relative to the terminal. The user's orientation relative to the terminal may be the direction facing the screen, and the corresponding vector information may be provided to the audio localization algorithm 322.

[0055] Figure 6 shows a schematic flowchart of a dynamically updated spatial audio processing procedure 600 according to an embodiment of the present disclosure. Procedure 600 can be implemented by, for example, the server 110 shown in Figure 1. It should be understood that procedure 600 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0056] In box 610, you confirm the start of live streaming and multi-person live chat, which means that spatial audio processing is required for the audio streams of multiple speakers.

[0057] In frame 620, spatial audio parameters are calculated based on the live stream layout and set to the RTC merge server. Then, based on the spatial audio parameters, spatial audio processing and mixing are performed on the speaker's audio stream.

[0058] In box 630, it is determined whether anyone has entered or left the live stream. If someone enters or leaves the live stream, it means that the layout will change, so process 600 proceeds to box 640, updates the layout, recalculates the spatial audio parameters, and sets them to the RTC merge server.

[0059] At frame 650, determine whether the multi-person live stream has ended. If so, proceed to frame 660 to end it. Otherwise, return to frame 630 and continue to determine whether anyone has entered or left the live stream.

[0060] Exemplary embodiments of the present disclosure have been described above with reference to Figures 1 to 6. Compared to existing technologies, embodiments of the present disclosure can perform spatial audio processing on multiple audio streams in a real-time session, realizing the correspondence between the direction of the speaker's voice and the speaker's position displayed on the screen, thereby enhancing the audience's sense of presence, the speaker's recognizability, and the audio experience.

[0061] Figure 7 shows a schematic block diagram of an audio processing apparatus 700 according to an embodiment of the present disclosure. The apparatus 700 can be used to implement the method 200 described with reference to Figure 2. As shown in Figure 7, the apparatus 700 includes an audio stream receiving unit 710, a spatial audio processing unit 720, and a mixing unit 730. The audio stream receiving unit 710 is configured to receive multiple audio streams for a real-time session, each audio stream being associated with a corresponding speaker among a plurality of speakers. The spatial audio processing unit 720 is configured to perform spatial audio processing on the multiple audio streams based on a layout of the multiple speakers presented on a terminal. The mixing unit 730 is configured to mix the spatially audio-processed multiple audio streams, wherein the mixed audio streams are transmitted to the terminal.

[0062] In some embodiments, performing spatial audio processing on multiple audio streams includes: determining a layout of multiple speakers presented on a terminal based on the number of speakers; determining spatial audio parameters for each of the multiple speakers based on the layout; and performing spatial audio processing on the associated audio streams based on the spatial audio parameters.

[0063] In some embodiments, determining the spatial audio parameters of each of the multiple speakers based on the layout may include: determining the virtual spatial position of the speaker relative to the user of the terminal based on the position of the speaker in the layout among the multiple speakers.

[0064] In some embodiments, determining the spatial audio parameters of each of the multiple speakers based on the layout may further include determining the user's orientation relative to the terminal.

[0065] In some embodiments, determining the virtual spatial position of a speaker among a plurality of speakers relative to the user of the terminal may include: determining the virtual spatial position of the speaker relative to the user based on the speaker's position in the layout and the user's virtual listening angle.

[0066] In some embodiments, the speaker's position in the layout includes one of the following: top left, left side, bottom left, top right, right side, and bottom right.

[0067] In some embodiments, in response to the speaker being the main speaker of the terminal, the speaker's position is determined to be on the left, upper left, or lower left.

[0068] In some embodiments, the audio streams of a speaker among a plurality of speakers are spatially processed to produce a plurality of processed audio streams, each of which corresponds to a different position of the speaker in the layout. Mixing the plurality of spatially processed audio streams may include: selecting one of the plurality of processed audio streams for mixing based on the layout.

[0069] In some embodiments, the layout may be updated in response to determining that a new speaker has joined or a speaker has left the live session; and spatial audio processing may be re-executed on the audio stream of the live room based on the updated layout.

[0070] It should be noted that further actions or steps shown in Figures 1 to 6 can be implemented using the device 700 shown in Figure 7. For example, device 700 may include more modules or units to implement the actions or steps described above, or some of the units or modules shown in Figure 7 may be further configured to implement the actions or steps described above. This will not be repeated here.

[0071] Figure 8 shows a schematic block diagram of an example device 800 that can be used to implement embodiments of the present disclosure. As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 806 into random access memory (RAM) 803. Various programs and data required for the operation of device 800 may also be stored in RAM 803. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0072] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0073] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).

[0074] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0075] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0076] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media within the respective computing / processing device.

[0077] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​and conventional procedural programming languages. The computer-readable program instructions may execute entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0078] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0079] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0081] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An audio processing method, comprising: Receive multiple audio streams for a real-time session, each audio stream being associated with a corresponding speaker among multiple speakers; Based on the layout of the multiple speakers presented on the terminal, spatial audio processing is performed on the multiple audio streams; as well as The plurality of audio streams that have undergone spatial audio processing are mixed, and the mixed audio streams are transmitted to the terminal.

2. The method according to claim 1, wherein, Performing spatial audio processing on the plurality of audio streams includes: Based on the number of the plurality of speakers, determine the layout of the plurality of speakers presented on the terminal; Based on the aforementioned layout, the spatial audio parameters of each of the plurality of speakers are determined; and Spatial audio processing is performed on the associated audio stream based on the spatial audio parameters.

3. The method according to claim 2, wherein, Based on the aforementioned layout, determining the spatial audio parameters of each of the plurality of speakers includes: Based on the position of the speaker among the plurality of speakers in the layout, the virtual spatial position of the speaker relative to the user of the terminal is determined.

4. The method according to claim 3, wherein, Based on the aforementioned layout, determining the spatial audio parameters of each of the multiple speakers further includes: Determine the user's orientation relative to the terminal.

5. The method according to claim 3, wherein, Determining the virtual spatial location of one of the plurality of speakers relative to the user of the terminal includes: Based on the speaker's position in the layout and the user's virtual listening angle, the speaker's virtual spatial position relative to the user is determined.

6. The method according to any one of claims 3 to 5, wherein the speaker's position in the layout includes one of the following: upper left, left side, lower left, upper right, right side, and lower right.

7. The method according to claim 6, wherein, In response to the speaker being the main speaker of the terminal, the speaker's position is determined to be on the left, upper left, or lower left.

8. The method according to any one of claims 1 to 7, wherein, The audio streams of the speakers from among the multiple speakers are processed by spatial audio processing to generate multiple processed audio streams, each of which corresponds to a different position of the speaker in the layout. The mixing of the plurality of spatially processed audio streams includes: selecting one processed audio stream from the plurality of processed audio streams for mixing based on the layout.

9. The method according to any one of claims 1 to 8, further comprising: The layout is updated in response to the determination that a new speaker has joined or a speaker has left the real-time session; as well as Based on the updated layout, spatial audio processing is re-executed on the audio stream of the live broadcast room.

10. An audio processing apparatus, comprising: The audio stream receiving unit is configured to receive multiple audio streams for a real-time session, each audio stream being associated with a corresponding speaker among multiple speakers; A spatial audio processing unit is configured to perform spatial audio processing on the multiple audio streams based on the layout of the multiple speakers presented on the terminal; as well as A mixing unit is configured to mix the plurality of audio streams that have undergone spatial audio processing, wherein the mixed audio streams are transmitted to the terminal.

11. An electronic device, comprising: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform an audio processing method, the method comprising: Receive multiple audio streams for a real-time session, each audio stream being associated with a corresponding speaker among multiple speakers; Based on the layout of the multiple speakers presented on the terminal, spatial audio processing is performed on the multiple audio streams; and The plurality of audio streams that have undergone spatial audio processing are mixed, and the mixed audio streams are transmitted to the terminal.

12. The electronic device according to claim 11, wherein, Performing spatial audio processing on the plurality of audio streams includes: Based on the number of the plurality of speakers, determine the layout of the plurality of speakers presented on the terminal; Based on the aforementioned layout, the spatial audio parameters of each of the plurality of speakers are determined; and Spatial audio processing is performed on the associated audio stream based on the spatial audio parameters.

13. The electronic device according to claim 12, wherein, Based on the aforementioned layout, determining the spatial audio parameters of each of the plurality of speakers includes: Based on the position of the speaker among the plurality of speakers in the layout, the virtual spatial position of the speaker relative to the user of the terminal is determined.

14. The method according to claim 13, wherein, Determining the virtual spatial location of one of the plurality of speakers relative to the user of the terminal includes: Based on the speaker's position in the layout and the user's virtual listening angle, the speaker's virtual spatial position relative to the user is determined.

15. A computer program product comprising machine-executable instructions that, when executed by a device, cause the device to perform the method as described in any one of claims 1 to 9.