Spatial audio controller

By optimizing audio rendering through a spatial audio controller and adjusting the position of virtual sound sources based on voice activity detection and visual representation, the problem of excessive audio processing burden in video communication is solved, achieving a more efficient audio rendering method and providing a more realistic spatial listening experience.

CN115442556BActive Publication Date: 2026-03-20APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In video communication sessions, existing technologies place a heavy processing burden on the electronics (such as the CPU) of the local device, making it difficult to efficiently manage and render audio data from multiple remote devices to provide a realistic spatial listening experience.

Method used

A spatial audio controller is used to determine whether the audio stream is rendered as an individual virtual sound source or mixed into a single sound source based on speech activity detection and visual representation. The position of the virtual sound source is adjusted according to the visual representation and device orientation to reduce the amount of computational processing.

Benefits of technology

By optimizing the audio rendering method, the computational burden is reduced, providing a more realistic spatial listening experience, while also reducing the resource requirements of local devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115442556B_ABST
    Figure CN115442556B_ABST
Patent Text Reader

Abstract

The present disclosure relates to "spatial audio controllers." The present disclosure provides a method performed by a local device communicatively coupled with a number of remote devices, the method comprising: receiving, from each remote device in a communication session with the local device, an input audio stream; receiving, for each remote device, a set of parameters; determining, for each input audio stream, based on the set of parameters, whether the input audio stream is to be 1) rendered individually, or 2) rendered as a mix of input audio streams; for each input audio stream determined to be rendered individually, spatially rendering the input audio stream as an individual virtual sound source containing only the input audio stream; and for input audio streams determined to be rendered as a mix of input audio streams, spatially rendering the mix of input audio streams as a single virtual sound source containing the mix of input audio streams.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One aspect of the disclosure relates to a system that includes a spatial audio controller that controls how audio is spatialized in a communication session. Other aspects are also described. BACKGROUND

[0002] Many devices today, such as smartphones, are capable of various types of telecommunication activities with other devices. For example, a smartphone can make a phone call with another device. In this case, when a phone number is dialed, the smartphone connects to a cellular network, which can then connect the smartphone with another device (e.g., another smartphone or a landline). Additionally, a smartphone is also capable of making a video conference call, in which video data and audio data are exchanged with another device. SUMMARY

[0003] When a local device is in a communication session, such as a video conference call, with several remote devices, the local device can receive video data and audio data of the session from each remote device. The local device can use the video data to display a dynamic video representation of each remote participant and can use the audio data to create a spatialized audio rendering of each remote participant. To do so, the local device can perform spatial rendering operations on the audio data of each remote device so that the local user perceives each remote participant from different locations. However, performing these video processing and audio processing operations places a heavy processing burden on the local device’s electronics (e.g., central processing unit (CPU)). Thus, there is a need for a spatial audio controller that creates and manages audio spatial renderings during a communication session with remote devices that takes into account both complexity and video presentation while maintaining audio quality.

[0004] To overcome these deficiencies, the present disclosure describes a local device with a spatial audio controller to perform audio signal processing operations to efficiently and effectively spatially render input audio streams from one or more remote devices during a communication session. One aspect of the present disclosure is a method performed by an electronic device (e.g., a local device) that is communicatively coupled with one or more remote devices and engaged in a communication session. While engaged in the session, the local device receives an input audio stream from each remote device and receives a set of communication session parameters for each remote device. For example, these parameters can include one or more voice activity detection (VAD) parameters based on a VAD signal received from each remote device that indicates at least one of voice activity and voice intensity of a remote participant of the respective remote device. Further, when the devices are engaged in a video communication session (e.g., a video conference call) in which input video streams are received and visual representations (or tiles) of the video streams are placed in a graphical user interface (GUI) (view window) on a display screen of the local device, the session parameters can indicate how to arrange the visual representations within the GUI (e.g., in larger per-user tiled canvas areas or smaller per-user tiled roster areas), the size of the visual representations, etc. The local device determines, for each input audio stream, based on the set of communication session parameters, whether the input audio stream is to be 1) rendered separately relative to other received input audio streams or 2) rendered in a mix of input audio streams with one or more other input audio streams. For example, an input audio stream can be rendered separately when at least one of the VAD parameters (such as voice activity) is above a voice activity threshold (e.g., indicating that the remote participant is actively speaking), the visual representation associated with the input audio stream is contained within a highlighted area of the GUI (e.g., in a canvas area of the GUI), and / or the size of the visual representation (e.g., the size of the representation showing the remote participant’s video) is above a threshold size. Thus, for each input audio stream determined to be rendered separately, the local device spatially renders the input audio stream as an individual virtual sound source containing only the input audio stream, and for input audio streams determined to be rendered as a mix of input audio streams, the local device spatially renders the mix of input audio streams as a single virtual sound source containing the mix of input audio streams. Thus, by spatially rendering certain input audio streams that are more important to the local participant separately, while spatially rendering a mix of other input audio streams that are less critical, the local device can reduce the amount of computational processing required to spatially render all of the streams in the communication session.

[0005] In one aspect, the local device can determine an arrangement of locations of individual virtual sound sources in front of, behind, or to a side of a display screen of the local device based on the communication session parameters. For example, for each individual virtual sound source, the local device determines a location of a visual representation of a respective input video stream (associated with a respective input audio stream to be spatially rendered as the individual virtual sound source) within a GUI, and determines spatial parameters (e.g., azimuth, elevation, distance, and even reverb level) based on the determined location of the visual representation, which indicate a location of the individual virtual sound source within the arrangement (e.g., relative to a reference point in space), whereby spatially rendering the input audio stream as the individual virtual sound source can include spatially rendering the input audio stream at the location as the individual virtual sound source using the determined spatial parameters. Such locations can also be included with a virtual room model, and the spatial parameters include model aspects, such as reverb levels of a room reverb model as a function of distance and / or location in the room. Thus, the local device can be considered to be located within the virtual room. In one aspect, the arrangement also includes a location of a single virtual sound source of a grouping (e.g., mix) of input audio sources, where determining the location includes using one or more of a location of a visual representation of each of the respective input video streams associated with each of the mixed input audio streams within the GUI and / or a level of each of the mixed input audio streams, and determining the location and spatial parameters indicating the location of the virtual (grouped or mixed) sound source of the mixed input audio streams from this information. This grouped location is also used to determine new spatial parameters. For example, determining the new spatial parameters includes determining a weighted combination of spatial position data of all of the mixed input audio streams, where the weighting can be a function of the energy level of the individual streams. In one aspect, the streams selected for grouping in such a joint single location include those whose visual representations are less prominent (e.g., video tiles that are smaller than tiles within a prominent area of the GUI), or those whose visual representations are not visible (e.g., are currently implicitly off the display screen). Thus, a single grouped audio stream that is rendered to a single location in a spatial sense can be used when some of the visual representations are less prominent or not visible. Thus, in general, the individual or grouped virtual sound sources can be arranged by function, and in commensurate relationship to the arrangement of visual representations to provide an ideal spatial experience for the local user, while controlling complexity and considering the most prominent aspects in the video communication session.

[0006] According to another aspect of the disclosure, a method is performed by a local device that provides another arrangement of virtual sound source locations. For example, the local device receives an input audio stream and an input video stream from each remote device, displays the input video streams as visual representations in a GUI on a display screen, and spatially renders at least one input audio stream to output an individual virtual sound source comprising only that stream. In response to determining that an additional remote device has joined the video communication session, an input audio stream and an input video stream are received for each of the additional devices. The local device determines whether the local device supports an additional individual virtual sound source for one or more input audio streams of the additional devices. In response to determining that the local device does not support an additional individual virtual sound source, the local device defines a number of user interface (UI) zones in the GUI, each UI zone comprising one or more visual representations of one or more video streams, and spatially renders a mix of one or more input audio streams associated with the one or more visual representations included within each UI zone for each UI zone.

[0007] According to another aspect of the disclosure, a method performed by a local device provides adjustment of locations of virtual sound sources based on changes in a range (or limit) of panning due to the local device rotating to a different orientation. Specifically, a remote device receives an input audio stream and determines a first orientation (e.g., portrait orientation) of a local device. The local device determines a range of panning for a number of speakers for the first orientation of the local device spanning along a horizontal axis. The local device spatially renders the input audio stream as a virtual sound source at a location along the first horizontal axis and within the range of panning using the speakers. The local device can also pan jointly in horizontal and vertical directions. The panning limit or joint range of horizontal and vertical panning directions can be determined based on the orientation of the device. The orientation of the device can imply how the audio is positioned relative to the horizontal and vertical span of the device (e.g., the rectangular screen of the device viewed by the user). In response to determining that the local device is in a second orientation (e.g., the device has been rotated 90° to a landscape orientation), the local device determines an adjusted range of panning for the speakers that spans wider along the horizontal axis than the initial range of panning, and adjusts the location of the virtual sound source along the horizontal axis based on the adjusted range of panning. Thus, as the local device is rotated, the virtual sound source is perceived by the local user to be in a wider location than when the local device is in the previous orientation. In one aspect, the joint horizontal and vertical panning limit can be a function of the orientation. Individual or mixed sound sources will have a virtual azimuth and elevation within this range, where the mapping from the location of the visual representation uses this range to define a function for the mapping.

[0008] According to another aspect of the disclosure, a method performed by a local device determines a pan range of a number of loudspeakers based on an aspect ratio of a GUI of a communication session displayed on a display screen (e.g., a window displayed on the GUI). The local device receives an input audio stream and an input video stream, and displays a visual representation of the input video stream within a GUI of a video communication session displayed on a display screen (e.g., the display screen can be integrated within the local device, and on which can be a window containing the communication session). The local device determines an aspect ratio of the GUI, and based on the aspect ratio, determines an azimuth pan range that is at least a portion of a total azimuth pan range of the number of loudspeakers, and an elevation pan range that is at least a portion of a total elevation pan range of the loudspeakers. The local device spatially renders the input audio stream to output a virtual sound source within the azimuth and elevation pan ranges.

[0009] The above summary does not include an exhaustive list of all aspects of the present disclosure. It is contemplated that the present disclosure includes all systems and methods that can be practiced with the various aspects disclosed above, as well as modifications and permutations of these aspects. Such modifications and permutations can not have been previously considered but can be associated with particular aspects within the scope of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0010] Numerous aspects are illustrated by way of example, and not by way of limitation, in the accompanying drawings. In the drawings, like reference numerals refer to similar elements. It should be noted that "a" or "an" as used in this disclosure includes "one or more" unless the context clearly dictates otherwise. Additionally, for the purpose of clarity and the ease of understanding, one or more of the drawings can have been simplified. Further, the figures can not be drawn to scale and may

[0011] Figure 1 A system including one or more remote devices in a communication session with a local device that includes a spatial audio controller for spatially rendering audio from the one or more remote devices is shown in accordance with one aspect.

[0012] Figure 2 A block diagram of a local device 2 that spatially renders an input audio stream from a remote device in a communication session with the local device to output a virtual sound source is shown in accordance with one aspect.

[0013] Figure 3 An exemplary graphical user interface (GUI) of a communication session being displayed by a local device and an arrangement of virtual sound positions of spatial audio output by the local device during the communication session is shown in accordance with one aspect.

[0014] Figure 4A block diagram of a local device including an audio space controller that performs spatial rendering operations during a communication session is shown, according to one aspect.

[0015] Figure 5 is a flowchart of one aspect of a process for determining whether input audio streams are to be rendered individually as individual virtual sound sources or are to be mixed and rendered as a single virtual sound source, and for determining the arrangement of the virtual sound source locations.

[0016] Figure 6 is a flowchart of one aspect of a process for defining UI regions and spatially rendering input audio streams associated with the UI regions.

[0017] Figure 7 is shown are several stages in which a local device defines user interface (UI) regions for one or more visual representations of input video streams of one or more remote devices, and spatially renders one or more input audio streams associated with the defined UI regions.

[0018] Figure 8 is shown is a panning angle for rendering input audio streams at respective virtual sound sources of the input audio streams, according to one aspect.

[0019] Figure 9 is a flowchart of one aspect of a process for determining spatial parameters based on locations of respective visual representations, the spatial parameters indicating locations at which input audio streams are to be spatially rendered as virtual sound sources, according to one aspect.

[0020] Figure 10 is shown is an example of determining spatial parameters by using one or more linear functions to map locations of visual representations to angles at which input audio streams are to be rendered, according to one aspect.

[0021] Figure 11A and Figure 11B is shown is an example of determining spatial parameters by using one or more functions to map locations of estimated or real viewing angles of visual representations (e.g., using assumed plane sizes and viewing locations) to audio panning angles at which input audio streams are to be rendered, according to one aspect.

[0022] Figure 12 is shown are several stages in which a panning range of one or more loudspeakers is adjusted based on a local device rotating from a portrait orientation to a landscape orientation.

[0023] Figure 13 is a flowchart of one aspect of a process for adjusting locations of one or more virtual sound sources based on a change in one or more panning ranges of one or more loudspeakers, the change based on a change in an orientation of a local device.

[0024] Figure 14A and Figure 14B Several stages are shown in accordance with some aspects, in which a panning range is adjusted based on an aspect ratio of a communication session-based GUI.

[0025] Figure 15 is a flowchart of one aspect of a process 170 for adjusting a position of one or more virtual sound sources based on a change in an aspect ratio of a communication session-based GUI. DETAILED DESCRIPTION

[0026] Aspects of the disclosure will now be explained with reference to the accompanying drawings. The scope of the disclosure is not limited to the aspects shown, but only by the claims. The aspects shown are for illustrative purposes only and, where the shape, relative position, and other aspects of components described in an aspect are not explicitly defined, the scope of the disclosure herein is not limited only to the aspects shown, which are for illustrative purposes only. Additionally, although a number of details are set forth, it is understood that some embodiments can be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description. Also, where the context permits, singular or plural forms of any word or phrase can be used. For example, where a singular form of a word is used, this includes the plural unless the context clearly dictates otherwise. Furthermore, the word "or" is used in the inclusive sense (and not the exclusive sense) so that when used, for example, in a conjunctive context, it represents any of the possibilities.

[0027] Figure 1 A system 1 is shown that includes one or more remote devices in communication session with a local device that includes a spatial audio controller for spatially rendering audio from the one or more remote devices, in accordance with one aspect. This can allow the local device to simulate a more realistic listening experience during a (e.g., video) communication session (e.g., video conference call) with the remote devices of a local user, in which the audio is perceived by the local user as coming from sound sources in a separate frame of reference (such as the physical environment around or in front of the user), and, if there is a visual representation, moving from or around the screen, as described herein. The audio system includes a local (or first electronic) device 2, a remote (or second electronic) device 3, a network 4 (e.g., a computer network, such as the Internet), and an audio output device 6. In one aspect, the system can include more or fewer elements. For example, the system can have one or more remote devices, in which all devices are in communication session with the local device, as described herein. In another aspect, the audio system can include one or more remote (electronic) servers communicatively coupled with at least some of the devices of the audio system 1, and can be configured to perform at least some of the operations described herein. In another aspect, the system can not include an audio output device. In this case, the local device can perform audio output operations, such as driving one or more speakers using one or more audio driver signals, which can be integrated within the local device or separate from the local device (e.g., a standalone speaker communicatively coupled with the local device, such as a set of headphones).

[0028] In one aspect, the local devices (and / or remote devices) can be any electronic devices (e.g., with electronic components such as processors, memories, etc.) capable of participating in a communication session, such as a video (conference) call. For example, the local devices can be desktop computers, laptop computers, digital media players, etc. In one aspect, the devices can be portable electronic devices (e.g., capable of being hand-held), such as tablet computers, smart phones, etc. In another aspect, the devices can be head-mounted devices, such as smart glasses, or wearable devices, such as smart watches. In one aspect, the remote devices can be the same type of device as the local devices (e.g., both devices are smart phones). In another aspect, at least some of the remote devices can be different, such as some being desktop computers and others being smart phones.

[0029] As shown, the local device 2 is coupled (e.g., communicatively) to the remote device 3 via a computer network (e.g., the Internet) 4. In particular, the local and remote devices can be configured to establish and conduct a video conference call, in which the devices conducting the call exchange audio and video data. For example, the local device can use any signaling protocol (e.g., Session Initiation Protocol (SIP)) to establish the communication session and any communication protocol (e.g., Transmission Control Protocol (TCP), Real-time Transport Protocol (RTP), etc.) to exchange audio and video data during the session. For example, when a session is initiated (e.g., by a communication session application executable within the local device), the local device can capture one or more microphone signals (using one or more microphones of the local device), encode the audio data (e.g., using any audio codec), and transmit the audio data (e.g., as IP packets) to one or more remote devices, and receive audio data (e.g., as an input audio stream) from each of these remote devices via the network for driving one or more loudspeakers of the local device.

[0030] Further, the local device can transmit video data (captured by one or more cameras of the device) to each of the remote devices conducting the call, and receive video data (as an input video stream) from each of the remote devices as an output video stream, and receive at least one video signal (or input video stream) for display on one or more display screens. In one aspect, when transmitting the video data, the local device can encode the video data using any video codec (e.g., H.264), which can then be decoded and rendered on each of the remote devices to which the local device transmits the encoded data.

[0031] In one aspect, network 4 can be any type of network that enables a local device to be communicatively coupled with one or more remote devices. In another aspect, the network can include a telecommunications network with one or more cell towers that can be part of a communication network (e.g., a 4G Long Term Evolution (LTE) network) that supports data transmission (and / or voice calls) for electronic devices such as mobile devices (e.g., smartphones).

[0032] In one aspect, audio output device 6 can be any electronic device that includes at least one speaker and is configured to perform outputting sound by driving the speaker. For example, as shown, the device is a wireless headset (e.g., an in-ear headphone or earbud) that is designed to be positioned on (or in) a user’s ear and is designed to output sound into the user’s ear canal. In some aspects, the earbud can be of a sealed type with a flexible earbud tip that is used to acoustically seal the entrance of the user’s ear canal relative to the surrounding environment by blocking or occluding in the ear canal. As shown, the output device includes a left earbud for the user’s left ear and a right earbud for the user’s right ear. In this case, each earbud can be configured to output at least one audio channel of media content (e.g., the right earbud outputs the right audio channel of a two-channel input of a stereo recording (such as a musical work) and the left earbud outputs the left audio channel). In another aspect, the output device can be any electronic device that includes at least one speaker and is arranged to be worn by a user and is arranged to output sound by driving the speaker with an audio signal. As another example, the output device can be any type of headphone, such as an over-ear (or supra-aural) headphone that at least partially covers a user’s ear and is arranged to direct sound into the user’s ear.

[0033] In some aspects, the audio output device can be a headset device, as explained herein. In another aspect, the audio output device can be any electronic device that is arranged to output sound into the surrounding environment. Examples can include a standalone speaker, a smart speaker, a home theater system, or an infotainment system integrated within a vehicle.

[0034] In one aspect, the audio output device 6 can be a wireless device communicably coupled to the local device to exchange audio data. For example, the local device can be configured to establish a wireless connection with the audio output device via a wireless communication protocol (e.g., a Bluetooth protocol or any other wireless communication protocol). During the established wireless connection, the local device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the audio output device, which can include audio digital data in any audio format. In particular, the local device can be configured to establish and communicate with the audio output device over a bidirectional wireless audio connection (e.g., which allows two devices to exchange audio data), such as to make a hands-free call or use voice commands. Examples of bidirectional wireless communication protocols include, but are not limited to, Hands-Free Profile (HFP) and Headset Profile (HSP), both of which are Bluetooth communication protocols. In another aspect, the local device can be configured to establish and communicate with the output device via a unidirectional wireless audio connection, such as the Advanced Audio Distribution Profile (A2DP) protocol, which allows the local device to transmit audio data to one or more audio output devices.

[0035] In another aspect, the local device 2 can be communicably coupled with the audio output device 6 via other methods. For example, both devices can be coupled via a wired connection. In this case, one end of the wired connection can be (e.g., fixedly) connected to the audio output device, while the other end can have a connector that plugs into a jack of the audio source device, such as a media jack or a Universal Serial Bus (USB) connector. Once connected, the local device can be configured to drive one or more speakers of the audio output device with one or more audio signals via the wired connection. For example, the local device can transmit the audio signals as digital audio (e.g., PCM digital audio). In another aspect, the audio can be transmitted in an analog format.

[0036] In some aspects, the local device 2 and the audio output device 6 can be different (separate) electronic devices, as illustrated herein. In another aspect, the local device can be a component of (or integrated with) the audio output device. For example, as described herein, at least some of the components of the local device, such as the controller, can be part of the audio output device, and / or at least some of the components of the audio output device can be part of the local device. In this case, each device can be communicably coupled via traces that are part of one or more printed circuit boards (PCBs) within the audio output device.

[0037] Figure 2A block diagram of a local device 2 that spatially renders input audio streams from remote devices in a communication session with the local device to output virtual sound sources is shown in accordance with an aspect. The local device 2 includes a controller 10, a network interface 11, a speaker 12, a microphone 14, a camera 15, a display screen 13, an inertial measurement unit (IMU) 16, and (optionally) one or more additional sensors 17. In an aspect, the local device can include more or fewer elements, as described herein. For example, the device can include two or more of at least some of these elements, such as having two or more speakers, two or more microphones, two or more cameras, and two or more display screens.

[0038] The controller 10 can be a special purpose processor such as an application specific integrated circuit (ASIC), a general purpose microprocessor, a field programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). The controller is configured to perform audio signal processing operations and / or networking operations. For example, the controller 10 can be configured to conduct a video communication session with one or more remote devices via the network interface 11. In another aspect, the controller can be configured to perform audio signal processing operations on audio data (e.g., input audio streams) associated with a conducted communication session, such as spatially rendering the streams to output them as virtual sound sources to provide a more realistic listening experience for a local user. More is described herein regarding the operations performed by the controller 10.

[0039] In an aspect, the one or more sensors 17 are configured to detect an environment (e.g., in which the local device is located) and generate sensor data based on the environment. In some aspects, the controller can be configured to perform operations based on the sensor data generated by the one or more sensors 17. For example, the sensors can include a (e.g., optical) proximity sensor designed to generate sensor data indicative of an object being a particular distance from the sensor (or local device), such as detecting a viewing distance between the local device and a local user. The sensors can also include an accelerometer arranged and configured to receive (detect or sense) vibrations (e.g., speech vibrations generated when a user speaks) and generate an accelerometer signal representative of (or containing) the vibrations. The IMU is designed to measure a position and / or orientation of the local device. For example, the IMU can generate sensor data indicative of a change in orientation (e.g., with respect to any X, Y, Z axis) of the local device and / or a change in position of the device.

[0040] The speaker 12 can be, for example, an electrodynamic driver that can be specifically designed for sound output in a particular frequency band, such as a woofer, a tweeter, or a midrange driver. In one aspect, the speaker 12 can be a “full-range” (or “full-band”) electrodynamic driver that reproduces as much of the audible frequency range as possible. The microphone 14 can be any type of microphone (e.g., a differential pressure gradient microelectromechanical system (MEMS) microphone) configured to convert acoustic energy caused by sound waves propagating in an acoustic environment into an input microphone signal (or audio signal).

[0041] In one aspect, the camera 15 is a complementary metal-oxide-semiconductor (CMOS) image sensor that is capable of capturing digital images that include image data representing a field of view of the camera 15, where the field of view includes a scene of an environment in which the local device 2 is located. In some aspects, the camera can be a charge-coupled device (CCD) camera type. The camera is configured to capture still digital images and / or video represented by a series of digital images. In one aspect, the camera can be positioned anywhere near the local device. In some aspects, the device can include multiple cameras (e.g., where each camera can have a different field of view).

[0042] The display screen 13 is designed to present (or display) video of digital image or video (or image) data. In one aspect, the display screen can use liquid crystal display (LCD) technology, light-emitting polymer display (LPD) technology, or light-emitting diode (LED) technology, although other display technologies can be used in other aspects. In some aspects, the display can be a touch-sensitive display screen configured to sense user input as an input signal. In some aspects, the display can use any touch-sensing technology, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies.

[0043] In one aspect, any of the elements described herein can be part of (or integrated into) the local device (e.g., integrated into a housing of the local device). In another aspect, at least some of the elements can be separate electronic devices communicatively coupled with the local device (e.g., via a controller of a network interface of the local device) (via a Bluetooth connection). For example, the speaker can be integrated into another electronic device that is configured to receive audio data from the local device for driving the speaker. As another example, the display screen 13 can be integrated with the local device, or the display screen can be a separate electronic device (e.g., a monitor) communicatively coupled with the local device.

[0044] In one aspect, the controller 10 is configured to perform audio signal processing operations and / or networking operations, as described herein. For example, the controller can be configured to conduct a communication session with one or more remote devices, and obtain (or receive) audio / video data from the remote devices. The controller is configured to output (display) the video data on a display screen and spatially render the audio data. More is described herein regarding spatially rendering the audio data. In one aspect, the operations performed by the controller can be implemented in software (e.g., as instructions stored in memory and executed by the controller) and / or can be implemented by hardware logic structures, as described herein.

[0045] In one aspect, the controller 10 can be configured to perform (additional) audio signal processing operations based on elements coupled to the controller. For example, when the output device includes two or more “off-ear” speakers arranged to output sound into the acoustic environment rather than speakers arranged to output sound into a user’s ear (e.g., as speakers of an in-ear headphone), the controller can include a sound output beamformer configured to produce speaker driver signals that, when driving the two or more speakers, produce a spatially selective sound output. Thus, when used to drive the speakers, the output device can produce a directional beam pattern that can be pointed to a location within the environment.

[0046] In some aspects, the controller 10 can include a sound pickup beamformer that can be configured to process audio (or microphone) signals produced by two or more external microphones of the output device to form a directional beam pattern (as one or more audio signals) for spatially selective sound pickup in certain directions so as to be more sensitive to one or more sound source locations. In some aspects, the controller can perform audio processing operations (e.g., perform spectral shaping) on the audio signals containing the directional beam pattern.

[0047] Figure 3An exemplary graphical user interface (GUI) (or GUI window) of a communication session being displayed by a local device and an arrangement of virtual sound locations of spatial audio output by the local device during the communication session is shown in accordance with one aspect. Specifically, the figure shows a GUI 41 of a communication session being displayed on a display screen 13 of a local device and shows an arrangement 50 of virtual sound source locations in which spatial audio associated with one or more remote devices (participants of the remote devices) is located and perceived by a local user 40. The GUI 41 includes an arrangement 51 of visual (or video) representations (or tiles), in which five visual representations 44-48 are each associated with a different remote participant and one tile 49 associated with the local user 40. Specifically, each remote participant’s visual representation displays an input video stream received from the remote device of each respective participant that is in communication with the local device in the communication session. In one aspect, at least some of the representations can be dynamic in which live video is displayed during at least a portion of the length of the communication session. In another aspect, one or more visual representations can display a static image (e.g., a static image of a remote participant). The visual representation 49 displays video data of the local user 40 that can be captured by the camera 15 (and transmitted to at least some of the remote devices for display on their respective devices during the communication session).

[0048] As shown, the GUI is divided into two regions: a canvas (or primary) region 42 and a roster (or secondary) region 43. The canvas region includes three visual representations 44-46 and the roster region includes two visual representations 47 and 48. As shown, the visual representations within the canvas region can be visually larger than the visual designations in the roster region, in which a remote participant with prominent speech activity (e.g., relative to other remote participants in the session) is positioned within the canvas region. As described herein, remote participants can move between these regions, e.g., a remote participant can be adaptively moved from the roster region to the canvas region based on speech activity. As described herein, the local device can spatially render audio data differently based on the region in which the visual representation is located. For example, the local device can render audio data associated with remote participants in the canvas region separately, while in contrast, audio data associated with remote participants in the roster region can be mixed and then the mix can be spatially rendered as one virtual sound source. More is described herein regarding spatial rendering and regions of the GUI.

[0049] In a conventional (video) communication session between multiple remote participants and a local user of a local device 2, audio from all remote participants is perceived to originate from the same location relative to the local user (e.g., as if the remote participants were all directly talking to each other at the same point in space). Such interactions can result in many interruptions because it is difficult for the local user to follow the conversation when more than one remote participant is speaking at a time. Moreover, it is difficult to maintain eye contact or follow the conversation between different remote participants within different areas in the GUI if the audio from all remote participants comes from the same point in space (more precisely as perceived by the local user), which can not have anything to do with where the remote participants are visually located. To overcome this, the local device spatially renders the audio data received from the remote devices differently so that the sound of at least some of the remote participants originates from different locations (in space) relative to the local user 40. Specifically, this figure illustrates an arrangement 50 of virtual sound source locations in which the local device spatially renders the audio data from the remote participants associated with visual representations 44-46 individually as individual virtual sound sources 54-56, and spatially renders the audio data from one or more remote devices associated with representations 47 and 48 as a single virtual sound source 53 that includes a mix of the audio received from those devices. In one aspect, the arrangement of virtual sound source locations can be similar to the arrangement of visual representations in order to give the local user the impression that the voices from the remote participants originate from the same general locations as their respective visual representations on the screen, as shown. For example, as shown, the individual virtual sound sources are arranged (or positioned) on top of their respective visual representations, while the single virtual sound source 53 for the roster visual representations 47 and 48 is centered between those visual representations. Thus, during the communication session, when remote participant 44 speaks, the local user will perceive the remote participant's voice to originate from (or above) that participant's visual representation. In another aspect, the arrangement of virtual sound source locations can be different. For example, the arrangement can be similar to the arrangement of visual representations but scaled larger, so that although the voices do not originate from each opposite visual representation, they can originate from a general location. This can have advantages if the screen is small or if the user chooses to minimize the size of the GUI. In both cases, a larger, more natural acoustic space representation can be more comfortable and useful to the local user. More is described herein regarding how the audio data for the remote participants is spatially rendered.

[0050] As shown in the figure, a local user engages in a communication session with five remote participants. In one aspect, more remote participants can join the communication session, in which case they (or more precisely, their associated visual representations) can be placed within a canvas area or a list area. As more remote participants are placed within the canvas area, the local device performs additional spatial rendering, which may require more resources and computational processing. This additional processing can place a heavy burden on electronics such as controller 10. Therefore, as described herein, controller 10 performs spatial rendering operations that manage audio spatial rendering during the communication session with the remote devices. Further details regarding these operations are described herein.

[0051] Figure 4 A block diagram of a local device 2, including an audio spatial controller, is shown according to one aspect. This audio spatial controller performs spatial rendering operations during a communication session. Specifically, the diagram shows a controller 10 having several operation blocks for performing audio signal processing operations to spatially render input audio streams from one or more remote devices communicating with the local device. As shown, the controller includes an audio spatial controller 20, a video communication session manager 21, a video renderer 22, and an audio spatial renderer 23.

[0052] The video communication session manager 21 is configured to initiate (and perform) a communication session between a local device 2 (e.g., via network interface 11) and one or more remote devices 3. For example, the session manager may be part of (or receive instructions from) a communication session application (e.g., a telephone application) executed by the local device 2 (e.g., the controller 10 of the local device). For example, the application may display a GUI on the local device's display screen 13, which may provide the local user 40 with the ability to initiate a session (e.g., using a simulated keyboard, a contact list, etc.). Once the GUI receives user input (e.g., dialing a remote user's phone number using a keyboard), the session manager 21 may communicate with the network 4 to establish a communication session, as described herein.

[0053] Once initiated, the video communication session manager 21 can receive communication session data from each ("N") remote device in communication session with the local device. In one aspect, the received data can include an input audio stream from each remote device including audio data of the respective remote participant (e.g., captured by one or more microphones of the remote device) and an input video stream including video data of the respective remote participant (e.g., captured by one or more cameras of the remote device). Thus, the session manager can receive N input audio streams and N input video streams. In one aspect, the session manager can assign each input audio stream to a particular input audio channel of a predefined number of input audio channels. In one aspect, the assigned channel can remain assigned to the particular input audio stream (e.g., remote device) for the duration of the communication session. In one aspect, the session manager can dynamically assign input audio channels to remote participants as they join the communication session. In one aspect, the manager can receive more or fewer streams (e.g., the manager can only receive audio streams when the remote device has disabled its camera).

[0054] In one aspect, the session manager 21 can receive one or more audio signals from each remote device. For example, the input audio stream can be one audio channel (e.g., a mono audio signal). In another aspect, the input audio stream can include two or more audio channels, such as a stereo recording or an audio recording in a multi-channel format, such as a 5.1 surround format.

[0055] In one aspect, the session data can include additional data from at least some of the N remote devices, such as a voice activity detection (VAD) signal. For example, a remote device can generate a VAD signal (e.g., using a microphone signal captured by the remote device) that indicates whether voice is contained within the respective input audio stream of the remote device. For example, the VAD signal can have a high signal level (e.g., high signal level 1) when voice is detected to be present, and a low signal level (e.g., low signal level 0) when voice is not detected (or at least not detected within a threshold level). In another aspect, the VAD signal need not be a binary decision (voice / non-voice); but can be a probability of voice presence. In some aspects, the VAD signal can also indicate a signal energy level (e.g., sound pressure level (SPL)) of the detected voice.

[0056] In one aspect, the video communication session manager 21 can be configured to transmit data (e.g., audio data, video data, VAD signals, etc.) to one or more of N remote devices. For example, the manager can receive microphone signals generated by microphone 14 (which may include the voice of local user 4) and video data generated by camera 15. Once received, the session manager 21 can distribute this data to at least some of the N remote devices.

[0057] Video communication session manager 21 is configured to transmit N video streams and one or more VAD values ​​(or parameters) for each video stream to video renderer 22 based on VAD signals. For example, each (or at least one) video stream may be associated with one or more VAD parameters indicating the speech activity of a remote participant within an input audio stream associated with the input video stream (e.g., the value may be zero or 1, as described herein). In one aspect, the manager may transmit VAD parameters indicating the duration for which the remote participant is speaking. Specifically, the VAD parameter may indicate the duration for which the remote participant is currently speaking (e.g., the duration may correspond to the duration since the VAD value changed from zero to 1). In another aspect, the VAD parameter may indicate the total time during the entire session in which the remote participant has spoken (e.g., twenty minutes in a thirty-minute communication session). In yet another aspect, the VAD parameter may indicate the speech intensity (e.g., signal energy level) of the remote participant (e.g., in SPL). In yet another aspect, the session manager may transmit (initial) VAD signals received from a remote device to the video renderer.

[0058] The video renderer is configured to render the input video stream into a visual representation (such as...) Figure 3 The arrangement shown is 51). In one aspect, the renderer is configured to arrange the visual representation based on one or more VAD parameters received from the session manager 21. For example, refer to Figure 3For example, when the speech activity is above (or equal to) a speech activity threshold (e.g., equal to one), the renderer can position the visual representation within the canvas region, which can indicate that the remote participant is actively speaking. However, if the speech activity is below the threshold, the renderer can position the remote participant's visual representation in the roster region. Further, the renderer can position the visual representations within the canvas region 42 in a particular order based on the speech activity and / or speech intensity. For example, remote participants that speak frequently (and / or that are currently speaking) can be positioned higher in the arrangement. For example, remote participant 44 can be currently speaking and / or can have spoken during a majority (e.g., above a threshold) of the communication session. In contrast, remote participants that speak less frequently can be positioned closer to the roster region 43 (e.g., visual representation 46), and those that speak less regularly (e.g., below a threshold) can be positioned in the roster region. In another aspect, remote participants that have a high speech intensity (e.g., have a signal energy level above a threshold) can be positioned higher than those that do not have a high speech intensity.

[0059] With positioning the visual representations (e.g., based on the VAD parameters), the renderer can define the size of the representations based on one or more criteria. In particular, the one or more similar criteria mentioned above in relation to the position of the visual representations can apply to the size of the visual representations. For example, remote participants that speak longer and more frequently (than other participants) during the communication session can have larger visual representations than those that speak less (e.g., remote participant 44 can speak more often than participants 45 and 46). In another aspect, the size of the visual representations can be based on the speech intensity. For example, when a remote participant's speaking voice is louder (e.g., above a signal threshold), the renderer can increase the size of the participant's representation. In another aspect, the size (and / or position) can be based on the signal-to-noise ratio (SNR) of the remote participant's input audio stream, whereby participants with a higher SNR can have larger representations than those with a lower SNR (e.g., below a threshold). In one aspect, the representations can appear larger and be positioned higher in order to provide a visual significance to the local user of who should be given during the communication session. In other aspects, the (e.g., vertical) position within the canvas region can be based on the size of the representations. For example, the largest (with more corresponding surface area) representation 44 is positioned above representation 45, which is larger than representation 46.

[0060] In one aspect, the renderer 22 can position the representations in the roster based on any of the criteria described herein that are below one or more respective thresholds. For example, a remote participant can be positioned in the roster area when speech activity is infrequent and the time period is short. In some aspects, the position of the remote participant and their size can also be based on speech intensity. For example, a remote participant with high intensity (e.g., with an energy level above a threshold) can be positioned in the canvas area and / or can have a representation size that is larger than a remote participant with lower speech intensity, as described herein. In one aspect, the visual representations in the roster area can all have the same size, as shown in FIG. 4B. Figure 3

[0061] In one aspect, the rendering operations performed by the renderer 22 can be dynamic throughout the communication session. For example, the renderer can continuously and dynamically rearrange and / or resize the visual representations based on changes in one or more VAD parameters. For example, when a remote participant is speaking less, the position of the visual representation 44 can change (e.g., lower along the arrangement) and / or its size can change within the canvas area (e.g., size decreases). If speech activity drops (e.g., below a speech activity threshold) for a period of time, the renderer can eventually place that visual representation in the roster. Conversely, the renderer can adjust the position of one or more roster representations based on the criteria set forth herein. For example, when the remote participant associated with representation 47 begins to speak more frequently, the renderer can move that representation into the canvas area. In one aspect, this determination can be based on whether the VAD parameters indicate that the remote participant has been speaking for a period of time (e.g., consecutively).

[0062] In one aspect, by moving visual representations into and out of the two areas, the renderer can adjust the arrangement as needed. For example, if the renderer moves representation 47 into the canvas area, the canvas representations can be moved around the canvas area in order to accommodate the addition. In particular, the renderer can distribute the representations evenly within the canvas area. In another aspect, the renderer can arrange the representations such that none overlap each other. In some aspects, the renderer can adjust the arrangement 51 of visual representations as additional remote participants join the communication session, adding their respective visual representations to the canvas area 42 or the roster area 43, according to the criteria set forth herein. The renderer 22 transmits the N video streams to the display screen 13 for display in the GUI 41, as described herein.

[0063] ​In one aspect, the renderer is configured to generate N sets of communication session parameters based on the criteria set forth herein, one set of the parameters for each (input audio stream received from a respective remote device) remote device. For example, the set of communication session parameters can: 1) indicate a size of the GUI (e.g., an orientation and / or a size of the GUI relative to the display screen), 2) indicate a position of the visual representation (e.g., X, Y coordinates), 3) indicate a size of the visual representation of the respective input video stream in the GUI, and 4) can include one or more of the VAD parameters as described herein. In one aspect, the position indicated by the parameters can be a specific position within the visual representation. For example, the position can be a center point of the visual representation. In another aspect, the position can be based on the video displayed within the visual representation. For example, the position can be a specific portion of the remote participant that is displayed within the visual representation, such as the mouth of the remote participant. In one aspect, to determine the position, the renderer can execute an object recognition algorithm to identify the mouth of the remote participant.

[0064] In one aspect, the position of the visual representation can be relative to the size of the GUI and / or can be relative to the display screen of the local device. As another example, the coordinates can be relative to a component of the local device, such as when the local device is a multimedia handheld device (e.g., a smartphone), the coordinates can be based on the dimensions of the housing of the device or can be based on the display screen size that displays the GUI. In this case, the parameters can also include boundary conditions of the GUI and / or component (e.g., a width of the GUI (in the X direction) and a height of the GUI (in the Y direction)). In one aspect, the renderer can translate one or more of the VAD parameters into another domain to generate a prominence value that is a rank ordering from most prominent "1" to least prominent "N" from the audio streams. Thus, in Figure 3 In this case, each of the visual representations 44-48 will be assigned a prominence value between 1 and 5. In one aspect, the prominence value can correspond to the position and / or size of the visual representation within the GUI (e.g., the visual representation has a higher prominence value than the other visual representations).

[0065] The renderer 22 transmits the N sets of communication session parameters to the video communication session manager 21, which transmits the N sets of parameters and the N input audio streams associated with the parameters to the audio spatial controller 20. In one aspect, the video renderer can directly transmit the sets of communication session parameters to the audio spatial controller. The audio spatial controller is configured to receive the parameters and the input audio streams and is configured to perform one or more audio signal processing operations to spatially render at least some of the input audio streams in accordance with the communication session parameters. As described herein, spatially rendering audio streams can require a large amount of processing power. When a local user is in a communication session with a small number of remote participants, the audio spatial controller 20 can have the resources to individually spatially render audio data from each of the remote participants. However, as the number of remote participants grows, the controller can not be able to process all of the data as individual virtual sound sources. Accordingly, the controller is configured to determine, for each input audio stream, whether the input audio stream is to be 1) individually rendered relative to other received input audio streams, or 2) rendered as a mix of input audio streams with one or more other input audio streams based on the set of communication session parameters. Once determined, the audio spatial controller can be configured to determine how to spatially render (e.g., where to output a virtual sound source of an input audio stream) the input audio stream. Accordingly, the audio spatial controller can manage the output of input audio streams, thereby ensuring that there are sufficient computing resources.

[0066] The audio spatial controller 20 includes an individual audio stream selector 27, a spatial parameter generator 25, and a matrix / router 26. The selector is configured to select one or more of the N input audio streams to be individually rendered, and to select one or more of the remaining N input audio streams to be rendered as a mix. In one aspect, the audio spatial renderer 23 can be configured to spatially render a limited number of output audio channels as virtual sound sources (e.g., due to resource limitations as described herein). When the audio spatial controller spatially renders input audio streams, the controller can assign one or more input audio streams to one of the output audio channels (which the audio spatial renderer renders as a virtual sound source). Accordingly, affected by the limited number of output audio channels, the controller can be limited to output a predefined number of virtual sound sources. In one aspect, the predefined number is between three and six output audio channels. Accordingly, of the input audio streams, the selector can only select a number of input audio streams for individual spatial rendering that is equal to or less than the predefined number. In some aspects, one of the predefined number of output audio channels can be reserved for spatially rendering a mix of input audio streams for remote participants of a roster region 43. More is described herein regarding spatially rendering input audio streams.

[0067] In one aspect, selector 27 is configured to determine, based on N sets of communication session parameters received from session manager 21, which of the N input audio streams should be spatially rendered individually (e.g., M individual audio streams less than or equal to N audio streams), and / or rendered as a mixture. For example, the selector may determine whether an input audio stream should be rendered individually based on speech activity, speech intensity, and / or prominence values ​​indicated (or determined) by one or more VAD values ​​of a remote participant. Specifically, the selector may determine that an input audio stream should be rendered individually when speech activity is above a threshold (e.g., indicating that a remote participant is currently and / or periodically speaking during a communication session). In some aspects, where a ranking of prominence values ​​is available, the selector may select the highest “M” streams as those to be spatially rendered individually. In one aspect, the M streams are less than or equal to a predefined number of audio output channels that can be used for spatial rendering as virtual sound sources, as described herein.

[0068] On the other hand, the selector may determine whether the input audio stream should be rendered separately based on characteristics (such as the position and size of the representation) of the visual representation of the input video stream (e.g., a video containing remote participants) associated with the input audio stream displayed in the GUI. For example, the selector determines that the input audio stream should be rendered separately when the position of the visual representation is within the canvas area of ​​the GUI. Similarly, the selector may determine that the input audio stream should be rendered separately when the size of the visual representation is above a threshold size (e.g., larger than the size of all or some of the representations displayed in the list area). In one aspect, the selector may determine whether the input audio stream should be rendered as a blend based on the same criteria. For example, the selector determines that the input audio stream should be rendered as a blend when the associated visual representation has a size below a threshold size, is within the list, and / or has speech activity below a threshold. In another aspect, the determination made by the selector may be based on one or more criteria. For example, the selector may determine that the input audio stream should be rendered separately when one or more of the criteria proposed herein are met.

[0069] Spatial parameter generator 25 is configured to receive N sets of communication session parameters from session manager 21, and is configured to determine the arrangement of individual virtual sound sources for the input audio stream to be rendered individually based on these communication session parameters (e.g., Figure 3and / or the position of one or more individual virtual sound sources, each virtual sound source having a mix of one or more input audio streams that would not be rendered individually within the physical environment (e.g., with respect to the local user 40). In one aspect, the arrangement of virtual sound source positions that will be rendered by the local device can be the same (or similar) as the arrangement of visual representations. For example, with reference to Figure 3 the arrangement 50 of virtual sound source positions (e.g., approximately) maps to the arrangement of visual representations such that the sound source 54 is perceived by the local user 40 to originate from the remote participant 44. In another aspect, the arrangement of sound source positions can be different. For example, the virtual positions can have the same arrangement as the visual representations, but can be proportionally larger or smaller. By having the arrangement of sound sources be proportional or commensurate in a relative but not exact sense, advantages are had as described herein. In another aspect, such an arrangement is beneficial where the positions of the visual representations are not fixed or certain with respect to the local user. Thus, with reference to Figure 3 When rendered by the local device, instead of the local user perceiving the individual virtual sources at (or approximately) the same positions as the corresponding visual representations, these positions can be wider and higher than the GUI when scaled up proportionally.

[0070] In another aspect, the arrangement 50 of virtual sound source positions can not be similar to the arrangement of visual representations. For example, instead of being distributed around a two-dimensional (2D) XY plane, the virtual sound sources can be distributed along one axis (e.g., along a vertical axis). In another aspect, although shown as a 2D planar distribution, the virtual sound sources can be part of a three-dimensional (3D) sound field, where the virtual sound sources appear to originate from various distances from the local user. More is described herein regarding producing a 3D sound field.

[0071] In one aspect, the generator determines, for each input audio stream, one or more spatial parameters (or spatial data) indicative of spatial characteristics for spatially rendering the input audio stream as a virtual sound source using a respective set of communication session parameters.

[0072] For example, the spatial parameters can include the position of the virtual sound source within the arrangement of virtual sound source positions based on the determined positions of the visual representations. In particular, the spatial parameters can map the virtual sound sources onto their respective visual representations (above, behind, or adjacent). For example, the spatial parameters can indicate that the position of the virtual sound source is positioned on its corresponding visual representation such that the sound of the remote participant is perceived by the local user to originate from the participant’s visual representation.

[0073] In one aspect, spatial parameters indicating the location of the virtual sound source may be relative to (e.g., predefined) a reference point located in front of the display screen 13 of the local device. In some aspects, the reference point is a pre-determined viewing position within the physical environment (e.g., positioned at the head of the local user while viewing the display screen 13). In another aspect, the reference point may be determined by a spatial parameter generator. For example, the generator may use sensor data to determine the location of the local user, such as proximity sensor data generated by one or more proximity sensors. In yet another aspect, the generator may receive user input (e.g., via a GUI displayed on the display screen) indicating the location of the local user (e.g., the distance the user is positioned from the local device (the display screen)).

[0074] To map the location of virtual sound sources (e.g., individuals) associated with the canvas visual representation, a spatial parameter generator can be configured to determine one or more translation ranges of one or more speakers 12 within which the virtual sound sources can be positioned. In one aspect, the translation ranges can be predefined. Using the translation ranges, the spatial parameter generator generates one or more spatial parameters comprising one or more angles within one or more translation ranges, the one or more translation ranges being positions within an arrangement corresponding to the location of the virtual sound source relative to the visual representation. For example, a reference... Figure 8 The visual representation 44 may include azimuth angles (e.g., -θ) between the position of the individual virtual sound source (e.g., L1) along a first axis (e.g., the X-axis) and a second (reference or 0°) axis (e.g., the Z-axis). L1 And may include an elevation angle (e.g., +β) along the third axis (e.g., the Y-axis) between L1 and the second axis. L1 This article describes more about determining the location of virtual sound sources.

[0075] On the other hand, the generator can determine one or more additional spatial parameters. For example, the generator can determine the distance (e.g., distance in the Z direction) at which the virtual sound source will be perceived by the local user. In one aspect, this distance can be determined based on the size and / or position of the visual representation within the GUI. For example, the generator can assign a first distance (e.g., from a reference point) to a first visual representation of a first size and a second distance to a second visual representation of a second size. In one aspect, the first distance can be shorter than the second distance when the first visual representation is larger than the second visual representation. For example, it can be... Figure 3The visual representation 44 is assigned a distance to the reference point that is shorter than the distance assigned to the representations 45 and 46. In another aspect, the distance can be based on the position of the canvas visual representation, whereby a higher positioned visual representation can have a shorter distance than a lower positioned visual representation within the canvas area. In another aspect, the distance can be defined based on the area in which the representation is located. For example, representations within the roster area 43 can have a greater distance than any representation within the canvas.

[0076] In another aspect, the spatial parameters can include a reverb level that is used by the audio spatial renderer to apply reverb to the input audio streams. In one aspect, the addition of reverb to the streams can be modeled based on the virtual audio position being located as a position in a virtual room around the device. Specifically, the spatial parameter generator can produce (or receive) a reverb model for the room as a function of distance and / or position in the room, such that when applied to one or more input audio streams, gives the local user an audible impression that the communication session is occurring within the virtual room (e.g., the conversation between the remote participant and the local user is occurring within a conference room). In some aspects, the generator can determine a reverb level to apply to one or more input audio streams based on one or more communication session parameters. For example, the reverb level can be based on the position and / or size of the visual representation, such as having a greater reverb level for a smaller size visual representation to give the listener the impression that the virtual sound source associated with that visual representation is farther away than a virtual sound source that can be associated with a larger size visual representation having a lower applied reverb level. In one aspect, the application of reverb (by the renderer 23) can provide spatial depth to the virtual sound sources. Figure 8 to Figure 11B More is described regarding the spatial parameters and how these parameters are generated.

[0077] Once the one or more spatial parameters for all (or most) of the virtual sound sources are determined, the generator determines whether the spatial parameters for the input audio streams are to be used for rendering the individual sources. Specifically, the generator receives one or more control signals from the selector 27 that indicate which set of communication parameters are associated with input audio streams to be individualized, and which set of parameters are associated with streams that are not to be individualized. The generator 25 passes the spatial parameters to the audio spatial renderer 23, which are used to spatially render the respective input audio streams as individual virtual sound sources at the positions indicated by the one or more spatial parameters.

[0078] However, for non-personalized streams, the generator can generate different (or a new set) of one or more spatial parameters for the mix of one or more input audio streams that indicate a particular location based on at least some of the spatial parameters of the streams in the mix. For example, the different spatial parameters can include a weighted combination of the spatial parameters of all (or some) of the input audio streams of the mix. For example, the single virtual sound source 53 associated with the roster visual representations 47 and 48 is positioned in the middle of the visual representations within the roster area (or can be perceived to originate from the center of the roster area). Thus, this location can be determined by averaging the spatial parameters of the two roster representations.

[0079] In another aspect, the location of the single virtual sound source comprising the mix of one or more input streams can be based on whether a remote participant within the mix is speaking. For example, the generator can analyze the VAD parameters of a set of conversation parameters, and when these parameters indicate that a participant is speaking (e.g., speech activity is above a threshold), the generator can generate one or more spatial parameters for the mix such that the virtual sound source of the mix is at the location of the visual representation of that participant, similar to the placement of individual virtual sound sources. To illustrate, refer to Figure 3 , the single virtual sound source can be positioned at a location based on which of the two remote participants 47 or 48 is speaking. In this case, the virtual sound source 53 can move (or switch) from the two locations (e.g., along the X-axis) depending on which of the two participants is speaking.

[0080] In another aspect, if the weighting of the visual representations of a particular stream is a function of the energy level of the audio of the stream, then the virtual sound source can be determined by the location data of the audio sources in the mix that are more dominant (e.g., have the highest speech activity, etc.). For example, when the visual representations of the roster area are in a row (as shown in Figure 3 , the sound location can be closer to the representation of the remote participant that has the greatest (or highest) speech activity relative to the other remote participants. Thus, even though the visual representations themselves are not moved, the audio location can be moved to accommodate the speech activity within the row (or column) of visual representations. Moreover, as the visual representations move between the canvas area and the roster area, the determination of which streams are rendered individually and which are mixed can be adjusted based on one or more of the same criteria set forth herein. Thus, for both individual virtual sound sources and the single sound source of the mix, the generation of spatial parameters can be dynamic throughout the communication session. As with the personalized virtual sound sources, the generator passes the spatial parameters of the single virtual sound source to the Tenderer 23.

[0081] In one aspect, the difference in the arrangement can be based on the panning range of the loudspeakers. More is described herein with respect to panning ranges.

[0082] The matrix / router 26 is configured to receive N input audio streams (e.g., as N input audio channels) from the manager 21 and one or more control signals from the stream selector 27 indicating which stream is personalized and which stream is non-personalized, and to route the audio streams to the audio spatial renderer 23 via (e.g., predefined) output audio streams. In particular, the matrix / router 26 is configured to route M individual audio streams, each of which is selected by the stream selector to need to be spatially rendered individually via its own output audio stream. In other words, the router allocates an output audio stream (or channel) to each of the N audio streams to be spatially rendered individually. In addition, the matrix / router mixes (e.g., by performing a matrix mixing operation) the non-personalized audio streams into a mix of audio streams (e.g., as a single audio output stream).

[0083] The audio spatial renderer 23 is configured to receive a plurality of input audio streams. As shown, this can correspond to M individual audio streams and one or more single audio streams corresponding to a mix of one or more audio streams. In one aspect, there can be more than one input audio stream corresponding to a mix, as described herein, for each input audio stream the renderer also receives spatial parameters directing how to render the input stream. The renderer receives the input stream and spatial parameters for each stream and is configured to spatially render these streams according to the spatial parameters to produce a placement of virtual sound sources (e.g., a three-dimensional (3D) sound field including the same) using one or more loudspeakers 12. In particular, for each input stream, the audio spatial renderer 23 internally creates a spatial rendering of that stream. These spatial renderings are then combined, e.g., by merging, to create an output audio stream (e.g., which can include two or more driver signals for driving loudspeakers). In one aspect, the audio spatial rendering can use any spatial rendering method, such as vector-based amplitude panning (VBAP), to render each of the audio streams to output individual virtual sound sources and a single sound source by two or more loudspeakers, where each individual virtual sound source includes a separate individual audio stream and the single sound source includes a mix of the audio streams at a location indicated by the corresponding spatial parameters. In another aspect, the renderer can apply other spatial operations, such as upmixing one or more individual audio streams using the spatial parameters to produce a multi-channel for driving two or more loudspeakers. For example, the renderer can produce a multi-channel audio in a surround sound multi-channel format (e.g., 5.1, 7.1, etc.), where each channel is for driving a particular loudspeaker 12.

[0084] In one aspect, the renderer can apply other spatial operations to produce binaural, two-channel output signals (e.g., which can be used to drive headphones, as described herein). In one aspect, the renderer spatially renders the input audio stream, which can apply one or more spatial filters, such as head-related transfer functions (HRTFs). For example, using the spatial parameters (which can include azimuth, elevation, distance, reverberation level, etc., as described herein), the renderer can determine one or more HRTFs, and can apply the HRTFs to the received input audio stream to produce binaural audio signals that provide spatial audio. In one aspect, the spatial filters can be generic or predetermined spatial filters (e.g., determined in a controlled setting, such as a laboratory), which can be applied by the renderer to predetermined locations based on the distance indicated by the spatial parameters (e.g., typically optimized for one or more listener and / or optimal locations in front of the audio device). In another aspect, the spatial filters can be user-specific (e.g., can be determined based on user input or can be automatically determined by the local device) according to one or more measurements of the listener’s head. For example, the system can determine HRTFs, or equivalently head-related impulse responses (HRIRs) based on anthropometric measurements of the listener. For example, the renderer can receive sensor data (e.g., image data produced by camera 15), and use the data to determine anthropometric measurements of the listener.

[0085] In another aspect, the renderer can perform a cross-talk cancellation (XTC) algorithm. For example, the renderer can perform the algorithm by mixing and / or delaying (e.g., by applying one or more XTC filters to the audio stream) the audio stream to produce one or more XTC signals (or driver signals). In one aspect, when used to drive one or more of the speakers, the renderer can produce one or more first XTC audio signals containing audio content (at least a portion of which) of the audio stream that is primarily heard at one ear (e.g., the left ear) of a listener within the optimal location (e.g., which can be in front of or facing the local device), and produce one or more second XTC audio signals containing audio content of the audio stream that can be primarily heard at the other ear (e.g., the right ear) of the user.

[0086] In some aspects, the Tenderer 23 can perform one or more additional audio signal processing operations. For example, the Tenderer can apply reverb to one or more of the input audio streams (or renderings of streams) based on the received reverb level and the received spatial parameters. In another aspect, the Tenderer can perform one or more equalization operations (e.g., to spectral shaping) on one or more of the streams, such as by applying one or more filters (such as low-pass filters, band-pass filters, high-pass filters, etc.). In another aspect, the Tenderer can apply one or more scalar gain values to one or more of the streams. In some aspects, the application of equalization and scalar gain values can be based on the distance values of the spatial parameters, such that the application of these operations provides spatial (or depth) for the respective virtual sound sources.

[0087] As a result of spatial rendering of each individual input audio stream and one or more mixtures of input audio streams, the Tenderer produces a single set of (e.g., one or more) driver signals for driving one or more loudspeakers 12, which can be part of the local device or separate from the local device, in order to produce a 3D sound field including each of the individual virtual sound sources and one or more single virtual sound sources, as described herein.

[0088] In some aspects, the Controller 10 can perform one or more additional audio signal processing operations. For example, the Controller can be configured to perform an active noise cancellation (ANC) function to cause the one or more loudspeakers to produce anti-noise in order to reduce ambient noise from the environment that leaks into the user’s ear (e.g., when the loudspeakers are part of headphones worn by a local user). The ANC function can be implemented as one of feed-forward ANC, feedback ANC, or a combination thereof. Thus, the Controller can receive a reference microphone signal from a microphone that captures external environmental sounds, such as microphone 14. In another aspect, the Controller can perform any ANC method to produce anti-noise. In another aspect, the Controller can perform a transparency function in which the sound reproduced by the device is a reproduction of the ambient sound captured by an external microphone of the device in a “transparent” manner (e.g., as if the headphones were not being worn by the user). The Controller processes at least one microphone signal captured by at least one microphone and filters the signal through a transparency filter, which can reduce acoustic occlusion due to the audio output device being on, in, or above the user’s ear, while also preserving the spatial filtering effects of the wearer’s anatomical features (e.g., head, pinna, shoulders, etc.). The filter also helps preserve timbre and spatial cues associated with the actual ambient sound. In one aspect, the filter of the transparency function can be user-specific according to specific measurements of the user’s head. For example, the Controller can determine the transparency filter according to an HRTF or an equivalent to an HRIR based on anthropometric measurements of the user.

[0089] In another aspect, the controller 10 can perform de-correlation on one or more audio streams in order to provide a more (or less) diffuse 3D sound field. In some aspects, de-correlation can be activated based on whether the local device is outputting the 3D sound field via headphones or via one or more (extra-aural) speakers that can be integrated within the local device. In another aspect, the controller can perform an echo cancellation operation. Specifically, the controller can determine a linear filter based on the transfer path between the one or more microphones 14 and the one or more speakers 12 and apply the filter to the audio stream to generate an echo estimate that will be subtracted from the microphone signal captured by the one or more microphones. In some aspects, the controller can use any method of echo cancellation.

[0090] As noted above, the audio spatial controller 20 is configured to spatially render one or more input audio streams associated with the canvas visual representation as individual virtual sound sources. However, in another aspect, the controller can spatially render a mix of one or more input audio streams of the canvas remote participants as one virtual sound source. In one aspect, it can have only one such grouped mix, or it can use multiple mixes. As described herein, the spatial audio controller can have a predefined number of output audio channels for individualized spatial rendering. However, in some cases, the controller can determine that more virtual sound sources are needed than the local device outputs. For example, additional remote participants can join the communication session, and the controller can determine that one or more of their respective input audio streams are to be rendered individually, as described herein. As another example, the controller can determine that existing remote participants that were not previously rendered individually can now require their output individual virtual sound source. For example, during the communication session, the controller can determine that a roster remote participant is moving from the roster into the canvas region based on the criteria set forth herein. Thus, if an output audio channel is needed, but there are not enough available channels (e.g., the existing spatial rendering of existing virtual sound sources has reached the total number of output audio channels, etc.), the audio spatial controller can begin spatially rendering a mix of the canvas input audio streams to a single virtual sound source. In one aspect, the determination can be based on the location of the visual representations within the canvas region. For example, the controller can perform a vector quantization operation with respect to the visual representation locations within the GUI to group (or as a mix) one or more input audio streams of neighboring visual representations. In another aspect, the controller can group streams based on the distance between the visual representations within the canvas region (e.g., within a threshold distance). In another aspect, the controller can assign predefined regions within the GUI (e.g., its canvas region), whereby streams associated with visual representations within the region are mixed. More is described herein with respect to regions.

[0091] Once the grouping of the one or more input audio streams is determined, the controller (e.g., its spatial parameter generator 25) can determine one or more spatial parameters for the mix (e.g., in a similar manner to determining spatial parameters for the list representation, such as determining a weighted combination of the spatial parameters of the streams in the mix, etc.), and transmit the spatial parameters to the matrix / router 26 to mix the input audio streams and transmit the mix as one of the output audio channels. In one aspect, the mixing of the audio streams can be dynamic as the visual representation within the GUI changes.

[0092] In another aspect, the output audio channels can be predefined as to whether the channels support individual input audio streams or a mix of input audio streams. In this case, the controller can be configured to accommodate a plurality of individual virtual sound sources and a plurality of mix sound sources. In one aspect, the determination of which input audio streams to mix and how much to mix can be based on these quantities.

[0093] In one aspect, the audio spatial controller 20 can dynamically render the input audio streams such that the controller adjusts the rendering based on changes to the audio system 1 (e.g., its local device). As described herein, the spatial parameters can be generated based on the panning range of the speakers of the local device. In one aspect, the panning range can change based on certain criteria. For example, the panning range can be based on the physical arrangement of the speakers integrated within the local device. Thus, as the local device changes its position and / or orientation, the panning range can also change. Accordingly, the audio spatial controller can be configured to adjust the spatial rendering based on any changes to the local device, such as changes in orientation and / or changes to the aspect ratio of the GUI relative to the display screen. Figure 12 to Figure 15 More is described regarding adjusting the spatial rendering based on changes to the aspect ratio of the local device and / or GUI.

[0094] Figure 5 、 Figure 6 、 Figure 9 、 Figure 13 and Figure 15 are flowcharts of processes 30, 90, 80, 140, and 170, respectively, for performing one or more audio signal processing operations to spatially render input audio streams of a communication session. In one aspect, the processes can be performed by one or more devices of the audio system 1, as shown in Figure 1 For example, at least some of the operations of the processes can be performed by the local device 2 (e.g., its controller 10). In another aspect, at least some of the operations can be performed by another device, such as a remote server communicatively coupled with the local device.

[0095] Regarding Figure 5This diagram is a flowchart of one aspect of the process 30 used to determine whether the input audio stream should be rendered individually as individual virtual sound sources or mixed and rendered as individual virtual sound sources, and to determine the arrangement of these virtual sound source locations. In one aspect, the operations described in the process are performed by one or more operation blocks of the controller 10, such as... Figure 4 As described herein. In one aspect, prior to the start of the process, local device 2 has established a communication session with one or more remote devices 3, as described herein. The process begins with controller 10 receiving communication session data (e.g., input audio streams, input video streams, and / or VAD signals) from each of the one or more remote devices that have established a communication session with local device (at box 31). Controller 10 determines a set of communication session parameters for each remote device (e.g., based on the input video streams and / or VAD signals) (at box 32). For example, a video renderer may determine one or more VAD parameters, the size of the visual representation of each input video stream within the GUI, the position of the visual representation, reverberation level, etc., as parameters. For each input audio stream, controller 10 determines, based on the set of communication session parameters, whether the input audio stream is to be 1) rendered separately relative to other received input audio streams, or 2) rendered as a mixture of input audio streams with one or more other input audio streams (at box 33). Controller 10 determines, based on the set of communication session parameters, the arrangement of 1) individual virtual sound sources, each comprising a separately rendered input audio stream, and 2) the locations of one or more individual virtual sound sources, each comprising a mixture of input audio signals (in box 34). Specifically, in determining the arrangement, the controller determines the spatial parameters of each input audio stream (e.g., based on the location of the visual representation of the corresponding input video stream displayed in the communication session GUI on display screen 13). Additionally, based on the arrangement of the visual representation, the controller (e.g., its matrix / router 26) can perform a matrix mixing operation to mix the one or more input audio streams to produce a mixture of input audio streams. Controller 10 renders the space of each input audio stream determined to be rendered separately as containing only the individual virtual sound sources of that input audio stream (in box 35). Controller 10 also renders the mixing space of each input audio stream as containing a single virtual sound source of the mixture (in box 36).

[0096] As described above, the controller 10 manages the allocation of output audio channels to individual input audio streams based on one or more criteria (e.g., whether a remote participant is actively speaking). This is to ensure that the number of allocated output audio channels, some of which include individualized input audio streams and / or a mix of one or more input audio streams, does not exceed a predefined number in order to optimize computing resources. However, in some instances, the controller can determine that the number of input audio streams determined to be individually rendered as individual virtual sound sources exceeds the predefined number of audio output channels. This can be due to a large number of remote participants actively speaking (e.g., having a VAD parameter value above a threshold). Accordingly, as described herein, the controller can allocate one or more input audio streams as a mix to a single virtual sound source instead of exceeding the predefined number of output audio streams. Additionally, the controller can adjust the arrangement of virtual sound source locations. As described herein, the controller can group one or more input audio streams of a canvas region and render the group to a virtual sound source. In another aspect, the controller is able to adjust the arrangement of virtual sound source locations and visual representations in a grid-like manner in order to accommodate more remote participants. Figure 6 A process is described for determining whether to adjust the arrangement based on adding a remote participant within a communication session, which can result in exceeding the predefined number of output audio channels if their respective input audio stream is individually spatially rendered as an individual virtual sound source.

[0097] In particular, Figure 6 is a flowchart of one aspect of a process 90 for defining user interface (UI) regions (e.g., in a grid-like manner) and spatially rendering input audio streams associated with representations within the UI regions, where each UI region includes one or more visual representations. The process 90 begins with the controller receiving an input audio stream and an input video stream for each remote device of a first set of remote devices in a video communication session with the local device (at block 91). For each input video stream, the controller displays a visual representation of the input video stream in a GUI on the display screen 13 (at block 92). In particular, the video renderer 22 receives the input video stream (and VAD parameter) and determines an arrangement of visual representations. Once determined, the renderer displays the video stream on the display. The controller spatially renders the input audio stream to output one or more individual virtual sound sources and / or a single virtual sound source that includes a mix of the input audio streams (at block 93). Thus, at this point, the local user 40 can perceive the audio of remote participants at virtual sound source locations within the physical environment as shown in Figure 3 For example, the controller can spatially render at least one input audio stream to output a (e.g., individual) virtual sound source that includes the input audio stream through one or more loudspeakers 12. In one aspect, the number of individual virtual sound sources can be below the predefined number.

[0098] The controller 10 determines that a second group (one or more additional) remote devices have joined the video communication session (at block 94). In one aspect, this determination can be based on one or more requests received by the session manager 21 for the additional remote devices to participate in a pre-existing communication session. In response, the session manager can accept the requests and establish communication channels with the remote devices to begin receiving session data. In another aspect, this determination can be based on the session manager receiving session data from newly added remote devices of the communication session. For example, the session can be an "open" session in which remote participants are free to join the session (e.g., without authorization from the local device and / or other remote devices already participating in the session). In response to determining that the second group has joined the call, the controller receives input audio and video streams from each of the remote devices.

[0099] The controller 10 determines whether the local device supports an additional individual virtual sound source for one or more input audio streams of the second group of remote devices (at decision block 95). Specifically, as described herein, the local device (e.g., its controller 10) can be configured to spatially render a predefined number of input audio streams as individual virtual sound sources. In the existing configuration (e.g., the video communication session with the first group of remote devices), the controller 10 has been rendering a number of individual input audio streams that can be below the predefined number. Accordingly, the individual audio stream selector 27 can receive the additional group communication session parameters and determine whether the additional input audio streams are to be rendered as individual virtual sound sources. Specifically, the selector determines whether the number of input audio streams of the first group of remote devices and the second group of remote devices that are determined to be spatially rendered individually is greater than the predefined number.

[0100] If so, the controller determines that the local device does not support the aggregation of the individualized streams that the communication session can require. In response, the controller 10 defines a number (or one or more) of UI zones in the GUI, each UI zone including one or more visual representations of one or more input video streams of the first group of remote devices, the second group of remote devices, or a combination thereof, to be displayed in the UI zone (at block 97). For example, the controller 10 (e.g., its video Tenderer 22) can display all of the visual representations associated with the first group of remote devices and the second group of remote devices in a grid fashion (e.g., in one or more rows and one or more columns). In one aspect, the visual representations can be evenly spaced between the edges of the display screen and each other, and / or can have the same size. The controller 10 establishes a virtual grid of UI zones on the GUI, where each UI zone contains one or more visual representations. For example, the controller can assign one or more adjacent visual representations within the GUI to each UI zone when establishing the grid. In another aspect, the controller can define the UI zones based on the number of visual representations and / or input audio streams received from the two groups of remote devices. For example, if the predefined number of individual virtual sound sources (or output channels) is four and there are eight input audio streams, the controller can evenly assign (or distribute) two input audio streams to each UI zone such that the number of defined UI zones does not exceed the predefined number.

[0101] Once these UI zones are defined, the controller spatially renders, for each UI zone, a mix of one or more input audio streams associated with one or more visual representations included within the UI zone as a virtual sound source through the speakers 12 (at block 98). In particular, the controller can spatially render the mix such that each zone is associated with its own virtual sound source. For example, the virtual sound source of each zone can be located at a position within the UI zone displayed on the display screen. In particular, the virtual sound source can be positioned at the center of the UI zone such that the audio of the input audio streams from the zone are perceived by the local user as originating from the zone. In one aspect, the controller can dynamically position the virtual sound source based on speech activity of one or more remote participants associated with the zone. For example, the controller can place the virtual sound source on the visual representation associated with one of the input audio streams in the mix of the zone that has a signal energy level above a threshold (e.g., associated with speech activity greater than the threshold, indicating that a remote participant is speaking).

[0102] However, if the local device does support rendering the additional input audio stream separately, the controller spatially renders the additional input audio stream to output one or more additional individual virtual sound sources (at block 96). Specifically, the controller can add visual representations to the GUI of the communication session (e.g., the canvas region) and output the additional stream as an individual source. Otherwise, the controller can rearrange the arrangement of virtual sound source positions. In addition to adding individual virtual sound sources, the controller can also add one or more input audio streams to the mix of input audio streams that are rendered as a single virtual sound source of the roster region of the GUI.

[0103] In one aspect, the controller 10 can redefine the UI regions based on whether a remote device is added to or removed from the communication session. For example, in response to a third group of remote devices having joined the session, the controller can redefine the UI regions by at least one of: 1) adding visual representations of input audio streams from the third group to the already defined UI regions, 2) creating one or more new UI regions (e.g., possibly including at least one input video stream of the third group), or 3) a combination thereof. Thus, the controller can dynamically redefine the UI regions as needed. In another aspect, the controller can define the UI regions based on user input of the local device (e.g., the user selects a menu option that, when selected, instructs the controller to define the UI regions as described herein). In some aspects, the controller 10 can switch between defining the UI regions for spatially rendering the input audio streams and providing the canvas region and the roster region to the GUI (e.g., based on whether a predefined number of output audio channels has been exceeded).

[0104] Figure 7 Several stages 70 and 71 are shown in accordance with one aspect, in which the local device 2 defines a user interface (UI) region for one or more visual representations of input video streams and spatially renders one or more input audio streams associated with the defined UI region. Specifically, the first stage 70 shows the communication session GUI 41 when the local user is participating in the session, as well as the positions 62 of the virtual sound sources of the session. Specifically, the figure shows the visual representations 44-48 of the remote participants and the visual representation of the local user 49 in the arrangement 61 in the GUI 41 displayed on the local device display screen 13. Specifically, due to the orientation of the local device, this arrangement is different from the arrangement 51 of Figure 3 For example, the local device shown in Figure 3 is in a portrait orientation, while the device shown in this figure is in a landscape orientation (e.g., rotated 90° about the center axis of the local device). In this arrangement 61, the distribution of the representations within the canvas region 42 within the GUI is different than their distribution in the arrangement 51 of Figure 3The positions in arrangement 51 are wider (e.g., wider along the X-axis). Additionally, list area 43 is displayed on one side (right side) of the GUI, where visual representations are stacked in a column rather than a row. Furthermore, the arrangement 62 of virtual sound source positions is similarly arranged to the arrangement 61 of visual representations. For example, sound source 55 corresponding to visual representation 45 is higher in the vertical direction and positioned between sound sources 54 and 56 corresponding to visual representations 44 and 46, respectively, which are lower in the vertical direction and on either side of representation 45. As described herein, the arrangement 62 of virtual sound sources can be proportionally larger than the arrangement of visual representations (e.g., wider and higher relative to the representations). However, on the other hand, these sound sources can be located on (or near) their respective representations, whereby sound sources 54-56 are centered on their respective visual representations 44-46, and the virtual sound source 53 of the list is positioned at the center of both list representations 47 and 48.

[0105] In one respect, the position of list area 43 in this arrangement can be consistent with that in Figure 3 The same positioning is used in the arrangement 51. For example, the list visual representations are not arranged in a stacked column, but can be arranged in a row, for example, at the bottom of the GUI 41. On the other hand, the visual representations within the list area can be optional, such that remote participants may not be placed in the list area if not needed. For example, when there are enough output audio channels and / or each of these remote participants meets the criteria for being within the canvas area, the controller 10 can position all remote participants within the canvas area. In this case, the GUI may not include the list area.

[0106] Phase 71 illustrates the result of more remote participants joining the communication session, and in response, the controller 10 of the local device defines UI areas, each containing one or more visual representations. Specifically, as shown, three new remote participants 73-75 have joined the communication session. In one aspect, the controller may have determined that one or more of these new participants should be spatially rendered as individual virtual sound sources. However, in another aspect, the controller may have determined that the local device might exceed the optimal (predefined) number for individual rendering if rendered individually. Therefore, the controller has defined four (simulated) areas 76-79 as a grid, each area in which is a UI area. The arrangement 63, including the visual representations of the three newly added remote participants, is also arranged in a grid manner, with two visual representations assigned to each area. Additionally, the sizes of the existing visual representations have been adjusted so that all representations have the same size.

[0107] With the rearrangement of the visual representations, the controller is outputting four different virtual sound sources 65-68 that are in the new arrangement 64. Specifically, the virtual sound sources 65-68 have been arranged in a grid that is similar to the grid of the simulated regions 76-79. In one aspect, the arrangement 64 of virtual sound source locations can be proportional to the arrangement of the UI regions 76-79. In another aspect, the virtual sound source locations can be on (or adjacent to) their respective UI regions. In this case, each virtual sound source can be positioned at the center of its respective region. In one aspect, the arrangement 64 of virtual sound source locations can be static during the communication session, such that the remote participants sharing a virtual sound source (e.g., remote participants 74 and 75 sharing sound source 68) have the same spatial cues when they speak. In another aspect, the virtual sound source locations of the UI regions can dynamically change their locations based on which remote participant within that region is speaking. For example, virtual sound source 68 can move horizontally depending on whether remote participant 74 or 75 is speaking.

[0108] In one aspect, the controller can arrange the visual representations and / or virtual sound sources differently. For example, the controller can define regions of different sizes within the GUI, each region being associated with one or more virtual sound sources that include one or more input audio streams.

[0109] As described above, the controller determines spatial parameters based on the locations of the visual representations within the GUI. In one aspect, the controller can spatially render input audio streams that are not associated with visual representations that are displayed (or visible) within the GUI. Specifically, the video Tenderer can not display one or more visual representations associated with the rendered input audio streams. For example, the video Tenderer can determine that the GUI does not have enough white space to support displaying one or more additional visual representations (e.g., without overcrowding the display). As another example, the communication session can have more candidate remote participants than can be displayed within the roster region. For example, referring to Figure 3 , the roster region includes two visual representations 47 and 48. However, if several additional remote participants are added, the visual representations (at that size) would not fit within the width of the GUI. As another example, the video Tenderer can not receive input audio streams from one or more remote devices. In this case, the spatial parameter generator 25 can determine the location of a virtual sound source that does not include an associated visual representation to be positioned outside (or to one side) of the GUI (and / or to one side of other virtual sound sources). Thus, referring to Figure 3 , when rendering a "non-visible" virtual sound source in the arrangement 50, the sound source can be positioned to the right of the GUI 41.

[0110] As described herein, the controller 10 (e.g., its audio spatial controller 20) is configured to determine spatial parameters indicative of the locations of the virtual sound sources of the one or more input audio streams based on the visual representation displayed in the GUI 41. Such spatial parameters can include panning angles, such as azimuth panning angles and elevation panning angles relative to at least one reference point in the space (e.g., the location of the local user or the user’s head), the distance between the virtual sound source and the reference point, and the reverb level. Figure 8 A panning angle for rendering an input audio stream at the respective virtual sound source location of the input audio stream in a 3D sound field is shown. Specifically, this figure illustrates an arrangement 62 of virtual sound source locations corresponding to the arrangement 61 of visual representations as shown in Figure 7 The arrangement 62 is bounded by the panning ranges of the one or more loudspeakers used to output the virtual sound sources, as shown. Specifically, these boundaries include an azimuth panning range that spans the width of the arrangement along the X-axis, and an elevation panning range that spans the height of the arrangement along the Y-axis. In one aspect, the loudspeakers can produce virtual sound sources at any location within the bounded ranges. In some aspects, the boundaries of the arrangement correspond to the locations of the loudspeakers. For example, these boundaries can correspond to the dimensions of the local device, in which case the local device can be configured to produce virtual sound sources in front of the device (e.g., its display screen). Specifically, the azimuth panning range can span the width of the local device (e.g., its display screen), and the elevation panning range can span the height of the local device (e.g., its display screen). Thus, the locations of the virtual sound sources 54-56 can be located on (or in front of) their corresponding visual representations 44-46, respectively, as described herein. In another aspect, the controller can limit these panning ranges (relative to the wider possible panning ranges) so as to position the virtual sound sources within a particular region of the display screen of the local device (e.g., in front of, to the side of, and / or behind the display screen of the local device).

[0111] Additionally, this figure illustrates the panning angles relative to a reference point 99 in the space (e.g., within the physical environment). For example, the azimuth panning range 100 includes the azimuths of the four virtual sound sources 53-54 (or L1-L4, respectively) relative to (or at) the reference point 99 along the horizontal X-axis. Specifically, the reference point is the vertex of each angle, and each azimuth panning angle extends away from the 0° reference axis Z-axis (e.g., or toward +ω) along the horizontal X-axis. Similarly, the elevation panning range 101 illustrates each of the elevations of the four virtual sound sources relative to the reference point along the vertical Y-axis. Again, the reference point is the vertex of each angle, and each elevation extends away from the 0° reference axis (e.g., or toward or + β) extension. Additionally, the distance (along the Z-axis) between the reference point and each of these virtual sound sources is shown as D L1 - D L4 Thus, when spatially rendered, the virtual sound source corresponding to the spatial parameter will be perceived by the local user as originating at a certain azimuth, elevation, and distance from the local user in order to provide a more stereoscopic (3D) spatial experience.

[0112] Figure 9 is a flowchart of one aspect of a process 80 for determining one or more spatial parameters in accordance with an aspect, the one or more spatial parameters indicating a (3D) location at which an input audio stream is to be spatially rendered as a virtual sound source based on the location of the corresponding visual representation. In one aspect, the spatial parameter generator 25 of the controller 20 can perform at least some of these operations to determine the spatial parameters for each input audio stream of a communication session. This process will be described with reference to Figure 10 to Figure 11B each of which illustrates different examples of how the location of a virtual sound source is mapped to a corresponding visual representation.

[0113] The process 80 begins with the generator 25 selecting a set of communication session parameters for an input audio stream (at block 81). For example, the generator can receive all N sets in a data structure from the session manager and can select the first set. As described herein, the session parameters can include information about the visual representations, such as the size, location, prominence values, and associated VAD parameters of the visual representations. The generator uses the set of communication session parameters to determine the location of the visual representation of the input audio stream associated with the input audio stream (at block 82). For example, the session parameters can include location information (e.g., X, Y coordinates) of the visual representation relative to the GUI and / or relative to the display screen on which the GUI is displayed. The generator determines one or more panning ranges (e.g., azimuth panning ranges, elevation panning ranges, etc.) for one or more loudspeakers (at block 83). In particular, these angular ranges can correspond to the maximum (or minimum) range within which a virtual sound source can be (e.g., optimally) located within the space, such as the azimuth panning range and the elevation panning range The ranges are shown in Figure 8 In one aspect, the panning ranges can be based on the physical location and / or orientation of the loudspeaker (and / or the device in which the loudspeaker is housed). For example, (e.g., when the loudspeaker is integrated within the local device), the panning ranges can be based on the orientation of the device. For example, when the local device is in a portrait orientation (e.g., as shown in Figure 3 the loudspeakers of the device can have a narrow azimuth panning range that spans along the width of the device, while when the local device is in a landscape orientation (e.g., as shown in Figure 7As noted above, the speakers of the device can have a wide azimuthal panning range (e.g., relative to a reference point in space) that spans along the width of the device (e.g., as shown in FIG. 1). Thus, the generator can determine the orientation of the speaker (e.g., the local device that houses it) and determine the azimuthal panning range (e.g., spanning along the horizontal X-axis) and the elevation panning range (e.g., spanning along the vertical, Y-axis) relative to the orientation. For example, upon determining the orientation (e.g., based on IMU data from the IMU 16), the generator can perform a table lookup on a data structure that stores the predefined panning ranges associated with one or more orientations of the device. Figure 12 and Figure 13 More on determining the panning ranges based on the orientation of the device is described in

[0114] For example, the generator can perform one or more methods for determining the spatial parameters of each virtual sound source, which can be based on the physical location of the local user relative to the orientation (or position) of the local device. For example, the generator can use sensor data (e.g., image data captured by the camera 15) to determine the position and / or orientation of the local user (e.g., the head of the local user) relative to the display screen. The generator can determine at least one of the azimuth and elevation from the local user to each visual representation displayed on the GUI of the local device.

[0115] In another aspect, the generator can determine the spatial parameters based on linearly mapping the angles by the position of the visual representation relative to the size of the GUI within the display screen. For example, referring to Figure 10 According to one aspect, the generator can determine the spatial parameters by using one or more functions to map the position of the visual representation to the angles at which the input audio stream is to be rendered. Specifically, the generator can map the position of the virtual sound source based on the position of the visual representation displayed in the GUI relative to the size of the GUI. This diagram shows a GUI 41 (displayed on the display screen of the local device) of a communication session that includes two visual representations 44 and 45, each representing a remote participant with whom the local user is in a session. The GUI has a width X along the X-axis and a height Y along the Y-axis. In one aspect, the size of the GUI can be based on the size of the GUI displayed on (or relative to) the display screen. In one aspect, when the GUI covers the entire display screen, the size of the GUI will be equal to the size of the display screen. Displayed on each visual representation is a simulated center point of the representation, which represents the position (e.g., X, Y coordinates) of the representation within the GUI. For example, the position of the representation 44L1 is (X L1Y L1 ), and indicates that the position of 45L2 is (X L2 ,Y L2 ). In another aspect, other positions can be defined. For example, the simulated point can be positioned over a particular portion of the visual representation, such as a portion showing the mouth of a remote participant (which can be identified using an object recognition algorithm).

[0116] In one aspect, the function used to map the position of the virtual sound source (e.g., its panning angle) to the visual representation is a linear function of the panning angle relative to the dimensions of the GUI. For example, the azimuth function 111 is a linear function of the azimuth panning range -0<0<0+0 relative to the fractional relationship between the X position and the total width of the GUI, X. Thus, the azimuth panning range begins at the left side of the GUI (e.g., where X = 0) and ends at the right side of the GUI. The elevation function 113 is a linear function of the elevation panning range -0<0<0+0 relative to the fractional relationship between the Y position and the total height of the GUI, Y. Thus, the elevation panning range begins at the bottom of the GUI (e.g., where Y = 0) and ends at the top of the GUI. These relationships between the panning ranges and the dimensions of the GUI allow the generator to map the position relative to the GUI regardless of the size of the GUI and / or the size of the panning range.

[0117] To determine the panning angles of the visual representation, the generator can apply the fractional relationship of the position of the visual representation as input into one or both of the linear functions. For example, to determine the azimuth of the virtual sound source, the generator can use the x-coordinate of the position of the visual representation within the GUI as input to the azimuth panning range function. Specifically, as shown, the fractional position relationship of the visual representation (X L1 / X and X L2 / X) is mapped to an azimuth panning angle at X L1 / X and X L2 / X that intersects the linear function, as shown in FIG. 111. The resulting mapping of these fractional relationships to azimuth is shown by the azimuth panning range 112, which shows the azimuth of LI at the reference point 99 to be -0 L1 , and the azimuth of L2 to be +0 L2 . Similarly, to determine the elevation panning angle of the virtual sound source, the generator can use the y-coordinate of the position of the visual representation as input to the (e.g., separate) elevation panning range function. Specifically, the fractional position relationship of the visual representation (Y L1 / Y and Y L2 / Y) is mapped to an elevation panning angle at Y L1 / Y and Y L2 ​The elevation angle translation angle at / Y intersects with the linear function 113. The resulting mapping of these fractional relationships to the elevation angle is shown by the elevation angle translation range 114, which shows the elevation angle L1 at the reference point. L2 elevation angle + β L2 Side view.

[0118] On the other hand, spatial parameters can be determined based on the viewing angle of the local user's (predefined) location (e.g., in the case where the reference point in the user or space is a vertex relative to its defined angle). Figure 11A and Figure 11B This illustrates an example of determining spatial parameters based on one aspect by using one or more functions to map the viewing angle position of a visual representation to the translation angle of the input video stream to be rendered. Specifically, the generator can map the position of the visual representation as the viewing angle at a reference point to one or more translation angles within one or more translation ranges. (Reference) Figure 11A This figure illustrates GUI 41, which includes two visual representations 44 and 45, as shown. Figure 9 As shown in the diagram. However, here, instead of determining a fractional relationship between the position of the visual representation and the total width / height of the GUI, the generator determines an estimated viewing angle of the GUI relative to a reference point 69 (e.g., at a predefined location in space). For example, the GUI has an estimated azimuth viewing range 115 between -θ' and +ω', spanning the width W' of the GUI and along the X-axis. This estimated azimuth viewing range is shown as a top view, where the reference point 69 is located in front of the GUI (or display screen) at a distance D' from the GUI (or display screen), where the GUI has a width of W'. In one aspect, this estimated azimuth viewing range is a predefined viewing range for the local user when the user is looking at the GUI while it is displayed on the display screen. In some aspects, the positions of W', D', and / or the reference point can be predefined; for example, D' could be the optimal viewing position for the local user. In other aspects, these dimensions can be determined by the generator. For example, the generator can obtain sensor data from one or more sensors (e.g., image data from camera 15, proximity sensor data, etc.) and determine the distance the local user is positioned relative to this sensor data. The generator can determine the width of the GUI currently displayed on the screen, relative to its width. Knowing the positions of D', W', and the reference point, the generator determines the azimuth viewing angle of L1 as –θ'. L1 The azimuth viewing angle of L2 is +ω' L2 These angles are the angles from their corresponding visual representations on the GUI to the reference point.

[0119] To determine the (actual) azimuthal pan angles, the generator can apply the viewing angles as input to one or more linear functions. For example, the figure shows an azimuthal function 116, which is a linear function of the azimuthal pan range -θ - +ω versus the estimated azimuthal viewing range -θ' - +ω'. The generator maps the viewing angles -θ L1 and +ω' L2 to the actual azimuthal angles that intersect the linear function. The mapping of these angles is shown by the azimuthal pan range 117, which shows -θ L1 and +ω L2 at the reference point 99. In one aspect, the reference point 99 can be the same as the reference point 69 (e.g., at the same location in space relative to the local device).

[0120] Referring Figure 11B to FIG. 6, this figure is concerned with determining the elevation pan angles based on the estimated elevation viewing angles relative to the reference point 69. Specifically, the controller can perform similar operations as described in Figure 11A to determine the elevation pan angles. For example, as shown, this estimated elevation viewing range 118 is between and spans the height H' of the GUI and along the Y-axis. Specifically, this viewing range is a side view of the GUI (or display screen) and the reference point 69 at the (predefined) distance D'. In one aspect, H' can be a predefined height of the GUI, or can be the current height of the GUI relative to the display screen. From this reference point, the generator determines the elevation viewing angle of L1 to be +β'L1 and the elevation viewing angle of L2 to be To determine the actual elevations, the generator applies the elevation viewing angles as input to an elevation function 119 of the elevation pan range versus the elevation viewing range . The generator maps the viewing angles +β'L1 and to the actual elevations that intersect the function 119. The mapping of these angles is shown by the elevation pan range 120. In one aspect, the generator can perform any of the methods described in Figure 10 to Figure 11B to map the position of the visual representation to the position of the virtual sound source.

[0121] Additionally, the generator can determine other spatial parameters based on the communication session parameters, such as a distance between the virtual sound source and the reference point. As described herein, the distance between the virtual sound source and the local user can be based on the size and / or position of the visual representation. For example, similar to a physical conversation in which a closer person sounds louder than a farther person, the generator can assign a smaller visual representation a shorter distance than a farther visual representation. In one aspect, the distance can be based on the position of the visual representation within the GUI. For example, a visual representation that is higher within the canvas region (e.g., along the Y-axis) can be given a shorter distance than a visual representation that is farther below the Y-axis. In some aspects, the visual representation within the roster can be assigned a farthest distance relative to all canvas visual representations. In another aspect, the distance can also be based on the VAD parameters. For example, a remote participant associated with a VAD parameter indicating a high signal energy level (e.g., above a threshold) can be assigned a closer distance than a remote participant with a lower VAD parameter. In some aspects, the generator can define a reverb value for each of the input audio streams based on the same criteria described above. For example, a remote participant within the roster region can be assigned a high reverb value to make the sound more diffuse to the local user.

[0122] As described herein, the controller can apply one or more linear functions to determine the pan angle. In some aspects, one or more of the functions can be a more general non-linear function or a piecewise linear function of the pan angle (e.g., relative to a fractional relationship and / or a viewing pan angle, as described herein).

[0123] Returning to Figure 9The controller determines whether there are any additional sets of communication session parameters for which one or more spatial parameters have not yet been determined (at decision box 85). If so, the controller selects another set of communication session parameters for another input audio stream that has not yet been analyzed to determine one or more spatial parameters. Otherwise, the controller determines whether any input audio stream is to be output as a blend (at decision box 87). For example, the controller may determine whether any input audio stream is associated with a remote participant in the list area. In some aspects, this determination may be based on whether selector 27 has assigned one or more input audio streams to a single output audio stream (e.g., this could be when there are no more individual output audio streams for the canvas remote participant to have individual virtual sound sources). In other aspects, this determination may be based on the output audio stream to which the input audio stream has been assigned, as described herein. If so, the controller determines new (or different) spatial parameters for the input audio stream to be blended, indicating a specific location, based on at least some of the spatial parameters of the input audio stream to be blended (at box 87). As described herein, the controller spatially renders the blend to output a single virtual sound source that includes the blend. Therefore, the controller determines a set of spatial parameters for a single virtual sound, which can be based on at least some of the determined spatial parameters. For example, the new spatial parameters can be a weighted combination of at least some of the spatial parameters. In this case, a single virtual sound source can be located at the center of an analog virtual sound source, generating analog virtual sound sources that are mapped to multiple locations based on corresponding visual representations within the list area if each of these input audio streams is to be spatially rendered as an individual virtual audio stream. On the other hand, the determined spatial parameters can be based on a set of predetermined data, rather than new spatial parameters. On the other hand, the spatial parameters can not be based on predetermined data, but rather the spatial parameters can be associated with a specific location in space. For example, the spatial parameters can indicate azimuth and elevation angles of 0°, such that the spatial audio of the list is located exactly in front of the local user.

[0124] In some respects, the controller can determine the spatial parameters of the mixed input audio streams differently. For example, instead of determining the spatial parameters relative to individually determined spatial parameters, the controller can combine communication session parameters (e.g., position / size, distance, VAD parameters, salience values, etc.) for at least some of the mixed input audio streams. For example, the controller can determine the average of at least some of these parameters. Once the combined (or jointed) communication session parameters are determined, the controller can determine the spatial parameters as described herein.

[0125] Once the controller determines the spatial parameters, the input audio streams are spatialized according to the data to output one or more virtual sound sources, each virtual sound source comprising one or more input audio streams, as described herein.

[0126] In one aspect, the position of the virtual sound source can change based on one or more criteria. For example, as described herein, to spatially render the input audio streams, spatial parameters are determined that indicate the position of the resulting virtual sound source. As described herein, the determined spatial parameters can depend on the panning range of the local device. Thus, when the panning range changes, the local device can adjust the virtual sound source that is currently being output to accommodate the change. For example, the panning range can change due to the local device changing orientation. For another example, the panning range can be based on the aspect ratio of the GUI of the communication session displayed on the display screen of the local device. The following figures describe adjusting the virtual sound source based on changes to the panning range of the local device.

[0127] Figure 12 Several stages 130 and 131 are shown in which the panning range of one or more loudspeakers is adjusted based on the local device rotating from a portrait orientation to a landscape orientation. For example, each stage shows the GUI 41 of the local device 2 while engaged in a communication session, and a corresponding arrangement of virtual sound source positions, where each arrangement shows several panning angle ranges. Specifically, each stage shows an azimuthal panning range -θ - +ω along the X-axis, and an elevation panning range As described herein, one or more of these ranges can change based on the orientation of the device.

[0128] The first stage 130 shows the local device oriented in a portrait orientation, where the height along the Y-axis is greater than the width along the X-axis. Also shown is an arrangement 50 of virtual sound source positions, which shows four virtual sound sources 53-56. Specifically, as described herein, for each of the sound sources 54-56, the local device outputs an input audio stream as a virtual sound source at a position within the arrangement 50 of positions relative to a reference point outside of the local device (e.g., a point at which the local user is located, or a predefined point as described herein). In addition, the local device is outputting a mix of the input audio streams as a single virtual sound source 53. In addition to showing the positions of the virtual sound sources, the arrangement 50 also shows the panning range of the local device when in this portrait orientation. Specifically, the azimuthal panning range is -θ P - +ω P and the elevation panning range is

[0129] The second stage 131 illustrates the result of the local device being rotated 90° about the Z-axis. In particular, the local device has been rotated to a landscape orientation, in which the width along the X-axis is greater than the height along the Y-axis. Additionally, the panning range of the local device has also changed. As shown, the azimuth panning range is -θ L - + ω L , - θ P - + ω P is wider (e.g., has a greater range), and the elevation panning range is , - θ is narrower (e.g., has a reduced range). In one aspect, the change in panning range can be based on the components or design of the local device. For example, the panning range can be defined based on the number and / or location of the speakers of the local device. The panning range can also be rotated when the device is rotated to a new orientation. In another aspect, the panning range can be defined by the controller and can be adjusted by the controller in response to determining that the orientation of the local device has changed. For example, upon determining that the device is now in a landscape orientation, the controller can determine the panning range for this orientation (e.g., by performing a table lookup on a data structure that associates one or more panning ranges with orientations), and then use the determined panning range for spatial rendering, as described herein.

[0130] Additionally, in response to the orientation of the local device changing to a new landscape orientation, one or more positions of virtual sound sources are adjusted relative to the reference point along one or more axes. In particular, due to the rotation of the local device, the virtual sound source positions are in the arrangement 62. As described herein, the positions of the virtual sound sources have been adjusted such that they are distributed more widely along the X-axis and more narrowly along the Y-axis compared to the virtual sound sources that were in the arrangement 50 when the local device was in a portrait orientation. Thus, the local user can perceive the virtual sound sources differently based on the orientation of the local device.

[0131] Figure 13 is a flowchart of one aspect of a process 140 for adjusting positions of one or more virtual sound sources based on a change in one or more panning ranges of one or more speakers, the change being based on a change in an orientation of a local device. In one aspect, this process will be described with reference to Figure 12 The process 140 begins with the controller 10 of the local device 2 receiving one or more input audio streams (and one or more input video streams) from one or more remote devices with which the local device is in a (video) communication session (at block 141). The controller determines a first orientation of the local device (at block 142). For example, the controller 10 can determine the orientation of the device based on sensor data, such as IMU data from the IMU 16. In one aspect, the IMU data can indicate that the local device is in a portrait orientation, as shown in the arrangement 50 of FIG. 5A. Figure 12As shown in the first phase 130, the controller determines one or more translation ranges (at box 143) for one or more speakers of a local device (communically coupled to or part of it). As described herein, the controller may perform a table lookup on a data structure that associates the orientation and / or position of the speakers and / or the local device with one or more translation ranges. In response, the controller may determine an azimuth (e.g., horizontal) translation range of -θ when the device is in a longitudinal orientation. P -+ω P Determine the range of translation angle (e.g., vertical) as follows: The controller determines one or more spatial parameters for each input audio stream, which indicate the position of the virtual sound source within one or more determined translation ranges including the input audio stream, and spatially renders these streams as virtual sound sources at one or more positions within the determined translation ranges (in box 144). For example, the controller can perform... Figure 9 At least some of the operations described in process 80 are used to render the input audio stream space as one or more individual virtual sound sources, and / or to render the mixing space of one or more streams as a single virtual sound source.

[0132] The controller determines whether the local device has a changed orientation (e.g., at decision box 145). Specifically, the controller may determine whether the orientation has changed (e.g., from longitudinal orientation to lateral orientation) based on IMU data, as described herein. If so, the controller determines one or more adjusted translation ranges for the one or more speakers based on the changed orientation (at box 146). For example, re-referencing... Figure 12 The adjusted translation range corresponds to the local device in a laterally oriented position. The controller then adjusts one or more positions of the virtual sound source (at box 147) based on this adjusted translation range. For example, the controller can adjust the azimuth and / or elevation angle of the virtual sound source relative to a reference point. For example, as... Figure 12 As shown, in response to the local device rotating to landscape orientation, the azimuth angle of the virtual sound source 54 is closer to the lower boundary -θ from the azimuth angle position when the local device is in portrait orientation. L Widening. In one respect, to adjust the position, the controller can perform the operation of process 80 to adjust the position. For example, in response to performing the operation based on the local device rotating to a lateral orientation, Figure 12 The virtual sound source position 62 in the diagram spans a wider range along the azimuth angle and a narrower range along the elevation angle. On the other hand, the controller can adjust the position without determining new spatial parameters, as described in process 80. For example, the controller can adjust the virtual sound source position by rotating the position based on the rotation of the device. For instance, the controller can rotate the virtual sound source position by 90° in response to determining that the local device has rotated from longitudinal to lateral.

[0133] In another aspect, the controller can adjust the position based on a difference in the pan angle between the two or more orientations. For example, the controller can adjust the position in proportion to a difference between the pan angle of the first orientation and the pan angle of the second orientation. For example, referring to Figure 12 , the azimuth pan angle can increase by 50% and the elevation angle can decrease by 50% from portrait to landscape. Thus, to adjust the position, the azimuth and elevation angles can change in proportion to the difference between their respective pan ranges of the two orientations.

[0134] In one aspect, as the local device rotates, the position of the virtual sound source can remain in its same position relative to the GUI (e.g., on the display screen). For example, referring to Figure 12 , although the position of the virtual sound source has changed relative to the pan range, the position of the virtual sound source can remain the same relative to the rotated local device due to the range increase / decrease, such that the virtual sound source rotates with the local device. Thus, as the device rotates, the local user can perceive the virtual sound source to maintain its position relative to the display screen, such that both the visual representation and its associated virtual sound source move together.

[0135] In another aspect, once the local device rotates in the opposite direction, the virtual sound source can move back to its initial position. For example, if the local user rotates the local device -90°, the virtual sound source can return to its initial position, as shown in the first stage 130 of Figure 12 .

[0136] In one aspect, the pan range can be attached (or correspond) to the edges of the display screen, and then applied to the GUI window within the screen by considering the maximum zoomed version of the GUI window that fills the screen as much as possible (e.g., where at least two opposite edges of the GUI window coincide or are adjacent to the corresponding edges of the display screen). Such a system has the advantage that the pan range is primarily a function of the aspect ratio of the GUI window, and not a function of the GUI size, thus preserving the audio spatial image that does not change with the GUI size, position, and / or location or folding if the window is minimized or if the window enters a picture-in-picture mode. Figure 14A and Figure 14B illustrate several stages in which the pan range is based on the aspect ratio of the communication session GUI, according to some aspects.

[0137] In particular, as described herein, the local device can define the pan range of the one or more speakers based on the aspect ratio of the GUI. For example, Figure 14ATwo stages 150 and 151 are shown, in which the azimuthal panning range is less than the total (possible) azimuthal panning range of the local device loudspeakers based on the aspect ratio of the GUI 41. The first stage 150 shows a communication session GUI displayed on the display screen 13 of the local device (when the local device is in communication session with five remote participants, as described herein). Specifically, the GUI is overlaid on top of a home screen GUI 152 displayed on the display screen, and provides the local user with an interface to execute and / or terminate one or more computer program applications, which can include a communication session application, as described herein. As shown, the communication session GUI is smaller (or has a smaller surface area) than the home screen GUI. In one aspect, the local device can receive input (e.g., via any input device such as a mouse, or via the display screen, which can be a touch-sensitive display screen, as described herein) to adjust the size and / or position of the GUI 41 within the home screen GUI 152. As shown, the GUI 41 has a current aspect ratio of 4:3.

[0138] Additionally, this stage also shows the (e.g., azimuthal and elevation) panning range of the loudspeakers of the local device. As shown, the total panning range (e.g., the maximum angle at which a virtual sound source can be positioned when spatially rendering a corresponding input audio stream using the loudspeakers of the local device) extends to the edges of the display screen. For example, the azimuthal panning range - θ - + ω of the loudspeakers spans the total width of the display screen 13 (along the X-axis), and the elevation panning range spans the total height of the display screen 13 (along the Y-axis). Thus, the local device can position a virtual sound source at any location on (or in front of) the display screen.

[0139] The second stage 151 shows the result of zooming in on the simulated communication session GUI until the two sides of the GUI hit the respective edges of the display screen. As shown in this figure, the display screen is displaying a simulated GUI 153 that has been fully zoomed in in the Y-direction (e.g., the visible portion of the session GUI cannot be expanded any further in the Y-direction), while the edges along the width of the GUI are separated from the edges of the display screen. Since the height of the session GUI extends the height of the display screen, the elevation panning angle remains unchanged, however, since the width of the session GUI is less than the width of the display screen, the azimuthal panning range is reduced to - θ w - + ω w which is less than the total azimuthal panning range. Thus, the controller 10 of the local device can limit the azimuthal and elevation panning ranges accordingly.

[0140] Figure 14B Two stages 160 and 161 are shown, in which the elevation panning range is less than the total elevation panning range of the local device loudspeakers based on the aspect ratio of the GUI 41. This stage is similar to Figure 14Athe second stage 151, the second stage 161 of this figure shows that the simulated GUI 154 has fully expanded along the width of the display screen 13, but has not yet fully expanded along the height of the display screen. Thus, when the GUI 41 has a larger aspect ratio, the defined azimuthal panning range can equal the total panning range, while the elevation panning range is reduced to which is less than the total elevation panning range.

[0141] As noted above, the local device (whose controller 10) can limit the panning range based on whether the GUI is a 4:3 or 16:9 aspect ratio. In another aspect, the panning range can be limited for any aspect ratio. In one aspect, the controller can determine and / or adjust the spatial parameters of one or more input audio streams based on the aspect ratio or whether the aspect ratio has changed (e.g., in response to a user input). Figure 15 More is described in this regard.

[0142] Figure 15 is a flowchart of one aspect of a process 170 for adjusting the location of one or more virtual sound sources based on a change in the aspect ratio of a GUI for a communication session. The process 170 begins with the controller 10 receiving input audio streams and input video streams from a remote device in a video communication session with the local device (at block 171). In one aspect, the controller can also receive other data (e.g., a VAD signal), as described herein. The controller displays a visual representation of the input video streams within a GUI (having an aspect ratio) for the video communication session displayed on a display screen (at block 172). The controller determines the aspect ratio of the GUI for the video communication session (at block 173). In particular, the video renderer 22 can determine the aspect ratio in which the GUI is being displayed. The controller determines an azimuthal panning range that is at least a portion of a total azimuthal panning range, and an elevation panning range that is at least a portion of a total elevation panning range for the speakers based on the aspect ratio of the GUI (at block 174). In particular, the controller can perform the operations described in Figure 14A and Figure 14B For example, the controller (e.g., its video renderer) can simulate expanding (or expanding) the GUI until at least one of the width of the GUI has fully expanded to the width of the display screen or the height of the GUI has fully expanded to the height of the display screen, while maintaining the aspect ratio. In one aspect, this can mean expanding the GUI until at least two sides (or edges) of the GUI are in contact with two comparable sides of the display screen. For example, the renderer can expand the GUI until the top and bottom edges of the GUI are in contact with the edges of the display screen (as shown in Figure 14A ), or can expand the GUI until the side edges of the GUI are in contact with the edges of the display screen (as shown in Figure 14BOnce zoomed in, the controller can limit the panning range based on the size of the simulated GUI relative to the size of the display screen. Specifically, the panning range is limited to span the width and height of the GUI as a function of the width and height of the display screen. In one aspect, if the GUI has the same aspect ratio as the display screen, the determined panning range can be the total panning range of the loudspeakers.

[0143] As described herein, when determining the panning range while zooming in on the GUI, the range can span the width and height of the display screen. In another aspect, these panning ranges can extend beyond the boundaries of the display screen. In this case, the panning range can be determined based on a percentage of the zoomed-in simulated GUI. For example, with reference to Figure 14B , the determined azimuth panning range can be 100% of the total azimuth panning range because the simulated GUI has expanded to the side edges of the display screen. In contrast, the determined elevation panning range can be 70% of the total elevation panning range because the simulated GUI has only expanded along 70% of the total height of the display screen.

[0144] The controller determines spatial parameters indicating a location of the virtual sound source within the determined azimuth and elevation panning ranges based on the location of the visual representation within the GUI, as described herein (at block 175). The controller then spatially renders the input audio stream using the loudspeakers based on these spatial parameters to output the virtual sound source at that location within the azimuth and elevation panning ranges (e.g., at that location) (at block 176). Thus, along with displaying the visual representation, the controller outputs the input audio signal as a virtual sound source at a location within the environment (e.g., a location at which the local device is located).

[0145] The controller determines whether the aspect ratio of the GUI has changed (at decision box 177). In one aspect, this determination may be based on whether user input has been received via one or more input devices to change the width or height of the GUI. For example, the controller may receive an indication that the user has performed a click-drag operation with the mouse to manually zoom in (or stretch) the GUI in one or more directions (e.g., by selecting one side and performing a dragging motion away from or towards the GUI) (e.g., via video renderer 22). In another aspect, when the display is a touch-sensitive display, user input may be received when the user performs a touch-drag motion with one or more fingers to resize the GUI. In some aspects, the controller may (e.g., periodically) perform the operations described in box 173 to determine whether the aspect ratio has changed. In response to determining that the aspect ratio has changed, the controller adjusts one or more translation ranges based on the changed aspect ratio (at box 178). For example, the controller may perform the operations described in box 174 to determine the changed (or new) azimuth translation range. For example, when the aspect ratio increases, the adjusted azimuth translation range can extend the width of the display screen, while the adjusted elevation translation range will not fully extend the height of the display screen. Figure 14B As shown in the diagram. The controller then adjusts the position of the virtual sound source based on the adjusted translation range (at box 179). For example, the controller can determine the spatial parameters according to any of the methods described herein, and then spatially render the input audio signal according to the new spatial parameters.

[0146] On the other hand, the controller can adjust its position based on changes in the adjusted translation range. Specifically, the controller does not recalculate the position of the virtual sound source (e.g., as...). Figure 9 (As described above), but the existing position can be adjusted based on the adjusted translation range. For example, by increasing the aspect ratio (e.g., from 4:3 to 16:9), the azimuth translation range can be increased from 60% to 100% of the total azimuth translation range, and the elevation translation range can be decreased from 100% to 70%, as... Figure 14A and Figure 14B As shown in the diagram. In response, the controller can adjust the azimuth angle of the virtual sound source by increasing the angle by 40% and the elevation angle of the virtual sound source by decreasing the angle by 30%. Specifically, the controller can increase the azimuth angle to move the position of the virtual sound source to a wider azimuth angle position and decrease the elevation angle to move the position to a narrower elevation angle position (e.g., relative to the 0° reference Z-axis).

[0147] In one respect, the panning range is determined based on the aspect ratio of the GUI, providing consistent spatial audio for the local listener regardless of the GUI's position relative to the display screen. For example, by defining the panning range using aspect ratio, the position of the virtual sound source is independent of the orientation, location, and / or size of the GUI displayed on the screen. Therefore, the local user can move the GUI throughout the display screen during a communication session without adversely affecting the spatial cues of the remote participant. Furthermore, spatial cues (e.g., the position of the virtual sound source) are also independent of the display screen's position and / or orientation, which may differ between different users.

[0148] In one aspect, these panning ranges can extend beyond the edges of the display screen. In this case, the panning range can be a function of the size of the analog magnified GUI relative to the size of the display screen (or the area of ​​the display screen showing video data). Therefore, when at a lower aspect ratio, such as in Figure 14A As shown, the controller can utilize the full omnidirectional translation range to spatially render virtual sound sources, while using only a portion of the full elevation translation range (e.g., it may be based on the difference between the height of the magnified GUI and the height of the area that the display can show, as described herein).

[0149] Some aspects can be discussed separately in Figure 5 , Figure 6 , Figure 9 , Figure 13 and Figure 15 The processes 30, 90, 80, 140, and 170 described herein may be modified. For example, at least some of the specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects.

[0150] As described above, one or more speakers that output one or more virtual sound sources can be arranged to output sound to the surrounding environment, such as external speakers that can be integrated into local devices, displays, or any electronic devices, as described herein. On the other hand, the speakers can be part of headphones, such as… Figure 1 The headset 6. In this case, the controller can perform similar operations as described herein to determine spatial parameters and use data to spatially render one or more input audio streams. However, in one aspect, the controller can differently define one or more translation ranges. For example, the azimuth translation range and the elevation translation range can extend 360° around the local user. Therefore, the virtual sound source can be positioned anywhere within the sound field. On the other hand, the external speaker of System 1 can have a similar translation range.

[0151] As described above, controller 10 determines various parameters and data for spatial rendering of the input audio streams, such as: one or more azimuth translation ranges and one or more elevation translation ranges (e.g., a set of ranges when the local device is in longitudinal orientation and a set of ranges when the local device is in lateral orientation), one or more translation angles for each input audio stream, distances (e.g., between the local user and the virtual sound source, between the local user and the display screen, etc.), reverberation, device orientation, GUI dimensions (e.g., size, shape, positioning, and aspect ratio), and display screen dimensions (e.g., width, height, and aspect ratio). On the other hand, this data may also include predefined data, such as predefined dimensions of the GUI, predefined distances between the display screen and reference points, etc. In one aspect, the local user can change any of these parameters or values. For example, the local device may display a menu (e.g., based on user selections within the UI of a communication session GUI). Once displayed, the user can adjust any of these parameters or values. For example, the user can adjust the translation range based on a specific situation. Specifically, the user may want to reduce the translation range defined by the display screen, rather than extending it beyond the display screen (e.g., to minimize sound leakage within the acoustic environment).

[0152] As is widely recognized, the use of personally identifiable information should comply with privacy policies and practices that are generally accepted to meet or exceed industry or governmental requirements for protecting user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly explained to users.

[0153] In one aspect, the size of at least one visual representation of a respective input video stream associated with an input audio stream that is rendered separately is greater than the size of a visual representation of an input video stream associated with a mixed input audio stream of input audio streams that are rendered together, where all visual representations of respective input video streams associated with all input audio streams that are rendered together as a mix have the same size. In other words, the roster visual representations can all have the same size. In one aspect, the arrangement of locations of individual virtual sound sources is in front of a display screen of the local device based on the communication session parameters. In some aspects, the arrangement of locations of individual virtual sound sources includes determining, for each individual virtual sound source, a location within the GUI of a visual representation of a respective input video stream associated with a respective input audio stream to be spatially rendered as the individual virtual sound source using the set of communication session parameters, and determining one or more spatial parameters indicative of a location of the individual virtual sound source within the arrangement based on the determined location of the visual representation, where spatially rendering the input audio stream as the individual virtual sound source includes spatially rendering the input audio stream as the individual virtual sound source at the location using the determined spatial data. In one aspect, the spatial parameters include an azimuth angle along a first axis between the location of the individual virtual sound source and a second axis, and an elevation angle along a third axis between the location of the individual virtual sound source and the second axis. In some aspects, the arrangement further includes a location of a single virtual sound source, where determining the locations further includes, for each input audio stream of the mix, determining a location within the GUI of a visual representation of a respective input video stream associated with the input audio stream of the mix using the set of communication session parameters, determining spatial parameters indicative of a location of a virtual sound source of the input audio stream based on the determined location, and determining new spatial parameters indicative of a particular location based on at least some of the spatial parameters, where spatially rendering the mix of input audio streams includes spatially rendering the mix of input audio streams as a single virtual sound source at the particular location using the new spatial parameters. In some aspects, the new spatial parameters are determined by determining a weighted combination of the spatial parameters of all input audio streams of the mix, where the particular location is different from the location of the virtual sound source of the input audio streams of the mix. In one aspect, the visual representation of the input video stream associated with the input audio stream of the mix is arranged in a row or column based on an orientation of the local device, where the different location is at a center of the row or column on the display screen.

[0154] According to one aspect of the disclosure, the local device 2 (e.g., its controller 10) can perform a method comprising one or more operations, such as receiving, for each remote device of a first plurality of remote devices engaged in a video communication session with the local device, an input audio stream and an input video stream; for each input video stream, displaying a visual representation of the input video stream in a graphical user interface (GUI) on a display screen; for at least one input audio stream, spatially rendering the input audio stream to output, via a plurality of speakers, only an individual virtual sound source of the input audio stream; in response to determining that a second plurality of remote devices have joined the video communication session, receiving, for each remote device of the second plurality of remote devices, an input audio stream and an input video stream; determining whether the local device supports additional individual virtual sound sources for one or more input audio streams of the second plurality of remote devices; in response to determining that the local device does not support additional individual virtual sound sources that are confined to a plurality of user interface (UI) zones located in the GUI, each UI zone comprising one or more visual representations of one or more input video streams of the first plurality of remote devices, the second plurality of remote devices, or a combination thereof, displayed in the UI zone; and for each UI zone, spatially rendering, via the plurality of speakers, a mix of one or more input audio streams associated with the one or more visual representations included within the UI zone as a virtual sound source.

[0155] In one aspect, the local device is configured to spatially render a predefined number of input audio streams as individual virtual sound sources, where determining whether the local device supports additional individual virtual sound sources comprises determining whether a number of input audio streams of the first plurality of remote devices and the second plurality of remote devices determined to be spatially rendered individually is greater than the predefined number. In another aspect, a number of the confined UI zones does not exceed the predefined number of input audio streams that can be spatially rendered as individual virtual sound sources. In one aspect, confining the plurality of UI zones comprises: displaying all visual representations associated with the first plurality of remote devices and the second plurality of remote devices in a grid manner; and establishing a virtual grid of UI zones on the GUI, where each UI zone contains one or more visual representations. In some aspects, establishing the virtual grid of UI zones comprises assigning one or more adjacent visual representations to each UI zone.

[0156] In one aspect, defining the plurality of UI zones includes determining a number of input audio streams received from the first plurality of remote devices and the second plurality of remote devices, wherein the number of UI zones is defined based on the number of input audio streams. In one aspect, the input audio streams of the first plurality of remote devices and the second plurality of remote devices are evenly distributed among the plurality of UI zones. In some aspects, the method further includes: in response to determining that a third plurality of remote devices has joined the video communication session, receiving an input audio stream and an input video stream for each remote device of the third plurality of remote devices; and redefining the plurality of UI zones by: 1) adding a visual representation of the input audio stream from the third plurality of remote devices to a defined UI zone, 2) creating one or more new UI zones, or 3) a combination thereof. In one aspect, for each UI zone, a respective UI zone virtual sound source is located at a position on the UI zone displayed on the display screen. In some aspects, the position is at a center of the UI zone. In another aspect, the position is on a visual representation associated with a mixed input audio stream having a signal energy level above a threshold value.

[0157] According to one aspect of the disclosure, the local device 2 (controller 10 thereof) can perform a method comprising one or more operations, such as receiving an input audio stream from a remote device in a communication session with the local device; determining a first orientation of the local device; determining a panning range of the plurality of speakers for the first orientation of the local device along a horizontal axis; spatially rendering the input audio stream as a virtual sound source at a position along the horizontal axis and within the panning range using the plurality of speakers; in response to determining that the local device is in a second orientation, determining an adjusted panning range of the plurality of speakers that spans wider along the horizontal axis than the panning range; and adjusting the position of the virtual sound source along the horizontal axis based on the adjusted panning range.

[0158] In one aspect, the first orientation is a portrait orientation of the local device and the second orientation is a landscape orientation. In another aspect, the position of the virtual sound source is proportionally adjusted along the horizontal axis relative to the adjusted panning range. In one aspect, the panning range is a horizontal panning range and the adjusted panning range is an adjusted horizontal panning range, wherein the method further includes, while the local device is oriented in the first orientation, determining a vertical panning range of the plurality of speakers that spans along a vertical axis along which the position of the virtual sound source lies; in response to determining that the local device has been oriented in the second orientation, determining an adjusted vertical panning range of the plurality of speakers that spans along the vertical axis that is less than the vertical panning range; and adjusting the position of the virtual sound source along the vertical axis based on the adjusted vertical panning range and adjusting the position of the virtual sound source along the horizontal axis based on the adjusted horizontal panning range.

[0159] In one aspect, the method further includes receiving an input audio stream and an input video stream from a remote device for display as a visual representation in a graphical user interface (GUI) on a display screen of the local device. In some aspects, when in the first orientation, a position along a horizontal axis at which the virtual sound source is located is the same position at which the visual representation is displayed relative to the display screen; and in response to determining that the local device has been oriented into the second orientation, maintaining the position of the visual representation relative to the display screen while adjusting the position of the virtual sound source such that the virtual sound source and the visual representation remain in the same position relative to the display screen. In one aspect, the method further includes receiving individual input audio streams from each of a plurality of remote devices engaged in a communication session with the local device; and spatially rendering a mix of the individual input audio streams as a single virtual sound source that contains a mix of the individual input audio streams. In another aspect, the method further includes receiving a plurality of input video streams, each input video stream from a different remote device of the plurality of remote devices; displaying a plurality of visual representations as a row along a horizontal axis within a graphical user interface (GUI) on a display screen of the local device when the local device is in the first orientation, each visual representation for a different input video stream of the plurality of input video streams, with the single virtual sound source rendered at a position of one of the visual representations. In one aspect, the single virtual sound source is rendered at the position of one of the visual representations in response to an individual input audio stream associated with an input video stream displayed in one of the visual representations having an energy level that is greater than a remaining portion of the individual input audio streams in the mix of the individual input audio streams. In some aspects, the position at which the single virtual sound source is rendered changes along the horizontal axis, but not along a vertical axis based on which individual audio stream associated with a respective visual representation of the row has a greater energy level. In another aspect, the method further includes displaying the plurality of visual representations as a column along a vertical axis within the GUI on the display screen of the local device in response to determining that the local device has been oriented into the second orientation; and adjusting the single virtual sound source to continue to render the signal virtual sound source at a position of one of the visual representations within the column. In some aspects, the position at which the signal virtual sound source is rendered changes along the vertical axis, but not along the horizontal axis based on which individual audio stream associated with a respective visual representation of the column has a greater energy level than a remaining portion of the individual audio streams.

[0160] In one aspect, prior to determining that the local device has been oriented in the second orientation, the method further includes determining spatial data for spatially rendering the input audio stream, the spatial data indicating a position of the individual virtual sound source as an angle between a reference point and a position along a horizontal axis. In another aspect, adjusting the position of the virtual sound source includes determining adjusted spatial data indicating an adjusted position as an adjusted angle between the reference point and the adjusted position along the horizontal axis and within an adjusted range of translation; and spatially rendering the input audio stream as the virtual sound source at the adjusted position using the adjusted spatial data.

[0161] According to another aspect, the local device 2 (controller 10 thereof) can perform a method including one or more operations, such as receiving an input audio stream and an input video stream from a remote device in a video communication session with the local device; displaying a visual representation of the input video stream within a graphical user interface (GUI) of the video communication session displayed on a display screen; determining an aspect ratio of the GUI of the video communication session; determining an azimuthal range of translation and an elevation range of translation based on the aspect ratio of the GUI of the video communication session, the azimuthal range of translation being at least a portion of a total azimuthal range of translation of a plurality of speakers, the elevation range of translation being at least a portion of a total elevation range of translation of the plurality of speakers; and spatially rendering the input audio stream using the plurality of speakers to output a virtual sound source of the input audio stream included within the azimuthal and elevation ranges of translation.

[0162] In another aspect, the GUI of the video communication session is smaller than the display screen on which it is displayed, wherein the azimuthal and elevation ranges of translation are independent of a position of the GUI within the display screen. In one aspect, the azimuthal and elevation ranges of translation are independent of a position and orientation of the display screen. In some aspects, the display screen is integrated within the local device, wherein the total azimuthal range of translation spans a width of the display screen and the total elevation range of translation spans a height of the display screen. In another aspect, determining the azimuthal and elevation ranges of translation includes expanding the GUI of the video communication session until at least one of a width of the GUI has fully expanded to a width of the display screen or a height of the GUI has fully expanded to a height of the display screen while maintaining the aspect ratio; and defining the azimuthal range of translation to span the width of the GUI and the elevation range of translation to span the height of the GUI. In another aspect, when the width of the GUI is the width of the display screen, the azimuthal range of translation is the total azimuthal range of translation, the elevation range of translation is less than the total elevation range of translation, and when the height of the GUI is the height of the display screen, the azimuthal range of translation is less than the total azimuthal range of translation and the elevation range of translation is the total elevation range of translation.

[0163] In one aspect, the method further includes determining a spatial parameter indicating a position of a virtual sound source within a range of azimuthal and elevation panning based on a position of the visual representation within the GUI, wherein the input audio stream is spatially rendered using the spatial parameter. In some aspects, the spatial parameter includes an azimuthal angle along the range of azimuthal panning and an elevation angle relative to a reference point in front of the display screen along the range of elevation panning. In one aspect, the method further includes: determining that the aspect ratio of the GUI has changed; adjusting at least one of the range of azimuthal panning and the range of elevation panning based on the changed aspect ratio; adjusting the spatial parameter such that the position is within one of the adjusted range of azimuthal panning and the range of elevation panning; and spatially rendering the input audio stream using the adjusted spatial parameter.

[0164] In some aspects, determining the spatial parameter includes: determining the azimuthal angle of the virtual sound source using an x-coordinate of a center point position of the visual representation within the GUI as an input to a first linear function of the range of azimuthal panning; determining the elevation angle of the virtual sound source using a y-coordinate of the center point position of the visual representation within the GUI as an input to a second linear function of the range of elevation panning; and spatially rendering the input audio stream according to the azimuthal angle and the elevation angle to output the virtual sound source. In another aspect, determining the spatial parameter includes: estimating a range of azimuthal viewing angles for the GUI having a predefined width and a range of elevation viewing angles for the GUI having a predefined height; determining a reference point located at a predefined distance from a front of a display screen displaying the GUI; determining a viewing azimuthal angle from the visual representation on the GUI to the reference point and a viewing elevation angle from the visual representation on the GUI to the reference point; determining the azimuthal angle of the virtual sound source using the viewing azimuthal angle as an input to a first linear function of the range of azimuthal panning relative to the estimated range of azimuthal viewing; and determining the elevation angle of the virtual sound source using the viewing elevation angle as an input to a second linear function of the range of elevation panning relative to the estimated range of elevation viewing.

[0165] In one aspect, the position includes an azimuthal angle and an elevation angle relative to a reference point in front of the local device, wherein the adjusted position has a lower azimuthal angle and a higher elevation angle in response to determining that the aspect ratio has decreased. In another aspect, the lower azimuthal angle is an angle that does not extend fully across a width of the display screen within a range of azimuthal panning of the plurality of speakers and the higher elevation angle is an angle that extends across a height of the display screen within a range of elevation panning of the plurality of speakers when the GUI has a decreased aspect ratio.

[0166] As previously noted, one aspect of the present disclosure can be a non-transitory machine-readable medium (such as a microelectronic memory) having instructions stored thereon that program one or more data processing components (here generally referred to as "processors") to perform network operations and audio signal processing operations as described herein. In other aspects, some of these operations can be performed by specific hardware components containing hardwired logic. Alternatively, those operations can be performed by any combination of programmed data processing components and fixed hardwired circuit components. In one aspect, the operations of the methods described herein can be performed by a local device when the one or more processors execute the instructions stored within the non-transitory machine-readable medium.

[0167] While certain aspects have been described and shown with reference to the attached drawing figures, it will be appreciated that such aspects are merely illustrative of the present disclosure and are not intended to limit the disclosure from that which can be derived from the description, drawings and claims. Therefore, it is contemplated to cover any and all modifications that would fall within the scope of the present disclosure.

[0168] In some aspects, the present disclosure can include language such as "at least one of [element A] and [element B]." This language can mean that one or more of these elements are present. For example, "at least one of A and B" can mean A, B, or A and B. In particular, "at least one of A and B" can mean at least one of A and at least one of B or at least one of A or B. In some aspects, the present disclosure can include language such as "[element A], [element B], and / or [element C]." This language can mean that any one of these elements can be present or any combination of these elements can be present. For example, "A, B, and / or C" can mean A, B, C, A and B, A and C, B and C, or A, B, and C.

Claims

1. A method executed by a programmable processor of a local device, the method comprising: While maintaining communication sessions with multiple remote devices, it can receive multiple input audio streams; Based on the speech activity detection (VAD) parameters associated with the multiple input audio streams, it is determined that the first set of input audio streams from the first group of remote devices will be rendered as individual virtual sound sources, and the mixture of the second set of input audio streams from the second group of remote devices will be rendered as a single virtual sound source. Using multiple speakers and while conducting the communication session with the multiple remote devices, virtual sound sources are output according to the arrangement of virtual sound source locations, wherein the arrangement of virtual sound source locations includes both: different positions corresponding to the individual virtual sound sources relative to a reference point outside the local device, and positions corresponding to the single virtual sound source outside the local device relative to the reference point; and In response to a change in orientation of the local device to a new orientation, at least one of the positions of at least one of the virtual sound sources is adjusted relative to the reference point along the horizontal axis.

2. The method of claim 1, wherein the at least one virtual sound source in the virtual sound source was located at at least one position before the adjustment having a first azimuth angle relative to the reference point along the horizontal axis, and the adjusted position has a second azimuth angle greater than the first azimuth angle.

3. The method of claim 2, wherein when the local device is in the orientation, the first azimuth angle is within the first azimuth angle translation range of the plurality of speakers, and when the local device is in the new orientation, the second azimuth angle is within the second azimuth angle translation range of the plurality of speakers that is greater than the first azimuth angle translation range, wherein the second azimuth angle is proportionally greater than the first azimuth angle relative to the difference between the first azimuth angle translation range and the second azimuth angle translation range.

4. The method of claim 1, further comprising adjusting the at least one position of the at least one virtual sound source among the virtual sound sources relative to the reference point along a vertical axis in response to a change in orientation of the local device to the new orientation, wherein the at least one position has a first elevation angle along the vertical axis and the adjusted position has a second elevation angle less than the first elevation angle.

5. The method of claim 1, wherein the orientation is a longitudinal orientation and the new orientation is a transverse orientation.

6. The method of claim 1, further comprising, in response to the new orientation of the local device changing back to the orientation, adjusting the position of at least one of the virtual sound sources along the horizontal axis back to its original position relative to the reference point.

7. The method according to claim 1, further comprising: Receive 1) a first set of input video streams from the first group of remote devices, and 2) a second set of input video streams from the second group of remote devices; as well as A visual representation of each of the first set of input video streams and the second set of input video streams is displayed in the graphical user interface (GUI) of the communication session on the display screen of the local device.

8. The method of claim 1, wherein the local device includes a handheld device.

9. An electronic device, comprising: At least one processor; as well as A memory having instructions that, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1 to 8.

10. A non-transitory machine-readable medium having instructions that, when executed by at least one processor, cause a local device to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for realizing video communication

    CN102209225A

  • Image display unit with speaker

    JP2006217307A