Spatial audio controller
By utilizing the spatial audio controller and the arrangement of VAD parameters and visual representations, the high load problem of audio rendering in video communication sessions is solved, resulting in a more natural listening experience and reduced CPU load.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-06-02
- Publication Date
- 2026-06-02
AI Technical Summary
In video communication sessions, the central processing unit (CPU) of existing devices suffers from high load issues due to processing video and audio data, especially in the case of multiple remote participants, making it difficult to perform efficient audio spatial rendering.
A spatial audio controller is used to determine whether to render the input audio stream individually or in combination by using speech activity detection (VAD) parameters and the arrangement of visual representations in the graphical user interface (GUI). The spatial parameters of the virtual sound source are determined based on the position of the visual representations and the GUI layout, reducing the amount of computational processing.
It effectively reduces the computational complexity of audio spatial rendering, providing a more natural listening experience while reducing the burden on the central processing unit.
Smart Images

Figure CN122137932A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on June 2, 2022, with application number 202210623716.1 and invention title "Spatial Audio Controller". Technical Field
[0002] One aspect of this disclosure relates to a system including a spatial audio controller that controls how audio is spatialized in a communication session. Other aspects are also described. Background Technology
[0003] Many devices today, such as smartphones, are capable of various types of telecommunications activities with other devices. For example, a smartphone can make a phone call with another device. In this case, when a phone number is dialed, the smartphone connects to a cellular network, which can then connect the smartphone to another device (such as another smartphone or a landline). Furthermore, smartphones are also capable of making video conferencing calls, in which video and audio data are exchanged with another device. Summary of the Invention
[0004] When a local device engages in a communication session with several remote devices (such as a video conferencing call), the local device can receive video and audio data for that session from each remote device. The local device can use the video data to display a dynamic video representation of each remote participant and the audio data to create a spatialized audio rendering for each remote participant. To this end, the local device can perform spatial rendering operations on the audio data of each remote device, allowing the local user to perceive each remote participant from different locations. However, performing these video and audio processing operations places a heavy processing burden on the local device's electronics, such as the central processing unit (CPU). Therefore, a spatial audio controller is needed that creates and manages the audio spatial rendering during the communication session with the remote devices, taking into account both complexity and video presentation while maintaining audio quality.
[0005] To overcome these shortcomings, this disclosure describes a local device with a spatial audio controller for performing audio signal processing operations to efficiently and effectively spatially render input audio streams from one or more remote devices during a communication session. One aspect of this disclosure is a method performed by an electronic device (e.g., a local device) communicatively coupled to one or more remote devices and conducting a communication session. During the session, the local device receives an input audio stream from each remote device and receives a set of communication session parameters for each remote device. For example, these parameters may include one or more Voice Activity Detection (VAD) parameters based on VAD signals received from each remote device, indicating at least one of voice activity and voice intensity of a remote participant at the respective remote device. Furthermore, when these devices conduct video communication sessions (e.g., video conferencing calls), where input video streams are received and visual representations (or tiles) of these video streams are placed in a graphical user interface (GUI) (window) on the local device's display, these session parameters may indicate how these visual representations are arranged within the GUI (e.g., in a larger per-user tiled canvas area or a smaller per-user tiled list area), the size of the visual representations, etc. The local device determines for each input audio stream, based on the set of communication session parameters, whether the input audio stream should be rendered 1) separately relative to other received input audio streams, or 2) in a mixture with one or more other input audio streams. For example, an input audio stream may be rendered separately when at least one of the following conditions is met: a VAD parameter (such as speech activity) is above a speech activity threshold (e.g., indicating that a remote participant is actively speaking), a visual representation associated with the input audio stream is contained within a prominent area of the GUI (e.g., the canvas area of the GUI), and / or the size of the visual representation (e.g., the size of a representation of a video showing a remote participant) is above a threshold size. Therefore, for each input audio stream determined to be rendered separately, the local device renders the input audio stream space as an individual virtual sound source containing only that input audio stream, and for an input audio stream determined to be rendered as a mixture of input audio streams, the local device renders the mixture space of the input audio streams as a single virtual sound source containing the mixture of the input audio streams. Therefore, by spatially rendering certain input audio streams that are more important to local participants separately, while spatially rendering a mixture of other input audio streams that are less critical, the local device can reduce the amount of computational processing required to spatially render all streams in a communication session.
[0006] In one aspect, the local device can determine the arrangement of individual virtual sound sources in front of, behind, or to the side of the local device's display screen based on communication session parameters. For example, for each individual virtual sound source, the local device determines the location within the GUI of a visual representation of the corresponding input video stream (associated with the corresponding input audio stream to be spatially rendered as the individual virtual sound source), and determines spatial parameters (e.g., azimuth, elevation, distance, and even reverberation level) based on the determined location of the visual representation. These spatial parameters indicate the location of the individual virtual sound source within the arrangement (e.g., relative to a reference point in space), thereby spatially rendering the input audio stream as the individual virtual sound source may include using the determined spatial parameters to spatially render the input audio stream as the individual virtual sound source at that location. Such locations may also be included along with a virtual room model, and the spatial parameters include model aspects such as the reverberation level of a room reverberation model as a function of distance and / or location within the room. Thus, the local device can be considered to be located within the virtual room. In one aspect, the arrangement also includes the location of a single virtual sound source for grouping (e.g., mixing) of input audio sources, wherein determining the location involves using one or more locations within the GUI of the visual representation of each video stream in the corresponding input video stream associated with each input audio stream in the mixed input audio stream and / or the level of each input audio stream in the mixed input audio stream, and determining the location and spatial parameters indicating the location of the virtual (grouped or mixed) sound source of the mixed input audio stream based on this information. This grouping location is also used to determine new spatial parameters. For example, determining these new spatial parameters includes determining a weighted combination of spatial location data for all the mixed input audio streams, where the weighting may be a function of the energy level of the individual streams. In one aspect, streams selected for grouping in such a joint single location include those streams whose visual representation is less prominent (e.g., video tiles smaller than tiles in a prominent area of the GUI), or those streams whose visual representation is invisible (e.g., currently implicitly off the display). Thus, when some of these visual representations are less prominent or invisible, a single grouped audio stream rendered spatially to a single location can be used. Therefore, in general, individual or grouped virtual sound sources can be arranged according to function and in proportion to the arrangement of visual representations to provide local users with an ideal spatial experience, while controlling complexity and taking into account the most prominent aspects of the video communication session.
[0007] According to another aspect of this disclosure, a method is performed by a local device that provides an alternative arrangement of virtual sound source locations. For example, the local device receives input audio and video streams from each remote device, displays these input video streams as visual representations in a GUI on a display screen, and spatially renders at least one input audio stream to output an individual virtual sound source comprising only that stream. In response to determining that an additional remote device has joined the video communication session, input audio and video streams are received for each of the additional devices. The local device determines whether the device supports additional individual virtual sound sources for one or more input audio streams from the additional device. In response to determining that the local device does not support additional individual virtual sound sources, the local device defines several user interface (UI) areas located in the GUI, each UI area comprising one or more visual representations of one or more video streams, and spatially renders a mixture of one or more input audio streams associated with one or more visual representations included within the UI area for each UI area.
[0008] According to another aspect of this disclosure, a method performed by a local device provides adjustment of the position of a virtual sound source based on changes in translation range (or limitation) caused by the local device rotating to a different orientation. Specifically, a remote device receives an input audio stream and determines a first orientation of the local device (e.g., a vertical orientation). The local device determines a translation range of several speakers for the first orientation of the local device spanning along a horizontal axis. The local device uses the speakers to spatially render the input audio stream as a virtual sound source at a position along the first horizontal axis and within the translation range. The local device may also translate jointly in the horizontal and vertical directions. Translation limitations or a combined range of horizontal and vertical translation directions may be determined based on the device's orientation. The device's orientation may suggest how the audio is positioned relative to the horizontal and vertical span of the device (e.g., the rectangular screen of the device viewed by a user). In response to determining that the local device is in a second orientation (e.g., the device has rotated 90° to a horizontal orientation), the local device determines an adjusted translation range of the speakers that spans wider along the horizontal axis than the initial translation range, and adjusts the position of the virtual sound source along the horizontal axis based on this adjusted translation range. Therefore, when the local device rotates, the virtual sound source is perceived by the local user to be in a wider position than when the local device was in a previous orientation. In one aspect, the combined horizontal and vertical translation constraints can be a function of orientation. Individual or mixed sound sources will have virtual azimuth and elevation angles within this range, where the mapping of position from the visual representation uses this range to define the function used for mapping.
[0009] According to another aspect of this disclosure, a method performed by a local device determines the translation range of a plurality of speakers based on the aspect ratio of a GUI (e.g., a window displayed on the GUI) of a communication session displayed on a display screen. The local device receives an input audio stream and an input video stream, and displays a visual representation of the input video stream within the GUI of the video communication session displayed on the display screen (e.g., the display screen may be integrated within the local device and may be a window containing the communication session). The local device determines the aspect ratio of the GUI and, based on that aspect ratio, determines an azimuth translation range and an elevation translation range, the azimuth translation range being at least a portion of the total azimuth translation range of the plurality of speakers, and the elevation translation range being at least a portion of the total elevation translation range of these speakers. The local device spatially renders the input audio stream to output a virtual sound source within the azimuth and elevation translation ranges.
[0010] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed embodiments below and specifically pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description
[0011] Multiple aspects are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals in the drawings indicate similar elements. It should be noted that references to "a" or "an" aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a single drawing may be used to illustrate features of more than one aspect, and for a particular aspect, not all elements in that drawing may be necessary.
[0012] Figure 1 A system is shown according to one aspect, comprising one or more remote devices communicating with a local device, the local device including a spatial audio controller for spatially rendering audio from the one or more remote devices.
[0013] Figure 2 A block diagram of a local device 2 is shown, which renders an input audio stream from a remote device with which it is communicating to output a virtual sound source.
[0014] Figure 3 An exemplary graphical user interface (GUI) of a communication session being displayed by a local device according to one aspect is shown, along with the arrangement of virtual sound locations of spatial audio output by the local device during the communication session.
[0015] Figure 4A block diagram of a local device including an audio spatial controller is shown according to one aspect, which performs spatial rendering operations during a communication session.
[0016] Figure 5 It is a flowchart of one aspect of the process used to determine whether an input audio stream should be rendered individually as a single virtual sound source or mixed and rendered as a single virtual sound source, and to determine the arrangement of the virtual sound source's location.
[0017] Figure 6 It is a flowchart of one aspect of the process used to define UI areas and spatially render the input audio stream associated with these UI areas.
[0018] Figure 7 The diagram illustrates several stages according to one aspect, wherein a local device defines a user interface (UI) area for one or more visual representations of input video streams from one or more remote devices, and spatially renders one or more input audio streams associated with the defined UI area.
[0019] Figure 8 The diagram illustrates the translation angle used to render these input audio streams at the corresponding virtual sound sources, according to one aspect.
[0020] Figure 9 It is a flowchart of one aspect of the process for determining spatial parameters based on the location of the corresponding visual representation, which indicates the location where the input audio stream is to be spatially rendered as a virtual sound source.
[0021] Figure 10 An example is shown of determining spatial parameters based on one aspect by using one or more linear functions to map the location of the visual representation to the angle at which the input audio stream is to be rendered.
[0022] Figure 11A and Figure 11B An example is shown of determining spatial parameters based on one aspect by using one or more functions to map the estimated or actual viewing angle of the visual representation (e.g., using assumed plane size and viewing position) to the audio translation angle of the input audio stream to be rendered.
[0023] Figure 12 Several stages are shown, in which the translation range of one or more speakers is adjusted based on the local device rotating from longitudinal orientation to lateral orientation.
[0024] Figure 13 This is a flowchart of one aspect of a process for adjusting the position of one or more virtual sound sources based on a change in one or more translation ranges of one or more speakers, said change being based on a change in the orientation of a local device.
[0025] Figure 14A and Figure 14B The diagram illustrates several stages based on various aspects, where the pan range is adjusted according to the aspect ratio of the GUI based on the communication session.
[0026] Figure 15 This is a flowchart of one aspect of the process 170 for adjusting the position of one or more virtual sound sources by changing the aspect ratio of a GUI based on a communication session. Detailed Implementation
[0027] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Unless the shape, relative position, and other aspects of the components described in any aspect are explicitly defined, the scope of this disclosure is not limited to the components shown, which are for illustrative purposes only. Furthermore, while numerous details have been set forth, it should be understood that some embodiments may be implemented without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description. Moreover, unless the meaning is explicitly contrary, all scopes shown herein are to be considered to include the endpoints of each scope.
[0028] Figure 1 A system 1 is illustrated according to one aspect, comprising one or more remote devices communicating with a local device, the local device including a spatial audio controller for spatially rendering audio from the one or more remote devices. As described herein, this allows the local device to simulate a more realistic listening experience during communication sessions (e.g., video) with these remote devices of a local user, where the audio is perceived by the local user as a sound source from a separate frame of reference (such as the physical environment around or in front of the user) and, if a visual representation is present, moves from or around a screen. The audio system includes a local (or first electronic) device 2, a remote (or second electronic) device 3, a network 4 (e.g., a computer network, such as the Internet), and an audio output device 6. In one aspect, the system may include more or fewer elements. For example, the system may have one or more remote devices, all of which simultaneously communicate with the local device, as described herein. In another aspect, the audio system may include one or more remote (electronic) servers communicatively coupled to at least some of the devices in the audio system 1 and configured to perform at least some of the operations described herein. In yet another aspect, the system may not include an audio output device. In this scenario, the local device can perform audio output operations, such as driving one or more speakers using one or more audio driver signals. These speakers may be integrated into the local device or separate from it (e.g., independent speakers, such as a set of headphones, that are communicatively coupled to the local device).
[0029] In one aspect, the local device (and / or remote device) can be any electronic device (e.g., with electronic components such as processors, memory, etc.) capable of participating in a communication session (such as a video (conference) call). For example, the local device can be a desktop computer, a laptop computer, a digital media player, etc. In another aspect, the device can be a portable electronic device (e.g., capable of being handheld), such as a tablet computer, a smartphone, etc. In another aspect, the device can be a head-mounted device, such as smart glasses, or a wearable device, such as a smartwatch. In one aspect, the remote device can be a device of the same type as the local device (e.g., both devices are smartphones). In another aspect, at least some of the remote devices can be different, such as some being desktop computers and others being smartphones.
[0030] As shown in the figure, local device 2 is coupled to remote device 3 via a computer network (e.g., the Internet) 4 (e.g., communicatively). Specifically, the local device and the remote device can be configured to establish and conduct video conferencing calls, in which the calling devices exchange audio and video data. For example, the local device can use any signaling protocol (e.g., Session Initiation Protocol (SIP)) to establish a communication session and use any communication protocol (e.g., Transmission Control Protocol (TCP), Real-Time Transport Protocol (RTP), etc.) to exchange audio and video data during the session. For example, when a session is initiated (e.g., by a communication session application that can execute within the local device), the local device can capture one or more microphone signals (using one or more microphones of the local device), encode audio data (e.g., using any audio codec), and transmit the audio data (e.g., as IP packets) to one or more remote devices, and receive audio data (e.g., as an input audio stream) from each of these remote devices via the network to drive one or more speakers of the local device.
[0031] Furthermore, the local device can transmit video data (captured by one or more cameras of the device) to each remote device making the call, and receive video data (as an input video stream) from each remote device as an output video stream, and receive at least one video signal (or input video stream) for display on one or more displays. In one aspect, when transmitting video data, the local device can encode the video data using any video codec (e.g., H.264), and then decode and render it on each of the remote devices to which the local device transmits the encoded data.
[0032] In one aspect, network 4 can be any type of network that enables a local device to communicate with one or more remote devices. In another aspect, the network can include a telecommunications network with one or more cell towers, which can be part of a communications network (e.g., a 4G Long Term Evolution (LTE) network) that supports data transmission (and / or voice calls) of electronic devices such as mobile devices (e.g., smartphones).
[0033] In one aspect, the audio output device 6 can be any electronic device including at least one speaker and configured to output sound by driving the speaker. For example, as shown, the device is a wireless headset (e.g., in-ear headphones or earbuds) designed to be positioned on (or in) a user's ear and designed to output sound into the user's ear canal. In some aspects, the headphones can be a sealed type with flexible earpiece ends designed to acoustically seal the entrance to the user's ear canal relative to the surrounding environment by blocking or occluding it within the ear canal. As shown, the output device includes a left earpiece for the user's left ear and a right earpiece for the user's right ear. In this case, each earpiece can be configured to output at least one audio channel of media content (e.g., the right earpiece outputs the right audio channel of a stereo recording (such as a musical work) with two-channel input and the left earpiece outputs the left audio channel). In another aspect, the output device can be any electronic device including at least one speaker and arranged for wear by a user and arranged to output sound by driving the speaker with an audio signal. For example, the output device can be any type of headset, such as over-ear (or on-ear) headphones that at least partially cover the user's ears and are arranged to direct sound into the user's ears.
[0034] In some respects, the audio output device can be a headset, as illustrated herein. In other respects, the audio output device can be any electronic device arranged to output sound to the surrounding environment. Examples may include standalone speakers, smart speakers, home theater systems, or infotainment systems integrated into vehicles.
[0035] In one aspect, the audio output device 6 can be a wireless device communicatively coupled to a local device for exchanging audio data. For example, the local device can be configured to establish a wireless connection with the audio output device via a wireless communication protocol (e.g., Bluetooth or any other wireless communication protocol). During the established wireless connection, the local device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the audio output device, which can include audio digital data of any audio format. Specifically, the local device can be configured to establish and communicate with the audio output device via a two-way wireless audio connection (e.g., which allows two devices to exchange audio data), such as making hands-free calls or using voice commands. Examples of two-way wireless communication protocols include, but are not limited to, Hands-Free Mode (HFP) and Headset Mode (HSP), both of which are Bluetooth communication protocols. In another aspect, the local device can be configured to establish and communicate with the output device via a one-way wireless audio connection (e.g., the Advanced Audio Distribution Profile (A2DP) protocol), which allows the local device to transfer audio data to one or more audio output devices.
[0036] On the other hand, local device 2 can be communicatively coupled to audio output device 6 via other methods. For example, both devices can be coupled via a wired connection. In this case, one end of the wired connection can be (e.g., fixedly) connected to the audio output device, while the other end can have a connector, such as a media jack or a Universal Serial Bus (USB) connector, that inserts into the jack of the audio source device. Once connected, the local device can be configured to drive one or more speakers of the audio output device with one or more audio signals via the wired connection. For example, the local device can transmit the audio signals as digital audio (e.g., PCM digital audio). On the other hand, the audio can be transmitted in an analog format.
[0037] In some respects, local device 2 and audio output device 6 may be different (separate) electronic devices, as illustrated herein. In other respects, the local device may be a component of the audio output device (or integrated with the audio output device). For example, as described herein, at least some components of the local device (such as a controller) may be part of the audio output device, and / or at least some components of the audio output device may be part of the local device. In this case, each device may be communicatively coupled via traces that are part of one or more printed circuit boards (PCBs) within the audio output device.
[0038] Figure 2A block diagram of a local device 2 according to one aspect is shown, which spatially renders input audio streams from a remote device communicating with the local device to output a virtual sound source. Local device 2 includes a controller 10, a network interface 11, a speaker 12, a microphone 14, a camera 15, a display screen 13, an inertial measurement unit (IMU) 16, and (optionally) one or more additional sensors 17. In one aspect, the local device may include more or fewer elements as described herein. For example, the device may include two or more of at least some of these elements, such as having two or more speakers, two or more microphones, two or more cameras, and two or more displays.
[0039] Controller 10 may be a dedicated processor such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). The controller is configured to perform audio signal processing operations and / or networking operations. For example, controller 10 may be configured to conduct video communication sessions with one or more remote devices via network interface 11. Alternatively, the controller may be configured to perform audio signal processing operations on audio data (e.g., input audio streams) associated with the communication session, such as spatially rendering these streams to output them as virtual sound sources, thereby providing a more realistic listening experience for local users. Further details regarding the operations performed by controller 10 are described herein.
[0040] In one aspect, one or more sensors 17 are configured to detect an environment (e.g., in which a local device is located) and generate sensor data based on the environment. In other aspects, a controller may be configured to perform operations based on sensor data generated by one or more sensors 17. For example, these sensors may include proximity sensors (e.g., optical) designed to generate sensor data indicating a specific distance between an object and the sensor (or local device), such as detecting the viewing distance between the local device and a local user. These sensors may also include accelerometers arranged and configured to receive (detect or sense) vibrations (e.g., voice vibrations generated when a user speaks) and generate accelerometer signals representing (or containing) the vibrations. The IMU is designed to measure the position and / or orientation of the local device. For example, the IMU may generate sensor data indicating changes in the orientation of the local device (e.g., with respect to any X, Y, Z axis) and / or changes in the position of the device.
[0041] The speaker 12 may be, for example, an electrically driven driver specifically designed for sound output in a particular frequency band, such as a woofer, tweeter, or midrange driver. In one aspect, the speaker 12 may be a “full-range” (or “full-band”) electrically driven driver that reproduces as much of the audible frequency range as possible. The microphone 14 may be any type of microphone (e.g., a differential pressure gradient microelectromechanical system (MEMS) microphone) configured to convert acoustic energy caused by sound waves propagating in an acoustic environment into an input microphone signal (or audio signal).
[0042] In one aspect, camera 15 is a complementary metal-oxide-semiconductor (CMOS) image sensor capable of capturing digital images including image data representing the field of view of camera 15, wherein the field of view includes a scene of the environment in which the local device 2 is located. In some aspects, the camera may be a charge-coupled device (CCD) camera type. The camera is configured to capture still digital images and / or video represented by a series of digital images. In one aspect, the camera may be positioned anywhere near the local device. In some aspects, the device may include multiple cameras (e.g., each camera may have a different field of view).
[0043] Display screen 13 is designed to present (or display) digital image or video (or image) data. In one aspect, the display screen may use liquid crystal display (LCD) technology, light-emitting polymer display (LPD) technology, or light-emitting diode (LED) technology, although other display technologies may be used in other aspects. In some aspects, the display screen may be a touch-sensitive display screen configured to sense user input as an input signal. In some aspects, the display screen may use any touch sensing technology, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies.
[0044] In one aspect, any of the elements described herein may be part of (or integrated into) a local device (e.g., integrated into the housing of the local device). In another aspect, at least some of these elements may be separate electronic devices communicatively coupled to the local device (e.g., a controller via the local device's network interface) (via Bluetooth connection). For example, a speaker may be integrated into another electronic device configured to receive audio data from the local device for driving the speaker. As another example, display screen 13 may be integrated with the local device, or the display screen may be a separate electronic device (e.g., a monitor) communicatively coupled to the local device.
[0045] In one aspect, controller 10 is configured to perform audio signal processing operations and / or networking operations, as described herein. For example, the controller may be configured to establish communication sessions with one or more remote devices and to acquire (or receive) audio / video data from these remote devices. The controller is configured to output (display) video data on a display screen and spatially render audio data. Further details regarding spatially rendered audio data are described herein. In one aspect, the operations performed by the controller may be implemented in software (e.g., as instructions stored in memory and executed by the controller) and / or may be implemented by hardware logic structures as described herein.
[0046] In one aspect, controller 10 may be configured to perform (additional) audio signal processing operations based on elements coupled to the controller. For example, when the output device includes two or more “outside-ear” speakers arranged to output sound into an acoustic environment rather than speakers arranged to output sound into a user’s ears (e.g., as speakers in in-ear headphones), the controller may include an audio output beamformer configured to generate speaker driver signals that produce spatially selective sound output when driving two or more speakers. Thus, when used to drive speakers, the output device may generate a directional beam pattern that can be directed to a location within the environment.
[0047] In some aspects, the controller 10 may include a sound pickup beamformer configured to process audio (or microphone) signals generated by two or more external microphones of the output device to form a directional beam pattern (as one or more audio signals) for spatially selective sound pickup in certain directions, thereby increasing sensitivity to the location of one or more sound sources. In some aspects, the controller may perform audio processing operations (e.g., perform spectral shaping) on the audio signals containing the directional beam pattern.
[0048] Figure 3An exemplary graphical user interface (GUI) (or GUI window) of a communication session being displayed by a local device according to one aspect is illustrated, along with the arrangement of virtual sound locations of spatial audio output by the local device during the communication session. Specifically, the illustration shows a GUI 41 of the communication session being displayed on a display screen 13 of the local device, and shows an arrangement 50 of virtual sound source locations, in which spatial audio associated with one or more remote devices (participants of the remote devices) is located and perceived by the local user 40. GUI 41 includes an arrangement 51 of visual (or video) representations (or tiles), wherein five visual representations 44-48 are each associated with a different remote participant and a tile 49 associated with the local user 40. Specifically, the visual representation for each remote participant displays an input video stream received from the remote device of each respective participant, which is in a communication session with the local device. In one aspect, at least some of these representations may be dynamic, displaying live video for at least a portion of the length of the communication session. In another aspect, one or more visual representations may display static images (e.g., static images of the remote participants). Visual representation 49 displays video data of the local user 40 that can be captured by camera 15 (and transmits the video data to at least some of the remote devices for display on their respective devices during a communication session).
[0049] As shown in the figure, the GUI is divided into two regions: a canvas (or primary) region 42 and a list (or secondary) region 43. The canvas region includes three visual representations 44-46, and the list region includes two visual representations 47 and 48. As shown, the visual representations within the canvas region can be visually larger than the visual representations in the list region, where remote participants with prominent speech activity (e.g., relative to other remote participants in the session) are positioned within the canvas region. As described herein, remote participants can move between these regions; for example, a remote participant can adaptively move from the list region to the canvas region based on speech activity. As described herein, the local device can spatially render audio data differently based on the region where the visual representation is located. For example, the local device can render audio data associated with remote participants in the canvas region separately, while, in contrast, audio data associated with remote participants in the list region can be mixed, and this mixture can then be spatially rendered as a virtual sound source. This document describes more about spatial rendering and regions of the GUI.
[0050] In a traditional (video) communication session between multiple remote participants and a local user on local device 2, audio from all remote participants is perceived as originating from the same location relative to the local user (e.g., as if these remote participants were speaking directly to each other from the same point in space). Because it is difficult for the local user to keep up with the session when more than one remote participant is speaking at a time, such interactions lead to numerous interruptions. Furthermore, if the audio from all remote participants originates from the same point in space (more precisely, as perceived by the local user), maintaining eye contact or following the session between different remote participants from different areas within the GUI becomes difficult, potentially regardless of the visual location of the remote participants. To overcome this, the local device spatially renders the audio data received from the remote devices differently, such that the voices of at least some of the remote participants originate from different (spatial) locations relative to the local user 40. Specifically, the figure illustrates an arrangement 50 of virtual sound source locations, in which the local device spatially renders audio data from remote participants associated with visual representations 44-46 individually as individual virtual sound sources 54-56, and spatially renders audio data from one or more remote devices associated with representations 47 and 48 as a single virtual sound source 53 comprising a mixture of audio received from those devices. In one aspect, the arrangement of virtual sound source locations may resemble the arrangement of the visual representations to give the local user the impression that the speech sources from the remote participants originate from the same general location as their corresponding visual representations on the screen, as shown. For example, as shown, individual virtual sound sources are positioned on top of (or positioned above) their respective visual representations, while the single virtual sound source 53 of the list of visual representations 47 and 48 is centered between these visual representations. Thus, during a communication session, when remote participant 44 speaks, the local user will perceive the remote participant's speech source as originating from (or above) that participant's visual representation. In another aspect, the arrangement of virtual sound source locations may be different. For example, the arrangement could resemble the arrangement of a visual representation, but on a proportionally larger scale, so that although the speech does not originate from every relative visual representation, it can originate from the general location. This could be advantageous if the screen is small or if the user chooses to minimize the size of the GUI. In both cases, a larger, more natural acoustic spatial representation may be more comfortable and useful for local users. This article describes more about how audio data from remote participants is spatially rendered.
[0051] As shown in the figure, a local user engages in a communication session with five remote participants. In one aspect, more remote participants can join the communication session, in which case they (or more precisely, their associated visual representations) can be placed within a canvas area or a list area. As more remote participants are placed within the canvas area, the local device performs additional spatial rendering, which may require more resources and computational processing. This additional processing can place a heavy burden on electronics such as controller 10. Therefore, as described herein, controller 10 performs spatial rendering operations that manage audio spatial rendering during the communication session with the remote devices. Further details regarding these operations are described herein.
[0052] Figure 4 A block diagram of a local device 2, including an audio spatial controller, is shown according to one aspect. This audio spatial controller performs spatial rendering operations during a communication session. Specifically, the diagram shows a controller 10 having several operation blocks for performing audio signal processing operations to spatially render input audio streams from one or more remote devices communicating with the local device. As shown, the controller includes an audio spatial controller 20, a video communication session manager 21, a video renderer 22, and an audio spatial renderer 23.
[0053] The video communication session manager 21 is configured to initiate (and perform) a communication session between a local device 2 (e.g., via network interface 11) and one or more remote devices 3. For example, the session manager may be part of (or receive instructions from) a communication session application (e.g., a telephone application) executed by the local device 2 (e.g., the controller 10 of the local device). For example, the application may display a GUI on the local device's display screen 13, which may provide the local user 40 with the ability to initiate a session (e.g., using a simulated keyboard, a contact list, etc.). Once the GUI receives user input (e.g., dialing a remote user's phone number using a keyboard), the session manager 21 may communicate with the network 4 to establish a communication session, as described herein.
[0054] Once started, the video communication session manager 21 can receive communication session data from each (“N”) remote device in a communication session with the local device. In one aspect, the received data may include input audio streams and input video streams from each remote device, the input audio stream including audio data of the corresponding remote participant (e.g., captured by one or more microphones of the remote device), and the input video stream including video data of the corresponding remote participant (e.g., captured by one or more cameras of the remote device). Thus, the session manager can receive N input audio streams and N input video streams. In one aspect, the session manager can assign each input audio stream to a specific input audio channel within a predefined number of input audio channels. In one aspect, the assigned channels may remain assigned to the specific input audio stream (e.g., the remote device) for the duration of the communication session. In one aspect, the session manager can dynamically assign input audio channels to remote participants as they join the communication session. In one aspect, the manager can receive more or fewer streams (e.g., when a remote device has disabled its camera, the manager may only receive the audio stream).
[0055] In one respect, the session manager 21 may receive one or more audio signals from each remote device. For example, the input audio stream may be a single audio channel (e.g., a single audio signal). In another respect, the input audio stream may include two or more audio channels, such as stereo recordings in multi-channel formats or audio recordings such as 5.1 surround sound formats.
[0056] In one aspect, the session data may include additional data from at least some of N remote devices, such as Voice Activity Detection (VAD) signals. For example, a remote device may generate a VAD signal (e.g., using a microphone signal captured by the remote device) indicating whether speech is present in the corresponding input audio stream of the remote device. For example, the VAD signal may have a high signal level (e.g., high signal level 1) when speech is detected, and a low signal level (e.g., low signal level 0) when no speech is detected (or at least not within a threshold level). In another aspect, the VAD signal need not be a binary decision (speech / non-speech); rather, it may be a probability of speech presence. In some aspects, the VAD signal may also indicate the signal energy level (e.g., sound pressure level (SPL)) of the detected speech.
[0057] In one aspect, the video communication session manager 21 can be configured to transmit data (e.g., audio data, video data, VAD signals, etc.) to one or more of N remote devices. For example, the manager can receive microphone signals generated by microphone 14 (which may include the voice of local user 4) and video data generated by camera 15. Once received, the session manager 21 can distribute this data to at least some of the N remote devices.
[0058] Video communication session manager 21 is configured to transmit N video streams and one or more VAD values (or parameters) for each video stream to video renderer 22 based on VAD signals. For example, each (or at least one) video stream may be associated with one or more VAD parameters indicating the speech activity of a remote participant within an input audio stream associated with the input video stream (e.g., the value may be zero or 1, as described herein). In one aspect, the manager may transmit VAD parameters indicating the duration for which the remote participant is speaking. Specifically, the VAD parameter may indicate the duration for which the remote participant is currently speaking (e.g., the duration may correspond to the duration since the VAD value changed from zero to 1). In another aspect, the VAD parameter may indicate the total time during the entire session in which the remote participant has spoken (e.g., twenty minutes in a thirty-minute communication session). In yet another aspect, the VAD parameter may indicate the speech intensity of the remote participant (e.g., signal energy level) (e.g., in SPL). In yet another aspect, the session manager may transmit (initial) VAD signals received from a remote device to the video renderer.
[0059] The video renderer is configured to render the input video stream into a visual representation of the arrangement (such as...). Figure 3 (As shown in arrangement 51). In one aspect, the renderer is configured to arrange the visual representation based on one or more VAD parameters received from the session manager 21. For example, refer to Figure 3The renderer can locate the visual representations of remote participants 44-48 in their respective regions based on their corresponding speech activity and / or speech intensity. For example, when speech activity is above (or equal to) a speech activity threshold (e.g., equal to one), the renderer can locate the visual representation within the canvas region, indicating that the remote participant is actively speaking. However, if speech activity is below the threshold, the renderer can locate the visual representation of the remote participant in the list region. Furthermore, the renderer can locate the visual representations within canvas region 42 in a specific order based on speech activity and / or speech intensity. For example, remote participants who speak frequently (and / or are currently speaking) may be located higher in the arrangement. For example, remote participant 44 may be currently speaking and / or may have spoken for most of the communication session (e.g., above the threshold). In contrast, remote participants who speak less frequently may be located closer to list region 43 (e.g., visual representation 46), and those who speak less regularly (e.g., below the threshold) may be located in the list region. On the other hand, remote participants with high speech intensity (e.g., signal energy levels above a threshold) can be positioned higher than those without high speech intensity.
[0060] Along with positioning visual representations (e.g., based on VAD parameters), the renderer can define the size of these representations based on one or more criteria. Specifically, one or more similar criteria mentioned above related to the position of a visual representation can be applied to the size of the visual representation. For example, a remote participant who speaks longer and more frequently during a communication session (compared to other participants) may have a larger visual representation than two remote participants who speak less (e.g., remote participant 44 may speak more frequently than participants 45 and 46). On the other hand, the size of the visual representation can be based on speech intensity. For example, when a remote participant speaks louder (e.g., above a signal threshold), the renderer can increase the size of the participant's representation. On another hand, size (and / or position) can be based on the signal-to-noise ratio (SNR) of the remote participant's input audio stream, whereby a participant with a higher SNR may have a larger representation than a participant with a lower SNR (e.g., below a threshold). In one aspect, the representation may appear larger and be positioned higher to provide the local user with the visual meaning that should be given during the communication session. In other aspects, the position (e.g., vertical) within the canvas area may be based on the size of the representation. For example, the largest (with more corresponding surface area) representation 44 is positioned above representation 45, which is greater than representation 46.
[0061] In one aspect, renderer 22 may locate representations within a list based on any criterion described herein below one or more corresponding thresholds. For example, when speech activity is infrequent and short-duration, remote participants may be located within a list region. In some aspects, the location of remote participants and their size may also be based on speech intensity. For example, remote participants with high intensity (e.g., energy levels above a threshold) may be located within a canvas region and / or may have a larger representation size than remote participants with lower speech intensity, as described herein. In one aspect, visual representations within a list region may all have the same size, such as... Figure 3 As shown in the image.
[0062] In one aspect, the rendering operations performed by renderer 22 can be dynamic throughout the communication session. For example, the renderer can continuously and dynamically rearrange and / or resize the visual representation based on changes in one or more VAD parameters. For example, when a remote participant speaks less, the position of visual representation 44 can change (e.g., decrease along the arrangement), and / or its size can change within the canvas area (e.g., decrease in size). If speech activity decreases over a period of time (e.g., below a speech activity threshold), the renderer can eventually place the visual representation in the list. Conversely, the renderer can adjust the position of one or more list representations based on criteria proposed herein. For example, when the remote participant associated with representation 47 begins to speak more frequently, the renderer can move that representation to the canvas area. In one aspect, this determination can be based on whether the VAD parameters indicate that the remote participant has spoken for a period of time (e.g., continuously).
[0063] In one aspect, the renderer can adjust the arrangement as needed by moving visual representations into and out of two areas. For example, if the renderer moves representation 47 to the canvas area, the canvas representation can be moved around the canvas area to accommodate additions. Specifically, the renderer can distribute the representations evenly within the canvas area. In another aspect, the renderer can arrange the representations so that they do not overlap. In some aspects, according to the criteria presented herein, when additional remote participants join the communication session, the renderer can adjust the arrangement 51 of the visual representations, adding their respective visual representations to canvas area 42 or list area 43. Renderer 22 transmits N video streams to display screen 13 for display in GUI 41, as described herein.
[0064] In one aspect, the renderer is configured to generate N sets of communication session parameters based on the criteria presented herein, one set of which is for each remote device (the input audio stream received from the corresponding remote device). For example, this set of communication session parameters may: 1) indicate the size of the GUI (e.g., the orientation of the GUI and / or the size relative to the display screen), 2) indicate the location of the visual representation (e.g., X, Y coordinates), 3) indicate the size of the visual representation of the corresponding input video stream in the GUI, and 4) may include one or more VAD parameters, as described herein. In one aspect, the location indicated by these parameters may be a specific location within the visual representation. For example, the location may be the center point of the visual representation. In another aspect, the location may be based on the video displayed within the visual representation. For example, the location may be a specific part of the remote participant displayed within the visual representation, such as the mouth of the remote participant. In one aspect, to determine this location, the renderer may perform an object recognition algorithm to identify the mouth of the remote participant.
[0065] In one aspect, the position of the visual representation can be relative to the size of the GUI and / or relative to the display screen of the local device. Similarly, coordinates can be relative to components of the local device; for example, when the local device is a multimedia handheld device (e.g., a smartphone), these coordinates can be based on the dimensions of the device's casing or on the size of the display screen showing the GUI. In this case, parameters may also include boundary conditions of the GUI and / or components (e.g., the width of the GUI (in the X direction) and the height of the GUI (in the Y direction)). In another aspect, the renderer can transfer one or more of these VAD parameters to another domain to produce emphase values ordered from the most prominent "1" to the least prominent "N" in the audio stream. Therefore, in Figure 3 In this context, each visual representation in visual representations 44-48 will be assigned a salient value between 1 and 5. In one aspect, this salient value may correspond to the position and / or size of the visual representation within the GUI (e.g., the visual representation has a higher salient value than other visual representations).
[0066] Renderer 22 transmits N sets of communication session parameters to video communication session manager 21, which in turn transmits the N sets of parameters and N input audio streams associated with those parameters to audio spatial controller 20. In one aspect, the video renderer may directly transmit these sets of communication session parameters to the audio spatial controller. The audio spatial controller is configured to receive these parameters and the input audio streams, and is configured to perform one or more audio signal processing operations based on the communication session parameters to spatially render at least some of the input audio streams. As described herein, spatially rendering audio streams may require significant processing power. When a local user is in a communication session with a small number of remote participants, audio spatial controller 20 may have the resources to spatially render audio data from each of these remote participants individually. However, as the number of remote participants increases, the controller may not be able to process all the data as individual virtual sound sources. Therefore, the controller is configured, for each input audio stream, to determine, based on the set of communication session parameters, whether the input audio stream is to be rendered separately relative to other received input audio streams, or 2) to be rendered as a mixture with one or more other input audio streams. Once determined, the audio spatial controller can be configured to determine how to spatially render the input audio stream (e.g., where to output the virtual sound source of the input audio stream). Therefore, the audio spatial controller manages the output of the input audio stream, thereby ensuring sufficient computational resources are available.
[0067] The audio spatial controller 20 includes an individual audio stream selector 27, a spatial parameter generator 25, and a matrix / router 26. The selector is configured to select one or more input audio streams from N input audio streams for individual rendering, and select one or more of the remaining N input audio streams for rendering as a mix. In one aspect, the audio spatial renderer 23 can be configured to spatially render a limited number of output audio channels as virtual sound sources (e.g., due to resource constraints as described herein). When the audio spatial controller spatially renders the input audio streams, it can assign one or more input audio streams to one of the output audio channels (which the audio spatial renderer renders as a virtual sound source). Therefore, due to the limited number of output audio channels, the controller can be limited to outputting a predefined number of virtual sound sources, as the number of virtual sound sources may be limited by the number of output audio channels. In one aspect, this predefined number is between three and six output audio channels. Therefore, in the input audio streams, the selector can select only the number of input audio streams used for individual spatial rendering, i.e., equal to or less than the predefined number. In some respects, one of a predefined number of output audio channels can be reserved for spatial rendering of the input audio streams for remote participants in list region 43. This article describes more about spatial rendering of the input audio streams.
[0068] In one aspect, selector 27 is configured to determine, based on N sets of communication session parameters received from session manager 21, which of the N input audio streams should be spatially rendered individually (e.g., M individual audio streams less than or equal to N audio streams), and / or rendered as a mixture. For example, the selector may determine whether an input audio stream should be rendered individually based on speech activity, speech intensity, and / or prominence values indicated (or determined) by one or more VAD values of a remote participant. Specifically, the selector may determine that an input audio stream should be rendered individually when speech activity is above a threshold (e.g., indicating that a remote participant is currently and / or periodically speaking during a communication session). In some aspects, where a ranking of prominence values is available, the selector may select the highest "M" streams as those to be spatially rendered individually. In one aspect, the M streams are less than or equal to a predefined number of audio output channels that can be used for spatial rendering as virtual sound sources, as described herein.
[0069] On the other hand, the selector may determine whether the input audio stream should be rendered separately based on characteristics (such as the position and size of the representation) of the visual representation of the input video stream (e.g., a video containing remote participants) associated with the input audio stream displayed in the GUI. For example, the selector determines that the input audio stream should be rendered separately when the position of the visual representation is within the canvas area of the GUI. Similarly, the selector may determine that the input audio stream should be rendered separately when the size of the visual representation is above a threshold size (e.g., larger than the size of all or some representations displayed within the list area). In one aspect, the selector may determine whether the input audio stream should be rendered as a blend based on the same criteria. For example, the selector determines that the input audio stream should be rendered as a blend when the associated visual representation has a size below a threshold size, is within the list, and / or has speech activity below a threshold. In another aspect, the determination made by the selector may be based on one or more criteria. For example, the selector may determine that the input audio stream should be rendered separately when one or more of the criteria proposed herein are met.
[0070] Spatial parameter generator 25 is configured to receive N sets of communication session parameters from session manager 21, and is configured to determine the arrangement of individual virtual sound sources for the input audio stream to be rendered individually based on these communication session parameters (e.g., Figure 3the arrangement 50) and / or the positions of one or more individual virtual sound sources, each having a mix of one or more input audio streams that are not rendered individually within the physical environment (e.g., relative to the local user 40). In one aspect, the arrangement of virtual sound source positions to be rendered by the local device can be the same (or similar) to the arrangement of the visual representation. For example, referring to Figure 3 , the arrangement 50 of virtual sound source positions (e.g., approximately) maps to the arrangement of the visual representation such that the sound source 54 is perceived by the local user 40 as originating from the remote participant 44. In another aspect, the arrangement of the sound source positions can be different. For example, the virtual positions can have the same arrangement as the visual representation, but can be proportionally larger or smaller. There are advantages to making the arrangement of the sound sources proportional or commensurate in a relative but not exact sense, as described herein. In another aspect, such an arrangement is beneficial when the position of the visual representation is not fixed or is indeterminate relative to the local user. Thus, referring to Figure 3 , when rendered by the local device, instead of the local user perceiving the individual virtual sources at (or approximately) the same position as the corresponding visual representation, when scaled up proportionally, these positions can be wider and taller than the GUI.
[0071] In another aspect, the arrangement 50 of virtual sound source positions may not be similar to the arrangement of the visual representation. For example, instead of being distributed around a two-dimensional (2D) XY plane, the virtual sound sources can be distributed along one axis (e.g., along the vertical axis). In another aspect, although shown as a 2D plane distribution, the virtual sound sources can be part of a three-dimensional (3D) sound field, where the virtual sound sources appear to originate from various distances from the local user. More about generating a 3D sound field is described herein.
[0072] In one aspect, the generator uses a corresponding set of communication session parameters for each input audio stream to determine one or more spatial parameters (or spatial data) indicative of the spatial characteristics for spatially rendering the input audio stream as a virtual sound source.
[0073] For example, these spatial parameters can include the position of the virtual sound source within the arrangement of virtual sound source positions based on the determined position of the visual representation. Specifically, these spatial parameters can map the virtual sound sources to their corresponding visual representations (above, behind, or adjacent). For example, these spatial parameters can indicate that the position of the virtual sound source is located on its corresponding visual representation such that the voice of the remote participant is perceived by the local user as originating from the visual representation of that participant.
[0074] In one aspect, spatial parameters indicating the location of the virtual sound source may be relative to (e.g., predefined) a reference point located in front of the display screen 13 of the local device. In some aspects, the reference point is a pre-determined viewing position within the physical environment (e.g., positioned at the head level of the local user while viewing the display screen 13). In another aspect, the reference point may be determined by a spatial parameter generator. For example, the generator may use sensor data to determine the location of the local user, such as proximity sensor data generated by one or more proximity sensors. In yet another aspect, the generator may receive user input (e.g., via a GUI displayed on the display screen) indicating the location of the local user (e.g., the distance the user is positioned from the local device (the display screen)).
[0075] To map the location of virtual sound sources (e.g., individuals) associated with a visual representation of the canvas, a spatial parameter generator can be configured to determine one or more translation ranges of one or more speakers 12 within which the virtual sound sources can be positioned. In one aspect, the translation ranges can be predefined. Using the translation ranges, the spatial parameter generator generates one or more spatial parameters comprising one or more angles within one or more translation ranges, the one or more translation ranges being positions within an arrangement corresponding to the location of the virtual sound source relative to the visual representation. For example, a reference... Figure 8 The visual representation 44 may include azimuth angles (e.g., -θ) between the position of the individual virtual sound source (e.g., L1) along a first axis (e.g., the X-axis) and a second (reference or 0°) axis (e.g., the Z-axis). L1 ), and may include an elevation angle (e.g., +β) along the third axis (e.g., the Y-axis) between L1 and the second axis. L1 This article describes more about determining the location of virtual sound sources.
[0076] On the other hand, the generator can determine one or more additional spatial parameters. For example, the generator can determine the distance (e.g., distance in the Z direction) at which the virtual sound source will be perceived by the local user. In one aspect, this distance can be determined based on the size and / or position of the visual representation within the GUI. For example, the generator can assign a first distance (e.g., from a reference point) to a first visual representation of a first size and a second distance to a second visual representation of a second size. In one aspect, the first distance can be shorter than the second distance when the first visual representation is larger than the second visual representation. For example, it can be... Figure 3Visual representation 44 is assigned a shorter distance to the reference point than the distances assigned to representations 45 and 46. Alternatively, the distance can be based on the position of the visual representation on the canvas, whereby a visual representation positioned higher may have a shorter distance than a visual representation positioned lower within the canvas area. Alternatively, the distance can be defined based on the region in which the representation is located. For example, a representation within list area 43 may have a greater distance than (any) representation within the canvas.
[0077] On the other hand, spatial parameters may include reverberation levels, which the audio spatial renderer uses to apply reverberation to the input audio streams. In one aspect, adding reverberation to the stream can be based on the virtual audio location being modeled as a location within a virtual room surrounding the device. Specifically, the spatial parameter generator can generate (or receive) a reverberation model of that room as a function of distance and / or location within the room, such that when applied to one or more input audio streams, it gives the local user the audible impression that a communication session is taking place within the virtual room (e.g., a session between a remote participant and a local user taking place within a conference room). In some aspects, the generator can determine the reverberation level to be applied to one or more input audio streams based on one or more communication session parameters. For example, the reverberation level can be based on the location and / or size of a visual representation, such as having a larger reverberation level for a smaller-sized visual representation to give the listener the impression that the virtual sound source associated with that visual representation is farther away than a virtual sound source with a lower applied reverberation level (which may be associated with a larger-sized visual representation). In one aspect, the application of reverberation (by renderer 23) can provide spatial depth to the virtual sound source. Figures 8 to 11B More details are described regarding spatial parameters and how these parameters are generated.
[0078] Once one or more spatial parameters of all (or most) of the virtual sound sources have been determined, the generator determines whether the spatial parameters of the input audio stream will be used to render individual sources. Specifically, the generator receives one or more control signals from selector 27, indicating which set of communication parameters is associated with the input audio stream to be personalized, and which set of parameters is associated with the stream that will not be personalized. Generator 25 passes the spatial parameters to audio spatial renderer 23, which renders the corresponding input audio stream spatially as an individual virtual sound source at the location indicated by the one or more spatial parameters.
[0079] However, for non-personalized streams, the generator can generate one or more distinct (or new sets) spatial parameters for mixing one or more input audio streams, which indicate a specific location based on at least some of the spatial parameters of the streams in the mix. For example, the distinct spatial parameters could include a weighted combination of spatial parameters that determine all (or some) of the input audio streams in the mix. For instance, a single virtual sound source 53 associated with list visual representations 47 and 48 is located in the middle of the visual representation within the list region (or can be perceived as originating from the center of the list region). Therefore, this location can be determined by averaging the spatial parameters of the two list representations.
[0080] On the other hand, the location of a single virtual sound source that includes a mixture of one or more input streams can be based on whether a remote participant within the mixture is speaking. For example, the generator can analyze a set of conversational parameters, specifically VAD parameters, and when these parameters indicate that a participant is speaking (e.g., vocal activity is above a threshold), the generator can generate one or more spatial parameters for the mixture such that the virtual sound source of the mixture is positioned in the participant's visual representation, similar to the placement of an individual virtual sound source. For illustration, please refer to... Figure 3 A single virtual sound source can be positioned based on which of the two remote participants 47 or 48 is speaking. In this case, the virtual sound source 53 can move (or switch) from two positions (e.g., along the X-axis) depending on which of the two participants is speaking.
[0081] On the other hand, if the weighting of the visual representation of a particular stream is a function of the energy level of the stream's audio, the generator can use a weighted combination of the positions of the visual representations, and the virtual sound sources can be determined by the positional data of the audio sources that are more dominant in the mix (e.g., those with the highest speech activity, etc.). For example, when the visual representation of a list region is in a single line (e.g., ... Figure 3 As shown in the diagram, the sound location can be closer to the representation of the remote participant with the greatest (or highest) speech activity relative to other remote participants. Therefore, even if the visual representation itself does not move, the audio location can move to accommodate the speech activity within that row (or column) of the visual representation. Furthermore, as the visual representation moves between the canvas area and the list area, one or more of the same criteria presented herein can be used to adjust the determination of which streams are rendered individually and which streams are mixed. Therefore, the generation of spatial parameters can be dynamic throughout the communication session for both individual virtual sound sources and mixed single sound sources. Similar to personalized virtual sound sources, the generator passes the spatial parameters of individual virtual sound sources to the renderer 23.
[0082] In one respect, differences in arrangement can be based on the speaker's translation range. This article describes more about translation range.
[0083] Matrix / Router 26 is configured to receive N input audio streams (e.g., as N input audio channels) from Manager 21, and to receive one or more control signals from Stream Selector 27 indicating which streams are personalized and which are non-personalized, and is configured to route audio streams to Audio Spatial Renderer 23 via (e.g., predefined) output audio streams. Specifically, Matrix / Router 26 is configured to route M individual audio streams, each of which is selected by the Stream Selector as needing to be spatially rendered individually via its own output audio stream. In other words, the router assigns the output audio stream (or channels) to each of the N audio streams to be spatially rendered individually. Furthermore, the matrix / router mixes non-personalized audio streams (e.g., by performing a matrix mixing operation) into the mix of audio streams (e.g., as a single audio output stream).
[0084] The audio spatial renderer 23 is configured to receive multiple input audio streams. As shown, this may correspond to M individual audio streams and one or more individual audio streams corresponding to a mixture of one or more audio streams. In one aspect, there may be more than one input audio stream corresponding to the mixture, and as described herein, for each input audio stream, the renderer also receives spatial parameters instructing how to render the input stream. The renderer receives the input stream and spatial parameters for each stream and is configured to spatially render these streams according to the spatial parameters to generate an arrangement of virtual sound sources (e.g., including their three-dimensional (3D) sound field) using one or more speakers 12. Specifically, for each input stream, the audio spatial renderer 23 internally creates a spatial rendering of that stream. These spatial renderings are then combined, for example by merging, to create an output audio stream (e.g., which may include two or more driver signals for driving the speakers). In one aspect, audio spatial rendering can use any spatial rendering method (such as vector-based amplitude translation (VBAP)) to render each audio stream in the audio stream to output individual virtual sound sources and a single sound source through two or more speakers, where each individual virtual sound source comprises a separate individual audio stream, and the single sound source comprises a mixture of audio streams at a location indicated by corresponding spatial parameters. In another aspect, the renderer can apply other spatial operations, such as using spatial parameters to upmix one or more individual audio streams to produce multichannel audio for driving two or more speakers. For example, the renderer can produce multichannel audio in a surround sound multichannel format (e.g., 5.1, 7.1, etc.), where each channel is used to drive a specific speaker 12.
[0085] In one aspect, the renderer may apply other spatial operations to generate a binaural, two-channel output signal (e.g., which can be used to drive headphones, as described herein). In another aspect, the renderer spatially renders the input audio stream, applying one or more spatial filters, such as head-related transfer functions (HRTFs). For example, using spatial parameters (which may include azimuth, elevation, distance, reverberation level, etc., as described herein), the renderer may determine one or more HRTFs and may apply HRTFs to the received input audio stream to generate a binaural audio signal providing spatial audio. In one aspect, the spatial filter may be a generic or pre-determined spatial filter (e.g., determined in a controlled setting, such as a laboratory), which may be applied by the renderer to a pre-determined location based on a distance indicated by the spatial parameters (e.g., typically optimized for optimal positions in front of one or more listeners and / or audio equipment). In another aspect, the spatial filter may be user-specific (e.g., determined based on user input or automatically determined by a local device), based on one or more measurements of the listener's head. For example, the system may determine an HRTF, or equivalently, a head-related impulse response (HRIR) based on anthropometric measurements of the listener. For example, the renderer may receive sensor data (e.g., image data generated by camera 15) and use the data to determine anthropometric measurements of the listener.
[0086] On the other hand, the renderer may perform a crosstalk cancellation (XTC) algorithm. For example, the renderer may perform the algorithm by mixing and / or delaying (e.g., by applying one or more XTC filters to the audio stream) the audio stream to generate one or more XTC signals (or driver signals). In one aspect, when used to drive one or more speakers in a loudspeaker, the renderer may generate one or more first XTC audio signals containing audio content (at least a portion of which) of the audio stream, which is primarily heard at one ear (e.g., the left ear) of a listener in an optimal position (e.g., in front of or facing the local device), and generate one or more second XTC audio signals containing audio content of the audio stream, which is primarily heard at the user's other ear (e.g., the right ear).
[0087] In some respects, renderer 23 may perform one or more additional audio signal processing operations. For example, renderer may apply reverberation to one or more input audio streams (or the rendering of streams) based on received reverberation levels and received spatial parameters. In other respects, renderer may perform one or more equalization operations (e.g., spectral shaping) on one or more streams, such as by applying one or more filters (e.g., low-pass filters, band-pass filters, high-pass filters, etc.). In yet another respect, renderer may apply one or more scalar gain values to one or more streams. In some respects, the application of equalization and scalar gain values may be based on distance values of spatial parameters, such that the application of these operations provides space (or depth) for the corresponding virtual sound sources.
[0088] Because spatial rendering is performed on each individual input audio stream and one or more mixtures of the input audio streams, the renderer generates a single set (e.g., one or more) of driver signals to drive one or more speakers 12, which may be part of a local device or separate from the local device, in order to generate a 3D sound field including each individual virtual sound source in the individual virtual sound source and one or more individual virtual sound sources, as described herein.
[0089] In some aspects, controller 10 may perform one or more additional audio signal processing operations. For example, the controller may be configured to perform active noise cancellation (ANC) to make one or more speakers generate noise immunity in order to reduce ambient noise from the environment leaking into the user's ears (e.g., when the speakers are part of headphones worn by the local user). The ANC function may be implemented as feedforward ANC, feedback ANC, or a combination thereof. Thus, the controller may receive a reference microphone signal from a microphone such as microphone 14 that captures external ambient sound. In another aspect, the controller may perform any ANC method to generate noise immunity. In yet another aspect, the controller may perform a transparency function, in which the sound played back by the device is a reproduction of ambient sound captured in a "transparent" manner (e.g., as if the headphones were not being worn by the user) by the device's external microphone. The controller processes at least one microphone signal captured by at least one microphone and filters the signal through a transparency filter, which reduces acoustic blockage caused by the audio output device being located on, in, or above the user's ears, while also preserving the spatial filtering effect of the wearer's anatomical features (e.g., head, auricle, shoulders, etc.). The filter also helps to preserve the timbre and spatial cues associated with the actual ambient sound. In one respect, the filter for the transparency function can be user-specific, depending on specific measurements of the user's head. For example, the controller can determine the transparency filter based on the HRTF or equivalent HRIR based on the user's anthropometric measurements.
[0090] On the other hand, controller 10 may perform decorrelation on one or more audio streams to provide a more (or less) dispersed 3D sound field. In some aspects, decorrelation may be activated based on whether the local device is outputting the 3D sound field via headphones or via one or more (out-of-ear) speakers that may be integrated into the local device. On the other hand, controller may perform echo cancellation. Specifically, controller may determine a linear filter based on the transmission path between one or more microphones 14 and one or more speakers 12 and apply the filter to the audio stream to generate an echo estimate subtracted from the microphone signals captured by the one or more microphones. In some aspects, controller may use any method of echo cancellation.
[0091] As described above, the audio spatial controller 20 is configured to render one or more input audio streams associated with the canvas visual representation spatially as individual virtual sound sources. However, on the other hand, the controller can render a mixed spatial representation of one or more input audio streams from a remote participant in the canvas as a single virtual sound source. In one aspect, it may have only one such grouped mix, or it may use multiple mixes. As described herein, the spatial audio controller may have a predefined number of output audio channels for personalized spatial rendering. However, in some cases, the controller may determine that more virtual sound sources are needed than those output by the local device. For example, additional remote participants may join a communication session, and the controller may determine that one or more of their respective input audio streams need to be rendered individually, as described herein. As well, the controller may determine that existing remote participants who were not previously rendered individually may now require their individual virtual sound sources. For example, during a communication session, the controller may determine that a roster of remote participants is moving from the roster to the canvas area based on the criteria presented herein. Therefore, if output audio channels are required, but there are not enough available channels (e.g., existing spatial rendering of existing virtual sound sources has reached the total number of output audio channels), the audio spatial controller can begin to spatially render the canvas input audio streams to a single virtual sound source. In one aspect, this determination can be based on the location of visual representations within the canvas area. For example, the controller can perform vector quantization relative to the location of visual representations within the GUI to group (or as a blend) one or more input audio streams of adjacent visual representations. In another aspect, the controller can group the streams based on the distance between visual representations within the canvas area (e.g., within a threshold distance). In yet another aspect, the controller can allocate predefined regions within the GUI (e.g., its canvas area), thereby blending the streams associated with visual representations within those regions. Further details regarding regions are described herein.
[0092] Once the grouping of one or more input audio streams is determined, the controller (e.g., its spatial parameter generator 25) can determine one or more spatial parameters for the mix (e.g., in a manner similar to determining spatial parameters for a list representation, such as determining a weighted combination of spatial parameters of the streams in the mix), and transfer the space to the matrix / router 26 to mix these input audio streams and transmit the mix as one of the output audio channels. In one aspect, the mixing of audio streams can be dynamic as the visual representation within the GUI changes.
[0093] On the other hand, output audio channels can be predefined, specifying whether the channel supports individual input audio streams or a mixture of input audio streams. In this case, the controller can be configured to accommodate multiple individual virtual sound sources and multiple mixed sound sources. In one aspect, these quantities can be used to determine which input audio stream to mix and how much to mix.
[0094] In one aspect, the audio spatial controller 20 can dynamically render the input audio stream, such that the controller adjusts the rendering based on changes in the audio system 1 (e.g., its local devices). As described herein, spatial parameters can be generated based on the translation range of the speakers of the local device. In one aspect, the translation range can be varied based on certain criteria. For example, the translation range can be based on the physical arrangement of the speakers integrated within the local device. Therefore, the translation range can also change when the local device changes its position and / or orientation. Thus, the audio spatial controller can be configured to adjust the spatial rendering based on any changes to the local device, such as changes in orientation and / or changes in the aspect ratio relative to the GUI of the display screen. Figures 12 to 15 More details are described regarding adjusting spatial rendering based on changes in the aspect ratio of the local device and / or the GUI.
[0095] Figure 5 , Figure 6 , Figure 9 , Figure 13 and Figure 15 These are flowcharts for processes 30, 90, 80, 140, and 170, which are used to perform one or more audio signal processing operations to spatially render the input audio stream of a communication session. In one aspect, the execution of these processes can be performed by one or more devices of the audio system 1, such as... Figure 1 As shown in the diagram. For example, at least some of the operations of these processes can be performed by local device 2 (e.g., its controller 10). On the other hand, at least some of these operations can be performed by another device, such as a remote server communicatively coupled to the local device.
[0096] about Figure 5This diagram is a flowchart of one aspect of the process 30 used to determine whether the input audio stream should be rendered individually as individual virtual sound sources or mixed and rendered as individual virtual sound sources, and to determine the arrangement of these virtual sound source locations. In one aspect, the operations described in the process are performed by one or more operation blocks of the controller 10, such as... Figure 4 As described herein. In one aspect, prior to the start of the process, local device 2 has established a communication session with one or more remote devices 3, as described herein. The process begins with controller 10 receiving communication session data (e.g., input audio streams, input video streams, and / or VAD signals) from each of the one or more remote devices that have established a communication session with local device (at box 31). Controller 10 determines a set of communication session parameters for each remote device (e.g., based on the input video streams and / or VAD signals) (at box 32). For example, a video renderer may determine one or more VAD parameters, the size of the visual representation of each input video stream within the GUI, the position of the visual representation, reverberation level, etc., as parameters. For each input audio stream, controller 10 determines, based on the set of communication session parameters, whether the input audio stream is to be 1) rendered separately relative to other received input audio streams, or 2) rendered as a mixture of input audio streams with one or more other input audio streams (at box 33). Controller 10 determines, based on the set of communication session parameters, the arrangement of 1) individual virtual sound sources, each comprising a separately rendered input audio stream, and 2) the locations of one or more individual virtual sound sources, each comprising a mixture of input audio signals (in box 34). Specifically, in determining the arrangement, the controller determines the spatial parameters of each input audio stream (e.g., based on the location of a visual representation of the corresponding input video stream displayed in the communication session GUI on display screen 13). Additionally, based on the arrangement of the visual representation, the controller (e.g., its matrix / router 26) can perform a matrix mixing operation to mix the one or more input audio streams to produce a mixture of input audio streams. Controller 10 renders the space of each input audio stream determined to be rendered separately as containing only the individual virtual sound sources of that input audio stream (in box 35). Controller 10 also renders the mixing space of each input audio stream as containing a single virtual sound source of the mixture (in box 36).
[0097] As described above, controller 10 manages the allocation of output audio channels to individual input audio streams based on one or more criteria (e.g., whether a remote participant is actively speaking). This is to ensure that the number of allocated output audio channels (some of which include personalized input audio streams and / or mixtures of one or more input audio channels) does not exceed a predefined number, in order to optimize computational resources. However, in some cases, the controller may determine that the number of input audio streams determined to be rendered individually as individual virtual sound sources exceeds the predefined number of audio output channels. This could be due to a large number of remote participants actively speaking (e.g., having VAD parameter values above a threshold). Therefore, as described herein, the controller may allocate one or more input audio streams as a mixture to a single virtual sound source, rather than exceeding the predefined number of output audio streams. Additionally, the controller can adjust the arrangement of virtual sound source locations. As described herein, the controller can group one or more input audio streams within a canvas area and render that group to a virtual sound source. On the other hand, the controller is able to adjust the arrangement of virtual sound source locations and visual representations in a grid-like manner to accommodate more remote participants. Figure 6 The process described is based on adding remote participants within a communication session to determine whether to adjust the layout, which could result in exceeding the predefined number of output audio channels if their respective input audio streams are spatially rendered individually as virtual sound sources.
[0098] Specifically, Figure 6 This is a flowchart of one aspect of a process 90 for defining user interface (UI) areas (e.g., in a grid pattern) and spatially rendering input audio streams associated with representations within the UI areas, where each UI area includes one or more visual representations. Process 90 begins with the controller receiving input audio and video streams from each of a first set of remote devices in a video communication session with the local device (at box 91). For each input video stream, the controller displays a visual representation of the input video stream in a GUI on display screen 13 (at box 92). Specifically, video renderer 22 receives the input video stream (and VAD parameters) and determines the arrangement of the visual representations. Once determined, the renderer displays the video stream on the display. The controller spatially renders the input audio streams to output one or more individual virtual sound sources and / or a single virtual sound source that includes a mixture of the input audio streams (at box 93). Thus, at this point, the local user 40 can perceive the audio of remote participants at the location of the virtual sound sources within the physical environment, as in… Figure 3 As shown in the diagram. For example, the controller can spatially render at least one input audio stream to output (e.g., individual) virtual sound sources including the input audio stream through one or more speakers 12. In one aspect, the number of individual virtual sound sources may be less than a predefined number.
[0099] Controller 10 determines that a second group (one or more additional) of remote devices has joined the video communication session (at box 94). In one aspect, this determination may be based on requests received by session manager 21 from one or more additional remote devices to participate in a pre-existing communication session. In response, the session manager may accept the request and establish a communication channel with the remote devices to begin receiving session data. In another aspect, this determination may be based on the session manager receiving session data from newly added remote devices in the communication session. For example, the session may be an "open" session, where remote participants are free to join (e.g., without authorization from the local device and / or other remote devices already participating in the session). In response to determining that the second group has joined the call, the controller receives input audio and video streams from each of these remote devices.
[0100] Controller 10 determines whether the local device supports additional individual virtual sound sources for one or more input audio streams from the second group of remote devices (at decision box 95). Specifically, as described herein, the local device (e.g., its controller 10) can be configured to spatially render a predefined number of input audio streams as individual virtual sound sources. In the existing configuration (e.g., during a video communication session with the first group of remote devices), controller 10 is already rendering a number of individual input audio streams that may be less than the predefined number. Therefore, individual audio stream selector 27 can receive additional group communication session parameters and determine whether additional input audio streams should be rendered as individual virtual sound sources. Specifically, the selector determines whether the number of input audio streams from the first and second groups of remote devices that are determined to be spatially rendered individually is greater than the predefined number.
[0101] If so, the controller determines that the local device does not support the aggregation of personalized streams that the communication session may require. In response, the controller 10 defines several (or one or more) UI areas located in the GUI, each UI area including one or more visual representations of one or more input video streams from a first set of remote devices, a second set of remote devices, or combinations thereof to be displayed in the UI area (at box 97). For example, the controller 10 (e.g., its video renderer 22) may display all visual representations associated with the first set of remote devices and the second set of remote devices in a grid manner (e.g., in one or more rows and one or more columns). In one aspect, the visual representations may be evenly spaced between the edges of the display and between each other, and / or may have the same size. The controller 10 establishes a virtual grid of UI areas on the GUI, where each UI area contains one or more visual representations. For example, when establishing the grid, the controller may assign one or more adjacent visual representations within the GUI to each UI area. In another aspect, the controller may define UI areas based on the number of visual representations and / or input audio streams received from the two sets of remote devices. For example, if the predefined number of individual virtual sound sources (or output channels) is four and there are eight input audio streams, the controller can evenly distribute (or distribute) the two input audio streams to each UI area, so that the number of limited UI areas does not exceed the predefined number.
[0102] Once these UI zones are defined, the controller, for each UI zone, renders a mixing space (at box 98) via speaker 12 of one or more input audio streams associated with one or more visual representations included within the UI zone, as a virtual sound source. Specifically, the controller can render the mixing space such that each zone is associated with its own virtual sound source. For example, the virtual sound source for each zone can be located at a position within the UI zone displayed on the screen. Specifically, the virtual sound source can be positioned at the center of the UI zone such that audio from the input audio streams from that zone is perceived by the local user as originating from that zone. In one aspect, the controller can dynamically position the virtual sound source based on the speech activity of one or more remote participants associated with that zone. For example, the controller can place the virtual sound source on a visual representation associated with one of the input audio streams in the mixture of the zone's input audio streams, which has a signal energy level above a threshold (e.g., associated with speech activity above a threshold, indicating that a remote participant is speaking).
[0103] However, if the local device does support rendering additional input audio streams separately, the controller space renders the additional input audio streams to output one or more additional individual virtual sound sources (at box 96). Specifically, the controller can add a visual representation to the GUI of the communication session (e.g., a canvas area) and output the additional stream as an individual source. Otherwise, the controller may rearrange the arrangement of the virtual sound source positions. In addition to adding individual virtual sound sources, the controller can also add one or more input audio streams to a mix of input audio streams that are rendered as a single virtual sound source in a list area of the GUI.
[0104] In one aspect, controller 10 can requalify the UI area based on whether a remote device is added to or removed from the communication session. For example, in response to a third group of remote devices joining the session, the controller can requalify the UI area by at least one of the following: 1) adding a visual representation of the input audio stream from the third group to the already qualified UI area; 2) creating one or more new UI areas (e.g., possibly including at least one input video stream from the third group); or 3) a combination thereof. Thus, the controller can dynamically requalify the UI area as needed. In another aspect, the controller can qualify the UI area based on user input from a local device (e.g., the user selects a menu option that, when selected, instructs the controller to qualify the UI area as described herein). In some aspects, controller 10 can switch between qualifying a UI area for spatial rendering of the input audio stream and providing a canvas area and a list area to the GUI (e.g., based on whether a predefined number of output audio channels has been exceeded).
[0105] Figure 7 The diagram illustrates several stages 70 and 71 according to one aspect, wherein the local device 2 defines a user interface (UI) area for one or more visual representations of the input video stream, and spatially renders one or more input audio streams associated with the defined UI area. Specifically, the first stage 70 illustrates a communication session GUI 41 when the local user participates in the session, and the location 62 of the virtual sound source for the session. Specifically, the diagram illustrates the visual representations 44-48 of the remote participant and the visual representation of the local user 49 in arrangement 61 of the GUI 41 displayed on the local device display 13. Specifically, this arrangement differs from the orientation of the local device. Figure 3 Arrangement 51. For example, Figure 3 The local device shown is portrait-oriented, while the device shown in this figure is landscape-oriented (e.g., rotated 90° about the central axis of the local device). In this arrangement 61, the representations within canvas area 42 are distributed more within the GUI than they are in... Figure 3The positions in arrangement 51 are wider (e.g., wider along the X-axis). Additionally, list area 43 is displayed on one side (right side) of the GUI, where visual representations are stacked in a column rather than a row. Furthermore, the arrangement 62 of virtual sound source positions is similarly arranged to the arrangement 61 of visual representations. For example, sound source 55 corresponding to visual representation 45 is higher in the vertical direction and positioned between sound sources 54 and 56 corresponding to visual representations 44 and 46, respectively, which are lower in the vertical direction and on either side of representation 45. As described herein, the arrangement 62 of virtual sound sources can be proportionally larger than the arrangement of visual representations (e.g., wider and higher relative to the representations). However, on the other hand, these sound sources can be located on (or near) their respective representations, whereby sound sources 54-56 are centered on their respective visual representations 44-46, and the virtual sound source 53 of the list is positioned at the center of both list representations 47 and 48.
[0106] In one respect, the position of list area 43 in this arrangement can be consistent with that in Figure 3 The same positioning is used in the arrangement 51. For example, the list visual representations are not arranged in a stacked column, but can be arranged in a row, for example, at the bottom of the GUI 41. On the other hand, the visual representations within the list area can be optional, such that remote participants may not be placed in the list area if not needed. For example, when there are enough output audio channels and / or each of these remote participants meets the criteria for being within the canvas area, the controller 10 can position all remote participants within the canvas area. In this case, the GUI may not include the list area.
[0107] Phase 71 illustrates the result of more remote participants joining the communication session, and in response, the controller 10 of the local device defines UI areas, each containing one or more visual representations. Specifically, as shown, three new remote participants 73-75 have joined the communication session. In one aspect, the controller may have determined that one or more of these new participants should be spatially rendered as individual virtual sound sources. However, in another aspect, the controller may have determined that the local device might exceed the optimal (predefined) number for individual rendering if rendered individually. Therefore, the controller has defined four (simulated) areas 76-79 as a grid, each of which is a UI area. The arrangement 63, including the visual representations of the three newly added remote participants, is also arranged in a grid manner, with two visual representations assigned to each area. Additionally, the sizes of the existing visual representations have been adjusted so that all representations have the same size.
[0108] Along with the rearrangement of the visual representation, the controller is outputting four distinct virtual sound sources 65-68, which are positioned within a new arrangement 64. Specifically, the virtual sound sources 65-68 have been arranged in a grid similar to the grid of the simulation areas 76-79. In one aspect, the arrangement 64 of the virtual sound source locations can be scaled to the arrangement of the UI areas 76-79. In another aspect, the virtual sound source locations can be on (or adjacent to) their respective UI areas. In this case, each virtual sound source can be positioned at the center of its respective area. In one aspect, the arrangement 64 of the virtual sound source locations can be static during a communication session, such that remote participants sharing the virtual sound sources (e.g., remote participants 74 and 75 sharing sound source 68) have the same spatial cues when they speak. In another aspect, the virtual sound source locations in the UI areas can dynamically change their positions based on which remote participant in that area is speaking. For example, virtual sound source 68 can move horizontally depending on whether remote participant 74 or 75 is speaking.
[0109] In one respect, the controller can arrange the visual representation and / or virtual sound sources differently. For example, the controller can define areas of different sizes within the GUI, each associated with one or more virtual sound sources that include one or more input audio streams.
[0110] As described above, the controller determines spatial parameters based on the location of the visual representation within the GUI. In one aspect, the controller can spatially render an input audio stream that is not associated with a visual representation displayed (or visible) within the GUI. Specifically, the video renderer may not display one or more visual representations associated with the rendered input audio stream. For example, the video renderer may determine that the GUI does not have sufficient white space to support the display of one or more additional visual representations (e.g., without overcrowding the display). As another example, a communication session may have more candidate remote participants than can be displayed within the list area. For example, refer to... Figure 3 The list area includes two visual representations 47 and 48. However, if several additional remote participants are added, the visual representations (at this size) will not fit the width of the GUI. Similarly, the video renderer may not be able to receive input audio streams from one or more remote devices. In this case, the spatial parameter generator 25 can determine the location of a virtual sound source that does not include an associated visual representation to be positioned outside (or to one side of) the GUI (and / or to one side of other virtual sound sources). Therefore, refer to... Figure 3 When rendering an "invisible" virtual sound source in layout 50, the sound source can be positioned to the right of GUI 41.
[0111] As described herein, controller 10 (e.g., its audio spatial controller 20) is configured to determine spatial parameters indicating the location of virtual sound sources that indicate one or more input audio streams based on a visual representation displayed in GUI 41. These spatial parameters may include translation angles, such as azimuth and elevation translation angles relative to at least one reference point in space (e.g., the location of a local user or the user's head), the distance between the virtual sound source and the reference point, and the reverberation level. Figure 8 This diagram illustrates the translation angle used to render the input audio stream at the corresponding virtual sound source location within the 3D sound field. Specifically, this figure shows the translation angle corresponding to, for example, the virtual sound source location of the input audio stream in the 3D sound field. Figure 7 The arrangement 62 shows the virtual sound source positions of the visual representation 61, as well as the azimuth and elevation translation angles of each sound source and the distance between the sound source and the reference point. As shown, the arrangement 62 is defined by the translation range of one or more speakers used to output the virtual sound sources. Specifically, these boundaries include an azimuth translation range -φ - +ω, which spans the width of the arrangement along the X-axis, and an elevation translation range -φ - +β, which spans the height of the arrangement along the Y-axis. In one aspect, the speakers can generate virtual sound sources at any location within the defined range. In some aspects, the boundaries of the arrangement correspond to the positions of the speakers. For example, these boundaries can correspond to the size of a local device, in which case the local device can be configured to generate virtual sound sources in front of the device (e.g., its display). Specifically, the azimuth translation range can span the width of the local device (e.g., its display), and the elevation translation range can span the height of the local device (e.g., its display). Thus, the positions of virtual sound sources 54-56 can be located on (or in front of) their corresponding visual representations 44-46, as described herein. On the other hand, the controller can limit these pan ranges (relative to a wider range of possible pans) in order to position the virtual sound source within a specific area of the local device's display screen (e.g., in front of, beside, and / or behind the local device's display screen).
[0112] Additionally, the figure illustrates the translation angles relative to a reference point 99 in space (e.g., within a physical environment). For example, the azimuth translation range 100 includes the azimuth angles of four virtual sound sources 53-54 (or L1-L4 respectively) relative to / or at the reference point 99 along the horizontal X-axis. Specifically, the reference point is the vertex of each angle, and each azimuth translation angle extends along the horizontal X-axis away from the 0° reference axis Z-axis (e.g., or toward -φ or +ω). Similarly, the elevation translation range 101 shows each of the elevation angles of the four virtual sound sources relative to the reference point along the vertical Y-axis. Again, the reference point is the vertex of each angle, and each elevation angle extends along the vertical Y-axis away from the 0° reference axis (e.g., or toward -φ or +β). Furthermore, for the four sound sources L1-L4, the distance (along the Z-axis) between the reference point and each of these virtual sound sources is shown as D. L1 -D L4 Therefore, when rendering a space, the virtual sound source corresponding to the spatial parameters will be perceived by the local user as originating from a certain azimuth angle, elevation angle, and distance from the local user, in order to provide a more three-dimensional (3D) spatial experience.
[0113] Figure 9 This is a flowchart of one aspect of a process 80 for determining one or more spatial parameters based on a corresponding visual representation of location, which indicate the location of the input audio stream to be spatially rendered as a virtual sound source (3D). In one aspect, the spatial parameter generator 25 of the controller 20 can perform at least some of these operations to determine the spatial parameters of each input audio stream of the communication session. This process will refer to... Figures 10 to 11B The descriptions, each of which illustrates a different example of how the location of a virtual sound source is mapped to a corresponding visual representation.
[0114] Process 80 begins with generator 25 selecting a set of communication session parameters for the input audio stream (at box 81). For example, the generator may receive all N groups from a data structure from the session manager and may select the first group. As described herein, session parameters may include information about visual representations, such as the size, position, salience value, and associated VAD parameters of these visual representations. The generator uses this set of communication session parameters to determine the position of the visual representations of the input audio stream associated with it (at box 82). For example, session parameters may include positional information (e.g., X, Y coordinates) of the visual representations relative to the GUI and / or relative to the display on which the GUI is displayed. The generator determines one or more translation ranges (e.g., azimuth translation range, elevation translation range, etc.) for one or more speakers (at box 83). Specifically, these angular ranges may correspond to the maximum (or minimum) range within which the virtual sound source can (e.g., optimally) lie in space, such as the azimuth translation range -φ - +ω and the elevation translation range -φ - +β, said ranges being within... Figure 8 As shown in the diagram. In one aspect, the translation range can be based on the physical location and / or orientation of the speaker (and / or the device in which the speaker is housed). For example, (e.g., when the speaker is integrated within a local device), the translation range can be based on the orientation of the device. For example, when the local device is in a longitudinal orientation (e.g., as shown in the diagram). Figure 3 As shown), the device's speaker can have a narrow azimuth translation range spanning the width of the device, while when the local device is in a lateral orientation (e.g., as shown). Figure 7 As shown in the diagram, the speaker of the device can have a wide azimuth translation range (e.g., relative to a reference point in space) spanning the width of the device. Therefore, the generator can determine the orientation of the speaker (e.g., the local device housing it) and, relative to that orientation, determine the azimuth translation range (e.g., spanning the horizontal X-axis) and the elevation translation range (e.g., spanning the vertical Y-axis). For example, after determining the orientation (e.g., based on IMU data from IMU 16), the generator can perform a table lookup on a data structure that stores predefined translation ranges associated with one or more orientations of the device. Figure 12 and Figure 13 More details are described regarding determining the translation range based on device orientation. The generator determines spatial parameters (e.g., azimuth and elevation) (at box 84) that indicate the position of the virtual sound source within that one or more translation ranges of the input audio stream (e.g., relative to a reference point in space) based on the determined visual representation's position within the GUI.
[0115] For example, the generator can execute one or more methods for determining the spatial parameters of each virtual sound source, which can be determined based on the physical location of the local user relative to the orientation (or position) of the local device. For example, the generator can use sensor data (e.g., image data captured by camera 15) to determine the position and / or orientation of the local user (e.g., the local user's head) relative to the display screen. The generator can determine at least one of the azimuth and elevation angles of each visual representation displayed on the GUI of the local device from the local user.
[0116] On the other hand, the generator can determine spatial parameters based on a linear mapping of angles to the position of the visual representation relative to the size of the GUI within the display screen. For example, referencing Figure 10 According to one aspect, the generator can determine spatial parameters by using one or more functions to map the position of the visual representation to the angle at which the input audio stream is to be rendered. Specifically, the generator can map the position of the virtual sound source based on the position of the visual representation displayed in the GUI relative to the size of the GUI. This figure illustrates a GUI 41 (displayed on a local device's display screen) comprising a communication session of two visual representations 44 and 45, each with a remote participant having a session with a local user. The GUI has a width X along the X-axis and a height Y along the Y-axis. In one aspect, the size of the GUI can be based on the size of the GUI displayed on (or relative to) the display screen. In another aspect, the size of the GUI will be equal to the size of the display screen when the GUI covers the entire display screen. Displayed on each visual representation are the analog center points of these representations, which represent the position of the representation within the GUI (e.g., X, Y coordinates). For example, the position of representation 44 L1 is (X... L1 , Y L1 ), and indicates that the position of 45 L2 is (X L2 , Y L2 On the other hand, other locations can be defined. For example, a simulated point can be positioned above a specific part of the visual representation, such as the part showing the mouth of a remote participant (which can be identified using an object recognition algorithm).
[0117] In one aspect, the function used to map the position of a virtual sound source (e.g., its translation angle) to a visual representation is a linear function of the translation angle relative to the size of the GUI. For example, the azimuth function 111 is a linear function of the fractional relationship between the azimuth translation range -θ - +ω and the X position with respect to the total width X of the GUI. Thus, the azimuth translation range begins on the left side of the GUI (e.g., in the case of X=0) and ends on the right side of the GUI. The elevation function 113 is a linear function of the fractional relationship between the elevation translation range -φ - +β and the Y position with respect to the total height Y of the GUI. Thus, the elevation translation range begins at the bottom of the GUI (e.g., in the case of Y=0) and ends at the top of the GUI. These relationships between the translation range and the size of the GUI allow the generator to map the position relative to the GUI, regardless of the size of the GUI and / or the size of the translation range.
[0118] To determine the translation angle of the visual representation, the generator can apply a fractional relationship of the visual representation's position as input to one or both of a linear function. For example, to determine the azimuth of a virtual sound source, the generator can use the x-coordinate of the visual representation's position within the GUI as input to an azimuth translation range function. Specifically, as shown in the figure, the fractional positional relationship of the visual representation (X... L1 / X and X L2 / X) is mapped to X L1 / X and X L2 The azimuth translation angle at / X, where the linear function intersects, is shown in Figure 111. The resulting mapping of these fractional relationships to the azimuth is shown by the azimuth translation range 112, which illustrates the azimuth angle -θ of L1 at reference point 99. L1 And the azimuth of L2 + ω L2 Similarly, to determine the elevation translation angle of the virtual sound source, the generator can use the y-coordinate of the visual representation's location as input to (e.g., a separate) elevation translation range function. Specifically, the fractional positional relationship of the visual representation (Y... L1 / Y and Y L2 / Y) is mapped to Y L1 / Y and Y L2 The elevation angle translation angle at / Y intersects with the linear function 113. The resulting mapping of these fractional relationships to the elevation angle is shown by the elevation angle translation range 114, which shows the elevation angle of L1 at the reference point - φL1, the elevation angle of L2 + β. L2 Side view.
[0119] On the other hand, spatial parameters can be determined based on the viewing angle of the local user's (predefined) location (e.g., in the case where the reference point in the user or space is a vertex relative to its defined angle). Figure 11A and Figure 11BThis illustrates an example of determining spatial parameters based on one aspect by using one or more functions to map the viewing angle position of a visual representation to the translation angle of the input video stream to be rendered. Specifically, the generator can map the position of the visual representation as the viewing angle at a reference point to one or more translation angles within one or more translation ranges. (Reference) Figure 11A This figure illustrates GUI 41, which includes two visual representations 44 and 45, as shown. Figure 9 As shown in the diagram. However, here, instead of determining a fractional relationship between the position of the visual representation and the total width / height of the GUI, the generator determines an estimated viewing angle of the GUI relative to a reference point 69 (e.g., at a predefined location in space). For example, the GUI has an estimated azimuth viewing range 115 between -θ' and +ω', spanning the width W' of the GUI and along the X-axis. This estimated azimuth viewing range is shown as a top view, where the reference point 69 is located in front of the GUI (or display screen) at a distance D' from the GUI (or display screen), where the GUI has a width of W'. In one aspect, this estimated azimuth viewing range is a predefined viewing range for the local user when the user is looking at the GUI while it is displayed on the display screen. In some aspects, the positions of W', D', and / or the reference point can be predefined; for example, D' could be the optimal viewing position for the local user. In other aspects, these dimensions can be determined by the generator. For example, the generator can obtain sensor data from one or more sensors (e.g., image data from camera 15, proximity sensor data, etc.) and determine the distance the local user is positioned relative to this sensor data. The generator can determine the width of the GUI currently displayed on the screen, relative to its width. Knowing the positions of D', W', and the reference point, the generator determines the azimuth viewing angle of L1 as –θ'. L1 The azimuth viewing angle of L2 is +ω' L2 These angles are the angles from their corresponding visual representations on the GUI to the reference point.
[0120] To determine the (actual) azimuth translation angle, the generator can apply the viewing angle as input to one or more linear functions. For example, this figure shows azimuth function 116, which is a linear function of the azimuth translation range -θ - +ω relative to the estimated azimuth viewing range -θ' - +ω'. The generator will use the viewing angle -θ' L1 and +ω' L2 Mapped to the actual azimuth angles intersecting the linear function. The mapping of these angles is shown by the azimuth angle translation range 117, which is represented at reference point 99 as -θ. L1 and +ω L2 In one respect, reference point 99 may be the same as reference point 69 (e.g., at the same location in space relative to the local device).
[0121] refer to Figure 11B This figure relates to determining the elevation translation angle based on the estimated elevation viewing angle relative to reference point 69. Specifically, the controller can perform actions such as... Figure 11A Similar operations are described in the text to determine the elevation angle translation angle. For example, as shown in the figure, the estimated elevation viewing range 118 is between –φ' - +β', spanning the height H' of the GUI and along the Y-axis. Specifically, this viewing range is a side view of the GUI (or display screen) and a reference point 69 at a (predefined) distance D'. In one aspect, H' can be a predefined height of the GUI, or it can be the current height of the GUI relative to the display screen. Based on this reference point, the generator determines the elevation viewing angle of L1 as +β'L1 and the elevation viewing angle of L2 as –φ'L2. To determine the actual elevation angle, the generator applies the elevation viewing angle as input to the elevation angle translation range –φ - +β relative to the elevation viewing range –φ' - +β' elevation angle function 119. The generator maps the viewing angles +β'L1 and –φ'L2 to the actual elevation angles intersecting with the function 119. The mapping of these angles is shown by the elevation angle translation range 120. In one aspect, the generator can perform Figures 10 to 11B Any of the methods described herein can be used to map the location of a visual representation to the location of a virtual sound source.
[0122] Additionally, the generator can determine other spatial parameters, such as the distance between the virtual sound source and the reference point, based on communication session parameters. As described herein, the distance between the virtual sound source and the local user can be based on the size and / or location of the visual representation. For example, in a physical dialogue where a closer person's voice is compared to a farther person's, the generator can assign a shorter distance to a smaller visual representation that is farther from a larger visual representation. In one aspect, the distance can be based on the location of the visual representation within the GUI. For example, a higher visual representation within the canvas area (e.g., along the Y-axis) might be given a shorter distance than a visual representation that is farther below the Y-axis. In some aspects, visual representations within the list can be assigned the farthest distance relative to all canvas visual representations. In another aspect, the distance can also be based on the VAD parameter. For example, a remote participant associated with a VAD parameter indicating a high signal energy level (e.g., above a threshold) can be assigned a closer distance than a remote participant with a lower VAD parameter. In some aspects, the generator can define the reverberation value for each of these input audio streams based on the same criteria described above. For example, a remote participant within the list area can be assigned a high reverberation value to make the sound more diffuse to the local user.
[0123] As described herein, the controller may apply one or more linear functions to determine the translation angle. In some respects, one or more of these functions may be a more general nonlinear function of the translation angle or a piecewise linear function (e.g., relative to a fractional relationship and / or the viewed translation angle, as described herein).
[0124] Return to Figure 9 The controller determines whether there are any additional sets of communication session parameters for which one or more spatial parameters have not yet been determined (at decision box 85). If so, the controller selects another set of communication session parameters for another input audio stream that has not yet been analyzed to determine one or more spatial parameters. Otherwise, the controller determines whether any input audio stream is to be output as a mix (at decision box 87). For example, the controller may determine whether any input audio stream is associated with a remote participant in the list area. In some aspects, this determination may be based on whether selector 27 has assigned one or more input audio streams to a single output audio stream (e.g., this could be when there are no more individual output audio streams for the canvas remote participant to have individual virtual sound sources). In other aspects, this determination may be based on the output audio stream to which the input audio stream has been assigned, as described herein. If so, the controller determines new (or different) spatial parameters for the input audio stream to be rendered as a mix, indicating a specific location, based on at least some of the spatial parameters of the input audio stream to be rendered as a mix (at box 87). As described herein, the controller spatially renders the mix to output a single virtual sound source that includes the mix. Therefore, the controller determines a set of spatial parameters for a single virtual sound, which can be based on at least some of the determined spatial parameters. For example, the new spatial parameters can be a weighted combination of at least some of the spatial parameters. In this case, a single virtual sound source can be located at the center of an analog virtual sound source, generating analog virtual sound sources that are mapped to multiple locations based on corresponding visual representations within the list area if each of these input audio streams is to be spatially rendered as an individual virtual audio stream. On the other hand, the determined spatial parameters can be based on a set of predetermined data, rather than new spatial parameters. On the other hand, the spatial parameters can not be based on predetermined data, but rather the spatial parameters can be associated with a specific location in space. For example, the spatial parameters can indicate azimuth and elevation angles of 0°, such that the spatial audio of the list is located exactly in front of the local user.
[0125] In some respects, the controller can determine the spatial parameters of the mixed input audio streams differently. For example, instead of determining the spatial parameters relative to individually determined spatial parameters, the controller can combine communication session parameters (e.g., position / size, distance, VAD parameters, salience values, etc.) for at least some of the mixed input audio streams. For example, the controller can determine the average of at least some of these parameters. Once the combined (or jointed) communication session parameters are determined, the controller can determine the spatial parameters as described herein.
[0126] Once the controller determines the spatial parameters, the input audio stream is spatialized according to the data in order to output one or more virtual sound sources, each of which includes one or more input audio streams, as described herein.
[0127] In one respect, the position of the virtual sound source can be changed based on one or more criteria. For example, as described herein, spatial parameters indicating the position of the resulting virtual sound source are determined for spatial rendering of the input audio stream. As described herein, the determined spatial parameters can depend on the translation range of the local device. Therefore, when the translation range changes, the local device can adjust the currently output virtual sound source to adapt to the change. For example, the translation range can change due to a change in orientation of the local device. As another example, the translation range can be based on the aspect ratio of the GUI of the communication session displayed on the local device's screen. The following figures illustrate adjusting the virtual sound source based on changes in the local device's translation range.
[0128] Figure 12 Several stages 130 and 131 are illustrated, in which the translation range of one or more speakers is adjusted based on the local device rotating from a longitudinal orientation to a lateral orientation. For example, each stage illustrates the GUI 41 of the local device 2 during a communication session, and the corresponding arrangement of the virtual sound source positions, where each arrangement shows several translation angle ranges. Specifically, each stage shows an azimuth translation range of -θ - +ω along the X-axis, and an elevation translation range of -φ - +β along the Y-axis. As described herein, one or more of these ranges can be changed based on the device's orientation.
[0129] Phase 130 shows a local device oriented longitudinally, with its height along the Y-axis greater than its width along the X-axis. An arrangement 50 of virtual sound source locations is also shown, illustrating four virtual sound sources 53-56. Specifically, as described herein, for each of the sound sources 54-56, the local device outputs an input audio stream as a virtual sound source at a location within the arrangement 50 relative to a reference point outside the local device (e.g., the point where the local user is located, or a predefined point as described herein). Additionally, the local device is outputting a mixture of the input audio streams as a single virtual sound source 53. Besides showing the locations of the virtual sound sources, arrangement 50 also shows the translation range of the local device in this longitudinal orientation. Specifically, the azimuth translation range is -θ. P - +ω P And the range of elevation angle translation is -φ P - +β P .
[0130] The second stage 131 shows the result of rotating the local device 90° around the Z-axis. Specifically, the local device has been rotated to a lateral orientation, where the width along the X-axis is greater than the height along the Y-axis. Additionally, the translation range of the local device has also changed. As shown in the figure, the azimuth translation range is -θ. L - +ω L , compared to -θ P - +ω P Wider (e.g., with a larger range), elevation angle range of -φ L - +β L , compared to -φ P - +β P Narrower (e.g., with a reduced range). In one aspect, the change in translation range can be based on the components or design of the local device. For example, the translation range can be defined based on the number and / or position of the speakers in the local device. The translation range can also rotate when the device is rotated to a new orientation. In another aspect, the translation range can be defined by a controller and can be adjusted by the controller in response to determining that the orientation of the local device has changed. For example, upon determining that the device is now in a lateral orientation, the controller can determine the translation range for this orientation (e.g., by performing a table lookup on a data structure that associates one or more translation ranges with orientations), and then use the determined translation range for spatial rendering, as described herein.
[0131] Additionally, in response to a change in the orientation of the local device to a new lateral orientation, one or more positions of the virtual sound sources are adjusted relative to a reference point along one or more axes. Specifically, due to the rotation of the local device, the virtual sound sources are positioned in arrangement 62. As described herein, the positions of the virtual sound sources have been adjusted such that they are distributed wider along the X-axis and narrower along the Y-axis compared to the virtual sound sources that were in arrangement 50 when the local device was in a longitudinal orientation. Therefore, the local user can perceive the virtual sound sources differently based on the orientation of the local device.
[0132] Figure 13 This is a flowchart of one aspect of a process 140 for adjusting the position of one or more virtual sound sources based on a change in one or more translation ranges of one or more speakers, said change being based on a change in the orientation of a local device. In one aspect, reference will be made to... Figure 12 To describe this process: Process 140 begins with the controller 10 of local device 2 receiving one or more input audio streams (and one or more input video streams) from one or more remote devices with which the local device is having a (video) communication session (at box 141). The controller determines a first orientation of the local device (at box 142). For example, the controller 10 may determine the device's orientation based on sensor data, such as IMU data from IMU 16. In one aspect, the IMU data may indicate that the local device is in a longitudinal orientation, such as... Figure 12 As shown in the first phase 130, the controller determines one or more translation ranges (at box 143) for one or more speakers of a local device (communically coupled to or part of it). As described herein, the controller may perform a table lookup on a data structure that associates the orientation and / or position of the speakers and / or the local device with one or more translation ranges. In response, the controller may determine an azimuth (e.g., horizontal) translation range of -θ when the device is in a longitudinal orientation. P - +ω P, The range of translation angle (e.g., vertical) is determined to be -φ. P - +β P The controller determines one or more spatial parameters for each input audio stream, which indicate the position of the virtual sound source within one or more determined translation ranges including the input audio stream, and spatially renders these streams as virtual sound sources at one or more positions within the determined translation ranges (in box 144). For example, the controller can perform... Figure 9 At least some of the operations described in process 80 are used to render the input audio stream space as one or more individual virtual sound sources, and / or to render the mixing space of one or more streams as a single virtual sound source.
[0133] The controller determines whether the local device has a changed orientation (e.g., at decision box 145). Specifically, the controller may determine whether the orientation has changed (e.g., from longitudinal orientation to lateral orientation) based on IMU data, as described herein. If so, the controller determines one or more adjusted translation ranges for the one or more speakers based on the changed orientation (at box 146). For example, re-referencing... Figure 12 The adjusted translation range corresponds to the local device in a laterally oriented position. The controller then adjusts one or more positions of the virtual sound source (at box 147) based on this adjusted translation range. For example, the controller can adjust the azimuth and / or elevation angle of the virtual sound source relative to a reference point. For example, as... Figure 12 As shown, in response to the local device rotating to landscape orientation, the azimuth angle of the virtual sound source 54 is closer to the lower boundary -θ from the azimuth angle position when the local device is in portrait orientation. L Widening. In one respect, to adjust the position, the controller can perform the operation of process 80 to adjust the position. For example, in response to performing the operation based on the local device rotating to a lateral orientation, Figure 12 The virtual sound source position 62 in the diagram spans a wider range along the azimuth angle and a narrower range along the elevation angle. On the other hand, the controller can adjust the position without determining new spatial parameters, as described in process 80. For example, the controller can adjust the virtual sound source position by rotating the position based on the rotation of the device. For instance, the controller can rotate the virtual sound source position by 90° in response to determining that the local device has rotated from longitudinal to lateral.
[0134] On the other hand, the controller can adjust the position based on the difference in translation angles between two or more orientations. For example, the controller can adjust the position proportionally to the difference between the translation angles of a first orientation and a second orientation. For example, referencing... Figure 12 The azimuth angle shift from longitudinal to lateral may increase by 50%, while the elevation angle may decrease by 50%. Therefore, to adjust position, the azimuth and elevation angles can be changed proportionally to the difference between the translation ranges of their respective two orientations.
[0135] In one respect, the position of the virtual sound source can remain in the same position relative to the GUI (e.g., on the display screen) as the local device rotates. For example, see reference. Figure 12 Although the position of the virtual sound source has changed relative to the translation range, the position of the virtual sound source can remain the same relative to the rotated local device due to the increase / decrease of the range, allowing the virtual sound source to rotate with the local device. Therefore, as the device rotates, the local user can perceive the virtual sound source to maintain its position relative to the display screen, causing both the visual representation and its associated virtual sound source to move together.
[0136] On the other hand, once the local device is rotated in the opposite direction, the virtual sound source can move back to its initial position. For example, if the local user rotates the local device -90°, the virtual sound source can return to its initial position, as shown in [the image / video]. Figure 12 As shown in the first phase 130.
[0137] In one aspect, the pan range can be appended (or corresponded) to the edges of the display screen, and then such a pan range is applied to the GUI window within the screen by taking into account the largest magnified version of the GUI window that fills as much of the screen as possible (e.g., where at least two opposite edges of the GUI window coincide or are adjacent to corresponding edges of the display screen). An advantage of such a system is that the pan range is primarily a function of the aspect ratio of the GUI window, rather than the GUI size, thus preserving an audio-spatial image that does not change with GUI size, position, and / or positioning or collapsing if the window is minimized or if it enters picture-in-picture mode. Figure 14A and Figure 14B The diagram illustrates several stages where the translation range is based on the aspect ratio of the communication session GUI, according to some aspects.
[0138] Specifically, as described herein, the local device can define the translation range of one or more speakers based on the aspect ratio of the GUI. For example, Figure 14A Two phases 150 and 151 are shown, where the azimuth translation range is less than the total (possible) azimuth translation range of the local device speaker based on the aspect ratio of GUI 41. Phase 150 shows a communication session GUI displayed on the local device's display screen 13 (when the local device is in a communication session with five remote participants, as described herein). Specifically, the GUI overlays on top of the main screen GUI 152 displayed on the display screen and provides the local user with an interface to execute and / or terminate one or more computer program applications (which may include communication session applications, as described herein). As shown, the communication session GUI is smaller (or has a smaller surface area) than the main screen GUI. In one aspect, the local device can receive input (e.g., via any input device such as a mouse, or via the display screen, which may be a touch-sensitive display screen, as described herein) to adjust the size and / or position of GUI 41 within the main screen GUI 152. As shown, GUI 41 has a current aspect ratio of 4:3.
[0139] Additionally, this stage also illustrates the translation range (e.g., azimuth and elevation) of the local device's speakers. As shown, the total translation range (e.g., the maximum angle at which a virtual sound source can be positioned when spatially rendering the corresponding input audio stream using the local device's speakers) extends to the edge of the display screen. For example, the speaker's azimuth translation range -θ- +ω spans the total width of display screen 13 (along the X-axis), and the elevation translation range -φ - +β spans the total height of display screen 13 (along the Y-axis). Therefore, the local device can position the virtual sound source anywhere on (or in front of) the display screen.
[0140] The second stage 151 illustrates the result of magnifying the analog communication session GUI until both sides of the GUI reach the corresponding edges of the display screen. As shown in the figure, the display screen is showing the analog GUI 153, which has been fully magnified in the Y direction (e.g., the visible portion of the session GUI can no longer be extended further in the Y direction), while separating from the edges of the display screen along the width of the GUI. Since the height of the session GUI extends the height of the display screen, the elevation translation angle remains unchanged; however, since the width of the session GUI is smaller than the width of the display screen, the azimuth translation range is reduced to -θ. w - +ω w This is less than the total azimuth translation range. Therefore, the controller 10 of the local device can accordingly limit the azimuth and elevation translation ranges.
[0141] Figure 14B Two stages, 160 and 161, are shown, where the elevation translation range is smaller than the total elevation translation range of the local device speaker based on the aspect ratio of GUI 41. This stage is similar to... Figure 14A Phase 150 differs in that the aspect ratio of the GUI is 16:9 instead of 4:3. Therefore, unlike Phase 151, Phase 161 of this figure shows that the simulated GUI 154 has fully extended along the width of the display 13, but not yet fully extended along the height of the display. Thus, when the GUI 41 has a larger aspect ratio, the defined azimuth translation range can be equal to the total translation range, while the elevation translation range decreases to -φw - +βw, which is less than the total elevation translation range.
[0142] As described above, the local device (its controller 10) can define the pan range based on whether the GUI aspect ratio is 4:3 or 16:9. Alternatively, the pan range can be defined for any aspect ratio. In one aspect, the controller can determine and / or adjust the spatial parameters of one or more input audio streams (e.g., in response to user input) based on the aspect ratio or whether the aspect ratio has changed. Figure 15 More details are described here.
[0143] Figure 15This is a flowchart of one aspect of a process 170 for adjusting the position of one or more virtual sound sources based on a change in the aspect ratio of a GUI for a communication session. Process 170 begins with controller 10 receiving an input audio stream and an input video stream from a remote device having a video communication session with a local device (at box 171). In one aspect, the controller may also receive other data (e.g., VAD signals), as described herein. The controller displays a visual representation of the input video stream within the GUI (with aspect ratio) of the video communication session displayed on a screen (at box 172). The controller determines the aspect ratio of the GUI for the video communication session (at box 173). Specifically, the video renderer 22 may determine the aspect ratio at which the GUI is being displayed. Based on the aspect ratio of the GUI, the controller determines an azimuth translation range for at least a portion of the total azimuth translation range, and an elevation translation range for at least a portion of the total elevation translation range of the speakers (at box 174). Specifically, the controller may perform… Figure 14A and Figure 14B The operation described herein. For example, a controller (e.g., its video renderer) can simulate zooming (or expanding) the GUI until the width of the GUI has been fully expanded to the width of the display or the height of the GUI has been fully expanded to the height of the display, while maintaining the aspect ratio. In one aspect, this might mean expanding the GUI until at least two sides (or edges) of the GUI are in contact with two comparable sides of the display. For example, the renderer can zoom in on the GUI until the top and bottom edges of the GUI are in contact with the edges of the display (as described in the image). Figure 14A (as shown in the image), or the GUI can be expanded until the side edges of the GUI touch the edge of the display (as shown in the image). Figure 14B (As shown in the diagram). Once zoomed in, the controller can define the translation range based on the size of the simulated GUI relative to the size of the display screen. Specifically, the translation range is defined as spanning the width and height of the GUI, as a function of the width and height of the display screen. In one aspect, if the GUI and the display screen have the same aspect ratio, the determined translation range can be the total translation range of the speaker.
[0144] As described in this article, when determining the pan range while zooming in on the GUI, the range can span both the width and height of the display. On the other hand, these pan ranges can extend beyond the boundaries of the display. In this case, the pan range can be determined based on a percentage of the zoomed-in analog GUI. For example, refer to... Figure 14B The determined azimuth translation range can be 100% of the total azimuth translation range because the simulated GUI has been extended to the side edges of the display. In contrast, the determined elevation translation range can be 70% of the total elevation translation range because the simulated GUI only extends along 70% of the total height of the display.
[0145] The controller determines spatial parameters, as described herein (at box 175), indicating the location of the virtual sound source within a defined azimuth and elevation translation range, based on the position of the visual representation within the GUI. The controller then uses a speaker to spatially render the input audio stream based on these spatial parameters to output the virtual sound source (at box 176) within that azimuth and elevation translation range (e.g., at that location). Thus, in conjunction with the displayed visual representation, the controller outputs the input audio signal as a virtual sound source at a location within the environment (e.g., the location of the local device).
[0146] The controller determines whether the aspect ratio of the GUI has changed (at decision box 177). In one aspect, this determination may be based on whether user input has been received via one or more input devices to change the width or height of the GUI. For example, the controller may receive an indication that the user has performed a click-drag operation with the mouse to manually zoom in (or stretch) the GUI in one or more directions (e.g., by selecting one side and performing a dragging motion away from or towards the GUI) (e.g., via video renderer 22). In another aspect, when the display is a touch-sensitive display, user input may be received when the user performs a touch-drag motion with one or more fingers to resize the GUI. In some aspects, the controller may (e.g., periodically) perform the operations described in box 173 to determine whether the aspect ratio has changed. In response to determining that the aspect ratio has changed, the controller adjusts one or more translation ranges based on the changed aspect ratio (at box 178). For example, the controller may perform the operations described in box 174 to determine the changed (or new) azimuth translation range. For example, when the aspect ratio increases, the adjusted azimuth translation range can extend the width of the display screen, while the adjusted elevation translation range will not fully extend the height of the display screen. Figure 14B As shown in the diagram. The controller then adjusts the position of the virtual sound source (at box 179) based on the adjusted translation range. For example, the controller can determine the spatial parameters according to any of the methods described herein, and then spatially render the input audio signal according to the new spatial parameters.
[0147] On the other hand, the controller can adjust its position based on changes in the adjusted translation range. Specifically, the controller does not recalculate the position of the virtual sound source (e.g., as...). Figure 9 (As described above), but the existing position can be adjusted based on the adjusted translation range. For example, by increasing the aspect ratio (e.g., from 4:3 to 16:9), the azimuth translation range can increase from 60% to 100% of the total azimuth translation range, and the elevation translation range can decrease from 100% to 70%, as... Figure 14A and Figure 14BAs shown in the diagram. In response, the controller can adjust the azimuth angle of the virtual sound source by increasing the angle by 40% and the elevation angle of the virtual sound source by decreasing the angle by 30%. Specifically, the controller can increase the azimuth angle to move the virtual sound source to a wider azimuth position and decrease the elevation angle to move the position to a narrower elevation position (e.g., relative to the 0° reference Z-axis).
[0148] In one aspect, the panning range is determined based on the aspect ratio of the GUI, providing consistent spatial audio for the local listener regardless of the GUI's position relative to the display screen. For example, by defining the panning range using aspect ratio, the position of the virtual sound source is independent of the orientation, location, and / or size of the GUI displayed on the screen. Therefore, the local user can move the GUI throughout the display screen during a communication session without adversely affecting the spatial cues of the remote participant. Furthermore, spatial cues (e.g., the position of the virtual sound source) are also independent of the display screen's position and / or orientation, which may differ between different users.
[0149] In one aspect, these panning ranges can extend beyond the edges of the display screen. In this case, the panning range can be a function of the size of the analog magnified GUI relative to the size of the display screen (or the area of the display screen showing video data). Therefore, when at a lower aspect ratio, such as in Figure 14A As shown, the controller can utilize the full omnidirectional translation range to spatially render virtual sound sources, while using only a portion of the full elevation translation range (e.g., it may be based on the difference between the height of the magnified GUI and the height of the area that the display can show, as described herein).
[0150] Some aspects can be discussed separately in Figure 5 , Figure 6 , Figure 9 , Figure 13 and Figure 15 The processes 30, 90, 80, 140, and 170 described herein may be modified. For example, at least some of the specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects.
[0151] As described above, one or more speakers that output one or more virtual sound sources can be arranged to output sound to the surrounding environment, such as external speakers that can be integrated into local devices, displays, or any electronic devices, as described herein. On the other hand, the speakers can be part of headphones, such as… Figure 1The headset 6. In this case, the controller can perform similar operations as described herein to determine spatial parameters and use data to spatially render one or more input audio streams. However, in one aspect, the controller can differently define one or more translation ranges. For example, the azimuth translation range and the elevation translation range can extend 360° around the local user. Therefore, the virtual sound source can be positioned anywhere within the sound field. On the other hand, the external speaker of System 1 can have a similar translation range.
[0152] As described above, controller 10 determines various parameters and data for spatial rendering of the input audio streams, such as: one or more azimuth translation ranges and one or more elevation translation ranges (e.g., a set of ranges when the local device is in longitudinal orientation and a set of ranges when the local device is in lateral orientation), one or more translation angles for each input audio stream, distances (e.g., between the local user and the virtual sound source, between the local user and the display screen, etc.), reverberation, device orientation, GUI dimensions (e.g., size, shape, positioning, and aspect ratio), and display screen dimensions (e.g., width, height, and aspect ratio). On the other hand, this data may also include predefined data, such as predefined dimensions of the GUI, predefined distances between the display screen and reference points, etc. In one aspect, the local user can change any of these parameters or values. For example, the local device may display a menu (e.g., based on user selections within the UI of a communication session GUI). Once displayed, the user can adjust any of these parameters or values. For example, the user can adjust the translation range based on a specific situation. Specifically, the user may want to reduce the translation range defined by the display screen, rather than extending it beyond the display screen (e.g., to minimize sound leakage within the acoustic environment).
[0153] As is widely recognized, the use of personally identifiable information should comply with privacy policies and practices that are generally accepted to meet or exceed industry or governmental requirements for protecting user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly explained to users.
[0154] In one aspect, the size of at least one visual representation of a corresponding input video stream associated with an input audio stream rendered separately is larger than the size of the visual representation of an input video stream associated with a mixed input audio stream, wherein all visual representations of corresponding input video streams associated with all input audio streams rendered as a mixture have the same size. In other words, the list of visual representations may all have the same size. In one aspect, the location of individual virtual sound sources is arranged in front of the display screen of a local device based on communication session parameters. In some aspects, the arrangement of the location of individual virtual sound sources includes determining, for each individual virtual sound source, the location within the GUI of the visual representation of the corresponding input video stream associated with the corresponding input audio stream to be spatially rendered as an individual virtual sound source using the set of communication session parameters; and determining one or more spatial parameters indicating the location of the individual virtual sound source within the arrangement based on the determined location of the visual representation, wherein spatially rendering the input audio stream as an individual virtual sound source includes spatially rendering the input audio stream as an individual virtual sound source at that location using the determined spatial data. In one aspect, the spatial parameters include an azimuth angle along a first axis between the position of the individual virtual sound source and a second axis, and an elevation angle along a third axis between the position of the individual virtual sound source and the second axis. In some aspects, the arrangement also includes the position of the individual virtual sound sources, wherein determining these positions further includes, for each input audio stream being mixed, using a set of communication session parameters to determine the position within the GUI of the visual representation of the corresponding input video stream associated with the mixed input audio stream; determining spatial parameters indicating the position of the virtual sound sources of the input audio streams based on the determined positions; and determining new spatial parameters indicating a specific position based on at least some of these spatial parameters, wherein spatial rendering of the mixed input audio streams includes rendering the mixed space of the input audio streams as a single virtual sound source at the specific position using the new spatial parameters. In some aspects, the new spatial parameters are determined by determining a weighted combination of the spatial parameters of all the mixed input audio streams, wherein the specific position differs from the position of the virtual sound sources of the mixed input audio streams. In one aspect, the visual representation of the input video streams associated with the mixed input audio streams is arranged in rows or columns based on the orientation of the local device, wherein the different position is located at the center of the row or column on the display screen.
[0155] According to one aspect of this disclosure, a local device 2 (e.g., its controller 10) may perform a method comprising one or more operations, such as receiving an input audio stream and an input video stream for each of a first plurality of remote devices having a video communication session with the local device; for each input video stream, displaying a visual representation of the input video stream in a graphical user interface (GUI) on a display screen; for at least one input audio stream, spatially rendering the input audio stream to output an individual virtual sound source comprising only the input audio stream via a plurality of speakers; receiving an input audio stream and an input video stream for each of the second plurality of remote devices in response to determining that a second plurality of remote devices has joined the video communication session; determining whether the local device supports additional individual virtual sound sources for one or more input audio streams of the second plurality of remote devices; in response to determining that the local device does not support additional individual virtual sound sources defining a plurality of user interface (UI) areas in the GUI, each UI area including one or more visual representations of one or more input video streams of the first plurality of remote devices, the second plurality of remote devices, or combinations thereof displayed in the UI area; and for each UI area, spatially rendering a mixture of one or more input audio streams associated with the one or more visual representations included in the UI area as a virtual sound source via the plurality of speakers.
[0156] In one aspect, the local device is configured to spatially render a predefined number of input audio streams as individual virtual sound sources, wherein determining whether the local device supports additional individual virtual sound sources includes determining whether the number of input audio streams from a first plurality of remote devices and a second plurality of remote devices determined to be spatially rendered individually is greater than a predefined number. In another aspect, the number of defined UI areas does not exceed a predefined number of input audio streams that can be spatially rendered as individual virtual sound sources. In one aspect, defining a plurality of UI areas includes: displaying all visual representations associated with the first plurality of remote devices and the second plurality of remote devices in a grid manner; and establishing a virtual grid of UI areas on the GUI, wherein each UI area contains one or more visual representations. In some aspects, establishing a virtual grid of UI areas includes assigning one or more adjacent visual representations to each UI area.
[0157] In one aspect, defining the plurality of UI zones includes determining the number of input audio streams received from a first plurality of remote devices and a second plurality of remote devices, wherein the number of the plurality of UI zones is defined based on the number of input audio streams. In one aspect, the input audio streams from the first plurality of remote devices and the second plurality of remote devices are uniformly distributed across the plurality of UI zones. In some aspects, the method further includes: in response to determining that a third plurality of remote devices has joined the video communication session, receiving input audio streams and input video streams for each of the third plurality of remote devices; and redefining the plurality of UI zones by: 1) adding a visual representation of the input audio streams from the third plurality of remote devices to the already defined UI zones, 2) creating one or more new UI zones, or 3) a combination thereof. In one aspect, for each UI zone, the corresponding UI zone virtual sound source is located at a position on the UI zone displayed on the screen. In some aspects, this position is located at the center of the UI zone. In another aspect, this position is located on a visual representation associated with a mixed input audio stream having a signal energy level above a threshold.
[0158] According to one aspect of this disclosure, a local device 2 (its controller 10) may perform a method including one or more operations, such as receiving an input audio stream from a remote device having a communication session with the local device; determining a first orientation of the local device; determining a translation range of the plurality of speakers for the first orientation of the local device along a horizontal axis; using the plurality of speakers to spatially render the input audio stream as a virtual sound source at a position along the horizontal axis and within the translation range; in response to determining that the local device is in a second orientation, determining an adjusted translation range of the plurality of speakers, the adjusted translation range being wider along the horizontal axis than the translation range; and adjusting the position of the virtual sound source along the horizontal axis based on the adjusted translation range.
[0159] In one aspect, the first orientation is a longitudinal orientation of the local device, and the second orientation is a lateral orientation. In another aspect, the position of the virtual sound source is adjusted proportionally along a horizontal axis relative to the adjusted translation range. In one aspect, the translation range is a horizontal translation range, and the adjusted translation range is an adjusted horizontal translation range, wherein the method further includes, when the local device is oriented to the first orientation, determining a vertical translation range of the plurality of speakers that spans along a vertical axis along which the position of the virtual sound source is located; in response to determining that the local device has been oriented to the second orientation, determining an adjusted vertical translation range of the plurality of speakers that spans along the vertical axis less than the vertical translation range; and adjusting the position of the virtual sound source along the vertical axis based on the adjusted vertical translation range, and adjusting the position of the virtual sound source along the horizontal axis based on the adjusted horizontal translation range.
[0160] In one aspect, the method further includes receiving input audio streams and input video streams from remote devices for display as a visual representation on a graphical user interface (GUI) of a local device. In some aspects, in a first orientation, the virtual sound source is positioned along a horizontal axis at the same position relative to the display of the visual representation; and in response to determining that the local device has been oriented to a second orientation, the position of the visual representation relative to the display is maintained while the position of the virtual sound source is adjusted such that the virtual sound source and the visual representation remain in the same position relative to the display. In one aspect, the method further includes receiving individual input audio streams from each of a plurality of remote devices communicating with the local device; and rendering a mixing space of these individual input audio streams into a single virtual sound source, the single virtual sound source comprising a mixture of the individual input audio streams and the mixed input audio streams. In another aspect, the method further includes: receiving a plurality of input video streams, each input video stream originating from a different remote device among the plurality of remote devices; and, when the local device is in a first orientation, displaying a plurality of visual representations in a row along a horizontal axis within a graphical user interface (GUI) on the display of the local device, each visual representation for a different input video stream among the plurality of input video streams, wherein a single virtual sound source is rendered at the location of one of the visual representations. In one aspect, a single virtual sound source is rendered at the location of one of the visual representations in response to an individual input audio stream associated with an input video stream displayed in one of the visual representations having an energy level greater than the remainder of an individual input audio stream in a mixture of individual input audio streams. In some aspects, the location of the rendered single virtual sound source changes along the horizontal axis, but not along the vertical axis based on which individual audio stream associated with a corresponding visual representation in the row has a greater energy level. In another aspect, the method further includes: in response to determining that the local device has been oriented to a second orientation, displaying the plurality of visual representations in a column along a vertical axis within the GUI on the display of the local device; and adjusting the single virtual sound source to continue rendering a signal virtual sound source at the location of one of the visual representations in the column. In some respects, the position of the virtual sound source in the rendering signal changes along the vertical axis, but not along the horizontal axis based on which individual audio stream associated with the corresponding visual representation of that column has a higher energy level than the remainder of those individual audio streams.
[0161] In one aspect, before determining that the local device has been oriented to the second orientation, the method further includes determining spatial data for spatially rendering the input audio stream, the spatial data indicating the position of an individual virtual sound source as an angle between a reference point and the position along a horizontal axis. In another aspect, adjusting the position of the virtual sound source includes determining adjusted spatial data indicating the adjusted position as an adjusted angle between the reference point and the adjusted position along a horizontal axis and within an adjusted translation range; and using the adjusted spatial data to spatially render the input audio stream as a virtual sound source at the adjusted position.
[0162] According to another aspect, the local device 2 (its controller 10) can perform a method including one or more operations, such as receiving an input audio stream and an input video stream from a remote device having a video communication session with the local device; displaying a visual representation of the input video stream within a graphical user interface (GUI) of the video communication session displayed on a screen; determining the aspect ratio of the GUI of the video communication session; determining an azimuth translation range and an elevation translation range based on the aspect ratio of the GUI of the video communication session, the azimuth translation range being at least a portion of the total azimuth translation range of a plurality of speakers, the elevation translation range being at least a portion of the total elevation translation range of the plurality of speakers; and rendering the input audio stream using the plurality of speaker space to output a virtual sound source including the input audio stream within the azimuth and elevation translation ranges.
[0163] On the other hand, the GUI of the video communication session is smaller than the display screen on which it is displayed, wherein the azimuth and elevation translation ranges are independent of the GUI's position within the display screen. In one aspect, the azimuth and elevation translation ranges are independent of the display screen's position and orientation. In some aspects, the display screen is integrated within a local device, wherein the total azimuth translation range spans the width of the display screen, and the total elevation translation range spans the height of the display screen. In another aspect, determining the azimuth and elevation translation ranges includes extending the GUI of the video communication session until at least one of the GUI's width or height has been fully extended to the width of the display screen, while maintaining the aspect ratio; and defining the azimuth translation range as spanning the width of the GUI and the elevation translation range as spanning the height of the GUI. On the other hand, when the width of the GUI is the same as the width of the display screen, the azimuth translation range is the same as the total azimuth translation range, and the elevation translation range is smaller than the total elevation translation range. When the height of the GUI is the same as the height of the display screen, the azimuth translation range is smaller than the total azimuth translation range, and the elevation translation range is the same as the total elevation translation range.
[0164] In one aspect, the method further includes determining spatial parameters indicating the position of a virtual sound source within an azimuth and elevation translation range based on the position of the visual representation within the GUI, wherein the input audio stream uses these spatial parameters for spatial rendering. In some aspects, the spatial parameters include an azimuth angle along the azimuth translation range and an elevation angle relative to a reference point in front of the display along the elevation translation range. In one aspect, the method further includes: determining that the aspect ratio of the GUI has changed; adjusting at least one of the azimuth and elevation translation ranges based on the changed aspect ratio; adjusting the spatial parameters such that the position is within one of the adjusted azimuth and elevation translation ranges; and spatially rendering the input audio stream using the adjusted spatial parameters.
[0165] In some aspects, determining spatial parameters includes: using the x-coordinate of the center point position of the visual representation within the GUI as input to a first linear function of the azimuth translation range to determine the azimuth of the virtual sound source; using the y-coordinate of the center point position of the visual representation within the GUI as input to a second linear function of the elevation translation range to determine the elevation of the virtual sound source; and spatially rendering an input audio stream to output the virtual sound source based on the azimuth and elevation angles. In other aspects, determining spatial parameters includes: estimating the azimuth viewing angle range of the GUI with a predefined width and the elevation viewing angle range of the GUI with a predefined height; determining a reference point located at a predefined distance in front of the display screen showing the GUI; determining the viewing azimuth from the visual representation on the GUI to the reference point and the viewing elevation from the visual representation on the GUI to the reference point; using the viewing azimuth as input to a first linear function of the azimuth translation range relative to the estimated azimuth viewing range to determine the azimuth of the virtual sound source; and using the viewing elevation as input to a second linear function of the elevation translation range relative to the estimated elevation viewing range to determine the elevation of the virtual sound source.
[0166] In one aspect, the position includes an azimuth and an elevation angle relative to a reference point in front of the local device, wherein the adjusted position is determined to have a lower azimuth and a higher elevation angle in response to a smaller aspect ratio. In another aspect, when the GUI has a reduced aspect ratio, the lower azimuth is an angle that does not fully extend the width of the display screen within the azimuth translation range of the plurality of speakers, and the higher elevation is an angle that extends the height of the display screen within the elevation translation range of the plurality of speakers.
[0167] As previously described, one aspect of this disclosure may be a non-transitory machine-readable medium (such as microelectronic memory) on which instructions are stored, programming one or more data processing units (generally referred to herein as "processors") to perform network operations and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components containing hard-wired logic. Alternatively, those operations may be performed by any combination of programmed data processing units and fixed hard-wired circuit components. In one aspect, when the one or more processors execute instructions stored in the non-transitory machine-readable medium, the operations described herein may be performed by a local device.
[0168] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that such aspects are merely illustrative of the broad disclosure and not limiting, and that this disclosure is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.
[0169] In some aspects, this disclosure may include the language "[element A] and [element B] at least one". This language may refer to one or more of these elements. For example, "at least one of A and B" may refer to "A", "B", or "A and B". Specifically, "at least one of A and B" may refer to "at least one of A and at least one of B" or "at least either A or B". In some aspects, this disclosure may include the language "[element A], [element B], and / or [element C]". This language may refer to any of these elements or any combination thereof. For example, "A, B, and / or C" may refer to "A", "B", "C", "A and B", "A and C", "B and C", or "A, B, and C".
Claims
1. A method executed by a programmable processor of a local device communicatively coupled to a plurality of remote devices, the method comprising: Receives an input audio stream from each of the plurality of remote devices that are in a communication session with the local device; Each remote device receives a set of communication session parameters; For each input audio stream, the set of communication session parameters determines whether the input audio stream is to be rendered separately relative to other received input audio streams, or to be rendered as an input audio stream mixed with one or more other input audio streams. For each input audio stream determined to be rendered separately, the input audio stream space is rendered as an individual virtual sound source containing only that input audio stream; as well as For an input audio stream that is determined to be rendered as a mixture of the aforementioned input audio streams, the input audio stream mixing space is rendered as a single virtual sound source containing the mixture of the input audio streams.
2. The method of claim 1, wherein spatial rendering of each of the input audio stream and the input audio stream mixture comprises generating a single set of driver signals for driving a plurality of speakers of the local device.
3. The method of claim 1, wherein the local device is configured to output a predefined number of virtual sound sources, wherein the number of input audio streams determined to be rendered individually is less than the predefined number of virtual sound sources.
4. The method of claim 1, further comprising receiving a Voice Activity Detection (VAD) signal from each of the plurality of remote devices, wherein each set of communication session parameters includes at least one VAD parameter based on the VAD signal, the VAD signal indicating at least one of voice activity and voice intensity of a remote participant at the respective remote device, wherein determining for each input audio stream whether the input audio stream is to be rendered individually or as a mixture of input audio streams includes: When the VAD parameter is higher than the threshold, it is determined that the input audio stream should be rendered separately; and When the VAD parameter is lower than the threshold, it is determined that the input audio stream should be rendered as the mix.
5. The method of claim 1, further comprising receiving the input audio stream and the input video stream from each of the plurality of remote devices, the input video stream being used to display as a visual representation in a graphical user interface (GUI) on a display of the local device.
6. The method of claim 5, wherein each set of communication session parameters indicates the position of the visual representation of the corresponding input video stream in the GUI, wherein determining for each input audio stream whether the input audio stream is to be rendered individually or as a mixture of input audio streams comprises: When the visual representation of the corresponding input video stream associated with the input audio stream is within the canvas area of the GUI, it is determined that the input audio stream should be rendered separately. and When the visual representation of the corresponding input video stream associated with the input audio stream is within a list area of the GUI that is separate from the canvas area, it is determined that the input audio stream is to be rendered as the mix.
7. The method of claim 5, wherein each set of communication session parameters indicates the size of the visual representation of the corresponding input video stream in the GUI, wherein determining whether the input audio stream is to be rendered separately or as a mix of input audio streams comprises: When the size of the visual representation of the corresponding input video stream associated with the input audio stream is greater than a threshold size, it is determined that the input audio stream should be rendered separately. and When the size of the visual representation of the corresponding input video stream associated with the input audio stream is less than the threshold size, it is determined that the input audio stream should be rendered as the mix.
8. The method of claim 5, wherein the visual representation of the input video stream is in a first arrangement in the GUI, wherein the method further comprises a second arrangement for determining, based on the communication session parameters, the positions of the individual virtual sound sources and the single virtual source, respectively, in front of, behind, or to the side of the display on the local device.
9. The method of claim 8, wherein the second arrangement is the same as the first arrangement, such that each of the individual virtual sound sources is located at a corresponding visual representation.
10. The method of claim 8, wherein the second arrangement is the same as the first arrangement and is proportionally larger than the first arrangement.
11. A local electronic device, the local electronic device comprising: At least one processor; and The memory has instructions that, when executed by the at least one processor, cause the local electronic device to: Receive input audio streams from each of a plurality of remote devices that are in communication sessions with the local electronic device; Each remote device receives a set of communication session parameters; For each input audio stream, the set of communication session parameters determines whether the input audio stream is to be rendered separately relative to other received input audio streams, or to be rendered as an input audio stream mixed with one or more other input audio streams. For each input audio stream determined to be rendered separately, the input audio stream space is rendered as an individual virtual sound source containing only that input audio stream; as well as For an input audio stream determined to be rendered as a mix of the input audio streams, the input audio stream mix space is rendered as a single virtual sound source containing the mix of the input audio streams.
12. The local electronic device of claim 11, further comprising a display, wherein the memory has further instructions for receiving the input audio stream and the input video stream from each of the plurality of remote devices, the input video stream being displayed as a visual representation in a graphical user interface (GUI) on the display.
13. The local electronic device of claim 12, wherein each set of communication session parameters indicates the location of the visual representation of the corresponding input video stream in the GUI, wherein the instructions for determining, for each input audio stream, whether the input audio stream is to be rendered individually or as a mixture of input audio streams include: When the visual representation of the corresponding input video stream associated with the input audio stream is within the canvas area of the GUI, it is determined that the input audio stream should be rendered separately. and When the visual representation of the corresponding input video stream associated with the input audio stream is within a list area of the GUI that is separate from the canvas area, it is determined that the input audio stream is to be rendered as the mix.
14. The local electronic device of claim 12, wherein each set of communication session parameters indicates the size of the visual representation of the corresponding input video stream in the GUI, wherein determining whether the input audio stream is to be rendered separately or as a mixture of input audio streams comprises: When the size of the visual representation of the corresponding input video stream associated with the input audio stream is greater than a threshold size, it is determined that the input audio stream should be rendered separately. and When the size of the visual representation of the corresponding input video stream associated with the input audio stream is less than the threshold size, it is determined that the input audio stream should be rendered as the mix.
15. The local electronic device of claim 12, wherein the visual representation of the input video stream is in a first arrangement in the GUI, wherein the memory has further instructions for determining, based on the communication session parameters, a second arrangement for determining the positions of the individual virtual sound sources and the single virtual source, respectively, in front of, behind, or to the side of the display of the local electronic device.
16. An application program in a non-transitory machine-readable medium, the application program being configured to execute on a local device, the application program operating as follows: Receive input audio streams from each of a plurality of remote devices that are in a communication session with the local device; Each remote device receives a set of communication session parameters; For each input audio stream, the set of communication session parameters determines whether the input audio stream is to be rendered separately relative to other received input audio streams, or to be rendered as an input audio stream mixed with one or more other input audio streams. For each input audio stream determined to be rendered separately, the input audio stream space is rendered as an individual virtual sound source containing only that input audio stream; as well as For an input audio stream determined to be rendered as a mix of the input audio streams, the input audio stream mix space is rendered as a single virtual sound source containing the mix of the input audio streams.
17. The application of claim 16, wherein spatial rendering of each of the input audio stream and the input audio stream mix comprises generating a single set of driver signals for driving a plurality of speakers of the local device.
18. The application of claim 16, further comprising receiving a Voice Activity Detection (VAD) signal from each of the plurality of remote devices, wherein each set of communication session parameters includes at least one VAD parameter based on the VAD signal, the VAD signal indicating at least one of voice activity and voice intensity of a remote participant of the respective remote device, wherein determining for each input audio stream whether the input audio stream is to be rendered individually or as a mixture of input audio streams includes: When the VAD parameter is higher than the threshold, it is determined that the input audio stream should be rendered separately; and When the VAD parameter is lower than the threshold, it is determined that the input audio stream should be rendered as the mix.
19. The application of claim 16, further comprising receiving the input audio stream and the input video stream from each of the plurality of remote devices, the input video stream being used to display as a visual representation in a graphical user interface (GUI) on a display of the local device.
20. The application of claim 19, wherein each set of communication session parameters indicates the location of the visual representation of the corresponding input video stream in the GUI, wherein determining for each input audio stream whether the input audio stream is to be rendered individually or as a mixture of input audio streams comprises: When the visual representation of the corresponding input video stream associated with the input audio stream is within the canvas area of the GUI, it is determined that the input audio stream should be rendered separately. and When the visual representation of the corresponding input video stream associated with the input audio stream is within a list area of the GUI that is separate from the canvas area, it is determined that the input audio stream is to be rendered as the mix.
21. A method executed by a programmable processor of a local device, the method comprising: Using multiple speakers, at a location within the environment where the local device is located, the input audio stream from a remote device having a video communication session with the local device is output as a virtual sound source; A graphical window of the video communication session is displayed on the monitor, the graphical window having a visual representation of the remote participants of the remote device; It has been determined that the size of the graphics window has changed; as well as The position of the virtual sound source is adjusted based on the changed size.
22. The method of claim 21, further comprising determining the aspect ratio of the graphics window based on the changed size, wherein the position is adjusted according to the aspect ratio.
23. The method of claim 22, wherein the position includes an azimuth and an elevation angle relative to a reference point in front of the local device, wherein in response to determining that the aspect ratio has increased, the adjusted position has a higher azimuth and a lower elevation.
24. The method according to claim 23, wherein, When the graphics window has an increased aspect ratio, the higher azimuth angle is within the range of azimuth angle translations of the plurality of speakers that extend the width of the display, and the lower elevation angle is within the range of elevation angle translations of the plurality of speakers that do not fully extend the height of the display.
25. The method of claim 21, wherein the adjusted position of the virtual sound source is independent of the position and orientation of the display.
26. The method of claim 21, wherein determining that the size of the graphics window has changed includes receiving user input that changes the width or height of the graphics window.
27. The method of claim 21, wherein the adjusted position of the virtual sound source is independent of the position of the graphics window within the display.
28. A local device, the local device comprising: monitor; At least one processor; and A memory, wherein the memory stores instructions that, when executed by the at least one processor, cause the local device to: Using multiple speakers, the input audio stream from a remote device having a video communication session with the local device is output as a virtual sound source at a location within the environment where the local device is located; A graphical window of the video communication session is displayed on the monitor, the graphical window having a visual representation of the remote participants of the remote device; It has been determined that the size of the graphics window has changed; as well as The position of the virtual sound source is adjusted based on the changed size.
29. The local device of claim 28, wherein the memory includes further instructions for determining the aspect ratio of the graphics window based on the changed size, wherein the position is adjusted according to the aspect ratio.
30. The local device of claim 29, wherein the position includes an azimuth and an elevation angle relative to a reference point in front of the local device, wherein in response to determining that the aspect ratio has increased, the adjusted position has a higher azimuth and a lower elevation.
31. The local device according to claim 30, wherein, When the graphics window has an increased aspect ratio, the higher azimuth angle is within the range of azimuth angle translations of the plurality of speakers that extend the width of the display, and the lower elevation angle is within the range of elevation angle translations of the plurality of speakers that do not fully extend the height of the display.
32. The local device of claim 28, wherein the adjusted position of the virtual sound source is independent of the position and orientation of the display.
33. The local device of claim 28, wherein the instruction for determining that the size of the graphics window has changed includes an instruction for receiving user input that changes the width or height of the graphics window.
34. The local device of claim 28, wherein the adjusted position of the virtual sound source is independent of the position of the graphics window within the display.
35. The local device of claim 28, wherein the local device includes a desktop or laptop computer in which the plurality of speakers are integrated.
36. An application program in a non-transitory machine-readable medium, the application program being configured to execute on a local device, the application program operating as follows: Using multiple speakers, at a location within the environment where the local device is located, the input audio stream from a remote device having a video communication session with the local device is output as a virtual sound source; A graphical window of the video communication session is displayed on the monitor, the graphical window having a visual representation of the remote participants of the remote device; It has been determined that the size of the graphics window has changed; as well as The position of the virtual sound source is adjusted based on the changed size.
37. The application of claim 36, wherein the application is further configured to determine the aspect ratio of the graphics window based on the changed size, wherein the position is adjusted according to the aspect ratio.
38. The application of claim 37, wherein the position includes an azimuth and an elevation angle relative to a reference point in front of the local device, wherein in response to determining that the aspect ratio has increased, the adjusted position has a higher azimuth and a lower elevation.
39. The application of claim 36, wherein the application determines that the size of the graphics window has changed by receiving user input that changes the width or height of the graphics window.
40. The application of claim 36, wherein the adjusted position of the virtual sound source is independent of at least one of: 1) the position and orientation of the display, or 2) the position of the graphics window within the display.