A multi-participant, spatial audio service
A dual rendering pipeline system with separate latency settings for local and remote participants addresses high processing demands in spatial audio, ensuring efficient and immersive multi-participant audio services by managing spatial audio separation and rendering.
Patent Information
- Application Number
- GB2024003237
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-10
AI Technical Summary
Existing spatial audio processing systems face high processing demands, making it challenging to efficiently render multi-participant audio services with low latency and high fidelity, especially when participants are in different physical locations.
Implementing a dual rendering pipeline system with a lower latency pipeline for local participants and a higher latency pipeline for remote participants, utilizing a point-to-point wireless link for local communication and a network link for remote communication to manage spatial audio rendering in a virtual space.
The system effectively separates and renders spatial audio sources for local and remote participants with controlled locations, providing low latency for local interactions and manageable latency for remote interactions, enhancing the immersive experience in multi-participant audio services.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Proposed multi-participant audio services will use spatial audio to spatially separate a first sound source associated with a first participant from a second sound source associated with a second participant when the first and second sound sources are rendered to a third participant within a virtual space. Spatial audio provides, via digital processing, for the controllable location (direction and / or distance) of different sound source so that each has a different, independently controllable location (direction and / or distance) in the virtual space. Because of the very high processing demands, the spatial audio processing is expected to be performed as a remote service. BRIEF SUMMARY According to various, but not necessarily all, examples there is provided an apparatus comprising: means for audio rendering comprising: at least a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and at least a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. In some but not necessarily all examples, the means for audio rendering is configured to provide first-person-perspective-mediated-reality wherein a listening participant’s real position determines, overtime, a position a virtual listening participant within a virtual space for the first-person perspective tracked binaural audio. In some but not necessarily all examples, the lower latency rendering pipeline and the higher latency rendering pipeline are configured to operate simultaneously. In some but not necessarily all examples, the apparatus comprises means for using a point-to-point, wireless communication link between the lower latency pipeline and an apparatus of the local, visible, or potentially visible participant. In some but not necessarily all examples, the point-to-point, wireless communication link is a point-to-point, wireless, bi-directional communication link between the lower latency pipeline and the apparatus of the local, visible, or potentially visible participant. In some but not necessarily all examples, the apparatus comprises: means for detecting a local apparatus of the local, visible, or potentially visible participant; and means for coupling the lower latency rendering pipeline to the local apparatus via a direct, point-to-point, wireless communication link, to receive at least position information for a remote apparatus and audio captured by the remote apparatus. In some but not necessarily all examples, the apparatus comprises means for coupling the higher latency rendering pipeline to a remote service via a network link, to receive at least partially rendered spatial audio content. In some but not necessarily all examples, the apparatus comprises means for coupling the higher latency rendering pipeline to a remote service via a network link, to receive split-rendered spatial audio content. In some but not necessarily all examples, the lower latency rendering pipeline receives, for the local, visible, or potentially visible, participant, a position of the local participant and audio associated with the participant; and the higher latency rendering pipeline receives, for the remote participant at least partially rendered audio representing a virtual space comprising multiple positioned sound sources. In some but not necessarily all examples, the apparatus is configured to adjust a position of the virtual space is response to adjustment of a position of the apparatus to provide first-person-perspective tracked binaural audio. In some but not necessarily all examples, the apparatus comprises means for positioning the apparatus. In some but not necessarily all examples, the lower latency rendering pipeline is configured for single sound source spatial rendering. In some but not necessarily all examples, the higher latency rendering pipeline is configured for multiple sound source spatial rendering. In some but not necessarily all examples, the apparatus is configured to receive uncompressed audio for the lower latency rendering pipeline and the apparatus is configured to receive encoded audio for the higher latency rendering pipeline. In some but not necessarily all examples, the apparatus comprises means for coupling the apparatus to a lower latency rendering pipeline of a local apparatus of the local, visible, or potentially visible participant to provide at least position information for the apparatus and audio captured by the apparatus to the local apparatus. In some but not necessarily all examples, the apparatus comprises means for coupling the apparatus, via a network link, to a remote service to provide to the remote service at least position information for the apparatus and audio captured by the apparatus. In some but not necessarily all examples, the lower latency rendering pipeline is for rendering audio with a latency below 20ms and the higher latency rendering pipeline is for rendering audio with a latency above 20ms. In some but not necessarily all examples, the apparatus is configured as a head mounted apparatus comprising means for tracking a position of a head of a user. According to various, but not necessarily all, examples there is provided a computer program comprising instructions that when executed by at least one processor of an apparatus cause the apparatus to: provide a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and provide a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. According to various, but not necessarily all, examples there is provided a method comprising: rendering audio content received directly from an apparatus of a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and rendering a virtual space comprising a positioned audio source originating from a remote participant as first-person perspective tracked binaural audio. According to various, but not necessarily all, examples there is provided an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: provide a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and provide a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. According to various, but not necessarily all, examples there is provided a system comprising: means for rendering first-person-perspective-mediated-reality* comprising: at least a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and at least a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. According to various, but not necessarily all, examples there is provided examples as claimed in the appended claims. While the above examples of the disclosure and optional features are described separately, it is to be understood that their provision in all possible combinations and permutations is contained within the disclosure. It is to be understood that various examples of the disclosure can comprise any or all the features described in respect of other examples of the disclosure, and vice versa. Also, it is to be appreciated that any one or more or all the features, in any combination, may be implemented by / comprised in / performable by an apparatus, a method, and / or computer program instructions as desired, and as appropriate. The description of a function should additionally be considered to also disclose any means suitable for performing that function BRIEF DESCRIPTION Some examples will now be described with reference to the accompanying drawings in which: FIG. 1 illustrates a system suitable for the rendering of the audio service in a virtual space (with or without participant virtual visual objects); FIG. 2 illustrates an example of a rendering part 100 of the system illustrated in FIG. 1; FIG. 3 illustrates an example of an apparatus 10 comprised in the rendering part 100; FIG. 4 illustrates the system previously discussed in relation to FIGs 1 and 2 using the apparatus 10 as described in FIG 3; FIG. 5A illustrates the information flow (downlink flow) when an apparatus 10 is being used by a listening participant; FIG. 5B illustrates the information flow (uplink flow) when an apparatus 10 is being used by a contributing participant; FIG. 6 illustrates an example based on FIG 4 with an additional remote apparatus used by a participant. FIG 7 illustrates, compared to FIG 6, change in position (orientation) of some participants; FIG 8 illustrates, compared to FIG 6, change in position (orientation and location) of some participants; FIG. 9A and FIG 9B illustrate an effect of latency when there is relative movement between participants; FIG. 10 illustrates an example of method suitable for set-up and configuration, or reconfiguration; FIG. 11 illustrates an example of a controller for the apparatus 10; FIG. 12 illustrates an example of executable code for the apparatus 10. The figures are not necessarily to scale. Certain features and views of the figures can be shown schematically or exaggerated in scale in the interest of clarity and conciseness. For example, the dimensions of some elements in the figures can be exaggerated relative to other elements to aid explication. Similar reference numerals are used in the figures to designate similar features. For clarity, all reference numerals are not necessarily displayed in all figures. In the following description a class (or set) can be referenced using a reference number without a subscript index (e.g. 10) and a specific instance of the class (member of the set) can be referenced using the reference number with a numerical type subscript index (e.g. 10_1) and a non-specific instance of the class (member of the set) can be referenced using the reference number with a variable type subscript index (e.g. 10J). DETAILED DESCRIPTION In examples in the following description, an apparatus 10 comprises: means for rendering first-person-perspective-mediated-reality 120 comprising: at least a lower latency rendering pipeline 20 for rendering audio content originating from a local, visible, or potentially visible participant 50 as first-person perspective tracked binaural audio; at least a higher latency rendering pipeline 30 for rendering audio content originating from a remote participant 50 as first-person perspective tracked binaural audio. In at least some examples, the first-person-perspective-mediated-reality is used to provide a multi-participant audio service that uses spatial audio to spatially separate sound sources within a virtual space when rendered to a third participant using the apparatus 10. In at least some examples, the multi-participant audio service uses spatial audio to spatially separate a first sound source associated with a first participant from a second sound source associated with a second participant when the first and second sound sources are rendered to the third participant. In some examples, the spatial separation is via direction only. In other examples, the spatial separation is via direction and distance. Spatial audio provides, via digital processing, for the controllable location (direction and / or distance) of different sound source so that each has a different, independently controllable location (direction and / or distance) in a virtual space. The user of the apparatus 10 is a participant to the multi-participant audio service where the multiple participants of the audio service share a virtual space. Some participants may be local to each other (local participants) and occupy a common, shared physical space. Some participants (remote participants) may be remote to each other and occupy distinct, separate physical spaces. The distinct, separate physical spaces can be distant from each other. In at least some examples, a participant is representable in the virtual space to the other participants using a collection of one or more sound sources, for example sound objects. Different participants are representable in the virtual space to the other participants using different collections of one or more sound sources, that are spatially separated in the virtual space. The sound source(s) representing a participant can for example provide captured speech of the participant or other audio associated with the participant (e.g. ambient sound or selected audio content). The sound source(s) representing speech of a participant will be referred to as a speech sound source. In at least some examples, a speech sound source associated with a participant (a contributing participant) can be reoriented in real-time in the virtual space in response to a reorientation of that participant in that participant’s real space. In at least some examples, a speech sound source associated with a contributing participant can be relocated in real-time in the virtual space in response to a re-location of that participant in that participant’s real space. A speech sound source associated with a contributing participant can therefore be re-positioned (orientation and / or location) in real-time in response to re-positioning (orientation and / or location) of that participant in real space. This can occur in real-time (contemporaneously) for multiple participants. Any of multiple participants can therefore experience, as listening participants, a virtual space in which a speech sound source associated with another contributing participant is re-positioned (orientation and / or location) in real-time in response to re-positioning (orientation and / or location) of the other contributing participant in the contributing participant’s real space. This can occur contemporaneously for multiple of the other contributing participants. The virtual space is therefore a dynamic sound space in which sound sources, for different participants, start (appear) and stop (disappear) and in which, for multiple participants, a sound source associated with a participant (e.g. a speech sound source) can be re-positioned (orientation and / or location), for example, in response a repositioning (orientation and / or location) of that participant in real space. The virtual space provides spatial audio as there are one or more sound sources that have distinct controllable positions within the virtual space. Spatial audio refers to 3D audio, i.e., it can provide a perception where sound sources are heard at least from different directions. Spatial audio can be reproduced, e.g., using a loudspeaker setup or via headphones, for example with head-tracking capability. If a listening participant changes position (orientation and / or location) within their real space, this has a consequential change in their virtual position (orientation and / or location) in the virtual space. A position (orientation and / or location) of a virtual listener in the virtual space determines what is rendered to the listening participant and the position (orientation and / or location) of the virtual listeners changes in real-time with the position (orientation and / or location) of the listening participant in real space. The position (orientation and / or location) of the virtual listener in the virtual space tracks the position (orientation and / or location) of the listening participant in real space. The change in position can be a change in orientation, for example in three (angular) degrees-of-freedom (3DoF). The three angular degrees-of-freedom can for example be spherical polar co-ordinates. The change in position can be a change in orientation, for example in three (angular) degrees-of-freedom, with small translational movements, for example, by leaning (3DoF+). The change in position can be a change in orientation and location, for example in six degrees-of-freedom comprised of 3 angular and 3 translational degrees-of-freedom (6DoF). The three translational degrees-of-freedom can for example be Cartesian co-ordinates. A change in a participant’s virtual position (orientation and / or location) in the virtual space has, as described above, a consequential change in a position (orientation and / or location) their associated sound source (e.g. their speech sound source) is rendered in the virtual space to other participants. A change in a listening participant’s virtual position (orientation and / or location) in the virtual space also has a consequential change in a position (orientation and / or location) of the virtual listener and therefore a change in position (orientation and / or location) at which sound source(s) (e.g. speech sound source(s)) associated with other participants are rendered in the virtual space to the listening participant who has changed position. The position of a participant can be expressed as an orientation, a location or a combination of a position and a location in their own real space. The virtual position of a virtual participant can be expressed as an orientation, a location or a combination of a location and an orientation in the virtual space. An orientation of a participant (or virtual participant) can be referred to a point-of-view (POV) or as a pose. The term “first person perspective-mediated” as applied to mediated reality, augmented reality or virtual reality means mediated with the constraint that a listening participant’s real position (location and / or orientation) determines the position (location and / or orientation) of the virtual listening participant within the virtual space. In at least some examples, the audio service in the virtual space is augmented with visual objects. For example, in virtual reality, another participant can be represented visually to a viewing participant as a virtual visual object (a participant virtual visual object), for example a visual avatar. The position (orientation and / or location) of the participant virtual visual object in the visual space is the virtual position (orientation and / or location) that tracks the participants real position (orientation and / or location) within their real space. Thus, the virtual visual object of a participant and the sound source associated with that participant (e.g. the speech sound source) are copositioned (orientation and / or location) and move together tracking changes in position (orientation and / or location) of that participant. If participants do (or do not) share a common physical space, then virtual reality is used for rendering the associated sound sources of the participants and their participant virtual visual objects. Virtual reality is used for rendering the associated sound sources of local and remote participants. If participants share a common physical space, then augmented reality can be used. The associated sound sources of the local participants are rendered as in virtual reality and as described above, however, instead of rendering participant virtual visual objects for the other local participants, a see-through display can be used so that the listening / viewing participant (the user of the apparatus 10) can see the other local participants. The sound sources of the local participants track the positions of the local participants. It should be appreciated that any one or more of the participants can each use an apparatus 10. In the following description, however, a single apparatus 10 will be described for conciseness. In addition, the description will focus on user of the apparatus 10 being a listening participant. However, the user of the apparatus 10 can also be a contributing participant. In at least some examples, the user of the apparatus 10 is contemporaneously a listening participant and a contributing participant. FIG. 1 illustrates a system suitable for the rendering of the audio service in the virtual space (with or without participant virtual visual objects). The system comprises a rendering part 100 for rendering the virtual space to a listening participant. The rendering part 100 comprises a position estimator 110 that tracks a position (orientation and / or location) of the listening participant in real space so that listening participant can, as a virtual participant, change their position (orientation and / or location) in the virtual space as described above. The rendering part 100 comprises a renderer 120 for rendering the dynamically changing virtual space to the listening participant, as the corresponding virtual participant changes their position (orientation and / or location) in the virtual space. The virtual space dynamically changes as a consequence of other participants actions. For example, sound sources appear and disappear, and sound sources can move with movement of contributing participants. The position (orientation and / or location) of the virtual participant in the virtual space dynamically changes as a consequence of changes in position (orientation and / or location) of the listening participant. A network part 200 interconnects the rendering part 100 of the system to a remote service part 202 of the system. The service part 202 enable the rendering of the multi-participant virtual space (with or without participant virtual visual objects) to the user of the rendering part 100. This can, for example, be a cloud service. In some examples, the remote service part 202 performs some or all of rendering to account for virtual space dynamic changes as a consequence of other participants actions and changes in rendering the virtual space (first-person perspective mediated reality) as a consequence of changes in position (orientation and / or location) of the listening participant. In some examples, the rendering part 100 performs some of the rendering to account for virtual space dynamic changes as a consequence of other participants actions and changes in rendering the virtual space (first-person perspective mediated reality) as a consequence of changes in position (orientation and / or location) of the listening participant. This can be described as split rendering. Split rendering is a way to provide audio for the listener using an apparatus 10 that has constrained resources. For example, 3DoF, 3DoF+, 6DoF audio processing can be done at the remote service part 202. The remote service part 202 generates pre-rendered audio. For example, the remote service part 202 generates an audio virtual space for the latest participant positions known to it, and this pre-rendered audio is transmitted downstream. There is latency when audio is processed and transmitted. Also, transmission of participant positions and participant audio to the remote service part 202 takes time adding to latency. During the latency, a listening participant may for example change position (orientation and / or location). This change in position can be compensated for downstream at the rendering part 100 or the network part 200 based on the listening participant position. Thus, the rendering part 100 and / or the network part 200 performs post rendering. The rendering to produce rendered audio is performed as split-rendering comprising a pre-rendering stage that produces split-rendered audio and a post-rendering stage (a correction stage) that produces rendered audio. There exist several methods how this post-correction can be done. In some examples of split-rendering, the remote service part 202 performs all of rendering to account for virtual space dynamic changes as a consequence of other participants actions and the rendering part 100 performs all of rendering to account for changes in rendering the virtual space (first-person perspective mediated reality) as a consequence of changes in position (orientation and / or location) of the listening participant (user of the rendering part 100). Any one of the parts (rendering part 100, network 200 or remote service part 202) can perform some aspect of the required rendering. The information transferred between parts is encoded for transmission and decoded on reception, if it is to be rendered or processed, otherwise it can pass-through. The virtual space can be encoded in immersive audio data which is streamed directly to the rendering part 100 (or network 200), which is responsible for decoding, rendering, and synchronizing the audio with the corresponding visual content. The rendering part 100 (or network 200) can process the user position information locally and adjust the audio rendering accordingly to create a convincing immersive experience that provides user head-tracked immersive audio. Immersive audio decoding, head-tracked binaural rendering based on pose information, playback, and pose estimation can be performed on one device (standalone) or across many devices (non-standalone) including different devices. Head-tracked binaural rendering can be split between pre-rendering at a presentation engine and post-rendering at an end device which may be power constrained or limited in computational power. Pre-rendering is, for example, the process of decoding / rendering an original coded immersive audio format to an intermediate immersive audio representation suitable to be transmitted to an end device. Postrendering is, for example, the process of decoding / rendering an intermediate immersive audio representation in an end device. The pre-rendering and the postrendering can be performed on the same device or on different devices. The audio decoding and head-tracked binaural rendering can be performed by one device (local decoding and rendering) and the pose estimation and playback can be performed by the one different device or by multiple different devices. The head-tracked binaural rendering can be split. Post-rendering can be performed where pose estimation occurs or where playback occurs or separately. Pre-rendering can be performed where decoding occurs. Split rendering uses pre-rendering, atone part, to an intermediate representation, followed by coding and transmitting that representation for decoding and rendering on another part. Head-tracked immersive audio attributes and especially DoF (Degrees of Freedom) attributes of the audio formats of the original coding format are preferably retained during transcoding. For example, if the immersive audio of the original coding format is head-trackable in 3DoF, split rendering preferably retains this property. The input at the remote service part 202 can, for example, be audio and position information from the participants. The remote service part 202 can create a virtual audio space as immersive audio and then encode it for transmission. A downstream part (the rendering part 100 or network 200), decodes the received data, and performs additional rendering. The result is that binaural audio is rendered to the listening participant representing the virtual space with a head-tracking response for the listening participant. The listening participant can change position (orientation and / or location) within the virtual space. FIG. 2 illustrates an example of the rendering part 100. In this example, the rendering part 100 comprises an apparatus 10 comprising means 124 for rendering audio to a user. In this example, but not necessarily all examples, the rendering to the listening participant is provided via isolated ear speakers (headphones) in a binaural format. The isolated ear speakers (headphones) can be provided as in-ear, on-ear, or over-ear speakers. For example, a left-ear bud and a right-ear bud can be used. Alternatively, an over-head headset can be used. The apparatus 10 can therefore be configured as a head-mounted apparatus worn by a user. In this example, but not necessarily all examples, the apparatus 10 comprises a position estimator 110 for estimating a position (or information dependent upon a position) of the listening participant (the user of the apparatus 10). The position is defined by orientation and / or location. The position can be defined using three orientation angles (3DoF). The position can be defined using three orientation angles and three translation axes (3DoF+, 6DoF). The three orientation angles span a three-dimension space and can be orthogonal (e.g. spherical co-ordinate system). The three translation axes span the three-dimension space and can be orthogonal (e.g. Cartesian coordinate system). In the example illustrated but not necessarily all examples, the rendering part 100 comprises means 122 for rendering visual objects at different positions in the virtual space that correspond to different participants. The rendering of the visual objects can be performed via a head-mounted display. This can be an opaque display for virtual reality and a see-through display for augmented reality. In some examples, the display can be switched between opaque (for VR) and see-through (for AR). In the example illustrated but not necessarily all examples, the rendering part 100 comprises a sub-part, apparatus 10, for rendering the spatial audio content and a different sub-part 12 rendering the visual content. In this example, the rendering part 100 comprises an apparatus 10 comprising means 124 for rendering audio to a user. The rendering is generally provided via isolated ear speakers in a binaural format. The isolated ear speakers can be provided as in-ear, on-ear, or over-ear speakers. For example, a left-ear bud and a right-ear bud can be used. Alternatively, a headset can be used. The apparatus 10 can therefore be configured as a head-mounted apparatus. In this example, but not necessarily all examples, the apparatus 10 comprises a position estimator 110 for estimating a position (or information dependent upon a position) of the user. The position is defined by orientation and / or location. The position can be defined using three orientation angles (3DoF). The position can be defined using three orientation angles and three translation axes (3DoF+, 6DoF). The three orientation angles span a three-dimension space and can be orthogonal (e.g. spherical co-ordinate system). The three translation axes span the three-dimension space and can be orthogonal (e.g. Cartesian coordinate system). FIG. 3 illustrates an example of an apparatus 10 comprising: means 124 for rendering first-person-perspective-mediated-reality comprising: at least a lower latency rendering pipeline 20 for rendering audio content originating from a local, visible, or potentially visible participant 50 as first-person perspective tracked binaural audio; and at least a higher latency rendering pipeline 30 for rendering audio content originating from a remote participant 50 as first-person perspective tracked binaural audio. There are two separate rendering pipelines: a lower latency pipeline and a higher latency pipeline. Some or all of the pipelines can be local, for example, at the apparatus 10. Some of the pipelines can be remote, for example, at a remote service part 202. It is more probable that a pipeline from a remote service part 202 will have greater latency than a wholly local pipeline. First-person perspective tracked binaural audio is characterized by movement of the virtual space (including its sound sources and virtual visual objects if any) relative to the user (listening / viewing participant) in a direction opposite to physical movement of the user in a physical space. This gives the impression that the virtual space is fixed with respect to the physical space. Thus, if a user rotates their head to the left, the virtual space is rotated to the right equally, such that it appears fixed to the physical space as diegetic sound. The lower latency rendering pipeline 20 and the higher latency rendering pipeline 30 are configured to operate simultaneously. This does not necessarily mean that they always operate simultaneously. For example, the lower latency rendering pipeline 20 is operable when there is audio content originating from the local, visible, or potentially visible participant. For example, the higher latency rendering pipeline 30 is operable when there is at least one remote participant. In some but not necessarily all examples, the higher latency rendering pipeline 30 is operable when there is audio content originating from (or in association with) at least one remote participant. The lower latency rendering pipeline 20 has an associated latency (a delay) that is less than the associated latency of the higher latency rendering pipeline 30. The higher latency of the higher latency rendering pipeline 30 arises because it uses the remote service 202. The output of the higher latency rendering pipeline 30 is delayed relative to the output from the lower latency rendering pipeline 20. In some but not necessarily all examples, the lower latency rendering pipeline 20 has a latency below 20ms. In some but not necessarily all examples, the higher latency rendering pipeline 30 has a latency above 20ms. In some but not necessarily all examples, the lower latency rendering pipeline 20 has a latency below 20ms and the higher latency rendering pipeline 30 has a latency above 20ms. In some but not necessarily all examples, the lower latency rendering pipeline 20 has a latency below 45-100ms to maintain lip synchronization, or below 12-15ms to render multi-participant speech well or below 5ms to render multi-participant music well. FIG. 4 illustrates the system previously discussed in relation to FIGs 1 and 2 using the apparatus 10 as described in FIG 3. For clarity of description, it will be assumed that the user of a first apparatus 10_1 is a listening participant, that the user of a second apparatus 10_2 is a contributing participant and that the user of a third apparatus 10_3 is a contributing participant. However, the user of any of the apparatus 10 can be both a listening participant and / or a contributing participant, and this can change over time. The user of the first apparatus 10_1 and the user of the second apparatus 10_2 are local participants. They occupy a common, shared physical space 40. The user of the remote third apparatus 10_3 occupies a different physical space 42. From the perspective of the user of the first apparatus 10_1 (the listening participant), the user of the second apparatus 10_2 (a contributing participant) is a local participant because they occupy the common, shared physical space 40. From the perspective of the user of the first apparatus 10_1 (the listening participant), the user of the third apparatus 10_3 (a contributing participant) is a remote participant because they occupy different, distinct physical spaces 40, 42. The distinct, separate physical spaces 40, 42 can be distant from each other. The first apparatus 10_1 provides a point-to-point, wireless, bi-directional link 70 between the lower latency rendering pipeline 20 of the first apparatus 10_1 and the second apparatus 10_2. The link 70 provides a channel for a separate audio and position metadata connection (stream) between the apparatuses 10_1, 10_2 sharing the same shared physical space 40. The link 70 carries both audio and position data. The audio data transmitted by the second apparatus 10_2, received at the first apparatus 10_1 is immediately rendered by the first apparatus 10_1 via the lower latency rendering pipeline 20. Likewise, any audio data transmitted by the first apparatus 10_1, received at the second apparatus 10_2 is immediately rendered by the second apparatus 10_2 via its lower latency rendering pipeline 20. The first apparatus 10_1 provides networked link 72 for the higher latency rendering pipeline 30 of the first apparatus 10_1 to the remote third apparatus 10_3. The remote network link 72 is a normal split-rendering path for the service 202 via the network 200. This has previously been described with reference to FIGs 1 and 2. The higher latency rendering pipeline 30 can for example be configured to perform split rendering for audio content originating from one or more remote participants as first-person perspective tracked binaural audio. The first apparatus 10_1 therefore receives two or more rendering streams that are rendered at the same time. The stream from the service 202, received via network link 72, defines the virtual space and will be processed via the higher latency rendering pipeline 30 to render the remote contributing participants (e.g. the user of the third apparatus 10_3). The stream direct from the second apparatus 10_2, received via link 70, defines the audio and position for the contributing participant using the second apparatus 10_2 is processed via the lower latency rendering pipeline 20 to render the local contributing participant (e.g. the user of the second apparatus 10_2). The same-space rendering performed by the lower latency rendering pipeline 20 is of low complexity compared to the remote rendering performed by the service 202 to create the virtual space. No complex audio coding is needed since full-band audio will only take few megabits at maximum and fit any local wireless link. For example, an uncompressed PCM stream can be transmitted, since in the local environment bandwidth is not an issue. Also, low frame size can be used to reduce the latency. The minimum achievable latency is the frame size Even if very low latency signal processing methods and fast hardware is used, a longer frame may significantly affect the achievable minimum latency. Typical digital wireless microphones already achieve latency in sub 5ms domain (~3ms being quite typical). Rendering complexity is manageable. No complex virtual spaces need to be rendered. The direction of arrival of the sound is calculated from the received position information and fast headtracking performed for binaural rendering. The amount of metadata required for head-tracked binaural rendering is small. Simple mono object audio rendering can be performed at the determined position of the local participant sound source. The lower latency rendering pipeline 20 can for example be configured to perform rendering for audio content originating from one or more local participants, as first-person perspective tracked binaural audio by positioning the audio content from a local participant as a sound source within the virtual space. The virtual space may be a common virtual space (with remote participants) and may not match the physical space where a listening low latency participant is. In at least some examples, in addition to panning direction and distance of a participant sound source, there is additional processing, which adds for example reverberation / echo and / or makes the rendered audio match better the virtual space and / or makes the rendered audio match better the physical space. In at least some examples, the lower latency pipeline 20 uses metadata about positions and relative poses between participants. These can be captured using some generic VR / AR / MR pose capturing methods running e.g. some laser / radio beacons or are captured from a headset using video processing, gyroscopes etc. The core concept described is about audio rendering, and it is assumed that relevant position / pose metadata will be provided, perhaps from gear worn by the participant. For example, a sound source for a local participant, for example a speech sound source, can be located using metadata that defines a position of the sound source in the virtual space. An absolute real-time position of the local participant in the participant’s real space can be converted to a real-time position of the local participant in the virtual space by comparing the real-time position of the local participant in the participant’s real space to a reference position of the participant’s real space. The apparatus 10 can, for example, receive an absolute real-time position of the local participant in the participant’s real space and perform a conversion to a position in the virtual space. The apparatus 10 can, for example, receive a position of the local participant in the virtual space. The conversion occurring elsewhere. The apparatus 10 can, for example, receive a change in real-time position of the local participant in the participant’s real space and use this to change a virtual position of the participant in the virtual space. The sound source for a local participant, for example a speech sound source, can be located at a position of the sound source in the virtual space using a head-related transfer function (HRTF) and the metadata that defines a position of the sound source in the virtual space. A HRTF is a function that is based on a sound source position in the virtual space and takes into account cues humans use to localize sounds. The HRTF creates a pair of finite impulse response (FIR) filters for a specific sound position in the virtual space. There is one FIR filter for the left ear and one FIR filter for the right ear. The pair of FIR filters that correspond to the position of the participant in the virtual space is applied to the sound source associated with that participate, producing a spatially located sound. A simple method to render binaural audio is pan a single audio source to a specific direction, which can be done using HRTF’s. If some distance (gain I equalization) is needed, that can also be performed. More advanced audio effects such as spatial reverb / echo increase complexity. In some cases direction panning (with, optionally, gain, equalization, and non-spatial reverb / echo provides a low latency rendering pipeline). In at least some examples, the lower latency rendering pipeline 20 locally adjusts rendering for a local participant sound source for a position of the local participant. The audio is received at the lower latency rendering pipeline 20 is without spatial audio rendering. In at least some examples, the higher latency rendering pipeline 30 uses remotely rendered spatial audio comprising positioned remote participant sound sources. The audio received at the higher latency rendering pipeline 30 is with spatial audio rendering. FIG. 5A illustrates the information flow (downlink flow) when a first apparatus 10_1 is being used by a listening participant. The lower latency rendering pipeline 20 receives audio and position information from a local apparatus for a local contributing participant. The Higher latency rendering pipeline 30 receives at least partially rendered audio. For example, the higher latency rendering pipeline 30 receives spatial rendered audio (for example pre-rendered spatial audio if split-rendering is used). In at least some 19 examples, the higher latency rendering pipeline 30 does not receive position information for the remote contributing participants. FIG. 5B illustrates the information flow (uplink flow) when a first apparatus 10_1 is being used by a contributing participant. Spatial capture is possible using various means. For example, a multi-microphone device such as a smartphone or professional ambisonics microphone. The first apparatus 10_1 sends audio and position information to local apparatus of local listening participants. The first apparatus 10_1 sends audio and position information to remote service 202. The communication remains the same for uplink / downlink with a local apparatus. The communication is different for uplink / downlink with the remote service 202. If a participant using the first apparatus 10_1 shares the same physical space with an apparatus 10J of another participant, then an additional stream / link between an apparatus 10_1, 10_2 is established. The streams consist of both audio and position. FIG. 6 illustrates an example based on FIG 4 with an additional remote apparatus (e.g. 10_4, not illustrated) used by a participant. A first participant 50_1 uses a first apparatus 10_1 (illustrated in FIG 3 but not FIG 6) at position 60_1 (orientation and / or potion) in a shared space 40. A second participant 50_2 uses a second apparatus 10_2 (illustrated in FIG 3 but not FIG 6) at position 60_2 (orientation and / or potion) in the shared space 40. A third participant 50_3 uses a third apparatus 10_3 (illustrated in FIG 3 but not FIG 6) at position 60_3 (orientation and / or potion) in a first unshared (remote) physical space 42_1. A fourth participant 50_4 uses a fourth apparatus (not illustrated in FIG 3 or FIG 6) at position 60_4 (orientation and / or potion) in a second unshared (remote) physical space 42_2. For clarity of description, it will be assumed that the first participant 50_1 is a listening participant, and that the other participants (the second participant 50_2, the third participant 50_3, the fourth participant 50_4) are contributing participants. However, the user of any of the apparatus 10 can be both a listening participant and / or a contributing participant, and this can change over time. The virtual space rendered to the first participant 50_1 is illustrated as an overlay on the shared physical space 40. The circles in FIG 6 illustrate the participants 50_i and their respective positions 60J. The local participant sound source(s) for local participant 50_2 is rendered at the position (orientation and / or location) of the local participant 50_2. The triangles in FIG 6 illustrate virtual participants 52J for remote participants 50_3, 50_4. The virtual participants 52J have respective positions 62_i. The remote participant sound source(s) for the third participant 50_3 is rendered at the position 62_3 (orientation and / or location) of the virtual participant 52_3. The remote participant sound source(s) for the fourth participant 50_4 is rendered at the position 62_4 (orientation and / or location) of the virtual participant 52_4. A change in the position 60_2 (orientation and / or location) of the second participant 50_2 results in a corresponding change in the position 60_2 (orientation and / or location) of the local participant sound source(s). The change in the position (orientation and / or location) of the local participant sound source(s) (virtual participant) tracks the change in position 60_2 (orientation and / or location) of the second participant 50_2. The local participant sound source(s) for the second participant 50_2 remain co-positioned with second participant 50_2 as the position 60_2 of the second participant 50_2 changes. A change in the position 60_3 (orientation and / or location) of the third participant 50_3 results in a corresponding change in the position 62_3 (orientation and / or location) of the virtual third participant 52_3. The change in the position 62_3 (orientation and / or location) of the virtual third participant 52_3 tracks the change in position 60_3 (orientation and / or location) of the third participant 50_3. A change in the position 60_4 (orientation and / or location) of the fourth participant 50_4 results in a corresponding change in the position 62_4 (orientation and / or location) of the virtual fourth participant 52_4. The change in the position 62_4 (orientation and / or location) of the virtual fourth participant 52_4 tracks the change in position 60_4 (orientation and / or location) of the fourth participant 50_4. The virtual space for the first participant comprises the contributing participant sound sources for the participants 50_2, 50_3, 50_4. A change in the position 60_1 (orientation and / or location) of the first participant 50_1 results in a corresponding change (orientation and / or location) of the virtual space relative to the first participant and the virtual listener. FIG. 7 illustrates, compared to FIG 6, a change in position 60_3 (orientation) of the third participant 50_3, a change in position 60_2 (orientation) of the second participant 50_2, and a change in position 60_1 (orientation) of the first participant 50_1. This represents a 3DoF implementation for both listening participant and for contributing participant. The change in position 60_3 (orientation) of the third participant 50_3 (a remote participant) results in a corresponding change in position 62_3 (orientation) of the remote participant sound source represented by the virtual third participant 52_3 in the virtual space. The change in position 60_2 (orientation) of the second participant 50_2 (a local participant) results in a corresponding change in the co-located position (orientation) of the local participant sound source in the virtual space. The change in position 60_1 (orientation) of the first participant 50_1 results in a re-positioning (reorientation) of the virtual space relative to the notional virtual listener used for rendering the virtual space to the first participant 50_1. FIG. 8 illustrates, compared to FIG 6, a change in position 60_3 (orientation and location) of the third participant 50_3, a change in position 60_2 (orientation and location) of the second participant 50_2, and a change in position 60_1 (orientation and location) of the first participant 50_1. This represents a 6DoF implementation for both listening participant and for contributing participant. The change in position 60_3 (orientation and location) of the third participant 50_3 (a remote participant) results in a corresponding change in position 62_3 (orientation and location) of the remote participant sound source represented by the virtual participant 52_3 in the virtual space. The change in position 60_2 (orientation and position) of the second participant 50_2 (a local participant) results in a corresponding change in the co-located position (orientation and location) of the local participant sound source in the virtual space. The change in position 60_1 (orientation and location) of the first participant 50_1 results in a re-positioning (re-orientation and location) of the virtual space relative to the notional virtual listener used for rendering the virtual space to the first participant 50_1. In some examples, there may be a 3DoF implementation for one or more listening participants and a 6Dof implementation for one or more contributing participants. In some examples, there may be a 6DoF implementation for one or more listening participants and a 3DoF implementation for one or more contributing participants. In some examples, there may be a 3DoF implementation for one or more listening participants and a 3DoF implementation for one or more contributing participants. In some examples, there may be a 6DoF implementation for one or more listening participants and a 6DoF implementation for one or more contributing participants. In some examples, the implementation can be controlled to be 3DoF, 3DoF+, 6DoF for any one or more contributing participants and can be controlled to be 3DoF, 3DoF+, 6DoF for any one or more listening participants. In the examples of FIGs 6, 7, 8 four participants are sharing the same virtual space but not the same physical space. Two participants 50_1, 50_2 share the same physical space 40. Two additional meeting participants 50_3, 50_4 are at two different remote physical spaces 42_1,42_2 and presented as avatars 52_3, 52_4 in the virtual space. For users alone in their location, the rendered virtual space (visual and audio) based on their position is the one provided by the remote service (full split-rendering scene). For a first user sharing the same physical space with a second user, the rendered virtual space (visual and audio) based on the first user’s position is a combination of the virtual space or scene provided by the remote service 202 (full split-rendering scene) but for remote participants only; and audio and position of the second user, provided locally. The stream coming from the remote service 202 does not contain the audio or metadata for the other user in the same space or it is ignored by the renderer and local stream lower latency stream is preferred. If no low latency stream is established between space sharing users, they have possibility to use acoustic coupling for audio, if open headphones are used and users are close enough to hear each other directly. Alternatively or in addition, as a backup high latency split rendered full stream from the cloud / edge can be consumed especially if users are relative far from each other and they cannot otherwise hear or see the other user clearly. FIGs. 9A &9B illustrate an effect of latency. The apparent angle of arrival changes due to latency, when there is relative movement between participants 50 sharing the same physical space. In FIG 9A, a participant 50_2 is moving 51_2 in the same space as the person 50_1 (stationary) In FIG 9B, a participant 50_2 is stationary in the same space as the person 50_1 who is moving 51_1. In other examples all or some of participants 50 can be moving 51. If position information as well as uplink audio is sent all the way to the service 202 for processing and then the pre-rendered signal is transmitted back, there will be significant latency. If two users are in same room this will cause visible and audible discrepancies, especially when there is line of sight between the users sharing the same space. Since local users may be directly visible to each other it would be desirable to have low latency. If normal 6DoF audio rendering is done at the remote service 202, the latency will be significant for the users in the same physical space. Audio latency will cause immediate lip-sync problems, if users are in line of sight. Position latency will cause unnatural separation between the local participant audio source and the local participant position because of delay. As illustrated in FIG 9A and FIG 9B, direction of audio arrival will not match the current position. If two users are sharing the same physical and virtual space, the apparatuses 10 allows them to communicate with each other and other remote participants with high perceived quality. The direct audio and metadata stream between local participants sharing the same space overcomes the latency problem. A direct link is locally established between two or more users in the same room for this purpose. FIG. 10 illustrates an example of method 500 suitable for set-up and configuration, or re-configuration. For clarity of description, it will be assumed that the first participant 50_1 is a listening participant, and that the other participants (the second participant 50_2, the third participant 50_3, the fourth participant 50_3) are contributing participants. However, the user of any of the apparatus 10 can be both a listening participant and / or a contributing participant, and this can change over time. At block 502, an apparatus 10_1 used by the first participant 50_1 determines whether any other participants, for example actual or possible contributing participants, are ‘local’. A minimum required for a participant to be local to the listening participant is that they share the same physical space, for example the same room. This can be determined for example via positioning information or user input from the listening participant or the other participant. Additional requirements for a determination of local can, for example, be that the participant and the listening participant are in line-of-sight. This can be determined for example via positioning information or via signal exchange or via user input from the listening participant or the other participant. Further additional requirements for a determination of local can, for example, be that the participant and the listening participant are within a defined distance. This can be determined for example via positioning information or via signal exchange or via user input from the listening participant or the other participant. At block 504, the first apparatus 10_1 configures a direct point-to-point communication link 70 with each of the local apparatuses 10_i used by the local contributing participants. At block 506, the first apparatus 10_1 and the local apparatus 10_i exchange parameters to allow the bi-directional exchange of audio and position information via the link 70. The first apparatus 10_1 informs the service 202 that it does not require centrally provided audio for the local contributing participants. At block 508, the first apparatus 10_1 receives from the service 202 the immersive audio representing the virtual space and including the properly positioned participant sound sources of the remote contributing participants (but not including participant sound sources of the local contributing participants) and processes this data via the higher latency rendering pipeline 30 as described above. The first apparatus 10_1 receives the audio and position information from the local apparatuses 10_i of the contributing participants via the created direct point-to-point communication links and processes this data via the lower latency rendering pipeline 20 as described above. The first apparatus 10_1 performs final rendering taking into account the current position of the listening participant and renders a virtual space that is correctly positioned relative to the listening participant’s position (orientation and / or location) and which includes correctly position sound sources for the local contributing participants (via the lower latency rendering pipeline 20) and the remote contributing participants (via the higher latency rendering pipeline 30). In some examples, at block 506, the first apparatus 10_1 and the local apparatus 10_i exchange parameters to allow the bi-directional exchange of audio and position information but the first apparatus 10_1 does not inform the service 202 that it does not require centrally provided audio for the local contributing participants. Then at block 508, the first apparatus 10_1 receives the immersive audio representing the virtual space and including the properly positioned participant sound sources of the remote contributing participants and including positioned participant sound sources of the local contributing participants and processes this data via the higher latency rendering pipeline 30 as described. This processing includes removing the participant sound sources of the local contributing participants. The first apparatus 10_1 receives the audio and position information from the local apparatuses 10J of the local contributing participants via the created direct point-to-point communication link 70 and processes this data via the lower latency rendering pipeline 20 as described above. The first apparatus 10_1 performs final rendering taking into account the current position (orientation and / or location) of the listening participant and renders a virtual space that is correctly positioned relative to the listening participant’s position and which includes correctly position sound sources for the local contributing participants (via the lower latency rendering pipeline 20) and the remote contributing participants (via the higher latency rendering pipeline 30). The above method can be adapted to accommodate more, or less, local participants and more or less remote participants. For a local participant relative to the listening participant, the audio and position metadata from the local participant is rendered locally for the listening participant. For a remote participant relative to the listening participant, the audio and position metadata from the remote participant is rendered remotely for the listening participant. It may be partially rendered remotely if split-rendering is used. For a local participant relative to the listening participant, the audio and position metadata for the listening participant is sent locally, for local rendering before being listen to by the local participant. For a remote participant relative to the listening participant, the audio and position metadata from the listening participant is sent to the service 202 for remote rendering before being sent to the local participant. It may be partially rendered remotely if splitrendering is used. Fig 11 illustrates an example of a controller 400 suitable for use in an apparatus 10. Implementation of a controller 400 may be as controller circuitry. The controller 400 may be implemented in hardware alone, have certain aspects in software including firmware alone or can be a combination of hardware and software (including firmware). As illustrated in Fig 11 the controller 400 may be implemented using instructions that enable hardware functionality, for example, by using executable instructions 406 in a general-purpose or special-purpose processor 402 that may be stored on a machine-readable storage medium (disk, memory etc.) to be executed by such a processor 402. The processor 402 is configured to read from and write to the memory 404. The processor 402 may also comprise an output interface via which data and / or commands are output by the processor 402 and an input interface via which data and / or commands are input to the processor 402. The memory 404 stores instructions, program, or code 406 that controls the operation of the apparatus 10 when loaded into the processor 402. The computer program instructions, program or code am 406, provide the logic and routines that enables the apparatus 10 to perform the methods illustrated in the accompanying FIGs. The processor 402 by reading the memory 404 is configured to load and execute the instructions, program, or code 406. The apparatus 10 comprises: at least one processor 402; and at least one memory 404 storing instructions that, when executed by the at least one processor 402, cause the apparatus at least to: provide a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and provide a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. In some examples, there is a (computer implemented) system comprising: means for rendering first-person-perspective-mediated-reality* comprising: at least a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and at least a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. As illustrated in Fig 12, the instructions, program, or code 406 may arrive at the apparatus 10 via any suitable delivery mechanism 408. The delivery mechanism 408 may be, for example, a machine readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a record medium such as a Compact Disc Read-Only Memory (CD-ROM) or a Digital Versatile Disc (DVD) or a solid-state memory, an article of manufacture that comprises or tangibly embodies the computer program 406. The delivery mechanism may be a signal configured to reliably transfer the computer program 406. The apparatus 10 may propagate or transmit the computer program 406 as a computer data signal. The term “non-transitory”, as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM). Computer program instructions for causing an apparatus to perform at least the following or for performing at least the following: provide a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; provide a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio. The computer program instructions may be comprised in a computer program, a non-transitory computer readable medium, a computer program product, a machine readable medium. In some but not necessarily all examples, the computer program instructions may be distributed over more than one computer program. Although the memory 404 is illustrated as a single component / circuitry it may be implemented as one or more separate components / circuitry some or all of which may be integrated / removable and / or may provide permanent / semi-permanent / dynamic / cached storage. Although the processor 402 is illustrated as a single component / circuitry it may be implemented as one or more separate components / circuitry some or all of which may be integrated / removable. The processor 402 may be a single core or multi-core processor. References to ‘computer-readable storage medium’, ‘computer program product’, ‘tangibly embodied computer program’ etc. or a ‘controller’, ‘computer’, ‘processor’ etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device whether instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device etc. As used in this application, the term ‘circuitry’ may refer to one or more or all the following: (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii)any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory or memories that work together to cause an apparatus, such as a mobile phone or server, to perform various functions and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (for example, firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the claim element, a baseband integrated circuit for a mobile device or a similar integrated circuit in a server, a cellular network device, or other computing or network device. The blocks illustrated in the accompanying Figs may represent steps in a method and / or sections of code in the computer program 406. The illustration of a particular order to the blocks does not necessarily imply that there is a required or preferred order for the blocks and the order and arrangement of the block may be varied. Furthermore, it may be possible for some blocks to be omitted. As used here ‘module’ refers to a unit or apparatus that excludes certain parts / components that would be added by an end manufacturer or a user. The apparatus 10 can, for example be a module. A controller 400 of the apparatus 10 can, for example be a module. Where a structural feature has been described, it may be replaced by means for performing one or more of the functions of the structural feature whether that function or those functions are explicitly or implicitly described. The above-described examples find application as enabling components of: automotive systems; telecommunication systems; electronic systems including consumer electronic products; distributed computing systems; media systems for generating or rendering media content including audio, visual and audio visual content and mixed, mediated, virtual and / or augmented reality; personal systems including personal health systems or personal fitness systems; navigation systems; user interfaces also known as human machine interfaces; networks including cellular, non- cellular, and optical networks; ad-hoc networks; the internet; the internet of things; virtualized networks; and related software and services. The apparatus can be provided in an electronic device, for example, a mobile terminal, according to an example of the present disclosure. It should be understood, however, that a mobile terminal is merely illustrative of an electronic device that would benefit from examples of implementations of the present disclosure and, therefore, should not be taken to limit the scope of the present disclosure to the same. While in certain implementation examples, the apparatus can be provided in a mobile terminal, other types of electronic devices, such as, but not limited to: mobile communication devices, hand portable electronic devices, wearable computing devices, portable digital assistants (PDAs), pagers, mobile computers, desktop computers, televisions, gaming devices, laptop computers, cameras, video recorders, GPS devices and other types of electronic systems, can readily employ examples of the present disclosure. Furthermore, devices can readily employ examples of the present disclosure regardless of their intent to provide mobility. The term ‘comprise’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising Y indicates that X may comprise only one Y or may comprise more than one Y. If it is intended to use ‘comprise’ with an exclusive meaning then it will be made clear in the context by referring to ‘comprising only one...’ or by using ‘consisting.’ In this description, the wording ‘connect’, ‘couple’ and ‘communication’ and their derivatives mean operationally connected / coupled / in communication. It should be appreciated that any number or combination of intervening components can exist (including no intervening components), i.e., to provide direct or indirect connection / coupling / communication. Any such intervening components can include hardware and / or software components. As used herein, the term "determine / determining" (and grammatical variants thereof) can include, not least: calculating, computing, processing, deriving, measuring, investigating, identifying, looking up (for example, looking up in a table, a database, or another data structure), ascertaining and the like. Also, "determining" can include receiving (for example, receiving information), accessing (for example, accessing data in a memory), obtaining and the like. Also," determine / determining" can include resolving, selecting, choosing, establishing, and the like. In this description, reference has been made to various examples. The description of features or functions in relation to an example indicates that those features or functions are present in that example. The use of the term ‘example’ or ‘for example’ or ‘can’ or ‘may’ in the text denotes, whether explicitly stated or not, that such features or functions are present in at least the described example, whether described as an example or not, and that they can be, but are not necessarily, present in some of or all other examples. Thus ‘example’, ‘for example’, ‘can’, or ‘may’ refers to a particular instance in a class of examples. A property of the instance can be a property of only that instance or a property of the class or a property of a sub-class of the class that includes some but not all the instances in the class. It is therefore implicitly disclosed that a feature described with reference to one example but not with reference to another example, can where possible be used in that other example as part of a working combination but does not necessarily have to be used in that other example. As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements. Although examples have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the claims. Features described in the preceding description may be used in combinations other than the combinations explicitly described above. Although functions have been described with reference to certain features, those functions may be performable by other features whether described or not. Although features have been described with reference to certain examples, those features may also be present in other examples whether described or not. The term ‘a’, ‘an’ or ‘the’ is used in this document with an inclusive not an exclusive meaning. That is any reference to X comprising a / an / the Y indicates that X may comprise only one Y or may comprise more than one Y unless the context clearly indicates the contrary. If it is intended to use ‘a’, ‘an’ or ‘the’ with an exclusive meaning then it will be made clear in the context. In some circumstances the use of ‘at least one’ or ‘one or 32 more’ may be used to emphasis an inclusive meaning but the absence of these terms should not be taken to infer any exclusive meaning. The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and to features that achieve substantially the same technical effect (equivalent features). The equivalent features include, for example, features that are variants and achieve substantially the same result in substantially the same way. The equivalent features include, for example, features that perform substantially the same function, in substantially the same way to achieve substantially the same result. In this description, reference has been made to various examples using adjectives or adjectival phrases to describe characteristics of the examples. Such a description of a characteristic in relation to an example indicates that the characteristic is present in some examples exactly as described and is present in other examples substantially as described. The above description describes some examples of the present disclosure however those of ordinary skill in the art will be aware of possible alternative structures and method features which offer equivalent functionality to the specific examples of such structures and features described herein above and which for the sake of brevity and clarity have been omitted from the above description. Nonetheless, the above description should be read as implicitly including reference to such alternative structures and method features which provide equivalent functionality unless such alternative structures or method features are explicitly excluded in the above description of the examples of the present disclosure. Whilst endeavoring in the foregoing specification to draw attention to those features believed to be of importance the Applicant may seek protection via the claims in respect of any patentable feature or combination of features hereinbefore referred to and / or shown in the drawings whether or not emphasis has been placed thereon. l / we claim:
Claims
1. An apparatus comprising:means for audio rendering comprising:at least a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; andat least a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio.
2. An apparatus as claimed in claim 1, wherein the means for audio rendering is configured to provide first-person-perspective-mediated-reality wherein a listening participant’s real position determines, overtime, a position a virtual listening participant within a virtual space for the first-person perspective tracked binaural audio.
3. An apparatus as claimed in claim 1 or 2, wherein the lower latency rendering pipeline and the higher latency rendering pipeline are configured to operate simultaneously.
4. An apparatus as claimed in claim 1, 2 or 3, comprising means for using a point-to-point, wireless communication link between the lower latency pipeline and an apparatus of the local, visible, or potentially visible participant.
5. An apparatus as claimed in claim 4, wherein the point-to-point, wireless communication link is a point-to-point, wireless, bi-directional communication link between the lower latency pipeline and the apparatus of the local, visible, or potentially visible participant.
6. An apparatus as claimed in any preceding claim comprising:means for detecting a local apparatus of the local, visible, or potentially visible participant; andmeans for coupling the lower latency rendering pipeline to the local apparatus via a direct, point-to-point, wireless communication link, to receive at least position information for a remote apparatus and audio captured by the remote apparatus.
7. An apparatus as claimed in claim 6, comprising means for coupling the higher latency rendering pipeline to a remote service via a network link, to receive at least partially rendered spatial audio content.
8. An apparatus as claimed in claim 6 or 7, comprising means for coupling the higher latency rendering pipeline to a remote service via a network link, to receive split-rendered spatial audio content.
9. An apparatus as claimed in any preceding claim wherein the lower latency rendering pipeline receives, for the local, visible, or potentially visible, participant, a position of the local participant and audio associated with the participant; andwherein the higher latency rendering pipeline receives, for the remote participant at least partially rendered audio representing a virtual space comprising multiple positioned sound sources.
10. An apparatus as claimed in claim 9, configured to adjust a position of the virtual space is response to adjustment of a position of the apparatus to provide first-person-perspective tracked binaural audio.
11. An apparatus as claimed in claim 10, comprising means for positioning the apparatus.
12. An apparatus as claimed in any preceding claim wherein the lower latency rendering pipeline is configured for single sound source spatial rendering.
13. An apparatus as claimed in claim 12 wherein the higher latency rendering pipeline is configured for multiple sound source spatial rendering.
14. An apparatus as claimed in any preceding claim, whereinthe apparatus is configured to receive uncompressed audio for the lower latency rendering pipeline and the apparatus is configured to receive encoded audio for the higher latency rendering pipeline.
15. An apparatus as claimed in any preceding claim comprisingmeans for coupling the apparatus to a lower latency rendering pipeline of a local apparatus of the local, visible, or potentially visible participant to provide at least position information for the apparatus and audio captured by the apparatus to the local apparatus.
16. An apparatus as claimed in claim 15, comprisingmeans for coupling the apparatus, via a network link, to a remote service to provide to the remote service at least position information for the apparatus and audio captured by the apparatus.
17. An apparatus as claimed in any preceding claim wherein the lower latency rendering pipeline is for rendering audio with a latency below 20ms and the higher latency rendering pipeline is for rendering audio with a latency above 20ms.
18. An apparatus as claimed in any preceding claim, configured as a head mounted apparatus comprising means for tracking a position of a head of a user.
19. A computer program comprising instructions that when executed by at least one processor of an apparatus cause the apparatus to:provide a lower latency rendering pipeline for rendering audio content originating from a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; andprovide a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio.
20. A method comprising:rendering audio content received directly from an apparatus of a local, visible, or potentially visible participant as first-person perspective tracked binaural audio; and rendering a virtual space comprising a positioned audio source originating from a remote participant as first-person perspective tracked binaural audio.
21. An apparatus comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to:provide a lower latency rendering pipeline for rendering audio content originating froma local, visible, or potentially visible participant as first-person perspective tracked binaural audio; andprovide a higher latency rendering pipeline for rendering audio content originating from5 a remote participant as first-person perspective tracked binaural audio.
22. A system comprising:means for rendering first-person-perspective-mediated-reality* comprising:at least a lower latency rendering pipeline for rendering audio content originating from a 10 local, visible, or potentially visible participant as first-person perspective trackedbinaural audio; andat least a higher latency rendering pipeline for rendering audio content originating from a remote participant as first-person perspective tracked binaural audio.
Citation Information
Patent Citations
Object prioritisation of virtual content
GB2568726A
Audio spatialization and reinforcement between multiple headsets
US10873825B2
Spatial audio teleconferencing
US20140016793A1
Audio apparatus, audio distribution system and method of operation therefor
US20220137916A1
System and method for spatially projected audio communication
WO2021044419A1