Audio signal processing
By determining a small number of reference head orientations and processing associated spatial audio signals and metadata in the pre-rendering device, the problems of high processing complexity and linearly increasing resource requirements in the prior art are solved, thereby improving the system's processing efficiency and resource utilization.
Patent Information
- Application Number
- CN202510505939.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2025-04-22
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies, when processing spatial audio signals, especially in multi-device rendering scenarios, suffer from high processing complexity and resource requirements that increase linearly with the number of devices, resulting in low efficiency.
A small number of reference head orientations are determined by a pre-rendering device, and spatial audio signals and metadata associated with these orientations are processed and sent to multiple post-rendering devices, which adjust according to the actual head orientation.
It reduces the processing load of the pre-rendering device, improves the system's processing efficiency and resource utilization, and adapts to the needs of more post-rendering devices.
Smart Images

Figure CN120835265A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Example embodiments relate to audio signal processing. BACKGROUND
[0002] Spatial audio refers to audio that, when output to a user device such as a pair of headphones, enables a user to perceive audio sources as if they came from respective directions relative to the user’s position. For example, one audio source can be perceived as coming from a position in front of the user, while one or more other audio sources can be perceived as coming from positions to the left and / or right of the user. Spatial audio can be more complex to transmit, decode and render than other audio formats, and therefore partitioned rendering approaches have been proposed whereby the rendering operation can be divided into different stages, with different stages being performed by different devices. SUMMARY
[0003] The scope of protection sought for various embodiments of the present invention is set forth by the independent claims. The embodiments and features that have not been fallen within the scope of the independent claims, if any, will be interpreted as examples useful for understanding various embodiments of the present invention.
[0004] According to a first aspect, an apparatus is described, comprising: means for receiving an audio signal; means for receiving, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests are indicative of respective user head orientations; means for determining a plurality, N, of reference head orientations, wherein N < M; means for processing the received audio signal to obtain a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of metadata sets respectively associated with the plurality of spatial audio signals, wherein a metadata set comprises information indicative of how an associated spatial audio signal should be adjusted to take into account a change in head orientation relative to an associated reference head orientation; and means for sending, to at least one of the devices, a selected spatial audio signal and an associated metadata set, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
[0005] In some example embodiments, the apparatus can further comprise means for detecting that the plurality, M, of devices comprises a number greater than a threshold number, G, wherein the processing is performed in response to the detection.
[0006] In some example embodiments, the plurality, N, of reference head orientations can comprise a number equal to the threshold number, G.
[0007] In some example embodiments, the plurality, N, of reference head orientations can be determined prior to receiving the respective requests.
[0008] In some example embodiments, the plurality, N, of reference head orientations can be determined based at least in part on the respective user head orientations indicated in the respective requests.
[0009] In some example embodiments, the plurality, N, of reference head orientations can be determined by: arranging the plurality, M, of user head orientations into N groups of one or more user head orientations, wherein at least one group comprises two or more user head orientations; and determining, for the N groups, a respective reference head orientation, wherein the reference head orientation for the at least one group comprising two or more reference head orientations is determined based on at least one of the two or more reference head orientations.
[0010] In some example embodiments, the reference head orientation for the at least one group can comprise one of the respective user head orientations.
[0011] In some example embodiments, the reference head orientation for the at least one group can comprise an average of the respective user head orientations.
[0012] In some example embodiments, the at least one group can comprise two or more respective head orientations that are most similar or within a similarity threshold.
[0013] In some example embodiments, the arranging can comprise: determining direction vectors associated with the respective user head orientations received from the plurality, M, of devices; identifying which spatial partition in a first set of spatial partitions corresponds to a maximum number of direction vectors; dividing the identified spatial partition into two or more spatial partitions; if the number of spatial partitions is not equal to N, re-performing the identifying and dividing operations until the number of spatial partitions is equal to N.
[0014] In some example embodiments, the selected spatial audio signal for a particular device can be a spatial audio signal associated with the reference head orientation of the group into which the respective user head orientation from the particular device is arranged.
[0015] In some example embodiments, the reference head orientation can comprise an orientation for at least one of: a yaw axis; a yaw axis and a pitch axis; or a yaw axis, a pitch axis, and a roll axis.
[0016] In some example embodiments, the apparatus can comprise at least one of: a user device or a server.
[0017] In some example embodiments, the plurality, M, of devices can comprise at least one of: a headphone device, a loudspeaker device, or a user device.
[0018] According to a second aspect, a method is described, comprising: receiving an audio signal; receiving, from a plurality, M, of devices, respective requests for spatial audio signals, wherein the respective requests are indicative of respective user head orientations; determining a plurality, N, of reference head orientations, wherein N < M; processing the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of metadata sets respectively associated with the plurality of spatial audio signals, wherein a metadata set comprises information indicative of how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and sending, to at least one of the devices, a selected spatial audio signal and an associated metadata, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
[0019] In some example embodiments, the method can further comprise detecting that the plurality, M, of devices comprises a number greater than a threshold number, G, wherein the processing is performed in response to the detection.
[0020] In some example embodiments, the plurality, N, of reference head orientations can comprise a number equal to the threshold number, G.
[0021] In some example embodiments, the plurality, N, of reference head orientations can be determined prior to receiving the respective requests.
[0022] In some example embodiments, the plurality, N, of reference head orientations can be determined based at least in part on the respective user head orientations indicated in the respective requests.
[0023] In some example embodiments, the plurality, N, of reference head orientations can be determined by: arranging the plurality, M, of user head orientations into N groups of one or more user head orientations, wherein at least one group comprises two or more user head orientations; and determining, for the N groups, respective reference head orientations, wherein the reference head orientation for the at least one group comprising two or more reference head orientations is determined based on at least one of the two or more reference head orientations.
[0024] In some example embodiments, the reference head orientation for the at least one group can comprise one of the respective user head orientations.
[0025] In some example embodiments, the reference head orientation for the at least one group can comprise an average of the respective user head orientations.
[0026] In some example embodiments, the at least one group can comprise two or more respective head orientations that are most similar or within a similarity threshold.
[0027] In some example embodiments, the arrangement can comprise: determining directional vectors associated with the respective user head orientations received from the plurality, M, of devices; identifying which spatial partition in a first set of spatial partitions corresponds to a maximum number of directional vectors; dividing the identified spatial partition into two or more spatial partitions; and if the number of spatial partitions is not equal to N, re-performing the identifying and dividing operations until the number of spatial partitions is equal to N.
[0028] In some example embodiments, the selected spatial audio signal for a particular device can be a spatial audio signal associated with the reference head orientation of the group into which the respective user head orientation from the particular device is arranged.
[0029] In some example embodiments, the reference head orientation can comprise an orientation for at least one of: a yaw axis; a yaw axis and a pitch axis; or a yaw axis, a pitch axis, and a roll axis.
[0030] In some example embodiments, the method can be performed by an apparatus comprised in at least one of a user device or a server.
[0031] In some example embodiments, the plurality, M, of devices can be comprised in at least one of: a headphone device, a loudspeaker device, or a user device.
[0032] According to a third aspect, there is provided a computer program product comprising a set of instructions which, when executed on an apparatus, is configured to cause the apparatus to perform a method comprising: receiving an audio signal; receiving, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests are indicative of respective user head orientations; determining a plurality, N, of reference head orientations, wherein N < M; processing the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of sets of metadata respectively associated with the plurality of spatial audio signals, wherein a set of metadata comprises information indicative of how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and sending, to at least one of the devices, a selected spatial audio signal and associated metadata, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
[0033] The third aspect can further comprise any feature described in relation to the second aspect.
[0034] According to a fourth aspect, there is provided a non-transitory computer readable medium comprising program instructions stored thereon for performing a method comprising: receiving an audio signal; receiving, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests are indicative of respective user head orientations; determining a plurality, N, of reference head orientations, wherein N < M; processing the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of sets of metadata respectively associated with the plurality of spatial audio signals, wherein a set of metadata comprises information indicative of how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and sending, to at least one of the devices, a selected spatial audio signal and associated metadata, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
[0035] The fourth aspect can further comprise any feature described in relation to the second aspect.
[0036] According to a fifth aspect, there is provided an apparatus comprising: at least one processor; and at least one memory including computer program code, the computer program code, when executed by the at least one processor, causing the apparatus to: receive an audio signal; receive, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests are indicative of respective user head orientations; determine a plurality, N, of reference head orientations, wherein N < M; process the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of metadata sets respectively associated with the plurality of spatial audio signals, wherein a metadata set comprises information indicative of how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and send, to at least one of the devices, a selected spatial audio signal and an associated metadata, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
[0037] The fifth aspect can further comprise any feature described in relation to the second aspect. BRIEF DESCRIPTION OF DRAWINGS
[0038] Example embodiments will now be described, by way of example only, with reference to the accompanying drawings in which:
[0039] Figure 1 A system that can be used to understand example embodiments is shown;
[0040] Figure 2 A user listening to a spatial audio scene that can be useful to understand example embodiments is shown;
[0041] Figure 3 A user listening to a spatial audio scene when head tracking is not used is shown; Figure 2
[0042] Figure 4 A user listening to a spatial audio scene when head tracking is used is shown; Figure 2
[0043] Figure 5 A split-rendering system that can be used to understand example embodiments is shown;
[0044] Figure 6 is a flowchart showing operations according to one or more example embodiments;
[0045] Figure 7 A split-rendering system according to one or more example embodiments is shown;
[0046] Figure 8 Another split-rendering system is shown in accordance with one or more example embodiments;
[0047] Figure 9 is a flowchart showing operations in accordance with one or more example embodiments;
[0048] FIG. 10 graphically shows clustering operations in accordance with one or more example embodiments;
[0049] Figure 11 An apparatus that can be configured to operate in accordance with example embodiments is shown; and
[0050] Figure 12 A non-transitory computer-readable medium for storing computer-readable instructions for causing an apparatus to Figure 11 An apparatus to operate in accordance with example embodiments. DETAILED DESCRIPTION
[0051] Example embodiments relate to audio signal processing.
[0052] An audio signal can include spatial audio data. Spatial audio refers to audio that enables a user to perceive audio sources as if they came from respective directions relative to the user’s position when output to a user device such as a pair of headphones. For example, one audio source can be perceived as coming from a position in front of the user, while one or more other audio sources can be perceived as coming from positions to the left and / or right of the user.
[0053] Example formats of spatial audio data can include, but are not limited to, multi-channel mix, high- fidelity stereo sound reproduction, parametric spatial audio (e.g., metadata- assisted spatial audio (MASA)), object-based audio, or any combination thereof. Spatial audio data can be encoded and decoded using a codec, which can include, but is not limited to, the 3GPP Immersive Video and Audio Service (IVAS) standard.
[0054] In the case where, for example, the audio output device includes a head-worn device that includes a pair of speakers (examples are a pair of headphones, earbuds, earphones, or an extended reality (XR) headset), binaural rendering can be used to render the spatial audio data. In binaural rendering, a binaural rendering module of the audio output device or an associated media player can use various algorithms based on head-related impulse responses (HRIRs) or frequency-domain equivalents to provide spatial reproduction such that the user perceives audio sources as if they were positioned within the spatial audio scene.
[0055] Head tracking can also be performed as part of the rendering process. This can involve tracking the position, such as orientation, of the user's head and compensating or correcting the binaural rendering so that the audio sources are perceived as remaining stationary even if the user's head moves or rotates. This can be performed by providing a reference or "0" orientation, where one or more audio sources are perceived from corresponding spatial positions relative to the reference orientation. In response to a change in the tracked user orientation from the reference orientation to a new orientation, the spatial audio scene can be modified to compensate for the tracked change, so that the user's perception is that one or more audio sources remain stationary in the spatial audio scene, which mimics how the user perceives sounds in the real world and provides a strong cue for spatial audio perception. It is well known in spatial audio research that reliable head tracking significantly improves the quality of binaural rendering and allows for good perceptual quality even when the rendering might otherwise lack accuracy. In some cases, the spatial audio signal representing the spatial audio scene can be modified prior to rendering, and in some cases, the binaural rendering itself can be modified.
[0056] Figure 1 is a block diagram of system 100 that may be useful in understanding example embodiments.
[0057] System 100 may include a server 110 , a media player 120 , a network 130 , and an audio output device, which in this example comprises a set of headphones 140 worn by a user 150 .
[0058] The server 110 can be connected to the media player 120 over the network 130 in order to send spatial audio data to the media player 120. The server 110 may, for example, comprise an Internet Protocol (IP) telecommunication server that sends spatial audio data comprising part of a voice call to the media player 120. The spatial audio data can represent one or more other users that are participants of the voice call, such that when their respective audio data is rendered and output to the headset set 140 by the media player 120, their respective audio data will be perceived from respective directions. Alternatively, the spatial audio data can represent a music track or an audio track of a movie or a mix of speech and music. The sending can be by means of any suitable streaming data protocol. Alternatively or additionally, the server 110 can provide one or more files representing the spatial audio data to the media player 120 for storage and processing there. At the media player 120, the spatial audio data can be processed and rendered to the headset set 140. In example embodiments, the headset set 140 can comprise head tracking sensors for providing head tracking data to the media player 120, using any suitable method to indicate the head orientation of the user or changes in the head orientation of the user, in order to determine how the spatial audio data is to be rendered at the headset 140. One or more known head tracking methods can be used, such as by using one or more inertial sensors (e.g. gyroscopes and / or accelerometers) within or attached to the headset 140 to determine the head orientation of the user in real time or near real time. Alternative or additional examples can include the use of one or more cameras that can identify facial features in real time.
[0059] In some example embodiments, the media player 120 can comprise one of a mobile phone, a tablet computer, a game console, a laptop computer, a personal computer, a wearable device or any device comprising or including a bitstream decoder. In some example embodiments, the media player 120 can comprise part of the headset set 140.
[0060] The network can be any suitable data communication network, including for example one or more of a Radio Access Network (RAN) that communicates through one or more base stations, a WiFi network that communicates through one or more access points, or a short range network such as a short range network using Bluetooth or Zigbee protocols.
[0061] Figure 2 、 3 Figures 1, 2, 3 and 4 are representative diagrams of the user 150 when wearing the headset set 140, which can also be used to understand example embodiments.
[0062] Reference is made to Figure 2Fig. 1 shows a user 150 listening to a rendered spatial audio field comprising first to fourth audio sources (indicated collectively by reference 220) corresponding to different respective sounds labelled "1", "2", "3" and "4". With reference to Fig. 1(a), the respective perceived spatial positions of the first to fourth audio sources 220 are indicated relative to the user's head. It can be seen that a clockwise rotation of the user's head does not result in a modification of the spatial audio scene and the respective perceived spatial positions of the first to fourth audio sources 220 follow the user's movement. Figure 3 With reference to Fig. 1(b), the respective perceived spatial positions of the first to fourth audio sources 220 are indicated relative to the user's head with head tracking rendering. It can be seen that a clockwise rotation of the user's head results in a modification of the spatial audio scene and the respective perceived spatial positions of the first to fourth audio sources 220 do not follow the user's movement but remain stationary. Figure 4 With reference to Fig. 1(b), the respective perceived spatial positions of the first to fourth audio sources 220 are indicated relative to the user's head with head tracking rendering. It can be seen that a clockwise rotation of the user's head results in a modification of the spatial audio scene and the respective perceived spatial positions of the first to fourth audio sources 220 do not follow the user's movement but remain stationary.
[0063] Spatial audio data can be more complex to transmit, decode and render compared to other audio formats, and so a split rendering approach is proposed whereby the rendering operation can be split into different stages, with different stages being performed by different apparatus. The concept of split rendering can for example comprise a first or pre-rendering apparatus for receiving audio signals and creating a pre-rendered intermediate format of the spatial audio signals, which can be further rendered into a consumable format by a second apparatus. As will become clear, this can provide advantages in terms of lower complexity at the second apparatus, lower motion-to-sound latency at the second apparatus when using head tracking rendering, and can also provide a wider compatibility for different types of second apparatus, which can support the pre-rendered intermediate format without necessarily being able to support the audio signals as received by the pre-rendering apparatus.
[0064] Example embodiments relate to a split rendering system comprising first and second types of apparatus.
[0065] The first type of apparatus can be referred to as a pre-rendering apparatus. The second type of apparatus, which is separate from the first type of apparatus, can be referred to as a post-rendering device or simply a device.
[0066] The pre-rendering apparatus can be in communication with a plurality of post-rendering devices.
[0067] For example, the pre-rendering device can comprise a user device, a server, or a telecommunication room server. For example, the post-rendering devices can comprise a user device, a headphone set with head tracking capabilities, or one or more of a speaker system. Any combination of the described embodiments can be used. In this context, the user device can comprise one of a mobile phone, a tablet computer, a laptop computer, a personal computer, or a wearable device. The headphone set can comprise headphones mounted on or near a user’s ear, an earbud headphone that can be at least partially located within a user’s ear, or a speaker of an XR headset. The pre-rendering device can communicate with the plurality of post-rendering devices via wired or wireless channels. The wireless channels can comprise one or more of Bluetooth, Zigbee, or WiFi channels to give some non-limiting examples.
[0068] Figure 5 A split-rendering system 500 is shown that comprises a pre-rendering device 502 in communication with a plurality of post-rendering devices, in particular first, second, and third post-rendering devices 504, 506, 508. The first, second, and third post-rendering devices 504, 506, 508 can be associated with respective first, second, and third users 544, 546, 548 that can be at the same or different locations. The pre-rendering device 502 and the first, second, and third post-rendering devices 504, 506, 508 can comprise any combination of the examples described above.
[0069] The pre-rendering device 502 can receive an audio signal that can represent a spatial audio scene via an input line 510. The spatial audio scene can comprise a plurality of sound sources, such as those indicated in the above examples. The input line 510 can be connected to an antenna 511, for example, for receiving the audio signal from a remote source. Figures 2-4
[0070] The pre-rendering device 502 can be configured to receive respective requests for the spatial audio signal from the first, second, and third post-rendering devices 504, 506, 508, wherein the respective requests indicate respective user head orientations 514, 516, 518.
[0071] In the examples described herein, a user head orientation can refer to a head orientation that can comprise at least one of a yaw, a pitch, and / or a roll orientation.
[0072] For example, the first post-rendering device 504 can transmit a request for the spatial audio signal via a signal line 512 that indicates the first user head orientation 514. The second and third post-rendering devices 506, 508 can transmit their own respective requests that indicate the respective second and third user head orientations 516, 518.
[0073] The first, second, and third user head orientations 514, 516, 518 can represent respective reference or "0" head orientations. The first, second, and third user head orientations 514, 516, 518 may, for example, represent current head orientations of the respective first, second, and third users 544, 546, 548, or alternatively can comprise respective inverse directions, average head orientations over a time period, or default orientations associated with the first, second, and third post-rendering devices 504, 506, 508.
[0074] The pre-rendering apparatus 502 can process the audio signal to obtain first, second, and third spatial audio signals (alternatively referred to as intermediate spatial audio signals) associated with the first user head orientation 514, the second user head orientation 516, and the third user head orientation 518, respectively.
[0075] For example, the pre-rendering apparatus 502 can perform binaural rendering for each of the first user head orientation 514, the second user head orientation 516, and the third user head orientation 518 to obtain first, second, and third intermediate spatial audio signals. The pre-rendering apparatus 502 can also obtain first, second, and third metadata sets associated with the first, second, and third intermediate spatial audio signals, respectively. The first, second, and third metadata sets can comprise information indicative of how to locally modify the associated first, second, and third intermediate spatial audio signals at the respective first, second, and third post-rendering devices 504, 506, 508 to account for tracked user head positions, in this case changes in orientation, which can differ from the respective first, second, and third user head orientations 514, 516, 518.
[0076] The pre-rendering apparatus 502 can send the first intermediate spatial audio signal via signal line 524 and the first metadata set via signal line 525 to the first post- rendering device 504.
[0077] Similarly, the pre-rendering apparatus 502 can send the second intermediate spatial audio signal and the second metadata set to the second post-rendering device 506, and the third intermediate spatial audio signal and the third metadata set to the third post- rendering device 508.
[0078] As indicated by reference numerals 534, 536, 538, the binaural rendering for the first, second, and third intermediate spatial audio signals uses orientations corresponding to the respective first, second, and third user head orientations 514, 516, 518.
[0079] Thus, the first post-rendering device 504 can render the first intermediate spatial audio signal based on the first user head orientation 514, and the first metadata set can be used to locally modify the first intermediate spatial audio signal for other head orientations once they are tracked locally, e.g., when the user rotates his head away or towards the reference or “0” head orientation. In other words, the first post-rendering device 504 can use the metadata and tracked changes in the user head orientation to locally correct the binaural cues. The same process can be applied to the second and third spatial audio signals using the second and third metadata sets received by the second and third post-rendering devices 506, 508. This split approach is mentioned in the 3GPP TSG-SA WG4 Meeting #127Bis-e CR document, which relates to the IVAS standard specification TS 26.253.
[0080] Although the first, second, and third post-rendering devices 504, 506, 508 can perform less processing, because the pre-rendering apparatus 502 needs to obtain and send the intermediate spatial audio signal and the associated metadata set for each of the received first, second, and third user head orientations 514, 516, the amount of power and processing resources will linearly increase with the number of post-rendering devices. Also, as the number of post-rendering devices increases, there can be multiple post-rendering devices requesting or requiring substantially the same spatial audio signal and associated metadata because their respective user head orientations can be the same or substantially the same. Thus, the pre-rendering apparatus 502 can repeat at least some of the processing.
[0081] Example embodiments can avoid or mitigate such problems.
[0082] Figure 6 is a flowchart showing operations 600 according to one or more example embodiments. The operations 600 can be performed in hardware, software, firmware, or a combination thereof. For example, the operations 600 can be performed by the components individually or collectively, where the components can include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause performance of the operations. The operations 600 can be performed, for example, by a pre-processing apparatus.
[0083] A first operation 601 can include receiving an audio signal.
[0084] A second operation 602 can include receiving, from a plurality of, i.e., M, devices, respective requests for the spatial audio signal, where the respective requests indicate respective user head orientations.
[0085] The devices can include any of the above-described types of post-processing devices.
[0086] A third operation 603 can include determining a plurality of, i.e., N, reference head orientations, where N < M.
[0087] The third operation 603 can be performed before, during or after the execution of the first and second operations 602, 603.
[0088] The fourth operation 604 can comprise processing the received audio signal to obtain a plurality of spatial audio signals associated with a plurality of, i.e. N, reference head orientations, respectively, and a plurality of metadata sets associated with the plurality of spatial audio signals, respectively.
[0089] The term "obtaining" can relate to generating the plurality of spatial audio signals.
[0090] The metadata set can comprise information indicative of how the associated spatial audio signal should be adjusted to take into account a change in head position from the associated reference head orientation. The metadata set can further comprise an indication of the reference head orientation for which the associated spatial audio signal was obtained.
[0091] The fifth operation 605 can comprise sending the selected spatial audio signal and the associated metadata to the at least one device, wherein the selection is based on which reference head orientation is most closely associated with a user head orientation indicated in a request from the at least one device.
[0092] By determining a plurality of, i.e. N, reference head orientations, where N is a number smaller than the number M of devices from which respective requests for spatial audio signals are received, the amount of processing required by the pre-processing apparatus is reduced and the pre-processing apparatus can cater for a potentially large number of post-processing devices.
[0093] In some example embodiments, the term user head orientation can comprise something other than the current actual head orientation of the user, such as a predicted head orientation of the user or some other orientation of or related to the user. For example, one or more of the plurality of, i.e. M, devices can predict a transmission delay for which the respective request is valid, and instead request a predicted user head orientation for the time instance at which the respective device will receive the spatial audio signal.
[0094] In some example embodiments, another operation can comprise detecting that the plurality of, i.e. M, devices from which requests are received comprises a number larger than a threshold number G. The fourth and fifth operations 604, 605, and possibly the third operation 603, can be performed in response to the detection. If the plurality of, i.e. M, devices comprises a number equal to or smaller than the threshold number G, the described procedure can be performed, whereby M spatial audio signals associated with M user head orientations indicated in M respective requests from the plurality of, i.e. M, devices, respectively, are obtained. Figure 5 The described procedure, whereby M spatial audio signals associated with M user head orientations indicated in M respective requests from the plurality of, i.e. M, devices, respectively, are obtained.
[0095] In some example embodiments, N can comprise any number up to the threshold number G.
[0096] In some example embodiments, the plurality, N, of reference head orientations can be determined based at least in part on respective user head orientations indicated in respective requests. For example, the plurality, N, of reference head orientations can be determined in a grouping or clustering operation. Any suitable clustering algorithm can be used, for example k-means clustering, where the aim is to arrange the orientations into N groups while minimizing the variation within the groups. For example, the grouping or clustering operation can comprise arranging the plurality, M, of user head orientations into N groups of user head orientations, where at least one group comprises two or more user head orientations, determining a respective reference head orientation for the N groups.
[0097] A reference head orientation for the at least one group comprising two or more reference head orientations can be determined based on at least one of the two or more reference head orientations. For example, the reference head orientation for the at least one group can comprise one of the respective user head orientations. Alternatively, the reference head orientation for the at least one group can comprise an average of the respective user head orientations. A group can comprise two or more respective head orientations that are most similar or within a similarity threshold, for example within a predetermined angular range of each other.
[0098] In some example embodiments, and as will be explained in further detail below, the plurality, N, of reference head orientations can be predetermined, for example prior to receiving the respective requests.
[0099] The operation 600 will be understood with reference to the following non-limiting example embodiments Figure 6 of the operation 600.
[0100] Figure 7 A split-rendering system 700 is illustrated that comprises a pre-rendering device 702 in communication with first, second and third post-rendering devices 704, 706, 708. In this case, M = 3, but can be a larger number. A further post-rendering device 709 associated with a further user 749 is shown to indicate that the operations described herein can be extended to any number, M, of post-rendering devices.
[0101] The first, second and third post-rendering devices 704, 706, 708 can be associated with respective first, second and third users 744, 746, 748 that can be at the same or different locations.
[0102] The pre-rendering device 702 can receive an audio signal that can represent a spatial audio scene via an input line 710. The spatial audio scene can comprise a plurality of sound sources, such as those indicated in Figures 2-4 For example, the input line 710 can be connected to an antenna 711 for receiving the audio signal from a remote source.
[0103] The pre-rendering device 702 can be configured to receive respective requests for the spatial audio signal from the first, second, and third post-rendering devices 704, 706, 708, wherein the respective requests indicate respective first, second, and third user head orientations 714, 716, 718.
[0104] For example, the first user head orientation 714 can equal 280 degrees, the second user head orientation 716 can equal 310 degrees, and the third user head orientation 718 can equal 45 degrees. The head orientations 714, 716, 718 can refer to yaw orientations, and in other embodiments can refer to at least one of yaw, pitch, and / or roll orientations.
[0105] For example, the first post-rendering device 704 can send a request for the spatial audio signal via the signal line 712 that indicates the first user head orientation 714 of 280 degrees. The second and third post-rendering devices 706, 708 can transmit their own respective requests that indicate the respective second and third user head orientations 716, 718 of 320 and 45 degrees, respectively.
[0106] The first, second, and third user head orientations 714, 716, 718 can represent respective reference or “0” head orientations as described above with respect to Figure 5
[0107] The pre-rendering device 702 can responsively determine a plurality, i.e., N, of reference head orientations, where N < M.
[0108] For example, this can be performed in response to the pre-rendering device 702 detecting that M > G, which would be the case, for example, if G = 2, because M = 3.
[0109] The pre-rendering device 702 can determine a set of two reference orientations 750 that includes first and second reference head orientations 751, 752.
[0110] In this example, the first and second reference head orientations 751, 752 can be determined based on arranging the first and second user head orientations 714, 716 into a first group and the third user head orientation 718 into a second group. The arranging can be performed using any suitable clustering algorithm as described above, for example, based on the first user head orientation 714 and the second user head orientation 716 (which are 280 and 310 degrees, respectively) being closer to each other than the third user head orientation 718 (which is 45 degrees).
[0111] The first reference orientation 751 can include one or an average of the first user head orientation 714 and the second user head orientation 716.
[0112] Assuming the latter, the first reference orientation 751 can comprise an average of 280 and 310 degrees, which is 295 degrees.
[0113] The second reference orientation 752 can comprise the third user head orientation 718, which is 45 degrees.
[0114] The pre-rendering device 702 can process the audio signal to obtain first and second intermediate spatial audio signals respectively associated with the first and second reference orientations 751, 752 of 295 and 45 degrees (yaw orientation) respectively.
[0115] The pre-rendering device 702 can also obtain first and second sets of metadata respectively associated with the first and second intermediate spatial audio signals. The first and second sets of metadata can comprise an indication of the respective first and second reference orientations 751, 752 for which the first and second intermediate spatial audio signals were generated (as these can be different from the requested orientations) and metadata for correcting or compensating the spatial audio signal based on tracked changes in the user orientation relative to said respective reference orientations. In this way, less processing is required compared to the case where three intermediate spatial audio signals and three associated sets of metadata are obtained. Figure 5
[0116] The pre-rendering device 702 can then transmit, according to a fifth operation 605, a selected one of the first and second intermediate spatial audio signals and its associated metadata to the first, second and third post-rendering devices 704, 706, 708.
[0117] For example, the pre-rendering device 702 can select to transmit the first intermediate spatial audio signal and its associated metadata to the first and second post-rendering devices 704, 706. For example, the pre-rendering device 702 can transmit the first intermediate spatial audio signal via signal line 724 and the first set of metadata via signal line 744 to the first post-rendering device 704. This selection is performed based on the first and second user head orientations 714, 716 (280 and 310 degrees respectively) being most closely associated with the first reference orientation 751 of 295 degrees.
[0118] For example, the pre-rendering device 702 can select to transmit the second intermediate spatial audio signal and its associated metadata to the third post-rendering device 708. This selection is performed based on the third user head orientation 718 of 45 degrees being identical to the second reference orientation 752.
[0119] The first, second and third post-rendering devices 704, 706, 708 can then render the intermediate spatial audio signal they received to provide a rendered audio.
[0120] The first, second and third post-rendering devices 704, 706, 708 can correct the binaural rendering for tracked changes in the user orientation by using the received metadata sets in the same manner as described for Figure 5
[0121] Figure 8 A split rendering system 800 according to another embodiment is illustrated.
[0122] Figure 8 The system 800 is similar to the system 700 of Figure 7 and comprises a pre-rendering arrangement 802 in communication with at least first, second and third post-rendering devices 704, 706, 708. Another post-rendering device 709 associated with another user 749 is shown to indicate that the operations described herein can be extended to any number M of post-rendering devices.
[0123] The pre-rendering arrangement 802 in this example comprises a set of reference orientations 850 comprising a plurality (N=4) of reference head orientations 851, 852, 853, 854 that can be determined in advance of receiving the request.
[0124] For example, the first reference head orientation 851 can comprise 180 degrees, the second reference head orientation 852 can comprise 0 degrees, the third reference head orientation 853 can comprise 90 degrees, and the fourth reference head orientation 854 can comprise 270 degrees. The reference head orientations 851, 852, 853, 854 can refer to yaw orientations, and in other embodiments can refer to at least one of yaw, pitch and / or roll orientations.
[0125] The pre-rendering arrangement 802 can thus process the audio signals received on the input line 710 to obtain first, second, third and fourth intermediate spatial audio signals associated with the first, second, third and fourth reference head orientations 851, 852, 853, 854, respectively. The pre-rendering arrangement 802 can also obtain four associated metadata sets as previously described, wherein the metadata sets can comprise an indication of the respective first, second, third and fourth reference head orientations 851, 852, 853, 854 for which the first, second, third and fourth intermediate spatial audio signals were generated and metadata for correcting or compensating the spatial audio signals based on tracked changes in the user orientation relative to said respective reference orientations.
[0126] For example, the processing can be performed in response to the pre-rendering arrangement 802 detecting that M>G.
[0127] The number of, N, reference head orientations 851, 852, 853, 854 should be less than M, so in this example G can equal at least 5, and assume that requests have been received from six or more post-rendering devices, although only three post-rendering devices are considered for ease of explanation.
[0128] The processing can be performed before, during or after receiving the requests from the first, second and third post-rendering devices 704, 706, 708.
[0129] In response to receiving the first request from the first post-rendering device 704, the pre-rendering apparatus 802 can determine which of the first, second, third and fourth intermediate spatial audio signals to send in response.
[0130] For example, if the first request comprises a first user head orientation 714 of 280 degrees, the pre-rendering apparatus 802 can determine that this is most closely associated with the fourth reference head orientation 854 of 270 degrees. Accordingly, the pre-rendering apparatus 802 can select to send the fourth intermediate spatial audio signal to the first post-rendering device 704.
[0131] Similarly, in response to receiving the second request from the second post-rendering device 706, the pre-rendering apparatus 802 can determine that the second user head orientation 716 of 200 degrees is most closely associated with the first reference head orientation 851 of 180 degrees. Accordingly, the pre-rendering apparatus 802 can select to send the first intermediate spatial audio signal to the second post-rendering device 706. Similarly, in response to receiving the third request from the third post-rendering device 708, the pre-rendering apparatus 802 can determine that the third user head orientation 718 of 45 degrees is most closely associated with the second reference head orientation 852 of 0 degrees. Accordingly, the pre-rendering apparatus 802 can select to send the second intermediate spatial audio signal to the third post-rendering device 708. For example, the pre-rendering apparatus 802 can send the first intermediate spatial audio signal and the first set of metadata to the first post-rendering device 704 via signal line 724.
[0132] The first, second and third post-rendering devices 704, 706, 708 can then render the intermediate spatial audio signals they receive to provide rendered audio. The first, second and third post-rendering devices 704, 706, 708 can correct for changes in the tracked user direction by using the received sets of metadata in the same way as described above for the first, second and third post-rendering devices 704, 706, 708. Figure 5 The described same way as described above for the first, second and third post-rendering devices 704, 706, 708.
[0133] As previously described, since the number of intermediate spatial audio signals N and associated sets of metadata is less than the number of post-rendering devices M making the requests, less processing is required at the pre-rendering apparatus 802.
[0134] Figure 9is a flowchart illustrating operations 900 according to one or more example embodiments. The operations 900 can be performed in hardware, software, firmware, or a combination thereof. For example, the operations 900 can be performed by the components alone or collectively, where the components can include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause performance of the operations. The operations 900 can be performed, for example, by the pre-processing apparatus.
[0135] A first operation 901 can comprise receiving an audio signal.
[0136] A second operation 902 can comprise receiving, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests are indicative of respective user head orientations.
[0137] The apparatus can comprise any of the above-described types of post-processing apparatuses.
[0138] A third operation 603 can comprise determining whether M > G, wherein G is a threshold number.
[0139] If M > G, a fourth operation 904 can comprise determining a plurality, N, of reference head orientations, wherein N < M.
[0140] A fifth operation 905 can comprise processing the received audio signal to obtain a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations, and a plurality of metadata sets respectively associated with the plurality of spatial audio signals.
[0141] The term “obtaining” can relate to generating the plurality of spatial audio signals.
[0142] The metadata set can comprise information indicative of how the associated spatial audio signal should be adjusted to take into account a change in head orientation relative to the associated reference head orientation. The metadata set can further comprise an indication of the reference head orientation for which the associated spatial audio signal was obtained.
[0143] A sixth operation 906 can comprise sending the selected spatial audio signal and the associated metadata to at least one of the devices, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
[0144] If M ≤ G, a seventh operation 907 can comprise processing the received audio signal to obtain a plurality of spatial audio signals respectively associated with a plurality, M, of user head orientations, and a plurality of metadata sets respectively associated with the plurality of spatial audio signals.
[0145] The eighth operation 908 can comprise transmitting the spatial audio signal and the associated metadata associated with the user head orientation indicated in the request from the at least one device to at least one of the devices.
[0146] With respect to Figure 6 The described features can also apply to Figure 9 .
[0147] In some example embodiments, the term user head orientation can comprise something other than the current actual head orientation of the user, such as a predicted head orientation of the user or some other orientation of or related to the user. For example, one or more of the plurality, i.e. M, devices can predict a transmission delay for which its respective request is valid, and instead request a predicted user head orientation for the respective device for a time instance at which the spatial audio signal will be received.
[0148] In some example embodiments, the described certain operations, e.g. whenever a new pre-rendering operation is to be performed based on the received audio signal, can be repeated periodically. Figures 6 to 9 However, in order to allow for processing time of the described operations, in particular the grouping or clustering operation associated with determining the plurality, i.e. N, reference head orientations, the grouping or clustering can be performed at a larger interval, e.g. every 500 ms, or even longer periods if no significant change of the indicated user head orientation is observed. A different schedule can be used to update the determination of the plurality, i.e. N, reference head orientations, respectively, which can be shorter than the schedule for the grouping or clustering, and possibly once per 20 ms frame.
[0149] An orientation can be defined as a rotation from a reference orientation to another orientation. Thus, an orientation can be represented similar to a rotation, with examples being Tait-Bryan angles (yaw, pitch, roll), a rotation matrix, a quaternion, Euler angles, or a direction cosine matrix. These representations can be converted from one to another using known methods. Furthermore, the above examples assume a single orientation axis, such as one of yaw, pitch, and roll. However, when considering head orientations for binaural rendering, there are limitations in how a user can move their head. In normal use, most of the orientation changes are in the yaw and pitch axes, with yaw being the most important one. Although the above examples can be extended to use a reference orientation corresponding to each of the yaw, pitch, and roll axes, a simplified implementation from a computational point of view can consider only one or two axes, such as only the yaw axis or the yaw and pitch axes, with the roll axis being set to zero.
[0150] With respect to grouping or clustering the user head orientations into two or more groups to determine a respective reference head orientation, any suitable clustering algorithm can be used. As noted above, this can involve k-means clustering, where the goal is to arrange the respective user head orientations into N groups while minimizing the variation within the groups. The number of groups N can be any number up to a threshold number G, which can be performed by considering the corresponding vectors in the yaw axis alone or in the yaw and pitch axes. In some example embodiments, if the number M of post-processing devices from which the respective requests are received is greater than G but below a second threshold number G2, a simple grouping or clustering algorithm can include finding pairs of user head orientations, e.g., orientations having a difference below a predetermined threshold angle Al. Such pairs of user head orientations can be arranged into a common group, and the reference head orientation for that group can include one of the user head orientations or an average of the pair of user head orientations. However, if the number M of post-processing devices from which the respective requests are received is greater than G and equal to or greater than the second threshold number G2, an alternative grouping or clustering method can be used. For example, the method can include (i) determining directional vectors associated with the respective user head orientations received from the plurality of M devices, (ii) identifying which of the first set of spatial partitions corresponds to the largest number of directional vectors, (iii) dividing the identified spatial partition into two or more spatial partitions, and (iv) if the number of spatial partitions is not equal to N, re-executing the identifying and dividing operations until the number of spatial partitions is equal to N. Thus, the plurality of M user head orientations are arranged into N groups, and a reference head orientation can be determined for each of the N groups.
[0151] Figure 10A An example case is shown where M = 6 and N is set to 5.
[0152] Six vectors 1001-1006 corresponding to the M respective user head orientations mapped to the unit circle 1000 can be determined. In the case of considering two or more axes, e.g., the yaw and pitch axes, the vectors 1001-1006 can be mapped to a unit sphere. The unit circle 1000 includes a first set of four spatial partitions R1-R4, which in this example correspond to the quadrants of the unit circle. The first spatial partition R1 can be identified as including the largest number of directional vectors, i.e., the first, second, and third directional vectors 1001, 1002, 1003. If two or more spatial partitions include the same maximum number of directional vectors, a predetermined rule can determine which is divided, or alternatively, if no more than N, each spatial partition can be divided. In this case, the first spatial partition can be divided into two smaller spatial partitions, e.g., R1A and R1B, as shown in FIG. 10B. The second spatial partition R2 can be identified as including the largest number of directional vectors, i.e., the fourth and fifth directional vectors 1004, 1005. The second spatial partition R2 can be divided into two smaller spatial partitions, e.g., R2A and R2B, as shown in FIG. 10B. The third spatial partition R3 can be identified as including the largest number of directional vectors, i.e., the sixth directional vector 1006. The third spatial partition R3 can be divided into two smaller spatial partitions, e.g., R3A and R3B, as shown in FIG. 10B. The reference head orientation for each of the five groups can be determined, e.g., as the average of the user head orientations in the group. Figure 10B, where the first spatial partition R1 is replaced by two smaller spatial partitions R1A and R1B. The total number of spatial partitions is now N=5, so the grouping or clustering process can be stopped. The above-described method can be used to determine a reference user head position for each of the five spatial partitions R1A, R1B, R2, R3, and R4. For example, the first reference head position of the first spatial partition R1A may include the average of the first and second user head orientations represented by the first and second vectors 1001 and 1002. For example, the second to fifth reference head positions may include the third to sixth user head orientations represented by the third to sixth vectors 1001-1006.
[0153] Another method for grouping or clustering can involve applying angle quantization in a manner that increases the level of quantization error until at most G different orientations remain. For example, we can first quantize the orientations with 16 bits, and if that is not enough, we can quantize with 15 bits, 14 bits, etc. When at most G different orientations are left, the original orientations can be assigned to groups based on the resulting quantized orientations.
[0154] In some example embodiments, other implementations may be used.
[0155] For example, even if the number M of devices from which requests are received is not greater than G, or when G is small, e.g., less than 10, all pairs of user head orientations can be compared, and if any pairs are within an inaudible tolerance, e.g., within 5 degrees of each other, the pairs can be grouped, and one of them can be selected for the reference head orientation. This can be done by comparing reference head orientation pairs to determine if any pairs are within an inaudible tolerance, and if so, combining them, and this can be easily extended to larger groups. In this way, power and computing resources are used efficiently without sacrificing perceptual quality.
[0156] For example, according to Figure 8 For example, in situations where audio signals are provided along with video content and / or all audio signals are in relatively narrow directions, using predefined reference orientations may be useful. For example, based on other information, assigning a predetermined reference orientation to the relatively narrow directions by default and assigning at least one other predetermined reference orientation to unintended orientations may be more practical and effective.
[0157] The example embodiments have been described with respect to binaural rendering, but are not limited to such an output format and may be used with other output formats.
[0158] For the reasons stated above, example embodiments enable improved power and computing performance.
[0159] Example device
[0160] Figure 11 An example apparatus 1100 capable of supporting at least some embodiments is shown. A device 1100 is shown, which can be the pre-rendering apparatus 702, 802 described above. Included in the device 1100 is a processor 1110, which can include, for example, a single-core or multiple-core processor, where a single-core processor includes one processing core and a multiple-core processor includes more than one processing core. The processor 1110 can generally include control of the device. The processor 1110 can include more than one processor. The processor 1110 can be control of the device. The processing core can include, for example, a Cortex-A8 processing core manufactured by ARM Holdings or a Steamroller processing core manufactured by Advanced Micro Devices, Inc. The processor 1110 can include at least one Qualcomm Snapdragon and / or Intel Atom processor. The processor 1110 can include at least one application-specific integrated circuit (ASIC). The processor 1110 can include at least one field-programmable gate array (FPGA). The processor 1110 can be a means for performing method steps in the device 1100. The processor 1110 can be configured, at least in part, by computer instructions to perform actions.
[0161] The processor can include circuitry, or be structured as one or more circuits, configured to perform stages of the method of the application in accordance with embodiments described herein. As used in this application, the term “circuitry” can refer to one or more or all of: (a) solely hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of hardware circuits and software, such as, as applicable: (i) combinations of analog and / or digital hardware circuit(s) with software / firmware
[0162] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers implementations including, but not limited to, only hardware circuit implementations or processor(s) or processor core(s) or portions thereof, and their software and / or firmware implementations (with such software and / or firmware implementations comprising portions of the software and / or firmware that are necessary in order to render the device or apparatus capable of implementing functionalities described herein). The term circuitry also covers, for example and if applicable to particular claim elements, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in a server, cellular network device, or other computing or network device.
[0163] Device 1100 can include memory 1120. Memory 1120 can include random access memory and / or permanent memory. Memory 1120 can include at least one RAM chip. Memory 1120 can include, for example, solid state, magnetic, optical, and / or holographic memory. Memory 1120 can be at least partially accessible by processor 1110. Memory 1120 can be at least partially included in processor 1110. Memory 1120 can be a component for storing information. Memory 1120 can include computer instructions that processor 1110 is configured to execute. When computer instructions configured to cause processor 1110 to perform certain actions are stored in memory 1120, and device 1100 is generally configured to run using the computer instructions from memory 1120 under the direction of processor 1110, processor 1110 and / or at least one processing core thereof can be said to be configured to perform the certain actions. Memory 1120 can be at least partially included in processor 1110. Memory 1120 can be at least partially external to device 1100, but accessible by device 1100.
[0164] Device 1100 can include transmitter 1130. Device 1100 can include receiver 1140. Transmitter 1130 and receiver 1140 can be configured to transmit and receive information, respectively, according to at least one cellular or non-cellular standard.
[0165] Transmitter 1130 can include more than one transmitter. Receiver 1140 can include more than one receiver. Transmitter 1130 and / or receiver 1140 can be configured to operate according to, for example, the following standards: Global System for Mobile Communications, GSM, Wideband Code Division Multiple Access, WCDMA, 5G / NR, 5G Advanced, i.e., NR Rel-18, 19, and beyond, Long Term Evolution, LTE, IS-95, Wireless Local Area Network, WLAN, Ethernet, and / or Worldwide Interoperability for Microwave Access, WiMAX.
[0166] Device 1100 can include near field communication, NFC, transceiver 1150. NFC transceiver 1150 can support at least one NFC technology, such as NFC, Bluetooth, Wibree, or similar technologies.
[0167] The device 1100 can comprise a user interface, UI, 1160. The UI 1160 can comprise at least one of a display, a keyboard, a touch screen, a vibrator arranged to signal to a user by causing the device 1100 to vibrate, a loudspeaker, and a microphone. A user is able to operate the device 1100 via the UI 1160, e.g. to accept an incoming telephone call, to initiate a telephone call or a video call, to browse the internet, to manage digital files stored in the memory 1120 or on a cloud accessible via the transmitter 1130 and the receiver 1140 or via the NFC transceiver 1150, and / or to play a game.
[0168] The device 1100 can comprise or be arranged to accept a user identity module 1170. The user identity module 1170 can comprise, e.g., a subscriber identity module, SIM, card installable in the device 1100. The user identity module 1170 can comprise information identifying a subscription of a user of the device 1100. The user identity module 1170 can comprise cryptographic information usable to verify an identity of a user of the device 1100 and / or to facilitate encryption of information communicated and billing of a user of the device 1100 for communications implemented via the device 1100.
[0169] The processor 1110 can be equipped with a transmitter arranged to output information from the processor 1110 to other devices comprised in the device 1100 via electrical leads inside the device 1100. Such a transmitter can comprise a serial bus transmitter arranged to output information, e.g., to the memory 1120 for storage therein, via at least one electrical lead. As an alternative to a serial bus, the transmitter can comprise a parallel bus transmitter.
[0170] Likewise, the processor 1110 can comprise a receiver arranged to receive information in the processor 1110 from other devices comprised in the device 1100 over electrical leads inside the device 1100. Such a receiver can comprise a serial bus receiver arranged to receive information, e.g., from the receiver 1140 for processing in the processor 1110, via at least one electrical lead. As an alternative to a serial bus, the receiver can comprise a parallel bus receiver.
[0171] The device 1100 can comprise Figure 11The other devices not shown can be included in device 1100. For example, where device 1100 comprises a smartphone, it can include at least one digital camera. Some devices 1100 can include a back-facing camera and a front-facing camera, where the back-facing camera can be intended for digital photography and the front-facing camera can be intended for video telephony. Device 1100 can include a fingerprint sensor arranged to at least partially authenticate a user of device 1100. In some embodiments, device 1100 lacks at least one of the devices described above. For example, some devices 1100 can lack NFC transceiver 1150 and / or user identity module 1170.
[0172] Processor 1110, memory 1120, transmitter 1130, receiver 1140, NFC transceiver 1150, UI 1160, and / or user identity module 1170 can be interconnected by electrical leads inside device 1100 in a variety of different ways. For example, each of the devices described above can be individually connected to a main bus inside device 1100 to allow the devices to exchange information. However, as will be appreciated by those skilled in the art, this is merely one example, and various ways of interconnecting at least two of the devices described above can be selected according to an embodiment without departing from the scope of the invention.
[0173] Figure 12 A non-transitory medium 1200 is shown in accordance with some embodiments. Non-transitory medium 1200 is a computer-readable storage medium. It can be, for example, a CD, a DVD, a USB stick, a Blu-ray disc, etc. Non-transitory medium 1200 stores computer program instructions such that an apparatus performs a method of, for example, any of the foregoing processes disclosed in relation to the flowcharts in this specification and their related features.
[0174] The described features, structures, or characteristics can be combined in one or more embodiments. In the preceding description, numerous specific details are provided, such as examples of lengths, widths, shapes, etc., to provide a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the application.
[0175] While the forgoing examples are illustrative of the principles of the application in one or more particular applications, it will be apparent to those of ordinary skill in the art that numerous modifications in form, usage and details of implementation can be made without the departure from the principles and concepts underlying the present application. Accordingly, no limitation is placed on the scope of the application by the recitation of the claims that follow and / or the recitation of the summary of the application.
[0176] The verbs "comprise" and "comprising" are used as open-ended limitations that neither exclude nor require the existence of also un-recited features. The features recited in dependent claims are mutually freely combinable unless otherwise explicitly stated. Furthermore, it is to be understood that the use of "a" or "an", i.e. a singular form, throughout the document does not exclude a plurality.
Claims
1. An apparatus comprising: means for receiving an audio signal; means for receiving, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests are indicative of respective user head orientations; means for determining a plurality, N, of reference head orientations, wherein N < M; means for processing the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of sets of metadata respectively associated with the plurality of spatial audio signals, wherein a set of metadata comprises information indicative of how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and means for sending, to at least one of the devices, a selected spatial audio signal and associated metadata, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
2. The apparatus of claim 1, further comprising: means for detecting that the plurality, M, of devices comprises a number greater than a threshold number, G, wherein the processing is performed in response to the detection.
3. The apparatus of claim 2, wherein: the plurality, N, of reference head orientations comprises a number equal to the threshold number, G.
4. The apparatus of any of claims 1 to 3, wherein: the plurality, N, of reference head orientations is determined prior to receiving the respective requests.
5. The apparatus of any of claims 1 to 3, wherein: the plurality, N, of reference head orientations is determined based at least in part on the respective user head orientations indicated in the respective requests.
6. The apparatus of claim 5, wherein: the plurality, N, of reference head orientations is determined by: arranging the plurality, M, of user head orientations into N groups of one or more user head orientations, wherein at least one group comprises two or more user head orientations; and determining, for the N groups, respective reference head orientations, wherein the reference head orientation for the at least one group comprising two or more reference head orientations is determined based on at least one of the two or more reference head orientations.
7. The apparatus of claim 6, wherein: the reference head orientation for the at least one group comprises one of the respective user head orientations.
8. The apparatus of claim 6, wherein: the reference head orientation for the at least one group comprises an average of the respective user head orientations.
9. The apparatus of any of claims 6 to 8, wherein: the at least one group comprises two or more respective head orientations that are most similar or within a similarity threshold.
10. The apparatus of any of claims 6 to 9, wherein: the arranging comprises: determining direction vectors associated with the respective user head orientations received from the plurality, M, of devices; identifying which spatial partition of the first set of spatial partitions corresponds to a maximum number of direction vectors; dividing the identified spatial partition into two or more spatial partitions; and re-performing the identifying and dividing operations if the number of spatial partitions is not equal to N until the number of spatial partitions is equal to N.
11. The apparatus according to any of claims 6 to 10, wherein, the selected spatial audio signal for a particular device is a spatial audio signal associated with the reference head orientation of the group into which the respective user head orientation from the particular device is arranged.
12. The device according to any of the preceding claims, wherein, the reference head orientation comprises an orientation for at least one of: a yaw axis; a yaw axis and a pitch axis; or a yaw axis, a pitch axis, and a roll axis.
13. The apparatus according to any of the preceding claims, wherein, the apparatus is comprised by at least one of: a user device, or a server.
14. The apparatus according to any of the preceding claims, wherein, the plurality, M, of devices is comprised by at least one of: a headphone device, a loudspeaker device, or a user device.
15. A method comprising: receiving an audio signal; receiving, from a plurality, M, of devices, respective requests for a spatial audio signal, wherein the respective requests indicate respective user head orientations; determining a plurality, N, of reference head orientations, wherein N < M; processing the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of sets of metadata respectively associated with the plurality of spatial audio signals, wherein a set of metadata comprises information indicating how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and sending, to at least one of the devices, a selected spatial audio signal and an associated set of metadata, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.
16. The method according to claim 15, further comprising: detecting that the plurality, M, of devices comprises a number greater than a threshold number, G, wherein the processing is performed in response to the detection.
17. The method according to claim 16, wherein, the plurality, N, of reference head orientations comprises a number equal to the threshold number, G.
18. The method according to any of claims 15 to 17, wherein, the plurality, N, of reference head orientations is determined prior to receiving the respective requests.
19. The method according to any of claims 15 to 17, wherein, the plurality, N, of reference head orientations is determined based at least in part on the respective user head orientations indicated in the respective requests.
20. The method according to claim 19, wherein, the plurality, N, of reference head orientations is determined by: arranging the plurality, M, of user head orientations into N groups of one or more user head orientations, wherein at least one group comprises two or more user head orientations; and determining, for the N groups, a respective reference head orientation, wherein the reference head orientation for the at least one group comprising two or more reference head orientations is determined based on at least one of the two or more reference head orientations.
21. The method of claim 20, wherein, the reference head orientation for the at least one group comprises one of the respective user head orientations.
22. The method of claim 20, wherein, the reference head orientation for the at least one group comprises an average of the respective user head orientations.
23. The method of any one of claims 20 to 22, wherein, the at least one group comprises two or more respective head orientations that are most similar or within a similarity threshold.
24. The method of any one of claims 20 to 23, wherein: the arranging comprises: determining direction vectors associated with the respective user head orientations received from the plurality, M, of devices; identifying which spatial partition in a first set of spatial partitions corresponds to a maximum number of direction vectors; dividing the identified spatial partition into two or more spatial partitions; and re-executing the identifying and dividing operations if the number of spatial partitions is not equal to N until the number of spatial partitions is equal to N.
25. A non-transitory computer-readable medium comprising program instructions stored thereon for performing a method, the method comprising: receiving an audio signal; receiving, from a plurality, M, of devices, a respective request for a spatial audio signal, wherein the respective request indicates a respective user head orientation; determining a plurality, N, of reference head orientations, wherein N < M; processing the received audio signal to obtain: a plurality of spatial audio signals respectively associated with the plurality, N, of reference head orientations; and a plurality of metadata sets respectively associated with the plurality of spatial audio signals, wherein a metadata set comprises information indicating how an associated spatial audio signal should be adjusted to account for a change in head orientation relative to an associated reference head orientation; and sending the selected spatial audio signal and associated metadata to at least one of the devices, wherein the selection is based on which reference head orientation is most closely associated with the user head orientation indicated in the request from the at least one device.