Sound field related rendering
Patent Information
- Application Number
- JP2024006056
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-11
- Filing Date
- 2024-01-18
- Publication Date
- 2026-02-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing spatial audio technologies lack the ability to control the focus shape in addition to focus direction and amount, which is essential for users to selectively emphasize or de-emphasize audio sources based on their preferences and viewing directions in media with multiple viewing angles.
The implementation of a system that processes spatial audio signals to define and control a focus shape, adjusting the emphasis of audio within and outside this shape through parameters such as focus direction, width, height, radius, distance, and depth, using ambisonic and parametric spatial audio formats to enhance user control over the audio experience.
Enables users to dynamically adjust the audio focus shape, allowing for personalized audio emphasis based on their preferences, improving the spatial audio experience by selectively highlighting or attenuating audio sources within a defined area.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus and method for sound field related audio representation and rendering, but is not limited to audio representation for audio decoders. [Background technology]
[0002] Spatial audio playback is known for presenting media with multiple viewing directions. Examples of this playback include (at least) a head-mounted display (or a head-mounted phone) that can track the head orientation, or a non-head-mounted phone screen that can track the view direction by changing the phone position / orientation, or with any user interface gesture, or on a surrounding screen.
[0003] Footage related to "media with multiple viewing directions" may include, for example, 360-degree footage, 180-degree footage, or other footage that has a substantially wider viewing angle than traditional footage. Traditional footage is video content that is typically presented in its entirety on a screen, with no option (or particular need) to change the viewing direction.
[0004] Audio associated with video with multiple viewing directions can be presented over headphones or a surround loudspeaker setup where the viewing direction is tracked and affects the spatial audio reproduction.
[0005] The spatial audio associated with video with multiple viewing directions can come from spatial audio capture from a microphone array (e.g., an array attached to a VR camera like an OZO, or a handheld mobile device) or from other sources such as a studio mix. It is also possible for the audio content to be a mixture of multiple content types, such as microphone-captured audio and an added commentator track.
[0006] Spatial audio associated with video with multiple viewing directions can take various forms, for example: Ambisonic signals (any order) consisting of spherical harmonic audio signal components. The spherical harmonics can be considered as a set of spatially selective beam signals. Currently, Ambisonics are used, for example, in the YouTube® 360VR video service. The advantage of Ambisonics is that it is a simple and well-defined signal representation. Surround speaker signals (e.g. 5.1). Currently, typical cinematic spatial audio is conveyed in this format. The advantage of surround loudspeaker signals is that it is simple and legacy compatible. Some audio formats similar to the format of surround loudspeaker signals contain audio objects that can be considered as audio channels with time-varying positions. The position can signal both the direction and distance or orientation of the audio object. Some state-of-the-art audio coding and spatial audio capture methods apply such signal representations, such as parametric spatial audio, i.e. audio signals of two audio channels in perceptually relevant frequency bands and associated spatial metadata. Spatial metadata essentially determines how the audio signal should be spatially reproduced at the receiver (e.g. in which directions at different frequencies). The advantages of parametric spatial audio are its versatility, quality, and the ability to use lower bitrates for encoding. Summary of the Invention
[0007] According to a first aspect, there is provided an apparatus comprising means configured to obtain at least one focus parameter configured to define a focus shape, to process a spatial audio signal representative of an audio scene to control a relative emphasis of at least some of the parts of the spatial audio signal within the focus shape with respect to at least some of the other parts of the spatial audio signal outside the focus shape, to generate a processed spatial audio signal representative of a modified audio scene, and to output the processed spatial audio signal, wherein the modified audio scene enables a relative emphasis of at least some of the parts of the spatial audio signal within the focus shape with respect to at least some of the other parts of the spatial audio signal outside the focus shape.
[0008] The at least one focus parameter may be further configured to define a focus amount, and the means configured to process the spatial audio signal may be configured to process the spatial audio signal in accordance with the focus amount and further to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0009] The means configured to process the spatial audio signal may be configured to increase a relative emphasis, or decrease a relative emphasis, of at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.
[0010] The means configured to process the spatial audio signal may be configured to increase or decrease a relative sound level of at least a part of the spatial audio signal within the focus shape relative to at least a part of other parts of the spatial audio signal outside the focus shape.
[0011] The means configured to process the spatial audio signal may be configured to increase or decrease a relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some other portions of the spatial audio signal outside the focus shape according to the focus amount.
[0012] The means may be configured to obtain playback control information for controlling at least one aspect of outputting the processed spatial audio signal, and the means configured to output the processed spatial audio signal may be configured to perform one of: processing the processed spatial audio signal representing a modified audio scene to generate an output spatial audio signal in accordance with the playback control information; and, prior to the means configured to process the spatial audio signal representing the audio scene, processing the spatial audio signal in accordance with the playback control information to generate a processed spatial audio signal representing the modified audio scene and outputting the processed spatial audio signal as the output spatial audio signal.
[0013] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and the means configured to process the spatial audio signal to generate the processed spatial audio signal may be configured to: convert, for one or more frequency subbands, the Ambisonic signal associated with the spatial audio signal into a set of beam signals of a defined pattern; generate a set of modified beam signals based on the set of beam signals, a focus shape and a focus amount; and convert the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.
[0014] The defined pattern may consist of a defined number of beams equally spaced over a plane or over a volume.
[0015] The spatial audio signal and the processed spatial audio signal may be constructed from respective higher order Ambisonic signals.
[0016] The spatial audio signal and the processed spatial audio signal may be composed of a subset of Ambisonic signal components of any order.
[0017] The spatial audio signal and the processed spatial audio signal may comprise respective parametric spatial audio signals, which may comprise one or more audio channels and spatial metadata, which may comprise respective direction indications, energy ratio parameters, and potentially distance indications for a plurality of frequency subbands. The means configured to process the input spatial audio signal to generate a processed spatial audio signal may be configured to: calculate spectral adjustment coefficients for one or more frequency subbands based on the spatial metadata and the focus shape and focus amount; apply the spectral adjustment coefficients to one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels; calculate respective modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, focus amount, and spatial metadata; and configure a processed spatial audio signal consisting of the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.
[0018] The spatial audio signal and the processed spatial audio signal may include multi-channel loudspeaker channels and / or audio object channels. The means configured for processing the spatial audio signal into the processed spatial audio signal may be configured to calculate gain adjustment factors based on the respective audio channel direction indications, focus shapes and focus amounts, apply the gain adjustment factors to the respective audio channels and produce a processed spatial audio signal including one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels.
[0019] The multi-channel speaker channels and / or the audio object channels may further include respective audio channel distance indications, and the calculated gain adjustment factor may be further based on the audio channel distance indications.
[0020] The means may be further configured to determine a default respective audio channel distance, and the computing gain adjustment factor may be further configured based on the audio channel distance.
[0021] The at least one focus parameter configured to define the focus shape can include at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, and a focus shape characterizer.
[0022] The means may be further configured to obtain a focus input from a sensor arrangement comprising at least one directional sensor and at least one user input, the focus input may further include an indication of a focus direction of the focus shape based on a direction of the at least one directional sensor and an indication of a focus width based on the at least one user input, the focus input may further include an indication of a focus amount based on the at least one user input.
[0023] According to a second aspect, there is provided a method comprising the steps of obtaining at least one focus parameter configured to define a focus shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; and outputting the processed spatial audio signal, wherein the modified audio scene enables relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0024] The at least one focus parameter may be further configured to define a focus amount, and processing the spatial audio signal may include processing the spatial audio signal in accordance with the focus amount and to further control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0025] Processing the spatial audio signal may include increasing a relative emphasis or decreasing a relative emphasis in at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.
[0026] Processing the spatial audio signal may include increasing or decreasing a relative sound level in at least some of the portions of the spatial audio signal in the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.
[0027] Processing the spatial audio signal may include increasing or decreasing a relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape according to the focus amount.
[0028] The method may include obtaining playback control information for controlling at least one aspect of outputting the processed spatial audio signal, where outputting the processed spatial audio signal may include performing one of the following steps: processing the processed spatial audio signal representing the modified audio scene to generate an output spatial audio signal in accordance with the playback control information; and processing the spatial audio signal in accordance with the playback control information to generate a processed spatial audio signal representing the modified audio scene and outputting the processed spatial audio signal as the output spatial audio signal, prior to the means configured to process the spatial audio signal representing the audio scene.
[0029] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and processing the spatial audio signal to generate the processed spatial audio signal may include converting, for one or more frequency subbands, an Ambisonic signal associated with the spatial audio signal into a set of beam signals of a defined pattern; generating a set of modified beam signals based on the set of beam signals, a focus shape, and a focus amount; and converting the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.
[0030] The defined pattern may consist of a defined number of beams equally spaced over a plane or over a volume.
[0031] The spatial audio signal and the processed spatial audio signal may be constructed from respective higher order Ambisonic signals.
[0032] The spatial audio signal and the processed spatial audio signal may be composed of a subset of Ambisonic signal components of any order.
[0033] The spatial audio signal and the processed spatial audio signal may include respective parametric spatial audio signals, which may include one or more audio channels and spatial metadata, which may include respective direction indications, energy ratio parameters, and potentially distance indications for a plurality of frequency subbands. Processing the input spatial audio signal to generate the processed spatial audio signal may include calculating spectral adjustment coefficients for one or more frequency subbands based on the spatial metadata and a focus shape and a focus amount, and applying the spectral adjustment coefficients to one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels, calculating respective modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, focus amount, and spatial metadata, and configuring the processed spatial audio signal to include the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.
[0034] The spatial audio signal and the processed spatial audio signal may include multi-channel loudspeaker channels and / or audio object channels, and processing the spatial audio signal into the processed spatial audio signal may include calculating gain adjustment factors based on respective audio channel direction indications, focus shapes, and focus amounts, applying the gain adjustment factors to the respective audio channels, and constituting the processed spatial audio signal including one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels.
[0035] The multi-channel speaker channels and / or the audio object channels may further include respective audio channel distance indications, and the computing gain adjustment factor may further be performed based on the audio channel distance indications.
[0036] The method further includes determining a default respective audio channel distance, and the computing gain adjustment factor can be further determined based on the audio channel distance. The at least one focus parameter configured to define the focus shape can include at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, and a focus shape characterizer.
[0037] The method may further include obtaining focus input from a sensor arrangement comprising at least one directional sensor and at least one user input, the focus input may include an indication of a focus direction of the focus shape based on a direction of the at least one directional sensor, and an indication of a focus width based on the at least one user input.
[0038] The focus input may further include an indication of an amount of focus based on at least one user input.
[0039] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured, using the at least one processor, to cause the apparatus to at least perform the steps of obtaining at least one focus parameter configured to define a focus shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene to control a relative emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; and outputting the processed spatial audio signal, the modified audio scene enabling a relative emphasis in at least some of the portions of the spatial audio signal within the at least some of the focus shapes compared to at least some of the other portions of the spatial audio signal outside the focus shape.
[0040] The at least one focus parameter may be further configured to define a focus amount, and the device adapted to process the spatial audio signal may be adapted to process the spatial audio signal in accordance with the focus amount and further to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape. The device adapted to process the spatial audio signal may be adapted to increase the relative emphasis or decrease the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.
[0041] An apparatus adapted to process the spatial audio signal may be adapted to increase or decrease the relative sound level in at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.
[0042] An apparatus adapted to process the spatial audio signal may be adapted to increase or decrease the relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape according to the focus amount.
[0043] The apparatus may be adapted to obtain playback control information for controlling at least one aspect of outputting the processed spatial audio signal, and the apparatus adapted to output the processed spatial audio signal may be adapted to perform one of the following steps: processing the processed spatial audio signal representing the modified audio scene to generate an output spatial audio signal in accordance with the playback control information, processing the spatial audio signal in accordance with the playback control information to generate a processed spatial audio signal representing the modified audio scene, prior to the means configured to process the spatial audio signal representing the audio scene, and outputting the processed spatial audio signal as the output spatial audio signal.
[0044] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and the device for processing the spatial audio signal to generate the processed spatial audio signal may be configured to convert, for one or more frequency subbands, an Ambisonic signal associated with the spatial audio signal into a set of beam signals of a defined pattern, generate a set of modified beam signals based on the set of beam signals, a focus shape, and a focus amount, and convert the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.
[0045] The defined pattern may consist of a defined number of beams equally spaced over a plane or over a volume.
[0046] The spatial audio signal and the processed spatial audio signal may be constructed from respective higher order Ambisonic signals.
[0047] The spatial audio signal and the processed spatial audio signal may be composed of a subset of Ambisonic signal components of any order.
[0048] The spatial audio signal and the processed spatial audio signal may comprise respective parametric spatial audio signals, the parametric spatial audio signals may comprise one or more audio channels and spatial metadata, the spatial metadata may comprise respective directional indications, energy ratio parameters and potentially distance indications for a plurality of frequency subbands, and an apparatus adapted to process an input spatial audio signal to generate a processed spatial audio signal may further comprise: 1) the spatial audio signal may comprise respective directional indications for a portion of a plurality of frequency bands of a plurality of frequency bands of a plurality of frequency bands; 2) the spatial audio signal may comprise a plurality of directional indications for a portion of a plurality of frequency bands of a plurality of frequency bands; 3) the spatial metadata may comprise a portion of a plurality of frequency bands ... The method may further include the steps of: calculating spectral adjustment coefficients for one or more frequency subbands based on spatial metadata, which may include respective directional indications for some frequency bands, and the focus shape and the focus amount; applying the spectral adjustment coefficients to one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels; calculating respective modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, the focus amount, and the spatial metadata; and constructing a processed spatial audio signal consisting of the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.
[0049] The spatial audio signal and the processed spatial audio signal may include multi-channel loudspeaker channels and / or audio object channels, and the device for processing the spatial audio signal into a processed spatial audio signal may perform the steps of calculating gain adjustment factors based on respective audio channel directional indications, focus shapes and focus amounts, applying the gain adjustment factors to the respective audio channels, and constituting a processed spatial audio signal including one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels.
[0050] The multi-channel speaker channel and / or the audio object channel may further include a respective audio channel distance indication, and the computing gain adjustment factor may be further determined based on the audio channel distance indication. The device may be further caused to determine a default respective audio channel distance, and the computing gain adjustment factor may be further determined based on the audio channel distance. The at least one focus parameter configured to define the focus shape may include at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, and a focus shape characterizer.
[0051] The apparatus may further be caused to obtain focus input from a sensor arrangement comprising at least one directional sensor and at least one user input, the focus input including an indication of a focus direction of the focus shape based on a direction of the at least one directional sensor, and an indication of a focus width based on the at least one user input.
[0052] The focus input may further include an indication of an amount of focus based on at least one user input.
[0053] According to a fourth aspect, there is provided an apparatus comprising: a focus parameter acquisition circuit configured to acquire at least one focus parameter configured to define a focus shape; a spatial audio signal processing circuit configured to process a spatial audio signal representing an audio scene to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape to generate a processed spatial audio signal representing a modified audio scene; and an output control circuit configured to output the processed spatial audio signal, wherein the modified audio scene enables relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0054] According to a fifth aspect, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] to cause an apparatus to at least: obtain at least one focus parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to control a relative emphasis in at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape to generate a processed spatial audio signal representing a modified audio scene; and outputting the processed spatial audio signal, wherein the modified audio scene enables relative emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0055] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least: obtain at least one focus parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene to control emphasis relative to at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape; and output the processed spatial audio signal, wherein the modified audio scene enables relative emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0056] According to a seventh aspect, there is provided an apparatus comprising: means for obtaining at least one focus parameter configured to define a focus shape; means for processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene to control emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; and means for outputting the processed spatial audio signal, wherein the modified audio scene enables relative emphasis of at least some of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.
[0057] According to an eighth aspect, there is provided a computer readable medium comprising program instructions for causing an apparatus to at least obtain at least one focus parameter configured to define a focus shape, process a spatial audio signal representing an audio scene to control emphasis relative to at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape, and generate a processed spatial audio signal representing a modified audio scene, wherein the modified audio scene enables relative emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape. An apparatus comprising means for performing the actions of the method described above. An apparatus configured to perform the actions of the method described above. A computer program comprising program instructions for causing a computer to perform the method described above. A computer program product stored on the medium can cause an apparatus to perform the methods described herein.
[0058] The electronic device may include an apparatus as described herein.
[0059] The chipset may consist of the devices described herein.
[0060] SUMMARY OF THE PRESENT EMBODIMENTS Embodiments of the present invention aim to solve problems associated with the state of the art. [Brief description of the drawings]
[0061] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings, in which: [Figure 1a] 1a and 1b show an exemplary sound scene illustrating an audio focus region or area. [Figure 1b]1a and 1b show an exemplary sound scene illustrating an audio focus region or area. [Figure 2a] 2a and 2b illustrate generally an exemplary playback device and method of operating the playback device according to some embodiments. [Figure 2b] 2a and 2b illustrate generally an exemplary playback device and method of operating the playback device according to some embodiments. [Diagram 3] FIG. 3 is a schematic diagram illustrating spherical harmonic patterns and a selected subset of these spherical harmonic patterns as applied in some embodiments. [Figure 4] FIG. 4 shows a schematic of the beam pattern corresponding to the Ambisonic signal and the converted beam signal aligned with an exemplary focus direction of 20 degrees. [Figure 5a] 5a and 5b illustrate diagrammatically an exemplary focus processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operating the exemplary focus processor according to some embodiments. [Figure 5b] 5a and 5b illustrate diagrammatically an exemplary focus processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operating the exemplary focus processor according to some embodiments. [Figure 6] FIG. 6 is a schematic diagram showing the processing state in an example where the focus direction is 20 degrees and the width is 45 degrees. [Figure 7] FIG. 7 is a visualization diagram that shows a schematic processing of a further example with a focus direction of minus 90 degrees and a width of 90 degrees. [Figure 8a] 8A and 8B are schematic diagrams illustrating the example focus processor shown in FIG. 2A having a parametric spatial audio signal input and a method of operating the example focus processor according to some embodiments. [Figure 8b]8A and 8B are schematic diagrams illustrating the example focus processor shown in FIG. 2A having a parametric spatial audio signal input and a method of operating the example focus processor according to some embodiments. [Figure 9a] 9a and 9b are schematic diagrams of the exemplary focus processor shown in FIG. 2a having a multi-channel and / or audio object audio signal input and a method of operating the exemplary focus processor in accordance with some embodiments. [Figure 9b] 9a and 9b are schematic diagrams of the exemplary focus processor shown in FIG. 2a having a multi-channel and / or audio object audio signal input and a method of operating the exemplary focus processor in accordance with some embodiments. [Figure 10] FIG. 10 illustrates an exemplary focus width determination based on focus distance and radius inputs according to some embodiments. [Figure 11a] 11a and 11b illustrate schematic diagrams of an exemplary playback processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operation of the exemplary playback processor according to some embodiments. [Figure 11b] 11a and 11b illustrate schematic diagrams of an exemplary playback processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operation of the exemplary playback processor according to some embodiments. [Figure 12a] 12a and 12b are schematic diagrams illustrating an exemplary playback processor as shown in FIG. 2a having a parametric spatial audio signal input according to some embodiments and a method of operating the exemplary playback processor. [Figure 12b]12a and 12b are schematic diagrams illustrating an exemplary playback processor as shown in FIG. 2a having a parametric spatial audio signal input according to some embodiments and a method of operating the exemplary playback processor. [Figure 13] FIG. 13 illustrates an example implementation of some embodiments. [Figure 14] FIG. 14 illustrates an example controller for controlling focus direction, focus amount, and focus width, according to some embodiments. [Figure 15] FIG. 15 illustrates an example processing output based on processing of a high-order Ambisonics audio signal according to some embodiments. [Figure 16] FIG. 16 illustrates an exemplary apparatus suitable for implementing the illustrated apparatus. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0062] In the following, preferred apparatus and possible mechanisms for providing efficient rendering and playback of spatial audio signals are described in further detail.
[0063] Previous examples of spatial audio signal reproduction allowed the user to control the focus direction and focus amount. However, in some situations, such control of focus direction / amount may not be sufficient. In some situations, it may be desirable to allow a user with a control interface to control the focus shape. In a sound field, there may be many different characteristics, such as ambient sounds as well as multiple dominant sound sources in a particular viewing direction. Some users may prefer to hear certain characteristics of the sound field, while other users may prefer to hear alternative characteristics of the sound field depending on which viewing direction is desired. It is understood that such reproduced audio depends on one or more preferences and is configurable based on user-related preferences. A desired performance from a reproduction device is to configure spatial audio reproduction so that focus can be controlled to various shapes or regions (e.g., narrow, wide, shallow, deep, near, far).
[0064] As an example, there may be audio content of interest in a sector (or cone or another spatial span or range) rather than just in one direction. In particular, it may be useful to control the spatial span of focus. Figures 1a, 1b, described below, show what a user is intended to perceive when listening to a reproduced spatial audio signal. For example, as illustrated in Figure 1a, there may be a source of interest on one side of the user and a source of distraction on the other side of the user. Figure 1a shows a user 101 positioned with a defined orientation. Within the audio scene there is a source of interest 105, such as a speaker in a theatre play, that is within a desired focus region 103 defined by a focus direction and width. Additionally, there may be an audience or other ambient audio content 107 that is outside the view direction, such as behind the view direction.
[0065] Additionally, a user may wish to change the width of the sector over time, for example initially focusing on all sources in a play by keeping the focus sector relatively wide (as shown in Figure 1a), and then later focusing on a particular source by narrowing the focus sector.
[0066] As another example, desired or interesting audio content may be at a distance (relative to the listener or relative to another position). For example, there may be undesirable or uninteresting audio sources at a distance in one direction and desirable or interesting audio sources at another distance in the same direction (or approximately the same direction). This is illustrated in FIG. 1b. FIG. 1b shows a user 101 located at a defined orientation in an audio scene with a source of interest 105, e.g., a talker, around a table, within a desired focus region 103 defined by a center position and a radius. Additionally, there may be other ambient audio content, such as environmental audio content 151 on the left, music source audio components 155, and other talker audio content 153 beyond the source of interest that is outside the desired focus region. In such an embodiment, the audio focus region or shape is determined by the center focus position and the focus radius.
[0067] Thus, the embodiments as discussed herein seek to provide control of focus shape (in addition to focus direction and amount). The concept as discussed with respect to the embodiments described herein is relevant for spatial audio playback in media playback with multiple viewing directions by providing control of audio focus shape where the audio scene on the controlled audio focus shape changes but the signal format can remain the same.
[0068] In embodiments, at least one focus shape parameter corresponding to a selectable direction is provided by adjusting any (or a combination of two or all) of the following parameters: focus width, focus height, focus radius, focus distance, and focus depth, corresponding to the selected direction. This parameter set in some embodiments is comprised of parameters that define an arbitrary shape.
[0069] Spatial audio signal processing may, in some embodiments, be performed by obtaining a spatial audio signal associated with media having multiple viewing directions, obtaining focus direction and amount parameters, modifying the spatial audio signal to have at least one focus desired focus characteristic, modifying the spatial audio signal to have the desired focus characteristic, and playing back (using headphones or loudspeakers) the modified spatial audio signal.
[0070] The resulting spatial audio signal may be in a parametric spatial audio format, such as, for example, an Ambisonic signal, a loudspeaker signal, or spatial metadata associated with a set of audio channels.
[0071] The focus shape may, in some embodiments, depend on what parameters are available. For example, if it only has direction, width, and height, the shape may be an ellipsoidal cone-shaped volume. As another example, if it only has distance and depth, the focus shape may be a hollow sphere. If it does not have width / height and / or depth, they may be assumed to have some default value. Additionally, in some embodiments, any focus shape may be used.
[0072] The amount of focus may, in some embodiments, determine the "degree" or how much focus is applied. For example, the focus may be from 0% to 100%, where 0% means to keep the original sound scene unchanged and 100% means to focus maximally on the desired spatial shape.
[0073] In some embodiments, different users may wish to have different focus characteristics, and the original spatial audio signal may be modified and played back individually for each user based on their individual preferences.
[0074] FIG. 2a shows a block diagram of some components and / or entities of a spatial audio processing device 250 according to an example. It will be understood that the two separate steps (focus processor + playback processor) shown in this figure and further detailed later can be implemented as an integrated process or, in some examples, in the reverse order as described herein (where the playback processor operation then follows the focus processor operation). The spatial audio processing device 250 consists of an audio focus processor 201 configured to receive an input audio signal and further focus parameters 202 and derive an audio signal having a focus sound component 204 based on the input audio signal 200 and depending on the focus parameters 202 (which may include focus direction, focus amount, focus height, focus radius, focus distance, and focus depth). In some embodiments, the device may be configured to obtain a focus shape, where the focus shape includes at least one focus parameter (which may be configured to define the focus shape). The spatial audio processing device 250 may further comprise an audio playback processor 207 configured to receive the focus sound component 204 and playback control information 206 and configured to derive an output audio signal 208 in a predefined audio format based on the audio signal with the focus sound component, further depending on the playback control information 206 serving to control at least one aspect of the processing of the spatial audio signal with the focus sound component in the audio playback processor 207. The playback control information 206 may comprise an indication of the playback direction (or playback direction) and / or an indication of the applicable loudspeaker configuration. Considering the above-mentioned processing method of the spatial audio signal, the audio focus processor 201 may be arranged to implement an aspect of processing the spatial audio signal by modifying the audio scene to control the emphasis in at least a part of the spatial audio signal in the received focus region according to the received focus amount.The audio playback processor 207 may output the processed spatial audio signals based on the observed direction and / or position as a modified audio scene, the modified audio scene demonstrating emphasis according to the received amount of focus for at least the portion of the spatial audio signal in the focus region.
[0075] In the description of FIG. 2a, each of the input audio signal, the audio signal with the focus sound component, and the output audio signal is provided as a respective spatial audio signal in a predefined spatial audio format. Thus, these signals may be referred to as the input spatial audio signal, the spatial audio signal with the focus sound component, and the output spatial audio signal, respectively. Along the lines mentioned above, typically, the spatial audio signal conveys an audio scene that includes both one or more directional sound sources at respective specific positions in the audio scene and the ambience of the audio scene. However, in some scenarios, the spatial audio scene may include one or more directional sound sources without ambience, or ambience without directional sound sources. In this regard, the spatial audio signal includes information conveying one or more directional sound components representing distinct sound sources having a fixed position (e.g., a fixed direction of arrival and a fixed relative intensity with respect to the listening point) in the audio scene and / or environmental sound components representing environmental sounds in the audio scene. It should be noted that although the division of an audio scene into directional sound component(s) and ambient components is typically only a representation or approximation, real sound scenes may contain more complex features such as wide sound sources, coherent acoustic reflections, etc. However, even with such complex acoustic features, conceptualizing an audio scene as a combination of direct and ambient components is typically a fair representation or approximation, at least in a perceptual sense.
[0076] Generally, the input audio signal and the audio signal with the pickup component are provided in the same predefined spatial format, but the output audio signal can be provided in the same spatial format applied to the input audio signal (and the audio signal with the pickup component) or a different predefined spatial format can be adopted for the output audio signal. The spatial audio format of the output audio signal is selected taking into account the characteristics of the sound reproduction hardware applied for the reproduction of the output audio signal. Generally, the input audio signal may be provided in a first predefined spatial audio format and the output audio signal can be provided in a second predefined spatial audio format. Non-limiting examples of spatial audio formats suitable for use as the first and / or second spatial audio format are Ambisonics, surround loudspeaker signals according to a predefined loudspeaker configuration, predefined parametric spatial audio formats. More detailed non-limiting examples of the use of these spatial audio formats in the framework of the spatial audio processing device 250 as the first and / or second spatial audio format are provided later in the present disclosure.
[0077] The spatial audio processor 250 is typically applied to process the input spatial audio signal 200 as a sequence of input frames into a respective sequence of output frames, each input (output) frame including a respective segment of a digital audio signal for each channel of the input (output) spatial audio signal, provided as a respective time series of input (output) samples at a predefined sampling frequency. In some embodiments, the input signal to the spatial audio processor 250 may be in an encoded form, e.g. AAC, or AAC+embedded metadata. In such embodiments, the encoded audio input may first be decoded. Similarly, in some embodiments, the output from the spatial audio processor 250 may be encoded in any suitable manner.
[0078] In a typical example, the spatial audio processing device 250 employs a fixed, predefined frame length, such that each frame is composed of L samples for each channel of the input spatial audio signal, respectively, and corresponds to a corresponding duration in time at a predefined sampling frequency. As an example in this regard, the fixed frame length may be 20 milliseconds (ms), resulting in frames of L=160, L=320, L=640 and L=960 samples per channel, respectively, at sampling frequencies of 8, 16, 32 or 48 kHz. The frames may be non-overlapping or partially overlapping, depending on whether the processor applies filter banks and how these filter banks are configured. However, these values serve as non-limiting examples, and frame lengths and / or sampling frequencies different from these examples may be employed instead, depending on, for example, the desired audio bandwidth, the desired framing delay and / or the available processing capacity.
[0079] In the spatial audio processing device 250, the focus refers to a user-selectable spatial region of interest. The focus may be, for example, a certain direction, distance, radius, arc of the general audio scene. In another example, it is a focus region where a (directional) sound source of interest is currently located. In the former scenario, the user-selectable focus typically indicates an area that stays constant or does not change frequently, since the focus prevails in a particular spatial region, while in the latter scenario, the user-selected focus may change more frequently, since the focus is set on a particular sound source that may (or may not) change its position / shape / size in the audio scene over time. In one example, the focus can be defined, for example, as an azimuth angle that defines the spatial direction of interest with respect to a first predefined reference direction, and / or as an elevation angle that defines the spatial direction of interest with respect to a second predefined reference direction, and / or as a shape and / or distance and / or radius or shape parameters.
[0080] The functionality described above with reference to the components of the spatial audio processing device 250 may for example be provided according to a method 260 illustrated by the flowchart depicted in Fig. 2b. The method 260 may for example be provided by an apparatus arranged to implement the spatial audio processing system 250 described in the present disclosure through a number of examples. The method 260 serves as a method for processing an input spatial audio signal representative of an audio scene into an output spatial audio signal representative of a modified audio scene. The method 260 comprises receiving an indication of a focus region and an indication of focus intensity, as illustrated in block 261.
[0081] The method 260 further comprises processing the input spatial audio signals into intermediate spatial audio signals representing a modified audio scene in which the relative levels of sounds coming from the focus regions are modified according to the focus intensity, as shown in block 263.
[0082] The method 260 further comprises receiving playback control information to control the processing of the intermediate spatial signals into output spatial audio signals, as shown in block 265. The playback control information may, for example, define at least one of a playback direction (e.g., listening direction or line of sight direction) or a loudspeaker configuration for the output spatial audio signals.
[0083] The method 260 further includes processing the intermediate spatial audio signals into the output spatial audio signals according to the playback control information, as shown in block 267 .
[0084] The method 260 can be varied in a number of ways, for example according to examples of the functionality of each of the components of the spatial audio processor 250 described above and provided below.
[0085] In some embodiments, the input to the spatial audio processing device 250 is an Ambisonic signal. The device can be configured (and the method can be applied) to receive Ambisonic signals of any order. However, it is illustrated that a first-order Ambisonic (FOA) signal has a fairly broad spatial selectivity (specifically a first-order directivity), and therefore a higher-order Ambisonic (HOA) signal with a higher spatial selectivity is suitable for fine control of the focus shape. In particular, in the following examples, the method and device are configured to receive a third-order Ambisonic audio signal.
[0086] A third order Ambisonic audio signal has a total of 16 beam pattern signals (in 3D). However, for simplicity in the following example, only the seven more "horizontal" Ambisonic components (in other words, audio signals) are considered here, as shown in Fig. 3, to illustrate the implementation of the focus shape parameters. For example, Fig. 3 shows the 0th order spherical harmonic pattern 301, the 1st order spherical harmonic pattern 303, the 2nd order spherical harmonic pattern 305, and the 3rd order spherical harmonic pattern 307. Fig. 3 further shows subsets 309 and 311 for up to the more "horizontal" 3rd order spherical harmonic pattern.
[0087] With reference to FIG. 5a, an exemplary Ambisonic signal x HOA Illustrated is a focus processor 550 configured to receive (t) 500 and a focus direction 502. As mentioned above, the input to focus processor 550 in this example is a subset third order Ambisonic signal, e.g. subsets 309 and 311. Also, in the following, a third order Ambisonic signal x HOA For simplicity, (t) 500 is denoted as HOA. The signal x(t) arriving from the horizontal direction θ with t being the discrete sample index is
number
[0088] In some embodiments, the focus processor 550 is comprised of a matrix processor 501. The matrix processor 501 is configured in some embodiments to convert an Ambisonic (HOA) signal 500 (corresponding to an Ambisonic or Spherical Harmonic Pattern) into a set of seven equally spaced horizontal beam signals (corresponding to a beam pattern). This is done in some embodiments by using a transformation matrix T(θ f ), and θ f is the focus direction 502 parameter.
number
number
number
[0089] For example, θ f = 20 degrees, the transformed signal x cThe beam pattern corresponding to (t) 504 and the beam pattern corresponding to the original HOA signal are shown in Fig. 4. Fig. 4 shows an example of a beam pattern corresponding to an Ambisonic signal in the top row 401 and a transformed beam signal with a focus direction at 20 degrees in the bottom row 403. The transformed audio signal can then be output to a spatial beam (based on focus parameters) processor 503.
[0090] The focus processor 550 may further include a spatial beam (based on focus parameters) processor 503. The spatial beam processor 503 processes the converted Ambisonic signal x from the matrix processor 501. c (t) 504 and is further configured to receive focus amount and width focus parameters 508 .
[0091] The spatial beam processor 503 then processes the spatial beam signal x c (t) 504 to obtain the processed or modified spatial beam signal x' c (t) 506 is based on the focus amount and the shape parameters 508. c (t) 506 may then be output to a further matrix processor 505. The spatial beam processor 503 is configured to implement different processing methods based on the type of focus shape parameters. In this exemplary embodiment, the focus parameters are focus direction, focus width, and focus amount. The focus amount can be determined as a value a ranging between 0...1, where 1 indicates maximum focus. The focus width θ w (determined as the angle from the focus direction to the edge of the focus arc) is also a variable or controllable parameter. The spatial beam signal is
number
number
[0092] In this example, beam x c Note that (t) is formulated so that the first beam points in the focus direction and the second beam points in the focus direction +p. As a result, the matrix I(θ w When applying option a), the beam far from the focus direction will be attenuated depending on the focus width parameter.
[0093] The focus processor 201 further comprises a matrix processor 505. The further matrix processor 505 outputs the processed or modified spatial beam signal x' c (t) 506. The result of inversely transforming the focus direction 502 is generated as a focus-processed HOA signal. f ) is invertible, so the inversion process is
number
[0094] Regarding Fig. 6, the focus parameter is the maximum focus amount a=1, and the focus direction is θ f = 20 degrees, focus width θ w = 45 degrees. The top row 601 shows the focused transform domain signal x' c The lower row 603 shows the output signal x' HOA 7 shows the beam pattern corresponding to (t). With respect to FIG. 7, the focus parameter is the maximum focus amount a=1, and the focus direction parameter is θ f =-90 degrees, θ w= 90 degrees. The top row 701 shows the focused transform domain signal x' c The lower part 703 shows the beam pattern corresponding to the output signal x' HOA The corresponding beam pattern is shown in (t).
[0095] In the above examples, it has been shown that the HOA processing is only considered on a set of more "horizontal" beam pattern signals, it is understood that these operations can be extended to 3D using a set of 3D beam patterns.
[0096] With reference to FIG. 5b, there is shown a flow diagram of the operation 560 of the HOA focus processor as shown in FIG. 5a.
[0097] The first operation is to receive the HOA audio signal (and focus parameters such as direction, width, amount or other control information) as shown in FIG. 5b by step 561.
[0098] The next operation is to generate the converted HOA audio signal into a beam signal, as shown in FIG. 5b at step 563.
[0099] After converting the HOA audio signal into a beam signal, the next operation is one of spatial beam processing, as shown in FIG. 5b by step 565.
[0100] The processed beamed audio signals are then converted back to HOA format as shown in FIG. 5b by step 567.
[0101] The processed HOA audio signal is then output by step 569 as shown in FIG. 5b.
[0102] With reference to FIG. 8a, a focus processor is shown configured to receive a parametric spatial audio signal as input. The parametric spatial audio signal consists of an audio signal and spatial metadata such as direction(s) in frequency bands and direct-to-total energy ratio(s). The structure and generation of parametric spatial audio signals is known and its generation has been described from microphone arrays (e.g. mobile phones, VR cameras). Parametric spatial audio signals can also be generated from loudspeaker signals and Ambisonic signals. The parametric spatial audio signal in some embodiments may be generated from an IVAS (Immersive Voice and Audio Services) audio stream, which can be decoded and demultiplexed into the form of spatial metadata and audio channels. A typical number of audio channels of such a parametric spatial audio stream is a two audio channel audio signal, but in some embodiments the number of audio channels can be any number.
[0103] In these examples, the parametric information consists of depth / distance information, which may be implemented in six degrees of freedom (6DOF) playback, where distance metadata is used (along with other metadata) to determine how sound energy and direction should change in response to user movement.
[0104] Thus, in this example, each spatial metadata directional parameter is associated with both a direct-to-global energy ratio and a distance parameter. The estimation of distance parameters in the context of parametric spatial audio capture has been detailed in previous applications such as GB patent applications GB1710093.4 and GB1710085.0 and, for reasons of clarity, will not be discussed further.
[0105] A focus processor 850 configured to receive parametric (in this case 6DOF capable) spatial audio 800 is configured to use the focus parameters (in these examples, focus direction, amount, distance, and radius) to determine how much to attenuate or emphasize the direct and ambient components of the parametric spatial audio signal to effect a focus effect.
[0106] In the following examples, the methods (and formulas) are expressed without variation over time, however, it is understood that all parameters may vary over time.
[0107] In some embodiments, the focus processor comprises a ratio correction and spectral adjustment coefficient determiner 801 configured to receive focus parameters 808 and further spatial metadata consisting of direction 802, distance 822, and direct-to-total energy ratio 804 of frequency bands.
[0108] The ratio corrector and the spectral adjustment coefficient determiner are configured to implement the focus shape as a sphere in 3D space by first transforming the focus direction and distance into a Cartesian coordinate system (a 3x1 yzx vector f)
number
[0109] Similarly, for each frequency band k, the direction and distance of the spatial metadata are
number
[0110] The units of the spatial metadata distance and focus distance parameters should be the same (e.g., both in meters, or in some other scale). The mutual distance value d(k) of f and m(k) can be simply formulated as follows:
number
[0111] This mutual distance value d(k) is then used in a gain function along with a focus amount parameter a, which ranges from 0..1, and a focus radius parameter dr (with the same units as d(k)). When focusing, an example gain formula is:
number
[0112] In practice, it may be desirable to smooth the above focus gain function so that it transitions smoothly from high values in focus regions to low values in unfocused regions.
[0113] Then, the new direct part value D(k) of the parametric spatial audio signal is
number
number
number
number
[0114] For the numerically undetermined case D(k) = A(k) = 0, r'(k) can also be set to 0.
[0115] The direction and distance parameters of the spatial metadata may not be modified by the metadata adjustment and spectral adjustment coefficient determiner 801 and the modified and unmodified metadata output 810 in some embodiments.
[0116] The spatial processor 850 may include a spectral adjustment processor 803. The spectral adjustment processor 803 may be configured to receive an audio signal 806 and spectral adjustment coefficients 812. The audio signal may be in a time-frequency representation in some embodiments, or alternatively, first transformed to the time-frequency domain for the spectral adjustment process. The output 814 may also be in the time-frequency domain, or may be transformed back to the time domain before output. The domains of the input and output are implementation dependent.
[0117] The spectral adjustment processing unit 803 can be configured to multiply, for each band k, the frequency bins (of the time-frequency transform) of all channels in band k by a spectral adjustment coefficient s(k), i.e., perform spectral adjustment. The multiplication (i.e., the spectral correction) can be smoothed in time to avoid processing artifacts.
[0118] In other words, the processor is configured such that the spectral and spatial metadata of the signal is modified such that the procedure modifies the parametric spatial audio signal according to the focus parameters (in this case focus direction, amount, distance, radius).
[0119] With reference to FIG. 8b, there is shown a flow diagram 860 of the operation of a parametric spatial audio input processor such as that shown in FIG. 8a.
[0120] The first operation is to receive a parametric spatial audio signal (and focus parameters or other control information) as shown in FIG. 8b by step 861.
[0121] The next operation is the modification of the parametric metadata and the generation of spectral adjustment coefficients, as indicated in FIG. 8b by step 863.
[0122] The next operation is to perform spectral adjustments on the audio signal, as shown in FIG. 8b at step 865.
[0123] The spectrally adjusted audio signal and the modified (and unmodified) metadata may then be output by step 867 as shown in FIG. 8b.
[0124] With reference to FIG. 9a, a focus processor 950 is shown that is configured to receive a multi-channel or object audio signal as an input 900. The focus processor in such an embodiment may be comprised of a focus gain determiner 901. The focus gain determiner 901 is configured to receive focus parameters 908 and channel / object position / direction information, which may be static or time-varying. The focus gain determiner 901 is configured to generate direct gain f(k) parameters that are output as focus gains 912 for each channel based on the focus parameters 908 and the channel / object position / direction information 902 from the input signal 900. In some embodiments, the directions of the channel signals are signaled and in some embodiments, they are assumed. For example, when there are six channels, the directions can be assumed to be 5.1 audio channel directions. In some embodiments, there may be a look-up table that is used to determine the channel directions as a function of the number of channels.
[0125] For audio objects with direction and distance (i.e., position), the focus gain determiner 901 may utilize the same implementation process as expressed in the context of parametric audio processing to determine the direct gain f(k) 912 based on the spatial metadata and the focus parameters. In these embodiments, there is no filter bank, i.e., there is only one frequency band k.
[0126] The focus processor may also include a focus gain processor (for each channel) 903. The focus gain processor 903 is configured to receive a focus gain f(k) 912 for each audio channel and audio signal 906. The focus gain 912 may then be applied to the corresponding audio channel signal 906 (which may in some embodiments be further temporally smoothed). The output from the focus gain processor 903 may be a focus processed audio channel audio signal 914.
[0127] In these examples, the channel direction / position information 902 is unchanged and provided as channel direction / position information output 910 .
[0128] In some embodiments, when the input audio channels do not have distance information (e.g., loudspeakers or object sounds where the input is only direction and not distance), one option for processing such audio channels is to determine a fixed default distance for such signals and apply the same formula to determine f(k).
[0129] In some embodiments, determining the focus gain f(k) 912 for such an audio channel can be based on the angular difference between the focus direction and the direction of the audio channel. In some embodiments, this may first determine a focus width θ_w. For example, as shown in FIG. 10, the focus width θ_w 1005 may be determined based on trigonometry using the focus distance 1001 and the focus radius 1003, where the focus width is generated by the angle of a right triangle with the hypotenuse formed by the focus distance 1001 and the opposite side formed by the focus radius 1003. The focus width is simply
number
[0130] With reference to FIG. 9b, there is shown a flow diagram 960 of the operation of the multi-channel / object audio input processing device shown in FIG. 9a.
[0131] The first operation is to receive a multi-channel / object audio signal (and channel information such as focus parameters or other control information, and direction / distance) as shown in FIG. 9b by step 961.
[0132] The next operation is to generate focus gain coefficients, as shown in Figure 9b by step 963. The next operation is to apply a focus gain to each channel audio signal, as shown in Figure 9b by step 965. The processed audio signal and the unmodified channel directions (and distances) may then be output, as shown in Figure 9b by step 967.
[0133] In some embodiments, the focus shape may be defined using other parameters and other combinations of parameters, in which case the focus processor may be modified from the above example to use those parameters.
[0134] With reference to FIG. 11a, examples of playback processors 1150 based on Ambisonic audio input (which may be configured to receive, for example, an output from an example focus processor as shown in FIG. 5a) are shown. In these examples, the playback processors may consist of an Ambisonic rotation matrix processor 1101. The Ambisonic rotation matrix processor 1101 is configured to receive an Ambisonic signal having a focus process 1100 and a view direction 1102. The Ambisonic rotation matrix processor 1101 is configured to generate a rotation matrix based on the view direction parameter 1102. This may use any suitable method, such as that applied in head-tracked Ambisonic Ainauralization in some embodiments (or more generally, such rotation of spherical harmonics is used in many fields, including outside of audio). This rotation matrix is then applied to the Ambisonic audio signal. The result is a rotated Ambisonic signal with added focus 1104, which is output to the Ambisonic to binaural filter f1103. The Ambisonic to Binaural filter 1103 is configured to receive the focused rotated Ambisonic signal 1104 .
[0135] The Ambisonic to Binaural Filter 1103 may consist of a preformed 2xK matrix of Finite Impulse Response (FIR) filters that are applied to the K Ambisonic signals to generate the 2 binaural signals 1106. The FIR filters may be generated by least squares optimization with respect to a set of Head Related Impulse Responses (HRIRs). An example of such a design procedure is to transform the HRIR dataset into frequency bins (e.g., by FFT) to obtain an HRTF dataset, and to determine for each frequency bin a complex-valued processing matrix that least squares-approximates the available HRTF dataset at the data points of the HRTF dataset. When the complex-valued matrices for all frequency bins are so determined, the result may be inverted (e.g., by inverse FFT) as a time-domain FIR filter. The FIR filters may also be windowed, for example, by using a Hann window.
[0136] There are many known methods that can be used to render Ambisonic signals into loudspeaker outputs. As an example, the Ambisonic signals can be linearly decoded to the target loudspeaker configuration. This can be applied if the order of the Ambisonic signal is sufficiently high, for example at least third order, preferably fourth order. In an example of such linear decoding, an Ambisonic decoding matrix can be designed that, when applied to the Ambisonic signals (corresponding to the Ambisonic beam pattern), produces loudspeaker signals corresponding to a beam pattern that in a least squares sense approximates a vector-base amplitude panning (VBAP) beam pattern suitable for the target loudspeaker configuration. Processing the Ambisonic signals with such a designed Ambisonic decoding matrix can be configured to generate loudspeaker sound outputs. In such an embodiment, the playback processor is configured to receive information regarding the loudspeaker configuration.
[0137] With reference to FIG. 11b, there is shown a flow diagram 1160 of the operation of the Ambisonic input playback processor shown in FIG. 11a.
[0138] The first operation is to receive a focused Ambisonic audio signal (and a view direction) as shown in FIG. 11b by step 1161.
[0139] The next operation is to generate a rotation matrix based on the view direction, as shown in FIG. 11b by step 1163.
[0140] The next operation is to apply a rotation matrix to the Ambisonic audio signal to generate a focused rotated Ambisonic audio signal, as shown in FIG. 11b by step 1165.
[0141] The next operation is to convert the Ambisonic audio signal into a suitable audio output format, for example a binaural format (or a multi-channel audio format), as shown in FIG. 11b by step 1167.
[0142] Then, step 1169 outputs the output audio format as shown in FIG. 11b.
[0143] With reference to FIG. 12a, an example of a playback processor 1250 based on parametric spatial audio input (which may be configured to receive output from an example focus processor such as that shown in FIG. 8a) is shown.
[0144] In some embodiments, the playback processor comprises a filter bank 1201 configured to receive the audio signals in the audio channels 1200 and convert the audio channels into frequency bands (unless the input is already in the appropriate time-frequency domain). Examples of suitable filter banks include short-time Fourier transform (STFT) and complex quadrature mirror filter (QMF) banks. The time-frequency audio signal 1202 can be output to a parametric binaural synthesizer 1203.
[0145] In some embodiments, the playback processor is comprised of a parametric binaural synthesizer 1203 configured to receive the time-frequency audio signal 1202, the modified (and unmodified) metadata 1204, and further the view direction 1206 (or suitable playback related control or tracking information). In the 6DOF context, the user position can be provided along with the view direction parameter.
[0146] The parametric binaural synthesizer 1203 can be configured to implement any suitable known parametric spatial synthesis method configured to generate a binaural audio signal (frequency bands) 1208, since focus correction has already been performed on the signal and metadata before the parametric binauralization block. The binauralized time-frequency audio signal 1208 can then be passed to an inverse filterbank 1205. The embodiment may be further characterized in that the playback processor comprises an inverse filterbank 1205 configured to receive the binauralized time-frequency audio signal 1208 and generate the inverse of the applied forward filterbank, thus generating a time-domain binauralized audio signal 1210 with focus characteristics suitable for playback by headphones (not shown in Fig. 12a).
[0147] In some embodiments, the binaural audio signal output is replaced in loudspeaker channel audio signal output format from the parametric spatial audio signal using a suitable loudspeaker synthesis method. Any suitable approach may be used, for example, the view direction parameters may be replaced with information of the loudspeaker positions and the binaural processor may be replaced with a loudspeaker processor based on a suitable known method.
[0148] With reference to FIG. 12b, there is shown a flow diagram 1260 of the operation of a parametric spatial audio input playback processor such as that shown in FIG. 12a.
[0149] The first operation is to receive a focused parametric spatial audio signal (and view direction or other playback related control or tracking information) as shown in FIG. 12b by step 1261.
[0150] The next operation is to time-frequency transform the audio signal as shown in Fig. 12b by step 1263. The next operation is to apply a parametric binaural (or loudspeaker channel type) processor based on the time-frequency transformed audio signal, the metadata and the listening direction (or other information) as shown in Fig. 12b by step 1265.
[0151] The next operation is then to inverse transform the generated binaural or loudspeaker channel audio signals as shown in FIG. 12b by step 1267.
[0152] It then outputs the output audio format as shown in Figure 12b by step 1269. Considering the loudspeaker output of the playback processor when the audio signal is in the form of multi-channel audio and the focus processor 950 of Figure 9a is applied, in some embodiments the playback processor may configure a pass-through where the output loudspeaker configuration is the same as the format of the input signal.
[0153] In some embodiments where the output loudspeaker configuration differs from the input loudspeaker configuration, the playback processor can be comprised of a vector-based amplitude panning (VBAP) processor. Each focused audio channel can then be processed using VBAP, a known amplitude panning technique, and spatially reproduced using the target loudspeaker configuration. In this way, the output audio signal is adapted to the output loudspeaker setting.
[0154] In some embodiments, the transformation from the first loudspeaker configuration to the second loudspeaker configuration may be performed using any suitable amplitude panning technique. For example, the amplitude panning technique may consist of deriving an N×M matrix of amplitude panning gains that defines the transformation from the M channels of the first loudspeaker configuration to the N channels of the second loudspeaker configuration, and then multiplying the channels of the intermediate spatial audio signal provided as a multi-channel loudspeaker signal according to the first loudspeaker configuration with the matrix. The intermediate spatial audio signal may be understood to be similar to an audio signal having a focused sound component 204, as shown in FIG. 2a. As a non-limiting example, the derivation of the VBAP amplitude panning gains is described in Pulkki, Ville. "Virtual sound source positioning using vector base amplitude panning", Journal of the audio engineering society 45, no. 6 (1997), pp. 456-466.
[0155] For binaural output, any suitable binauralization of the multi-channel loudspeaker signal format (and / or objects) can be performed. For example, a typical binauralization might consist of processing the audio channels with Head Related Transfer Functions (HRTFs) and adding synthetic room reverberation to generate an auditory impression of a listening room. Distance + direction (i.e., position) information of the audio object sounds can be utilized for 6 degree of freedom reproduction with user movement, for example by employing the principles outlined in GB patent application GB1710085.0.
[0156] An example of an apparatus suitable for implementation is shown in Figure 13 in the form of a mobile phone or mobile device 1401 running suitable software 1403. Video may be played, for example, by attaching the mobile phone 1401 to a Daydream view type device (although for clarity, video processing is not described here).
[0157] The audio bitstream obtainer 1423 is configured to obtain an audio bitstream 1424, e.g., received / obtained from a storage. In some embodiments, the mobile device comprises a decoder 1425 configured to receive compressed audio and decode it. An example of a decoder is an AAC decoder in the case of AAC decoding. As a result, a decoded (e.g., Ambisonic, in which the embodiments shown in Figures 5a and 11a are implemented) audio signal 1426 may be forwarded to a focus processor 1427.
[0158] The mobile phone 1401 receives controller data 1400 from an external controller (e.g., via Bluetooth) at controller data receiver 1411 and passes the data to focus parameter (from controller data) determiner 1421. The focus parameter (from controller data) determiner 1421 determines focus parameters based, for example, on the orientation of the controller device and / or button events. The focus parameters may consist of any kind of combination of proposed focus parameters (e.g., focus direction, focus amount, focus height, and focus width). The focus parameters 1422 are forwarded to focus processor 1427.
[0159] Based on the Ambisonic audio signal and the focus parameters, the focus processor 1427 is configured to create modified Ambisonic signals 1428 having desired focus characteristics. These modified Ambisonic signals 1428 are forwarded to the Ambisonic-to-Binaural Processor 1429. The Ambisonic-to-Binaural Processor 1429 is also configured to receive head orientation information 1404 from the orientation tracker 1413 of the mobile phone 1401. Based on the modified Ambisonic signal 1428 and the head orientation information 1404, the Ambisonic-to-Binaural Processor 1429 is configured to create a head-tracked binaural signal 1430 that can be output from the mobile phone and played using, for example, headphones.
[0160] 14 shows an exemplary device (or focus parameter control device) 1550 that may be configured to control or generate appropriate focus parameters such as focus direction, focus amount, and focus width. A user of the device may be configured to select a focus direction by pointing the controller in a desired direction 1509 and pressing a focus direction selection button 1505. The controller has an orientation tracker 1501, and the orientation information may be used to determine the focus direction (e.g., in focus parameter (from controller data) determiner 1421, as shown in FIG. 13).
[0161] The focus direction in some embodiments can be visualized on a visual display while selecting the focus direction. In some embodiments, the focus amount can be controlled using focus amount buttons (shown as + and - in FIG. 14) 1507. Each press can increase / decrease the focus amount by, for example, 10 percentage points. The focus width can be controlled using focus width buttons (shown as + and - in FIG. 14) 1503. Each press can be configured to increase / decrease the focus width by a fixed amount, such as 10 degrees.
[0162] In some embodiments, the focus shape can be determined by drawing the desired shape with a controller (e.g., as depicted in FIG. 14). The user can initiate a drawing operation by pressing and holding the focus direction selection button, draw the desired shape with the controller, and finally accept the shape by releasing the press. A visual display of the drawn shape may be displayed while drawing. The drawn shape can be converted into focus direction, focus height, and focus width parameters. The focus amount may be selected with the "Focus Amount" button, as in the previous example.
[0163] In some embodiments, the focus controller as shown in FIG. 14 is modified such that the "focus width" control is replaced with a "focus radius" control, allowing for complex, content-adaptive control of the focus shape. In such an embodiment, the 360 video may be implemented as part of an advanced virtual reality playback system in which the 360 video is not only panoramic, but also includes depth information (i.e., is effectively a 3D video that can respond to user movements in 6 degrees of freedom). For example, the video content may be generated by computer graphics or by a VR video capture system that can detect visual depth and thus allows 6DOF as well as computer graphics.
[0164] For example, in a scene, there are two objects of interest (e.g., speakers). The user clicks "select focus direction" for these two sound sources, and the visual display indicates to the user that these sound sources (not only the auditory sources, but also the visual sources at a certain direction and distance) have been selected for audio focus. The user then selects the focus amount and focus radius parameters, where the focus radius indicates how much the auditory events from the source of interest will be contained within the determined focus shape. During control adjustment, the focus radius may be shown as a visual sphere around the visual source of interest.
[0165] While the field of view may respond to the user's movements, the source may also move in the scene, and its position is typically tracked visually. The focus shape may thus be represented in this case by two spheres in 3D space, and then its overall shape can be adaptively changed by moving the spheres. This means that a complex focus shape with depth focus is obtained. Depending on the form of spatial audio, the focus shape can then be reproduced exactly (provided that spatial audio has reliable distance information) or approximated in other ways, for example as illustrated above.
[0166] In some embodiments, it may be desirable to further specify the focus processing, for example, by determining a desired frequency range or spectral characteristics of the focused signal. In particular, it may be useful to emphasize the focused audio spectrum in audio frequency bands to improve intelligibility, for example, by attenuating low frequency content (e.g., below 200 Hz), high frequency content (e.g., above 8 kHz), and leaving particularly useful frequency bands associated with the audio.
[0167] It will be appreciated that the focus processed signal may be further processed with any known audio processing techniques, such as automatic gain control or enhancement techniques (eg, bandwidth expansion, noise suppression).
[0168] In some further embodiments, the focus parameters (including direction, amount, and at least one focus shape parameter) are generated by the content creator, and the parameters are transmitted together with the spatial audio signal. For example, the scene may be a VR video / audio recording of an unplugged music concert near the stage. The content creator may assume that a typical remote listener wants to determine a focus arc that extends towards the stage and also to the sides for room acoustics, but wants to remove at least to some extent the direct sound from the audience (behind the main direction of the VR camera). So we added a track of focus parameters to the stream and made it possible to set it as the default rendering mode. However, the sound of the audience is still present in the stream, and some users would prefer to discard the focus processing and be able to play the full sound scene including the sound of the audience.
[0169] That is, instead of the user selecting the focus direction or shape, preset dynamic focus parameters can be selected. The presets may be fine-tuned by the content creator to fit the show well, for example turning off focus at the end of each song and playing applause to the listener. The content creator can generate several expected preferred profiles for the focus parameters. This approach is beneficial because only one spatial audio signal needs to be conveyed, but it is also possible to add different preferred profiles. Legacy players that do not have focus enabled can decode the Ambisonic signal without the focus step.
[0170] In some further embodiments, the focus shape is controlled together with the visual zoom of a video with multiple viewing directions. The visual zoom can be conceptualized as a user controlling a set of virtual binoculars in a panoramic or 360 or 3D video. In such a use case, when the visual zoom function is enabled (e.g., at least 1.5x zoom is set), the audio focus of the spatial audio signal can also be enabled. At this time, since the user is clearly interested in that direction, the amount of focus can be set to a high value, for example 80%, and the focus width can be set to correspond to the arc of the visual field of the virtual binoculars. That is, the larger the visual zoom, the smaller the focus width. With the focus set to 80%, the user can hear the remaining spatial sound to some extent in the appropriate direction. In that way, the user can hear the occurrence of interesting new content and know to turn off the visual zoom and look in the new direction of interest. The zoom process can also be used in the context of an audio codec that allows such processing. An example of such a codec can be, for example, MPEG-I.
[0171] A user in the above-described embodiment can use the present invention to control the focus shape in a general purpose manner.
[0172] An example of the processing output based on the described embodiment for a Higher Order Ambisonics (HOA) signal is shown in Figure 15. This figure shows a spectrogram of a third order HOA signal, with a talker at 0°, a sine wave at -90°, and white noise at 110°, decoded output from 8 loudspeakers. The figure shows that narrowing the focus towards the talker reduces the relative energy of the sine wave and white noise, while a wider focus that includes both the talker and the sine wave significantly reduces the relative energy of only the white noise.
[0173] With respect to Figure 16, an example of an electronic device that can be used as an analysis device or synthesis device is shown. The device may be any suitable electronic device or device. For example, in some embodiments, the device 1700 is a mobile device, a user equipment, a tablet computer, a computer, an audio playback device, or the like.
[0174] In some embodiments, the device 1700 includes at least one processor or central processing unit 1707. The processor 1707 may be configured to execute various program code, such as the methods as described herein.
[0175] In some embodiments, the apparatus 1700 comprises a memory 1711. In some embodiments, at least one processor 1707 is coupled to the memory 1711. The memory 1711 may be any suitable storage means. In some embodiments, the memory 1711 constitutes a program code section for storing program code executable by the processor 1707. Furthermore, in some embodiments, the memory 1711 may further comprise a storage data section for storing data, for example data processed or to be processed according to the embodiments as described herein. The implementation program code stored in the program code section and the data stored in the storage data section can be retrieved by the processor 1707 whenever necessary via the memory-processor coupling.
[0176] In some embodiments, the apparatus 1700 comprises a user interface 1705. The user interface 1705 may, in some embodiments, be coupled to a processor 1707. In some embodiments, the processor 1707 may control the operation of the user interface 1705 and receive input from the user interface 1705. In some embodiments, the user interface 1705 may allow a user to input commands into the device 1700, for example via a keypad. In some embodiments, the user interface 1705 may allow a user to obtain information from the device 1700. For example, the user interface 1705 may include a display configured to display information from the device 1700 to the user. The user interface 1705 may, in some embodiments, be comprised of a touch screen or touch interface capable of both allowing information to be input into the device 1700 and further displaying information to the user of the device 1700.
[0177] In some embodiments, the apparatus 1700 includes an input / output port 1709. The input / output port 1709 in some embodiments comprises a transceiver. The transceiver in such embodiments may be coupled to the processor 1707 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may in some embodiments be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.
[0178] The transceiver may communicate with the further device by any suitable known communication protocol, for example in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short range radio frequency communication protocol such as Bluetooth, or an Infrared Data Path (IRDA).
[0179] The transceiver input / output port 1709 may be configured to receive signals and, in some embodiments, obtain focus parameters as described herein.
[0180] In some embodiments, device 1700 may be employed to generate appropriate audio signals using processor 1707 executing appropriate code. Input / output port 1709 may be coupled to any suitable audio output, such as to a multi-channel speaker system and / or headphones (which may be head tracked or non-tracked headphones).
[0181] In general, various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, although the invention is not limited thereto.
[0182] Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representations, it is to be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controllers or other computing devices, or any combination thereof.
[0183] The embodiments of the invention can be implemented by computer software executable by a data processor of a mobile device such as a processor entity, or by hardware, or by a combination of software and hardware. Further in this regard, it should be noted that any block of the logic flow as shown may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc.
[0184] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memories, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), gate level circuits, and processors based on multi-core processor architectures.
[0185] Embodiments of the present invention may be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic level designs into semiconductor circuit designs suitable for etching onto semiconductor substrates.
[0186] Programs such as those from Synopsys, Inc. of Mountain View, California, and Cadence Design, Inc. of San Jose, California, use established design rules and libraries of pre-stored design modules to automatically route the wires and place the components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the design may be sent in a standardized electronic format (e.g., Opus, GDSII) to a semiconductor manufacturing facility or "fab" for fabrication.
[0187] The foregoing description has provided a complete and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant art in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined in the appended claims.
Claims
1. 1. An apparatus for spatial audio reproduction, comprising: at least one processor; When executed by the at least one processor, the device includes at least: Obtaining a focus amount configured to define a focus area and an amount of focus; processing the spatial audio signal representing an audio scene to control attenuation of at least a portion of at least a portion of the spatial audio signal that is outside the acquired focus region in accordance with the focus amount relative to at least a portion of a portion of the spatial audio signal that is within the acquired focus region to generate a processed spatial audio signal that represents a modified audio scene; outputting the processed spatial audio signal, wherein the modified audio scene includes the attenuation of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region according to the focus amount relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; at least one memory storing instructions for executing the An apparatus comprising:
2. The obtained focus area is then transmitted to the device. Obtaining at least one of a focus direction and a focus width; The device of claim 1 .
3. The processed spatial audio signal is transmitted to the device: increasing the attenuation of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; or reducing the attenuation of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; The apparatus of claim 1 , further comprising:
4. The processed spatial audio signal is transmitted to the device: decreasing or increasing a relative sound level of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; The device of claim 1 .
5. The device of claim 3 , wherein the processed spatial audio signal causes the device to decrease or increase the relative sound levels according to the amount of focus.
6. The apparatus further comprises: obtaining playback control information for controlling at least one aspect of the output of the processed spatial audio signal; It was like this, The device comprises: - enabling the processed spatial audio signals, representing the audio scene modified according to the playback control information, to generate an output spatial audio signal; or processing the spatial audio signal in accordance with the playback control information, making available a modified audio scene, and outputting the processed spatial audio signal as the output spatial audio signal; and outputting the processed spatial audio signal, causing the device to further perform one of: The device according to claim 1 .
7. The spatial audio signal and the processed spatial audio signal each comprise an Ambisonic signal, and the processed spatial audio signal comprises, for one or more frequency subbands, converting the Ambisonic signal associated with the spatial audio signal into a set of beam signals in a predetermined pattern; generating a modified set of beam signals based on the set of beam signals, the focus region, and the focus amount; transforming the set of modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal; The apparatus of claim 1 , wherein the apparatus causes the execution of the following:
8. The apparatus of claim 7 , wherein the predetermined pattern comprises a defined number of beams equally spaced on a plane or volume.
9. The spatial audio signal and the processed spatial audio signal are Each higher-order Ambisonic signal, or A subset of Ambisonic signal components of any order, 8. The apparatus of claim 7, comprising at least one of:
10. the spatial audio signal and the processed spatial audio signal each comprise a parametric spatial audio signal, the parametric spatial audio signal comprising one or more audio channels and spatial metadata, the spatial metadata comprising at least one of a direction indication, an energy ratio parameter, and a distance indication for each of a plurality of frequency subbands, and the processed spatial audio signal is transmitted to the device; calculating spectral adjustment factors for one or more frequency subbands based on the spatial metadata, the focus region, and the focus amount; applying the spectral adjustment coefficients to the one or more frequency subbands of the one or more audio channels to generate processed one or more audio channels; calculating modified energy ratio parameters associated with the one or more frequency subbands of the processed spatial audio signal based at least in part on the focus region, the focus amount and the spatial metadata; constructing the processed spatial audio signal, the processed audio channel or channels being processed, the modified energy ratio parameter, and the spatial metadata other than the energy ratio parameter; The apparatus of claim 1 , further comprising:
11. The spatial audio signal and the processed spatial audio signal include multi-channel loudspeaker channels and / or audio object channels, and the processed spatial audio signal is transmitted to the device by: calculating a gain adjustment factor based on the directional indication of each audio channel, the focus region, and the focus amount; applying the gain adjustment factor to each of the audio channels; - constructing the processed spatial audio signal comprising the processed one or more multi-channel loudspeaker audio channels and / or the processed one or more audio object channels; The apparatus of claim 1 , further comprising:
12. The apparatus of claim 11 , wherein the multi-channel loudspeaker channels and / or the audio object channels each further comprise an audio channel distance indicator, and the calculated gain adjustment factor is further based on the audio channel distance indicator.
13. The apparatus of claim 11 , wherein the apparatus is further adapted to determine a respective default audio channel distance, and wherein the calculated gain adjustment factor is further based on the audio channel distance.
14. The obtained focus area is then transmitted to the device. Focus height, Focus radius, Focus distance, Depth of focus, Focus range, Focus diameter, or Focus Shape Characterizer, The apparatus of claim 2 , further comprising:
15. The device is further adapted to obtain focus input from a sensor arrangement including at least one directional sensor and at least one user input, the focus input comprising: an indication of the focus direction relative to the focus region based on a direction of the at least one direction sensor; an indication of the focus width based on the at least one user input; The apparatus of claim 2 , comprising:
16. The device of claim 1 , wherein the device is further adapted to obtain focus input comprising at least one user input, the focus input further comprising an indication of the amount of focus based on the at least one user input.
17. Obtaining a focus amount configured to define a focus area and an amount of focus; processing the spatial audio signal representing an audio scene to control attenuation of at least a portion of at least a portion of the spatial audio signal that is outside the acquired focus region in accordance with the focus amount relative to at least a portion of a portion of the spatial audio signal that is within the acquired focus region to generate a processed spatial audio signal that represents a modified audio scene; outputting the processed spatial audio signal, wherein the modified audio scene includes the attenuation of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region according to the focus amount relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; A method comprising:
18. The method of claim 17 , wherein obtaining the focus region includes obtaining at least one of a focus direction and a focus width.
19. Processing the spatial audio signal comprises: increasing the attenuation of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; or reducing the attenuation of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; 18. The method of claim 17, comprising:
20. Processing the spatial audio signal comprises: - decreasing or increasing the relative sound level of at least a portion of the at least one portion of the spatial audio signal that is outside the acquired focus region relative to at least a portion of the portion of the spatial audio signal that is within the acquired focus region; or decreasing or increasing the relative sound level according to the focus amount; 18. The method of claim 17, comprising at least one of:
21. 20. A non-transitory computer-readable medium comprising instructions that, when executed by an apparatus, cause the apparatus to perform the method of claim 17.