Sound field related rendering

The system addresses the limitation of fixed audio focus shapes by allowing users to customize spatial audio playback through adjustable focus parameters, improving the audio experience by emphasizing desired sound sources and ambient audio.

JP7764254B2Active Publication Date: 2025-11-05NOKIA TECHNOLOGIES OY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021573579
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-11
Filing Date
2020-06-03
Publication Date
2025-11-05
Estimated Expiration
2040-06-03

AI Technical Summary

Technical Problem

Existing spatial audio playback systems lack the ability to effectively control the shape of the audio focus region, which is crucial for users who want to prioritize specific sound sources or ambient audio based on their viewing direction and preferences.

Method used

A system that processes spatial audio signals to control the relative emphasis of audio within a defined focus shape by adjusting parameters such as focus direction, width, height, radius, distance, and depth, allowing for customizable audio focus regions.

Benefits of technology

Enables users to selectively emphasize audio within a specified focus area, enhancing the audio experience by allowing personalized audio focus shaping based on user preferences and viewing directions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764254000017
    Figure 0007764254000017
  • Figure 0007764254000018
    Figure 0007764254000018
  • Figure 0007764254000019
    Figure 0007764254000019
Patent Text Reader

Abstract

Apparatus and method for sound field related audio representation and rendering. [Solution] An apparatus for spatial audio reproduction, comprising means configured to obtain at least one focus parameter configured to define a focus shape, process a spatial audio signal representing an audio scene to control the relative emphasis of at least a portion of the spatial audio signal within the focus shape relative to at least a portion of other portions of the spatial audio signal outside the focus shape, generate a processed spatial audio signal representing a modified audio scene, and output the processed spatial audio signal, wherein the modified audio scene enables relative emphasis of at least a portion of a portion of the spatial audio signal within the focus shape relative to at least a portion of other portions of the spatial audio signal outside the focus shape.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for sound field related audio representation and rendering, but is not limited to audio representation for audio decoders. [Background technology]

[0002] Spatial audio playback is known for presenting media with multiple viewing directions. Examples of this playback include (at least) a head-mounted display (or head-mounted phone) that can track head orientation, or a non-head-mounted phone screen that can track view direction by changing the phone's position / orientation, or with any user interface gesture, or on a surrounding screen.

[0003] Video related to "media with multiple viewing directions" may be video that has a substantially wider viewing angle than traditional video, such as 360-degree video, 180-degree video, etc. Traditional video is video content that is typically displayed in its entirety on a screen, without the option (or particular need) to change the viewing direction.

[0004] Audio associated with video with multiple viewing directions can be presented over headphones or a surround loudspeaker setup where the viewing direction is tracked and affects the spatial audio reproduction.

[0005] The spatial audio associated with video with multiple viewing directions can come from spatial audio capture from a microphone array (e.g., an array attached to a VR camera like OZO, or a handheld mobile device) or from other sources such as a studio mix. The audio content can also be a mixture of multiple content types, such as microphone-captured audio and an added commentary track.

[0006] Spatial audio associated with video with multiple viewing directions can take various forms, including: Ambisonic signals (any order) consisting of spherical harmonic audio signal components. Spherical harmonics can be thought of as a set of spatially selective beam signals. Ambisonics are currently being used, for example, in YouTube® 360VR video services. The advantage of Ambisonics is its simple and well-defined signal representation. Surround speaker signals (e.g., 5.1). Currently, typical cinematic spatial audio is conveyed in this format. The advantage of surround loudspeaker signals is their simplicity and legacy compatibility. Some audio formats similar to surround loudspeaker signal formats contain audio objects that can be viewed as audio channels with time-varying positions. The positions can convey both the direction and distance of the audio object, or both. Some cutting-edge audio coding and spatial audio capture methods, such as parametric spatial audio, use this signal representation, i.e., audio signals for two audio channels in perceptually relevant frequency bands and associated spatial metadata. Spatial metadata essentially determines how the audio signal should be spatially reproduced at the receiver (e.g., in which directions at different frequencies). The advantages of parametric spatial audio are its versatility, quality, and the ability to use low bitrates for encoding. Summary of the Invention

[0007] According to a first aspect, there is provided an apparatus including means configured to obtain at least one focus parameter configured to define a focus shape, to process a spatial audio signal representing an audio scene to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape, to generate a processed spatial audio signal representing a modified audio scene, and to output the processed spatial audio signal, wherein the modified audio scene enables a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0008] The at least one focus parameter may be further configured to define a focus amount, and the means configured to process the spatial audio signal may be configured to process the spatial audio signal in accordance with the focus amount and to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0009] The means configured to process the spatial audio signal may be configured to increase the relative emphasis, or decrease the relative emphasis, of at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.

[0010] The means configured to process the spatial audio signal may be configured to increase or decrease the relative sound level of at least a part of the spatial audio signal within the focus shape relative to at least a part of another part of the spatial audio signal outside the focus shape.

[0011] The means configured to process the spatial audio signal may be configured to increase or decrease the relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some other portions of the spatial audio signal outside the focus shape according to the focus amount.

[0012] The means may be configured to obtain playback control information for controlling at least one aspect of outputting the processed spatial audio signal, and the means configured to output the processed spatial audio signal may be configured to perform one of: processing the processed spatial audio signal representing the modified audio scene to generate an output spatial audio signal in accordance with the playback control information; and, prior to the means configured to process the spatial audio signal representing the audio scene, processing the spatial audio signal in accordance with the playback control information to generate a processed spatial audio signal representing the modified audio scene and output the processed spatial audio signal as the output spatial audio signal.

[0013] The spatial audio signal and the processed spatial audio signal may comprise respective Ambisonic signals, and the means configured to process the spatial audio signal to generate the processed spatial audio signal may be configured to: transform, for one or more frequency subbands, the Ambisonic signal associated with the spatial audio signal into a set of beam signals of a defined pattern; generate a set of modified beam signals based on the set of beam signals, a focus shape, and a focus amount; and transform the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.

[0014] The defined pattern may consist of a defined number of beams equally spaced over a plane or over a volume.

[0015] The spatial audio signal and the processed spatial audio signal may be composed of respective higher-order Ambisonic signals.

[0016] The spatial audio signal and the processed spatial audio signal can be composed of a subset of Ambisonic signal components of any order.

[0017] The spatial audio signal and the processed spatial audio signal may comprise respective parametric spatial audio signals, which may comprise one or more audio channels and spatial metadata, where the spatial metadata may include respective direction indicators, energy ratio parameters, and potentially distance indicators for a plurality of frequency subbands. The means configured to process the input spatial audio signal to generate the processed spatial audio signal may be configured to: calculate spectral adjustment coefficients for one or more frequency subbands based on the spatial metadata and the focus shape and focus amount; apply the spectral adjustment coefficients to one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels; calculate respective modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, focus amount, and spatial metadata; and configure the processed spatial audio signal consisting of the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.

[0018] The spatial audio signal and the processed spatial audio signal may include multi-channel loudspeaker channels and / or audio object channels. The means configured to process the spatial audio signal into the processed spatial audio signal may be configured to calculate gain adjustment factors based on the respective audio channel directional indications, focus shapes, and focus amounts, apply the gain adjustment factors to the respective audio channels, and produce a processed spatial audio signal including one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels.

[0019] The multi-channel speaker channels and / or audio object channels may further include respective audio channel distance indicators, and the calculation gain adjustment factor may be further based on the audio channel distance indicators.

[0020] The means may be further configured to determine a default respective audio channel distance, and the computing gain adjustment factor may be further configured based on the audio channel distance.

[0021] The at least one focus parameter configured to define the focus shape can include at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, and a focus shape characterizer.

[0022] The means may be further configured to obtain focus input from a sensor arrangement comprising at least one directional sensor and at least one user input, wherein the focus input may further include an indication of a focus direction of the focus shape based on a direction of the at least one directional sensor and an indication of a focus width based on the at least one user input, and the focus input may further include an indication of a focus amount based on the at least one user input.

[0023] According to a second aspect, there is provided a method comprising the steps of obtaining at least one focus parameter configured to define a focus shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; and outputting the processed spatial audio signal, wherein the modified audio scene enables a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0024] The at least one focus parameter may be further configured to define a focus amount, and processing the spatial audio signal may include processing the spatial audio signal to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape in accordance with the focus amount.

[0025] Processing the spatial audio signal may include increasing or decreasing the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.

[0026] Processing the spatial audio signal may include increasing or decreasing the relative sound level in at least some of the portions of the spatial audio signal in the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.

[0027] Processing the spatial audio signal may include increasing or decreasing a relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape according to the focus amount.

[0028] The method may include obtaining playback control information for controlling at least one aspect of outputting the processed spatial audio signal, wherein outputting the processed spatial audio signal may include performing one of the following steps: processing the processed spatial audio signal representing the modified audio scene to generate an output spatial audio signal in accordance with the playback control information; and processing the spatial audio signal in accordance with the playback control information before the means configured to process the spatial audio signal representing the audio scene to generate a processed spatial audio signal representing the modified audio scene, and outputting the processed spatial audio signal as the output spatial audio signal.

[0029] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and processing the spatial audio signal to generate the processed spatial audio signal may include converting, for one or more frequency subbands, an Ambisonic signal associated with the spatial audio signal into a set of beam signals of a defined pattern; generating a set of modified beam signals based on the set of beam signals, a focus shape, and a focus amount; and converting the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.

[0030] The defined pattern may consist of a defined number of beams equally spaced over a plane or over a volume.

[0031] The spatial audio signal and the processed spatial audio signal may be composed of respective higher-order Ambisonic signals.

[0032] The spatial audio signal and the processed spatial audio signal can be composed of a subset of Ambisonic signal components of any order.

[0033] The spatial audio signal and the processed spatial audio signal may include respective parametric spatial audio signals, which may include one or more audio channels and spatial metadata, where the spatial metadata may include respective direction indicators, energy ratio parameters, and potentially distance indicators for a plurality of frequency subbands. Processing the input spatial audio signal to generate the processed spatial audio signal may include calculating spectral adjustment coefficients for one or more frequency subbands based on the spatial metadata and a focus shape and a focus amount, applying the spectral adjustment coefficients to one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels, calculating respective modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, focus amount, and spatial metadata, and configuring the processed spatial audio signal to include the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.

[0034] The spatial audio signal and the processed spatial audio signal may include multi-channel loudspeaker channels and / or audio object channels, and processing the spatial audio signal into the processed spatial audio signal may include calculating gain adjustment factors based on respective audio channel directional indications, focus shapes, and focus amounts, applying the gain adjustment factors to the respective audio channels, and configuring the processed spatial audio signal to include one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels.

[0035] The multi-channel speaker channels and / or audio object channels may further include respective audio channel distance indicators, and the computing gain adjustment factor may further be performed based on the audio channel distance indicators.

[0036] The method may further include determining a default respective audio channel distance, and the computing gain adjustment factor may be further determined based on the audio channel distance. The at least one focus parameter configured to define the focus shape may include at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, and a focus shape characterizer.

[0037] The method may further include obtaining focus input from a sensor arrangement comprising at least one directional sensor and at least one user input, wherein the focus input may include an indication of a focus direction of the focus shape based on a direction of the at least one directional sensor and an indication of a focus width based on the at least one user input.

[0038] The focus input may further include an indication of the amount of focus based on at least one user input.

[0039] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code configured, using the at least one processor, to cause the apparatus to at least: obtain at least one focus parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to control emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; and output the processed spatial audio signal, the modified audio scene enabling relative emphasis in portions of the spatial audio signal within at least some of the focus shapes compared to at least some of the other portions of the spatial audio signal outside the focus shape.

[0040] The at least one focus parameter may be further configured to define a focus amount, and the device adapted to process the spatial audio signal may be adapted to process the spatial audio signal in accordance with the focus amount and further to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape. The device adapted to process the spatial audio signal may be adapted to increase the relative emphasis of, or decrease the relative emphasis of, at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.

[0041] An apparatus adapted to process a spatial audio signal may be adapted to increase or decrease the relative sound level in at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape.

[0042] An apparatus adapted to process a spatial audio signal may be adapted to increase or decrease the relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape according to the focus amount.

[0043] The device may be adapted to obtain playback control information for controlling at least one aspect of outputting the processed spatial audio signal, and the device adapted to output the processed spatial audio signal may be adapted to perform one of the following steps: processing the processed spatial audio signal representing the modified audio scene to generate an output spatial audio signal in accordance with the playback control information; processing the spatial audio signal in accordance with the playback control information to generate a processed spatial audio signal representing the modified audio scene before the means configured to process the spatial audio signal representing the audio scene; and outputting the processed spatial audio signal as the output spatial audio signal.

[0044] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and the device for processing the spatial audio signal to generate the processed spatial audio signal may be configured to: convert, for one or more frequency subbands, the Ambisonic signals associated with the spatial audio signal into a set of beam signals of a defined pattern; generate a set of modified beam signals based on the set of beam signals, focus shape, and focus amount; and convert the modified beam signals to generate the modified Ambisonic signals associated with the processed spatial audio signal.

[0045] The defined pattern may consist of a defined number of beams equally spaced over a plane or over a volume.

[0046] The spatial audio signal and the processed spatial audio signal may be composed of respective higher-order Ambisonic signals.

[0047] The spatial audio signal and the processed spatial audio signal can be composed of a subset of Ambisonic signal components of any order.

[0048] The spatial audio signal and the processed spatial audio signal may comprise respective parametric spatial audio signals, which may comprise one or more audio channels and spatial metadata, which may comprise respective direction indications, energy ratio parameters and potentially distance indications for a plurality of frequency subbands, and an apparatus adapted to process an input spatial audio signal to generate a processed spatial audio signal, wherein: 1) the spatial audio signal may comprise respective direction indications for some of the plurality of frequency bands of the plurality of frequency bands; 2) the spatial audio signal may comprise a plurality of direction indications for some of the plurality of frequency bands of the plurality of frequency bands; 3) the spatial metadata may comprise a plurality of direction indications for some of the plurality of frequency bands of the plurality of frequency bands. The system may perform the steps of: calculating spectral adjustment coefficients for one or more frequency subbands based on spatial metadata, which may include respective directional indicators for some frequency bands, and the focus shape and the focus amount; applying the spectral adjustment coefficients to one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels; calculating respective modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, the focus amount, and the spatial metadata; and constructing a processed spatial audio signal consisting of the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.

[0049] The spatial audio signal and the processed spatial audio signal may include multi-channel loudspeaker channels and / or audio object channels, and the device for processing the spatial audio signal into a processed spatial audio signal may perform the steps of calculating gain adjustment factors based on the respective audio channel directional indications, focus shapes, and focus amounts, applying the gain adjustment factors to the respective audio channels, and configuring the processed spatial audio signal including one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels.

[0050] The multi-channel speaker channels and / or audio object channels may further include respective audio channel distance indicators, and the computing gain adjustment factor may be further determined based on the audio channel distance indicators. The device may be further caused to determine default respective audio channel distances, and the computing gain adjustment factor may be further determined based on the audio channel distances. The at least one focus parameter configured to define the focus shape may include at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, and a focus shape characterizer.

[0051] The device may further be caused to obtain focus input from a sensor configuration comprising at least one directional sensor and at least one user input, and the focus input may include an indication of a focus direction of the focus shape based on a direction of the at least one directional sensor, and an indication of a focus width based on the at least one user input.

[0052] The focus input may further include an indication of the amount of focus based on at least one user input.

[0053] According to a fourth aspect, there is provided an apparatus comprising: a focus parameter acquisition circuit configured to acquire at least one focus parameter configured to define a focus shape; a spatial audio signal processing circuit configured to process a spatial audio signal representing an audio scene to control a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape to generate a processed spatial audio signal representing a modified audio scene; and an output control circuit configured to output the processed spatial audio signal, wherein the modified audio scene enables a relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0054] According to a fifth aspect, there is provided a computer program comprising instructions (or a computer-readable medium comprising program instructions) to cause an apparatus to at least: obtain at least one focus parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to control relative emphasis in at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape to generate a processed spatial audio signal representing a modified audio scene; and output the processed spatial audio signal, wherein the modified audio scene enables relative emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0055] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus to at least: obtain at least one focus parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene to control emphasis relative to at least some of the portions of the spatial audio signal within the focus shape compared to at least some of the other portions of the spatial audio signal outside the focus shape; and output the processed spatial audio signal, wherein the modified audio scene enables relative emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0056] According to a seventh aspect, there is provided an apparatus comprising: means for obtaining at least one focus parameter configured to define a focus shape; means for processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to control emphasis in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; and means for outputting the processed spatial audio signal, wherein the modified audio scene enables relative emphasis of at least some of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

[0057] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions for causing an apparatus to perform at least the steps of obtaining at least one focus parameter configured to define a focus shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene to control emphasis in at least some portions of the spatial audio signal within the focus shape relative to at least some other portions of the spatial audio signal outside the focus shape; and outputting the processed spatial audio signal, wherein the modified audio scene enables relative emphasis in at least some portions of the spatial audio signal within the focus shape relative to at least some other portions of the spatial audio signal outside the focus shape. An apparatus comprising means for performing the actions of the above-described method. An apparatus configured to perform the actions of the above-described method. A computer program comprising program instructions for causing a computer to perform the above-described method. A computer program product stored on the medium can cause an apparatus to perform the methods described herein.

[0058] The electronic device may include an apparatus as described herein.

[0059] The chipset may consist of the devices described herein.

[0060] SUMMARY OF THE INVENTION Embodiments of the present invention aim to solve problems associated with the state of the art. [Brief explanation of the drawings]

[0061] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings, in which: [Figure 1a] 1a and 1b show an exemplary sound scene illustrating an audio focus region or area. [Figure 1b]1a and 1b show an exemplary sound scene illustrating an audio focus region or area. [Figure 2a] 2a and 2b schematically illustrate an exemplary playback device and method of operating the playback device according to some embodiments. [Figure 2b] 2a and 2b schematically illustrate an exemplary playback device and method of operating the playback device according to some embodiments. [Figure 3] FIG. 3 is a diagram illustrating schematic diagrams of spherical harmonic patterns and selected subsets of these spherical harmonic patterns as applied in some embodiments. [Figure 4] FIG. 4 shows a schematic representation of the beam pattern corresponding to the Ambisonic signal and the converted beam signal aligned with an exemplary focus direction of 20 degrees. [Figure 5a] 5a and 5b schematically illustrate an exemplary focus processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operating the exemplary focus processor according to some embodiments. [Figure 5b] 5a and 5b schematically illustrate an exemplary focus processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operating the exemplary focus processor according to some embodiments. [Figure 6] FIG. 6 is a diagram showing a typical processing state in an example where the focus direction is 20 degrees and the width is 45 degrees. [Figure 7] FIG. 7 is a visualization diagram that shows a schematic representation of the processing for a further example where the focus direction is minus 90 degrees and the width is 90 degrees. [Figure 8a] 8A and 8B are diagrams that schematically illustrate the example focus processor shown in FIG. 2A having a parametric spatial audio signal input and a method of operating the example focus processor, according to some embodiments. [Figure 8b]8A and 8B are diagrams that schematically illustrate the example focus processor shown in FIG. 2A having a parametric spatial audio signal input and a method of operating the example focus processor, according to some embodiments. [Figure 9a] 9a and 9b are diagrams illustrating the exemplary focus processor shown in FIG. 2a with multi-channel and / or audio object audio signal input and a method of operating the exemplary focus processor in accordance with some embodiments. [Figure 9b] 9a and 9b are diagrams illustrating the exemplary focus processor shown in FIG. 2a with multi-channel and / or audio object audio signal input and a method of operating the exemplary focus processor in accordance with some embodiments. [Figure 10] FIG. 10 illustrates an exemplary focus width determination based on focus distance and radius inputs, according to some embodiments. [Figure 11a] 11a and 11b schematically illustrate an exemplary playback processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operation of the exemplary playback processor according to some embodiments. [Figure 11b] 11a and 11b schematically illustrate an exemplary playback processor as shown in FIG. 2a having a high-order Ambisonic audio signal input and a method of operation of the exemplary playback processor according to some embodiments. [Figure 12a] 12a and 12b are diagrams that schematically illustrate an exemplary playback processor as shown in FIG. 2a having a parametric spatial audio signal input according to some embodiments, and a method of operating the exemplary playback processor. [Figure 12b]12a and 12b are diagrams that schematically illustrate an exemplary playback processor as shown in FIG. 2a having a parametric spatial audio signal input according to some embodiments, and a method of operating the exemplary playback processor. [Figure 13] FIG. 13 illustrates an exemplary implementation of some embodiments. [Figure 14] FIG. 14 illustrates an example controller for controlling focus direction, focus amount, and focus width, according to some embodiments. [Figure 15] FIG. 15 illustrates an example processing output based on processing a high-order Ambisonics audio signal according to some embodiments. [Figure 16] FIG. 16 shows an exemplary apparatus suitable for implementing the illustrated apparatus. DETAILED DESCRIPTION OF THE INVENTION

[0062] In the following, preferred apparatus and possible mechanisms for providing efficient rendering and reproduction of spatial audio signals are described in more detail.

[0063] Previous examples of spatial audio signal playback allowed users to control the focus direction and amount. However, in some situations, such control of focus direction / amount may not be sufficient. In some situations, it may be desirable for a user with a control interface to be able to control the focus shape. A sound field may have many different characteristics, such as ambient sound as well as multiple dominant sound sources in a particular viewing direction. Some users may prefer to hear specific characteristics of the sound field, while other users may prefer to hear alternative characteristics of the sound field depending on the desired viewing direction. It is understood that such playback audio depends on one or more preferences and is configurable based on user-related preferences. A desired performance from a playback device is to configure spatial audio playback so that focus can be controlled to various shapes or regions (e.g., narrow, wide, shallow, deep, near, far).

[0064] As an example, audio content of interest may exist within a sector (or cone or another spatial span or range) rather than simply in one direction. Specifically, controlling the spatial span of focus may be useful. Figures 1a and 1b, described below, illustrate what a user is intended to perceive when listening to a reproduced spatial audio signal. For example, as illustrated in Figure 1a, a source of interest may be present on one side of the user and a distracting source may be present on the other side of the user. Figure 1a shows a user 101 positioned with a defined orientation. Within the audio scene, there is a source of interest 105, such as a speaker in a theater play, that is within a desired focus region 103 defined by a focus direction and width. Additionally, there may be an audience or other ambient audio content 107 that is outside the view direction, such as behind the view direction.

[0065] Additionally, the user may wish to change the width of the sector over time, for example initially focusing on all sources in the play by keeping the focus sector relatively wide (as shown in Figure 1a), and then focusing on a particular source by narrowing the focus sector.

[0066] As another example, desired or interesting audio content may be at a distance (relative to the listener or to another location). For example, there may be an undesirable or uninteresting audio source at a distance in one direction, and a desirable or interesting audio source at another distance in the same direction (or approximately the same direction). This is illustrated in FIG. 1b. FIG. 1b shows a user 101 positioned at a defined orientation in an audio scene with a source of interest 105, such as a talker, around a table within a desired focus region 103 defined by a center position and a radius. Additionally, there may be other ambient audio content, such as environmental audio content 151 on the left, a music source audio component 155, and other speaker audio content 153 beyond the source of interest that is outside the desired focus region. In such an embodiment, the audio focus region or shape is determined by the center focus position and the focus radius.

[0067] Thus, the embodiments as discussed herein seek to provide control of focus shape (in addition to focus direction and amount). The concepts as discussed with respect to the embodiments described herein relate to spatial audio reproduction in media playback with multiple viewing directions by providing control of audio focus shape where the audio scene on the controlled audio focus shape may change but the signal format may remain the same.

[0068] In embodiments, at least one focus shape parameter corresponding to a selectable direction is provided by adjusting any (or a combination of two or all) of the following parameters: focus width, focus height, focus radius, focus distance, and focus depth corresponding to the selected direction. This parameter set in some embodiments is comprised of parameters that define an arbitrary shape.

[0069] In some embodiments, spatial audio signal processing can be performed by obtaining a spatial audio signal associated with media having multiple viewing directions, obtaining focus direction and amount parameters, modifying the spatial audio signal to have at least one desired focus characteristic, modifying the spatial audio signal to have the desired focus characteristic, and playing back (using headphones or loudspeakers) the modified spatial audio signal.

[0070] The resulting spatial audio signal may be in a parametric spatial audio format, such as an Ambisonic signal, a loudspeaker signal, or spatial metadata associated with a set of audio channels.

[0071] The focus shape, in some embodiments, may depend on which parameters are available. For example, if only direction, width, and height are present, the shape may be an ellipsoidal cone-shaped volume. As another example, if only distance and depth are present, the focus shape may be a hollow sphere. If width / height and / or depth are not present, they may be assumed to have some default values. Furthermore, in some embodiments, any focus shape may be used.

[0072] The amount of focus may, in some embodiments, determine the "degree" or how much focus is applied. For example, the focus may be between 0% and 100%, where 0% means keeping the original sound scene unchanged and 100% means maximum focus on the desired spatial shape.

[0073] In some embodiments, different users may desire to have different focus characteristics, and the original spatial audio signal may be modified and played individually for each user based on their individual preferences.

[0074] 2a illustrates a block diagram of some components and / or entities of a spatial audio processing device 250 according to one example. It will be understood that the two separate steps (focus processor + playback processor) illustrated in this figure and described in further detail below may be implemented as an integrated process or, in some examples, in the reverse order described herein (where the playback processor operation then follows the focus processor operation). The spatial audio processing device 250 comprises an audio focus processor 201 configured to receive an input audio signal and further focus parameters 202 and to derive, based on the input audio signal 200, an audio signal having a focus sound component 204 depending on the focus parameters 202 (which may include focus direction, focus amount, focus height, focus radius, focus distance, and focus depth). In some embodiments, the device may be configured to obtain a focus shape, where the focus shape includes at least one focus parameter (which may be configured to define the focus shape). The spatial audio processing device 250 may further include an audio playback processor 207 configured to receive the focus sound component 204 and playback control information 206, and configured to derive an output audio signal 208 in a predetermined audio format based on the audio signal having the focus sound component, further depending on the playback control information 206 serving to control at least one aspect of the processing of the spatial audio signal having the focus sound component in the audio playback processor 207. The playback control information 206 may include an indication of the playback direction (or playback orientation) and / or an indication of an applicable loudspeaker configuration. In view of the above-mentioned methods for processing spatial audio signals, the audio focus processor 201 may be arranged to perform an aspect of processing the spatial audio signal by modifying the audio scene to control emphasis in at least a portion of the spatial audio signal in the received focus region according to the received focus amount.The audio playback processor 207 may output the processed spatial audio signals based on the observed direction and / or position as a modified audio scene, the modified audio scene demonstrating emphasis according to the amount of focus received for at least the portion of the spatial audio signals in the focus region.

[0075] In the illustration of FIG. 2a, the input audio signal, the audio signal with the focus sound component, and the output audio signal are each provided as a respective spatial audio signal in a predefined spatial audio format. Therefore, these signals may be referred to as the input spatial audio signal, the spatial audio signal with the focus sound component, and the output spatial audio signal, respectively. Along the aforementioned lines, typically, a spatial audio signal conveys an audio scene that includes both one or more directional sound sources at each specific position in the audio scene and the ambience of the audio scene. However, in some scenarios, a spatial audio scene may include one or more directional sound sources without ambience, or ambience without directional sound sources. In this regard, the spatial audio signal includes information conveying one or more directional sound components representing distinct sound sources having a fixed position within the audio scene (e.g., a fixed direction of arrival and a fixed relative intensity relative to the listening point) and / or environmental sound components representing environmental sounds within the audio scene. It should be noted that dividing an audio scene into directional sound component(s) and ambient components is generally only a representation or approximation, but real sound scenes may contain more complex features, such as wide sound sources and coherent acoustic reflections. However, even with such complex acoustic features, conceptualizing an audio scene as a combination of direct and ambient components is typically a fair representation or approximation, at least in a perceptual sense.

[0076] Generally, the input audio signal and the audio signal having the sound pickup component are provided in the same predefined spatial format, but the output audio signal can be provided in the same spatial format applied to the input audio signal (and the audio signal having the sound pickup component), or a different predefined spatial format can be adopted for the output audio signal. The spatial audio format of the output audio signal is selected taking into account the characteristics of the sound reproduction hardware applied to reproduce the output audio signal. Generally, the input audio signal may be provided in a first predefined spatial audio format, and the output audio signal can be provided in a second predefined spatial audio format. Non-limiting examples of spatial audio formats suitable for use as the first and / or second spatial audio format are Ambisonics, surround loudspeaker signals according to a predefined loudspeaker configuration, and predefined parametric spatial audio formats. More detailed, non-limiting examples of the use of these spatial audio formats within the framework of the spatial audio processing device 250 as the first and / or second spatial audio format are provided later in this disclosure.

[0077] The spatial audio processor 250 is typically applied to process the input spatial audio signal 200 as a sequence of input frames into a respective sequence of output frames, each input (output) frame containing a respective segment of a digital audio signal for each channel of the input (output) spatial audio signal, provided as a respective time series of input (output) samples at a predetermined sampling frequency. In some embodiments, the input signal to the spatial audio processor 250 may be in an encoded form, such as AAC or AAC+embedded metadata. In such embodiments, the encoded audio input may first be decoded. Similarly, in some embodiments, the output from the spatial audio processor 250 may be encoded in any suitable manner.

[0078] In a typical example, the spatial audio processing device 250 employs a fixed, predetermined frame length, with each frame consisting of L samples for each channel of the input spatial audio signal, corresponding to a corresponding duration in time at a predetermined sampling frequency. As an example in this regard, the fixed frame length may be 20 milliseconds (ms), resulting in frames of L=160, L=320, L=640, and L=960 samples per channel at sampling frequencies of 8, 16, 32, or 48 kHz, respectively. The frames may be non-overlapping or partially overlapping, depending on whether the processor applies filter banks and how these filter banks are configured. However, these values ​​serve as non-limiting examples, and different frame lengths and / or sampling frequencies may be employed instead, depending on, for example, the desired audio bandwidth, the desired framing delay, and / or the available processing capacity.

[0079] In the spatial audio processing device 250, focus refers to a user-selectable spatial region of interest. The focus may be, for example, a direction, distance, radius, or arc of the overall audio scene. Another example is the focus region where a (directional) sound source of interest is currently located. In the former scenario, the user-selectable focus typically indicates a region that remains constant or does not change frequently, since the focus prevails in a particular spatial region, while in the latter scenario, the user-selected focus may change more frequently, since the focus is set on a particular sound source that may (or may not) change its position / shape / size in the audio scene over time. In one example, the focus can be defined, for example, as an azimuth angle defining the spatial direction of interest relative to a first predefined reference direction, and / or as an elevation angle defining the spatial direction of interest relative to a second predefined reference direction, and / or as a shape and / or distance and / or radius or shape parameter.

[0080] The functionality described above with reference to the components of the spatial audio processing device 250 may be provided, for example, according to a method 260 illustrated by the flowchart depicted in Fig. 2b. Method 260 may be provided, for example, by an apparatus arranged to implement the spatial audio processing system 250 described in this disclosure through numerous examples. Method 260 functions as a method for processing input spatial audio signals representing an audio scene into output spatial audio signals representing a modified audio scene. Method 260 comprises receiving an indication of a focus region and an indication of focus intensity, as shown in block 261.

[0081] The method 260 further comprises processing the input spatial audio signals into intermediate spatial audio signals representing a modified audio scene in which the relative levels of sounds coming from the focus region are modified according to the focus intensity, as shown in block 263.

[0082] Method 260 further comprises receiving playback control information that controls the processing of the intermediate spatial signals into output spatial audio signals, as shown in block 265. The playback control information may, for example, define at least one of a playback direction (e.g., listening direction or viewing direction) or a loudspeaker configuration for the output spatial audio signals.

[0083] The method 260 further includes processing the intermediate spatial audio signals into the output spatial audio signals according to the playback control information, as shown in block 267 .

[0084] The method 260 can be varied in a number of ways, for example, according to examples of the functionality of each of the components of the spatial audio processor 250 provided above and below.

[0085] In some embodiments, the input to the spatial audio processor 250 is an Ambisonic signal. The device can be configured to receive (and the method can be applied to) any order of Ambisonic signals. However, because first-order Ambisonic (FOA) signals have fairly broad spatial selectivity (specifically, first-order directivity), it is illustratively the case that higher-order Ambisonic (HOA) signals, which have greater spatial selectivity, are more suitable for finer control of the focus shape. In particular, in the following examples, the method and device are configured to receive third-order Ambisonic audio signals.

[0086] A third-order Ambisonic audio signal has a total of 16 beam pattern signals (in 3D). However, for simplicity in the following examples, only the seven more "horizontal" Ambisonic components (i.e., audio signals) are considered here, as shown in Figure 3, to illustrate the implementation of the focus shape parameters. For example, Figure 3 shows the zeroth-order spherical harmonic pattern 301, the first-order spherical harmonic pattern 303, the second-order spherical harmonic pattern 305, and the third-order spherical harmonic pattern 307. Figure 3 also shows subsets 309 and 311 of the more "horizontal" spherical harmonic patterns up to the third-order spherical harmonic pattern.

[0087] With reference to FIG. 5a, an exemplary Ambisonic signal x HOA Illustrated is a focus processor 550 configured to receive (t) 500 and focus direction 502. As noted above, the input to focus processor 550 in this example is a subset third-order Ambisonic signal, e.g., subsets 309 and 311. Also, in the following, a third-order Ambisonic signal x HOA For simplicity, the signal x(t) 500 is expressed as HOA. The signal x(t) arriving from the horizontal direction θ, where t is the discrete sample index, is expressed as follows:

number

[0088] In some embodiments, focus processor 550 comprises matrix processor 501. Matrix processor 501, in some embodiments, is configured to convert Ambisonic (HOA) signal 500 (corresponding to an Ambisonic or spherical harmonic pattern) into a set of seven equally spaced horizontal beam signals (corresponding to a beam pattern), which in some embodiments is represented by a transformation matrix T(θ f ) and θ f is the focus direction 502 parameter.

number

number

number

[0089] For example, θ f = 20 degrees, the transformed signal x cThe beam pattern corresponding to (t) 504 and the beam pattern corresponding to the original HOA signal are shown in Figure 4. Figure 4 shows an example of a beam pattern corresponding to an Ambisonic signal in the upper row 401 and a beam signal with a focus direction converted at 20 degrees in the lower row 403. The converted audio signal can then be output to a spatial beam (based on focus parameters) processor 503.

[0090] The focus processor 550 may further include a spatial beam (based on focus parameters) processor 503. The spatial beam processor 503 processes the converted Ambisonic signal x from the matrix processor 501. c (t) 504 and is further configured to receive focus amount and width focus parameters 508 .

[0091] The spatial beam processor 503 then processes the spatial beam signal x c (t) 504 to generate the processed or modified spatial beam signal x' c (t) 506 is based on the focus amount and the shape parameters 508. The processed or modified spatial beam signal x' c (t) 506 can then be output to a further matrix processor 505. The spatial beam processor 503 is configured to perform various processing methods based on the type of focus shape parameters. In this exemplary embodiment, the focus parameters are focus direction, focus width, and focus amount. The focus amount can be determined as a value a ranging between 0...1, where 1 indicates maximum focus. The focus width θ w (determined as the angle from the focus direction to the edge of the focus arc) is also a variable or controllable parameter. The spatial beam signal is

number

number

[0092] In this example, beam x c Note that (t) is formulated so that the first beam points in the focus direction and the second beam points in the focus direction +p. As a result, the matrix I(θ w When applying a), the beam far from the focus direction will be attenuated according to the focus width parameter.

[0093] The focus processor 201 further comprises a matrix processor 505. The further matrix processor 505 converts the processed or modified spatial beam signal x' c (t) 506. The result of inversely transforming the focus direction 502 is generated as a focus-processed HOA signal. f ) is invertible, so the inversion process is

number

[0094] Regarding Figure 6, the focus parameter is the maximum focus amount a=1, and the focus direction is θ f = 20 degrees, focus width θ w = 45 degrees. The top row 601 shows the focused transform domain signal x' c The lower row 603 shows the output signal x' HOA 7 shows the beam pattern corresponding to (t). With respect to FIG. 7, the focus parameter is the maximum focus amount a=1, and the focus direction parameter is θ f =-90 degrees, θ w= 90 degrees. The top row 701 shows the focused transform domain signal x' c The bottom row 703 shows the beam pattern corresponding to the output signal x' HOA (t) shows the corresponding beam pattern.

[0095] In the above examples, it has been shown that HOA processing is only considered on a set of more "horizontal" beam pattern signals. It is understood that these operations can be extended to 3D using a set of 3D beam patterns.

[0096] With reference to FIG. 5b, there is shown a flow diagram of the operation 560 of the HOA focus processor as shown in FIG. 5a.

[0097] The first operation is to receive the HOA audio signal (and focus parameters such as direction, width, amount or other control information) as shown in FIG. 5b by step 561.

[0098] The next operation is to generate the converted HOA audio signal into a beam signal, as shown in FIG. 5b at step 563.

[0099] After converting the HOA audio signal into a beam signal, the next operation is one of spatial beam processing, as shown in FIG. 5b by step 565.

[0100] The processed beamed audio signals are then converted back to HOA format as shown in FIG. 5b by step 567.

[0101] The processed HOA audio signal is then output by step 569 as shown in Figure 5b.

[0102] Referring to FIG. 8a, a focus processor configured to receive a parametric spatial audio signal as input is shown. The parametric spatial audio signal consists of an audio signal and spatial metadata, such as direction(s) in a frequency band and direct-to-total energy ratio(s). The structure and generation of parametric spatial audio signals are known, and their generation has been described from microphone arrays (e.g., mobile phones, VR cameras). Parametric spatial audio signals can also be generated from loudspeaker signals and Ambisonic signals. In some embodiments, the parametric spatial audio signal may be generated from an IVAS (Immersive Voice and Audio Services) audio stream, which can be decoded and demultiplexed into the form of spatial metadata and audio channels. A typical number of audio channels in such a parametric spatial audio stream is a two-audio-channel audio signal, but in some embodiments, the number of audio channels can be any number.

[0103] In these examples, the parametric information consists of depth / distance information, which may be implemented in six degrees of freedom (6DOF) playback, where distance metadata is used (along with other metadata) to determine how sound energy and direction should change in response to user movement.

[0104] Thus, in this example, each spatial metadata directional parameter is associated with both a direct-to-global energy ratio and a distance parameter. The estimation of distance parameters in the context of parametric spatial audio capture has been detailed in previous applications such as GB patent applications GB1710093.4 and GB1710085.0 and, for reasons of clarity, will not be discussed further.

[0105] A focus processor 850 configured to receive parametric (in this case 6DOF capable) spatial audio 800 is configured to use focus parameters (which in these examples are focus direction, amount, distance, and radius) to determine how much to attenuate or emphasize the direct and ambient components of the parametric spatial audio signal to enable the focus effect.

[0106] In the examples below, the methods (and formulas) are presented as being constant over time, but it should be understood that all parameters may vary over time.

[0107] In some embodiments, the focus processor comprises a ratio correction and spectral adjustment coefficient determiner 801 configured to receive focus parameters 808 and further spatial metadata consisting of direction 802, distance 822, and frequency band direct-to-total energy ratio 804.

[0108] The ratio corrector and the spectral adjustment coefficient determiner are configured to implement the focus shape as a sphere in 3D space by first transforming the focus direction and distance into a Cartesian coordinate system (a 3x1 yzx vector f):

number

[0109] Similarly, for each frequency band k, the direction and distance of the spatial metadata are

number

[0110] The units of the spatial metadata distance and focus distance parameters should be the same (e.g., both in meters, or in some other scale). The mutual distance value d(k) of f and m(k) can be simply formulated as follows:

number

[0111] This mutual distance value d(k) is then used in a gain function along with a focus amount parameter a ranging from 0 to 1 and a focus radius parameter dr (in the same units as d(k)). When focusing, an example gain formula is:

number

[0112] In practice, it may be desirable to smooth the focus gain function so that it transitions smoothly from high values ​​in focus regions to low values ​​in unfocused regions.

[0113] Then the new direct part value D(k) of the parametric spatial audio signal is

number

number

number

number

[0114] In the numerically undetermined case D(k) = A(k) = 0, r'(k) can also be set to 0.

[0115] The direction and distance parameters of the spatial metadata may not be modified by the metadata adjustment and spectral adjustment coefficient determiner 801 and modified and unmodified metadata output 810 in some embodiments.

[0116] The spatial processor 850 may include a spectral adjustment processor 803. The spectral adjustment processor 803 may be configured to receive an audio signal 806 and spectral adjustment coefficients 812. The audio signal may, in some embodiments, be in a time-frequency representation, or alternatively, be first transformed into the time-frequency domain for spectral adjustment processing. The output 814 may also be in the time-frequency domain, or may be transformed back to the time domain before output. The domains of the input and output are implementation dependent.

[0117] The spectral adjustment processing unit 803 can be configured to multiply, for each band k, the frequency bins (of the time-frequency transform) of all channels in band k by a spectral adjustment coefficient s(k), i.e., perform spectral adjustment. The multiplication (i.e., spectral correction) can be smoothed in time to avoid processing artifacts.

[0118] In other words, the processor is configured to modify the parametric spatial audio signal such that the spectral and spatial metadata of the signal is modified according to the focus parameters (in this case, focus direction, amount, distance, radius).

[0119] With reference to Figure 8b, there is shown a flow diagram 860 of the operation of a parametric spatial audio input processor such as that shown in Figure 8a.

[0120] The first operation is to receive a parametric spatial audio signal (and focus parameters or other control information) as shown in FIG. 8b via step 861.

[0121] The next operation is the modification of the parametric metadata and generation of spectral adjustment coefficients, as shown in FIG. 8b by step 863.

[0122] The next operation is to perform spectral adjustments on the audio signal, as shown in FIG. 8b at step 865.

[0123] The spectrally adjusted audio signal and the modified (and unmodified) metadata may then be output by step 867 as shown in Figure 8b.

[0124] 9a, a focus processor 950 is shown configured to receive a multi-channel or object audio signal as input 900. The focus processor in such an embodiment may comprise a focus gain determiner 901. The focus gain determiner 901 is configured to receive focus parameters 908 and channel / object position / direction information, which may be static or time-varying. The focus gain determiner 901 is configured to generate direct gain f(k) parameters, which are output as focus gains 912 for each channel, based on the focus parameters 908 and the channel / object position / direction information 902 from the input signal 900. In some embodiments, the directions of the channel signals are signaled, and in some embodiments, they are assumed. For example, when there are six channels, the directions can be assumed to be 5.1 audio channel directions. In some embodiments, there may be a lookup table used to determine the channel directions as a function of the number of channels.

[0125] For audio objects with direction and distance (i.e., position), the focus gain determiner 901 may utilize the same implementation process as described in the context of parametric audio processing to determine the direct gain f(k) 912 based on the spatial metadata and the focus parameters. In these embodiments, there is no filter bank; i.e., there is only one frequency band k.

[0126] The focus processor may also include a focus gain processor (for each channel) 903. The focus gain processor 903 is configured to receive a focus gain f(k) 912 for each audio channel and audio signal 906. The focus gain 912 may then be applied to the corresponding audio channel signal 906 (which may in some embodiments be further temporally smoothed). The output from the focus gain processor 903 may be a focus-processed audio channel audio signal 914.

[0127] In these examples, the channel direction / position information 902 is unchanged and provided as the channel direction / position information output 910 .

[0128] In some embodiments, if the input audio channels do not have distance information (e.g., loudspeaker or object sounds where the input is only direction and not distance), one option for processing such audio channels is to determine a fixed default distance for such signals and apply the same formula to determine f(k).

[0129] In some embodiments, determining the focus gain f(k) 912 for such an audio channel can be based on the angular difference between the focus direction and the direction of the audio channel. In some embodiments, this may first determine a focus width θ_w. For example, as shown in FIG. 10, the focus width θ_w 1005 may be determined trigonometrically using the focus distance 1001 and the focus radius 1003, where the focus width is generated by the angle of a right triangle with the hypotenuse formed by the focus distance 1001 and the opposite side formed by the focus radius 1003. The focus width is simply

number

[0130] With reference to FIG. 9b, there is shown a flow diagram 960 of the operation of the multi-channel / object audio input processing device shown in FIG. 9a.

[0131] The first operation is to receive the multi-channel / object audio signal (and focus parameters or other control information, and channel information such as direction / distance) as shown in FIG. 9b by step 961.

[0132] The next operation is to generate focus gain coefficients, as shown in Figure 9b by step 963. The next operation is to apply a focus gain to each channel audio signal, as shown in Figure 9b by step 965. The processed audio signal and unmodified channel direction (and distance) can then be output, as shown in Figure 9b by step 967.

[0133] In some embodiments, the focus shape may be defined using other parameters and other combinations of parameters, in which case the focus processor may be modified from the above example to use those parameters.

[0134] 11a, an example of a playback processor 1150 based on Ambisonic audio input is shown (e.g., it may be configured to receive the output from an example focus processor such as that shown in FIG. 5a). In these examples, the playback processor may comprise an Ambisonic rotation matrix processor 1101. The Ambisonic rotation matrix processor 1101 is configured to receive an Ambisonic signal having a focus process 1100 and a view direction 1102. The Ambisonic rotation matrix processor 1101 is configured to generate a rotation matrix based on the view direction parameter 1102. This may use any suitable method, such as that applied in head-tracked Ambisonic A / B inauralization in some embodiments (or more generally, such rotation of spherical harmonics is used in many fields, including outside of audio). This rotation matrix is ​​then applied to the Ambisonic audio signal. The result is a rotated Ambisonic signal with added focus 1104, which is output to an Ambisonic-to-binaural filter f1103. The Ambisonic to Binaural filter 1103 is configured to receive the focused rotated Ambisonic signal 1104 .

[0135] The Ambisonic-to-Binaural Filter 1103 can consist of a pre-formed 2xK matrix of finite impulse response (FIR) filters that are applied to the K Ambisonic signals to generate the 2 binaural signals 1106. The FIR filters may be generated by least-squares optimization with respect to a set of head-related impulse responses (HRIRs). An example of such a design procedure is to transform the HRIR dataset into frequency bins (e.g., by FFT) to obtain an HRTF dataset, and then determine, for each frequency bin, a complex-valued processing matrix that least-squares-approximates the available HRTF dataset at the data points of the HRTF dataset. When complex-valued matrices are so determined for all frequency bins, the result can be inverted (e.g., by inverse FFT) as a time-domain FIR filter. The FIR filters can also be windowed, for example, by using a Hann window.

[0136] There are many known methods that can be used to render Ambisonic signals into loudspeaker outputs. As an example, the Ambisonic signals can be linearly decoded to the target loudspeaker configuration. This can be applied if the order of the Ambisonic signals is sufficiently high, e.g., at least third order, and preferably fourth order. In such an implementation of linear decoding, an Ambisonic decoding matrix can be designed that, when applied to the Ambisonic signals (corresponding to Ambisonic beam patterns), produces loudspeaker signals corresponding to beam patterns that approximate, in a least-squares sense, a vector-base amplitude panning (VBAP) beam pattern appropriate for the target loudspeaker configuration. Processing the Ambisonic signals with such a designed Ambisonic decoding matrix can be configured to generate loudspeaker sound outputs. In such an embodiment, the playback processor is configured to receive information about the loudspeaker configuration.

[0137] With reference to FIG. 11b, there is shown a flow diagram 1160 of the operation of the Ambisonic input playback processor shown in FIG. 11a.

[0138] The first operation is to receive the focused Ambisonic audio signal (and view direction) as shown in FIG. 11b by step 1161.

[0139] The next operation is to generate a rotation matrix based on the view direction, as shown in FIG. 11b by step 1163.

[0140] The next operation is to apply a rotation matrix to the Ambisonic audio signal to generate a focused rotated Ambisonic audio signal, as shown in Figure 11b by step 1165.

[0141] The next operation is to convert the Ambisonic audio signal into a suitable audio output format, for example a binaural format (or a multi-channel audio format), as shown in FIG. 11b by step 1167.

[0142] Then, step 1169 outputs the output audio format as shown in FIG. 11b.

[0143] With reference to Figure 12a, an example playback processor 1250 based on parametric spatial audio input is shown (which may be configured to receive output from an example focus processor such as that shown in Figure 8a, for example).

[0144] In some embodiments, the playback processor comprises a filter bank 1201 configured to receive the audio signals in the audio channels 1200 and transform the audio channels into frequency bands (unless the input is already in the appropriate time-frequency domain). Examples of suitable filter banks include short-time Fourier transform (STFT) and complex quadrature mirror filter (QMF) banks. The time-frequency audio signal 1202 can be output to a parametric binaural synthesizer 1203.

[0145] In some embodiments, the playback processor consists of a parametric binaural synthesizer 1203 configured to receive the time-frequency audio signal 1202, the modified (and unmodified) metadata 1204, and also the view direction 1206 (or appropriate playback-related control or tracking information). In a 6DOF context, the user position can be provided along with the view direction parameter.

[0146] The parametric binaural synthesizer 1203 can be configured to implement any suitable known parametric spatial synthesis method configured to generate a binaural audio signal (frequency bands) 1208, as focus correction has already been performed on the signal and metadata prior to the parametric binauralization block. The binauralized time-frequency audio signal 1208 can then be passed to an inverse filterbank 1205. An embodiment may further be characterized in that the playback processor comprises an inverse filterbank 1205 configured to receive the binauralized time-frequency audio signal 1208 and to generate the inverse of the applied forward filterbank, thus generating a time-domain binauralized audio signal 1210 with focus characteristics suitable for playback via headphones (not shown in FIG. 12 a).

[0147] In some embodiments, the binaural audio signal output is replaced in loudspeaker channel audio signal output format from the parametric spatial audio signal using a suitable loudspeaker synthesis method. Any suitable approach may be used, for example, where the view direction parameters are replaced with information about the positions of the loudspeakers and the binaural processor is replaced with a loudspeaker processor based on a suitable known method.

[0148] With reference to Figure 12b, there is shown a flow diagram 1260 of the operation of a parametric spatial audio input reproduction processor such as that shown in Figure 12a.

[0149] The first operation is to receive a focused parametric spatial audio signal (and view direction or other playback related control or tracking information) as shown in FIG. 12b by step 1261.

[0150] The next operation is to time-frequency transform the audio signal, as shown in Figure 12b by step 1263. The next operation is to apply a parametric binaural (or loudspeaker channel type) processor based on the time-frequency transformed audio signal, the metadata and the listening direction (or other information), as shown in Figure 12b by step 1265.

[0151] The next operation is then to inverse transform the generated binaural or loudspeaker channel audio signals as shown in FIG. 12b by step 1267.

[0152] The output audio format is then output as shown in Figure 12b by step 1269. Considering the loudspeaker output of the playback processor when the audio signal is in the form of multi-channel audio and the focus processor 950 of Figure 9a is applied, in some embodiments the playback processor may configure a pass-through where the output loudspeaker configuration is the same as the format of the input signal.

[0153] In some embodiments where the output loudspeaker configuration differs from the input loudspeaker configuration, the playback processor can be configured with a vector-based amplitude panning (VBAP) processor. Each focused audio channel can then be processed using VBAP, a known amplitude panning technique, and spatially reproduced using the target loudspeaker configuration. In this way, the output audio signal is adapted to the output loudspeaker setting.

[0154] In some embodiments, the transformation from the first loudspeaker configuration to the second loudspeaker configuration may be performed using any suitable amplitude panning technique. For example, the amplitude panning technique may consist of deriving an N × M matrix of amplitude panning gains that define the transformation from the M channels of the first loudspeaker configuration to the N channels of the second loudspeaker configuration, and then multiplying the channels of the interspatial audio signal provided as a multichannel loudspeaker signal according to the first loudspeaker configuration with the matrix. The interspatial audio signal can be understood to be similar to an audio signal having a focused sound component 204, as shown in FIG. 2a. As a non-limiting example, the derivation of VBAP amplitude panning gains is described in Pulkki, Ville. "Virtual sound source positioning using vector base amplitude panning," Journal of the audio engineering society 45, no. 6 (1997), pp. 456-466.

[0155] For binaural output, any suitable binauralization of multi-channel loudspeaker signal formats (and / or objects) can be implemented. For example, typical binauralization may consist of processing audio channels with head-related transfer functions (HRTFs) and adding synthetic room reverberation to create the auditory impression of a listening room. The distance and direction (i.e., location) information of audio object sounds can be utilized for six-degree-of-freedom playback with user movement, for example, by employing the principles outlined in GB patent application GB1710085.0.

[0156] An example of a suitable apparatus for implementation is shown in Figure 13 in the form of a mobile phone or mobile device 1401 running suitable software 1403. Video may be played, for example, by attaching the mobile phone 1401 to a Daydream view type device (although for clarity the video processing will not be described here).

[0157] The audio bitstream obtainer 1423 is configured to obtain an audio bitstream 1424, for example received / obtained from storage. In some embodiments, the mobile device includes a decoder 1425 configured to receive compressed audio and decode it. An example of a decoder is an AAC decoder in the case of AAC decoding. The resulting decoded (e.g., Ambisonic, in which embodiments such as those shown in FIGS. 5a and 11a are implemented) audio signal 1426 may be forwarded to a focus processor 1427.

[0158] The mobile phone 1401 receives controller data 1400 from an external controller (e.g., via Bluetooth) at controller data receiver 1411 and passes the data to focus parameter (from controller data) determiner 1421. Focus parameter (from controller data) determiner 1421 determines focus parameters based, for example, on the orientation of the controller device and / or button events. The focus parameters may consist of any type of combination of proposed focus parameters (e.g., focus direction, focus amount, focus height, and focus width). The focus parameters 1422 are forwarded to focus processor 1427.

[0159] Based on the Ambisonic audio signal and the focus parameters, the focus processor 1427 is configured to create modified Ambisonic signals 1428 having desired focus characteristics. These modified Ambisonic signals 1428 are forwarded to the Ambisonic-to-binaural processor 1429. The Ambisonic-to-binaural processor 1429 is also configured to receive head orientation information 1404 from the orientation tracker 1413 of the mobile phone 1401. Based on the modified Ambisonic signals 1428 and the head direction information 1404, the Ambisonic-to-binaural processor 1429 is configured to create a head-tracked binaural signal 1430 that can be output from the mobile phone and played back using, for example, headphones.

[0160] 14 shows an exemplary device (or focus parameter control device) 1550 that may be configured to control or generate appropriate focus parameters such as focus direction, focus amount, and focus width. A user of the device may be configured to select a focus direction by pointing the controller in a desired direction 1509 and pressing a focus direction selection button 1505. The controller has an orientation tracker 1501, and orientation information may be used to determine the focus direction (e.g., in focus parameter (from controller data) determiner 1421, as shown in FIG. 13).

[0161] The focus direction in some embodiments can be visualized on a visual display while selecting the focus direction. In some embodiments, the focus amount can be controlled using focus amount buttons (shown as + and - in FIG. 14) 1507. Each press can increase or decrease the focus amount by, for example, 10 percentage points. The focus width can be controlled using focus width buttons (shown as + and - in FIG. 14) 1503. Each press can be configured to increase / decrease the focus width by a fixed amount, such as 10 degrees.

[0162] In some embodiments, the focus shape can be determined by drawing the desired shape with a controller (e.g., as depicted in FIG. 14). The user can initiate the drawing operation by pressing and holding the focus direction selection button, draw the desired shape with the controller, and finally accept the shape by releasing the button. A visual representation of the drawn shape may be displayed while drawing. The drawn shape can be converted into parameters for focus direction, focus height, and focus width. The focus amount may be selected with the "Focus Amount" button, as in the previous example.

[0163] In some embodiments, the focus controller, such as that shown in FIG. 14, is modified so that the "focus width" control is replaced with a "focus radius" control, allowing for complex, content-adaptive control of focus shapes. In such embodiments, the 360 ​​video may be implemented as part of an advanced virtual reality playback system in which the 360 ​​video is not only panoramic but also includes depth information (i.e., is effectively 3D video that can respond to user movement in six degrees of freedom). For example, the video content may be generated by computer graphics or by a VR video capture system that can detect visual depth and therefore enables 6DOF, similar to computer graphics.

[0164] For example, in a scene, there are two objects of interest (e.g., speakers). The user clicks "Select Focus Direction" for these two sound sources, and the visual display indicates to the user that these sound sources (not only the auditory sound sources but also the visual sound sources at a certain direction and distance) have been selected for audio focus. The user then selects the focus amount and focus radius parameters, and the focus radius indicates how much the auditory events from the sources of interest will be contained within the determined focus shape. During control adjustment, the focus radius may be shown as a visual sphere around the visual source of interest.

[0165] While the field of view may respond to user movement, sources may also move within the scene, and their locations are typically tracked visually. Thus, the focus shape, in this case, may be represented by two spheres in three-dimensional space, and then its overall shape can be adaptively changed by moving the spheres. This results in a complex focus shape that also has depth focus. Depending on the form of spatial audio, the focus shape can then be reproduced exactly (provided that spatial audio has reliable distance information) or approximated in other ways, for example, as illustrated above.

[0166] In some embodiments, it may be desirable to further specify the focus processing, for example, by determining the desired frequency range or spectral characteristics of the focused signal. In particular, it may be useful to emphasize the focused audio spectrum in the audio frequency band to improve intelligibility, for example, by attenuating low-frequency content (e.g., below 200 Hz), high-frequency content (e.g., above 8 kHz), and leaving particularly useful frequency bands associated with the audio.

[0167] It will be appreciated that the focus processed signal may be further processed with any known audio processing technique, such as automatic gain control or enhancement techniques (eg, bandwidth extension, noise suppression).

[0168] In some further embodiments, the focus parameters (including direction, amount, and at least one focus shape parameter) are generated by the content creator, and the parameters are transmitted along with the spatial audio signal. For example, the scene may be a VR video / audio recording of an unplugged music concert near the stage. The content creator may assume that a typical remote listener will want to determine a focus arc that extends toward the stage and also extends to the sides due to room acoustics, but will also want to remove at least some direct sound from the audience (behind the main direction of the VR camera). Therefore, a focus parameter track is added to the stream and can be set as the default rendering mode. However, because audience sounds are still present in the stream, some users may prefer to discard the focus processing and be able to play the full sound scene, including the audience sounds.

[0169] That is, instead of the user selecting the focus direction or shape, they can select pre-defined dynamic focus parameters. The presets may be fine-tuned by the content creator to suit the program, for example, by turning off focus at the end of each song and playing applause to the listener. The content creator can generate several expected preferred profiles for the focus parameters. This approach is beneficial because it only requires transmitting one spatial audio signal, but it is also possible to add different preferred profiles. Legacy players that do not have focus enabled can decode Ambisonic signals without the focus step.

[0170] In some further embodiments, the focus shape is controlled along with the visual zoom of a video with multiple viewing directions. Visual zoom can be conceptualized as a user controlling a set of virtual binoculars for a panoramic, 360, or 3D video. In such a use case, enabling the visual zoom function (e.g., setting at least 1.5x zoom) can also enable audio focus for spatial audio signals. Since the user is clearly interested in that direction, the focus amount can be set to a high value, e.g., 80%, and the focus width can be set to correspond to the arc of the visual field of the virtual binoculars. That is, increasing the visual zoom reduces the focus width. With the focus set to 80%, the user can hear the remaining spatial sound to some extent in the appropriate direction. This allows the user to hear the onset of interesting new content and know to turn off the visual zoom and look in the new direction of interest. Zoom processing can also be used in the context of an audio codec that enables such processing. An example of such a codec is MPEG-I.

[0171] A user in the above-described embodiment can use the present invention to control the focus shape in a general purpose manner.

[0172] An example of the processing output based on the described embodiment for a higher-order Ambisonics (HOA) signal is shown in Figure 15. This figure shows the spectrogram of a third-order HOA signal, with a talker at 0°, a sine wave at -90°, and white noise at 110°, as decoded from eight loudspeakers. The figure shows that narrowing the focus toward the talker reduces the relative energy of the sine wave and white noise, while a wider focus that includes both the talker and the sine wave significantly reduces the relative energy of only the white noise.

[0173] 16, an example of an electronic device that can be used as an analysis device or synthesis device is shown. The device may be any suitable electronic device or device. For example, in some embodiments, device 1700 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc.

[0174] In some embodiments, the device 1700 includes at least one processor or central processing unit 1707. The processor 1707 may be configured to execute various program code, such as the methods described herein.

[0175] In some embodiments, the apparatus 1700 comprises a memory 1711. In some embodiments, at least one processor 1707 is coupled to the memory 1711. The memory 1711 may be any suitable storage means. In some embodiments, the memory 1711 constitutes a program code section for storing program code executable by the processor 1707. Furthermore, in some embodiments, the memory 1711 may further comprise a storage data section for storing data, for example, data that has been processed or is to be processed according to embodiments as described herein. The implementation program code stored in the program code section and the data stored in the storage data section can be retrieved by the processor 1707 whenever needed via the memory-processor coupling.

[0176] In some embodiments, apparatus 1700 comprises a user interface 1705. User interface 1705, in some embodiments, may be coupled to processor 1707. In some embodiments, processor 1707 may control the operation of user interface 1705 and receive input from user interface 1705. In some embodiments, user interface 1705 may allow a user to input commands into device 1700, for example, via a keypad. In some embodiments, user interface 1705 may allow a user to retrieve information from device 1700. For example, user interface 1705 may include a display configured to display information from device 1700 to a user. User interface 1705, in some embodiments, may comprise a touchscreen or touch interface capable of both allowing information to be input into device 1700 and displaying information to a user of device 1700.

[0177] In some embodiments, apparatus 1700 includes an input / output port 1709. In some embodiments, input / output port 1709 comprises a transceiver. The transceiver in such embodiments may be coupled to processor 1707 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may, in some embodiments, be configured to communicate with other electronic devices or apparatuses via a wire or wired coupling.

[0178] The transceiver may communicate with the further device by any suitable known communication protocol, for example in some embodiments the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®, or an infrared data channel (IRDA).

[0179] The transceiver input / output port 1709 may be configured to receive signals and, in some embodiments, obtain focus parameters as described herein.

[0180] In some embodiments, device 1700 may be employed to generate appropriate audio signals using processor 1707 executing appropriate code. Input / output ports 1709 may be coupled to any suitable audio output, such as to a multi-channel speaker system and / or headphones (which may be head-tracked or non-tracked headphones).

[0181] In general, various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, although the invention is not limited thereto.

[0182] Although various aspects of the present invention may be illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing device, or any combination thereof.

[0183] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any block of the logic flow as shown may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. Software can be stored on physical media, such as memory chips or memory blocks implemented within a processor, magnetic media, such as hard disks or floppy disks, and optical media, such as DVDs and their data variants, CDs, etc.

[0184] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memory, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0185] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting logic-level designs into semiconductor circuit designs suitable for etching onto semiconductor substrates.

[0186] Programs such as those from Synopsys, Inc. of Mountain View, California, and Cadence Design, Inc. of San Jose, California, use established design rules and libraries of pre-stored design modules to automatically route wires and place components on semiconductor chips. Once the semiconductor circuit design is complete, the results may be sent in a standardized electronic format (e.g., Opus, GDSII) to a semiconductor manufacturing facility, or "fab," for fabrication.

[0187] The foregoing description provides a complete and informative description of exemplary embodiments of the present invention, by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant art in light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention, as defined by the appended claims.

Claims

1. 1. An apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the computer program code being configured to, using the at least one processor, cause the apparatus to perform at least: obtaining a spatial audio signal for audio playback, the spatial audio signal including ambient sound and at least one directional sound source; obtaining at least one focus parameter based on the spatial audio signal, the focus parameter being configured to define a focus shape; processing the spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to control emphasis relative to at least some of portions of the spatial audio signal within the focus shape relative to at least some other portions of the spatial audio signal outside the focus shape; outputting the processed spatial audio signal, the processed spatial audio signal including at least one of the at least one directional sound source or the environmental sound, and the modified audio scene enabling the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; An apparatus configured to cause a The at least one focus parameter is further configured to define an amount of focus; the apparatus is adapted to process the spatial audio signal to control a relative emphasis of at least a part of the portion of the spatial audio signal in the focus shape according to the focus amount. Device.

2. The step of processing the spatial audio signal includes: increasing the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; or reducing the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; Execute 10. The apparatus of claim 1.

3. 2. The apparatus of claim 1, wherein processing the spatial audio signal comprises causing the apparatus to increase or decrease a relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

4. The apparatus of claim 3 , wherein processing the spatial audio signal comprises increasing or decreasing the relative sound levels depending on the amount of focus.

5. the apparatus is further adapted to perform the step of obtaining playback control information for controlling at least one aspect of the processed spatial audio signal; the device is adapted to output the processed spatial audio signal; The step of processing the spatial audio signal further comprises the device: processing the processed spatial audio signal representing the modified audio scene to generate an output spatial audio signal in accordance with the playback control information; - processing the spatial audio signals according to playback control information before processing the spatial audio signals representing the modified audio scene; outputting the processed spatial audio signal as an output spatial audio signal; Execute one of the following:

10. The apparatus of claim 1.

6. the spatial audio signal and the processed spatial audio signal comprise respective Ambisonic signals; The step of processing the spatial audio signal may include, for one or more frequency subbands, causing the device to: converting an Ambisonic signal associated with the spatial audio signal into a set of beam signals of a defined pattern; generating a set of modified beam signals based on the set of beam signals, the focus shape, and the focus amount; or transforming the modified beam signals to generate modified Ambisonic signals related to the processed spatial audio signals; Execute 10. The apparatus of claim 1.

7. The apparatus of claim 6 , wherein the defined pattern comprises a defined number of beams spaced over a plane or volume.

8. The spatial audio signal and the processed spatial audio signal are Each higher-order Ambisonic signal, or a subset of first-order Ambisonic signal components, at least one of:

7. The apparatus of claim 6.

9. the spatial audio signal and the processed spatial audio signal comprise respective parametric spatial audio signals; the parametric spatial audio signal comprises one or more audio channels and spatial metadata; the spatial metadata includes a direction indication, an energy ratio parameter, and a distance indication for each of a plurality of frequency subbands; The step of processing the spatial audio signal further comprises the step of: calculating spectral adjustment factors for one or more frequency subbands based on the spatial metadata, the focus shape, and the focus amount; applying the spectral adjustment coefficients for the one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels; calculating respective modified energy ratio parameters associated with the one or more frequency subbands of the processed spatial audio signal based at least in part on the focus shape, the focus amount, and the spatial metadata; or constructing the processed spatial audio signal comprising the one or more processed audio channels, the modified energy ratio parameters, and the spatial metadata other than the energy ratio parameters; Execute 10. The apparatus of claim 1.

10. the spatial audio signals and the processed spatial audio signals include multi-channel loudspeaker channels and / or audio object channels; The step of processing the spatial audio signal further comprises the step of: calculating a gain adjustment factor based on each audio channel directional indication, the focus shape, and the focus amount; applying the gain adjustment factor to each of the audio channels; or constructing a processed spatial audio signal comprising one or more processed multi-channel loudspeaker audio channels and / or one or more processed audio object channels; Execute 10. The apparatus of claim 1.

11. the multi-channel loudspeaker audio channels and / or audio object channels further comprise respective audio channel distance indicators; the calculated gain adjustment factor is further based on an audio channel distance indicator; 11. The apparatus of claim 10.

12. The apparatus is further adapted to determine a default respective audio channel distance; The calculated gain adjustment factor is further based on the audio channel distance.

11. The apparatus of claim 10.

13. 10. The apparatus of claim 1, wherein the at least one focus parameter configured to define the focus shape comprises at least one of a focus direction, a focus width, a focus height, a focus radius, a focus distance, a focus depth, a focus range, a focus diameter, or a focus shape characterizer.

14. the device is further adapted to obtain focus input from a sensor arrangement comprising at least one direction sensor and at least one user input; wherein the focus input includes an indication of a focus direction of the focus shape based on a direction of the at least one direction sensor, and an indication of a focus width based on the at least one user input.

10. The apparatus of claim 1.

15. The device is further adapted to obtain focus input including at least one user input; the focus input further comprising an indication of an amount of focus based on the at least one user input.

10. The apparatus of claim 1.

16. 1. A method of an apparatus comprising: obtaining a spatial audio signal for audio playback, the spatial audio signal including ambient sound and at least one directional sound source; Based on the spatial audio signal, obtaining at least one focus parameter configured to define a focus shape; processing the spatial audio signal representing an audio scene to generate a processed spatial audio signal representing modified audio so as to control the relative emphasis of at least some portions of the spatial audio signal within the focus shape relative to at least some other portions of the spatial audio signal outside the focus shape; outputting the processed spatial audio signal, the processed spatial audio signal including at least one of the at least one directional sound source or the environmental sound, and the modified audio scene enabling relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; A method comprising: the at least one focus parameter defining an amount of focus; The method of claim 1, wherein processing the spatial audio signal comprises controlling relative emphasis of at least a portion of the portion of the spatial audio signal in the focus shape according to the focus amount.

17. The step of processing the spatial audio signal comprises: increasing the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; or reducing the relative emphasis of at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape.

17. The method of claim 16.

18. The step of processing the spatial audio signal comprises: increasing or decreasing a relative sound level in at least some of the portions of the spatial audio signal within the focus shape relative to at least some of the other portions of the spatial audio signal outside the focus shape; increasing or decreasing the relative sound level according to the amount of focus; at least one of:

17. The method of claim 16.

Citation Information

Patent Citations

  • Focusing on a portion of an audio scene for an audio signal

    EP2613564A2

  • Portable terminal, sound source position control method, and sound source position control program

    JP2013207759A

  • Apparatus and method for converting a first parametric spatial audio signal to a second parametric spatial audio signal.

    JP2013514696A

  • Sound collection system and sound emitting system

    JP2015198413A

  • Screen-related adaptation of higher-order ambisonic (hoa) content

    JP2018534853A