Sound field related rendering
The system processes spatial audio signals to modify focus shape and intensity based on user preferences, addressing the challenge of controlling audio emphasis in multi-directional media, resulting in enhanced audio quality and user engagement.
Patent Information
- Application Number
- CN202080043343.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-11
- Filing Date
- 2020-06-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-06-03
AI Technical Summary
In the spatial audio playback of multiple viewing directions, it is difficult to effectively control the focus shape and focus amount of the audio scene, resulting in poor user experience.
By defining focus parameters, the spatial audio signal is processed to generate a modified audio scene, controlling the aggravation of the audio signal within the focus shape, and achieving flexible control of the focus shape and focus amount of the audio scene.
It realizes precise focus on audio scenes, enhances the user's audio playback experience in multiple viewing directions, and meets the personalized preferences of different users.
Smart Images

Figure CN114009065B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to apparatuses and methods for audio representation and rendering related to a sound field, but non-exclusively to apparatuses and methods for audio representation for an audio decoder. Background Art
[0002] Spatial audio playback of media presented with multiple viewing directions is known. Examples of such playback include viewing visual content of such media, including playback on a head-mounted display (or head-mounted phone) with (at least) head orientation tracking; or on a non-head-mounted phone screen, where the viewing direction can be tracked by changing the position / orientation of the phone or by any user interface gesture; or on a surrounding screen.
[0003] Video associated with "media with multiple viewing directions" can be, for example, 360-degree video, 180-degree video, or other video with a much wider viewing angle than traditional video. Traditional video refers to video content that is typically displayed in its entirety on a screen without the option (or any particular need) to change the viewing direction.
[0004] Audio associated with video having multiple viewing directions can be presented on headphones, where the viewing direction is tracked and affects spatial audio playback; or can be presented with a surround speaker setup.
[0005] Spatial audio associated with video having multiple viewing directions can be sourced from spatial audio captured by a microphone array (e.g., an array mounted on a VR camera similar to OZO or a handheld mobile device), or from other sources such as a studio mix. The audio content can also be a mixture of several content types such as sounds captured by a microphone and an added narrator track.
[0006] Spatial audio associated with a video having multiple viewing directions can take various forms, such as: an Ambisonic signal (of any order) composed of spherical harmonic audio signal components. Spherical harmonics can be considered a set of spatially selective beam signals. Currently, Ambisonics is used, for example, in the YouTube 360VR video service. The advantage of Ambisonics is that it is a simple and well-defined signal representation; surround speaker signals, such as 5.1. Currently, the spatial audio of typical movies is transmitted in this form. The advantage of surround speaker signals is simplicity and legacy compatibility. Some audio formats similar to the surround speaker signal format include audio objects, which can be considered audio channels with time-varying positions. The position can inform both the direction and distance of the audio object, or just the direction; parametric spatial audio, such as two-channel audio signals and associated spatial metadata in the perceptually relevant frequency bands. Some state-of-the-art audio coding methods and spatial audio capture methods apply this signal representation. The spatial metadata essentially determines how the audio signal should be spatially reproduced at the receiver end (e.g., to those directions at different frequencies). The advantage of parametric spatial audio is its versatility, quality, and ability to use low-bitrate coding. SUMMARY OF THE INVENTION
[0007] According to a first aspect, there is provided an apparatus comprising components configured to perform the following operations: obtain at least one focusing parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape; and output the processed spatial audio signal, wherein the modified audio scene at least partially enables the relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0008] The at least one focusing parameter may further be configured to define a focus amount, and the component configured to process the spatial audio signal may be configured to: process the spatial audio signal so as to further control, according to the focus amount, the relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0009] A component configured to process a spatial audio signal may be configured to at least partially increase the relative emphasis of at least a portion of the spatial audio signal within a focused shape relative to at least other portions of the spatial audio signal outside the focused shape, or at least partially decrease the relative emphasis of at least a portion of the spatial audio signal within a focused shape relative to at least other portions of the spatial audio signal outside the focused shape.
[0010] A component configured to process a spatial audio signal may be configured to at least partially increase or decrease the relative sound level of at least a portion of the spatial audio signal within a focused shape relative to at least other portions of the spatial audio signal outside the focused shape.
[0011] A component configured to process a spatial audio signal may be configured to at least partially increase or decrease the relative sound level of at least a portion of the spatial audio signal within a focused shape relative to at least other portions of the spatial audio signal outside the focused shape according to a focusing amount.
[0012] The component may be configured to obtain reproduction control information to control at least one aspect of outputting the processed spatial audio signal, and wherein the component configured to output the processed spatial audio signal may be configured to perform one of the following: process the processed spatial audio signal representing a modified audio scene according to the reproduction control information to generate an output spatial audio signal; process the spatial audio signal according to the reproduction control information before a component configured to process the spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene and output the processed spatial audio signal as an output spatial audio signal.
[0013] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and wherein the component configured to process the spatial audio signal to generate the processed spatial audio signal may be configured to, for one or more frequency subbands, perform the following operations: convert the Ambisonic signal associated with the spatial audio signal into a set of beam signals in a defined pattern; generate a set of modified beam signals based on the set of beam signals, the focused shape, and the focusing amount; and convert the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.
[0014] The defined pattern may include a defined number of beams evenly spaced on a plane or in a volume.
[0015] The spatial audio signal and the processed spatial audio signal may include respective high-order Ambisonic signals.
[0016] The spatial audio signal and the processed spatial audio signal may include a subset of Ambisonic signal components of any order.
[0017] The spatial audio signal and the processed spatial audio signal may include corresponding parametric spatial audio signals, where the parametric spatial audio signal may include one or more audio channels and spatial metadata, where the spatial metadata may include corresponding direction indications, energy ratio parameters, and possibly distance indications for a plurality of frequency subbands, and where the component configured to process the input spatial audio signal to generate the processed spatial audio signal may be configured to: for one or more frequency subbands, calculate a spectral adjustment factor based on the spatial metadata, the focusing shape, and the focusing amount; apply the spectral adjustment factor to one or more frequency subbands of one or more audio channels to generate one or more processed audio channels; calculate corresponding modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based on the focusing shape, the focusing amount, and at least a portion of the spatial metadata; and compose the processed spatial audio signal, which includes one or more processed audio channels, the modified energy ratio parameters, and the spatial metadata other than the energy ratio parameter.
[0018] The spatial audio signal and the processed spatial audio signal may include multi-channel speaker channels and / or audio object channels, and where the component configured to process the spatial audio signal into the processed spatial audio signal may be configured to: calculate a gain adjustment factor based on the corresponding audio channel direction indication, the focusing shape, and the focusing amount; apply the gain adjustment factor to each audio channel; and compose the processed spatial audio signal, which includes one or more processed multi-channel speaker audio channels and / or one or more processed audio object channels.
[0019] The multi-channel speaker channels and / or audio object channels may further include corresponding audio channel distance indications, and where calculating the gain adjustment factor may further be based on the audio channel distance indication.
[0020] The component may further be configured to determine a default corresponding audio channel distance, and where calculating the gain adjustment factor may further be based on the audio channel distance.
[0021] At least one focusing parameter configured to define the focusing shape may include at least one of the following: focusing direction; focusing width; focusing height; focusing radius; focusing distance; focusing depth; focusing range; focusing diameter; and focusing shape characterizer.
[0022] The component may be further configured to obtain a focusing input from a sensor device including at least one orientation sensor and at least one user input, wherein the focusing input may include: an indication of a focusing direction for a focusing shape based on the orientation of at least one orientation sensor; and an indication of a focusing width based on at least one user input.
[0023] The focusing input may further include an indication of a focusing amount based on at least one user input.
[0024] According to a second aspect, there is provided a method including: obtaining at least one focusing parameter configured to define a focusing shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape; and outputting the processed spatial audio signal, wherein the modified audio scene at least partially enables the relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0025] The at least one focusing parameter may be further configured to define a focusing amount, and processing the spatial audio signal may include: processing the spatial audio signal so as to further control, according to the focusing amount, the relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0026] Processing the spatial audio signal may include: at least partially increasing the relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape, or at least partially decreasing the relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0027] Processing the spatial audio signal may include: at least partially increasing or decreasing the relative sound level of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0028] Processing the spatial audio signal may include: at least partially increasing or decreasing, according to the focusing amount, the relative sound level of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0029] The method may include: obtaining reproduction control information to control at least one aspect of outputting a processed spatial audio signal, and wherein outputting the processed spatial audio signal may include performing one of the following: processing a processed spatial audio signal representing a modified audio scene according to the reproduction control information to generate an output spatial audio signal; processing the spatial audio signal according to the reproduction control information before a component configured to process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene and output the processed spatial audio signal as an output spatial audio signal.
[0030] The spatial audio signal and the processed spatial audio signal may include respective Ambisonic signals, and wherein processing the spatial audio signal to generate the processed spatial audio signal may include, for one or more frequency subbands: converting the Ambisonic signal associated with the spatial audio signal into a set of beam signals in a defined pattern; generating a set of modified beam signals based on the set of beam signals, a focusing shape, and a focusing amount; and converting the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.
[0031] The defined pattern may include a defined number of beams evenly spaced on a plane or in a volume.
[0032] The spatial audio signal and the processed spatial audio signal may include respective higher-order Ambisonic signals.
[0033] The spatial audio signal and the processed spatial audio signal may include a subset of Ambisonic signal components of any order.
[0034] Spatial audio signals and processed spatial audio signals may include corresponding parametric spatial audio signals, wherein the parametric spatial audio signals may include one or more audio channels and spatial metadata, and wherein the spatial metadata may include corresponding direction indications, energy ratio parameters, and possibly distance indications for a plurality of frequency subbands. Processing an input spatial audio signal to generate a processed spatial audio signal may include: for one or more frequency subbands, calculating a spectral adjustment factor based on the spatial metadata, a focusing shape, and a focusing amount; applying the spectral adjustment factor to one or more frequency subbands of one or more audio channels to generate one or more processed audio channels; calculating corresponding modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based on the focusing shape, the focusing amount, and at least a portion of the spatial metadata; and composing the processed spatial audio signal, which includes one or more processed audio channels, the modified energy ratio parameters, and the spatial metadata other than the energy ratio parameter.
[0035] Spatial audio signals and processed spatial audio signals may include multi-channel speaker channels and / or audio object channels. Processing the spatial audio signal into a processed spatial audio signal may include: calculating a gain adjustment factor based on corresponding audio channel direction indications, a focusing shape, and a focusing amount; applying the gain adjustment factor to each audio channel; and composing the processed spatial audio signal, which includes one or more processed multi-channel speaker audio channels and / or one or more processed audio object channels.
[0036] The multi-channel speaker channels and / or audio object channels may also include corresponding audio channel distance indications, and wherein calculating the gain adjustment factor may further be based on the audio channel distance indication.
[0037] The method may also include determining a default corresponding audio channel distance, and wherein calculating the gain adjustment factor may further be based on the audio channel distance.
[0038] At least one focusing parameter configured to define the focusing shape may include at least one of the following: a focusing direction; a focusing width; a focusing height; a focusing radius; a focusing distance; a focusing depth; a focusing range; a focusing diameter; and a focusing shape characterizer.
[0039] The method may also include obtaining a focusing input from a sensor device including at least one direction sensor and at least one user input, wherein the focusing input may include: an indication of a focusing direction for the focusing shape based on at least one direction sensor direction; and an indication of a focusing width based on at least one user input.
[0040] The focus input may also include an indication of a focus amount based on at least one user input.
[0041] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: obtain at least one focus parameter configured to define a focus shape; process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape; and output the processed spatial audio signal, wherein the modified audio scene at least partially enables a relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0042] The at least one focus parameter may further be configured to define a focus amount, and the apparatus configured to process the spatial audio signal may be caused to: process the spatial audio signal so as to further, according to the focus amount, at least partially control a relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0043] The apparatus configured to process the spatial audio signal may be caused to: at least partially increase a relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape, or at least partially decrease a relative emphasis of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0044] The apparatus configured to process the spatial audio signal may be caused to: at least partially increase or decrease a relative sound level of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0045] The apparatus configured to process the spatial audio signal may be caused to: according to the focus amount, at least partially increase or decrease a relative sound level of a portion of the spatial audio signal within the focus shape relative to other portions of the spatial audio signal outside the focus shape.
[0046] The apparatus can obtain reproduction control information to control at least one aspect of the output of the processed spatial audio signal, and wherein the apparatus that is caused to output the processed spatial audio signal can be caused to perform one of the following: process the processed spatial audio signal representing the modified audio scene according to the reproduction control information to generate an output spatial audio signal; process the spatial audio signal according to the reproduction control information before a component configured to process the spatial audio signal representing the audio scene to generate a processed spatial audio signal representing the modified audio scene and output the processed spatial audio signal as the output spatial audio signal.
[0047] The spatial audio signal and the processed spatial audio signal can include respective Ambisonic signals, and wherein the apparatus that is caused to process the spatial audio signal to generate the processed spatial audio signal can be caused to, for one or more frequency subbands, perform the following operations: convert the Ambisonic signal associated with the spatial audio signal into a set of beam signals in a defined pattern; generate a set of modified beam signals based on the set of beam signals, a focusing shape, and a focusing amount; and convert the modified beam signals to generate a modified Ambisonic signal associated with the processed spatial audio signal.
[0048] The defined pattern can include a defined number of beams evenly spaced on a plane or in a volume.
[0049] The spatial audio signal and the processed spatial audio signal can include respective higher-order Ambisonic signals.
[0050] The spatial audio signal and the processed spatial audio signal can include a subset of Ambisonic signal components of any order.
[0051] A spatial audio signal and a processed spatial audio signal may include corresponding parametric spatial audio signals, where a parametric spatial audio signal may include one or more audio channels and spatial metadata, where the spatial metadata may include corresponding direction indications, energy ratio parameters, and possibly distance indications for a plurality of frequency subbands, and where the apparatus configured to process an input spatial audio signal to generate a processed spatial audio signal may be configured to: for one or more frequency subbands, calculate a spectral adjustment factor based on the spatial metadata, a focus shape, and a focus amount; apply the spectral adjustment factor to one or more frequency subbands of one or more audio channels to generate one or more processed audio channels; calculate corresponding modified energy ratio parameters associated with one or more frequency subbands of the processed spatial audio signal based on the focus shape, the focus amount, and at least a portion of the spatial metadata; and compose the processed spatial audio signal, the processed spatial audio signal including one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameters.
[0052] A spatial audio signal and a processed spatial audio signal may include multi-channel speaker channels and / or audio object channels, and where the apparatus configured to process the spatial audio signal into a processed spatial audio signal may be configured to: calculate a gain adjustment factor based on corresponding audio channel direction indications, a focus shape, and a focus amount; apply the gain adjustment factor to each audio channel; and compose the processed spatial audio signal, the processed spatial audio signal including one or more processed multi-channel speaker audio channels and / or one or more processed audio object channels.
[0053] The multi-channel speaker channels and / or audio object channels may further include corresponding audio channel distance indications, and where calculating the gain adjustment factor may further be based on the audio channel distance indications.
[0054] The apparatus may further be configured to determine default corresponding audio channel distances, and where calculating the gain adjustment factor may further be based on the audio channel distances.
[0055] At least one focus parameter configured to define a focus shape may include at least one of the following: a focus direction; a focus width; a focus height; a focus radius; a focus distance; a focus depth; a focus range; a focus diameter; and a focus shape characterizer.
[0056] The apparatus may further be configured to obtain a focus input from a sensor device including at least one direction sensor and at least one user input, where the focus input may include: an indication of a focus direction for the focus shape based on at least one direction sensor direction; and an indication of a focus width based on at least one user input.
[0057] The focus input may further include an indication of a focus amount based on at least one user input.
[0058] According to a fourth aspect, there is provided an apparatus, comprising: an obtaining circuit configured to obtain at least one focusing parameter configured to define a focusing shape; a spatial audio signal processing circuit configured to process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape; and an output control circuit configured to output the processed spatial audio signal, wherein the modified audio scene at least partially enables a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape.
[0059] According to a fifth aspect, there is provided a computer program [or a computer-readable medium including program instructions] comprising instructions for causing an apparatus to at least perform the following operations: obtaining at least one focusing parameter configured to define a focusing shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape; and outputting the processed spatial audio signal, wherein the modified audio scene at least partially enables a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape.
[0060] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus to at least perform the following operations: obtaining at least one focusing parameter configured to define a focusing shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape; and outputting the processed spatial audio signal, wherein the modified audio scene at least partially enables a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape.
[0061] According to a seventh aspect, there is provided an apparatus, comprising: means for obtaining at least one focusing parameter, wherein the at least one focusing parameter is configured to define a focusing shape; means for processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape; and means for outputting the processed spatial audio signal, wherein the modified audio scene at least partially enables a relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0062] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions for causing an apparatus to at least perform the following operations: obtaining at least one focusing parameter, the at least one focusing parameter being configured to define a focusing shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape; and outputting the processed spatial audio signal, wherein the modified audio scene at least partially enables a relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
[0063] An apparatus comprises means for performing the actions of the method as described above.
[0064] An apparatus is configured to perform the actions of the method as described above.
[0065] A computer program comprises program instructions for causing a computer to perform the method as described above.
[0066] A computer program product stored on a medium can cause an apparatus to perform the method described herein.
[0067] An electronic device may comprise an apparatus as described herein.
[0068] A chipset may comprise an apparatus as described herein.
[0069] Embodiments of the present application are intended to solve problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings, in which:
[0071] Figure 1a andFigure 1b shows an example sound scene showing an audio focus area or region;
[0072] Figure 2a and Figure 2b schematically shows an example playback device and a method for operating the playback device according to some embodiments;
[0073] Figure 3 shows a schematic diagram of spherical harmonic modes applied in some embodiments and a selected subset of these spherical harmonic modes;
[0074] Figure 4 schematically shows a beam pattern corresponding to an Ambisonic signal and a transformed beam signal aligned with an example focus direction of 20 degrees;
[0075] Figure 5a and Figure 5b schematically shows, according to some embodiments, an example focus processor having a high - order Ambisonic audio signal input as shown in Figure 2a and a method for operating the example focus processor;
[0076] Figure 6 schematically shows a visualization of the processing of an example focus direction of 20 degrees and a focus width of 45 degrees;
[0077] Figure 7 schematically shows a visualization of the processing of another example focus direction of - 90 degrees and a focus width of 90 degrees;
[0078] Figure 8a and Figure 8b schematically shows, according to some embodiments, an example focus processor having a parametric spatial audio signal input as shown in Figure 2a and a method for operating the example focus processor;
[0079] Figure 9a and Figure 9b schematically shows, according to some embodiments, an example focus processor having a multi - channel and / or audio object audio signal input as shown in Figure 2a and a method for operating the example focus processor;
[0080] Figure 10 shows an example focus width determination based on focus distance and radius input according to some embodiments;
[0081] Figure 11a and Figure 11b schematically shows, according to some embodiments, an example focus processor having a high - order Ambisonic audio signal input as shown in Figure 2aThe example reproduction processor shown in and a method of operating the example reproduction processor;
[0082] Figure 12a and Figure 12b Schematically shows an example reproduction processor as shown in and a method of operating the example reproduction processor with a parametric spatial audio signal input according to some embodiments; Figure 2a The example reproduction processor shown in and a method of operating the example reproduction processor;
[0083] Figure 13 Shows an example implementation of some embodiments;
[0084] Figure 14 Shows an example controller for controlling the focus direction, focus amount, and focus width according to some embodiments;
[0085] Figure 15 Shows an example processing output based on processing a high - order Ambisonic audio signal according to some embodiments;
[0086] Figure 16 Shows an example device suitable for implementing the shown apparatus. Detailed Description
[0087] Suitable apparatuses and possible mechanisms for providing effective rendering and playback of spatial audio signals are further described in detail below.
[0088] Previous examples of spatial audio signal playback have allowed users to control the focus direction and focus amount. However, in some cases, this control of the focus direction / amount may be insufficient. In some cases, it may be desirable to enable the user to control the shape of the focus using a control interface. In the sound field, there may be many different features, such as multiple dominant sound sources in certain viewing directions and ambient sounds. Some users may prefer to hear certain features of the sound field, while some other users may prefer to hear alternative features of the sound field, depending on the desired viewing direction. It can be understood that such playback audio depends on one or more preferences and can be configured based on user - related preferences. The desired performance of the playback apparatus is to configure the playback of spatial sound such that the focus can be controlled for various shapes or regions (e.g., narrow, wide, shallow, deep, near, far).
[0089] As an example, there may be audio content of interest within a sector region (or a cone or another spatial span or range) rather than just in one direction. Specifically, controlling the spatial span of the focus can be useful. The Figure 1a and Figure 1b show what a user expects to perceive when listening to a reproduced spatial audio signal. For example, there may be a source of interest on one side of the user, while there may be a distracting source on the other side of the user, as Figure 1aas shown in Figure 1a shows user 101 positioned in a defined orientation. There is a source of interest 105 within the audio scene, e.g., a speaker in a theater performance within a desired focus region 103 defined by a focus direction and width. Additionally, there may be audience or other ambient audio content 107 outside the viewing direction (such as behind the viewing direction).
[0090] Additionally, the user may wish to change the width of the sector region over time. For example, first focus on all sources in the theater performance by keeping the focus sector region relatively wide (as Figure 1a shown), and then focus on a specific source by narrowing the focus sector region.
[0091] As another example, the desired or interesting audio content may be at a certain distance (relative to the listener or relative to another location). For example, there may be an undesired or uninteresting audio source at a certain distance in a certain direction, and a desired or interesting audio source at another distance in the same direction (or nearly the same direction). This is shown in Figure 1b For example, Figure 1b shows user 101 positioned in a defined orientation within an audio scene with a source of interest 105, where the source of interest is, for example, a speaker around a table, which is within a desired focus region 103 defined by a central position and a radius. Additionally, there may be other ambient audio content, such as ambient audio content 151 on the left, a music source audio component 155, and other speaker audio content 153 outside the desired focus region beyond the source of interest. In such an embodiment, the audio focus region or shape is determined by a central focus position and a focus radius.
[0092] Thus, the embodiments discussed herein attempt to provide control of the focus shape (in addition to the focus direction and amount). The concepts discussed in connection with the embodiments described herein relate to: performing spatial audio reproduction in media playback with multiple viewing directions by providing control of the audio focus shape, where the audio scene changes on the controlled audio focus shape but the signal format may remain the same.
[0093] The embodiments provide at least one focus shape parameter corresponding to an optional direction by adjusting any one (or a combination of two or all) of the following parameters corresponding to the selected direction: focus width; focus height; focus radius; focus distance; and focus depth. In some embodiments, this set of parameters includes parameters defining any arbitrary shape.
[0094] In some embodiments, spatial audio signal processing may be performed by: obtaining a spatial audio signal associated with media having multiple viewing directions; obtaining a focus direction and amount parameter; obtaining at least one focus shape parameter; modifying the spatial audio signal to have a desired focusing characteristic; and reproducing the modified spatial audio signal (using headphones or speakers).
[0095] The obtained spatial audio signal may be, for example: an Ambisonic signal; a loudspeaker signal; a parametric spatial audio format such as a set of audio channels and associated spatial metadata.
[0096] In some embodiments, the focus shape may depend on which parameters are available. For example, if only direction, width, and height are available, the shape may be an elliptical cone type volume. As another example, if only distance and depth are available, the focus shape may be a hollow sphere. If width / height and / or depth are not available, they may be considered to have some default values. Additionally, in some embodiments, any focus shape may be used.
[0097] In some embodiments, the focus amount may determine the "degree" of focus or how much to focus. For example, the focus may range from 0% to 100%, where 0% means keeping the original sound scene unchanged, while 100% means focusing maximally on the desired spatial shape.
[0098] In some embodiments, different users may desire different focusing characteristics, and the original spatial audio signal may be modified and reproduced individually for each user based on their personal preferences.
[0099] Figure 2aA block diagram showing some components and / or entities of a spatial audio processing device 250 according to an example is presented. It can be understood that the two separate steps (focus processor + reproduction processor) shown in this figure and further detailed subsequently can be implemented as an integrated process, or in some examples, implemented in the reverse order as described herein (where the reproduction processor operation is followed by the focus processor operation). The spatial audio processing device 250 includes an audio focus processor 201, which is configured to receive an input audio signal and also receive a focus parameter 202; and based on the input audio signal 200 and according to the focus parameter 202 (which may include a focus direction; a focus amount; a focus height; a focus radius; a focus distance; and a focus depth), obtain an audio signal 204 having a focused sound component. In some embodiments, the device may be configured to obtain a focus shape, where the focus shape includes at least one focus parameter (which may be configured to define the focus shape). Additionally, the spatial audio processing device 250 may further include an audio reproduction processor 207, which is configured to receive the audio signal 204 having a focused sound component and reproduction control information 206, and is configured to obtain an output audio signal 208 in a predefined audio format based on the audio signal 204 having a focused sound component and further according to the reproduction control information 206, where the reproduction control information 206 is used to control at least one aspect related to processing the spatial audio signal having a focused component in the audio reproduction processor 207. The reproduction control information 206 may include an indication of a reproduction orientation (or reproduction direction) and / or an indication of an applicable speaker configuration. Considering the method for processing the above-mentioned spatial audio signal, the audio focus processor 201 may be set to implement an aspect of processing the spatial audio signal by modifying the audio scene so as to control the emphasis of at least a part of the spatial audio signal in the received focus area according to the received focus amount. The audio reproduction processor 207 may output the processed spatial audio signal as a modified audio scene based on the observed direction and / or position, where the modified audio scene shows an emphasis in the focus area and at least for the said part of the spatial audio signal according to the received focus amount.
[0100] In Figure 2aIn the illustrated example, each of the input audio signal, the audio signal having a focused sound component, and the output audio signal is provided as a respective spatial audio signal in a predefined spatial audio format. Accordingly, these signals may be referred to as the input spatial audio signal, the spatial audio signal having a focused sound component, and the output spatial audio signal, respectively. Along the lines described above, generally, spatial audio signal transmission involves an audio scene of one or more directional sound sources at respective positions in the audio scene and the environment of the audio scene. However, in some cases, the spatial audio scene may involve one or more directional sound sources without an environment or an environment without any directional sound sources. In this regard, the spatial audio signal includes information transmitting one or more directional sound components and / or ambient sound components, where the one or more directional sound components represent different sound sources having a certain position within the audio scene (e.g., a certain arrival direction and a certain relative intensity with respect to the listening point), and the ambient sound component represents the ambient sound within the audio scene. It should be noted that the division of the audio scene into directional sound components and ambient components is generally just a representation or approximation, and the actual sound scene may involve more complex features such as wide sources and coherent sound reflections. Nevertheless, even with such complex acoustic features, it is generally a reasonable representation or approximation to conceptually divide the audio scene into direct and ambient components at least in the perceptual sense.
[0101] Generally, the input audio signal and the audio signal having a focused sound component are provided in the same predefined spatial format, and the output audio signal may be provided in the same spatial format as that applied to the input audio signal (and the audio signal having a focused sound component), or a different predefined spatial format may be used for the output audio signal. The spatial audio format of the output audio signal is selected in view of the characteristics of the sound reproduction hardware to which the output audio signal is applied for playback. Generally, the input audio signal may be provided in a first predefined spatial audio format, and the output audio signal may be provided in a second predefined spatial audio format. Non-limiting examples of spatial audio formats suitable for use as the first and / or second spatial audio format include Ambisonics, surround speaker signals according to a predefined speaker configuration, and predefined parametric spatial audio formats. More detailed non-limiting examples of using these spatial audio formats as the first and / or second spatial audio format within the framework of the spatial audio processing device 250 are provided subsequently in this disclosure.
[0102] The spatial audio processing device 250 is generally applied to process an input spatial audio signal 200 as a sequence of input frames into a corresponding sequence of output frames, where each input (output) frame includes corresponding digital audio signal segments for each channel of the input (output) spatial audio signal, and is provided as a corresponding series of input (output) samples in time at a predefined sampling frequency. In some embodiments, the input signal of the spatial audio processing device 250 may have an encoded form, e.g., AAC or AAC+ with embedded metadata. In such embodiments, the encoded audio input may initially be decoded. Similarly, in some embodiments, the output from the spatial audio processing device 250 may be encoded in any suitable manner.
[0103] In a typical example, the spatial audio processing device 250 uses a fixed predefined frame length such that each frame includes corresponding L samples for each channel of the input spatial audio signal, and this fixed predefined frame length maps to a corresponding duration at the predefined sampling frequency. As an example in this regard, the fixed frame length may be 20 milliseconds (ms), which results in frames with L = 160, L = 320, L = 640, and L = 960 samples per channel at sampling frequencies of 8, 16, 32, or 48 kHz, respectively. These frames may be non - overlapping, or they may be partially overlapping, depending on whether the processor applies a filter bank and how the filter bank is configured. However, these values are used as non - limiting examples and different frame lengths and / or sampling frequencies different from these examples may be used instead, depending on, for example, the desired audio bandwidth, the desired framing delay, and / or the available processing power.
[0104] In the spatial audio processing device 250, focus refers to a user - selectable spatial region of interest. The focus can generally be, for example, a certain direction, distance, radius, arc of an audio scene. In another example, the focus region is the region where the current (directional) sound source of interest is located. In the former case, the user - selectable focus typically denotes a region that remains constant or changes infrequently, as the focus is mainly in a specific spatial region, while in the latter case, the user - selected focus can change more frequently, as the focus is set to a sound source that may (or may not) change its position / shape / size in the audio scene over time. In an example, the focus can be defined, for example, as the azimuth angle defining the spatial direction of interest relative to a first predefined reference direction, and / or as the elevation angle defining the spatial direction of interest relative to a second predefined reference direction, and / or as a shape and / or distance and / or radius or shape parameter.
[0105] It can be, for example, according to that in Figure 2bThe method 260 depicted in the flowchart provides the functions described above with reference to the components of the spatial audio processing apparatus 250. The method 260 can be provided, for example, by an apparatus configured to implement the spatial audio processing system 250 described via multiple examples in the present disclosure. The method 260 serves as a method for processing an input spatial audio signal representing an audio scene into an output spatial audio signal representing a modified audio scene. The method 260 includes receiving an indication of a focus region and an indication of a focus intensity, as shown in block 261. The method 260 further includes processing the input spatial audio signal into an intermediate spatial audio signal representing a modified audio scene, wherein the relative level of the sound arriving from the focus region is modified according to the focus intensity, as shown in block 263. The method 260 further includes receiving reproduction control information that controls the processing of the intermediate spatial signal into the output spatial audio signal, as shown in block 265. The reproduction control information can, for example, define at least one of a reproduction orientation (e.g., a listening direction or a viewing direction) for the output spatial audio signal or a speaker configuration. The method 260 further includes processing the intermediate spatial audio signal into the output spatial audio signal according to the reproduction control information, as shown in block 267.
[0106] The method 260 can be varied in a variety of ways, for example, according to the examples related to the corresponding functions of the components of the spatial audio processing apparatus 250 provided above and below.
[0107] In some embodiments, the input to the spatial audio processing apparatus 250 is an Ambisonic signal. The apparatus can be configured to receive (and the method can be applied to) Ambisonic signals of any order. However, since the first-order Ambisonic (FOA) signal is rather broad in terms of spatial selectivity (specifically, first-order directivity), higher-order Ambisonic (HOA) with higher spatial selectivity can be better used to illustrate the fine control of the focus shape. In particular, in the following examples, the method and apparatus are configured to receive a third-order Ambisonic audio signal.
[0108] The third-order Ambisonic audio signal has a total of 16 beam pattern signals (in 3D). However, as Figure 3 shown, for simplicity, the following examples only consider the more "horizontal" 7 Ambisonic components (in other words, audio signals) here to illustrate the implementation of the focus shape parameters. For example, Figure 3 shows the 0th-order spherical harmonic mode 301, the first-order spherical harmonic mode 303, the second-order spherical harmonic mode 305, and the third-order spherical harmonic mode 307. In addition, Figure 3 shows subsets 309 and 311 related to the more "horizontal" third-order spherical harmonic modes.
[0109] Regarding Figure 5a , a focusing processor 550 is shown, which is configured to receive an example Ambisonic signal x HOA (t) 500 and a focusing direction 502. In this example, the input to the focusing processor 550 is a subset of a third-order Ambisonic signal, e.g., subsets 309 and 311, as described above. For simplicity, the third-order Ambisonic signal x HOA (t) 500 is also described hereinafter as HOA. The signal x(t) arriving from the horizontal azimuth angle θ (where t is the discrete sample index) can be represented as an HOA signal by the following formula:
[0110]
[0111] where a(θ) is a vector of Ambisonic weights for the azimuth angle θ. As can be seen in this equation, the selected subset of Ambisonic modes can be defined in the horizontal plane by these very simple mathematical expressions.
[0112] In some embodiments, the focusing processor 550 includes a matrix processor 501. In some embodiments, the matrix processor 501 is configured to convert an Ambisonic (HOA) signal 500 (corresponding to an Ambisonic or spherical harmonic mode) into a set of beam signals (corresponding to beam modes) in seven uniformly spaced horizontal directions. In some embodiments, this can be represented by a transformation matrix T(θ f ), where θ f is the focusing direction 502 parameter:
[0113] x c (t) = T(θ f )x HOA (t)
[0114] where,
[0115]
[0116] where, Note that this transformation includes processing based on the focusing direction θ f 502 parameter such that the first mode is aligned with the focusing direction while the other modes are aligned with other symmetrically spaced directions.
[0117] For example, when θ f = 20 degrees, the beam pattern corresponding to the transformed signal x c (t) 504 and the beam pattern corresponding to the original HOA signal are shown in Figure 4 . Figure 4For example, the top row 401 and the bottom row 403 are shown. The top row 401 shows an example beam pattern corresponding to an Ambisonic signal, and the bottom row 403 shows the transformed beam signal whose focusing direction is at 20 degrees. Further, the transformed audio signal can be output to the spatial beam (based on the focusing parameter) processor 503.
[0118] The focusing processor 550 may also include a spatial beam (based on the focusing parameter) processor 503. The spatial beam processor 503 is configured to receive the transformed Ambisonic signal x c (t)504 from the matrix processor 501, and also receive the focusing amount and the width focusing parameter 508.
[0119] The spatial beam processor 503 is configured to further modify the spatial beam signal x c (t)504 based on the focusing amount and the shape parameter 508 to generate a processed or modified spatial beam signal x′ c (t)506. The processed or modified spatial beam signal x′ c (t)506 can then be output to another matrix processor 505. The spatial beam processor 503 is configured to implement various processing methods based on the type of the focusing shape parameter. In this example embodiment, the focusing parameters are the focusing direction, the focusing width, and the focusing amount. The focusing amount can be determined as a value a ranging from 0..1, where 1 indicates maximum focus. The focusing width θ w (determined as the angle from the focusing direction to the edge of the focusing arc) is also a variable or controllable parameter. The spatial beam signal can be generated by the following formula:
[0120] x′ c (t) = I(θ w , a)x c (t)
[0121] where I(θ w , a) is a diagonal matrix whose diagonal elements are determined as i(θ w , a), where,
[0122]
[0123] It should be noted that in this example, the beam x c (t) is formulated in such a way that the first beam points to the focusing direction, the second beam points to the focusing direction + p, and so on. Therefore, when the matrix I(θ w , a) is applied, the beams farther from the focusing direction will be attenuated according to the focusing width parameter.
[0124] The focusing processor 201 includes another matrix processor 505. The another matrix processor 505 is configured to receive the processed or modified spatial beam signal x′ c (t)506 and the focusing direction 502, and perform an inverse transformation on the result to generate the focused HOA signal. The transformation matrix T(θ f ) is invertible, and thus, the inverse processing can be expressed as:
[0125] x′ HOA (t) = T -1 (θ f )x′ c (t)
[0126] where x′ HOA (t) is the focused HOA output 510.
[0127] Regarding Figure 6 shows an example where the focusing parameter has a maximum focusing amount a = 1, the focusing direction is θ f = 20 degrees, and has a focusing width θ w = 45 degrees. The top row 601 shows the beam pattern corresponding to the processed transform domain signal x′ c (t) and the focus effect region, and the bottom row 603 shows the beam pattern corresponding to the output signal x′ HOA (t). Regarding Figure 7 shows an example where the focusing parameter has a maximum focusing amount a = 1, the focusing direction parameter is θ f = -90 degrees, and θ w = 90 degrees. The top row 701 shows the beam pattern corresponding to the processed transform domain signal x′ c (t), and the bottom row 703 shows the beam pattern corresponding to the output signal x′ HOA (t).
[0128] In the above example, HOA processing is only considered when showing a set of more "horizontal" beam pattern signals. It can be understood that these operations can be extended to 3D using a set of beam patterns in 3D.
[0129] Regarding Figure 5b , shows a flowchart of the operation 560 of the HOA focusing processor as shown in Figure 5a .
[0130] As shown in step 561 of Figure 5b , the initial operation is to receive the HOA audio signal (and focusing parameters such as direction, width, amount, or other control information).
[0131] AsFigure 5b As shown in step 563, the next operation is to generate a beam signal from the transformed HOA audio signal.
[0132] As Figure 5b shown in step 565, after the HOA audio signal has been transformed into a beam signal, the next operation is to perform spatial beam processing.
[0133] Furthermore, as Figure 5b shown in step 567, the processed beam audio signal is then inverse-transformed back to the HOA format.
[0134] Furthermore, as Figure 5b shown in step 569, the processed HOA audio signal is output.
[0135] Regarding Figure 8a , a focus processor is shown, which is configured to receive a parametric spatial audio signal as input. The parametric spatial audio signal includes an audio signal and spatial metadata, such as direction and direct-to-total energy ratio in a frequency band. The structure and generation of the parametric spatial audio signal are known, and its generation has been described from microphone arrays (e.g., mobile phones, VR cameras). In addition, the parametric spatial audio signal can also be generated from speaker signals and Ambisonic signals. In some embodiments, the parametric spatial audio signal can be generated from an IVAS (immersive voice and audio service) audio stream, which can be decoded and demultiplexed into the form of spatial metadata and audio channels. The typical number of audio channels in such a parametric spatial audio stream is two audio channel audio signals; however, in some embodiments, the number of audio channels can be any number of audio channels.
[0136] In these examples, the parametric information includes depth / distance information, which can be implemented in 6 degrees of freedom (6DOF) reproduction. In 6DOF, the distance metadata (along with other metadata) is used to determine how the energy and direction of the sound should change according to user movement.
[0137] Thus, in this example, each spatial metadata direction parameter is associated with both the direct-to-total energy ratio and the distance parameter. The estimation of the distance parameter in the context of parametric spatial audio capture has been detailed in earlier applications such as GB patent applications GB1710093.4 and GB1710085.0, but is not further explored for reasons of clarity.
[0138] A focusing processor 850 configured to receive parametric (in this case supporting 6DOF) spatial audio 800 is configured to use focusing parameters (which in these examples are the focusing direction, amount, distance, and radius) to determine how much the directional and ambient components of the parametric spatial audio signal should be attenuated or boosted to enable a focusing effect.
[0139] In the following examples, the methods (and formulas) are presented without changing over time, but it should be understood that all parameters can change over time.
[0140] In some embodiments, the focusing processor includes a ratio modifier and a spectral adjustment factor determiner 801, which is configured to receive focusing parameters 808 and also receive spatial metadata consisting of the direction 802, distance 822, and directional-to-total energy ratio 804 in a frequency band.
[0141] The ratio modifier and spectral adjustment factor determiner are configured to implement the focusing shape as a sphere in 3D space. First, the focusing direction and distance are transformed into the Cartesian coordinate system (3x1 y - z - x vector f) by the following formula:
[0142]
[0143] Similarly, in each frequency band k, the spatial metadata direction and distance are transformed into the Cartesian coordinate system (3x1 y - z - x vector m(k)):
[0144]
[0145] The units of the spatial metadata distance and the focusing distance parameter should be the same (e.g., both in meters, or both in any other scale). The mutual distance value d(k) between F and m(k) can be simply formulated as:
[0146] d(k) = |f - m(k)|
[0147] Which here refers to the length of the vector (f - m(k)).
[0148] Furthermore, the mutual distance value d(k) is used in the gain function, together with the focusing amount parameter a between 0..1 and the focusing radius parameter d r (using the same unit as d(k)). When performing focusing, the example gain formula is:
[0149]
[0150] Where c is a gain constant for focusing, for example, with a value of 4.
[0151] In practice, it may be necessary to smooth the above function so that the focusing gain function smoothly transitions from a high value in the focused region to a low value in the non-focused region.
[0152] Furthermore, the new direct portion value D(k) of the parametric spatial audio signal can be formulated as:
[0153] D(k) = r(k) * f(k)
[0154] where r(k) is the ratio of the directionality to the total energy in frequency band k. The new ambient portion value A(k) can be formulated as:
[0155] A(k) = (1 - r(k)) * (1 - a)
[0156] The spectral correction factor s(k) output 812 to the spectral adjustment processor 803 is further formulated based on the overall modification of the acoustic energy, in other words:
[0157]
[0158] The new modified ratio of directionality to total energy parameter r′(k) is further formulated to replace r(k) in the spatial metadata:
[0159]
[0160] In case of numerical uncertainty, D(k) = A(k) = 0, and thus r′(k) can also be set to zero.
[0161] In some embodiments, the direction and distance parameters of the spatial metadata may not be modified by the metadata adjustment and spectral adjustment factor determiner 801, and a modified and unmodified metadata output 810 is obtained.
[0162] The spatial processor 850 may include a spectral adjustment processor 803. The spectral adjustment processor 803 may be configured to receive an audio signal 806 and a spectral adjustment factor 812. In some embodiments, the audio signals may be in a time-frequency representation, or alternatively they are first transformed into the time-frequency domain for spectral adjustment processing. The output 814 may also be in the time-frequency domain, or is inverse-transformed into the time domain before output. The domain of the input and output depends on the implementation.
[0163] The spectral adjustment processor 803 may be configured to multiply the (time-frequency transformed) frequency bins of all channels within frequency band k by the spectral adjustment factor s(k) for each frequency band k. In other words, spectral adjustment is performed. The multiplication (i.e., spectral correction) may be smoothed over time to avoid processing artifacts.
[0164] In other words, the processor is configured to modify the spectrum of the signal as well as the spatial metadata such that the process results in a parameterized spatial audio signal that has been modified according to focusing parameters (in this case: focusing direction, number, distance, radius).
[0165] Regarding Figure 8b , a flowchart 860 of the operation of a parameterized spatial audio input processor as shown in Figure 8a is shown.
[0166] As shown in step 861 in Figure 8b , the initial operation is to receive a parameterized spatial audio signal (and focusing parameters or other control information).
[0167] As shown in step 863 in Figure 8b , the next operation is to modify the parameterized metadata and generate a spectral adjustment factor.
[0168] As shown in step 865 in Figure 8b , the next operation is to perform spectral adjustment on the audio signal.
[0169] Furthermore, as shown in step 867 in Figure 8b , a spectrally adjusted audio signal and modified (and unmodified) metadata can be output.
[0170] Regarding Figure 9a , a focusing processor 950 configured to receive a multi-channel or object audio signal as input 900 is shown. In such an example, the focusing processor may include a focusing gain determiner 901. The focusing gain determiner 901 is configured to receive focusing parameters 908 and channel / object position / orientation information, which may be static or time-varying. The focusing gain determiner 901 is configured to generate a direct gain f(k) parameter for each channel based on the focusing parameters 908 and the channel / object position / orientation information 902 from the input signal 900, which is output as a focusing gain 912. In some embodiments, the channel signal directions are signaled, and in some embodiments, they are assumed. For example, when there are 6 channels, the directions may be assumed to be 5.1 audio channel directions. In some embodiments, there may be a look-up table that is used to determine the channel directions based on the number of channels.
[0171] For an audio object having a direction and a distance (i.e., a position), the focusing gain determiner 901 may use the same implementation process as described in the context of parameterized audio processing to determine the direct gain f(k) 912 based on the spatial metadata and the focusing parameters. In these embodiments, there is no filter bank. In other words, there is only one frequency band k.
[0172] In addition, the focusing processor may further include a focusing gain processor (for each channel) 903. The focusing gain processor 903 is configured to receive the focusing gain f(k) 912 for each audio channel and the audio signal 906. Further, the focusing gain 912 may be applied to the corresponding audio channel signal 906 (and in some embodiments, time smoothing is also performed). The output from the focusing gain processor 903 may be the focused audio channel audio signal 914.
[0173] In these examples, the channel orientation / position information 902 is not changed and is also provided as the channel orientation / position information output 910.
[0174] In some embodiments, when the input audio channel has no distance information (for example, the input is a speaker or object sound that only has a direction but no distance), one option for processing such an audio channel is to determine a fixed default distance for such a signal and apply the same formula to determine f(k).
[0175] In some embodiments, determining the focusing gain f(k) 912 for such an audio channel may be based on the angular difference between the focusing direction and the direction of the audio channel. In some embodiments, this may first determine the focusing width θ w . For example, as Figure 10 shown, the focusing width θ w 1005 can be determined based on trigonometry using the focusing distance 1001 and the focusing radius 1103, where the focusing width is generated by the angle of a right triangle having the hypotenuse formed by the focusing distance 1001 and the opposite side formed by the focusing radius 1003. The focusing width can be simply determined by the following formula:
[0176]
[0177] Further, determine the angle θ a between the focusing direction and the direction of the audio channel (individually for each audio channel). Further, a formula similar to the above can be used to determine f(k), where d r is replaced by θ w , and d(k) is replaced by θ a (when determining the focusing gain for an audio channel without distance information). In some embodiments, when the focusing radius is greater than the focusing distance, the above asin function is not defined, and a large value (for example, π) can be used for the focusing width θ w .
[0178] Regarding Figure 9b , it is shown as Figure 9aFlowchart 960 of the operation of the multi-channel / object audio input processor shown in
[0179] As Figure 9b shown in step 961, the initial operation is to receive a multi-channel / object audio signal (along with focus parameters or other control information and channel information such as direction / distance).
[0180] As Figure 9b shown in step 963, the next operation is to generate a focus gain factor.
[0181] As Figure 9b shown in step 969, the next operation is to apply the focus gain to each channel audio signal.
[0182] Furthermore, as Figure 9b shown in step 967, the processed audio signal and the unmodified channel direction (and distance) can be output.
[0183] In some embodiments, other parameters and other combinations of parameters can also be used to define the focus shape. In these cases, the focus processor can be modified according to the above examples to use these parameters.
[0184] Regarding Figure 11a , an example of a reproduction processor 1150 based on Ambisonic audio input is shown (e.g., which can be configured to receive the output from an example focus processor as Figure 5a shown in
[0185] In these examples, the reproduction processor can include an Ambisonic rotation matrix processor 1101. The Ambisonic rotation matrix processor 1101 is configured to receive the focused Ambisonic signal 1100 and the viewing direction 1102. The Ambisonic rotation matrix processor 1101 is configured to generate a rotation matrix based on the viewing direction parameter 1102. In some embodiments, this can be done using any suitable method, such as those methods applied to Ambisonic binauralization for head tracking (or more generally, the rotation of such spherical harmonic functions is used in many fields including other fields besides audio). Furthermore, the rotation matrix is applied to the Ambisonic audio signal. The result is a rotated Ambisonic signal with the added focus 1104, and these rotated Ambisonic signals are output to the Ambisonic to binaural filter 1103.
[0186] The Ambisonic to binaural filter 1103 is configured to receive a rotated Ambisonic signal with added focus / defocus 1104. The Ambisonic to binaural filter 1103 may include a pre-formulated 2xK finite impulse response (FIR) filter matrix, and these FIR filters are applied to K Ambisonic signals to generate 2 binaural signals 1106. The FIR filters have been generated by a least squares optimization method with respect to a set of head-related impulse responses (HRIR). An example of such a design process is to transform the HRIR data set into frequency intervals (e.g., by FFT) to obtain an HRTF data set, and determine a complex-valued processing matrix for each frequency interval, which approximately fits the available HRTF data set at the data points of the HRTF data set in the least squares sense. When the complex-valued matrices are determined for all frequency intervals in this way, the result can be inverse-transformed (e.g., by inverse FFT) into time-domain FIR filters. The FIR filters may also be windowed, for example, by using a Hann window.
[0187] There are many known methods that can be used to render Ambisonic signals as speaker outputs. An example could be to linearly decode the Ambisonic signal into a target speaker configuration. This can be applied when the order of the Ambisonic signal is high enough (e.g., at least 3rd order, but preferably 4th order). In a particular example of such linear decoding, the Ambisonic decoding matrix can be designed to generate a speaker signal corresponding to a beam pattern that approximately fits, in the least squares sense, a vector base amplitude panning (VBAP) beam pattern suitable for the target speaker configuration when applied to the Ambisonic signal (corresponding to the Ambisonic beam pattern). Processing the Ambisonic signal with this designed Ambisonic decoding matrix can be configured to generate a speaker sound output. In such an embodiment, the reproduction processor is configured to receive information about the speaker configuration.
[0188] Regarding Figure 11b , a flowchart 1160 showing the operation of the Ambisonic input reproduction processor as shown in Figure 11a is presented.
[0189] As shown in step 1161 of Figure 11b , the initial operation is to receive the focused Ambisonic audio signal (and the viewing direction).
[0190] As shown in step 1163 of Figure 11b , the next operation is to generate a rotation matrix based on the viewing direction.
[0191] As shown in Figure 11bAs shown in step 1165, the next operation is to apply a rotation matrix to the Ambisonic audio signal to generate a rotated and focused Ambisonic audio signal.
[0192] Furthermore, as Figure 11b shown in step 1167, the next operation is to convert the Ambisonic audio signal into a suitable audio output format, such as a binaural format (or a multi-channel audio format).
[0193] Furthermore, as Figure 11b shown in step 1169, the output audio format is output.
[0194] Regarding Figure 12a , an example of a reproduction processor 1250 based on parametric spatial audio input is shown (e.g., it can be configured to receive the output of an example focusing processor as shown in Figure 8a ).
[0195] In some embodiments, the reproduction processor includes a filter bank 1201, which is configured to receive the audio channel 1200 audio signals and transform these audio channels into frequency bands (unless the input is already in a suitable time-frequency domain). Examples of suitable filter banks include the short-time Fourier transform (STFT) and the complex quadrature mirror filter (QMF) bank. The time-frequency audio signal 1202 can be output to the parametric binaural synthesizer 703.
[0196] In some embodiments, the reproduction processor includes a parametric binaural synthesizer 1203, which is configured to receive the time-frequency audio signal 1202 and the modified (and unmodified) metadata 1204, and also receive the viewing direction 1206 (or suitable reproduction-related control or tracking information). In the context of 6DOF reproduction, the user position can be provided together with the viewing direction parameter.
[0197] The parametric binaural synthesizer 1203 can be configured to implement any suitable known parametric spatial synthesis method, which is configured to generate the binaural audio signal (in the frequency band) 1208, since the signal and metadata have been focused and modified before the parametric binauralization block. The binauralized time-frequency audio signal 1208 can then be passed to the inverse filter bank 1205. A further feature of the embodiment is that the reproduction processor including the inverse filter bank 1205 is configured to receive the binauralized time-frequency audio signal 1208 and apply inverse filtering for the applied forward filter bank, thus generating the time-domain binauralized audio signal 1210 with focused characteristics suitable for reproduction by headphones (not shown in Figure 12a ).
[0198] In some embodiments, a suitable loudspeaker synthesis method is used to replace the binaural audio signal output with a loudspeaker channel audio signal output format from a parametric spatial audio signal. Any suitable method can be used. For example, based on a suitable known method, the information of the positions of the loudspeakers is used to replace the viewing direction parameter, and a loudspeaker processor is used to replace the binaural processor.
[0199] Regarding Figure 12b , a flowchart 1260 showing the operation of a parametric spatial audio input reproduction processor as shown in Figure 12a is presented.
[0200] As shown in step 1261 of Figure 12b , the initial operation is to receive the focused parametric spatial audio signal (along with the viewing direction or other reproduction-related control or tracking information).
[0201] As shown in step 1263 of Figure 12b , the next operation is to perform a time-frequency conversion on the audio signal.
[0202] As shown in step 1265 of Figure 12b , the next operation is to apply a parametric binaural (or loudspeaker channel format) processor based on the time-frequency converted audio signal, metadata, and the viewing direction (or other information).
[0203] Furthermore, as shown in step 1267 of Figure 12b , the next operation is to perform an inverse transform on the generated binaural or loudspeaker channel audio signal.
[0204] Furthermore, as shown in step 1269 of Figure 12b , the output audio format is output.
[0205] Considering the loudspeaker output for the reproduction processor when the audio signal is in the form of multi-channel audio and the focusing processor 950 in Figure 9a is applied, in some embodiments the reproduction processor may include a pass-through, where the output loudspeaker configuration is the same as the format of the input signal. In some embodiments where the output loudspeaker configuration is different from the input loudspeaker configuration, the reproduction processor may include a Vector Base Amplitude Panning (VBAP) processor. Furthermore, VBAP (a known amplitude panning technique) can be used to process each focused audio channel to spatially reproduce them using the target loudspeaker configuration. Thus, the output audio signal matches the output loudspeaker settings.
[0206] In some embodiments, any suitable amplitude panning technique can be used to effect the transition from the first loudspeaker configuration to the second loudspeaker configuration. For example, the amplitude panning technique can include deriving an N×M matrix of amplitude panning gains that defines the transition from the M channels of the first loudspeaker configuration to the N channels of the second loudspeaker configuration, and then using that matrix to multiply the channels of an intermediate spatial audio signal that is provided as a multi-channel loudspeaker signal according to the first loudspeaker configuration. The intermediate spatial audio signal can be understood as an audio signal similar to the focused sound component 204 as shown in Figure 2a As a non-limiting example, the derivation of VBAP amplitude panning gains is provided in Pulkki, Ville, “Virtual sound source positioning using vector base amplitude panning” (Journal of the Audio Engineering Society, Vol. 45, No. 6 (1997): pp. 456-466).
[0207] For binaural output, any suitable binauralization of the multi-channel loudspeaker signal format (and / or object) can be achieved. For example, typical binauralization can include processing the audio channels with head-related transfer functions (HRTFs) and adding synthetic room reverberation to generate an auditory impression of the listening room. By adopting, for example, the principles outlined in GB patent application GB1710085.0, the distance + direction (i.e., position) information of the audio object sound can be used for 6DOF reproduction as the user moves.
[0208] In Figure 13 an example apparatus suitable for implementation in the form of a mobile phone or mobile device 1401 running suitable software 1403 is shown. Video can be reproduced, for example, by attaching the mobile phone 1401 to a Daydream View type device (however, video processing is not discussed here for clarity).
[0209] The audio bitstream acquirer 1423 is configured to acquire an audio bitstream 1424, for example, by receiving / retrieving from a storage device. In some embodiments, the mobile device includes a decoder 1425 that is configured to receive compressed audio and decode it. In the case of AAC decoding, an example of the decoder is an AAC decoder. The resulting decoded (e.g., Ambisonic, where example implementations are as shown in Figure 5a and 11a audio signal 1426 can be forwarded to the focus processor 1427.
[0210] The mobile phone 1401 receives controller data 1400 from an external controller at the controller data receiver 1411 (e.g., via Bluetooth) and passes the data to the focus parameter (from controller data) determiner 1421. The focus parameter (from controller data) determiner 1421 determines the focus parameters, for example, based on the orientation of the controller device and / or button events. The focus parameters can include any kind of combination of the proposed focus parameters (e.g., focus direction, focus amount, focus height, and focus width). The focus parameters 1422 are forwarded to the focus processor 1427.
[0211] Based on the Ambisonic audio signal and the focus parameters, the focus processor 1427 is configured to create a modified Ambisonic signal 1428 with the desired focus characteristics. These modified Ambisonic signals 1428 are forwarded to the Ambisonic-to-binaural processor 1429. The Ambisonic-to-binaural processor 1429 is also configured to receive the head orientation information 1404 from the orientation tracker 1413 of the mobile phone 1401. Based on the modified Ambisonic signal 1428 and the head orientation information 1404, the Ambisonic-to-binaural processor 1429 is configured to create a head-tracked binaural signal 1430, which can be output from the mobile phone and played back using, for example, headphones.
[0212] Figure 14 An example device (or focus parameter controller) 1550 is shown, which can be configured to control or generate suitable focus parameters, such as focus direction, focus amount, and focus width. The user of the device can be configured to select the focus direction by pointing the controller in the desired direction 1509 and pressing the select focus direction button 1505. The controller has an orientation tracker 1501, and the orientation information can be used to determine the focus direction (e.g., in the focus parameter (from controller data) determiner 1421 as shown in Figure 13 ). In some embodiments, the focus direction can be visualized in a visual display when the focus direction is selected.
[0213] In some embodiments, a focus amount button (shown as + and - in Figure 14 ) 1507 can be used to control the focus amount. Each press will increase / decrease the focus amount by a certain amount, for example, 10 percentage points. A focus width button (shown as + and - in Figure 14 ) 1503 can be used to control the focus width. Each press can be configured to increase / decrease the focus width by a fixed amount, such as 10 degrees.
[0214] In some embodiments, a controller (e.g., with Figure 14The controller depicted therein determines the focusing shape by drawing the desired shape. The user can initiate the drawing operation by holding down the select focus direction button, then draw the desired shape with the controller, and finally approve the shape by stopping the pressing. The drawn shape can be visualized in the visual display while being drawn. The drawn shape can be converted into focus direction, focus height, and focus width parameters. As described in the previous example, the "focus amount" button can be used to select the focus amount.
[0215] In some embodiments, as Figure 14 shown in, the focus controller is modified such that the "focus width" control is replaced by a "focus radius" control to enable control of complex, content-adaptive focusing shapes. In such an embodiment, it can be implemented as part of an advanced virtual reality reproduction system, where the 360-degree video is not only panoramic but also contains depth information (i.e., it is essentially a 3D video that can respond to 6-degree-of-freedom user movement). For example, the video content can have been generated by computer graphics methods or by a VR video capture system capable of detecting visual depth, and thus enables 6DOF similar to computer-generated content.
[0216] In an example scenario, there are two sources of interest, e.g., speakers. Then, the user points at and clicks on "select focus direction" for the two sources, and the visual display then indicates to the user that these sources (which are not only auditory sources but also visual sources at certain directions and distances) have been selected for audio focusing. Then, the user selects the focus amount and focus radius parameters, where the focus radius indicates how far from the source of interest an auditory event will be included within the determined focusing shape. During control adjustment, the focus radius can be indicated as a visual sphere around the visual source of interest.
[0217] The field of view can respond to user movement, but the sources can also move within the scene and are generally visually tracked for their positions. Thus, the focusing shape (which can be represented by two spheres in 3D space in this case) then adaptively changes its overall shape by moving these spheres.
[0218] In other words, a complex focusing shape with depth focusing is obtained. Then, depending on the spatial audio format, the focusing shape can be accurately reproduced (under the condition that the spatial audio has reliable distance information), or approximated in other ways such as those illustrated above.
[0219] In some embodiments, it may be desirable to further specify the focusing process, for example, by determining the desired frequency range or spectral characteristics of the focusing signal. In particular, the following operations may be useful: accentuating the focused audio spectrum within the speech frequency range to improve clarity, for example, by attenuating low-frequency content (e.g., below 200 Hz) and high-frequency content (e.g., above 8 kHz), thus leaving a particularly useful frequency range related to speech.
[0220] It should be understood that the focused signal can be further processed with any known audio processing techniques, such as automatic gain control or enhancement techniques (e.g., bandwidth expansion, noise suppression).
[0221] In some further embodiments, the focusing parameters (including direction, amount, and at least one focusing shape parameter) are generated by the content creator and are sent together with the spatial audio signal. For example, the scenario could be a VR video / audio recording of an unplugged concert near a stage. The content creator may assume that a typical distant listener would like to define a focusing arc that spans towards the stage and also towards the sides for the indoor acoustic effect, but at least to some extent removes the direct sound from the audience (behind the main direction of the VR camera). Thus, a focusing parameter track is added to the stream, and it can be set to the default rendering mode. However, the audience sound still exists in the stream, and some users may prefer to forego the focusing process and enable the reproduction of the full sound scene including the audience sound.
[0222] In other words, instead of the user needing to select the direction and shape of the focus, it is possible to select possible dynamic focusing parameter presets. The preset may already be fine-tuned by the content creator to follow the program well, for example, to turn off the focus at the end of each song and play back the applause to the listener. The content creator can generate some expected preference profiles as focusing parameters. This approach is beneficial because only one spatial audio signal needs to be transmitted, but different preference profiles can be added. Traditional players that do not support focusing can decode the Ambisonic signal without the focusing process.
[0223] In some further embodiments, the focus shape is controlled in conjunction with visual zooming in a video having multiple viewing directions. Visual zooming can be conceptualized as a user controlling a virtual pair of binoculars in a panoramic or 360 or 3D video. In such a use case, when the visual zoom feature is enabled (e.g., set to at least 1.5x zoom), audio focusing of the spatial audio signal can also be enabled. Since the user is then apparently interested in a particular direction, the focus amount can be set to a high value, e.g., 80%, and the focus width can be set to correspond to the arc of the visual view in the virtual binoculars. In other words, as the visual zoom increases, the focus width becomes smaller. Since the focus is set to 80%, the user can to some extent hear the remaining ambient sound in the appropriate direction. In this way, the user hears new content of interest emerging, knows to turn off the visual zoom, and looks in the new direction of interest. The zoom processing can also be used in the context of an audio codec that permits such processing. An example of such a codec can be, for example, MPEG-I.
[0224] A user in such an embodiment as described above can use the present invention to control the focus shape in a variety of ways.
[0225] Figure 15 An example processing output based on an implementation described for high-order Ambisonic (HOA) signals is shown. The figure shows an 8-channel speaker decoded output as a spectrogram of a third-order HOA signal with a speaker at 0°, a sine wave at -90°, and white noise at 110°. It illustrates how a narrow focus towards the speaker reduces the relative energy of the sine wave and white noise, and how a wider focus encompassing both the speaker and the sine wave only significantly reduces the relative energy of the white noise.
[0226] Regarding Figure 16 , an example electronic device that can be used as an analysis or synthesis device is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1700 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0227] In some embodiments, device 1700 includes at least one processor or central processing unit 1707. The processor 1707 can be configured to execute various program codes, such as the methods described herein.
[0228] In some embodiments, device 1700 includes a memory 1711. In some embodiments, at least one processor 1707 is coupled to the memory 1711. The memory 1711 can be any suitable storage component. In some embodiments, the memory 1711 includes a program code portion for storing program code that can be implemented on the processor 1707. Additionally, in some embodiments, the memory 1711 can further include a stored data portion for storing data (such as data that has been processed or is to be processed according to the embodiments described herein). When needed, the implemented program code stored in the program code portion and the data stored in the stored data portion can be obtained by the processor 1707 via the memory-processor coupling.
[0229] In some embodiments, interface device 1700 includes a user interface 1705. In some embodiments, the user interface 1705 can be coupled to the processor 1707. In some embodiments, the processor 1707 can control the operation of the user interface 1705 and receive input from the user interface 1705. In some embodiments, the user interface 1705 can enable a user to input commands to the device 1700, for example, via a keypad. In some embodiments, the user interface 1705 can enable a user to obtain information from the device 1700. For example, the user interface 1705 can include a display configured to display information from the device 1700 to the user. In some embodiments, the user interface 1705 can include a touch screen or touch interface that can both enable information to be input into the device 1700 and display information to the user of the device 1700.
[0230] In some embodiments, device 1700 includes an input / output port 1709. In some embodiments, the input / output port 1709 includes a transceiver. In such embodiments, the transceiver can be coupled to the processor 1707 and is configured to communicate with other devices or electronic apparatuses, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components can be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.
[0231] The transceiver can communicate with other devices through any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an infrared data communication path (IRDA).
[0232] The transceiver input / output port 1709 can be configured to receive signals and, in some embodiments, obtain the focusing parameters as described herein.
[0233] In some embodiments, device 1700 can be used to execute appropriate code using processor 1707 to generate appropriate audio signals. Input / output port 1709 can be coupled to any suitable audio output, such as being coupled to a multi-channel speaker system and / or headphones (which can be head-tracking or non-tracking headphones), etc.
[0234] Generally, various embodiments of the present invention can be implemented using hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software executable by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuits or logic, general hardware or controllers, or other computing devices, or some combination thereof.
[0235] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Additionally, in this regard, it should be noted that any block of the logical flow in the accompanying drawings can represent a program step, or interconnected logical circuits, blocks, and functions, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variant CD.
[0236] The memory can be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, as a non-limiting example, can include a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and one or more of the processors.
[0237] Embodiments of the present invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Sophisticated and powerful software tools can be used to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0238] Programs, such as those provided by Synopsys, Inc., of Mountain View, California and Cadence Design, of San Jose, California, can use well-established design rules and pre-stored libraries of design modules to automatically route conductors and place components on a semiconductor chip. Once the design of the semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transferred to a semiconductor manufacturing facility or "fab" for fabrication.
[0239] The foregoing description has provided a complete and useful description of exemplary embodiments of the invention by way of examples and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.
Claims
1. A device for spatial audio reproduction, comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the device to at least: Obtain at least one focusing parameter, the at least one focusing parameter being configured to define a focusing shape; Process a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape; and Output the processed spatial audio signal, wherein, The modified audio scene at least partially enables the relative emphasis of the part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape, Wherein the at least one focusing parameter is further configured to define a focusing amount, and the device is caused to: process the spatial audio signal so as to at least partially control the relative emphasis of the part of the spatial audio signal within the focusing shape according to the focusing amount.
2. The device according to claim 1, wherein Processing the spatial audio signal causes the device to: At least partially increase the relative emphasis of the part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape; Or At least partially reduce the relative emphasis of the part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape.
3. The device according to claim 1, wherein Processing the spatial audio signal causes the device to: at least partially increase or decrease a relative sound level of the part of the spatial audio signal within the focusing shape relative to other parts of the spatial audio signal outside the focusing shape.
4. The device according to claim 3, wherein Processing the spatial audio signal causes the device to: increase or decrease the relative sound level according to the focusing amount.
5. The apparatus according to claim 1, wherein The device is further caused to: Obtain reproduction control information to control at least one aspect of outputting the processed spatial audio signal, and wherein the device caused to output the processed spatial audio signal is further caused to perform one of the following: Process the processed spatial audio signal representing the modified audio scene according to the reproduction control information to generate an output spatial audio signal; or Process the spatial audio signal according to the reproduction control information before processing the spatial audio signal representing the modified audio scene, and output the processed spatial audio signal as the output spatial audio signal.
6. The device according to claim 1, wherein The spatial audio signal and the processed spatial audio signal include corresponding panoramic surround sound signals, and wherein processing the spatial audio signal causes the device to, for one or more frequency subbands: Convert the panoramic surround sound signal associated with the spatial audio signal into a set of beam signals in a defined pattern; Generate a set of modified beam signals based on the set of beam signals, the focusing shape, and the focusing amount; and Convert the modified beam signals to generate a modified panoramic surround sound signal associated with the processed spatial audio signal.
7. The apparatus according to claim 6, wherein, The defined pattern includes a defined number of beams evenly spaced on a plane or in a volume.
8. The apparatus according to claim 6, wherein The spatial audio signal and the processed spatial audio signal include at least one of the following: A corresponding high-order panoramic surround sound signal or A subset of panoramic surround sound signal components of any order.
9. The apparatus according to claim 1, wherein, The spatial audio signal and the processed spatial audio signal include corresponding parametric spatial audio signals, where the parametric spatial audio signal includes one or more audio channels and spatial metadata, where the spatial metadata includes corresponding direction indicators, energy ratio parameters, and distance indicators for a plurality of frequency subbands, and where processing the spatial audio signal further causes the device to: Calculate a spectral adjustment factor for one or more frequency subbands based on the spatial metadata, the focusing shape, and the focusing amount; Apply the spectral adjustment factor to the one or more frequency subbands of the one or more audio channels to generate one or more processed audio channels; Calculate corresponding modified energy ratio parameters associated with the one or more frequency subbands of the processed spatial audio signal based on the focusing shape, the focusing amount, and at least a portion of the spatial metadata; and Construct the processed spatial audio signal, the processed spatial audio signal including the one or more processed audio channels, the modified energy ratio parameters, and spatial metadata other than the energy ratio parameter.
10. The apparatus according to claim 1, wherein, The spatial audio signal and the processed spatial audio signal include multi-channel speaker channels and / or audio object channels, where processing the spatial audio signal further causes the device to: Calculate a gain adjustment factor based on corresponding audio channel direction indicators, the focusing shape, and the focusing amount; Apply the gain adjustment factor to each audio channel; and Construct the processed spatial audio signal, the processed spatial audio signal including the one or more processed multi-channel speaker audio channels and / or the one or more processed audio object channels.
11. The device according to claim 10, wherein The multi-channel speaker channels and / or audio object channels further include corresponding audio channel distance indicators, and where the calculated gain adjustment factor is further based on the audio channel distance indicator.
12. The apparatus according to claim 10, wherein, The device is further caused to: determine a default corresponding audio channel distance, and where the calculated gain adjustment factor is further based on the audio channel distance.
13. The device according to claim 1, wherein The at least one focusing parameter configured to define the focusing shape includes at least one of the following: Focusing direction; Focusing width; Focusing height; Focusing radius; Focusing distance; Focusing depth; Focusing range; Focusing diameter; Or Focusing shape descriptor.
14. The apparatus according to claim 1, wherein The apparatus is further caused to obtain a focusing input from a sensor device including at least one orientation sensor and at least one user input, wherein the focusing input includes: an indication of a focusing direction for the focusing shape based on the orientation of the at least one orientation sensor; and an indication of a focusing width based on the at least one user input.
15. The device according to claim 1, wherein The apparatus is further caused to obtain a focusing input including at least one user input, and wherein the focusing input further includes an indication of the focusing amount based on the at least one user input.
16. A method for spatial audio reproduction, comprising: obtaining at least one focusing parameter configured to define a focusing shape; processing a spatial audio signal representing an audio scene to generate a processed spatial audio signal representing a modified audio scene so as to at least partially control a relative emphasis of a portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape; and outputting the processed spatial audio signal, wherein the modified audio scene at least partially enables the relative emphasis of the portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape, wherein obtaining at least one focusing parameter includes defining a focusing amount, and processing the spatial audio signal includes at least partially controlling the relative emphasis of the portion of the spatial audio signal within the focusing shape according to the focusing amount.
17. The method according to claim 16, wherein, Processing the spatial audio signal includes: at least partially increasing the relative emphasis of the portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape; or at least partially decreasing the relative emphasis of the portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape.
18. The method according to claim 16, wherein Processing the spatial audio signal includes at least one of the following: at least partially increasing or decreasing the relative sound level of the portion of the spatial audio signal within the focusing shape relative to other portions of the spatial audio signal outside the focusing shape; or increasing or decreasing the relative sound level according to the focusing amount.
Citation Information
Patent Citations
Determination of targeted spatial audio parameters and associated spatial audio playback
GB201710085D0
Audio distance estimation for spatial audio processing
GB201710093D0
Method, system and article of manufacture for processing spatial audio
WO2016109065A1