Audio processing

By processing audio signals according to spatial metadata in a portable handheld device and adopting crosstalk cancellation processing, the problem of poor quality when the speaker reproduces the spatial audio signal is solved, and a higher quality spatial audio reproduction is achieved.

CN120075724APending Publication Date: 2025-05-30NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510248194.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2020-09-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, when reproducing spatial audio signals using speakers of portable handheld devices, it is difficult to achieve high-quality spatial audio reproduction as listening to as headphones, especially when the speakers are close.

Method used

By processing the input audio signal according to the spatial metadata, different playback processes are used to render different parts of the spatial audio signal, including the sound direction in the front area and the sound direction not in the front area, the crosstalk cancellation process is used to improve the spatial audio image reproduced by the speaker.

Benefits of technology

The quality of the speakers of portable handheld devices reproduce spatial audio signals is improved, making them closer to the effect of earphone listening, and enhancing the ability to reproduce spatial sound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075724A_ABST
    Figure CN120075724A_ABST
Patent Text Reader

Abstract

Example embodiments relate to audio processing. According to an example embodiment, there is provided a method for processing an input audio signal according to spatial metadata to play back the spatial audio signal in a device according to at least one sound reproduction characteristic of the device, the method comprising: obtaining the input audio signal and the spatial metadata; obtaining the at least one sound reproduction characteristic of the device; rendering, in accordance with the spatial metadata, a first portion of the spatial audio signal using a first type of playback process applied to the input audio signal, where the first portion includes a sound direction within a front region of the spatial audio signal; and rendering a second portion of the spatial audio signal using a second type of playback process applied by the input audio signal in accordance with the spatial metadata and in accordance with the at least one sound reproduction characteristic, where the second portion includes sound directions not included in the first portion, and where the second portion includes sound directions not included in the first portion. The second type of playback process is different from the first type of playback process and involves a crosstalk cancellation process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application for invention titled "Audio Processing" (application number: 202080066763.X, filing date: September 17, 2020). Technical Field

[0002] Examples and non-limiting embodiments of the present invention relate to audio signal processing. In particular, various embodiments of the present invention relate to device-specific rendering of spatial audio signals such as stereo signals with associated spatial metadata. Background Art

[0003] Many portable handheld devices (such as mobile phones, portable media player devices, tablet computers, laptop computers, etc.) have a pair of speakers enabling stereo playback. Generally, the two speakers are located at opposite ends or sides of the device to maximize the distance between them, thereby facilitating the reproduction of stereo audio. However, due to the small size of such devices, these two speakers are usually still relatively close to each other, thus in many cases resulting in impaired spatial audio images in the reproduced stereo audio. In particular, the perceived spatial audio image may be different from the spatial audio image that can be perceived by playing back the same stereo audio signal (e.g., via the speakers of a home stereo system), where the two speakers can be arranged in appropriate positions relative to each other (e.g., far enough apart from each other) to ensure the reproduction of the spatial audio image in its full width or via headphones that enable sound to be reproduced at substantially fixed positions relative to the listener's ears.

[0004] Although dual-channel stereo signals serve as a traditional example of multi-channel sound reproduction that somewhat involves spatial characteristics, more advanced spatial audio reproduction can be provided via parametric spatial audio signals. In the present disclosure, the term "parametric spatial audio signal" refers to an audio signal provided together with associated spatial metadata. The audio signal can include a single-channel audio signal or a multi-channel audio signal, and it can be provided as a time-domain audio signal (e.g., such as linear PCM with a given number of bits per sample and a given sample rate) or an encoded audio signal that has been encoded using an audio encoder known in the art (and thus requires decoding using a corresponding audio decoder before playback). The spatial metadata conveys information defining at least some characteristics of the spatial rendering of the audio signal, for example, provided as a set of spatial audio parameters. The spatial audio parameters can, for example, include one or more sound direction parameters defining the direction of sound in a corresponding one or more frequency subbands and one or more energy ratio parameters defining the ratio between the energy of the directional sound components in the corresponding frequency subbands and the total energy.

[0005] During the audio rendering stage, spatial metadata is applied to control the processing of the audio signal to form an output audio signal in a desired spatial audio rendering format. The applicable spatial audio rendering format depends on the audio hardware intended (and / or available) for rendering the spatial audio signal. Non-limiting examples of spatial audio rendering formats include (two-channel) binaural audio signals, Ambisonic (spherical harmonic) audio formats, or (specified) multi-speaker audio formats (such as 5.1-channel or 7.1 surround sound). The processes applicable to converting a parametric spatial audio signal into a spatial audio rendering format of interest are well known in the art. In this regard, see, for example, [1] Audio rendering using Ambisonic-based audio rendering, [2] Audio rendering for binaural output, and [3] Audio rendering for multi-speaker output. In a typical scenario, the audio signal is processed (individually) according to the spatial metadata in a plurality of frequency subbands (e.g., those frequency subbands for which the associated spatial metadata is provided). Various other audio processing processes may be applied to the parametric spatial audio signal before conversion into a spatial audio rendering format of interest, and / or such audio processing processes may be provided as part of the conversion from the parametric spatial audio signal to the spatial audio rendering format of interest. Non-limiting examples of such audio processing processes include (automatic) gain control, audio equalization, noise processing, audio focusing processing, and dynamic range processing.

[0006] For example, a parametric spatial audio signal can be derived based on two or more microphone signals obtained from the respective two or more microphones of a capture device or via conversion from a spatial audio signal provided in another audio format (e.g., in a spatial audio rendering format such as a given multi-speaker audio format). The derived parametric spatial audio signal may depend on the spatial metadata, which includes respective sound direction parameters and energy ratio parameters for a plurality of frequency subbands based on the two or more microphone signals obtained from the respective two or more microphones of the capture device. For example, for microphone signals from a microphone array in a portable consumer device (such as a mobile phone, a tablet computer, or a digital camera, where the size and / or shape of these devices impose limitations on the placement of two or more microphones within the device), obtaining such a parametric spatial audio signal can be an advantageous option. Practical experiments have shown that traditional "linear" audio capture techniques generally have significant limitations in capturing high-quality spatial audio from typical microphone arrays available in such devices, while audio capture techniques that operate (directly) on the microphone signals to record parametric spatial audio signals generally enable high-quality spatial audio.

[0007] Crosstalk cancellation is an audio processing technique that typically has advantages in binaural audio reproduction using a pair of loudspeakers, in order to enable the reproduction of controlled sounds to the listener's left and right ears, thus enabling binaural playback from loudspeakers rather than from headphones. Another application where crosstalk cancellation is typically applied is stereo widening, where the input audio signal is processed into a signal that conveys a widened stereo image, which typically spans the width of a physical loudspeaker setup, thus enabling enhanced spatial sound reproduction especially in devices where the loudspeakers used for stereo playback are placed close to each other. Crosstalk cancellation addresses the acoustic situation where sound reaches the listener's binaural ears from two loudspeakers: the crosstalk cancellation process aims to ensure that sound is reproduced from the loudspeakers in a controlled manner such that the acoustic signal cancellation occurs at least in a certain frequency range, so that the sound can be reproduced to the user's ears in a manner similar to the scenario where the user wears headphones to listen to binaural or stereo audio. For example, crosstalk cancellation techniques in the context of stereo widening have been proposed (e.g., in [4] and [5]), and crosstalk cancellation is also applicable to, for example, sound reproduction systems using more than two loudspeakers.

[0008] Returning to the reference audio rendering stage, spatial audio rendering formats (e.g., the binaural audio, Ambisonic, and multi-channel loudspeaker formats mentioned above) do not themselves take into account the audio reproduction characteristics specific to the audio hardware used for sound reproduction. However, this can be an important factor affecting the perceivable sound quality, especially when reproducing spatial sound via the loudspeakers of a mobile device (such as a mobile phone, portable media player device, tablet computer, laptop computer, etc.). Typically, when reproducing a parametric spatial audio signal using the loudspeakers of such a device, one of the following options can be applied.

[0009] - Convert the parametric spatial audio signal into a "conventional" two-channel stereo format for playback via a pair of loudspeakers of the mobile device, which typically results in a narrow spatial audio image limited by the width of the playback device.

[0010] - Convert the parametric spatial audio signal into a binaural audio signal and apply a crosstalk cancellation process known in the art to the binaural audio signal. While this method typically provides acceptable sound reproduction in devices with two loudspeakers having substantially the same sound reproduction characteristics and symmetrically arranged with respect to the (assumed) listening position, it results in poor sound quality in scenarios where the assumption of symmetry or similarity of sound reproduction characteristics, for example, does not apply - which is the case in many (multi-purpose) mobile devices whose primary use is not audio playback.

[0011] Therefore, there is still room for improvement in the spatial parameterization of spatial audio signals via two or more speakers to make its sound quality comparable more easily to the sound quality obtainable via headphone listening.

[0012] Reference:

[0013] [1] International Patent Publication WO 2018 / 060550A1;

[0014] [2] Laitinen, Mikko-Ville; Pulkki, Ville, "Binaural reproduction for directional audio coding", 2009, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, pp. 337-340;

[0015] [3] Vilkamo, Juha; Pulkki, Ville, "Minimization of decorrelator artifacts in directional audio coding by covariance domain rendering", Journal of the Audio Engineering Society, Vol. 61, No. 9, pp. 637-646;

[0016] [4] Kirkeby, O; Nelson, A; Hamada, H; Orduna-Bustamante, F, "Fast deconvolution of multichannel systems using regularization", IEEE Transactions on Speech and Audio Processing, Vol. 6, No. 2, pp. 189-194, 1998;

[0017] [5] Bharitkar, S; Kyriakis, C, "Immersive Audio Signal Processing", ch.4, Springer, 2006;

[0018] [6] Vilkamo, J; T; Kuntz, A, "Optimized covariance domain framework for time-frequency processing of spatial audio", Journal of the Audio Engineering Society, Vol. 61, No. 6, pp. 103 - 411, 2013. Summary of the Invention

[0019] According to an exemplary embodiment, there is provided a method for processing an input audio signal according to spatial metadata for playing back a spatial audio signal in a device according to at least one sound reproduction characteristic of the device. The method includes: obtaining the input audio signal and the spatial metadata; obtaining the at least one sound reproduction characteristic of the device; rendering a first portion of the spatial audio signal using a first type of playback process applied to the input audio signal according to the spatial metadata, wherein the first portion includes sound directions within a front region of the spatial audio signal; and rendering a second portion of the spatial audio signal using a second type of playback process applied to the input audio signal according to the spatial metadata and according to the at least one sound reproduction characteristic, wherein the second portion includes sound directions not included in the first portion, and wherein the second type of playback process is different from the first type of playback process and involves crosstalk cancellation processing.

[0020] According to another exemplary embodiment, there is provided an apparatus for processing an input audio signal according to spatial metadata for playing back a spatial audio signal in a device according to at least one sound reproduction characteristic of the device. The apparatus is configured to: obtain the input audio signal and the spatial metadata; obtain the at least one sound reproduction characteristic of the device; render a first portion of the spatial audio signal using a first type of playback process applied to the input audio signal according to the spatial metadata, wherein the first portion includes sound directions within a front region of the spatial audio signal; and render a second portion of the spatial audio signal using a second type of playback process applied to the input audio signal according to the spatial metadata and according to the at least one sound reproduction characteristic, wherein the second portion includes sound directions not included in the first portion, and wherein the second type of playback process is different from the first type of playback process and involves crosstalk cancellation processing.

[0021] According to another example embodiment, there is provided an apparatus for processing an input audio signal according to spatial metadata so as to play back a spatial audio signal in the apparatus according to at least one sound reproduction characteristic of the apparatus. The apparatus includes: means for obtaining the input audio signal and the spatial metadata; means for obtaining the at least one sound reproduction characteristic of the apparatus; means for rendering a first portion of the spatial audio signal according to the spatial metadata using a first type of playback process applied to the input audio signal, wherein the first portion includes sound directions within a front region of the spatial audio signal; and means for rendering a second portion of the spatial audio signal according to the spatial metadata and according to the at least one sound reproduction characteristic using a second type of playback process applied to the input audio signal, wherein the second portion includes sound directions not included in the first portion, and wherein the second type of playback process is different from the first type of playback process and involves crosstalk cancellation processing.

[0022] According to another example embodiment, there is provided an apparatus for processing an input audio signal according to spatial metadata so as to play back a spatial audio signal in the apparatus according to at least one sound reproduction characteristic of the apparatus. The apparatus includes at least one processor and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: obtain the input audio signal and the spatial metadata; obtain the at least one sound reproduction characteristic of the apparatus; render a first portion of the spatial audio signal according to the spatial metadata using a first type of playback process applied to the input audio signal, wherein the first portion includes sound directions within a front region of the spatial audio signal; and render a second portion of the spatial audio signal according to the spatial metadata and according to the at least one sound reproduction characteristic using a second type of playback process applied to the input audio signal, wherein the second portion includes sound directions not included in the first portion, and wherein the second type of playback process is different from the first type of playback process and involves crosstalk cancellation processing.

[0023] According to another example embodiment, there is provided a computer program including computer-readable program code configured to, when the program code is executed on a computing device, cause at least one method according to the example embodiments described above to be performed.

[0024] A computer program according to an example embodiment may be embodied on a volatile or non-volatile computer-readable recording medium, for example, as a computer program product including at least one computer-readable non-transitory medium having program code stored thereon, which when executed by a device causes the device to at least perform the operations described above for the computer program according to an example embodiment of the present invention.

[0025] The exemplary embodiments of the present invention presented in this patent application should not be construed as limiting the applicability of the appended claims. The verb "comprising" and its derivatives are used in this patent application as open-ended limitations, which do not exclude the existence of unrecited features. Unless otherwise expressly stated, the features described hereinafter may be combined with each other arbitrarily.

[0026] Some features of the present invention are set forth in the appended claims. However, when read in conjunction with the drawings, the aspects of the present invention regarding its construction and method of operation, along with its additional objects and advantages, will be best understood from the following description of some exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Embodiments of the present invention are illustrated by way of example and not limitation in the accompanying drawings, in which:

[0028] Figure 1 A block diagram illustrating some elements of an audio processing system according to an example;

[0029] Figure 2 A block diagram illustrating some elements of a device for implementing an audio processing system according to an example;

[0030] Figure 3 A block diagram illustrating some elements of a signal decomposer according to an example;

[0031] Figure 4 A block diagram illustrating some elements of a spatial section processor according to an example;

[0032] Figure 5 A block diagram illustrating some elements of an audio processing system according to an example;

[0033] Figure 6 A block diagram illustrating some elements of an audio processing system according to an example;

[0034] Figure 7 A flowchart illustrating a method for audio processing according to an example;

[0035] Figure 8 An example of performance obtainable via operation of an audio processing system according to an example;

[0036] Figure 9 A block diagram illustrating some elements of a device according to an example. DETAILED DESCRIPTION

[0037] Figure 1The block diagram illustrates some components and / or entities of an audio processing system 100, which can serve as a framework for various embodiments of the audio processing techniques described in the present disclosure. The audio processing system 100 receives an input audio signal 101 and spatial metadata 103, which together constitute a parameterized spatial audio signal. The audio processing system 100 further receives at least one sound reproduction characteristic 105, which serves as a control input for controlling some aspects of the audio processing in the audio processing system 100. The audio processing system 100 enables the processing of the parameterized spatial audio signal into an output audio signal 115 of the audio processing system 100.

[0038] The input audio signal 101 includes a single-channel audio signal or a multi-channel audio signal, and it can be provided as a time-domain audio signal (e.g., such as linear PCM with a given number of bits per sample and a given sample rate) or an encoded audio signal that has been encoded using an audio encoder known in the art. In scenarios where the input audio signal 101 includes a corresponding encoded audio signal, the audio processing system 100 operates using a corresponding audio decoder to decode the encoded audio signal into a corresponding time-domain audio signal.

[0039] The spatial metadata 103 conveys information defining at least some characteristics of the spatial reproduction of the input audio signal 101, e.g., provided as a set of spatial audio parameters. The following description assumes that the spatial audio parameters include one or more sound direction parameters defining the sound direction in a corresponding one or more frequency subbands and one or more energy ratio parameters defining the ratio of the energy of the directional sound component in the corresponding frequency subband to the total energy (or the ratio of the energies of multiple directional sound components to the total energy). However, this is a non-limiting example chosen for editorial clarity of the description, and in other examples, a different set of spatial audio parameters for conveying information defining the sound direction and / or the relationship between the directional sound component and the diffuse sound component can be applied.

[0040] The parameterized spatial audio signal defined by the input audio signal 101 and the spatial metadata 103 defines a spatial audio image representing a sound scene, which may include one or more directional sounds in certain sound directions relative to a presumed listening point and ambient sounds and reverberation around the presumed listening point. In this regard, the directional sound can, for example, represent a corresponding different sound source in a corresponding sound direction relative to the presumed listening point. In other examples, the directional sound can represent reflections or reverberations, combinations of multiple different sound sources, and / or ambient sounds around the presumed listening point. Thus, the sound direction indicated in the spatial metadata for a certain frequency subband indicates the main sound direction in that frequency subband, and it does not necessarily indicate the direction (or even the presence) of different sound sources in that certain frequency subband.

[0041] At least one sound reproduction characteristic 105 includes information defining at least some characteristics of the sound rendering capabilities of the device implementing the audio processing system 100 and / or another device intended to reproduce the output audio signal 115. Examples of information included in at least one sound reproduction characteristic 105 are information obtained based on acoustic measurements and / or acoustic simulations performed on the device (such as the (complex-valued) crosstalk cancellation gain for one or more frequency sub-bands), where the measurement or simulation can at least partially rely on the use of an emulated human head positioned at a typical listening (or viewing) distance relative to the device, and where the emulated human head has corresponding microphones arranged at positions corresponding to the respective ear positions. Other examples of information included in at least one sound reproduction characteristic 105 include an indication of the number of speakers in the device and an indication of the speaker positions of the device relative to a reference position with respect to the device. Here, the reference position refers to the assumed listening (or viewing) position of the user relative to the device when listening to the sound reproduced via the speakers of the device (or viewing visual content from the display of the device). The information defining the speaker positions can include one or more of the following for each speaker of the device:

[0042] - The respective speaker direction relative to a reference direction (e.g., the assumed front direction), e.g., defined as the respective speaker angle ∝ relative to the reference direction i .

[0043] - The respective speaker distance relative to the reference position with respect to the device.

[0044] As a non-limiting example in this regard, at least one sound reproduction characteristic 105 can define that the speakers are positioned at speaker angles ∝ 1 = -15 degrees and ∝ 2 = 15 degrees relative to the front direction at a reference point that is in the middle position between the speakers, assuming a listening distance of 30 cm.

[0045] The output audio signal 115 can include an audio signal that, when reproduced via the speaker arrangement defined in at least one reproduction characteristic 105, provides a sound with binaural characteristics to a listener (located or approximately located at the reference point with respect to the device). On the other hand, if reproduced via headphones, the reproduction of the output audio signal 115 does not provide a sound with suitable binaural characteristics to the listener.

[0046] The audio processing system 100 enables processing of an input audio signal 101 based on spatial metadata 103 so as to playback a spatial audio signal in the device in accordance with at least one sound reproduction characteristic 105 of the device. The processing performed by the audio processing system 100 includes: rendering a first portion of the spatial audio signal using a first type of playback process applied to the input audio signal 101 based on the spatial metadata 103, wherein the first portion includes sound directions within a front region; and rendering a second portion of the spatial audio signal using a second type of playback process applied to the input audio signal 101 based on the spatial metadata 103 and based on the at least one sound reproduction characteristic 105, wherein the second portion includes sound directions not included in the first portion, and wherein the second type of playback process is different from the first playback process and involves crosstalk cancellation processing.

[0047] According to non-limiting examples, the first type of playback process may include or it may be based on an amplitude panning process. In other non-limiting examples, instead of amplitude panning, the first type of playback process may include, for example, delay panning, Ambisonics panning, or any combination or sub-combination of amplitude panning, delay panning, and Ambisonics panning. In contrast, the first type of playback process does not involve any crosstalk cancellation processing, or provides a substantially smaller crosstalk cancellation effect compared to the crosstalk cancellation processing involved in the second type of playback process. In an example, the first type of playback process may further be performed based on at least one sound reproduction characteristic.

[0048] Hereinafter, without loss of generality, the operation of the audio processing system is described by way of example, wherein the first type of playback process involves an amplitude panning process further performed based on at least one sound reproduction characteristic. Each of the first and second types of playback processes may further involve a respective one or more audio signal processing techniques. As non-limiting examples in this regard, as described in more detail in the following examples, the first type of playback process may include audio equalization, while the second type of playback process may include binauralization.

[0049] As a brief overview, according to Figure 1The audio processing system 100 of the example shown includes: a transformation entity (or transducer) 102 for converting an input audio signal 101 from the time domain into a transformed-domain audio signal 107; a signal decomposer 104 for obtaining a first signal component 109-1 representing a first part of a spatial audio image and a second signal component 109-2 representing a second part of the spatial audio image based on the transformed-domain audio signal 107, according to spatial metadata 103 and according to at least one sound reproduction characteristic 105; a first part processor 106 for obtaining a modified first signal component 111-1 based on the first signal component 109-1 and according to at least one sound reproduction characteristic 105; a second part processor 108 for obtaining a modified second signal component 111-2 based on the second signal component 109-2 and according to at least one sound reproduction characteristic 105; a signal combiner 110 for combining the modified first signal component 111-1 and the modified second signal component 111-2 into a transformed-domain output audio signal 113 suitable for speaker reproduction; and an inverse transformation entity 112 for converting the transformed-domain output audio signal 113 into an (time-domain) output audio signal 115 to be used as the output audio signal of the audio processing system 100.

[0050] In other examples, the audio processing system 100 may include other entities in addition to Figure 1 those shown in Figure 1 and / or some of the entities depicted in Figure 1 may be combined with other entities while providing the same or corresponding functions. In particular, the entities shown in Figures 2 to 4 and the entities shown in subsequent Figures 1 to 4 are used to represent the logical components of the audio processing system 100, and these logical components are arranged to perform the corresponding functions without imposing structural limitations on the implementation of the corresponding entities. Thus, for example, corresponding hardware components, corresponding software components, or corresponding combinations of hardware components and software components may be applied to implement any one of the entities shown in Figures 1 to 4 corresponding to any one of them, implement any sub-combination of two or more of the entities shown in Figures 1 to 4 corresponding to any one of them, or implement all of the entities shown in

[0051] The audio processing system 100 can be configured to process an input audio signal 101 arranged as a sequence of input frames (in view of the spatial metadata 103), each input frame including a respective segment of the digital audio signal for each channel, provided as a respective time series of input samples at a predefined sampling frequency. In a typical example, the audio processing system 100 uses a fixed predefined frame length. In other examples, the frame length can be a selectable frame length selectable from a plurality of predefined frame lengths, or the frame length can be an adjustable frame length selectable from a range of predefined frame lengths. The frame length can be defined as the number of samples L included in the frame for each channel of the input audio signal 101, which is mapped to a corresponding duration at the predefined sampling frequency. As an example in this regard, the audio processing system 100 can use a fixed frame length of 20 milliseconds (ms), which produces frames of L = 160, L = 320, L = 640, and L = 960 samples per channel at sampling frequencies of 8, 16, 32, or 48 kHz, respectively. These frames can be non-overlapping, or they can be partially overlapping. However, these values are used as non-limiting examples, and different frame lengths and / or sampling frequencies than these examples can be used instead, e.g., depending on the required audio bandwidth, the required frame delay, and / or the available processing power.

[0052] The audio processing system 100 can be implemented by one or more computing devices, and the resulting output audio signal 115 can be provided for playback via a speaker of one of these devices. Generally, the audio processing system 100 is implemented in a portable handheld device (such as a mobile phone, a media player device, a tablet computer, a laptop computer, etc.), which is also used to playback the output audio signal 115 via a pair of speakers provided in the device. In another example, the audio processing system 100 is provided in a first device, while the playback of the output audio signal 115 is provided in a second device. In another example, a first part of the audio processing system 100 is provided in a first device, while a second part of the audio processing system 100 and the playback of the output audio signal 115 are provided in a second device. In the latter two examples, the second device can include a portable handheld device (such as a mobile phone, a media player device, a tablet computer, a laptop computer, etc.), while the first device can include any type of computing device, e.g., a portable handheld device, a desktop computer, a server device, etc.

[0053] Figure 2 A block diagram illustrates some components and / or entities of a device 50 that can be used to implement the audio processing system 100. The device 50 can be provided, for example, as a portable handheld device or other kind of mobile device. For the sake of brevity and clarity of description, in reference Figure 2In the following description, it is assumed that the elements of the audio processing system 100 and the playback of the resulting output audio signal 115 are provided in the device 50. The device 50 further includes: a microphone array 52 including two or more microphones; an audio pre-processor 54 for processing the respective microphone signals captured by the microphone array 52 into a parametric spatial audio signal including an input audio signal 101 and spatial metadata 103; a memory 56 for storing information, such as the parametric spatial audio signal and at least one sound reproduction characteristic 105; an audio driver 58; and a pair of speakers 60, wherein the audio driver 58 is arranged to drive the playback of the output audio signal 115 via the speakers 60.

[0054] In the device 50, the audio processing system 100 can receive the parametric spatial audio signal (including the input audio signal 101 and the spatial metadata 103) and at least one sound reproduction characteristic 105 by reading this information from the memory 56 in the device 50 or coupled to the device 50. In another example, the device 50 can receive the parametric spatial audio signal and / or at least one sound reproduction characteristic 105 from another device that stores one or both of these information pieces in a memory provided therein via a communication interface, such as a network interface. Instead of providing the output audio signal 115 for playback via the audio driver 58 and the speakers 60 or in addition thereto, the device 50 can be arranged to store the output audio signal 115 in the memory 56 and / or provide the output audio signal 115 to another device via the communication interface for rendering and / or storage therein.

[0055] Return reference Figure 1 ,the transformation entity 102 can be arranged to convert the input audio signal 101 from the time domain into a transformed-domain audio signal 107. Generally, the transformed domain involves the frequency domain. In an example, the transformation entity 102 uses a short-time discrete Fourier transform (STFT) to convert each channel of the input audio signal 101 into a corresponding channel of the transformed-domain audio signal 107 using a predefined analysis window length, such as 20 milliseconds. In another example, the transformation entity 102 uses an (analysis) complex modulated quadrature mirror filter (QMF) bank for time-domain to frequency-domain conversion. The STFT and the QMF bank serve as non-limiting examples in this regard, and any suitable transformation technique known in the art can be used in other examples to create the transformed-domain audio signal 107.

[0056] Part of the processing performed by the audio processing system 100 (e.g., at least some aspects of the processing performed by the signal decomposer 104) can be performed separately for multiple frequency sub-bands. Thus, the operation of the audio processing system 100 can include (at least conceptually) dividing or decomposing each channel of the transform-domain audio signal 107 into multiple frequency sub-bands, thereby providing a corresponding time-frequency representation for each channel of the input audio signal 101. According to a non-limiting example, if applicable, the (conceptual) frequency sub-band division can be performed by the transform entity 102 or by the signal decomposer 104.

[0057] A given frequency band in a given frame can be referred to as a time-frequency tile. The number of frequency sub-bands and the corresponding bandwidths of the frequency sub-bands can be selected, for example, according to the desired frequency resolution and / or the available computing power. In an example, according to the Bark scale, equivalent rectangular band scale, or 3rd octave band scale known in the art, the sub-band structure involves 24 frequency sub-bands. In other examples, different numbers of frequency sub-bands with the same or different bandwidths can be used. In this regard, specific examples are a single frequency sub-band covering the entire input spectrum or a single frequency sub-band covering a continuous subset of the input spectrum.

[0058] The time-frequency tile representing frequency bin b in time frame n of channel i of the transform-domain audio signal 107 can be labeled as x(i, b, n). The transform-domain audio signal 107 (e.g., the time-frequency tile x(i, b, n)) is passed to the signal decomposer 104 to be decomposed into a first signal component 109-1 and a second signal component 109-2 therein. As mentioned before, multiple consecutive frequency bins can be grouped into frequency sub-bands, thereby providing multiple frequency sub-bands k = 0, …, K-1. For each frequency sub-band k, the lowest bin (i.e., the frequency bin representing the lowest frequency in that frequency sub-band) can be labeled as b k,low , and the highest bin (i.e., the frequency bin representing the highest frequency in that frequency sub-band) can be labeled as b k,high . In the following example, it is (implicitly) assumed that the STFT is used in the transform entity 102. In such an example, the transform entity 102 can transform each frame n of the input audio signal 101 into a corresponding frame of the frequency-domain audio signal 107, where the frequency-domain audio signal 107 has one time sample per time frame (for each frequency bin b). In other examples, the transform can produce multiple samples in the transform-domain audio signal 107 for each time frame (for each frequency bin b).

[0059] Still referring to Figure 1, the signal decomposer 104 can be set to obtain a first signal component 109-1 representing a first part of the spatial audio image and a second signal component 109-2 representing a second part of the spatial audio image based on the transform-domain audio signal 107 and in accordance with at least one sound reproduction characteristic 105. In this regard, the first part may include a specified spatial part or region of the spatial audio image, while the second part may represent one or more spatial parts or regions of the spatial audio image that do not include the specified spatial part. In an example, the second part may include the remaining part of the spatial audio image, i.e., those parts of the spatial audio image that are not included in the first part.

[0060] According to a non-limiting example, the first part includes sound directions within the front region of the spatial audio image, while the second part includes sound directions that are not included in the first part (e.g., those sound directions that are not included within the front region). Generally but not necessarily, the second part includes the remaining region related to those parts of the spatial audio image that are not included in the front region. This remaining region may also be referred to as the "peripheral" region of the spatial audio image. Thus, in the context of this example, the first signal component 109-1 may also be referred to as the front region signal, and the second signal component 109-2 may also be referred to as the remaining signal. Therefore, the front region may represent those directional sounds of the spatial audio image within the range of predefined sound directions that define the front region in the spatial audio image, while the remaining region may represent the directional sounds of the spatial audio image outside of this predefined range and the ambient (non-directional) sounds of the spatial audio image.

[0061] In a non-limiting example, the first part consists of sound directions within the front region, while the second part does not include the sound directions within the front region but consists of sound directions outside of the front region and the ambient sounds of the spatial audio image. However, those skilled in the art can easily understand that due to the necessary constraints imposed by the actual implementation of the signal decomposer 104 operating on real-world audio signals, it may be impossible to strictly include only directional sounds within the front region in the first part and / or strictly remove these sounds from the second part. Therefore, in this non-limiting example, strictly including only directional sounds within the front region in the first part and strictly removing these directional sounds from the second part only states the processing purpose rather than the processing result in all real-life scenarios.

[0062] The signal decomposition process performed by the signal decomposer 104 may include: obtaining a first signal component 109-1 based on the transform domain audio signal 107 using an amplitude translation technique in view of the spatial metadata 103 and in view of at least one sound reproduction characteristic 105; and obtaining a second signal component 109-2 based on the transform domain audio signal 107 using a binauralization technique in view of the spatial metadata 103 and in view of at least one sound reproduction characteristic 105. Generally, the decomposition process results in each of the first signal component 109-1 and the second signal component 109-2 having a corresponding audio channel for each speaker of the device implementing the audio processing system 100 (e.g., the speaker 60 of the device 50). Thus, in the case of processing a parametric spatial audio signal for playback by two speakers, each of the first signal component 109-1 and the second signal component 109-2 has a corresponding two audio channels, regardless of the number of channels of the transform domain audio signal 107. Wherein, according to an example, the two channels of the first signal component 109-1 may be used to convey spatial sound, wherein any directional sound within the front region of the spatial audio image is arranged in a corresponding sound direction via the application of the amplitude translation technique, while the two channels of the second signal component 109-2 may be used to convey binaural spatial sound, including any directional sound outside this front region and any ambient sound of the spatial audio image. The signal decomposer 104 provides the first signal component 109-1 to the first part processor 106 for corresponding further processing therein in view of at least one sound reproduction characteristic 105, and provides the second signal component 109-2 to the second part processor 108 for corresponding further processing therein in view of at least one sound reproduction characteristic 105.

[0063] Figure 3FIG. illustrates a block diagram of some components and / or entities of a signal decomposer 104 according to an example, the signal decomposer 104 including: a covariance matrix estimator 114 configured to obtain a covariance matrix 119 and an energy measure 117 based on a transform domain audio signal 107; a target matrix estimator 116 configured to obtain a first target covariance matrix 121-1 and a second target covariance matrix 121-2 based on spatial metadata 103 and the energy measure 117, in view of at least one sound reproduction characteristic 105, wherein the first target covariance matrix 121-1 represents sounds included in a first portion of a spatial audio image and the second target covariance matrix 121-2 represents sounds included in a second portion of the spatial audio image; a mixing rule determiner 118 configured to obtain a first mixing matrix 123-1 and a second mixing matrix 123-2 based on the covariance matrix 119, the first target covariance matrix 121-1, and the second target covariance matrix 121-2; and a mixer 120 configured to obtain a first signal component 109-1 and a second signal component 109-2 based on the transform domain audio signal 107, in view of the mixing matrices 123-1, 123-2. Hereinafter, these (logical) entities of the signal decomposer 104 according to Figure 3 the example will be described in more detail. In other examples, the signal decomposer 104 may include other entities, and / or some of the entities depicted in Figure 3 may be omitted or combined with other entities.

[0064] The covariance matrix estimator 114 is configured to perform a covariance matrix estimation process that includes obtaining the covariance matrix 119 and the energy measure 117 based on the transform domain audio signal 107. The covariance matrix estimator 114 provides the covariance matrix 119 for the mixing rule determiner 118 and provides the energy measure 117 for the target matrix estimator for further processing therein accordingly. Assuming a two-channel frequency domain audio signal 107, it can be represented in vector form as:

[0065]

[0066] Based on this definition, according to the example, the covariance matrix 119 can be obtained as:

[0067]

[0068] Here, E[] denotes the expectation operator, and H denotes the Hermitian transpose. In an example, the expected value that can be derived via the expectation operator can be provided as an average over a number of (consecutive) time indices n, and in another example, the instantaneous value of x(b,n) can be directly applied as the expected value without performing time averaging over the time index n. The energy metric 117 can, for example, include a total energy metric e(k,n) calculated as the sum of the diagonal elements of the covariance matrix C x (k,n).

[0069] The target matrix estimator 116 can be set to obtain a first target covariance matrix 121-1 and a second target covariance matrix 121-2 based on the spatial metadata 103 and the energy metric 117, possibly in view of at least one sound reproduction characteristic 105. The target matrix estimator 116 provides the first and second target covariance matrices 121-1, 121-2 for further processing in the mixing rule determiner 118. For clarity of description, examples involving spatial audio parameters including one or more sound direction parameters defining corresponding sound directions in the horizontal plane for one or more frequency subbands are described below. This readily extends by analogy to further examples which, additionally or alternatively, involve spatial audio parameters including one or more sound direction parameters defining corresponding elevation angles of sound directions for one or more frequency subbands.

[0070] In the present example, the spatial audio parameters included in the spatial metadata 103 include one or more azimuth angles θ(k,n) which serve as corresponding sound direction parameters for one or more frequency subbands. In particular, the azimuth angle θ(k,n) denotes the azimuth angle relative to a predefined reference sound direction (e.g., the direction directly in front of the assumed listening point) for the frequency subband k at the time index n. Additionally, in the present example, the spatial audio parameters included in the spatial metadata 103 include one or more direct-to-total energy ratios r(k,n) which serve as corresponding energy ratio parameters for one or more frequency subbands. In particular, the direct-to-total energy ratio r(k,n) denotes the ratio of the directional energy to the total energy at the frequency subband k at the time index n.

[0071] In this regard, the calculation of the target covariance matrices 121-1, 121-2 can include: determining, based on the spatial metadata 103, an energy factor value d(k,n) for a time index n and a frequency sub-band k such that, for example, the energy factor value d(k,n) has a value of 1 for those time-frequency tiles for which the indicated direction is within a first part of the spatial audio image (e.g., the sound direction is within the sound direction range defining the front region in the spatial audio image), and has a value of 0 for other time-frequency tiles. According to a non-limiting example, the energy factor value d(k,n) can be defined as:

[0072]

[0073] where θ d denotes the absolute value of the angle defining the sound direction range around a predefined reference direction (e.g., the front direction) belonging to the front region in the spatial audio image. Thus, Equation (3a) assumes a front region symmetrically positioned around the reference direction such that the front region spans sound directions from -θ d to θ d . In another example, the front region is not symmetrically positioned around the reference direction, and the energy factor value d(k,n) can be defined as:

[0074]

[0075] where θ d1 , θ d2 denote the respective angles defining the sound direction range relative to a reference direction (e.g., the front direction) belonging to the front region of the spatial audio image.

[0076] According to an example, the angle θ d or the angle θ d1 , θ d2 can be obtained, for example, based on at least one sound reproduction characteristic 105. As a specific example, the angle θ d or the angle θ d1 , θ d2 can be obtained based on the speaker angle α i defined in at least one sound reproduction characteristic 105 such that, for example, θ d1 =α 1 and θ d2 =α 2 . In another example, the angle θ d or the angle θ d1 , θ d2 can be included in at least one sound reproduction characteristic 105. In another example, the angle θ d or the angle θ d1 , θ d2is predefined.

[0077] Generally, the energy factor value d(k,n) indicates the degree of directional sound inclusion of the first part of the spatial audio image. In this regard, equations (3a) and (3b) are used to provide a non-limiting example of providing the energy factor value d(k,n) as a "binary" value that indicates, for time index n and frequency subband k, one of inclusion in the first part (e.g., d(k,n)=1) and exclusion from the first part (e.g., d(k,n)=0). In another example, the transition between the first part and the second part of the spatial audio image can be made smooth, for example, by introducing a transition range around angle θ d or angle θ d1 , θ d2 such that the energy factor value d(k,n) is set to a value 0 < d(k,n) < 1, so that the energy factor value decreases as the distance from the reference direction (e.g., the front direction) increases. Thus, for time index n and frequency subband k, the contribution of the directional sound whose sound direction is within the transition range is divided between the first part and the second part according to the energy factor value d(k,n).

[0078] As described above, the signal decomposition process performed by the signal decomposer 104 can include using an amplitude translation technique to obtain the first signal component 109-1. Thus, the calculation of the first target covariance matrix 121-1 can further include: for each time-frequency map block, determining a corresponding translation gain vector g(k,n) based on the sound direction parameter defined for the corresponding time-frequency map block in the spatial metadata 103, where the translation gain vector g(k,n) includes the corresponding translation gain for each channel of the first signal component 109-1 subsequently obtained by the operation of the signal decomposer 104. In an example, this includes: for time index n and frequency subband k, determining the translation gain vector g(θ(k,n)) based on the azimuth angle θ(k,n) defined for the corresponding time-frequency map block. In this example, the translation gain vector g(θ(k,n)) includes a 2x1 corresponding real-valued gain vector, thus providing corresponding gains for the left and right channels. Any amplitude translation technique known in the art can be used in obtaining the translation gain g(k,n), such as vector-based amplitude panning (VBAP), tangent panning law, or sine panning law.

[0079] Given the known energy value e(k,n), energy factor value d(k,n), and translation gain g(k,n), the target matrix estimator 116 can continue to obtain the first target covariance matrix 121-1 based on the spatial audio parameters available in the spatial metadata, for example:

[0080] C 1 (k,n) = g(k,n)gT (k,n)d(k,n)r(k,n)e(k,n) (4)

[0081] Thus, according to the first target covariance matrix C of Equation (4) 1 (k,n) represents those directional sounds included in the first part of the spatial audio image (e.g., in the front region of the spatial audio image).

[0082] Along the lines described above, the signal decomposition process performed by the signal decomposer 104 may include: obtaining the second signal component 109-2 as a binaural audio signal using, for example, the binauralization technique described below. Thus, as an example, the calculation of the second target covariance matrix 121-2 may include: for each time-frequency bin, determining the corresponding head-related transfer function (HRTF) vector h(k,n) based on the sound direction parameter defined for the corresponding time-frequency bin in the spatial metadata 103. In an example, this includes: for time index n and frequency subband k, determining the HRTF vector n(k,θ(k,n)) based on the azimuth angle θ(k,n) defined for the corresponding time-frequency bin, whereby the HRTF vector h(k,θ(k,n)) includes a 2x1 corresponding complex-valued gain vector and provides corresponding gains for the left and right channels. The HRTF vector h(k,n) may be obtained, for example, from an HRTF database stored in the memory of the device implementing the audio processing system 100 (e.g., in the memory 56 of device 50).

[0083] The obtaining of the second target covariance matrix 121-2 may further include: obtaining the diffuse field covariance matrix C d (k), which may be predefined by assuming (preferably substantially uniformly) a set of sound direction values θ across a predefined range of sound directions m (where m = 1,..., M) as:

[0084]

[0085] Here, the set of sound direction values θ m may include, for example, from 20 to 60 sound directions, which are (pseudo) uniformly spaced to cover the desired spatial portion of the 3D space, thereby modeling the response to all directions from the desired spatial portion of the 3D space. As described above, the diffuse field covariance matrix C d (k) may be pre-computed, for example, according to Equation (5) and provided to the signal decomposer 104 for obtaining the second target covariance matrix 121-2.

[0086] Thus, for example, the second target covariance matrix 121-2 may be obtained as:

[0087] C 2 (k,n) = r(k,n)h(k,n)h H (k,n)(1 - d(k,n))e(k,n)+(1 - r(k,n))C d (k)e(k,n) (6)

[0088] Therefore, according to the second target covariance matrix C of Equation (6) 2 (k,n) represents those directional sounds included in the second part of the spatial audio image (e.g., in the remaining regions of the spatial audio image) and the non - directional (ambient) sounds of the spatial audio image.

[0089] The mixing rule determiner 118 can be set to: obtain the first mixing matrix 123 - 1 and the second mixing matrix 123 - 2 based on the covariance matrix 119, the first target covariance matrix 121 - 1, and the second target covariance matrix 121 - 2; and provide the first and second mixing matrices 123 - 1, 123 - 2 to the mixer 120 for further processing therein. The mixing rule determination process performed by the mixing rule determiner 118 can include: obtaining the first mixing matrix 123 - 1 based on the covariance matrix 119 and the first target covariance matrix 121 - 1; and obtaining the second mixing matrix 123 - 2 based on the covariance matrix 119 and the second target covariance matrix 121 - 2. As an example, the respective obtaining of the first mixing matrix 123 - 1 and the second mixing matrix 123 - 2 can be performed as described in [6].

[0090] In particular, the formulas provided in Appendix [6] can be used to obtain, for time index n and frequency sub - band k, based on the covariance matrix 119 (e.g., covariance matrix C x (k,n)) and the first target covariance matrix 121 - 1 (e.g., target covariance matrix C 1 (k,n)), the mixing matrix M 1 (k,n), which can be used as the first mixing matrix 123 - 1 for the corresponding time - frequency map block to obtain the corresponding time - frequency map block of the first signal component 109 - 1 such that it has the same or similar covariance matrix as the first target covariance matrix 121 - 1. Along similar lines, the process of [6] can be applied to obtain, for time index n and frequency sub - band k, based on the covariance matrix 119 (e.g., covariance matrix C x (k,n)) and the second target covariance matrix 121 - 2 (e.g., target covariance matrix C 2 (k,n)), the mixing matrix M 2(k, n), which can be used as a second mixing matrix 123-2 for the corresponding time-frequency map block to obtain the corresponding time-frequency map block of the second signal component 109-2 so that it has a covariance matrix that is the same as or similar to the second target covariance matrix 121-2.

[0091] For the purpose of this acquisition, a prototype matrix Q is defined to guide the mixing matrix M according to the procedure described in detail in [6]. 1 (k,n) and M 2 Generation of (k,n):

[0092]

[0093] The procedure explained in detail in [6] is applied based on the covariance matrix C x (k,n) and the first target covariance matrix C 1 (k,n) Get the mixing matrix M 1 (k,n), so that when the mixing matrix M 1 (k,n) is applied to the covariance matrix C x (k,n) signal, the resulting processed signal is approximated by the least squares optimization method with the covariance matrix C 1 (k,n) signal. Along similar lines, this process can be applied to obtain the mixing matrix M 2 (k,n), which is applied to the covariance matrix C x (k,n) produces a processed signal that is approximated in the sense of least squares optimization to a signal with a covariance matrix C 2 Therefore, a mixing matrix M is provided as the first mixing matrix 123-1. 1 (k, n) is used to obtain the first signal component 109-1 based on the transform domain audio signal 107, and the mixing matrix M used as the second mixing matrix 123-2 is provided. 2 (k, n) for obtaining the second signal component 109-2 based on the transform domain audio signal 107. In this document, the prototype matrix q is provided as a unit matrix so that the signal content in the channels of the first and second signal components 109-1, 109-2 is similar to the signal content of the corresponding channels of the transform domain audio signal 107 (and therefore similar to the signal content of the corresponding channels of the input audio signal 101).

[0094] The mixer 120 may be configured to: obtain a first signal component 109-1 and a second signal component 109-2 based on the transform-domain audio signal 107, in view of the mixing matrices 123-1, 123-2; and provide the first and second signal components 109-1, 109-2 to the first partial processor 106 and the second partial processor 108, respectively, for further processing therein.

[0095] The mixing process performed by the mixer 120 may include: obtaining the first signal component as the product of the first mixing matrix 123-1 and the transform-domain audio signal 107, for example, obtained as:

[0096]

[0097] where k denotes the frequency subband in which the frequency bin b resides, and where x 1 (x,b,n) and x 1 (2,b,n) denote the left and right channels of the first signal component 109-1, respectively. Along similar lines, the mixing process may include: obtaining the second signal component as the product of the second mixing matrix 123-2 and the transform-domain audio signal 107, for example, obtained as:

[0098]

[0099] where x 1 (1,b,n) and x 1 (2,b,n) denote the left and right channels of the second signal component 109-2, respectively.

[0100] Although the above examples provided in equations (8a) and (8b) apply the first and second mixing matrices M 1 (k,n) and M 2 (k,n) in such a way, in another example, one or both of the mixing matrices M 1 (k,n) and M 2 (k,n) may be subject to temporal smoothing (such as averaging over a predefined number of frames (e.g., four frames)).

[0101] Now return to reference Figure 1, the first part processor 106 can be set to: obtain a modified first signal component 111-1 based on the first signal component 109-1 and according to at least one sound reproduction characteristic 105; and provide the modified first signal component 111-1 to the signal combiner 110 for further processing therein. In this regard, the first part processor 106 can be set to: apply a set of equalization gains to obtain a modified first signal component 111-1 based on the first signal component 109-1, for example, obtained as:

[0102] x′ 1 (i,b,n) = g EQ (i,k)x 1 (i,b,n) (9)

[0103] wherein, x′ 1 (i,b,n) denotes the modified first signal component 111-1 for frequency bin b at time index n in channel i (e.g., obtained according to the above equation (8a)), g EQ (i,k) denotes the equalization gain in frequency sub-band k in which the frequency bin b for channel i resides. Thus, for each channel i, equation (9) applies the corresponding equalization gain g EQ (i,k) for each frequency bin b of the frequency sub-band k, resulting in equalization gains that can vary according to the frequency sub-band k and channel i.

[0104] The equalization gain g EQ (i,k) can include corresponding predefined gain values reflecting the characteristics of the device (e.g., device 50) implementing the audio processing system 100. According to an example, the equalization gain g EQ (i,k) can include corresponding predefined gain values provided as part of at least one sound reproduction characteristic 105. In another example, at least one sound reproduction characteristic 105 can include corresponding gains that can be used as a basis for obtaining the equalization gain g EQ (i,k), and / or at least one sound reproduction characteristic 105 can include other types of equalization information enabling the obtaining of the equalization gain g EQ (i,k).

[0105] According to an example, the equalization gain g EQ (i,k) can be obtained on an experimental basis, for example, by using a microphone located at a reference position relative to the device (e.g., device 50) implementing the audio processing system 100 to record a test signal, and obtaining the equalization gain g EQ (i,k) such that they equalize the spectrum of the test signal to a desired degree. In an example, the equalization gain g EQ(i,k) is set such that over - amplification of spectral portions where the signal level is relatively low is avoided. In another example, additionally or alternatively, at least some of the equalization gains g EQ (i,k) can be set to a unity value or a value close to the unity value.

[0106] The equalization gain g EQ (i,k) is intended to equalize the response of the loudspeaker of the device so that the timbre of the sound is less colored, while possibly providing different equalization gains g EQ (i,k) for the channels i of the first signal component 109 - 1. The equalization gain g EQ (i,k) is intended to reduce the differences in the corresponding response of the loudspeaker. Thus, applying the equalization gain g

[0107] In an example, the first - part processor 106 is set to delay the modified first signal component 111 - 1 by a predefined time delay so as to align the modified first signal component 111 - 1 in time with the modified second signal component 111 - 2. Thus, if a delay is applied, the predefined time delay is selected such that it matches or substantially matches the delay resulting from the process performed by the second - part processor 108. In an example, the time delay can be applied to the first signal component 109 - 1, for example, before performing the equalization process according to Equation (9), while in another example, the time delay can be applied to the modified first signal component 111 - 1 obtained, for example, by Equation (9) before providing the signal to the signal combiner 110.

[0108] Still referring to Figure 1, the second part processor 108 can be configured to: obtain a modified second signal component 111-1 based on the second signal component 109-2 and in accordance with at least one sound reproduction characteristic 105; and provide the modified second signal component 111-2 to the signal combiner 110 for further processing therein. In this regard, at least one sound reproduction characteristic 105 can include information specifying the respective acoustic propagation characteristics of each speaker (e.g., speaker 60 of device 50) of the device implementing the audio processing system 100. The processing performed by the second part processor 108 is for performing a crosstalk cancellation process for the second signal component 109-2. In this regard, the second signal component 109-2 obtained from the signal decomposer 104 can be provided as a binaural signal obtained according to equation (8b). Thus, the left channel of the second signal component 109-2 is intended to be played back to the listener's left ear, while the right channel of the second signal component 109-2 is intended to be played back to the listener's right ear. The crosstalk cancellation process performed by the second spatial processor 108 is intended to provide the modified second signal component 111-2 as an audio signal in which the "leakage" of audio signal content from the left channel of the second signal component 109-2 to the listener's right ear located at a reference position relative to the device is reduced (e.g., substantially eliminated), and vice versa, where the "leakage" of audio signal content from the right channel of the second signal component 109-2 to the listener's left ear located at a reference position relative to the device is reduced (e.g., substantially eliminated). Thus, when played back via a speaker, the spatial characteristics generated by the audio content conveyed by the modified second signal component 111-2 are substantially similar to the spatial characteristics obtained by listening via the (binaural) headphones of the second signal component 109-2, thereby enabling high-quality spatial audio reproduction via the speaker.

[0109] Figure 4 illustrates a block diagram of some components and / or entities of the second part processor 108 according to an example, including filter gains H LL (b), H RL (b), H LR (b) and H RR (b) and a filter gain determiner 122 for obtaining the respective filter gains H LL (b), H RL (b), H LR (b) and H RR (b). The filter gains H LL (b), H RL (b), H LR (b) and H RR (b) can also be labeled as crosstalk cancellation gains or crosstalk cancellation filters, which can be provided as respective complex-valued gains for a plurality of frequency bins b (e.g., for all frequency bins b). The filter gains HLL (b), H RL (b), H LR (b) and H RR (b) is typically at least partially based on measurements performed on the speakers of the device implementing the audio processing system 100, and thus they can, at least to some extent, further take into account the device-specific equalization of the speakers.

[0110] The second part processor 108 can be set to: create the left channel of the modified second signal component 111-1 as the left channel of the second signal component 109-2 multiplied by the filtering gain H LL (b) and the right channel of the second signal component 109-2 multiplied by the filtering gain H LR (b) summed; and create the right channel of the modified second signal component 111-2 as the left channel of the second signal component 109-2 multiplied by the filtering gain H RL (b) and the right channel of the second signal component 109-2 multiplied by the filtering gain H RR (b) summed. Herein, the left channel and the right channel of the second signal component 109-2 can respectively include, for example, x obtained according to Equation (8b) 2 (1, b, n) and x 2 (2, b, n), and the left channel and the right channel of the modified second signal component 111-2 for frequency bin b at time index n in channel i can be labeled as x′ 2 (i, b, n).

[0111] According to an example, the corresponding gains of the filtering gain H LL (b), H RL (b), H LR (b) and H RR (b) are predefined and provided as part of at least one sound reproduction characteristic 105, and the filtering gain determiner 122 can thus be configured to: read the filtering gains from the memory in the device implementing the audio processing system 100; and provide these filtering gains for use as the filtering gain H LL (b), H RL (b), H LR (b) and H RR (b) to implement crosstalk cancellation filtering.

[0112] According to another example, for each speaker, at least one sound reproduction characteristic 105 includes a corresponding transfer function (from the corresponding speaker to the left ear of the user and the right ear of the user located at a reference position relative to the device implementing the audio processing system 100), and the filtering coefficient determiner 122 can be set to obtain the corresponding filtering gain H based on the reference frequency response obtained in at least one sound reproduction characteristic 105LL (b), H RL (b), H LR (b) and H RR (b). As a non - limiting example in this regard, the filter gain determiner 122 can be set to obtain the corresponding filter gain H according to the techniques described in [4] LL (b), H RL (b), H LR (b) and H RR (b). An overview of the technique is provided below.

[0113] According to [4], the filter gain for frequency bin b can be obtained as:

[0114] H(b) = (D(b) H D(b) + βI) -1 D(b) H A(b)(10)

[0115] where H(b) denotes a 2x2 complex - valued filter gain matrix in the transform domain, D(b) denotes a 2x2 transfer function matrix obtained as part of at least one sound reproduction characteristic 105, β denotes a real - valued scalar regularization coefficient, I denotes a 2x2 identity matrix, and A(b) denotes a 2x2 target transfer function matrix. Equation (10) can be "expanded" to:

[0116]

[0117] where D LL (b) denotes the reference transfer function from the left speaker to the left ear, D LR (b) denotes the reference transfer function from the left speaker to the right ear, D RL (b) denotes the reference transfer function from the right speaker to the left ear, D RR (b) denotes the reference transfer function from the right speaker to the right ear, A LL (b) denotes the target transfer function from the left speaker to the left ear, A LR (b) denotes the target transfer function from the left speaker to the right ear, A RL (b) denotes the target transfer function from the right speaker to the left ear, A RR (b) represents the transfer function from the right speaker to the right ear.

[0118] As previously mentioned, the transfer functions D LL (b), D RL (b), D LR (b) and D RR(b)Available in at least one sound reproduction characteristic 105, and they can be obtained based on experimental data, for example, via a process involving the following operations: using a microphone arrangement located at a reference position relative to the device (e.g., device 50) implementing the audio processing system 100 to record a test signal, and based on the corresponding test-recorded test signal (e.g., such as the average or another linear combination of multiple recorded test signals corresponding to the respective one of transfer functions D LL (b), D RL (b), D LR (b) and D RR (b), obtain the transfer functions D LL (b), D RL (b), D LR (b) and D RR (b). Here, the corresponding test signals for each of the transfer functions D LL (b), D RL (b), D LR (b) and D RR (b) can be recorded by slightly changing the position and / or orientation of the microphone used to capture the test signal, so as to account for small differences in the user's orientation and / or posture relative to the device. The microphone arrangement mentioned above can, for example, include an artificial human head located at a reference position relative to the device, where the artificial human head has corresponding microphones arranged at positions corresponding to the respective ears.

[0119] According to a non-limiting example, for crosstalk cancellation, the second part processor 108 can be set to: set the target transfer function A LL (b) from the left speaker to the left ear to be equal to the reference transfer function D LL (b), i.e., A LL (b) = D LL (b), and set the target transfer function A RR (b) from the right speaker to the right ear to be equal to the reference transfer function D RR (b), i.e., A RR (b) = D RR (b). In another example, the second part processor 108 can provide crosstalk cancellation by setting each of the target transfer function A LL (b) from the left speaker to the left ear and the target transfer function A RR (b) from the right speaker to the right ear to be equal to the unit value, i.e., A LL (b) = A RR (b) = 1. To provide a crosstalk cancellation effect, the target transfer function A LR(b) is set to have an amplitude less than the amplitude of the reference transfer function from the left speaker to the right ear D LR (b), and / or the target transfer function A from the right speaker to the left ear RL (b) is set to have an amplitude less than the amplitude of the reference transfer function from the right speaker to the left ear A RL (b)'s amplitude, for example, such that |A LR (b)| < |D LR (b)| and / or |A RL (b)| < |D RL (b)|. In a non - limiting example, this can be done by setting the target transfer function A LR (b) = g LR D LR (b) for the transfer function from the left speaker to the right ear and / or by setting the target transfer function A LR (b) and / or according to A RL (b) = g RR D RL (b) for the transfer function from the right speaker to the left ear, where, at least in some frequency sub - bands b, 0 ≤ g RL (b) < 1 and / or 0 ≤ g LR < 1. As an example, the second - part processor 108 can be set to set the target transfer function A RL (b) from the left speaker to the right ear and the target transfer function A LR (b) from the right speaker to the left ear to be equal to "zero", i.e., A RL (b) = A LR (b) = A RL (b) = 0.

[0120] Still referring to equations (10) and (11), by example, the regularization coefficient β can be set to a predefined constant value that is the same across the frequency bins b. In another example, the regularization coefficient β can be set to predefined frequency - dependent values (e.g., according to a predefined frequency function) across the frequency bins b, thereby enabling crosstalk cancellation that avoids strong signal - level increases (e.g., "boost") or signal - level decreases (e.g., "cut") in certain frequency ranges of the corresponding frequency response due to the application of the filtering gains H LL (b), H RL (b), H LR (b) and H RR (b). In this case, the constant regularization coefficient β in equation (10) can be replaced by a frequency - bin - dependent regularization coefficient β(b), which has values for the application of the filtering gains H LL (b), H RL (b), H LR (b) and HRR (b) relatively high values of the frequency that causes excessive changes in the signal level (e.g., "surges" or "dips") and the application filtering gain H LL (b), H RL (b), H LR (b) and H RR (b) relatively low values of the frequency that does not cause excessive changes in the signal level (e.g., "surges" or "dips").

[0121] Return reference Figure 1 , the signal combiner 110 can be set to: combine the modified first signal component 111-1 and the modified second signal component 111-2 into a transformed-domain output audio signal 113 suitable for speaker reproduction; and provide the transformed-domain audio signal 113 to the inverse transform entity 112 for further processing therein. As an example in this regard, in the signal combiner 112, the transformed-domain output audio signal 113 can be obtained as the sum, average, or another linear combination of the modified first signal component 111-1 and the modified second signal component 111-2.

[0122] Still referring Figure 1 , the inverse transform entity 112 can be set to: convert the transformed-domain output audio signal 113 into a (time-domain) output audio signal 115; and provide the output audio signal 115 as the output audio signal of the audio processing system 100. In this regard, the inverse transform entity 112 is set to use an appropriate inverse transform that reverses the time-to-transformed-domain conversion performed in the transform entity 102. As a non-limiting example in this regard, the inverse transform entity 112 can apply an inverse STFT or a (synthesis) QMF bank to provide this inverse transform.

[0123] Figure 5 The block diagram illustrates some components and / or entities of the audio processing system 100', which can be used as a framework for various embodiments of the audio processing techniques described in the present disclosure. The audio processing system 100' is a variant of the audio processing system 100 described above via multiple non-limiting examples, so only the differences in its operation from the operation of the audio processing system 100 are described herein. The audio processing system 100' includes a first subsystem 100a and a second subsystem 100b, which can be provided and / or operated separately from each other. In this regard, the first and second subsystems 100a, 100b can be implemented in the same device (e.g., device 50), or they can be implemented in different devices.

[0124] The first subsystem 100a includes a transform entity 101 and a signal decomposer 104, each of which is arranged to operate as described previously in the context of the audio processing system 100. The first subsystem 100a further includes: a first inverse transform entity 112-1 for converting the first signal component 109-1 from the transform domain to the time domain, thereby providing a time-domain first signal component 109-1'; and an inverse transform entity 112-2 for converting the second signal component 109-2 from the transform domain to the time domain, thereby providing a time-domain second signal component 109-2'. Each of the inverse transform entities 112-1, 112-2 is arranged to operate in the manner described previously in the context of the inverse transform entity 112, with necessary modifications.

[0125] Along the lines described previously for the audio processing system 100, in the case of applying the audio processing system 100' to process parametric spatial audio signals for playback by two loudspeakers, each of the first signal component 109-1' and the second signal component 109-2' has a corresponding two audio channels, regardless of the number of channels of the transform-domain audio signal 107. Among them, according to an example, the two channels of the first signal component 109-1' can be used to convey spatial sound, where any directional sound within the front region of the spatial audio image is arranged in the corresponding sound directions via the application of amplitude translation techniques, while the two channels of the second signal component 109-2' can be used to convey binaural spatial sound, which includes any directional sound outside the front region and any ambient sound of the spatial audio image.

[0126] The device implementing the first subsystem 100a can further be arranged to transmit the first and second signal components 109-1', 109-2' to a second subsystem 100b for further processing therein. The first and second signal components 109-1', 109-2' can be accompanied by an audio format indicator for identifying the first and second signal components 109-1', 109-2' as originating from the first subsystem 100a. The transmission from the first subsystem 100a to the second subsystem 100b can include, for example: the device implementing the first subsystem 100a is arranged to send this information to the device implementing the second subsystem 100b via a communication network or communication channel, and / or the device implementing the first subsystem 100a is arranged to store the information in a memory, which can subsequently be read by the second subsystem 100b.

[0127] Accordingly, the device implementing the second subsystem 100b can be set to: receive the first and second signal components 109-1', 109-2' via a network interface, or read the first and second signal components 109-1', 109-2' from a memory. The second subsystem 100b includes: a first transformation entity 102-1 for transforming the first signal component 109-1' from the time domain to a transform domain to recover the first signal component 109-1 in the frequency domain; and a second transformation entity 102-2 for transforming the second signal component 109-2' from the time domain to a transform domain to recover the second signal component 109-2 in the frequency domain. Each of the transformation entities 102-1, 102-2 is set to operate in the manner described previously in the context of the transformation entity 102, with necessary modifications. The second subsystem 100b further includes a first partial processor 106, a second partial processor 108, a combiner 110, and a transformation entity 112, each of which is set to operate as described previously in the context of the audio processing system 100.

[0128] Figure 6 A block diagram illustrates some components and / or entities of an audio processing system 200, which can serve as a framework for various embodiments of the audio processing techniques described in the present disclosure. The audio processing system 200 is a variant of the audio processing system 100 described previously via multiple non-limiting examples. The audio processing system 200 receives an input audio signal 101 and spatial metadata 103 (which together constitute a parameterized spatial audio signal), and the audio processing system 200 further receives at least one sound reproduction characteristic 105, which serves as a control input for controlling some aspects of the audio processing in the audio processing system 200. As in the case of the audio processing system 100, the audio processing system 200 enables the processing of the parameterized spatial audio signal into an output audio signal 215 that constitutes the audio output signal of the audio processing system 200.

[0129] Like the audio processing system 100, the audio processing system 200 also enables the processing of the parameterized spatial audio signal for playback by the speakers of the device, where the processing is performed in accordance with at least one sound reproduction characteristic 105 of the device, and where the processing includes: rendering a first portion of the spatial audio image conveyed by the parameterized spatial audio signal using an amplitude translation process applied to the input audio signal, in accordance with the spatial metadata and the at least one sound reproduction characteristic 105; and rendering a second portion of the spatial audio image using a crosstalk cancellation process applied to the input audio signal, in accordance with the spatial metadata and the at least one sound reproduction characteristic 105.

[0130] As a brief overview, according to Figure 6The audio processing system 200 of the example shown in the figure includes: a transformation entity 102 for converting an input audio signal 101 from the time domain into a transformed-domain audio signal 107; a covariance matrix estimator 114 for obtaining a covariance matrix 119 and an energy metric 117 based on the transformed-domain audio signal 107 in view of spatial metadata 103; a target matrix estimator 216 for obtaining an extended target covariance matrix 221 based on the spatial metadata 103 and the energy metric 117 in view of at least one sound reproduction characteristic 105, wherein the extended target covariance matrix 221 serves as a target covariance matrix for both the sound included in the first part of the spatial audio image and the sound included in the second part of the spatial audio image; a mixing rule determiner 218 for obtaining an extended mixing matrix 223 based on the covariance matrix 119 and the extended target covariance matrix 221; a mixer 220 for obtaining a transformed-domain output audio signal 213 suitable for speaker reproduction based on the transformed-domain audio signal 107 in view of the extended mixing matrix 223; and an inverse transformation entity 112 for converting the transformed-domain output audio signal 213 into an (time-domain) output audio signal 215 to serve as the output audio signal of the audio processing system 200.

[0131] In other examples, the audio processing system 200 may include other entities in addition to those shown in Figure 6 and / or some of the entities depicted in Figure 6 may be combined with other entities while providing the same or corresponding functions. In particular, the entities shown in Figure 6 are used to represent the logical components of the audio processing system 200, and these logical components are set to perform corresponding functions without imposing structural limitations on the implementation of the corresponding entities. Thus, for example, corresponding hardware components, corresponding software components, or corresponding combinations of hardware components and software components may be applied to implement any one of the entities shown in Figure 6 separately from other entities, implement any sub-combination of two or more of the entities shown in Figure 6 or implement all of the entities shown in Figure 6 in combination.

[0132] The overall operation of the audio processing system 200, for example, regarding the characteristics of the input audio signal 101, the spatial metadata 103, and at least one sound reproduction characteristic 105, as well as regarding processing the input audio signal into a sequence of input frames and implementing the audio processing system 200 by the device 50, is similar to that described above for the audio processing system 100. Additionally, the corresponding operations of the transformation entity 102, the covariance matrix estimator 114, and the inverse transformation entity 112 are the same as those described above in the context of the audio processing system 100 (refer to Figure 1 and Figure 3) similar to that described in

[0133] Still referring to Figure 6 , the target matrix estimator 216 can be set to: based on the spatial metadata 103 and the energy metric 117, in view of at least one sound reproduction characteristic 105, obtain an extended target covariance matrix 221, where the extended target covariance matrix 22 is used as the target covariance matrix for both the sound included in the first part of the spatial audio image and the sound included in the second part of the spatial audio image. The target matrix estimator 216 can further be set to: provide the extended target covariance matrix 221 to the mixing rule determiner 218 for further processing therein. As in the case of the target matrix estimator 116, for the sake of clarity of description, examples of spatial audio parameters involving one or more sound direction parameters that define the respective sound directions in the horizontal plane for one or more frequency subbands are described below, which can be easily analogized to further examples. Additionally or alternatively, these examples involve spatial audio parameters including one or more sound direction parameters that define the elevation angle of the sound direction for one or more frequency subbands.

[0134] The target matrix determiner 216 can be set to: calculate the first and second target covariance matrices 121-1, 121-2 according to the process described above in the context of the target matrix determiner 116. As an example in this regard, the target matrix determiner 216 can obtain the first target covariance matrix C 1 (k,n) representing the sound included in the first part of the spatial audio image according to Equation (4), and obtain the second target covariance matrix C 2 (k,n) representing the sound included in the second part of the spatial audio image according to Equation (6). Additionally, the target matrix determiner 216 can further be set to obtain, based on the first target covariance matrix C 1 (k,n), an extended first covariance matrix C′ EQ (k,n) that further explains the characteristics of the device (e.g., device 50) of the audio processing system 200 implemented by using the equalization gain g 1 (i,k) in the context of the first part processor 106 of the audio processing system 100 described above, for example, obtained as:

[0135]

[0136] The target matrix determiner 216 can further be set to obtain, based on the second target covariance matrix C 2 (k), an extended second covariance matrix C′ 2 (k,n), for example, obtained as:

[0137] C′ 2 (k,n) = H(b k,mid )C 2 (k,n)H H (b k,mid ) (13)

[0138] where H(b k,mid ) denotes a 2x2 complex-valued filter coefficient matrix in the transform domain. In this regard, H(b k,mid ) is similar to H(b) defined in the context of Equation (10) above, where the index b k,mid refers to the frequency bin closest to the center frequency of the frequency subband k. The extended first and second target covariance matrices C′ 1 (k,n), C′ 2 (k,n)) can be used to obtain (a combined) extended target covariance matrix 221, for example obtained as:

[0139] C′ y (k,n) = C′ 1 (k,n) + C′ 2 (k,n) (14)

[0140] The mixing rule determiner 218 can be set to: obtain an extended mixing matrix 223 based on the covariance matrix 119 and the extended target covariance matrix 221; and provide the extended mixing matrix 223 to the mixer 220 for further processing therein. The operation of the mixing rule determiner 218 is similar to the operation of the mixing rule determiner 118 described above, except that a single mixing matrix is obtained, which is suitable for processing both the sound included in the first part of the spatial audio image and the sound included in the second part of the spatial audio image in the mixer 220. In this regard, the mixing rule determiner 218 can be set to: apply the formula provided in Appendix [6] to generate a mixing matrix M(k,n) for the time index n and the frequency subband k, based on the covariance matrix 119 (e.g., covariance matrix C x (k,n)) and the extended target covariance matrix 221 (e.g., extended target covariance matrix C′ y (k,n)), which can be used as the extended mixing matrix 223 for the corresponding time-frequency map block.

[0141] The mixer 220 may be configured to: obtain a transformed-domain output audio signal 213 based on the transformed-domain audio signal 107, in view of the extended mixing matrix 223; and provide the transformed-domain output audio signal 213 to the inverse transform entity 112 for further processing therein. The mixing process performed by the mixer 220 may include: obtaining the transformed-domain output audio signal 213 as the product of the extended mixing matrix 223 and the transformed-domain audio signal 107, for example obtaining as:

[0142]

[0143] where k denotes the frequency subband in which the frequency bin b resides.

[0144] The inverse transform entity 112 may be configured to: convert the transformed-domain output audio signal 213 into an (in the time domain) output audio signal 215; and provide the output audio signal 215 as the output audio signal of the audio processing system 200 as described above.

[0145] Above, the operations of the audio processing systems 100, 100', 200 have been described by (implicitly and / or explicitly) referring each of the first signal component 109-1, the second signal component 109-2, the modified first signal component 111-1, the modified second signal component 111-2, the transformed-domain output audio signal 213, and the output audio signal 215 (serving as the output audio signal) as a corresponding dual-channel signal to prepare for sound reproduction via two speakers. However, this is a non-limiting example chosen for clarity and conciseness of description, and the corresponding operations of each element of the audio processing systems 100, 100', 200 can be easily analogized to processing involving three or more channels to address a speaker arrangement including three or more speakers.

[0146] In addition, the above description relates to processing in multiple frequency sub-bands. For example, this can involve: performing the processing described above for audio processing systems 100, 100', 200 for a set of frequency sub-bands that cover or substantially cover the spectrum globally represented by the parametric spatial audio signal. In another example, the audio processing procedures described above with reference to audio processing systems 100, 100', 200 can be performed in a predefined portion of the spectrum represented by the parametric spatial audio signal, and for the remaining portion of this spectrum, output audio signal 215 can be obtained using audio rendering techniques known in the art. In this regard, the predefined portion of the spectrum can include one or more predefined frequency sub-bands, for example, such that certain frequency sub-bands at the low end and / or at the high end of this spectrum are processed using audio rendering mechanisms known in the art, while the frequency sub-bands in between are processed as described above with reference to audio processing systems 100, 100', 200.

[0147] In another example, some aspects of the audio processing described above with reference to audio processing systems 100, 100', 200 can be replaced with different audio rendering techniques in predefined frequency sub-bands (e.g., in certain frequency sub-bands at the low end and / or at the high end of the spectrum). As an example, in audio processing systems 100, 100', this can be achieved by omitting the crosstalk cancellation processing described above with reference to the second part processor 108 at certain frequency sub-bands at the low end and / or at the high end of the spectrum, while in audio processing system 200, this can be achieved by preparing an extended target covariance matrix C′ y (k,n) in the target matrix estimator 216 at certain frequency sub-bands at the low end and / or at the high end of the spectrum and omitting the contribution from the crosstalk cancellation filter H(b k,mid ) (e.g., by setting the filter gains H RL (b) and H LR (b) to "zero" and by setting the filter gains H LL (b) and H RR (b) to unity).

[0148] In another example, in the target matrix estimators 116, 216 regarding the generation of the second target covariance matrix 121-2 (e.g., C 2Binaural synthesis for (k,n) can be omitted at certain frequency subbands at the low end of the spectrum and / or at the high end of the spectrum and replaced by an amplitude translation technique. This can be achieved, for example, by replacing HRTFh(k,n) in Equation (6) with an appropriate amplitude translation gain and replacing the diffuse field covariance matrix in Equation (6) with an identity matrix. Similarly, in the case of the audio processing systems 100, 100' in this example, the crosstalk cancellation process described above with reference to the second part processor 108 should be omitted in certain frequency subbands at the low end of the spectrum and / or at the high end of the spectrum, while in the case of the audio processing system 200, an extended target covariance matrix C' should be prepared in the target matrix estimator 216 y when (k,n), the contribution from the crosstalk cancellation filter H(b k,mid ) should be omitted at specific frequency subbands at the low end of the spectrum and / or at the high end of the spectrum.

[0149] The logic elements of the audio processing systems 100, 100', 200 can be set to operate, for example, according to the method 300 shown in the flowchart depicted in Figure 7 . Method 300 serves as a method for processing the input audio signal 101 according to the spatial metadata 103 for playing back a spatial audio signal in the device, wherein the processing is performed based on at least one sound reproduction characteristic 105 of the device. For example, in view of the examples related to the operation of any one of the audio processing systems 100, 100' and / or 200 described above, method 300 can be varied in a variety of ways.

[0150] Method 300 includes: obtaining the input audio signal 101, the spatial metadata 103, and at least one sound reproduction characteristic 105 of the device, as shown in block 302. Method 300 further includes: rendering a first part of the spatial audio signal using a first type of playback process applied to the input audio signal 101 according to the spatial metadata 103, wherein the first part includes sound directions within the front region, as shown in block 304; and rendering a second part of the spatial audio image using a crosstalk cancellation process applied to the input audio signal 101 according to the spatial metadata 104 and based on at least one sound reproduction characteristic 105, wherein the second part includes sound directions not included in the first part, and wherein the second type of playback process is different from the first type of playback process and involves crosstalk cancellation processing.

[0151] Figure 8Illustrated is an example of the performance obtainable via the operation of audio processing systems 100, 100', 200 (labeled "Proposed Output" in the illustration) compared to a previously known audio processing technique involving binaural synthesis combined with a general crosstalk cancellation technique (labeled "HRTF+CTC Output" in the illustration). In Figure 8 the illustration of Figure 8 , the upper figure depicts the magnitude spectrum of the left channel and the lower figure depicts the magnitude spectrum of the right channel, which are obtained by processing an exemplary parametric audio signal (including a pulse as the input audio signal 101 and spatial metadata 103 defining a "zero" degree sound direction and a directivity-to-total energy ratio of "one" for all frequency subbands). Thus, the exemplary parametric audio signal models a sound source exactly in front of the assumed listening point under anechoic conditions. Figure 8 The magnitude response of the input audio signal is illustrated as a corresponding solid curve, the magnitude response of the output audio signals 115, 215 obtained via the processing of audio processing systems 100, 100', 200 is illustrated as a corresponding dashed curve, and the magnitude response of the processed audio signal obtained using the previously known audio processing technique is illustrated as a corresponding dotted curve.

[0152] As Figure 8 shown in Figure 8 , an input audio signal 101 containing a pulse exactly in front of the assumed listening point under anechoic conditions produces a flat magnitude spectrum in both the left and right channels. Due to the amplitude translation technique applied by audio processing systems 100, 100', 200 to the pulse in the first part of the spatial audio image (e.g., in the front region), the magnitude response of the output audio signals 115, 215 is substantially similar to the magnitude response of the input audio signal 101, except for a slight attenuation (about 3 dB) due to the application of the amplitude translation gain. Thus, there is no signal coloring, enabling a good timbre to be reproduced for the listener. In contrast, a previously known audio processing technique in which the input audio is processed into a binaural signal (via the use of HRTFs) that is further subjected to a crosstalk cancellation process results in significant distortion in the magnitude spectrum (especially in the high end of the spectrum in both channels), which leads to signal coloring and timbre degradation in the reproduced sound, which can be avoided by using audio processing systems 100, 100', 200.

[0153] Figure 9 Illustrated is a block diagram of some components of an exemplary apparatus 400. Apparatus 400 may include other components, elements, or parts not depicted in Figure 9 Figure 9 . For example, in implementing one or more of the components described above in the context of audio processing systems 100, 100', 200, apparatus 400 may be used. Apparatus 400 may implement, for example, device 50 or one or more of its components.

[0154] Device 400 includes a processor 416 and a memory 415 for storing data and computer program code 417. A part of the memory 415 and the computer program code 417 stored therein may further be configured to, together with the processor 416, implement at least some of the operations, processes, and / or functions described above in the context of the audio processing systems 100, 100', 200.

[0155] Device 400 includes a communication section 412 for communicating with other devices. The communication section 412 includes at least one communication device enabling wired or wireless communication with other devices. The communication devices of the communication section 412 may also be referred to as corresponding communication components.

[0156] Device 400 may further include a user I / O (input / output) component 418, which may be configured to (possibly together with the processor 416 and a part of the computer program code 417) provide a user interface for receiving input from a user of the device 400 and / or providing output to a user of the device 400 to control at least some aspects of the operation of the audio processing systems 100, 100', 200 implemented by the device 400. The user I / O component 418 may include hardware components such as a display, a touch screen, a touch pad, a mouse, a keyboard, and / or an arrangement of one or more keys or buttons, etc. The user I / O component 418 may also be referred to as a peripheral device. The processor 416 may be configured to control the operation of the device 400, for example, according to a part of the computer program code 417 and possibly further according to user input received via the user I / O component 418 and / or according to information received via the communication section 412.

[0157] Although the processor 416 is depicted as a single component, it may be implemented as one or more separate processing components. Similarly, although the memory 415 is depicted as a single component, it may be implemented as one or more separate components, some or all of which may be integrated / removable, and / or may provide permanent / semi-permanent / dynamic / cache storage devices.

[0158] The computer program code 417 stored in the memory 415 may include computer-executable instructions that, when loaded into the processor 416, control one or more aspects of the operation of the apparatus 400. As an example, the computer-executable instructions may be provided as one or more sequences of one or more instructions. By reading one or more sequences of one or more instructions contained therein from the memory 415, the processor 416 is able to load and execute the computer program code 417. The one or more sequences of one or more instructions may be configured to cause the apparatus 400 to perform at least some of the operations, processes, and / or functions described hereinabove in the context of the audio processing systems 100, 100', 200.

[0159] Accordingly, the apparatus 400 may include at least one processor 416 and at least one memory 415 including computer program code 417 for one or more programs, the at least one memory 415 and the computer program code 417 being configured to, with the at least one processor 416, cause the apparatus 400 to perform at least some of the operations, processes, and / or functions described hereinabove in the context of the audio processing systems 100, 100', 200.

[0160] The computer program stored in the memory 415 may, for example, be provided as a corresponding computer program product, which includes at least one computer-readable non-transitory medium having stored thereon the computer program code 417, the computer program code causing the apparatus 400 to perform at least some of the operations, processes, and / or functions described hereinabove in the context of the audio processing systems 100, 100', 200 when executed by the apparatus 400. The computer-readable non-transitory medium may include a memory device or a recording medium, such as a CD-ROM, a DVD, a Blu-ray disc, or another article of manufacture tangibly embodying the computer program. As another example, the computer program may be provided as a signal configured to reliably convey the computer program.

[0161] The reference to a processor should not be construed as encompassing only programmable processors, but also includes dedicated circuits such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), signal processors. The features described in the foregoing description may be used in combinations other than the explicitly described combinations.

[0162] Although some functions have been described above with reference to certain features and / or elements, these functions may be performed by other features and / or elements whether described or not. Although features have been described with reference to certain embodiments, these features may also be present in other embodiments whether described or not.

Claims

1. A method for processing an input audio signal according to spatial metadata to playback a spatial audio signal in a device according to at least one sound reproduction characteristic of the device, the method comprises: obtaining, by the device, the input audio signal; separately from obtaining the input audio signal, obtaining, by the device, the spatial metadata, wherein the input audio signal and the spatial metadata convey a spatial audio image; obtaining the at least one sound reproduction characteristic of the device; based on the spatial metadata, obtaining a first target covariance matrix, wherein the first target covariance matrix represents sounds included in a first part of the spatial audio image and includes sound directions within a front region of the spatial audio image; based on the spatial metadata and the at least one sound reproduction characteristic of the device, obtaining a second target covariance matrix, wherein the second target covariance matrix represents sounds included in a second part of the spatial audio image, wherein the second part includes sound directions not included in the first part; and rendering, by the device, the spatial audio signal based on the input audio signal and according to the first target covariance matrix and the second target covariance matrix.

2. The method according to claim 1, wherein rendering the spatial audio signal comprises: rendering the first part of the spatial audio signal using a first type of playback process applied to the input audio signal according to the first target covariance matrix; and rendering the second part of the spatial audio signal using a second type of playback process applied to the input audio signal according to the second target covariance matrix, wherein the second type of playback process is different from the first type of playback process.

3. The method according to claim 1, wherein rendering the spatial audio signal comprises: rendering the spatial audio signal using the first type of playback process and the second type of playback process applied to the input audio signal respectively according to the first target covariance matrix and the second target covariance matrix.

4. The method according to claim 1, wherein performing the processing separately in a plurality of frequency subbands.

5. The method according to claim 1, wherein for one or more frequency subbands, the spatial metadata includes: corresponding sound direction parameters, and corresponding energy ratio parameters.

6. The method according to claim 1, wherein the second part also represents non-directional sounds of the spatial audio image.

7. The method according to claim 1, wherein the at least one sound reproduction characteristic includes a corresponding definition of speaker positions related to a reference position relative to the device, and wherein the method further comprises: defining a range of sound directions belonging to the front region based on the speaker positions.

8. The method according to claim 3, wherein the first type of playback process includes an amplitude translation process.

9. The method according to claim 3, wherein the second type of playback process includes crosstalk cancellation processing.

10. The method according to claim 1, comprising: obtaining a covariance matrix and an energy metric based on the input audio signal; obtaining a combined target covariance matrix based on the first target covariance matrix and the second target covariance matrix; obtaining a mixing matrix based on the covariance matrix and the combined target covariance matrix, the mixing matrix when applied to the input audio signal producing a modified audio signal having a covariance matrix similar to the combined target covariance matrix; obtaining the spatial audio signal based on the input audio signal and the mixing matrix.

11. The method according to claim 10, wherein, obtaining the first target covariance matrix comprises: obtaining an energy factor value indicating the degree of inclusion in the first part of the spatial audio image based on the sound direction parameters included in the spatial metadata; determining a translation gain based on the sound direction parameters; and obtaining the first target covariance matrix based on the energy metric, the translation gain, the energy factor value, and the energy ratio parameter included in the spatial metadata.

12. The method according to claim 10, wherein, obtaining the second target covariance matrix comprises: obtaining an energy factor value indicating the degree of inclusion in the first part of the spatial audio image based on the sound direction parameters and the at least one sound reproduction characteristic included in the spatial metadata; determining a head-related transfer function (HRTF) based on the sound direction parameters included in the spatial metadata; obtaining a diffuse field covariance matrix based on the HRTF across a predefined range of sound directions; and obtaining the second target covariance matrix based on the energy metric, the HRTF, the diffuse field covariance matrix, and the energy ratio parameter included in the spatial metadata.

13. The method according to claim 1, wherein, obtaining the first target covariance matrix comprises: multiplying the first part of the spatial audio image by a gain value that is based on predefined equalization information included in the at least one sound reproduction characteristic.

14. The method according to claim 1, wherein, obtaining the second target covariance matrix comprises: obtaining a set of crosstalk cancellation gains based on a reference transfer function included in the at least one sound reproduction characteristic; and applying the set of crosstalk cancellation gains to the second part of the spatial audio signal.

15. The method according to claim 1, further comprising: obtaining a covariance matrix and an energy metric based on the input audio signal; obtaining an extended first target covariance matrix based on the first target covariance matrix and using a gain value based on predefined equalization information included in the at least one sound reproduction characteristic; obtaining an extended second target covariance matrix based on the second target covariance matrix and crosstalk cancellation gains; obtaining the target covariance matrix as a combination of the extended first target covariance matrix and the extended second target covariance matrix; Obtain a mixing matrix based on the covariance matrix and the target covariance matrix, the mixing matrix generating a modified audio signal having a covariance matrix similar to the target covariance matrix when applied to the input audio signal; and Obtain an output audio signal for playback by the device as the product of the input audio signal and the corresponding mixing matrix.

16. An apparatus for processing an input audio signal according to spatial metadata to playback a spatial audio signal in the device according to at least one sound reproduction characteristic of the device, the apparatus comprising at least one processor and at least one memory including computer program code which, when executed by the at least one processor, causes the apparatus to: Obtain the input audio signal; Separate from obtaining the input audio signal, obtain the spatial metadata, wherein The input audio signal and the spatial metadata convey a spatial audio image; Obtain the at least one sound reproduction characteristic of the device; Based on the spatial metadata, obtain a first target covariance matrix, wherein the first target covariance matrix represents sounds included in a first portion of the spatial audio image and includes sound directions within a front region of the spatial audio image; Based on the spatial metadata and the at least one sound reproduction characteristic of the device, obtain a second target covariance matrix, wherein the second target covariance matrix represents sounds included in a second portion of the spatial audio image, wherein the second portion includes sound directions not included in the first portion; and Based on the input audio signal and in accordance with the first target covariance matrix and the second target covariance matrix, render the spatial audio signal.

17. The apparatus according to claim 16, wherein The apparatus renders the spatial audio signal by: In accordance with the first target covariance matrix, using a first type of playback process applied to the input audio signal to render the first portion of the spatial audio signal; and In accordance with the second target covariance matrix, using a second type of playback process applied to the input audio signal to render the second portion of the spatial audio signal, wherein the second type of playback process is different from the first type of playback process.

18. The apparatus according to claim 16, wherein The apparatus renders the spatial audio signal by: respectively in accordance with the first target covariance matrix and the second target covariance matrix, using the first type of playback process and the second type of playback process applied to the input audio signal to render the spatial audio signal.

19. The apparatus according to claim 18, wherein The first type of playback process includes an amplitude translation process and the second type of playback process includes a crosstalk cancellation process.

20. The apparatus according to claim 16, wherein The apparatus is further caused to: Based on the input audio signal, obtain a covariance matrix and an energy metric; Obtain an extended first target covariance matrix based on the first target covariance matrix and using a gain value based on predefined equalization information included in the at least one sound reproduction characteristic; Obtain an extended second target covariance matrix based on the second target covariance matrix and a crosstalk cancellation gain; Obtain a target covariance matrix as a combination of the extended first target covariance matrix and the extended second target covariance matrix; Based on the covariance matrix and the target covariance matrix, obtain a mixing matrix that, when applied to the input audio signal, produces a modified audio signal having a covariance matrix similar to the target covariance matrix; And Obtain an output audio signal for playback by the device as a product of the input audio signal and the corresponding mixing matrix.

Citation Information

Patent Citations

  • Spatial audio signal format generation from a microphone array using adaptive capture

    WO2018060550A1