Sound field-related rendering

By determining the type of transport audio signal and applying tailored processing methods, the system addresses audio quality degradation issues in spatial audio rendering, achieving improved audio quality.

JP7704686B2Active Publication Date: 2025-07-08NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021557218
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-27
Filing Date
2020-03-19
Publication Date
2025-07-08
Estimated Expiration
2040-03-19

AI Technical Summary

Technical Problem

Existing audio processing systems face challenges in efficiently rendering spatial audio signals due to the use of inappropriate processing techniques for different types of transport audio signals, leading to audio quality degradation.

Method used

The system determines the type of transport audio signal by analyzing spatial metadata or the audio signal itself, and applies specific processing methods such as linear and parametric rendering to generate optimal prototype signals, minimizing artifacts.

Benefits of technology

This approach improves audio quality by ensuring that processing is tailored to the specific type of transport audio signal, reducing artifacts and enhancing the rendering of spatial audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704686000039
    Figure 0007704686000039
  • Figure 0007704686000040
    Figure 0007704686000040
  • Figure 0007704686000041
    Figure 0007704686000041
Patent Text Reader

Abstract

Sound field related rendering. [Solution] 1. An apparatus comprising: means configured to obtain at least two audio signals; determine types of the at least two audio signals; and process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and method for audio representation and rendering related to a sound field, but is not limited to audio representation for an audio decoder.

Background Art

[0002] Immersive audio codecs support a number of operating points, from low bitrate operation to transparency. An example of such a codec is an immersive voice and audio service (IVAS) codec designed to be suitable for use on communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of voice, music, and general-purpose audio. Furthermore, it is expected to support channel-based audio and scene-based audio inputs that include spatial information regarding the sound field and sound sources. Also, the codec is expected to support high error robustness under various transmission conditions and to operate with low latency to enable conversational services.

[0003] The input signal can be presented to the IVAS encoder in any of a number of supported formats (and also by combinations of possible formats). For example, a monaural audio signal (without metadata) can be encoded using an EVS (Enhanced Voice Service) encoder. Other input formats may utilize the IVAS encoding tools. At least some inputs can utilize the Metadata-Assisted Spatial Audio (MASA) tools or any suitable spatial metadata-based scheme. This is a parametric spatial audio format suitable for spatial audio processing. Parametric spatial audio processing is a field of audio signal processing in which the spatial aspects of sound (or sound fields) are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective choice to estimate a set of parameters such as the direction of sound in a frequency band and the ratio between the directional and omnidirectional parts of the captured sound in the frequency band from the microphone array signal. These parameters are known to well describe the perceptual spatial characteristics of the sound captured at the position of the microphone array. These parameters can then be utilized for spatial sound synthesis and other formats such as binaural headphones, loudspeakers, or ambisonics accordingly.

[0004] For example, there are two channels (stereo) of an audio signal and spatial metadata. The spatial metadata further includes a direction index (describing the direction of arrival of sound at time-frequency parameter intervals), a direction-to-total energy ratio (describing a direction index, i.e., the energy ratio for a time-frequency subframe), spread coherence (describing the energy ratio of omnidirectional sound with respect to the surrounding direction), a diffuse-to-total energy ratio (describing the coherence of omnidirectional sound with respect to the surrounding direction), surround coherence (describing the coherence of omnidirectional sound with respect to the surrounding direction), a remainder-to-total energy ratio (describing the energy ratio of residual (such as microphone noise) acoustic energy to meet the requirement that the sum of the energy ratios is 1), and distance (describing the distance of sound emitted from a direction index (i.e., a time-frequency subframe) on a logarithmic scale). Such parameters can be defined.

[0005] The IVAS stream can be decoded and rendered into various output formats such as binary, multi-channel, and Ambisonic (FOA / HOA) output.

[0006] Means configured to process at least two audio signals configured to be rendered based on a determined type of at least two audio signals can include converting at least two audio signals into an Ambisonic audio signal representation, converting at least two audio signals into a multi-channel audio signal representation, and downmixing at least two audio signals into fewer audio signals.

[0007] Means configured to process at least two audio signals configured to be rendered based on a determined type of at least two audio signals can be configured to generate at least one prototype signal based on the at least two audio signals and the type of the at least two audio signals.

[0008] According to a second aspect, there is provided a method comprising the steps of obtaining at least two audio signals, determining the type of the at least two audio signals, and processing at least two audio signals configured to be rendered based on the determined type of the at least two audio signals.

[0009] The at least two audio signals can be one of a carrier audio signal and a previously processed audio signal.

[0010] The method can further comprise obtaining at least one parameter associated with the at least two audio signals.

[0011] Determining the type of the at least two audio signals can include determining the type of the at least two audio signals based on at least one parameter associated with the at least two audio signals.

[0012] Determining the type of the at least two audio signals based on at least one parameter can include one of extracting and decoding at least one type of signal from the at least one parameter and, when the at least one parameter represents a spatial audio aspect associated with the at least two audio signals, analyzing the at least one parameter to determine the type of the at least two audio signals.

[0013] Analyzing at least one parameter to determine the types of the at least two audio signals may include determining a broadband left or right channel to total energy ratio based on the at least two audio signals, determining a higher frequency left or right channel to total energy ratio based on the at least two audio signals, determining a sum of the total to total energy ratio based on the at least two audio signals, determining a subtraction to target energy ratio based on the at least two audio signals, and determining the types of the at least two audio signals based on at least one of the broadband left or right channel to total energy ratio, the higher frequency left or right channel to total energy ratio based on the at least two audio signals, the sum of the total to total energy ratio based on the at least two audio signals, and the subtraction to target energy ratio.

[0014] The method of the present application may further include determining at least one type parameter related to the type of at least one audio signal.

[0015] Processing at least two audio signals configured to be rendered based on the determined types of the at least two audio signals may further include converting the at least two audio signals based on at least one type parameter related to the types of the at least two audio signals.

[0016] The types of the at least two audio signals may include at least one of a capture microphone arrangement, a capture microphone separation distance, a capture microphone parameter, a transport channel identifier, an interleaved audio signal type, a downmix audio signal type, an identical audio signal type, and a transport channel arrangement.

[0017] Processing the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals may include one of converting the at least two audio signals to an Ambisonic audio signal representation, converting the at least two audio signals to a multi-channel audio signal representation, and downmixing the at least two audio signals to fewer audio signals.

[0018] Processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals may include generating at least one prototype signal based on the at least two audio signals and the types of the at least two audio signals.

[0019] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory containing computer program code, the at least one memory and the at least one computer program code configured to cause, using the at least one processor, the apparatus to at least acquire at least two audio signals, determine types of the at least two audio signals, and process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.

[0020] The at least two audio signals may be one of a carrier audio signal and a previously processed audio signal.

[0021] The means may be configured to obtain at least one parameter relating to the at least two audio signals.

[0022] An apparatus configured to determine the types of at least two audio signals can determine the types of the at least two audio signals based on at least one parameter associated with the at least two audio signals.

[0023] An apparatus for determining the types of at least two audio signals based on the at least one parameter can extract and decode at least one type signal from the at least one parameter, and when the at least one parameter represents a spatial audio mode associated with the at least two audio signals, can perform one of analyzing the at least one parameter to determine the types of the at least two audio signals.

[0024] An apparatus for analyzing at least one parameter for determining the types of at least two acoustic signals determines a broadband left or right channel to total energy ratio based on the at least two acoustic signals, determines a higher frequency or right channel to total energy ratio based on the at least two acoustic signals, determines a sum to total energy ratio based on the at least two acoustic signals, determines a subtraction to total energy ratio based on the at least two acoustic signals, and can determine the types of the at least two acoustic signals based on at least one of the broadband left or right channel to total energy ratio, the higher frequency left or right channel to total energy ratio based on the at least two acoustic signals, the sum to total energy ratio based on the at least two acoustic signals, and the subtraction to target energy ratio.

[0025] The apparatus can determine at least one type parameter associated with the type of at least one audio signal.

[0026] An apparatus that processes at least two audio signals configured to be rendered based on a determined type of the at least two audio signals can convert the at least two audio signals based on at least one type parameter related to the type of the at least two audio signals.

[0027] The type of the at least two audio signals can include at least one of a capture microphone arrangement, a capture microphone separation distance, a capture microphone parameter, a transport channel identifier, an interleaved audio signal type, a downmix audio signal type, an identical audio signal type, and a transport channel arrangement.

[0028] The apparatus can process at least two audio signals configured to be rendered based on a determined type of the at least two audio signals, convert the at least two audio signals into an ambisonic audio signal representation, convert the at least two audio signals into a multi-channel audio signal representation, and downmix the at least two audio signals into fewer audio signals.

[0029] The apparatus of the present application can process at least two audio signals configured to be rendered based on a determined type of the at least two audio signals, and generate at least one prototype signal based on the at least two audio signals and the type of the at least two audio signals.

[0030] According to a fourth aspect, there is provided an apparatus including: a step of obtaining a circuit configured to obtain at least two audio signals; a determination circuit configured to determine the type of the at least two audio signals; and a processing circuit configured to process the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals.

[0031] According to a fifth aspect, there is provided a computer program (or a computer-readable medium including program instructions) including instructions for causing an apparatus to at least perform acquiring at least two audio signals, determining types of the at least two audio signals, and processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.

[0032] According to a sixth aspect, there is provided a non-transitory computer-readable medium including program instructions for causing an apparatus to at least perform acquiring at least two audio signals, determining types of the at least two audio signals, and processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.

[0033] According to a seventh aspect, there is provided an apparatus including means for acquiring at least two audio signals, means for determining types of the at least two audio signals, and means for processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.

[0034] According to an eighth aspect, there is provided a computer-readable medium including program instructions for causing an apparatus to perform acquiring at least two audio signals, determining types of the at least two audio signals, and processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.

[0035] An apparatus including means for performing the operations of the method described above.

[0036] An apparatus configured to perform the actions of the method described above.

[0037] A computer program including program instructions for causing a computer to execute the above-described method.

[0038] A computer program product stored on a medium can cause an apparatus to execute the method described herein.

[0039] An electronic device can include the apparatus described herein.

[0040] A chipset can include the apparatus described herein.

[0041] Embodiments of the present invention are aimed at addressing problems related to the state of the art.

Brief Description of the Drawings

[0042] To enhance the understanding of the present application, reference will be made here to the accompanying drawings by way of example.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

[0043] In the following, a suitable apparatus and possible mechanisms for providing efficient rendering of spatial metadata-aided audio signals will be described in more detail.

[0044] With respect to FIG. 1, an example of an apparatus and a system for realizing audio capture and rendering is shown. System 100 is shown as including an "analysis" unit 121 and a "demultiplexer / decoder / synthesizer" unit 133. The "analysis" unit 121 is the part from receiving a multi-channel loudspeaker signal to encoding metadata and a carrier signal, and the "demultiplexer / decoder / synthesizer" unit 133 is the part from decoding the encoded metadata and carrier signal to presenting the regenerated signal (e.g., forming a multi-channel loudspeaker).

[0045] The input to system 100 and the "analysis" part 121 is a multi-channel signal 102. In the following example, a microphone channel signal input is described, but in other embodiments, any suitable input (or synthesized multi-channel) format can be realized. For example, in some embodiments, the spatial analyzer and spatial analysis may be performed outside the encoder. For example, in one embodiment, the spatial metadata related to the audio signal may be provided to the encoder as a separate bitstream. In one embodiment, the spatial metadata may be provided as a set of spatial (direction) index values.

[0046] The multi-channel signal is passed to a carrier signal generator 103 and an analysis processor 105.

[0047] In some embodiments, the carrier signal generator 103 is configured to receive a multi-channel signal, generate an appropriate carrier signal including a determined number of channels, and output the carrier signal 104. For example, the transport signal generator 103 can be configured to generate a two-audio-channel downmix of the multi-channel signal. The determined number of channels can be any appropriate number of channels. The carrier signal generator in some embodiments is configured to select or combine the input audio signals into a determined number of channels and output them as carrier signals, for example, by beamforming techniques.

[0048] In some embodiments, the carrier signal generator 103 is optional, and the multi-channel signal is passed untreated to the "encoder / MUX" block 107, similar to this example where the carrier signal is present.

[0049] In some embodiments, the analysis processor 105 is also configured to receive the multi-channel signal, analyze the signal, and generate metadata 106 related to the multi-channel signal and thus related to the carrier signal 104. The analysis processor 105 can be configured to generate metadata including a direction parameter 108, an energy ratio parameter 110 (an example of which is a diffuseness parameter), and a coherence parameter 112 for each time-frequency analysis interval. The direction, energy ratio, and coherence parameters can be considered spatial audio parameters in an embodiment. In other words, the spatial audio parameters include parameters aimed at characterizing the sound field created by the multi-channel signal (or generally two or more reproduced audio signals).

[0050] In some embodiments, the generated parameters may vary for each frequency band. Thus, for example, in band X, all parameters are generated and transmitted, while in band Y, only one parameter is generated and transmitted, and in band Z, no parameters are generated and transmitted. As a practical example, in some frequency bands such as the highest band, some parameters may be unnecessary for perceptual reasons. The transport signal 104 and the metadata 106 can be passed to the "encoder / MUX" block 107.

[0051] In some embodiments, the spatial audio parameters may be grouped or separated into directional and non-directional (e.g., diffuse) parameters.

[0052] The "encoder / MUX" block 107 can be configured to receive the transport (e.g., downmix) signals 104 and generate appropriate encoding of these audio signals. The "encoder / MUX" block 107 can be, in some embodiments, a computer (executing appropriate software stored on a memory and at least one processor), or alternatively, a specific device that utilizes, for example, an FPGA or an ASIC. The encoding can be performed using any suitable scheme. The "encoder / MUX" block 107 can further be configured to receive the metadata and generate an encoded or compressed form of the information. In some embodiments, the "encoder / MUX" block 107 can interleave, multiplex, or embed the metadata by a dashed line into the single data stream 111 prior to the transmission or storage shown in FIG. 1. The multiplexing can be performed using any suitable scheme.

[0053] On the decoder side, the received or retrieved data (stream) may be received by the "demultiplexer / decoder / synthesizer" 133. The "demultiplexer / decoder / synthesizer" 133 can demultiplex the encoded stream, decode the audio signal, and obtain the transport signal. Similarly, the "demultiplexer / decoder / synthesizer" 133 may be configured to receive and decode the encoded metadata. In some embodiments, the "demultiplexer / decoder / synthesizer" 133 can be a computer (executing appropriate software stored in memory and on at least one processor), or alternatively, a specific device that utilizes, for example, an FPGA or an ASIC.

[0054] The "demultiplexer / decoder / synthesizer" portion 133 of system 100 may further be configured to recreate the synthesized spatial audio in the form of the multi-channel signal 110 in any suitable format based on the transport signal and the metadata (these can be in a multi-channel loudspeaker format, and in certain embodiments, depending on the use case, can be in any suitable output format such as a binaural signal or an Ambisonics signal for headphone listening).

[0055] Thus, first in summary, the system (analysis part) is set to receive a multi-channel audio signal.

[0056] Next, the system (analysis part) is set to generate an appropriate carrier audio signal (for example, by selecting or downmixing a part of the audio signal channels).

[0057] Next, the system is configured to encode for storing / transmitting the transport signal and the metadata.

[0058] After that, the system can store / transmit the encoded transport and metadata.

[0059] The system can retrieve / receive the encoded transport and metadata.

[0060] Next, the system is configured to extract the transport and metadata from the encoded transport and metadata parameters, for example, demultiplex, and decode the encoded transport and metadata parameters.

[0061] The system (synthesis unit) is configured to synthesize an output multi-channel audio signal based on the extracted transport audio signal and metadata. For the decoder (synthesis part), it is configured to receive the spatial metadata and transfer an audio signal (potentially a pre-processed version) that can be, for example, a downmix of a 5.1 signal, two spaced microphone signals from a mobile device, or two beam patterns from a matching microphone array.

[0062] The decoder may be configured to render spatial audio (such as ambisonics) from the spatial metadata and the transport audio signal. This is typically achieved by adopting one of two approaches, linear and parametric rendering, to render spatial audio from such inputs.

[0063] Assuming processing in the frequency band, linear rendering means utilizing some static mixing weights to generate the desired output. Parametric rendering is to modify the transport audio signal based on the spatial metadata and generate the target output.

[0064] Methods for generating ambisonics from various inputs are presented.

[0065] 5.1. For a carried audio signal and spatial metadata from a signal, ambisonics can be rendered using parametric processing.

[0066] When carrying an audio signal or spatial metadata from a distant microphone, a combination of linear processing and parametric processing can also be used.

[0067] For a carried audio signal and spatial metadata from simultaneous microphones, a combination of linear processing and parametric processing can be used.

[0068] Therefore, there are various ways to render ambisonics from various types of inputs. However, all certain ambisonic rendering methods assume a certain type of input. Some of the embodiments described below show apparatuses and methods for preventing the occurrence of the following problems.

[0069] Using linear rendering, the Y signal, which is the leftward first-order (8-digit) signal of ambisonics, can be created from two matching opposite cardioid microphones by Y(f) = S0(f) - S1(f). Here, f is the frequency. As another example, the Y signal can be created by Y(f) = -i(S0(f) - S1(f))g eq (f). Here, g eq (f) is a frequency-dependent equalizer (depending on the distance of the microphone), and i is the imaginary unit. The processing of spaced microphones (including a -90-degree phase shift and frequency-dependent equalization) is different from the processing of matching microphones, and incorrect processing techniques may degrade the sound quality.

[0070] To use parametric rendering with some rendering schemes, it is necessary to generate "prototype" signals using linear averaging. These prototype signals are then adaptively modified in the time-frequency domain based on spatial metadata. Optimally, the prototype signals should follow the target signal as closely as possible. This minimizes the need for parametric processing and thus minimizes potential artifacts due to parametric processing. For example, the prototype signal should contain all signal components related to the corresponding output channel within a sufficient range.

[0071] As an example, when an omnidirectional signal W is rendered (the same effect exists for other ambisonic signals), the prototype can be created from the stereo carrier audio signal in, for example, two simple approaches. Select one channel (such as the left channel), or the sum of the two channels.

[0072] Which one to choose depends largely on the type of the carrier audio signal. When the carrier signal is generated from a 5.1 signal, usually, the left signal is only the left carrier audio signal and the right signal is only the right carrier audio signal (when using a common downmix matrix). Therefore, if one channel is used for the prototype, the signal content of the other channel is lost and distinct artifacts are generated (for example, in the worst case, there is no signal in the selected single channel). Thus, in this case, it would be better to formulate the W prototype as the sum of both channels. On the other hand, when the carrier signal is generated from widely separated microphones, using the sum of the carrier audio signals as the prototype for the W signal results in severe comb filtering (due to the time delay between the signals). This generates artifacts similar to the above. In this case, it is better to select only one of the two channels as the W prototype, at least in the high-frequency range.

[0073] Therefore, there is no suitable option that is compatible with all transport audio signal types.

[0074] Therefore, applying spatial audio processing designed for one transport audio signal type to another transport audio signal type using both linear and parametric methods is expected to produce a distinct degradation in audio quality.

[0075] Concepts, as will be discussed in more detail with respect to the following embodiments and examples, relate to audio encoding and decoding when a decoder receives at least two transport audio signals from an encoder. Further, embodiments may have transport audio signals that can be at least two types, e.g., a downmix of a 5.1 signal, spaced microphones signals, or coincident microphone signals. Further, in some embodiments, the apparatus and method implement solutions for improving the quality of processing of transport audio signals and providing a determined output (e.g., ambisonic, 5.1, mono). The quality can be improved by determining the type of transport audio signal and performing audio processing based on the determined type of transport audio signal.

[0076] In some embodiments discussed in more detail herein, the transport audio signal type is determined either by obtaining metadata indicating the type of the transport audio signal or by determining the type of the transport audio signal based on the transport audio signal itself (and spatial metadata if available).

[0077] Metadata describing the transport audio signal type can include conditions such as spaced microphones (which may be associated with the position of the microphones), coincident microphones or columns that are substantially similar to coincident microphones (which may be accompanied by a microphone direction pattern), a downmix from a multi-channel audio signal (such as 5.1).

[0078] The determination of the type of the carried audio signal based on the analysis of the carried audio signal itself can be based on comparing the frequency bands or spectral effects to be combined (in different ways) with the expected spectral effects (partially based on the available spatial metadata if available).

[0079] Furthermore, in some embodiments, the processing of the audio signal can include the rendering of an Ambisonic signal, the rendering of a multi-channel audio signal (such as 5.1), and the transport of a downmix to a smaller number of audio signals:

[0080] FIG. 2 shows a schematic diagram of an example decoder suitable for implementing some embodiments. This embodiment can be implemented, for example, within the "demultiplexer / decoder / synthesizer" block 133. In this example, the input is a metadata-assisted spatial audio (MASA) stream including two audio channels and spatial metadata. However, as discussed herein, the input format can be any suitable metadata-assisted spatial audio format.

[0081] (MASA) The bitstream is transferred to the carried audio signal type determiner 201. The carried audio signal type determiner 201 is configured to determine the carried audio signal type 202 and, optionally, some additional parameters 204 (such as microphone distance) based on the bitstream. The determined parameters are transferred to the MASA-Ambisonic signal converter 203.

[0082] The MASA-Ambisonic signal converter 203 is configured to receive the bitstream and the carried audio signal type 202 (and optionally some additional parameters 204), and is configured to convert the MASA stream into an Ambisonic signal based on the determined carried audio signal type 202 (and possible additional parameters 204).

[0083] The operation of the example is summarized in the flow diagram shown in FIG. 3.

[0084] The first operation is one of receiving or obtaining a bit stream (MASA stream) as shown in FIG. 3 by step 301.

[0085] The next operation is one of determining the transport audio signal type based on the bit stream (and generating a type signal or indicator and possible other additional parameters) as shown in FIG. 3 by step 303.

[0086] The next operation after determining the transport audio signal type is to convert the bit stream (MASA stream) into an ambisonic signal based on the determined transport audio signal type as shown in FIG. 3 by step 305.

[0087] FIG. 4 shows a schematic diagram of an example of a transport audio signal type discriminator 201. In this example, an example of a transport audio signal type determiner is suitable when the transport audio signal type is available in the MASA stream.

[0088] The example of the transport audio signal type determiner 201 in this example includes a transport audio signal type extractor 401. The transport audio signal type extractor 401 is configured to receive a bit (MASA) stream and extract (i.e., read and / or decode) a type indicator from the MASA stream. This type of information is available, for example, in the "channel audio format" field of the MASA stream. In addition, if additional parameters are available, they are also extracted. This information is output from the transport audio signal type extractor 401. In one embodiment, the transport audio signal type can include "space", "downmix", "match". In some other embodiments, the transport audio signal type can include any suitable value.

[0089] FIG. 5 shows a schematic diagram of a transport audio signal type determiner 201 as a further example. In this example, the transport audio signal type cannot be directly extracted or decoded from the MASA stream. In this example, the transport audio signal type is estimated or determined from the analysis of the MASA stream. This determination in some embodiments is based on using a set of estimators / energy comparisons that reveal spectral effects for different transport audio signal types.

[0090] In one embodiment, the transport audio signal type determiner 201 includes a transport audio signal and spatial metadata extractor / decoder 501. The transport audio signal and spatial metadata extractor / decoder 501 is configured to receive the MASA stream and extract and / or decode the transport audio signal and spatial metadata from the MASA stream. The resulting transport audio signal 502 can be transferred to a time / frequency converter 503. The resulting spatial metadata 522 can further be transferred for subtraction to a target energy comparator 511.

[0091] In some embodiments, the transport audio signal type determiner 201 includes a time / frequency converter 503. The time / frequency converter 503 is configured to receive the transport audio signals 502 and convert them to the time-frequency domain. Suitable conversions include, for example, the short-time Fourier transform (STFT) and the complex modulated quadrature mirror filter bank (QMF). The resulting signal is represented as S i (i, b, n). Here, i is the channel index, b is the frequency bin index, and n is the time index. In situations where the transport audio signals (output from the extractor and / or decoder) are already in the time-frequency domain, this may be omitted or may include a conversion from one time-frequency domain representation to another. The T / F domain transport audio signal 504 can be transferred to a comparator.

[0092] In one embodiment, the transport audio signal type determiner 201 includes a broadband L / R total energy comparator 505. The broadband L / R pair total energy comparator 505 is configured to receive the T / F domain transport audio signal 504 and output the broadband L / R with respect to the total ratio parameter.

[0093] Within the broadband L / R to total energy comparator 505, the broadband left, right, and total energies are calculated.

Number

Number

Number

[0094] Next, the broadband L / R pair total energy comparator 505 can generate a broadband L / R pair total energy ratio 506 as follows.

Number

[0095] In some embodiments, the transport audio signal type determiner 201 includes a high-frequency L / R-total energy comparator 507. The high-frequency L / R-total energy comparator 507 is configured to receive the T / F domain transport audio signal 504 and output a high-frequency L / R-total ratio parameter.

[0096] Within the broadband L / R-total energy comparator 507, the left, right, and total energies in the high-frequency band are calculated.

Number

Number

[0097] Next, the high-frequency L / R-to-total energy comparator 507 can be configured to select the smaller of the left and right energies, and the result is multiplied by 2.

Number

[0098] Next, the high-frequency L / R-to-total energy comparator 507 can then generate a high-frequency L / R-to-total ratio 508.

Number

[0099] In some embodiments, the transport audio signal type determiner 201 includes a total energy comparator 509. The sum to the total energy comparator 509 is configured to receive the T / F domain transport audio signal 504 and output a sum for the total energy ratio parameter. The sum to the total energy comparator 509 is configured to detect, at several frequencies, a situation where two channels are out of phase, which is a typical phenomenon especially for spaced microphone recordings.

[0100] The sum to the total energy comparator 509 is configured to calculate the energy of the total signal and the total energy for each frequency bin.

Number

[0101] These energies are smoothed, for example,

Number

[0102] Next, the total energy comparator 509 is configured to calculate the minimum total ratio 510 as follows.

Number

[0103] Next, the sum to the total energy comparator 509 is configured to output the ratio χ(n) 510.

[0104] In some embodiments, the transport audio signal type determiner 201 includes subtraction to a target energy comparator 511. The subtraction to the target energy comparator 511 is configured to receive a T / F domain transport audio signal 504 and spatial metadata 522 and output subtraction to a target energy ratio parameter 512.

[0105] The subtraction to the target energy comparator 511 is configured to calculate the energy of the difference between the left and right channels.

Number

[0106] This can be considered as a “prototype” of the ambisonic Y signal for at least some input signal types (the Y signal has a dipole directional pattern with a positive lobe on the left and a negative lobe on the right).

[0107] Next, the subtraction to the target energy comparator 511 can be configured to calculate the target energy E target (b,n) for the Y signal. This is based on estimating how the total energy should be distributed among the spherical harmonics based on the spatial metadata. For example, in some embodiments, the subtraction to the target energy comparator 511 is configured to construct a target covariance matrix (channel energy and cross-correlation) based on the spatial metadata and the energy estimate. However, in some embodiments, only the energy of the Y signal is estimated, which is one entry of the target covariance matrix. Thus, the target energy E target (b,n) of Y consists of two parts.

Number

Number

[0108] Note that the spatial metadata can have a lower frequency and / or time resolution than per b,n so that the parameters can be the same for some frequency or time indices.

[0109] This E (target,dir) (b,n) is the energy of the more directional part. To formulate it, the spread coherence c of the spatial metadata spread The spread coherence distribution vector as a function of the (b,n) parameter from 0 to 1 is

Number

[0110] The subtraction to the target energy comparator 511 is a vector of azimuth values,

Number

Number

[0111] Therefore, E target (b,n) is obtained. These energies, in some embodiments, for example,

Number

[0112] Furthermore, the subtraction to the target energy comparator 511 is configured to calculate the subtraction to the target ratio 512 using the energy in the lowest frequency bin as follows.

Number

[0113] In one embodiment, the carrier audio signal type determiner 201 includes a carrier audio signal type (based on the estimated metric) determiner 513. The carrier audio signal type determiner 513 receives the broadband L / R with respect to the total ratio 506, the high frequency L / R with respect to the total ratio 508, the sum of the parts with respect to the total ratio 510, and the subtraction with respect to the target ratio 512, and is configured to determine the carrier audio signal type based on these estimated metrics.

[0114] The determination can be made in various ways, and the actual implementation can vary in many aspects, such as the T / F conversion used. An example of a non-limiting form is that the carrier audio signal type (based on the estimated metric) determiner 513 first calculates the change to the non-metric.

Number

[0115] The transport audio signal type (based on the estimated metric) determiner 513 can then be configured to calculate the change to the downmix metric.

Number

[0116] The transport audio signal type (based on the estimated metrics) determiner 513 can then, based on these metrics, determine whether the transport audio signal is generated from spaced microphones or is a downmix from a surround sound signal (such as 5.1). For example,

Number

[0117] In this example, the transport audio signal type (based on the estimated metrics) determiner 513 does not detect a matching microphone type. However, in practice, the processing according to the T(n) = "downmix" type can generally generate good audio in the case of a matched capture (for example, when using cardioid microphones directed left and right).

[0118] The transport audio signal type (based on the estimated metric) determiner 513 can then be configured to output the transport audio signal type as the transport audio signal type 202. In some embodiments, other parameters 204 may also be output.

[0119] FIG. 6 summarizes the operation of the apparatus shown in FIG. 5. Thus, in some embodiments, the first operation is, as shown in FIG. 6 by step 601, the operation of extracting and / or decoding the transport audio signal and metadata from the MASA stream (or bitstream).

[0120] The next operation can convert the transmitted audio signal into the time-frequency domain as shown in FIG. 6 by step 603.

[0121] Next, a series of comparisons can be performed. For example, by comparing the broadband L / R energy with the total energy value, a broadband L / R to total energy ratio can be generated as shown in FIG. 6 by step 605.

[0122] For example, by comparing the high-frequency L / R energy with the total energy value, a high-frequency L / R to total energy ratio can be generated as shown in FIG. 6 by step 607.

[0123] By comparing the total energy with the total energy value, the total to total energy ratio may be generated by step 609 as shown in FIG. 6.

[0124] Furthermore, a subtraction to target energy ratio may be generated as shown in FIG. 6 by step 611.

[0125] After determining these metrics, the method can determine the transmitted audio signal type by analyzing these metric ratios as shown in FIG. 6 by step 613.

[0126] FIG. 7 shows an example of the MASA to ambisonic converter 203 in more detail. The MASA to ambisonic converter 203 is configured to receive a MASA stream (bit stream) and a transmitted audio signal type 202 and possible additional parameters 204, and is configured to convert the MASA stream into an ambisonic signal based on the determined transmitted audio signal type.

[0127] The MASA-to-ambisonic converter 203 includes a transport audio signal and spatial metadata extractor / decoder 501. This is configured to receive a MASA stream and output a transport audio signal 502 and spatial metadata 522 in the same manner as seen within a transport audio signal type determiner, as shown in FIG. 5. In some embodiments, the extractor / decoder 501 is an extractor / decoder from a transport audio signal type determiner. The resulting transport audio signal 502 can be transferred to a time / frequency converter 503. The resulting spatial metadata 522 can further be transferred to a signal mixer 705.

[0128] In certain embodiments, the MASA-to-ambisonic converter 203 includes a time / frequency converter 503. The time / frequency converter 503 is configured to receive the transport audio signals 502 and convert them into the time-frequency domain. Suitable conversions include, for example, the short-time Fourier transform (STFT) and the complex modulated quadrature mirror filter bank (QMF). The resulting signals are represented as S i (i, b, n). Here, i is the channel index, b is the frequency bin index, and n is the time index. If the output of the audio extraction and / or decoding is already in the time-frequency domain, this block may be omitted or may include a conversion from one time-frequency domain representation to another. The T / F domain transport audio signal 504 can be transferred to a prototype signal creator 701. In some embodiments, the time / frequency converter 503 is the same time / frequency converter from a transport audio signal type determiner.

[0129] In one embodiment, the MASA-to-ambisonic converter 203 includes a prototype signal creator 701. The prototype signal creator 701 is configured to receive a T / F domain carrier audio signal 504, a carrier audio signal type 202, and optional additional parameters 204. The T / F prototype signal 702 can then be output to a signal mixer 705 and a decorrelator 703.

[0130] In one embodiment, the MASA-to-ambisonic converter 203 includes a decorrelator 703. The decorrelator 703 is configured to receive the T / F prototype signal 702, apply decorrelation (uncorrelation), and output a decorrelated T / F prototype signal 704 to the signal mixer 705. In some embodiments, the decorrelator 703 is optional.

[0131] In one embodiment, the MASA-to-ambisonic converter 203 includes a signal mixer 705. The signal mixer 705 is configured to receive the T / F prototype signal 702, the uncorrelated T / F prototype signal, and the spatial metadata 522.

[0132] The prototype signal creator 701 is configured to generate a prototype signal for each of the spherical harmonic functions of ambisonics (FOA / HOA) based on the carrier audio signal type.

[0133] In some embodiments, the prototype signal creator 701 is configured to operate as follows. If T(n) = "spaced", the prototype of the W signal can be created as

Equation

[0134] In practice, W proto(b,n) can be created as a means of carrying a low-frequency audio signal. The phases of the signals are approximately in-phase, and no comb filtering is performed. Also, one of the high-frequency channels is selected. The value of B3 varies depending on the T / F conversion and the distance between the microphones. If the distance is unknown, some default values may be used (such as the value corresponding to 1 kHz). If T(n) = "downmix" or T(n) = "coincident", the prototype of the W signal can be created as follows.

Number

[0135] Since the original audio signal is usually assumed to have no significant delay between these signal types, W proto (b,n) is created by summing the carrier audio signals.

[0136] Regarding the Y prototype signal, if T(n) = "spaced", the prototype of the Y signal can be created as follows.

Number

[0137] In the mid-frequency range (between B4 and B5), a dipole signal can be created by subtracting the transport signal, shifting the phase by -90 degrees, and equalizing. Therefore, especially if the microphone distance is known, it serves as a good prototype for the Y signal, and thus the equalization coefficient is appropriate. This is not achievable at low and high frequencies, and the prototype signal is generated in the same way as in the case of an omnidirectional W signal.

[0138] When the microphone distance is known exactly, the Y prototype may be used directly for Y at those frequencies (i.e., Y(b,n) = Y proto (b,n)). If the microphone spacing is unknown, g eq (b) = 1 can be used.

[0139] In some embodiments, signal mixer 705 applies gain processing in a frequency band to correct the energy of W(b,n) in the frequency band to the target energy in the frequency band using potential gain smoothing. proto (b,n) can correct the energy of (b,n). The target energy of an omnidirectional signal in a certain frequency band can be the sum of the carrier audio signal energies in that frequency band. As a result of this processing, an omnidirectional signal W(b,n) is obtained.

[0140] Y proto For Y signals where Y(b,n) cannot be used directly as Y(b,n), if the frequency is between B4 and B5, adaptive gain processing is performed. In this case, it is similar to the case of the above omnidirectional W. The prototype signal is already a Y dipole, except potentially for an incorrect spectrum. The signal mixer performs gain processing on the prototype signal in the frequency band. (Furthermore, in this particular context, non-correlation processing of the Y signal is not required). The gain processing uses spatial metadata (direction, ratio, other parameters) and an overall signal energy estimate in the frequency band (e.g., the sum of the carrier signal energies) to determine what the energy of the Y component should be within the frequency band, then corrects the energy of the prototype signal within the frequency band, which is the determined energy, by a gain, and then the result becomes the output Y(b,n).

[0141] The procedure for generating the aforementioned Y(b,n) is not valid for all frequencies in the current context T(n) = "spaced". Since the prototype signals are different at different frequencies, the signal mixer and decorrelator will have different configurations depending on the frequency with this transport signal type. To explain different types of prototype signals, we can consider a scenario where sound arrives from the negative gain direction of the Y dipole (with positive and negative lobes). At medium frequencies (between B4 and B5), the phase of the Y prototype signal should be opposite to that of the W prototype signal because of the direction of the incoming sound. At other frequencies (below B4 and above B5), the phase of the prototype Y signal is the same as that of the W prototype signal. The appropriate synthesis of phases (and correlation with energy) is then explained by the signal mixer and decorrelator at those frequencies.

[0142] At low frequencies (below B4) where the wavelength is large, the phase difference between audio signals captured by spaced microphones (usually slightly closer to each other) becomes small. Therefore, the creator of the prototype signal should not be set to generate the prototype signal in the same way as the frequencies between B4 and B5 for SNR reasons. Therefore, typically, a channel sum omnidirectional signal is used as the prototype signal instead. At high frequencies (above B5) where the wavelength is small, the beam pattern is severely distorted by spatial aliasing (when the same method as the frequencies between B4 and B5 is used). Therefore, it is better to use an omnidirectional prototype signal for channel selection.

[0143] Next, the configurations of the signal mixer and decorrelator at these frequencies (below B4 or above B5) will be described. In a simple example, the spatial metadata parameter settings are composed of the azimuth θ and ratio r of the frequency band. Apply the gain sin(θ)sqrt(r) to the prototype signal in the signal mixer to generate the Y dipole signal, and the result becomes the coherent partial signal. The prototype signal is also decorrelated (in the decorrelator), and the decorrelated result is received by the signal mixer. Here, the coefficient sqrt(1 - r)gorder is multiplied, and the result is an uncorrelated partial signal. The gain g order is the diffusion field gain at the spherical harmonic order according to the known SN3D normalization method. For example, in the case of the first order (in this case, the case of the Y dipole), it is sqrt(1 / 3), in the case of the second order, it is sqrt(1 / 5), and in the case of the third order, it is sqrt(1 / 7). The coherent partial signal and the incoherent partial signal are added. As a result, since the prototype signal energy may be incorrect, an incorrect energy is removed, and a synthesized Y signal is obtained. The same energy correction procedure in the frequency band described in the context of the intermediate frequency (between B4 and B5) is applied to correct the energy in the frequency band to a desired target, and the output is the signal Y(b,n).

[0144] For other spherical harmonics such as the X, Z components and components of the second order and above, the above procedure can be applied except that the gain regarding the azimuth (and other potential parameters) depends on which spherical harmonic signal is being synthesized. For example, the gain generated for the X dipole coherent part from the W prototype is cos(θ)sqrt(r). The uncorrelated, ratio - processing, and energy correction can be the same as those determined above for the Y component other than the frequency between B4 and B5.

[0145] Other parameters such as height, spread coherence, and surround coherence can be considered in the above procedure. For the spread coherence parameter, a value from 0 to 1 can be specified. A coherence spread value of 0 indicates a point sound source. In other words, when playing an audio signal using a multi-loudspeaker system, the sound should be reproduced by as few loudspeakers as possible (for example, only the central loudspeaker if the direction is central). As the value of the diffusion coherence increases, up to a value of 0.5, more energy is diffused by the other loudspeakers around the center loudspeaker, and the energy is evenly diffused between the center and adjacent loudspeakers. When the value of the diffusion coherence increases above 0.5, the energy of the center loudspeaker decreases until it reaches a value of 1, there is no energy in the center loudspeaker, and all the energy is in the adjacent loudspeakers. The value of the surround coherence parameter is from 0 to 1. A value of 1 means that there is coherence between all (or almost all) loudspeaker channels. A value of 0 means that there is no coherence between all (or almost all) loudspeaker channels. This is further described in GB application No. 1718341.9 and, in addition, in PCT application PCT / FI2018 / 050788.

[0146] For example, increased surround coherence can be implemented by a decrease in the combined ambience energy in the spherical harmonic components, and elevation can be added by adding an elevation-related gain according to the definition of the ambisonic pattern in the generation of the coherent part.

[0147] If T(n) = "downmix" or T(n) = "coincident", the prototype of the Y signal can be

Number

[0148] In this situation, since it can be assumed that the original audio signal usually has no significant delay between these signal types, there is no need for phase shift. Regarding the "mixed signal" block, when T(n) = "coincident", the prototypes of Y and W may, depending on the actual directivity pattern, sometimes be directly used for the outputs of Y and W after gain. When T(n) = "downmix", Y proto (b,n) and W proto (b,n) cannot be directly used for Y(b,n) and W(b,n). However, energy correction in the frequency band to the desired target determined when T(n) = "spaced" may be required (note that the omnidirectional component has a spatial gain of 1 regardless of the angle of the incoming sound).

[0149] For other spherical harmonic functions (such as X and Z), prototypes that can reproduce the target signal well cannot be created. This is because a typical downmix signal is oriented along the left - right axis rather than the front - back X - axis or the top - bottom Z - axis. Thus, in some embodiments, the approach is to utilize, for example, a prototype of an omnidirectional signal. [Number]

[0150] Similarly, W proto (b,n) is also used for higher harmonics for the same reason. Signal mixers and decorrelators in such situations can process signals for these spherical harmonic components in the same way as when T(n) = "spaced".

[0151] In some cases, the type of the transport audio signal may change during audio playback (e.g., due to an actual signal type change or an incompleteness in automatic type detection). To avoid artifacts due to rapidly changing types, the prototype signal in some embodiments can be interpolated. This may be achieved, for example, by linearly interpolating simply from a prototype signal corresponding to an older type to a prototype signal corresponding to a newer type.

[0152] The output of the signal mixer is the obtained time-frequency domain ambisonic signal, which is transferred to the inverse T / F transformer 707.

[0153] In some embodiments, the MASA-ambisonic signal converter 203 includes an inverse T / F transformer 707 configured to convert the signal into the time domain. The time domain ambisonic signal 906 is the output from the MASA-ambisonic signal converter.

[0154] Regarding FIG. 8, an overview of the operation of the apparatus shown in FIG. 7 is shown.

[0155] Thus, in one embodiment, the first operation is the operation of extracting and / or decoding the transport audio signal and the metadata from the MASA stream (or bit stream) as shown in FIG. 8 by step 801.

[0156] The next operation can transform the transport audio signal into the time-frequency domain as shown in FIG. 8 by step 803.

[0157] Then, the method includes creating a prototype audio signal based on the transport signal in the time-frequency domain, and further creating a prototype audio signal based on the type of the transport audio signal (and further based on additional parameters) as shown in FIG. 8 by step 805.

[0158] In some embodiments, the method includes applying decorrelation on a time-frequency prototype audio signal as shown in FIG. 8 by step 807.

[0159] Then, by step 809, as shown in FIG. 8, based on the spatial metadata and the carrier audio signal type, the uncorrelated time-frequency prototype audio signal and the time-frequency prototype audio signal can be mixed.

[0160] Then, the mixed signal may be inverse time-frequency transformed as shown in FIG. 8 by step 811.

[0161] Then, by step 813, as shown in FIG. 8, a time domain signal can be output.

[0162] FIG. 9 shows a schematic diagram of an example decoder suitable for implementing some embodiments. This example can be implemented, for example, within the "demultiplexer / decoder / synthesizer" block 133 shown in FIG. 1, where in this example, the input is a metadata-assisted spatial audio (MASA) stream including two audio channels and spatial metadata. However, as discussed herein, the input format can be any suitable metadata-assisted spatial audio format.

[0163] (MASA) The bitstream is transferred to the transport audio signal type determiner 201. The transport audio signal type determiner 201 is configured to determine a transport audio signal type 202 and, optionally, some additional parameters 204 (such as microphone distance) based on the bitstream. The determined parameters are transferred from MASA to the multi-channel audio signal converter 903. The transport audio signal type determiner 201 in some embodiments is the same transport audio signal type determiner 201 as described above with respect to FIG. 2, or a separate instance of a transport audio signal type determiner 201 configured to operate in the same manner as the transport audio signal type determiner 201 as described above with respect to the example shown in FIG. 2.

[0164] The MASA-to-multi-channel audio signal converter 903 is configured to receive the bitstream and the transport audio signal type 202 (and optionally some additional parameters 204), and based on the determined transport audio signal type 202 (and possible additional parameters 204), is configured to convert the MASA stream into a multi-channel audio signal (such as 5.1).

[0165] The operation of the example shown in FIG. 9 is summarized in the flowchart shown in FIG. 10.

[0166] The first operation is one of receiving or obtaining a bitstream (MASA stream) as shown in FIG. 10 by step 301.

[0167] The next operation is one of determining a transport audio signal type (and generating a type signal or indicator and possible other additional parameters) based on the bitstream as shown in FIG. 10 by step 303.

[0168] Once the transport audio signal type is determined, the next operation is to convert the bitstream (MASA stream) into a multi-channel audio signal (such as 5.1) based on the determined transport audio signal type, as shown in FIG. 10 by step 1005.

[0169] FIG. 11 shows an exemplary MASA - multi-channel audio signal converter 903 in more detail. The MASA-to-multi-channel audio signal converter 903 is configured to receive a MASA stream (bitstream) and a transport audio signal type 202 and possible additional parameters 204, and is configured to convert the MASA stream into a multi-channel audio signal based on the determined transport audio signal type.

[0170] The MASA-to-multi-channel audio signal converter 903 includes a transport audio signal and spatial metadata extractor / decoder 501. This is configured to receive the MASA stream and output a transport audio signal 502 and spatial metadata 522 in the same way as seen within the transport audio signal type determiner, as shown in FIG. 5 and as discussed. In some embodiments, the extractor / decoder 501 is the extractor / decoder from the previously described transport audio signal type determiner, or a separate instance of the extractor / decoder. The resulting transport audio signal 502 can be transferred to a time / frequency converter 503. The resulting spatial metadata 522 can further be transferred to a target signal characteristic determiner 1101.

[0171] In some embodiments, the MASA - multi-channel audio signal converter 903 includes a time / frequency converter 503. The time / frequency converter 503 is configured to receive the transport audio signals 502 and convert them into the time-frequency domain. Suitable conversions include, for example, the short-time Fourier transform (STFT) and the complex modulated quadrature mirror filter bank (QMF). As a result, the resulting signal is S iLet it be (b, n). Here, i represents the channel index, b represents the frequency bin index, and n represents the time index. Here, are the channel index, the frequency bin index, and the time index. If the output of audio extraction and / or decoding is already in the time-frequency domain, this block may be omitted, or it may include a conversion from one time-frequency domain representation to another time-frequency domain representation. The T / F domain carrier audio signal 504 can be transferred to the prototype signal creator 1111. In some embodiments, the time / frequency converter 503 is the same time / frequency converter from a carrier voice signal type determiner or a MASA-ambisonic converter or a separate instance. In one embodiment, the MASA-to-multi-channel audio signal converter 903 includes the prototype signal creator 1111.

[0172] The prototype signal creator 1111 is configured to receive the T / F domain carrier audio signal 504, the carrier audio signal type 202, and possible additional parameters 204. Then, the T / F prototype signal 1112 can be output to the signal mixer 1105 and the decorrelator 1103.

[0173] As an example regarding the operation of the prototype signal creator 1111a, rendering to a 5.1 multi-channel audio signal configuration will be described. In this example, the prototype signal for the left side (left front and left surround) output channels is

Number

Number

[0174] Therefore, for the output channels on both sides of the central plane, the prototype signal can directly utilize the corresponding carrier audio signal. In the case of the center output channel, the prototype audio signal needs to include energy from both the left and right. This is because it can be used for panning to either side. Therefore, the prototype signal can be created in the same way as the omnidirectional channels in the case of ambisonic rendering. That is, when T(n) = "spaced", [Number] In one embodiment, the prototype audio signal can generate a prototype center audio channel. When T(n) = "downmix" or T(n) = "coincident", [Number]

[0175] In one embodiment, the MASA-to-multi-channel audio signal converter 903 includes a decorrelator 1103. The decorrelator 1103 is configured to receive the T / F prototype signal 1112, apply decorrelation, and output the decorrelated T / F prototype signal 1104 to the signal mixer 1105. In some embodiments, the decorrelator 1103 is optional.

[0176] In one embodiment, the MASA-to-multi-channel audio signal converter 903 includes a target signal characteristic determiner 1101. The target signal characteristic determiner 1101 in some embodiments is configured to generate a target covariance matrix (target signal characteristics) within the frequency band based on the spatial metadata and the overall estimation of the signal energy within the frequency band. In some embodiments, this energy estimation value can be the total of the carrier signal energy in the frequency band. This determination of the target covariance matrix (target signal characteristics) can be performed in a similar way as provided by patent application GB 1718341.9.

[0177] Next, the target signal characteristic 1102 can be passed to the signal mixer 1105.

[0178] In one embodiment, the MASA-to-multi-channel audio signal converter 903 includes a signal mixer 1105. The signal mixer 1105 is configured to measure the covariance matrix of the prototype signal and formulate a mixing solution based on the estimated (prototype signal) covariance matrix and the target covariance matrix. In some embodiments, the mixing solution can be similar to that described in GB1718341.9. The mixing solution is applied to the prototype signal and the uncorrelated prototype signal, and the resulting signal is obtained in the frequency band characteristics based on the target signal characteristics. That is, it is based on the determined target covariance matrix. In some embodiments, the MASA-multi-channel audio signal converter 903 includes an inverse T / F transformer 707 configured to convert the signal into the time domain. The time-domain multi-channel audio signal is the output from the MASA to the multi-channel audio signal converter.

[0179] Regarding FIG. 12, an overview of the operation of the apparatus shown in FIG. 11 is shown.

[0180] Thus, in one embodiment, the first operation is the operation of extracting and / or decoding the carrier audio signal and metadata from the MASA stream (or bit stream) as shown in FIG. 12 by step 801.

[0181] The next operation can be to convert the carrier audio signal into the time-frequency domain as shown in FIG. 12 by step 803.

[0182] Next, the method creates a prototype audio signal based on the carrier signal in the time-frequency domain, and further includes, as shown in FIG. 12 by step 1205, creating a prototype audio signal based on the type of the carrier audio signal (and further based on additional parameters).

[0183] In some embodiments, the method includes, as shown in FIG. 12 by step 1207, applying decorrelation on the time-frequency prototype audio signal.

[0184] Next, by step 1208, as shown in FIG. 12, the target signal characteristics can be determined based on the time-frequency domain carrier audio signal and the spatial metadata (for generating the covariance matrix of the target signal).

[0185] The covariance matrix of the prototype audio signal can be measured as shown in FIG. 12 up to step 1209.

[0186] Next, by step 1209, as shown in FIG. 12, the uncorrelated time-frequency prototype audio signal and the time-frequency prototype audio signal can be mixed based on the target signal characteristics.

[0187] Next, the mixed signal may be inverse time-frequency transformed as shown in FIG. 12 by step 1211.

[0188] Next, the time domain signal can be output as shown in FIG. 12 by step 1213.

[0189] Figure 13 shows a schematic diagram of a decoder of a further example suitable for implementing some embodiments. In other embodiments, a similar method can be implemented in a device other than a decoder, for example, as part of an encoder. This example can be implemented, for example, within a (IVAS) demultiplexer / decoder / synthesizer block 133 as shown in FIG. 1, where in this example, the input is a metadata-assisted spatial audio (MASA) stream including two audio channels and spatial metadata. However, as discussed herein, the input format can be any suitable metadata-assisted spatial audio format.

[0190] (MASA) The bitstream is transferred to the carrier audio signal type determiner 201. The carrier audio signal type determiner 201 is configured to determine a carrier audio signal type 202 and, optionally, some additional parameters 204 (an example of such an additional parameter is the microphone distance) based on the bitstream. The determined parameters are transferred to the downmixer 1303. The carrier audio signal type determiner 201 in some embodiments can be the same carrier audio signal type determiner 201 as described above or a separate instance of a carrier audio signal type determiner 201 configured to operate similarly to the carrier audio signal type determiner 201 as described above.

[0191] The downmixer 1303 is configured to receive the bitstream and the carrier audio signal type 202 (and optionally some additional parameters 204) and is configured to downmix the MASA stream from two carrier audio signals to one carrier audio signal based on the determined carrier audio signal type 202 (and possible additional parameters 204). Next, an output MASA stream 1306 is output.

[0192] The operation of the example shown in FIG. 13 is summarized in the flowchart shown in FIG. 14.

[0193] The initial operation is to receive or acquire a bitstream (MASA stream) as shown in FIG. 14 by step 301.

[0194] The next operation is to determine the transport audio signal type based on the bitstream (and generate a type signal or indicator and possibly other additional parameters) as shown in FIG. 14 by step 303.

[0195] After determining the type of the transport audio signal, the next operation is to downmix the MASA stream from two transport audio signals to one transport audio signal based on the determined type 202 of the transport audio signal (and possibly additional parameter 204) as shown in FIG. 14 by step 1405.

[0196] FIG. 15 shows an example of the downmixer 1303 in more detail. The downmixer 1303 is configured to receive a MASA stream (bitstream) and the transport audio signal type 202 and possibly additional parameter 204, and to downmix two transport audio signals to one transport audio signal based on the determined transport audio signal type.

[0197] The downmixer 1303 includes a transport audio signal and a spatial metadata extractor / decoder 501. This is configured to receive the MASA stream and output the transport audio signal 502 and the spatial metadata 522 in the same way as seen in the transport audio signal type determiner discussed there. In one embodiment, the extractor / decoder 501 is the extractor / decoder described above, or a separate instance of the extractor / decoder. The obtained transport audio signal 502 can be transferred to the time / frequency converter 503. The obtained spatial metadata 522 can further be transferred to the signal multiplexer 1507.

[0198] In some embodiments, the downmixer 1303 includes a time / frequency converter 503. The time / frequency converter 503 is configured to receive the carrier audio signals 502 and convert them into the time-frequency domain. Suitable conversions include, for example, the short-time Fourier transform (STFT) and the complex modulated quadrature mirror filter bank (QMF). The resulting signals are represented as S i (b,n). Here, b, n, and m are the channel index, frequency bin index, and time index, respectively. If the output of the audio extraction and / or decoding is already in the time-frequency domain, this block may be omitted, or it may include a conversion from one time-frequency domain representation to another. The T / F domain carrier audio signal 504 can be transferred to the prototype signal creator 1511. In some embodiments, the time / frequency converter 503 is the same time / frequency converter as described above or a separate instance.

[0199] In some embodiments, the downmixer 1303 includes a prototype signal creator 1511. The prototype signal creator 1511 is configured to receive the T / F domain carrier audio signal 504, the carrier audio signal type 202, and the possible additional parameters 204. Then, the T / F prototype signal 1512 can be output to the proto-energy determiner 1503 to match the prototype signal to the target energy collator 1505.

[0200] The prototype signal creator 1511 in some embodiments is configured to create a prototype signal of the mono carrier audio signal using two carrier audio signals based on the received carrier audio signal type. For example, the following can be used. When T(n) = "spaced",

Equation

Number

[0201] In some embodiments, the downmixer 1303 includes a target energy determiner 1501. The target energy determiner 1501 receives the T / F domain carrier audio signal 504 and generates a target energy value as the total energy of the carrier audio signal

Number

[0202] The target energy value can then be passed to the proto to match the target equalizer 1505.

[0203] In some embodiments, the downmixer 1303 includes a proto energy determiner 1503. The proto energy determiner 1503 receives the T / F prototype signal 1512 and is configured to determine an energy value, for example,

Number

[0204] Next, the proto energy value can be passed to the proto to match the target equalizer 1505.

[0205] The downmixer 1303 in some embodiments includes a proto that matches the target energy equalizer 1505. The proto for matching the target energy equalizer 1505 in some embodiments is configured to receive the T / F prototype signal 1502, the proto energy value, and the target energy value. The equalizer 1505 in some embodiments first, for example,

Number

[0206] Next, the prototype signal can be equalized using these gains as follows. [Number] The equalized prototype signal is passed to the inverse T / F transformer 707.

[0207] In some embodiments, the downmixer 1303 includes an inverse number T / F transformer 707 configured to convert the output of the equalizer into a time-domain version. Next, the time-domain equalized audio signal (mono signal) 1510 is passed to the carrier audio signal and spatial metadata multiplexer 1507 (or multiplexer).

[0208] In some embodiments, the downmixer 1303 includes a carrier audio signal and spatial metadata multiplexer 1507 (or multiplexer). The carrier audio signal and spatial metadata multiplexer 1507 (or multiplexer) is configured to receive the spatial metadata 522 and the mono audio signal 1510, multiplex them, and regenerate an appropriate output format (for example, a MASA stream having only one carrier audio signal) 1506. In some embodiments, the input mono audio signal is in pulse code modulation (PCM) format. In such embodiments, the signal may be not only multiplexed but also encoded. In some embodiments, multiplexing may be omitted, and the mono carrier audio signal and spatial metadata are directly used by the audio encoder.

[0209] In one embodiment, the output of the apparatus shown in FIG. 15 is a mono PCM audio signal 1510 in which the spatial metadata is discarded.

[0210] In some embodiments, other parameters can be implemented. For example, in some embodiments, when the type is "spaced", the spaced microphone distance can be estimated.

[0211] With respect to FIG. 16, the operation of an example of the apparatus shown in FIG. 15 is shown.

[0212] Thus, in one embodiment, the first operation is the operation of extracting and / or decoding the transport audio signal and metadata from the MASA stream (or bitstream) as shown in FIG. 16 by step 1601.

[0213] The next operation can be the time-frequency domain conversion of the transport audio signal as shown in FIG. 16 by step 1603.

[0214] Then, the method includes creating a prototype audio signal based on the transport signal in the time-frequency domain, and further, creating a prototype audio signal based on the type of the transport audio signal (and further based on additional parameters) as shown in FIG. 16 by step 1605.

[0215] Furthermore, in some embodiments, the method is configured to generate, determine, or calculate a target energy value based on the converted transport audio signal as shown in FIG. 16 by step 1604.

[0216] Furthermore, in some embodiments, the method is configured to generate, determine, or calculate a prototype audio signal energy value based on the prototype audio signal energy value as shown in FIG. 16 by step 1606.

[0217] After determining the energy, the method can further equalize the prototype audio signal to match the target audio signal energy, as shown in FIG. 16, by step 1607.

[0218] The equalized prototype signal (mono signal) may then be inverse time-frequency domain transformed to generate a time domain mono signal, as shown in FIG. 16, by step 1609.

[0219] Next, the time domain monaural audio signal may be (optionally encoded and multiplexed with) the spatial metadata, as shown in FIG. 16, by step 1610.

[0220] Next, the multiplexed audio signal can be output (as a MASA data stream), as shown in FIG. 16, by step 1611.

[0221] As described above, the block diagrams shown are merely examples of possible implementations. Other practical implementations may differ from the above examples. For example, an implementation may not have an individual T / F converter.

[0222] Furthermore, in some embodiments, instead of having an input MASA stream as shown above, any suitable bitstream that utilizes audio channels and (spatial) metadata can be used. Additionally, in some embodiments, the IVAS codec can be replaced with any other suitable codec (e.g., one having an operating mode for audio channels and spatial metadata).

[0223] In some embodiments, a carrier audio signal type determiner can be used to estimate parameters other than the carrier audio signal type. For example, the microphone spacing can be estimated. The microphone spacing is an example of a possible additional parameter 204. This is, in some embodiments, E sum(b, n) and E sub It can be achieved by examining the maximum and minimum frequencies of (b, n), determining the time delay between microphones based on them, and estimating the distance based on the delay and the estimated arrival direction (available in the spatial metadata). There is also a method for estimating the delay between two signals.

[0224] Regarding FIG. 17, an example of an electronic device that can be used as an analysis device or a synthesis device is shown. This device can be any suitable electronic device or apparatus. For example, in one embodiment, device 1700 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.

[0225] In one embodiment, device 1700 includes at least one processor or central processing unit 1707. Processor 1707 can be configured to execute various program codes such as the methods described herein.

[0226] In one embodiment, device 1700 includes a memory 1711. In one embodiment, at least one processor 1707 is coupled to the memory 1711. Memory 1711 can be any suitable storage means. In one embodiment, memory 1711 includes a program code section for storing program codes that can be implemented on processor 1707. Further, in some embodiments, memory 1711 can further include a storage data section for storing data that has been processed or is to be processed, for example, according to the embodiments described herein. The implemented program codes stored in the program code section and the data stored in the stored data section can be retrieved by processor 1707 at any time when needed via the memory-processor coupling.

[0227] In some embodiments, apparatus 1700 includes user interface 1705. User interface 1705 can be coupled to processor 1707 in some embodiments. In some embodiments, processor 1707 can control the operation of user interface 1705 and receive inputs from user interface 1705. In some embodiments, user interface 1705 can enable a user to input commands to apparatus 1700, for example, via a keypad. In some embodiments, user interface 1705 can enable a user to obtain information from apparatus 1700. For example, user interface 1705 can include a display configured to display information to the user from apparatus 1700. In some embodiments, user interface 1705 can comprise a touch screen or touch interface capable of both enabling input of information to apparatus 1700 and further displaying information to a user of apparatus 1700. In some embodiments, user interface 1705 can be a user interface for communicating with a positioner, as described herein.

[0228] In some embodiments, apparatus 1700 includes input / output port 1709. Input / output port 1709 in some embodiments includes a transceiver. The transceiver in such embodiments can be coupled to processor 1707 and can be configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means can be configured in some embodiments to communicate with other electronic devices or apparatuses via a wire or wired connection.

[0229] The transceiver can communicate with additional devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth®), or an Infrared Data Association (IrDA) communication path.

[0230] The transceiver input / output port 1709 may be configured to receive signals and, in some embodiments, to determine parameters as described herein by using a processor 1707 that executes suitable code.

[0231] In some embodiments, the device 1700 may be employed as at least part of a synthesizer. The input / output port 1709 can be coupled to any suitable audio output, such as a multi-channel speaker system and / or headphones, which can be head-tracked headphones or non-tracked headphones, or the like.

[0232] In general, the various embodiments of the present invention can be realized in hardware or special-purpose circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, but the present invention is not limited thereto, and may be implemented in firmware or software that may be executed by a controller, a microprocessor, or other computing device. The various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or some other pictorial representation, but it is well understood that these blocks, devices, systems, techniques, or methods described herein can be implemented, by way of non-limiting example, in hardware, software, firmware, special-purpose circuitry or logic, general-purpose hardware or controllers, or other computing devices, or combinations thereof.

[0233] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device, such as within a processor entity, or by hardware, or by computer software executable by a combination of software and hardware. Further, note that any block of a logical flow as shown in the figures can represent a program step, or interconnected logical circuits, blocks, and functions, or a combination of program steps and logical circuits, blocks, and functions. This software can be stored on a physical medium such as a memory chip, or a memory block implemented within a processor, a magnetic medium such as a hard disk or a floppy (registered trademark) disk, and an optical medium such as a DVD and its data variants, for example.

[0234] The memory can be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and can include, by way of non-limiting example, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor, an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0235] Embodiments of the present invention can be implemented in various components, such as integrated circuit modules. The design of integrated circuits is by a highly automated process and is large-scale. Complex and powerful software tools are available for converting a design at the logic level into a complete semiconductor circuit design that is etched and ready to be formed on a semiconductor substrate.

[0236] Programs such as those provided by Synopsys, Inc. in Mountain View, California and Cadence Design in San Jose, California use well-established rules of design and a library of pre-stored design modules to automatically route conductors and identify the positions of components on a semiconductor chip. When the design of a semiconductor circuit is complete, the resulting design can be transmitted in a standardized electronic format (e.g., Opus, GDSII, etc.) to semiconductor manufacturing equipment or a "fab" for manufacturing.

[0237] The foregoing description has been provided by way of illustrative example and non-limiting example to a complete and reference description of an exemplary embodiment of the invention. However, various modifications and adaptations will become apparent to those skilled in the art upon a review of the foregoing description when considered in conjunction with the accompanying drawings and the appended claims. However, all such changes and similar changes to the teachings of this invention will continue to fall within the scope of the invention as defined by the appended claims.

Claims

1. obtaining at least two audio signals, obtaining at least one parameter related to the at least two audio signals, determining (201) the type of the at least two audio signals based on the at least one parameter related to the at least two audio signals, processing (203) the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals means configured as a device comprising.

2. The apparatus according to claim 1, wherein the at least two audio signals are a transmission audio signal (502) and one of the previously processed audio signals.

3. The means configured to determine (201) the type of the at least two audio signals based on the at least one parameter comprises: extracting and decoding at least one type signal from the at least one parameter (501); and analyzing the at least one parameter to determine the type of the at least two audio signals when the at least one parameter represents a spatial audio mode related to the at least two audio signals. configured to perform one of The apparatus according to claim 1 or 2.

4. The means configured to analyze the at least one parameter to determine the type of the at least two audio signals comprises: determining (505) a broadband left or right channel to total energy ratio based on the at least two audio signals, determining (507) a high frequency left or right channel to total energy ratio based on the at least two audio signals, determining (509) a sum to total energy ratio based on the at least two audio signals, determining (512) a subtraction to target energy ratio based on the at least two audio signals, and Determine the type of the at least two audio signals (513) based on at least one of the broadband left or right channel to total energy ratio, the high-frequency left or right channel to total energy ratio based on the at least two audio signals, the total to total energy ratio based on the at least two audio signals, and the subtraction to target energy ratio. configured as The apparatus according to claim 3.

5. The apparatus according to any one of claims 1 to 4, wherein the means is configured to determine at least one type parameter related to at least one type of the at least two audio signals.

6. The means configured to process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals is configured to convert the at least two audio signals based on the at least one type parameter related to the types of the at least two audio signals. The apparatus according to claim 5.

7. The types of the at least two audio signals are audio signal types at intervals, downmixed audio signal types, the same audio signal types, including at least one of The apparatus according to any one of claims 1 to 6.

8. The means configured to process the at least two audio signals is converting at least two audio signals into an ambisonic audio signal representation, converting the at least two audio signals into an ambisonic audio signal representation (903), downmixing the at least two audio signals into fewer audio signals (1303), configured to perform one of The apparatus according to any one of claims 1 to 7.

9. The means configured to process the at least two audio signals is configured to generate at least one prototype signal based on the at least two audio signals and the types of the at least two audio signals. The apparatus according to any one of claims 1 to 8.

10. Step (301) of obtaining at least two audio signals Step of obtaining at least one parameter related to the at least two audio signals Step (303) of determining the types of the at least two audio signals based on the at least one parameter related to the at least two audio signals Step (305) of processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals A method comprising the above steps [

11. ] The method according to claim 10, wherein the at least two audio signals are one of a transmission audio signal (502) and a pre-processed audio signal [

12. ] The step of determining the types of the at least two audio signals based on the at least one parameter includes Step (601) of extracting and decoding at least one type signal from the at least one parameter When the at least one parameter represents a spatial audio mode related to the at least two audio signals, step of analyzing the at least one parameter to determine the types of the at least two audio signals One of the above steps is included The method according to claim 10 or 11 [

13. ] The step of analyzing the at least one parameter to determine the types of the at least two audio signals includes Step (605) of determining a broadband left or right channel to total energy ratio based on the at least two audio signals Step (607) of determining a high-frequency left or right channel to total energy ratio based on the at least two audio signals Step (609) of determining a sum to total energy ratio based on the at least two audio signals Step (611) of determining a subtraction to target energy ratio based on the at least two audio signals The broadband left or right channel to total energy ratio The high-frequency left or right channel to total energy ratio based on at least two audio signals The total-to-total energy ratio based on at least two audio signals, and, the subtraction-to-target energy ratio determining, based on at least one of the foregoing, the type of the at least two audio signals (613); comprising The method according to claim 12. **Claim 14** determining at least one type parameter related to at least one type among the at least two audio signals The method according to any one of claims 10 to 13, further comprising. **Claim 15** The step of processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals comprises: converting the at least two audio signals based on at least one type parameter related to the types of the at least two audio signals (305) comprising The method according to claim 10.

Citation Information

Patent Citations

  • Coding of higher-order ambisonic coefficients between multiple transitions

    JP2018534617A

  • Binaural audio signal processing method and apparatus

    JP2019533404A

  • Method and device for audio data processing

    US20170162210A1

  • Binaural audio signal processing method and apparatus

    WO2018056780A1