Sound field related rendering
Patent Information
- Application Number
- JP2023200065
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-27
- Filing Date
- 2023-11-27
- Publication Date
- 2025-09-03
AI Technical Summary
Existing audio rendering technologies face challenges in accurately processing and rendering spatial audio signals due to the use of inappropriate processing techniques for different types of carrier audio signals, leading to degradation in audio quality.
The system determines the type of carrier audio signal by analyzing spatial metadata and the audio signal itself, using techniques such as energy ratio comparisons and spectral analysis, to apply appropriate linear and parametric processing for efficient conversion to ambisonic or multi-channel formats, thereby improving audio quality.
This approach enhances audio quality by ensuring that the processing methods align with the specific characteristics of the carrier audio signal type, minimizing artifacts and improving the accuracy of spatial audio rendering.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus and method for sound field related audio representation and rendering, but is not limited to audio representation for audio decoders. [Background technology]
[0002] Immersive audio codecs support a number of operating points, from low bitrate operation to transparency. One example of such a codec is the Immersive Speech and Audio Services (IVAS) codec, which is designed to be suitable for use over communication networks such as 3GPP 4G / 5G networks, including for use in immersive services such as immersive speech and audio for virtual reality (VR). This speech codec is expected to handle the encoding, decoding, and rendering of speech, music, and general-purpose speech. In addition, it is expected to support channel-based audio and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to support high error robustness under various transmission conditions, as well as operate with low latency to enable conversational services.
[0003] The input signal can be presented to the IVAS encoder in any of a number of supported formats (and possible combinations of formats). For example, a mono audio signal (without metadata) can be encoded using the Enhanced Voice Service (EVS) encoder. Other input formats may utilize the IVAS encoding tools. At least some inputs can utilize the Metadata Assisted Spatial Audio (MASA) tools or any suitable spatial meta-database scheme. This is a parametric spatial audio format suitable for spatial audio processing. Parametric spatial audio processing is a field of audio signal processing in which the spatial aspects of the sound (or sound field) are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, it is a typical and valid choice to estimate from the microphone array signal a set of parameters such as the direction of the sound in a frequency band and the ratio between the directional and omnidirectional parts of the captured sound in a frequency band. These parameters are known to well describe the perceived spatial characteristics of the sound captured at the location of the microphone array. These parameters can be used accordingly for spatial sound synthesis, binaural for headphones, loudspeakers, or other formats such as ambisonics.
[0004] For example, there are two channels (stereo) of audio signals and spatial metadata. The spatial metadata can further define parameters such as directional index (which describes the direction of arrival of the sound in the time-frequency parameter interval), directional-to-total energy ratio (which describes the directional index, i.e., the energy ratio for the time-frequency subframe), spread coherence (which describes the energy ratio of an omnidirectional sound to the surrounding directions), diffuse-to-total energy ratio (which describes the coherence of an omnidirectional sound to the surrounding directions), surround coherence (which describes the coherence of an omnidirectional sound to the surrounding directions), remainder-to-total energy ratio (which describes the energy ratio of the residual (e.g., microphone noise) acoustic energy to meet the requirement that the sum of the energy ratios is 1), and distance (which describes the distance of the sound emanating from the directional index, i.e., the time-frequency subframe, in a logarithmic scale).
[0005] IVAS streams can be decoded and rendered to a variety of output formats, including binary, multi-channel and Ambisonic (FOA / HOA) output.
[0006] The means configured to process the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals may be configured to convert the at least two audio signals into an Ambisonic audio signal representation, convert the at least two audio signals into a multi-channel audio signal representation, and downmix the at least two audio signals into fewer audio signals.
[0007] The means configured to process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals may be configured to generate at least one prototype signal based on the at least two audio signals and the types of the at least two audio signals.
[0008] According to a second aspect, there is provided a method comprising the steps of obtaining at least two audio signals, determining types of the at least two audio signals, and processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
[0009] The at least two audio signals may be one of a carrier audio signal and a previously processed audio signal.
[0010] The method may further include obtaining at least one parameter associated with the at least two audio signals.
[0011] Determining a type of the at least two audio signals may include determining a type of the at least two audio signals based on at least one parameter associated with the at least two audio signals.
[0012] Determining a type of the at least two audio signals based on the at least one parameter may include one of extracting and decoding the at least one type of signal from the at least one parameter and, if the at least one parameter represents a spatial audio aspect associated with the at least two audio signals, analyzing the at least one parameter to determine the type of the at least two audio signals.
[0013] Analyzing at least one parameter to determine a type of the at least two audio signals may include determining a broadband left or right channel to total energy ratio based on the at least two audio signals, determining a higher frequency left or right channel to total energy ratio based on the at least two audio signals, determining a sum to sum to total energy ratio based on the at least two audio signals, and determining a subtraction to target energy ratio based on the at least two audio signals, and determining a type of the at least two audio signals based on at least one of the broadband left or right channel to total energy ratio, the higher frequency left or right channel to total energy ratio based on the at least two audio signals, the sum to total energy ratio based on the at least two audio signals, and the subtraction to target energy ratio.
[0014] The method may further include determining at least one type parameter associated with a type of the at least one audio signal.
[0015] Processing the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals may further include converting the at least two audio signals based on at least one type parameter associated with the type of the at least two audio signals.
[0016] The at least two audio signal types may include at least one of a capture microphone placement, a capture microphone separation distance, a capture microphone parameter, a transport channel identifier, a spaced audio signal type, a downmix audio signal type, a same audio signal type, and a transport channel placement.
[0017] Processing the at least two audio signals configured to be rendered based on the determined type of the at least two audio signals may include one of converting the at least two audio signals to an Ambisonic audio signal representation, converting the at least two audio signals to a multi-channel audio signal representation, and downmixing the at least two audio signals to fewer audio signals.
[0018] Processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals may include generating at least one prototype signal based on the at least two audio signals and the types of the at least two audio signals.
[0019] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the at least one computer program code configured to cause, using the at least one processor, the apparatus to at least acquire at least two audio signals, determine types of the at least two audio signals, and process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
[0020] The at least two audio signals may be one of a carrier audio signal and a previously processed audio signal.
[0021] The means may be configured to obtain at least one parameter relating to the at least two audio signals.
[0022] The apparatus adapted to determine a type of the at least two audio signals may be adapted to determine the type of the at least two audio signals based on at least one parameter related to the at least two audio signals.
[0023] The apparatus for determining the type of the at least two audio signals based on the at least one parameter may perform one of: extracting and decoding at least one type signal from the at least one parameter; and, when the at least one parameter represents a spatial audio aspect associated with the at least two audio signals, analysing the at least one parameter to determine the type of the at least two audio signals.
[0024] The apparatus for analyzing at least one parameter to determine a type of the at least two acoustic signals may determine a broadband left or right channel to total energy ratio based on the at least two acoustic signals, determine a higher frequency or right channel to total energy ratio based on the at least two acoustic signals, determine a sum-to-total energy ratio based on the at least two acoustic signals, and determine a subtraction-to-total energy ratio based on the at least two acoustic signals, and may determine the type of the at least two acoustic signals based on at least one of the broadband left or right channel to total energy ratio, the high frequency left or right channel to total energy ratio based on the at least two acoustic signals, the sum-to-total energy ratio based on the at least two acoustic signals, and the subtraction-to-target energy ratio.
[0025] The apparatus is capable of determining at least one type parameter associated with a type of the at least one audio signal.
[0026] The device having processed the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals may cause the at least two audio signals to be transformed based on at least one type parameter associated with the types of the at least two audio signals.
[0027] The at least two audio signal types may include at least one of a capture microphone placement, a capture microphone separation distance, a capture microphone parameter, a transport channel identifier, a spaced audio signal type, a downmix audio signal type, a same audio signal type, and a transport channel placement.
[0028] The apparatus is capable of processing at least two audio signals configured to be rendered based on a determined type of the at least two audio signals, converting the at least two audio signals to an Ambisonic audio signal representation, converting the at least two audio signals to a multi-channel audio signal representation, and downmixing the at least two audio signals to fewer audio signals.
[0029] The present apparatus is capable of processing at least two audio signals configured to be rendered based on determined types of the at least two audio signals and generating at least one prototype signal based on the at least two audio signals and the types of the at least two audio signals.
[0030] According to a fourth aspect, there is provided an apparatus comprising: obtaining circuitry configured to obtain at least two audio signals; a determination circuit configured to determine a type of the at least two audio signals; and a processing circuit configured to process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
[0031] According to a fifth aspect, there is provided a computer program (or a computer readable medium comprising program instructions) comprising instructions for causing an apparatus to at least: obtain at least two audio signals; determine a type of said at least two audio signals; and process said at least two audio signals configured to be rendered based on the determined types of said at least two audio signals.
[0032] According to a sixth aspect, there is provided a non-transitory computer-readable medium comprising program instructions for causing an apparatus to at least: obtain at least two audio signals; determine a type of the at least two audio signals; and process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
[0033] According to a seventh aspect, there is provided an apparatus comprising: means for obtaining at least two audio signals; means for determining a type of the at least two audio signals; and means for processing the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
[0034] According to an eighth aspect, there is provided a computer-readable medium comprising program instructions for causing an apparatus to obtain at least two audio signals, determine a type of the at least two audio signals, and process the at least two audio signals configured to be rendered based on the determined types of the at least two audio signals.
[0035] An apparatus comprising means for performing the operations of the methods described above.
[0036] An apparatus configured to perform the actions of the above method.
[0037] A computer program comprising program instructions for causing a computer to carry out the method described above.
[0038] A computer program product stored on the medium can cause an apparatus to perform the methods described herein.
[0039] The electronic device may include the apparatus described herein.
[0040] The chipset may include the devices described herein.
[0041] SUMMARY OF THE PRESENT EMBODIMENTS Embodiments of the present invention aim to address problems associated with the state of the art. [Brief description of the drawings]
[0042] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 illustrates a schematic of a system of apparatus suitable for implementing some embodiments. [Diagram 2] FIG. 2 illustrates a schematic of an example of a decoder / renderer according to some embodiments. [Diagram 3] FIG. 3 illustrates a flow diagram of an example decoder / renderer operation according to some embodiments. [Figure 4] FIG. 4 illustrates a schematic diagram of an example carrier audio signal type determiner as shown in FIG. 2, according to some embodiments. [Diagram 5] FIG. 5 illustrates a schematic diagram of a second example carrier audio signal type determiner as shown in FIG. 2, according to some embodiments. [Figure 6] FIG. 6 illustrates a flow diagram of a second example carrier audio signal type determiner operation according to some embodiments. [Figure 7]FIG. 7 illustrates a schematic example of a metadata-assisted spatial audio signal to Ambisonics format converter, as shown in FIG. 2, according to some embodiments. [Figure 8] FIG. 8 shows a flow diagram of the operation of a sample metadata-assisted spatial audio signal to Ambisonics format converter, according to some embodiments. [Figure 9] FIG. 9 illustrates a schematic of a second example decoder / renderer according to some embodiments. [Figure 10] FIG. 10 shows a flow diagram of a further example decoder / renderer operation according to some embodiments. [Figure 11] FIG. 11 illustrates a schematic example of a metadata-assisted spatial audio signal to multi-channel audio signal format converter, as shown in FIG. 9, according to some embodiments. [Figure 12] FIG. 12 shows a flow diagram of the operation of a sample metadata-assisted spatial audio signal to multi-channel audio signal format converter according to some embodiments. [Figure 13] FIG. 13 illustrates a schematic of a third example decoder / renderer according to some embodiments. [Figure 14] FIG. 14 shows a flow diagram of a third example decoder / renderer operation according to some embodiments. [Figure 15] FIG. 15 illustrates an exemplary metadata-assisted spatial audio signal downmixer, such as that shown in FIG. 13, according to some embodiments. [Figure 16] FIG. 16 illustrates a flow diagram of the operation of an example metadata-assisted spatial audio signal downmixer in accordance with some embodiments. [Figure 17] FIG. 17 shows an example apparatus suitable for realising the apparatus shown in FIGS. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0043] In the following, suitable apparatus and possible mechanisms for providing efficient rendering of spatial metadata-aided audio signals are described in further detail.
[0044] With reference to Fig. 1, examples of devices and systems for implementing audio capture and rendering are shown. System 100 is shown with an "analysis" section 121 and a "demultiplexer / decoder / synthesizer" section 133. The "analysis" section 121 is responsible for receiving multi-channel loudspeaker signals and encoding metadata and carrier signals, and the "demultiplexer / decoder / synthesizer" section 133 is responsible for decoding the encoded metadata and carrier signals and presenting a regenerated signal (e.g., multi-channel loudspeaker formation).
[0045] The input to the system 100 and the "analysis" part 121 is a multi-channel signal 102. In the examples below, microphone channel signal inputs are described, but in other embodiments, any suitable input (or composite multi-channel) format can be realized. For example, in some embodiments, the spatial analyzer and spatial analysis may be performed outside the encoder. For example, in one embodiment, spatial metadata associated with the audio signal may be provided to the encoder as a separate bitstream. In one embodiment, the spatial metadata may be provided as a set of spatial (directional) index values.
[0046] The multi-channel signal is passed to a carrier signal generator 103 and an analysis processor 105 .
[0047] In some embodiments, the carrier signal generator 103 is configured to receive the multi-channel signal, generate a suitable carrier signal including the determined number of channels, and output the carrier signal 104. For example, the transport signal generator 103 may be configured to generate a two audio channel downmix of the multi-channel signal. The determined number of channels may be any suitable number of channels. The carrier signal generator in some embodiments is configured to select or combine the input audio signal into the determined number of channels, for example by beam forming techniques, and output these as the carrier signal.
[0048] In some embodiments, the carrier signal generator 103 is optional and the multi-channel signal is passed unprocessed to the "Encoder / MUX" block 107, just as the carrier signal is in this example.
[0049] In some embodiments, the analysis processor 105 is also configured to receive the multi-channel signal and analyze the signal to generate metadata 106 related to the multi-channel signal and thus to the carrier signal 104. The analysis processor 105 may be configured to generate metadata including a direction parameter 108 and an energy ratio parameter 110 (one example of which is the diffuseness parameter) and a coherence parameter 112 for each time-frequency analysis interval. The direction, energy ratio and coherence parameters may be considered in embodiments as spatial audio parameters. In other words, spatial audio parameters include parameters that aim to characterize the sound field created by the multi-channel signal (or generally two or more reproduced audio signals).
[0050] In some embodiments, the generated parameters may be different for each frequency band. Thus, for example, in band X, all parameters are generated and transmitted, whereas in band Y only one parameter is generated and transmitted, and in band Z no parameters are generated and transmitted. A practical example of this would be that in some frequency bands, such as the highest band, some parameters are not required for perceptual reasons. The transport signal 104 and metadata 106 can be passed to an "encoder / mux" block 107.
[0051] In some embodiments, the spatial audio parameters may be grouped or separated into directional and non-directional (eg, diffuse) parameters.
[0052] The "Encoder / MUX" block 107 may be configured to receive the transport (e.g., downmix) signal 104 and generate a suitable encoding of these audio signals. The "Encoder / MUX" block 107 may in one embodiment be a computer (executing suitable software stored on a memory and on at least one processor), or alternatively a specific device utilizing, for example, an FPGA or an ASIC. The encoding may be performed using any suitable scheme. The "Encoder / MUX" block 107 may further be configured to receive metadata and generate an encoded or compressed form of the information. In one embodiment, the "Encoder / MUX" block 107 may interleave, multiplex, or embed the metadata within the encoded downmix signal by dashed lines into a single data stream 111 before transmission or storage as shown in FIG. 1. The multiplexing may be performed using any suitable scheme.
[0053] On the decoder side, the received or retrieved data (stream) may be received by a "Demultiplexer / Decoder / Synthesizer" 133. The "Demultiplexer / Decoder / Synthesizer" 133 may demultiplex the encoded stream and decode the audio signal to obtain a transport signal. Similarly, the "Demultiplexer / Decoder / Synthesizer" 133 may be configured to receive and decode encoded metadata. In some embodiments, the "Demultiplexer / Decoder / Synthesizer" 133 may be a computer (executing appropriate software stored on a memory and on at least one processor), or alternatively a specific device utilizing, for example, an FPGA or an ASIC.
[0054] The “Demultiplexer / Decoder / Synthesizer” portion 133 of the system 100 may be further configured to recreate synthetic spatial audio in the form of multi-channel signal 110 in any suitable format based on the transport signal and metadata (which may be a multi-channel loudspeaker format or, in an embodiment, any suitable output format such as a binaural or Ambisonics signal for headphone listening, depending on the use case).
[0055] Thus, at the start of the overview, the system (the analysis part) is set up to receive a multi-channel audio signal.
[0056] The system (the analysis part) is then configured to generate an appropriate carrier audio signal (eg, by selecting or downmixing some of the audio signal channels).
[0057] The system is then configured to encode the transport signal and the metadata for storage / transmission.
[0058] After this, the system can store / transmit the encoded transport and metadata.
[0059] The system is capable of retrieving / receiving the encoded conveyance and metadata.
[0060] The system is then configured to extract the conveyance and metadata from the encoded conveyance and metadata parameters, eg, demultiplex and decode the encoded conveyance and metadata parameters.
[0061] The system (synthesizing part) is configured to synthesize an output multi-channel audio signal based on the extracted carrier audio signal and the metadata. As for the decoder (synthesizing part), it is configured to receive the spatial metadata and forward the (potentially pre-processed version) audio signal, which can be, for example, a downmix of a 5.1 signal, two spaced microphone signals from a mobile device, or two beam patterns from a matching microphone array.
[0062] A decoder may be configured to render spatial audio (such as Ambisonic) from the spatial metadata and the carrier audio signal. This is typically achieved by employing one of two approaches to rendering spatial audio from such input: linear and parametric rendering.
[0063] Assuming processing in frequency bands, linear rendering refers to utilizing some static blending weights to produce the desired output, whereas parametric rendering refers to modifying the carrying audio signal based on spatial metadata to produce the desired output.
[0064] Methods for generating Ambisonics from a variety of inputs are presented.
[0065] For carrying audio signals and spatial metadata from the 5.1 signal, parametric processing can be used to render Ambisonics.
[0066] A combination of linear and parametric processing can also be used when carrying audio signals and spatial metadata from remote microphones.
[0067] For carrying audio signals and spatial metadata from simultaneous microphones, a combination of linear and parametric processing can be used.
[0068] Thus, there are various methods for rendering Ambisonics from various kinds of input. However, all constant Ambisonic rendering methods assume a certain type of input. Some embodiments described below present apparatus and methods that prevent problems such as the following from occurring:
[0069] Using linear rendering, the Ambisonic left-facing primary (8th order) signal, the Y signal, can be created from two matching opposite cardioids by Y(f)=S0(f)-S1(f), where f is frequency. As another example, the Y signal can be expressed as Y(f)=-i(S0(f)-S1(f))g eq (f) where g eq (f) is a frequency-dependent equalizer (dependent on microphone distance), i in imaginary units. Processing of spaced microphones (including -90 degree phase shift and frequency-dependent equalization) differs from processing of coincident microphones, and using the wrong processing techniques can result in poor sound quality.
[0070] To use parametric rendering in some rendering schemes, "prototype" signals must be generated using linear averaging. These prototype signals are then adaptively modified in the time-frequency domain based on the spatial metadata. Optimally, the prototype signals should follow the target signal as closely as possible. This minimizes the need for parametric processing and therefore minimizes potential artifacts due to parametric processing. For example, the prototype signals should contain all signal components relevant to the corresponding output channels with sufficient range.
[0071] As an example, once an omnidirectional signal W is rendered (similar effects exist for other Ambisonic signals), a prototype can be created, for example, from a stereo carrier audio signal with two simple approaches: either selecting one channel (e.g. the left channel) or the sum of the two channels.
[0072] Which one to choose depends heavily on the type of carrier audio signal. If the carrier signal originates from a 5.1 signal, the left signal is usually only the left carrier audio signal and the right signal is only the right carrier audio signal (when using a typical downmix matrix). Therefore, using one channel for the prototype would result in the loss of signal content in the other channel, producing clear artifacts (e.g., in the worst case, no signal is present at all in the one selected channel). Therefore, in this case, it was better to formulate the W prototype as the sum of both channels. On the other hand, if the carrier signal originates from a distant microphone, using the sum of the carrier audio signals as the prototype for the W signal would result in severe comb filtering (because there is a time delay between the signals). This would result in similar artifacts as mentioned above. In this case, it is better to select only one of the two channels as the W prototype, at least in the high frequency range.
[0073] Therefore, there is no suitable option that suits all types of carried audio signals.
[0074] Therefore, applying spatial audio processing designed for one carrier audio signal type to another carrier audio signal type, using both linear and parametric methods, is expected to produce a clear degradation in audio quality.
[0075] The concept as discussed in more detail with respect to the following embodiments and examples relates to audio encoding and decoding where the decoder receives at least two carrier audio signals from the encoder. Furthermore, the embodiments indicate that the carrier audio signals may be of at least two types, for example, a downmix of a 5.1 signal, a spaced microphone signal, or a coincident microphone signal. Furthermore, in some embodiments, the apparatus and method implement a solution for improving the quality of the processing of the carrier audio signal and providing a determined output (e.g., Ambisonic, 5.1, mono). By determining the type of the carrier audio signal and performing the processing of the audio based on the determined type of the carrier audio signal, the quality can be improved.
[0076] In some embodiments discussed in further detail herein, the carrying audio signal type is determined either by obtaining metadata indicative of the type of carrying audio signal, or by determining the type of carrying audio signal based on the carrying audio signal (and spatial metadata, if available) itself.
[0077] Metadata describing the carried audio signal type may include, for example, terms such as spaced microphones (which may be accompanied by microphone positions), matching microphones or beams that are substantially similar to matching microphones (possibly with microphone directional patterns), downmix from a multi-channel audio signal (e.g., 5.1), etc.
[0078] Determining the carrier audio signal type based on analysis of the carrier audio signal itself can be based on comparing frequency bands or spectral effects that combine (in different ways) with an expected spectral effect (based in part on spatial metadata, if available).
[0079] Furthermore, in some embodiments, the processing of audio signals can include rendering of Ambisonic signals, rendering of multi-channel audio signals (such as 5.1), and transport of downmixes to a smaller number of audio signals:
[0080] 2 shows a schematic diagram of an example decoder suitable for implementing some embodiments. This embodiment may be implemented, for example, in the "Demultiplexer / Decoder / Synthesizer" block 133. In this example, the input is a Metadata-Assisted Spatial Audio (MASA) stream containing two audio channels and spatial metadata. However, as discussed herein, the input format may be any suitable metadata-assisted spatial audio format.
[0081] The (MASA) bitstream is forwarded to a carrier audio signal type determiner 201. The carrier audio signal type determiner 201 is configured to determine a carrier audio signal type 202 based on the bitstream, and possibly some additional parameters 204 (such as microphone distance). The determined parameters are forwarded to a MASA-to-Ambisonic signal converter 203.
[0082] The MASA-Ambisonic signal converter 203 is configured to receive the bitstream and the carrier audio signal type 202 (and possibly some additional parameters 204) and is configured to convert the MASA stream into an Ambisonic signal based on the determined carrier audio signal type 202 (and possible additional parameters 204).
[0083] The operation of the example is summarized in the flow diagram shown in Figure 3.
[0084] The first action is one of receiving or acquiring a bitstream (the MASA stream), as shown in FIG.
[0085] The next operation is one of determining the carried audio signal type based on the bitstream (and generating a type signal or indicator and possibly other additional parameters), as indicated in FIG. 3 by step 303.
[0086] The next operation having determined the carrier audio signal type is to convert the bitstream (the MASA stream) into an Ambisonic signal based on the determined carrier audio signal type, as illustrated in FIG. 3 by step 305.
[0087] 4 shows a schematic diagram of an example carrier audio signal type determiner 201. In this example, the example carrier audio signal type determiner is suitable for the case where the carrier audio signal type is available in the MASA stream.
[0088] An example of the carrier audio signal type determiner 201 in this example includes a carrier audio signal type extractor 401. The carrier audio signal type extractor 401 is configured to receive a bit (MASA) stream and extract (i.e., read and / or decode) a type indicator from the MASA stream. This type of information is available, for example, in a "Channel Audio Format" field of the MASA stream. In addition, if additional parameters are available, they are also extracted. This information is output from the carrier audio signal type extractor 401. In some embodiments, the carrier audio signal type can include "space", "downmix", "match". In some other embodiments, the carrier audio signal type can include any suitable value.
[0089] 5 shows a schematic diagram of a further example carrier audio signal type determiner 201. In this example, the carrier audio signal type cannot be directly extracted or decoded from the MASA stream. In this example, the carrier audio signal type is estimated or determined from an analysis of the MASA stream. This determination in some embodiments is based on using a set of estimators / energy comparisons that account for certain spectral effects of different carrier audio signal types.
[0090] In an embodiment, the carrier audio signal type determiner 201 comprises a carrier audio signal and spatial metadata extractor / decoder 501. The carrier audio signal and spatial metadata extractor / decoder 501 is configured to receive a MASA stream and to extract and / or decode the carrier audio signal and spatial metadata from the MASA stream. The resulting carrier audio signal 502 can be forwarded to a time-to-frequency converter 503. The resulting spatial metadata 522 can be further forwarded to a subtraction to a target energy comparator 511.
[0091] In some embodiments, the carrier audio signal type determiner 201 includes a time-to-frequency converter 503. The time-to-frequency converter 503 is configured to receive the carrier audio signals 502 and convert them into the time-frequency domain. Suitable transforms include, for example, the short-time Fourier transform (STFT) and the complex-modulated quadrature mirror filter bank (QMF). The resulting signal is represented as S i The T / F domain carrier audio signal 504 may be represented as (b,n), where i is the channel index, b is the frequency bin index, and n is the time index. In situations where the carrier audio signal (output from the extractor and / or decoder) is already in the time-frequency domain, this may be omitted or may involve a conversion from one time-frequency domain representation to another. The T / F domain carrier audio signal 504 may be forwarded to a comparator.
[0092] In an embodiment, the carrier audio signal type determiner 201 includes a broadband L / R total energy comparator 505. The broadband L / R vs. total energy comparator 505 is configured to receive the T / F domain carrier audio signal 504 and to output a broadband L / R vs. total ratio parameter.
[0093] In the Broadband L / R to Total Energy Comparator 505 the broadband left, right and total energy are calculated.
number
number
number
[0094] Broadband L / R vs. total energy comparator 505 may then generate broadband L / R vs. total energy ratio 506 as follows:
number
[0095] In some embodiments, the carrier audio signal type determiner 201 includes a high-frequency L / R-total energy comparator 507. The high-frequency L / R-total energy comparator 507 is configured to receive the T / F domain carrier audio signal 504 and to output a high-frequency L / R-total ratio parameter.
[0096] In the broadband L / R-total energy comparator 507 the left, right and total energies of the high frequency band are calculated.
number
number
[0097] High frequency L / R vs. total energy comparator 507 may then be configured to select the smaller of the left and right energies, and the result is multiplied by two.
number
[0098] High frequency L / R vs. total energy comparator 507 may then generate a high frequency L / R vs. total ratio 508 .
number
[0099] In some embodiments, the carrier audio signal type determiner 201 includes a total energy comparator 509. The sum to sum to total energy comparator 509 is configured to receive the T / F domain carrier audio signal 504 and output a sum to total energy ratio parameter. The sum to sum to total energy comparator 509 is configured to detect situations where, at some frequencies, the two channels are out of phase, which is a typical phenomenon especially for spaced microphone recordings.
[0100] The summation to sum-to-total energy comparator 509 is configured to calculate the energy of the total signal and the total energy for each frequency bin.
number
[0101] These energies are, for example,
number
[0102] The sum-to-total energy comparator 509 is then configured to calculate a minimum sum-to-total ratio 510 as follows:
number
[0103] A sum to sum energy comparator 509 is then configured to output the ratio χ(n) 510 .
[0104] In some embodiments, the carrier audio signal type determiner 201 includes a subtract to target energy comparator 511 configured to receive the T / F domain carrier audio signal 504 and the spatial metadata 522 and to output a subtract to target energy ratio parameter 512.
[0105] The subtraction to target energy comparator 511 is configured to calculate the difference energy between the left and right channels.
number
[0106] This can be thought of as a "prototype" of the Ambisonics Y signal, at least for some input signal types (the Y signal has a dipole directional pattern, with a positive lobe on the left and a negative lobe on the right).
[0107] Then, the target energy comparator 511 subtracts the target energy E for the Y signal. target (b,n), which is based on estimating how the total energy should be distributed among the spherical harmonics based on the spatial metadata. For example, in some embodiments, the subtraction to target energy comparator 511 is configured to build a target covariance matrix (channel energy and cross-correlation) based on the spatial metadata and the energy estimates. However, in some embodiments, only the energy of the Y signal is estimated, which is one entry of the target covariance matrix. Thus, the target energy E of Y target (b,n) consists of two parts.
number
number
[0108] Note that the spatial metadata may be at a lower frequency and / or time resolution than every b,n, such that the parameters may be the same for several frequency or time indices.
[0109] This E (target,dir) (b,n) is the energy of the more directional part. To formulate it, we use the spread coherence of the spatial metadata c spread (b,n) The spread coherence distribution vector as a function of parameters 0–1,
number
[0110] Subtraction to the target energy comparator 511 is a vector of azimuth angle values,
number
number
[0111] Therefore, E target (b,n) are obtained. In some embodiments, these energies can be calculated using, for example,
number
[0112] Additionally, the subtraction to target energy comparator 511 is configured to calculate the subtraction to target ratio 512 using the energy in the lowest frequency bin as follows:
number
[0113] In an embodiment, the carrier audio signal type determiner 201 includes a carrier audio signal type (based on estimated metrics) determiner 513. The carrier audio signal type determiner 513 is configured to receive the broadband L / R to total ratio 506, the high frequency L / R to total ratio 508, the minute sum to total ratio 510, and the subtraction to target ratio 512, and to determine the carrier audio signal type based on these estimated metrics.
[0114] The decision can be made in various ways, and the actual implementation may differ in many aspects, such as the T / F conversion used. One non-limiting example form is for the Carrier Audio Signal Type (Based on Estimated Metric) Decider 513 to first calculate the changes to the non-metrics.
number
[0115] The carrier audio signal type (based on the estimation metric) determiner 513 may then be configured to calculate the changes to the downmix metric.
number
[0116] A carrier audio signal type (based on estimated metrics) determiner 513 can then determine, based on these metrics, whether the carrier audio signal originates from spaced microphones or is a downmix from a surround sound signal (such as 5.1). For example,
number
[0117] In this example, the carrier audio signal type (based on estimated metric) determiner 513 does not detect a matching microphone type, however in practice processing according to T(n)=“downmix” type can generally produce good audio in case of matched capture (e.g. with cardioid steered left and right).
[0118] The carrier audio signal type (based on the estimated metric) determiner 513 can then be configured to output the carrier audio signal type as the carrier audio signal type 202. In some embodiments, other parameters 204 may be output.
[0119] FIG. 6 summarizes the operation of the apparatus shown in FIG. 5 and thus, in some embodiments, the first operation is to extract and / or decode the carrier audio signal and metadata from the MASA stream (or bitstream), as shown in FIG. 6 by step 601.
[0120] The next operation may be a time-to-frequency domain transformation of the carrier audio signal, as shown in FIG.
[0121] A series of comparisons can then be made, for example by comparing the broadband L / R energy to the total energy value to generate a broadband L / R to total energy ratio as shown in FIG.
[0122] For example, by comparing the high frequency L / R energy to the total energy value, step 607 can generate a high frequency L / R to total energy ratio, as shown in FIG.
[0123] By comparing the sum energy to the total energy value, a sum-to-total energy ratio may be generated by step 609, as shown in FIG.
[0124] Additionally, step 611 may generate a subtraction-to-target energy ratio, as shown in FIG.
[0125] After determining these metrics, the method may determine the carried audio signal type by analyzing the ratio of these metrics, as shown in FIG.
[0126] 7 shows in further detail an example of a MASA to Ambisonic converter 203. The MASA to Ambisonic converter 203 is configured to receive a MASA stream (bitstream) and a carrier audio signal type 202 and possible additional parameters 204, and is configured to convert the MASA stream into an Ambisonic signal based on the determined carrier audio signal type.
[0127] The MASA to Ambisonic converter 203 includes a carrier audio signal and spatial metadata extractor / decoder 501. It is configured to receive the MASA stream and output a carrier audio signal 502 and spatial metadata 522 in the same way as found in the carrier audio signal type determiner as shown in Fig. 5. In some embodiments, the extractor / decoder 501 is an extractor / decoder from the carrier audio signal type determiner. The resulting carrier audio signal 502 can be forwarded to a time / frequency converter 503. The resulting spatial metadata 522 can further be forwarded to a signal mixer 705.
[0128] In one embodiment, the MASA to Ambisonic converter 203 includes a time-to-frequency converter 503 configured to receive the carrier audio signals 502 and convert them into the time-frequency domain. Suitable transforms include, for example, the Short-Time Fourier Transform (STFT) and the Complex Modulation Quadrature Mirror Filter Bank (QMF). The resulting signal is i The time-frequency converter 503 is represented as (b,n), where i is the channel index, b is the frequency bin index, and n is the time index. If the output of the audio extraction and / or decoding is already in the time-frequency domain, this block may be omitted or may include a conversion from one time-frequency domain representation to another. The T / F domain carrier audio signal 504 may be forwarded to the prototype signal creator 701. In some embodiments, the time-frequency converter 503 is the same time-frequency converter from the carrier audio signal type determiner.
[0129] In one embodiment, the MASA to Ambisonic converter 203 includes a prototype signal creator 701. The prototype signal creator 701 is configured to receive the T / F domain carrier audio signal 504, the carrier audio signal type 202, and possible additional parameters 204. The T / F prototype signal 702 can then be output to a signal mixer 705 and a decorrelator 703.
[0130] In one embodiment, the MASA to Ambisonic converter 203 includes a decorrelator 703. The decorrelator 703 is configured to receive the T / F prototype signal 702, apply decorrelation, and output a decorrelated T / F prototype signal 704 to a signal mixer 705. In some embodiments, the decorrelator 703 is optional.
[0131] In one embodiment, the MASA to Ambisonic converter 203 includes a signal mixer 705. The signal mixer 705 is configured to receive the T / F prototype signal 702 and the decorrelated T / F prototype signal and spatial metadata 522.
[0132] The prototype signal creator 701 is configured to generate a prototype signal for each of the Ambisonic (FOA / HOA) spherical harmonic functions based on the carrier audio signal type.
[0133] In some embodiments, the prototype signal creator 701 is configured to operate as follows: If T(n)=“spaced”, create a prototype for the W signal:
number
[0134] In fact, W proto(b,n) can be created as a means to carry low frequency audio signals. The signals are roughly in phase, no comb filtering is performed, and one of the high frequency channels is selected. The value of B3 depends on the T / F conversion and the distance between the microphones. If the distance is unknown, some default value may be used (such as the value corresponding to 1kHz). If T(n)=“downmix” or T(n)=“coincident”, the W signal can be prototyped as follows:
number
[0135] The original audio signal can usually be assumed to have no significant delay between these signal types, so proto (b,n) is created by summing the carrier audio signals.
[0136] Regarding the Y prototype signal, if T(n)=“spaced”, then the prototype of the Y signal can be created as follows:
number
[0137] At mid-frequency (between B4 and B5) a dipole signal can be created by subtracting the transport signal, shifting the phase by -90 degrees and equalizing it, so it serves as a good prototype of the Y signal, especially if the microphone distance is known and the equalization coefficients are therefore appropriate. At low and high frequencies this is not feasible and a prototype signal is generated similar to that for the omnidirectional W signal.
[0138] If the microphone distances are precisely known, the Y prototypes can be calculated by dividing the frequency by the number of microphones at each frequency (i.e., Y(b,n)=Y proto (b,n)) may be used directly for Y. If the microphone spacing is not known, g eq (b)=1 can be used.
[0139] In some embodiments, the signal mixer 705 applies gain processing in the frequency bands to reduce the W in the frequency bands to a target energy in the frequency band with potential gain smoothing. proto The energy of (b,n) can be corrected. The target energy of the omnidirectional signal in a frequency band can be the sum of the carrier audio signal energies in that frequency band. The result of this processing is the omnidirectional signal W(b,n).
[0140] Y proto For Y signals where (b,n) cannot be used directly for Y(b,n), adaptive gain processing is performed if the frequency is between B4 and B5. This case is similar to the omnidirectional W case above. The prototype signal is already a Y dipole, except for the potentially incorrect spectrum. The signal mixer performs gain processing of the prototype signal in the frequency band. (Furthermore, in this particular context, decorrelation processing of the Y signal is not necessary). The gain processing uses spatial metadata (direction, ratio, other parameters) and an overall signal energy estimate in the frequency band (e.g., the sum of the carrier signal energies) to determine what the energy of the Y component should be in the frequency band, then gain corrects the prototype signal's energy in the frequency band, which is the determined energy, and then the result is the output Y(b,n).
[0141] The procedure to generate Y(b,n) mentioned above is not valid for all frequencies in the current context T(n) = “spaced”. Since the prototype signals are different at different frequencies, the signal mixers and decorrelators are configured differently depending on the frequency with this transport signal type. To illustrate the different kinds of prototype signals, we can consider a scenario where sound arrives from the negative gain direction of the Y dipole (having positive and negative lobes). At mid-frequencies (between B4 and B5), the phase of the Y prototype signal is opposite to that of the W prototype signal, as it should be for that direction of the arriving sound. At other frequencies (below B4 and above B5), the phase of the prototype Y signal is the same as that of the W prototype signal. The synthesis of the appropriate phase (and energy and correlation) is then accounted for by the signal mixers and decorrelators at those frequencies.
[0142] At low frequencies (below B4) where the wavelength is large, the phase difference between audio signals captured by spaced microphones (usually somewhat closer to each other) is small. Therefore, the creator of the prototype signal should not set it to generate the prototype signal in the same way as for frequencies between B4 and B5 for SNR reasons. Therefore, typically, a channel-sum omnidirectional signal is used instead as the prototype signal. At high frequencies (above B5) where the wavelength is small, spatial aliasing will severely distort the beam pattern (if a method like that for frequencies between B4 and B5 is used). Therefore, it is better to use a channel-select omnidirectional prototype signal.
[0143] The signal mixer and decorrelator configurations at these frequencies (B4 and below or B5 and above) are described below. In a simple example, the spatial metadata parameter settings consist of a frequency band orientation θ and a ratio r. A gain sin(θ)sqrt(r) is applied to the prototype signal in the signal mixer to generate the Y dipole signal, the result of which is the coherent partial signal. The prototype signal is also decorrelated (in the decorrelator) and the decorrelated result is received at the signal mixer, where the coefficient sqrt(1-r)gorder The result is a decorrelated partial signal with gain g order is the diffuse field gain in spherical harmonic order according to the known SN3D normalization scheme. For example, for the first order (in this case for the Y dipole) it is sqrt(1 / 3), for the second order it is sqrt(1 / 5), for the third it is sqrt(1 / 7) and so on. The coherent and incoherent partial signals are added together. The result is the combined Y signal, minus the erroneous energy, as the prototype signal energy may be wrong. The same energy correction procedure in the frequency bands described in the context of mid-frequencies (between B4 and B5) can be applied to correct the energy in the frequency bands to the desired target, and the output is the signal Y(b,n).
[0144] For other spherical harmonics, such as X, Z components and second and higher order components, the above procedure can be applied, except that the azimuth gain (and other potential parameters) depends on which spherical harmonic signals are being combined. For example, the gain to generate for the X dipole coherent part from the W prototype is cos(θ)sqrt(r). The decorrelation, proportion-processing and energy correction can be the same as determined above for the Y component other than the frequencies between B4 and B5.
[0145] Other parameters such as altitude, spread coherence and surround coherence can be taken into account in the above steps. The spread coherence parameter can have values between 0 and 1. A coherence spread value of 0 indicates a point source. In other words, when playing an audio signal using a multi-loudspeaker system, the sound should be played through as few loudspeakers as possible (for example, only the center loudspeaker if the direction is centered). As the value of spread coherence increases, more energy is spread to other loudspeakers around the center loudspeaker, until a value of 0.5, the energy is spread evenly between the center and neighboring loudspeakers. As the value of spread coherence increases above 0.5, the energy in the center loudspeaker decreases, until a value of 1, where there is no energy in the center loudspeaker and all the energy is in the neighboring loudspeakers. The surrounding coherence parameter has values between 0 and 1. A value of 1 means there is coherence between all (or nearly all) loudspeaker channels. A value of 0 means there is no coherence between all (or nearly all) loudspeaker channels. This is further explained in GB application no. 1718341.9 as well as PCT application no. PCT / FI2018 / 050788.
[0146] For example, increased surround coherence can be implemented by reducing the synthesis ambience energy in the spherical harmonic components, and elevation can be added by adding elevation-related gain according to the definition of the Ambisonic pattern in the generation of the coherent part.
[0147] If T(n)=“downmix” or T(n)=“coincident”, the prototype of the Y signal is
number
[0148] In this situation, it can be assumed that the original audio signal usually does not have a significant delay between these signal types, so no phase shift is necessary. For the "mixed signal" block, if T(n)=“coincident”, the Y and W prototypes may be used directly for the Y and W outputs, possibly after gaining (depending on the actual directional pattern). If T(n)=“downmix”, the Y proto (b,n) and W proto (b,n) cannot be used directly for Y(b,n) and W(b,n), except in cases where it is necessary to compensate for the energy in the frequency band to the desired target determined when T(n) = "spaced" (note that the omnidirectional component has a spatial gain of 1 regardless of the angle of sound arrival).
[0149] Other spherical harmonics (such as X or Z) cannot create prototypes that reproduce the target signal well because typical downmix signals are oriented on the left-right axis, rather than the front-back X axis or the top-bottom Z axis. Thus, in some embodiments, an approach is to utilize prototypes of, for example, omnidirectional signals.
number
[0150] Similarly, W proto (b,n) is also used for higher harmonics for the same reason. Signal mixers and decorrelators in this situation can process signals for these spherical harmonics in a similar way to the T(n)=“spaced” case.
[0151] In some cases, the type of the carrier audio signal may change during audio playback (e.g., due to changes in the actual signal type or imperfections in automatic type detection). To avoid artifacts due to abrupt type changes, the prototype signal in some embodiments may be interpolated. This may be achieved, for example, by simply linearly interpolating from the prototype signal corresponding to the old version to the prototype signal corresponding to the new version.
[0152] The output of the signal mixer is the resulting time-frequency domain Ambisonic signal, which is forwarded to an inverse T / F transformer 707 .
[0153] In some embodiments, the MASA-Ambisonic signal converter 203 includes an inverse T / F transformer 707 configured to convert the signal to the time domain. A time-domain Ambisonic signal 906 is output from the MASA-Ambisonic signal converter.
[0154] With reference to FIG. 8, an overview of the operation of the apparatus shown in FIG. 7 is given.
[0155] Thus, in one embodiment, the first operation is that of extracting and / or decoding the carrier audio signal and metadata from the MASA stream (or bitstream), as illustrated in FIG. 8 by step 801 .
[0156] The next operation may be a time-to-frequency domain transformation of the carrier audio signal, as illustrated in FIG.
[0157] The method then includes creating a prototype audio signal based on the time-frequency domain carrier signal, and further creating a prototype audio signal based on the type of the carrier audio signal (and further based on additional parameters), as shown in FIG. 8 by step 805.
[0158] In some embodiments, the method includes applying decorrelation on the time-frequency prototype audio signal, as illustrated in FIG. 8 by step 807 .
[0159] The uncorrelated time-frequency prototype audio signals and the time-frequency prototype audio signals may then be mixed based on the spatial metadata and the carried audio signal type, as shown in FIG. 8, via step 809 .
[0160] The mixed signal may then be inverse time-frequency transformed, as shown in FIG.
[0161] A time domain signal can then be output, as shown in FIG. 8, via step 813 .
[0162] 9 shows a schematic diagram of an example decoder suitable for implementing some embodiments. This example may be implemented, for example, in the "Demultiplexer / Decoder / Synthesizer" block 133 shown in FIG. 1, where the input is a Metadata-Assisted Spatial Audio (MASA) stream that includes two audio channels and spatial metadata. However, as discussed herein, the input format may be any suitable metadata-assisted spatial audio format.
[0163] The (MASA) bitstream is forwarded to a carrier audio signal type determiner 201. The carrier audio signal type determiner 201 is configured to determine a carrier audio signal type 202 and possibly some additional parameters 204 (such as microphone distance) based on the bitstream. The determined parameters are forwarded from the MASA to the multi-channel audio signal converter 903. The carrier audio signal type determiner 201 in some embodiments may be the same carrier audio signal type determiner 201 as described above with respect to FIG. 2 or may be a separate instance of the carrier audio signal type determiner 201 configured to operate similarly to the carrier audio signal type determiner 201 as described above with respect to the example shown in FIG. 2.
[0164] The MASA to multi-channel audio signal converter 903 is configured to receive the bitstream and the transport audio signal type 202 (and possibly some additional parameters 204) and is configured to convert the MASA stream into a multi-channel audio signal (e.g., 5.1) based on the determined transport audio signal type 202 (and possible additional parameters 204).
[0165] The operation of the example shown in FIG. 9 is summarized in the flow diagram shown in FIG. 10.
[0166] The first action is one of receiving or acquiring a bitstream (the MASA stream), as shown in FIG.
[0167] The next operation, as indicated in FIG. 10 by step 303, is one of determining the carried audio signal type based on the bitstream (and generating a type signal or indicator and possibly other additional parameters).
[0168] Once the carrier audio signal type has been determined, the next operation is to convert the bitstream (MASA stream) into a multi-channel audio signal (such as 5.1) based on the determined carrier audio signal type, as shown in FIG. 10 by step 1005.
[0169] 11 shows in further detail an exemplary MASA to multi-channel audio signal converter 903. The MASA to multi-channel audio signal converter 903 is configured to receive a MASA stream (bitstream) and a carrier audio signal type 202 and possible additional parameters 204, and is configured to convert the MASA stream into a multi-channel audio signal based on the determined carrier audio signal type.
[0170] The MASA to multi-channel audio signal converter 903 includes a carrier audio signal and spatial metadata extractor / decoder 501. It is configured to receive the MASA stream and output a carrier audio signal 502 and spatial metadata 522 in the same way as found in the carrier audio signal type determiner, as shown in Fig. 5 and discussed. In an embodiment, the extractor / decoder 501 is the extractor / decoder from the carrier audio signal type determiner described above, or a separate instance of the extractor / decoder. The resulting carrier audio signal 502 can be forwarded to a time-to-frequency converter 503. The resulting spatial metadata 522 can further be forwarded to a target signal characteristics determiner 1101.
[0171] In some embodiments, the MASA to multi-channel audio signal converter 903 includes a time / frequency converter 503 configured to receive the carrier audio signals 502 and convert them into the time-frequency domain. Suitable transforms include, for example, the short-time Fourier transform (STFT) and the complex-modulated quadrature mirror filter bank (QMF). The resulting signal is then denoted as S iLet (b,n) be the time index, where i is the channel index, b is the frequency bin index, and n is the time index. Here, is the channel index, the frequency bin index, and is the time index. If the output of the audio extraction and / or decoding is already in the time-frequency domain, this block may be omitted or may include a conversion from one time-frequency domain representation to another. The T / F domain carrier audio signal 504 may be forwarded to a prototype signal creator 1111. In some embodiments, the time / frequency converter 503 is the same time / frequency converter from the carrier audio signal type determiner or the MASA-Ambisonic converter or a separate instance. In some embodiments, the MASA to multichannel audio signal converter 903 includes the prototype signal creator 1111.
[0172] The prototype signal creator 1111 is configured to receive the T / F domain carrier audio signal 504, the carrier audio signal type 202, and possible additional parameters 204. The T / F prototype signal 1112 can then be output to the signal mixer 1105 and the decorrelator 1103.
[0173] As an example of the operation of the prototype signal creator 1111a, rendering to a 5.1 multi-channel audio signal configuration will be described. In this example, the prototype signal for the left (left front and left surround) output channel is
number
number
[0174] Therefore, for output channels to either side of the mid-plane, the prototype signal can directly use the corresponding carrier audio signal. For the center output channel, the prototype audio signal must contain energy from the left and right, since it can be used for panning to either side. Therefore, the prototype signal can be created in the same way as for the omnidirectional channels for Ambisonic rendering, i.e., if T(n)=“spaced”,
number
number
[0175] In one embodiment, the MASA to multi-channel audio signal converter 903 includes a decorrelator 1103. The decorrelator 1103 is configured to receive the T / F prototype signal 1112, apply decorrelation, and output a decorrelated T / F prototype signal 1104 to the signal mixer 1105. In some embodiments, the decorrelator 1103 is optional.
[0176] In an embodiment, the MASA to multi-channel audio signal converter 903 includes a target signal characteristic determiner 1101. The target signal characteristic determiner 1101 in some embodiments is configured to generate a target covariance matrix (target signal characteristic) in a frequency band based on spatial metadata and a global estimate of the signal energy in the frequency band. In some embodiments, this energy estimate can be the sum of the carrier signal energies in the frequency band. This target covariance matrix (target signal characteristic) determination can be performed in a similar manner as provided by patent application GB 1718341.9.
[0177] The target signal characteristics 1102 may then be passed to a signal mixer 1105 .
[0178] In some embodiments, the MASA to multi-channel audio signal converter 903 includes a signal mixer 1105. The signal mixer 1105 is configured to measure the covariance matrix of the prototype signal and formulates a mixing solution based on the estimated (prototype signal) covariance matrix and the target covariance matrix. In some embodiments, the mixing solution may be similar to that described in GB1718341.9. The mixing solution is applied to the prototype signal and the uncorrelated prototype signal, and the resulting signal is obtained with frequency band characteristics based on the target signal characteristics, i.e., based on the determined target covariance matrix. In some embodiments, the MASA to multi-channel audio signal converter 903 includes an inverse T / F transformer 707 configured to convert the signal to the time domain. The time domain multi-channel audio signal is output from the MASA to the multi-channel audio signal converter.
[0179] With reference to FIG. 12, an overview of the operation of the device shown in FIG. 11 is given.
[0180] Thus, in one embodiment, the first operation is that of extracting and / or decoding the carrier audio signal and metadata from the MASA stream (or bitstream), as illustrated in FIG. 12 by step 801.
[0181] The next operation may be a time-to-frequency domain transformation of the carrier audio signal, as illustrated in FIG.
[0182] The method then includes creating a prototype audio signal based on the time-frequency domain carrier signal and further creating a prototype audio signal based on the type of the carrier audio signal (and further based on additional parameters), as shown in FIG. 12, by step 1205.
[0183] In some embodiments, the method includes applying decorrelation on the time-frequency prototype audio signal, as shown in FIG. 12, by step 1207.
[0184] Then, via step 1208, target signal characteristics may be determined based on the time-frequency domain carrier audio signal and the spatial metadata (to generate a covariance matrix of the target signal), as shown in FIG.
[0185] The covariance matrix of the prototype audio signal can be determined by step 1209 as shown in FIG.
[0186] Then, via step 1209, the decorrelated time-frequency prototype audio signal and the time-frequency prototype audio signal may be mixed based on the target signal characteristics, as shown in FIG.
[0187] The mixed signal may then be inverse time-frequency transformed, as shown in FIG.
[0188] The time domain signal may then be output by step 1213 as shown in FIG.
[0189] FIG. 13 shows a schematic diagram of a further example decoder suitable for implementing some embodiments. In other embodiments, a similar method may be implemented in a device other than a decoder, for example as part of an encoder. This example may be implemented, for example, in an (IVAS) Demultiplexer / Decoder / Synthesizer block 133, as shown in FIG. 1, where in this example the input is a Metadata-Assisted Spatial Audio (MASA) stream that includes two audio channels and spatial metadata. However, as discussed herein, the input format may be any suitable metadata-assisted spatial audio format.
[0190] The (MASA) bitstream is forwarded to a carrier audio signal type determiner 201. The carrier audio signal type determiner 201 is configured to determine the carrier audio signal type 202 and possibly some additional parameters 204 (one example of such additional parameters is microphone distance) based on the bitstream. The determined parameters are forwarded to the downmixer 1303. The carrier audio signal type determiner 201 in some embodiments may be the same carrier audio signal type determiner 201 as described above or a separate instance of the carrier audio signal type determiner 201 configured to operate similarly to the carrier audio signal type determiner 201 as described above.
[0191] The downmixer 1303 is configured to receive the bitstream and the carrier audio signal type 202 (and possibly some additional parameters 204) and is configured to downmix the MASA stream from two carrier audio signals to one carrier audio signal based on the determined carrier audio signal type 202 (and possible additional parameters 204). An output MASA stream 1306 is then output.
[0192] The operation of the example shown in FIG. 13 is summarized in the flow diagram shown in FIG. 14.
[0193] The first action is to receive or acquire a bitstream (MASA stream), as shown in FIG.
[0194] The next action is to determine the carried audio signal type based on the bitstream (and generate a type signal or indicator and possibly other additional parameters), as indicated in FIG. 14 by step 303.
[0195] After determining the type of the carrier audio signal, the next operation is to downmix the MASA stream from two carrier audio signals to one carrier audio signal based on the determined type 202 of the carrier audio signal (and possible additional parameters 204), as shown in FIG. 14 by step 1405.
[0196] 15 shows in more detail an example of a downmixer 1303. The downmixer 1303 is configured to receive a MASA stream (bitstream) and a carrier audio signal type 202 and possible additional parameters 204, and is configured to downmix two carrier audio signals into one carrier audio signal based on the determined carrier audio signal type.
[0197] The downmixer 1303 includes a carrier audio signal and spatial metadata extractor / decoder 501, which is configured to receive the MASA stream and output a carrier audio signal 502 and spatial metadata 522 in the same manner as found in the carrier audio signal type determiner discussed therein. In an embodiment, the extractor / decoder 501 is the extractor / decoder described above, or a separate instance of an extractor / decoder. The resulting carrier audio signal 502 can be forwarded to a time-to-frequency converter 503. The resulting spatial metadata 522 can further be forwarded to a signal multiplexer 1507.
[0198] In some embodiments, the downmixer 1303 includes a time-to-frequency converter 503 configured to receive the carrier audio signals 502 and convert them into the time-frequency domain. Suitable transforms include, for example, the short-time Fourier transform (STFT) and the complex-modulated quadrature mirror filter bank (QMF). The resulting signal is denoted by S i The T / F domain carrier audio signal 504 may be forwarded to the prototype signal creator 1511. In some embodiments, the time / frequency converter 503 is the same time / frequency converter as described above, or a separate instance.
[0199] In some embodiments, the downmixer 1303 includes a prototype signal creator 1511. The prototype signal creator 1511 is configured to receive the T / F domain carrier audio signal 504, the carrier audio signal type 202, and possible additional parameters 204. The T / F prototype signal 1512 can then be output to a proto-energy determiner 1503, which can align the prototype signal to a target energy equalizer 1505.
[0200] The prototype signal creator 1511 in some embodiments is configured to create a prototype signal of a mono carrier audio signal using two carrier audio signals based on the received carrier audio signal type. For example, the following can be used: If T(n)=“spaced”, then:
number
number
[0201] In some embodiments, the downmixer 1303 includes a target energy determiner 1501. The target energy determiner 1501 receives the T / F domain carrier audio signal 504 and determines a target energy value as a sum of the energies of the carrier audio signals.
number
[0202] The target energy values can then be passed to a proto to match target equalizer 1505 .
[0203] In some embodiments, the downmixer 1303 includes a proto-energy determiner 1503. The proto-energy determiner 1503 receives a T / F prototype signal 1512, e.g.
number
[0204] The proto energy values can then be passed to the proto to match the target equalizer 1505 .
[0205] The downmixer 1303 in some embodiments includes a prototype for matching a target energy equalizer 1505. The prototype for matching a target energy equalizer 1505 in some embodiments is configured to receive the T / F prototype signal 1502, the prototype energy values, and the target energy values. The equalizer 1505 in some embodiments may first
number
number
[0206] The prototype signal can then be equalized with these gains as follows:
number
[0207] In some embodiments, the downmixer 1303 includes an inverse T / F transformer 707 configured to convert the output of the equalizer into a time-domain version. The time-domain equalized audio signal (mono signal) 1510 is then passed to the carrier audio signal and spatial metadata multiplexer 1507 (or multiplexer).
[0208] In some embodiments, the downmixer 1303 includes a carrier audio signal and spatial metadata multiplexer 1507 (or multiplexer). The carrier audio signal and spatial metadata multiplexer 1507 (or multiplexer) is configured to receive the spatial metadata 522 and the mono audio signal 1510 and multiplex them to regenerate the appropriate output format (e.g., a MASA stream having only one carrier audio signal) 1506. In some embodiments, the input mono audio signal is in pulse code modulation (PCM) format. In such embodiments, the signal may be encoded as well as multiplexed. In some embodiments, the multiplexing may be omitted and the mono carrier audio signal and the spatial metadata are used directly in the audio encoder.
[0209] In one embodiment, the output of the apparatus shown in FIG. 15 is a mono PCM audio signal 1510 in which the spatial metadata is discarded.
[0210] In some embodiments, other parameters may be implemented, for example, in some embodiments, if the type is "spaced", the spaced microphone distance may be estimated.
[0211] With reference to FIG. 16, the operation of one example of the apparatus shown in FIG. 15 is illustrated.
[0212] Thus, in one embodiment, the first operation is that of extracting and / or decoding the carrier audio signal and metadata from the MASA stream (or bitstream), as illustrated in FIG. 16 by step 1601.
[0213] The next operation may be a time-to-frequency domain transformation of the carrier audio signal, as indicated in FIG.
[0214] The method then includes creating a prototype audio signal based on the time-frequency domain carrier signal and further creating a prototype audio signal based on the type of the carrier audio signal (and further based on additional parameters), as shown in FIG. 16, by step 1605.
[0215] Further, in some embodiments, the method is configured to generate, determine, or calculate a target energy value based on the transformed carrier audio signal, as illustrated in FIG. 16 by step 1604.
[0216] Further, in some embodiments, the method is configured to generate, determine, or calculate, by step 1606, a prototype audio signal energy value based on the prototype audio signal energy value, as shown in FIG.
[0217] After determining the energy, the method may further equalize the prototype audio signal to match the target audio signal energy, as shown in FIG. 16, via step 1607 .
[0218] The equalized prototype signal (mono signal) may then be inverse time-to-frequency domain transformed to generate a time-domain mono signal, as shown in FIG. 16, by step 1609.
[0219] The time-domain mono audio signal is then (optionally) encoded and multiplexed with the spatial metadata, as shown in FIG. 16, via step 1610.
[0220] The multiplexed audio signal can then be output (as a MASA data stream) as shown in FIG.
[0221] As mentioned above, the illustrated block diagram is only one example of a possible implementation. Other practical implementations may differ from the above example. For example, an implementation may not have a separate T / F converter.
[0222] Furthermore, rather than having an input MASA stream as shown above, in some embodiments any suitable bitstream that makes use of audio channels and (spatial) metadata can be used. Furthermore, in some embodiments the IVAS codec can be replaced with any other suitable codec (e.g. one that has audio channels and spatial metadata modes of operation).
[0223] In some embodiments, the carrier audio signal type determiner can be used to estimate parameters other than the carrier audio signal type, for example, microphone spacing. Microphone spacing is one example of a possible additional parameter 204. This can be used in some embodiments to estimate the E sum(b,n) and E sub This can be achieved by examining the frequencies of maxima and minima in (b,n), determining the time delay between the microphones based on those, and estimating the spacing based on the delay and the estimated direction of arrival (available in the spatial metadata). There are also methods to estimate the delay between the two signals.
[0224] With reference to Figure 17, an example of an electronic device that may be used as an analysis or synthesis device is shown. The device may be any suitable electronic device or equipment. For example, in an embodiment, device 1700 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.
[0225] In one embodiment, the apparatus 1700 includes at least one processor or central processing unit 1707. The processor 1707 can be configured to execute various program code, such as the methods described herein.
[0226] In an embodiment, the apparatus 1700 includes a memory 1711. In an embodiment, at least one processor 1707 is coupled to the memory 1711. The memory 1711 can be any suitable storage means. In an embodiment, the memory 1711 includes a program code section for storing program code implementable on the processor 1707. Additionally, in some embodiments, the memory 1711 can further include a stored data section for storing, for example, data that has been processed or is to be processed according to the embodiments described herein. The implemented program code stored in the program code section and the data stored in the stored data section can be retrieved by the processor 1707 whenever needed via the memory-processor coupling.
[0227] In some embodiments, the device 1700 includes a user interface 1705. The user interface 1705 can be coupled to the processor 1707 in some embodiments. In some embodiments, the processor 1707 can control the operation of the user interface 1705 and receive input from the user interface 1705. In some embodiments, the user interface 1705 can allow a user to input commands to the device 1700, for example, via a keypad. In some embodiments, the user interface 1705 can allow a user to obtain information from the device 1700. For example, the user interface 1705 can include a display configured to display information from the device 1700 to the user. The user interface 1705 can comprise a touch screen or touch interface in some embodiments that can both allow information to be input into the device 1700 and further display information to a user of the device 1700. In some embodiments, the user interface 1705 can be a user interface for communicating with a position determiner, as described herein.
[0228] In an embodiment, the device 1700 includes an input / output port 1709. The input / output port 1709 in some embodiments includes a transceiver. The transceiver in such an embodiment may be coupled to the processor 1707 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver or transmitter and / or receiver means may in some embodiments be configured to communicate with other electronic devices or devices via a wire or wired coupling.
[0229] The transceiver may communicate with the further device by any suitable known communication protocol, for example, in some embodiments, the transceiver may use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol such as IEEE 802.X, a suitable short-range radio frequency communication protocol such as Bluetooth, or an Infrared Data Path (IRDA).
[0230] The transceiver input / output port 1709 may be configured to receive signals and, in some embodiments, to determine parameters as described herein by using a processor 1707 executing appropriate code.
[0231] In some embodiments, device 1700 may be employed as at least a portion of a synthesis device. Input / output port 1709 may be coupled to any suitable audio output, such as a multi-channel speaker system and / or headphones (which may be head-tracked or non-tracked headphones) or the like.
[0232] In general, various embodiments of the invention may be realized in hardware or special purpose circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, but the invention is not limited thereto, in firmware or software that may be executed by a controller, microprocessor, or other computing device. Although various aspects of the invention may be illustrated and described as block diagrams, flow diagrams, or some other pictorial representations, it is well understood that these blocks, devices, systems, techniques, or methods described herein may be implemented in, by way of non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller, or other computing device, or combinations thereof.
[0233] The embodiments of the present invention can be realized by computer software executable by a data processor of a mobile device, such as in a processor entity, or by computer software executable by hardware, or by a combination of software and hardware. Furthermore, it should be noted that any block of the logic flow as illustrated may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on physical media such as memory chips, or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media, e.g. DVDs and data variants thereof.
[0234] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed and removable memories, etc. The data processor may be of any type suitable for the local technology environment and may include, by way of non-limiting examples, one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors, application specific integrated circuits (ASICs), gate level circuits, and processors based on multi-core processor architectures.
[0235] Embodiments of the present invention can be implemented in a variety of components, such as integrated circuit modules. The design of integrated circuits is a highly automated process and is large scale. Complex and powerful software tools are available to convert logic level designs into semiconductor circuit designs ready to be etched and formed on semiconductor substrates.
[0236] Programs such as those available from Synopsys, Inc. of Mountain View, Calif. and Cadence Design, Inc. of San Jose, Calif., use well-established rules of design and libraries of pre-stored design modules to automatically route conductors and identify the locations of components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design can be transmitted in a standardized electronic format (e.g., Opus, GDSII, etc.) to a semiconductor manufacturing facility or "fab" for production.
[0237] The above description provides a complete and informative description of exemplary embodiments of the present invention by way of illustrative and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art in view of the foregoing description upon perusal of the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of the present invention as defined in the appended claims.
Claims
[Claim 1] An apparatus, comprising: at least one processor; When executed by the at least one processor, the apparatus acquiring at least two audio signals; determining the types of the at least two audio signals; processing the at least two audio signals to generate at least one prototype audio signal based at least in part on the determined types of the at least two audio signals; at least one memory storing instructions for performing at least the An apparatus comprising: