Apparatus and method for encoding a combined input format spatial audio signal
The apparatus and method for encoding combined input format spatial audio signals improve packet loss resilience by selectively using non-differential encoding for frequently changing audio objects, addressing the quality degradation issues in existing codecs.
Patent Information
- Application Number
- PCT/EP2024/080078
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-10-24
- Publication Date
- 2025-06-12
AI Technical Summary
Existing immersive audio codecs face challenges in encoding combined input format spatial audio signals, particularly in maintaining quality under packet loss conditions, where differential encoding can lead to significant degradation in audio object representation.
The proposed apparatus and method involve selecting one audio object from a combined input format spatial audio signal for separate encoding, comparing it to previous frames, and encoding parameters based on these comparisons. This includes using non-differential or absolute encoding when the selected object changes frequently, to improve packet loss resilience.
This approach enhances packet loss resilience and maintains the quality of audio object representation by using adaptive encoding methods that prioritize non-differential encoding during object changes, thereby reducing errors and artifacts.
Smart Images

Figure EP2024080078_12062025_PF_FP_ABST
Abstract
Description
[0001] APPARATUS AND METHOD FOR ENCODING A COMBINED INPUT FORMAT SPATIAL AUDIO SIGNAL
[0002] Field
[0003] The present application relates to apparatus and methods for frame erasure recovery, but not exclusively for frame erasure recovery for combined input format audio signals.
[0004] Background
[0005] Parametric spatial audio processing is a field of audio signal processing where the spatial aspect of the sound is described using a set of parameters. For example, in parametric spatial audio capture from microphone arrays, it is a typical and an effective choice to estimate from the microphone array signals a set of parameters such as directions of the sound in frequency bands, and the ratios between the directional and non-directional parts of the captured sound in frequency bands. These parameters are known to well describe the perceptual spatial properties of the captured sound at the position of the microphone array. These parameters can be utilized in synthesis of the spatial sound accordingly, for headphones binaurally, for loudspeakers, or to other formats, such as Ambisonics.
[0006] The directions and direct-to-total energy ratios in frequency bands are thus a parameterization that is particularly effective for spatial audio capture.
[0007] A parameter set consisting of a direction parameter in frequency bands and an energy ratio parameter in frequency bands (indicating the directionality of the sound) can be also utilized as the spatial metadata (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance etc) for an audio codec. For example, these parameters can be estimated from microphone-array captured audio signals, and for example a stereo or mono signal can be generated from the microphone array signals to be conveyed with the spatial metadata. The stereo signal could be encoded, for example, with an AAC encoder and the mono signal could be encoded with an EVS encoder. A decoder can decode the audio signals into PCM signals and process the sound in frequency bands (using the spatial metadata) to obtain the spatial output, for example a binaural output.
[0008] Immersive audio codecs are being implemented supporting a multitude of operating points ranging from a low bit rate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec which is being designed to be suitable for use over a communications network such as a 3GPP 4G / 5G network including use in such immersive services as for example immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding and rendering of speech, music and generic audio. It is furthermore expected to support channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services as well as support high error robustness under various transmission conditions.
[0009] The aforementioned immersive audio codecs are particularly suitable for encoding captured spatial sound from microphone arrays (e.g., in mobile phones, VR cameras, stand-alone microphone arrays). However, such an encoder can have other input types, for example, loudspeaker signals, audio object signals, Ambisonic signals.
[0010] Summary
[0011] According to a first aspect there an apparatus for encoding a combined input format spatial audio signal, wherein the apparatus comprises means configured to: obtain a first input spatial audio signal, the first input spatial audio signal comprises one or more audio objects; obtain a second input spatial audio signal; select one audio object from the one or more audio objects for a current frame, to be encoded separately from other objects of the first input format spatial audio signal; compare the selected current frame one audio object to at least one selected previous frame one audio object; and encode at least one parameter associated with the selected one audio object for the current frame based on the comparing.
[0012] The means configured to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing may be configured to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object.
[0013] The means configured to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing may be configured to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from at least one of the selected previous frame audio objects for a defined threshold number of frames.
[0014] The means configured to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing may be configured to encode either the at least one parameter or a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object when the selected current frame one audio object is the same as the selected previous frames one audio object for a defined threshold number of frames.
[0015] The means configured to encode either the at least one parameter or the difference may be based on a coding efficiency.
[0016] The means configured to encode the at least one parameter associated with the selected current frame one audio object may be configured to implement a nondifferential encoding or absolute encoding with respect to the at least one parameter associated with the selected current frame. The means configured to encode a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object may configured to implement a differential encoding with respect to the difference.
[0017] The at least one parameter associated with the selected current frame one audio object may be at least one of: an azimuth parameter associated with the identified one audio object for the current frame; an elevation parameter associated with the identified one audio object for the current frame; a direction parameter associated with the identified one audio object for the current frame; a distance parameter associated with the identified one audio object for the current frame; and a location parameter associated with the identified one audio object for the current frame.
[0018] The means may be further configured to: determine at least one audio object audio signal, the at least one object audio signal is associated with the selected current frame one audio object; and encode the determined at least one audio object audio signal.
[0019] The means may be further configured to: determine at least one remaining audio signal based on a removal of the at least one object audio signal associated with the selected current frame one audio object from the first input format spatial audio signal; and encode a combined audio signal based on a combination of the at least one remaining audio signal and the second input format spatial audio signal.
[0020] The second input format spatial audio signal may be a metadata assisted spatial audio (MASA) format.
[0021] According to a second aspect there is provided a method for encoding a combined input format spatial audio signal, wherein the method comprises: obtaining a first input spatial audio signal, the first input spatial audio signal comprises one or more audio objects; obtaining a second input spatial audio signal; selecting one audio object from the one or more audio objects for a current frame, to be encoded separately from other objects of the first input format spatial audio signal; comparing the selected current frame one audio object to at least one selected previous frame one audio object; and encoding at least one parameter associated with the selected one audio object for the current frame based on the comparing.
[0022] Encoding the at least one parameter associated with the selected one audio object for the current frame based on the comparing may comprise encoding the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object.
[0023] Encoding the at least one parameter associated with the selected one audio object for the current frame based on the comparing may comprise encoding the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from at least one of the selected previous frame audio objects for a defined threshold number of frames.
[0024] Encoding the at least one parameter associated with the selected one audio object for the current frame based on the comparing may comprise encoding either the at least one parameter or a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object when the selected current frame one audio object is the same as the selected previous frames one audio object for a defined threshold number of frames.
[0025] Encoding either the at least one parameter or the difference may comprise encoding either the at least one parameter or the difference based on a coding efficiency.
[0026] Encoding the at least one parameter associated with the selected current frame one audio object may comprise implementing a non-differential encoding or absolute encoding with respect to the at least one parameter associated with the selected current frame. Encoding a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object may comprise implementing a differential encoding with respect to the difference.
[0027] The at least one parameter associated with the selected current frame one audio object may be at least one of: an azimuth parameter associated with the identified one audio object for the current frame; an elevation parameter associated with the identified one audio object for the current frame; a direction parameter associated with the identified one audio object for the current frame; a distance parameter associated with the identified one audio object for the current frame; and a location parameter associated with the identified one audio object for the current frame.
[0028] The method may further comprise: determining at least one audio object audio signal, the at least one object audio signal is associated with the selected current frame one audio object; and encoding the determined at least one audio object audio signal.
[0029] The method may further comprise: determining at least one remaining audio signal based on a removal of the at least one object audio signal associated with the selected current frame one audio object from the first input format spatial audio signal; and encoding a combined audio signal based on a combination of the at least one remaining audio signal and the second input format spatial audio signal.
[0030] The second input format spatial audio signal may be a metadata assisted spatial audio (MASA) format.
[0031] There is according to a third aspect an apparatus for encoding a combined input format spatial audio signal, wherein the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to: obtain a first input spatial audio signal, the first input spatial audio signal comprises one or more audio objects; obtain a second input spatial audio signal; select one audio object from the one or more audio objects for a current frame, to be encoded separately from other objects of the first input format spatial audio signal; compare the selected current frame one audio object to at least one selected previous frame one audio object; and encode at least one parameter associated with the selected one audio object for the current frame based on the comparing.
[0032] The apparatus caused to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing may be caused to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object.
[0033] The apparatus caused to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing may be caused to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from at least one of the selected previous frame audio objects for a defined threshold number of frames.
[0034] The apparatus caused to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing may be caused to encode either the at least one parameter or a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object when the selected current frame one audio object is the same as the selected previous frames one audio object for a defined threshold number of frames.
[0035] The apparatus caused to encode either the at least one parameter or the difference may be caused to encode either the at least one parameter or the difference based on a coding efficiency. The apparatus caused to encode the at least one parameter associated with the selected current frame one audio object may be caused to implement a nondifferential encoding or absolute encoding with respect to the at least one parameter associated with the selected current frame.
[0036] The apparatus caused to encode a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object may be caused to implement a differential encoding with respect to the difference.
[0037] The at least one parameter associated with the selected current frame one audio object may be at least one of: an azimuth parameter associated with the identified one audio object for the current frame; an elevation parameter associated with the identified one audio object for the current frame; a direction parameter associated with the identified one audio object for the current frame; a distance parameter associated with the identified one audio object for the current frame; and a location parameter associated with the identified one audio object for the current frame.
[0038] The apparatus may be further caused to: determine at least one audio object audio signal, the at least one object audio signal is associated with the selected current frame one audio object; and encode the determined at least one audio object audio signal.
[0039] The apparatus may be further caused to: determine at least one remaining audio signal based on a removal of the at least one object audio signal associated with the selected current frame one audio object from the first input format spatial audio signal; and encode a combined audio signal based on a combination of the at least one remaining audio signal and the second input format spatial audio signal.
[0040] The second input format spatial audio signal may be a metadata assisted spatial audio (MASA) format. An apparatus configured to perform the actions of the method as described above.
[0041] A computer program comprising program instructions for causing a computer to perform the method as described above.
[0042] A computer program product stored on a medium may cause an apparatus to perform the method as described herein.
[0043] An apparatus comprising means configured to perform the actions of the method as described above.
[0044] An electronic device may comprise apparatus as described herein.
[0045] A chipset may comprise apparatus as described herein.
[0046] Embodiments of the present application aim to address problems associated with the state of the art.
[0047] Summary of the Figures
[0048] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings in which:
[0049] Figure 1 shows schematically a system or apparatus suitable for implementing some embodiments;
[0050] Figure 2 shows schematically an example encoding mode selector as shown in the system of apparatus as shown in Figure 1 according to some embodiments;
[0051] Figure 3 shows a flow diagram of the operation of the example encoding mode selector shown in Figure 2 according to some embodiments; Figure 4 shows a flow diagram of the operation of the example first, lowest, or only MASA bitrate encoding mode shown in Figure 3 according to some embodiments;
[0052] Figure 5 shows a flow diagram of the operation of the example second, lower, or object information encoding mode shown in Figure 3 according to some embodiments;
[0053] Figure 6 shows a flow diagram of the operation of the example third, higher, or single object encoding mode shown in Figure 3 according to some embodiments;
[0054] Figure 7 shows a flow diagram of the operation of the example fourth, highest, or independent object and multi-input encoding mode shown in Figure 3 according to some embodiments;
[0055] Figure 8 shows schematically an example audio encoder as shown in Figure 1 with respect to the second encoding mode according to some embodiments;
[0056] Figure 9 shows a flow diagram of the operation of the selection of one audio object and encoding the selected audio object (audio signal and metadata) shown in Figure 5 according to some embodiments;
[0057] Figures 10a and 10b shows example graphs demonstrating reduced errors when applying example embodiments; and
[0058] Figure 11 shows an example device suitable for implementing the apparatus shown in previous figures.
[0059] Embodiments of the Application
[0060] The following describes in further detail suitable apparatus and possible mechanisms for the encoding of parametric spatial audio signals comprising transport audio signals and spatial metadata. As indicated above immersive audio codecs (such as 3GPP IVAS) support a multitude of operating points ranging from a low bit rate operation to transparency. It supports channel-based audio and scene-based audio inputs including spatial information about the sound field and sound sources. In the following the example codec is configured to be able to receive multiple input formats. In particular the codec is configured to obtain or receive a multi audio signal (for example received from a microphone array, or as a multi-channel audio format input, an ambisonics format input) and one or more audio object signal (these can also be called an independent stream with metadata - ISM format). Furthermore, in some situations the codec is configured to handle more than one input format at a time. This combined (input) format mode can, for example, enable simultaneous encoding of two different audio input formats. An example of two different audio input formats which IVAS currently supports is the combination of the MASA format with audio object format. Metadata-Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.
[0061] MASA format can be considered as an audio representation consisting of ‘N channels + spatial metadata’. It is a scene-based audio format particularly suited for spatial audio capture on practical devices, such as smartphones. The idea is to describe the sound scene in terms of time- and frequency-varying sound source directions and, e.g., energy ratios. Sound energy that is not defined (described) by the directions, is described as diffuse (coming from all directions).
[0062] As discussed above spatial metadata associated with the audio signals may comprise multiple parameters (such as multiple directions and associated with each direction (or directional value) a direct-to-total energy ratio, spread coherence, distance, etc.) per time-frequency tile. The spatial metadata may also comprise other parameters or may be associated with other parameters which are considered to be non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio) but when combined with the directional parameters are able to be used to define the characteristics of the audio scene. For example, a reasonable design choice which is able to produce a good quality output is one where the spatial metadata comprises one or more directions for each time-frequency subframe (and associated with each direction direct-to- total ratios, spread coherence, distance values etc) are determined.
[0063] As described above, parametric spatial metadata representation can use multiple concurrent spatial directions. With MASA, the proposed maximum number of concurrent directions is two. For each concurrent direction, there may be associated parameters such as: Direction index; Direct-to-total ratio; Spread coherence; and Distance. In some embodiments other parameters such as Diffuse- to-total energy ratio; Surround coherence; and Remainder-to-total energy ratio are defined.
[0064] Furthermore, the IVAS coding model also provides for encoding of a combined format at various bitrates. This coding model enables combined format encoding with one instance of the codec instead of two. This enables more efficient encoding of the combined format in terms of better quality of decoded data thanks to the adaptive bit allocation between the formats. Additionally this combined format encoding offers the possibility to employ the combined format encoding at very low bit rates.
[0065] As will be seen below when the IVAS codec provides for the simultaneous encoding of audio object format and MASA format, this pairing of formats can be encoded using one of four separate modes depending on the available bit rate and the number of objects.
[0066] In two of the current modes of operation, one of the objects in the combined format encoding is separated and encoded separately from the other objects. However, for one of the two modes, the one separated object in MASA representation coding mode, this can lead to situations where under packet loss conditions the quality of audio object representation is drastically reduced. The degradation in quality can, for example, be generated from decoding the separated object direction metadata which can, after suffering frame loss, experience sudden shifts in the separated object direction following a change of the direction of the separated object at the transmitter and encoder. For example, the object direction may suddenly change significantly when the separated object is changed. If a frame loss happens during this time, the encoding and decoding of the object direction may not be able to follow the sudden changes in the object direction. As a result, the separated object may be reproduced from a very wrong direction for some time. The result of the reproduction of the separated object from the wrong direction would be perceived as an artifact following the reobtaining of the correct direction of the object as the direction estimate would likely produce a sudden move of the active object direction.
[0067] The concept as discussed in the embodiments later is apparatus and methods which aim to improve the packet loss resilience. In some embodiments which can be implemented by adapting a maximum streak of differential encoding for the separated object metadata based on whether the selected separated object changes or not.
[0068] Differential encoding is a type of coding where a parameter is encoded based on a difference in parameter values. For example a temporal differential encoding is one where a difference is determined between a current time or frame parameter value and a previous time or frame parameter value (or a future time or frame parameter value), and the difference is then encoded. A frequency differential encoding is one where a difference is determined between a current sub-band parameter value and a further sub-band (for example neighbouring) sub-band parameter value, and the difference is then encoded. In some situations these difference values can be encoded based on any suitable encoding method, for example the difference values can be entropy or otherwise encoded.
[0069] A non-differential or absolute encoding is a type of coding where the parameter is encoded directly (in other words without determining a difference value). The nondifferential encoding methods can however encode these values with reference to other values, for example using an entropy encoding method to signal a symbol formed from a group of absolute values based on the probability of the symbol. Embodiments of the present application aim to address this problem associated with the state of the art.
[0070] In this regard Figure 1 depicts a general apparatus 100 and system for providing the joint coding of the MASA format and the audio object format. The system is shown with an ‘analysis’ part. The ‘analysis’ part is the part from receiving the multichannel signals up to an encoding of the metadata and downmix signal.
[0071] The input to the system ‘analysis’ part is the multi-channel audio signals 102. In the following examples a microphone channel signal input is described, however any suitable input (or synthetic multi-channel) format may be implemented in other embodiments. For example, in some embodiments the spatial analyser and the spatial analysis which can correspond to the metadata generator 103 and transport signal generator 105 as shown in Figure 1 and may be implemented external to the encoder. For example, in some embodiments the spatial (MASA) metadata associated with the audio signals may be provided to an encoder as a separate bitstream.
[0072] Additionally, Figure 1 also depicts multiple audio objects 104 (each audio object comprising an audio object signal and audio object metadata) as a further input to the analysis part. As mentioned above these multiple audio objects (or audio object stream) 104 may represent various sound sources within a physical space. Each audio object may be characterized by an audio (object) signal and accompanying metadata comprising directional data (in the form of azimuth and elevation values) which indicate the position or direction of the audio object within a physical space on an audio frame basis.
[0073] The multi-channel signals 102 are passed to an analyser and encoder 101 , and specifically a transport signal generator 105 and to a metadata generator 103.
[0074] In some embodiments the metadata generator 103 is also configured to receive the multi-channel signals and analyse the signals to produce metadata 114 associated with the multi-channel signals and thus associated with the transport signals 106. The metadata generator 103 may be configured to generate the metadata which can comprise, for each time-frequency analysis interval, a direction parameter and an energy ratio parameter and a coherence parameter (and in some embodiments a diffuseness parameter). The direction, energy ratio and coherence parameters may in some embodiments be considered to be MASA spatial audio parameters (or MASA metadata). In other words, the spatial audio parameters comprise parameters which aim to characterize the sound-field created / captured by the multi-channel signals (or two or more audio signals in general).
[0075] In some embodiments the parameters generated can differ from frequency band to frequency band. Thus, for example in band X all of the parameters are generated and transmitted, whereas in band Y only one of the parameters is generated and transmitted, and furthermore in band Z no parameters are generated or transmitted. A practical example of this can be that for some frequency bands such as the highest band some of the parameters are not required for perceptual reasons. The transport signals 106 and the metadata 114 may be passed to a combined encoder core 109.
[0076] In some embodiments the transport signal generator 105 is configured to receive the multi-channel signals and generate a suitable transport signal comprising a determined number of channels and output the transport signals 106 (MASA transport audio signals). For example, the transport signal generator 105 can be configured to generate a 2-audio channel downmix of the multi-channel signals. The determined number of channels may be any suitable number of channels. The transport signal generator in some embodiments is configured to otherwise select or combine, for example, by beamforming techniques the input audio signals to the determined number of channels and output these as transport signals.
[0077] In some embodiments the transport signal generator 105 is optional and the multichannel signals are passed unprocessed to a combined encoder core 109 in the same manner as the transport signal are in this example. The audio objects 104 can be passed to the audio object analyser 107 for processing. In some embodiments the audio object analyser 107 analyses the object audio input stream 104 in order to produce suitable audio object transport signals and audio object metadata. For example, the audio object analyser can be configured to produce the audio object transport signals 128 by downmixing the audio signals of the audio objects into a stereo channel together using amplitude panning based on the associated audio object directions. Additionally, the audio object analyser can also be configured to produce the audio object metadata associated with the audio object input stream 104. The audio object metadata may comprise direction values which are applicable for all sub-bands. So, if there are 4 objects, there are 4 directions. In the examples described herein the direction values also apply across all of the subframes of the frame, but in some embodiments the temporal resolution of the direction values can differ and the directions values apply for one or more than one sub-frames of the frame. Furthermore, energy ratios (or ISM ratios) may be determined for each object. The energy ratio (ISM ratio) defines the contribution of the object within the object part of the total audio environment. In the following examples the energy ratios (or ISM ratios) are for each time-frequency tile for each object.
[0078] The audio object analyser 107 can be additionally configured to determine an audio object for separation from the stream of audio objects 104 and create the associated separated object audio signal 152. In embodiments this associated separated object audio signal 152 can be a mono audio signal. The separated audio object audio signal 152 may then be encoded by an audio object (metadata and audio signal) encoder 111 (which can be implemented a part of an EVS encoder) to form the encoded separated object audio signal 154. This encoded separated object audio signal 154 can then be forwarded to the Bitstream generator 113 to form part of the bitstream 118. In embodiments the audio object (metadata and audio signal) encoder 111 may comprise a mono channel IVAS encoder. Additionally, an identifier index of the separated audio object can form part of the bitstream 118. In some embodiments, the audio object analyser 107 and audio object (metadata and audio signal) encoder 111 may be sited elsewhere and the audio objects 104 input to the analyser and encoder 111 is audio object transport signals and audio object metadata.
[0079] The analyser and encoder 101 can comprise a combined encoder core 109 which is configured to receive the transport audio (for example downmix) signals 106 and (remainder) audio object transport signals 128 in order to generate a suitable encoding of these audio signals.
[0080] The analyser and encoder 101 can also comprise an audio object (metadata and audio signal) encoder 111 which is similarly configured to receive the audio object metadata 108 and output an encoded or compressed form of the input information as encoded audio object metadata 112.
[0081] In some embodiments the combined encoder core 109 can be configured to implement a stream separation metadata determiner and encoder which can be configured to determine the relative contributory proportions of the multi-channel signals 102 (MASA audio signals) and audio objects 104 to the overall audio scene. This measure of proportionality produced by the stream separation metadata determiner and encoder can be used to determine the proportion of quantizing and encoding “effort” expended for the input multi-channel signals 102 and the audio objects 104. In other words, the stream separation metadata determiner and encoder may produce a metric which quantifies the proportion of the encoding effort expended on the multichannel audio signals 102 compared to the encoding effort expended on the audio objects 104. This metric can be used to drive the encoding of the audio object metadata 108 and the metadata 114. Furthermore, the metric as determined by the separation metadata determiner and encoder can also be used as an influencing factor in the process of encoding the transport audio signals 106 and audio object transport audio signal 128 performed by the combined encoder core 109. The output metric from the stream separation metadata determiner and encoder can furthermore be represented as encoded stream separation metadata and be combined into the encoded metadata stream from the combined encoder core 109.
[0082] In some embodiments the analyser and encoder 101 comprises a bitstream generator 113 configured to obtain the encoded metadata 116, the encoded transport audio signals 138 and the encoded audio object metadata 112 and generate the bitstream 118 for potential transmission or storage.
[0083] In some embodiments the analyser and encoder 101 comprises an encoder controller 115. The encoder controller 115 can in some embodiments control the encoding implemented by the audio object (metadata and audio signal) encoder 111 and the combined encoder core 109. In some embodiments encoder controller 115 is configured to determine the bitrate for the bitstream 1 18 and based on the bitrate control the encoding. In some embodiments the encoder controller 115 is further configured to control at least one of the audio object analyser 107, transport signal generator 105 and metadata generator in generating parameters.
[0084] The analyser and encoder 101 can in some embodiments be a computer or mobile device (running suitable software stored on memory and on at least one processor), or alternatively a specific device utilizing, for example, FPGAs or ASICs. The encoding may be implemented using any suitable scheme. In some embodiments the encoder 107 may further interleave, multiplex to a single data stream or embed the encoded MASA metadata, audio object metadata and stream separation metadata within the encoded (downmixed) transport audio signals before transmission or storage shown in Figure 1 by the dashed line. The multiplexing may be implemented using any suitable scheme.
[0085] Furthermore, with respect to Figure 1 is shown an associated decoder and Tenderer 129 which is configured to obtain the bitstream 118 comprising encoded metadata 116, Encoded transport audio signals 138 and encoded audio object metadata 112 and from these generate suitable spatial audio output signals. With respect to Figure 2 is shown in further detail the encoder controller 115 according to some embodiments.
[0086] In this example the encoder controller 115 comprises a bitrate determ iner / monitor 201 configured to determine and / or monitor the available bitrate for the bandwidth for the encoded audio and metadata. This could be determined based on a transmission path bandwidth estimation (and for example be based on an estimated signal strength) or a bandwidth storage determination to maintain the file for a determined time to be below a required size or by any suitable manner.
[0087] The bitrate determ iner / monitor 201 can furthermore be configured to control an encoding mode selector 203. The encoder controller 115 can comprise an encoding mode selector 203 configured to select an encoding mode, for example based on the determined bandwidth, bitrate and / or number of audio objects since the overall bit rate is distributed between the encoding of the MASA metadata & transport signals and the audio objects. Once the encoding mode has been determined the encoder controller 115 then controls the encoders, for example the combined encoder core 109 and audio object metadata encoder 111 , according to the encoding mode. Bandwidth and bitrate information can also be passed onto the audio object analyser 107 so that this information can be used [by the audio object analyser 107] to determine whether audio objects are to be separated in accordance with the second and third modes of encoding. As will be seen later these modes are also referred to as mode B and mode C respectively.
[0088] With respect to Figure 3 is shown a flow diagram of an example operation of the encoder controller shown in Figure 2 In this example there is an initial operation of receiving or obtaining or otherwise determining the bitrate or bandwidth for encoded parameters and audio data as shown in Figure 3 by step 301 .
[0089] Having obtained the available bandwidth or bitrate then a check can be made to determine whether the bitrate is below a first (or lowest) threshold limit as shown in Figure 3 by step 303. Where the available bandwidth or bitrate is below the first (or lowest) threshold limit then the encoders can be controlled to encode the transport channels and MASA metadata only (also shown as Mode A) as shown in Figure 3 by step 304.
[0090] Where the available bandwidth or bitrate is above the first (or lowest) threshold limit then a further check can be made to determine whether the bitrate is below a second (or lower or one audio object) threshold limit as shown in Figure 3 by step 305.
[0091] Where the available bandwidth or bitrate is below the second (or lower or one object) threshold limit then the encoders can be controlled to encode the combined input of MASA and audio objects in a mode B. To that end Figure 8 depicts an encoder suitable for mode B where it can be that the encoder has some functional blocks which are the same as found in Figure 1 . The encoder 1101 is arranged to receive the multichannel audio signals 102 and produce the respective transport audio signals 1106 and multichannel MASA metadata 1104. This is performed using the MASA meta generator 103 and transport signal generator 105 whose functionality is the same as the functional blocks 103 and 105 Figure 1.
[0092] The encoder 1101 is also depicted as receiving the stream of audio objects 104 via the audio object separator & analyser 1107, which can be arranged for each frame to decide which audio object to separate from the audio object stream 104 in order to create a separated audio object 1114 comprising audio object metadata and an audio object signal. The separated audio object 114 may be selected based on the relative level of the audio objects, such as the loudest audio object. This is explained in detail in WO2022 / 214730. The audio object separator & analyser 1107 is also arranged to produce a suitable remaining audio objects transport audio signal 1126. This may be produced by downmixing the audio signals of the remaining audio objects as explained before using amplitude panning based on the associated audio object directions. Additionally, the audio object separator & analyser 1107 may also be configured to produce the audio object metadata associated with the remaining audio objects. In mode B this can take the form of generating an audio object based MASA stream 1108 from the remaining audio objects by the methods presented in WO2019086757.
[0093] The separated audio object 1114 comprising the separated audio object audio signal and separated audio object direction) may be encoded by an audio object encoder 1100 to form the encoded separated audio object 1117 which is then passed to the bitstream generator 1113 to form part of the bitstream 118. In addition, an identifier describing which audio object was separated may also be sent.
[0094] In Figure 8, the combined encoder core 109 as described previously is shown as receiving the remaining audio objects MASA metadata 1108, the remaining audio objects transport audio signals 1126, the transport audio signals 1106 and the multichannel MASA metadata 1104. The combined encoder core 109 can be arranged to combine the multichannel transport audio signals 1106 and the audio signals / channels from the remaining audio objects transport audio signals 1126 into a single combined stream which is encoded to produce the encoded transport audio signals stream 1138. Similarly, the combined encoder core 109 can be arranged to combine the multichannel MASA metadata 1104 with the remaining audio objects MASA metadata 1108 to produce the encoded MASA metadata stream 1116.
[0095] In Figure 8, the encoder controller 115 is shown as controlling the encoding implemented by the combined encoder core 109. In some embodiments the controller 115 also controls the audio object encoder 1110 and audio object signal separator & analyser 1107.
[0096] Therefore, in mode B the encoders can be controlled to encode in a combined manner the multichannel transport channels 1106, multichannel MASA metadata 1104, remaining audio objects MASA metadata 1108 remaining audio objects transport audio signals 1126, one audio object data (separated audio object metadata and separated audio object signal) and some in embodiments the one audio object identifier as shown Figure 3 by step 306. Where the available bandwidth or bitrate is above the second (or lower or one object) threshold limit then a further check can be made to determine whether the bitrate is below a third, higher or full object threshold limit as shown in Figure 3 by step 307.
[0097] Where the available bandwidth or bitrate is below the third, higher or full object threshold limit then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), MASA to total ratios, ISM ratios, and 1 object audio data, with 1 object identifier (also shown as Mode C) as shown in Figure 3 by step 308.
[0098] Where the available bandwidth or bitrate is above the third, higher or full object threshold limit then the encoders can be controlled to encode transport channels, MASA metadata, ISM metadata (all objects), All objects audio data (also shown as Mode D) as shown in Figure 3 by step 310.
[0099] The encoding modes and potential mode selection based on the available bitrates and total number of objects can for example be summarised by the following tables:
[0100] The bitrates shown herein are examples and it would be understood that they can be other specific values. Additionally the bitrate and number of object values based mode allocation shown here are examples only and the modes can be implemented with other suitable allocations.
[0101] With respect to Figures 4 to 7 is shown flow diagrams showing a first (or lowest or combined) encoding mode as shown in Figure 3 by step 304, a second (or lower or object metadata) encoding mode as shown in Figure 3 by step 306, a third (or higher or one object) encoding mode as shown in Figure 3 by step 308 and a fourth (or highest or all objects) encoding mode as shown in Figure 3 by step 310 respectively.
[0102] For example, Figure 4 the mode A encoding method, shows the first (or lowest or combined) encoding mode as shown in Figure 3 by step 304 in further detail. Thus, for very low total bitrates (for example less or equal to 32kbps) all encoding is implemented using a MASA representation.
[0103] Thus, for example there is an operation of receiving / obtaining the object based streams (independent streams with metadata) and multichannel based (MASA stream) transport audio signals and metadata as shown in Figure 4 by step 401 .
[0104] Then, as shown in Figure 4 by step 403, there is an operation of generating an object based MASA stream from the object streams (independent streams with metadata). This object based MASA stream can in some embodiments be created from the object stream using, for example, the methods presented in WO2019086757A1.
[0105] After this, as shown in Figure 4 by step 405, the object based MASA stream and multichannel based MASA stream are combined. In some embodiments the original MASA stream and the MASA stream created from the objects can be combined using the method presented in GB2574238. The decoder received the information about the audio objects and the MASA audio content in the MASA format. Then the combined stream is output as shown in Figure 4 by step 407. In such embodiments the object audio content (together with the MASA audio content) is present in the decoded audio scene, but the objects cannot be edited nor separated from the scene at the decoder based on the decoded parameters.
[0106] Figure 5 shows the mode B encoding method, the second (or lower or object metadata) encoding mode as shown in Figure 3 by step 306. Additionally Figure 8 shows an example encoder when operating in mode B in further detail. Thus, for low bit-rates (for example between 48kbps and 80kbps) and since there are a more bits available, there is a possibility to send the audio signal and metadata of one audio object independently. Thus, there is shown in Figure 5 501 which portrays the processing step of selecting an audio object for separation and encoding said audio object.
[0107] In addition, a downmix formed from the transport audio signals / channels 1106 and the remaining audio object signals (from the remaining audio objects 1118) giving the encoded combined transport audio signals 1138 are sent together with a encoded combined multichannel based MASA and audio object based MASA metadata stream 1116.
[0108] Thus, for example there is a method step of receiving / obtaining the remaining audio objects stream 1126, the (multichannel) transport audio signals 1106, remaining audio objects MASA metadata 1108 and the multichannel MASA metadata 1104 as shown in Figure 5 by step 503.
[0109] Then, as shown by step 505 in Figure 5, generate a combined multichannel transport and audio object signals downmixed (channel pair element) audio signals. In other words, the audio content of MASA 1106 and the audio signal channels of the remaining audio objects 1118 are downmixed to 2 channels (channel pair element CPE).
[0110] The remaining audio objects MASA metadata 1124 and the multichannel MASA metadata 1104 may be combined into a single combined MASA metadata stream by any suitable method. For example, such as the method presented in GB2574238. This step is shown in Figure 5 as step 507.
[0111] Furthermore, the combined MASA metadata can then be encoded based on any suitable MASA metadata encoding method as shown in Figure 5 by step 509.
[0112] The combined multichannel and audio object signals can then be encoded based on any suitable audio signal encoding method as shown in Figure 5 by step 511.
[0113] The encoder can then output encoded combined MASA metadata 1116, encoded combined transport audio signals 1138 and encoded separated audio object 1117 as shown in Figure 5 by step 513.
[0114] Thus, in summary, there are currently 4 coding modes in OMASA: (A) MASA representation, (B) One separated object and MASA representation, (C) One separated object and parametric representation, and (D) separate MASA and objects.
[0115] In modes B and C, the object to separate is decided based on which object is the most important.
[0116] In mode C, the object metadata for all objects is encoded, but in mode B, only the object metadata of the separated object is transmitted.
[0117] The object importance changes over time and, as a consequence, the separated object changes. This means that there may be significant changes in the object trajectory, and it is not as smooth as it would be expected from one individual same object.
[0118] The embodiments as discussed herein address the encoding mode B and describes a method for enforcing a non-differential encoding mode when the object change is detected. In some embodiments therefore a non-differential (or an absolute) coding mode for the direction angles is enforced by the encoder controller 115 and the audio object encoder 1110 when:
[0119] 1 . When the separated object changes and
[0120] 2. When the frames count since the last object change is less than a threshold.
[0121] In other words, one object is selected for a current frame to be encoded separately from other objects. This selected object is compared to at least one previous frame selected audio object. The comparison can be between the current frame selected object and one or more of the proceeding or previous frame selected objects. This can be a comparison not only between the immediately previous frame and the current frame but between the current frame and one or more proceeding frames.
[0122] This comparison can be any suitable comparison. For example the comparison can identify whether the object selected for the current frame has a different index or reference or identifier to a previous selected object. In some embodiments the comparison is configured to determine whether one or more of the parameters associated with the object changes more than an expected or threshold value. For example, in some embodiments whether the object direction moves more than would be expected.
[0123] Furthermore the encoding of the at least one parameter associated with the selected one audio object for the current frame is implemented based on the comparison.
[0124] For example, in some embodiments, the encoding operation is of encoding the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object. This encoding can be a non-differential or absolute encoding of the parameter. Additionally, in some embodiments, the encoding operation is configured to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object is different from any of the selected previous frame audio objects for a defined threshold number of frames. In other words that a change in the audio object occurring within a defined threshold number of frames prior to this current frame prevents the encoder from differentially encoding the parameter but encodes it using non-differential or absolute encoding.
[0125] These embodiments furthermore encode the at least one parameter associated with the selected current frame one audio object (in a non-differential or absolute encoding) when the selected current frame one audio object is different from one of the selected previous frame audio objects within a defined guard number of frames. In other words that a change in the audio object occurring within the guard period (frames) prior to this current frame prevents the encoder from differentially encoding the parameter but encodes it using non-differential or absolute encoding.
[0126] The embodiments furthermore can be configured to encode a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object when the selected current frame one audio object is the same as the selected previous frames one audio object for a defined threshold number of frames. In other words, differential encoding can be selected based on the same object being selected for a number of frames.
[0127] In some embodiments the selection of differential encoding is implemented based not only on the same audio object has been selected for a number of frames but furthermore based on a characterization of the audio object. For example, a non- differential encoding is selected unless the same object has been selected for a number of frames and a change in the object direction is less than a defined value. In some embodiments the selection of differential encoding is stopped when same audio object has been selected for a defined (maximum) number of frames. For example a reference frame is selected and a non-differential encoding generated after a defined number of frames have been differentially encoded. Then succeeding frames can be encoded using non-differential or differential methods based on whether the same object as the reference frame selected object is the same object based on the same rules as described herein.
[0128] An example C source code representation of the control of when to encode the object in absolute or non-differential encoding and when to encode the object using a suitable differential encoding can be as follows:
[0129] *force_absolute_coding = 0; if ( hOMasa->prev_selected_object != selected_object )
[0130] {
[0131] *force_absolute_coding = 1 ; hOMasa->since_obj_change_cnt = 0;
[0132] }else
[0133] { hOMasa->since_obj_change_cnt ++; hOMasa->since_obj_change_cnt = min( OMASA_FEC_MAX, hOMasa- >since_obj_change_cnt ); if ( hOMasa->since_obj_change_cnt < OMASA_FEC_MAX )
[0134] {
[0135] *force_absolute_coding = 1 ;
[0136] }
[0137] }
[0138] The initial operation is one of initializing the coding mode.
[0139] *force_absolute_coding = 0
[0140] The code method example above thus initially determines if (whether) the selected object is not equal to the previous selected object. if ( hOMasa->prev_selected_object != selected_object ) If the selected object is not equal to the previous selected object then the absolute coding mode is forced on (*force_absolute_coding = 1 ) and a counter, which in this example is labelled since_obj_change_cnt is reset (hOMasa- >since_obj_change_cnt = 0).
[0141] Otherwise if the selected object is equal to the previous selected object then the counter is incremented (hOMasa->since_obj_change_cnt ++) and furthermore in this example setting the counter to the lower of the current count value and a threshold value. hOMasa->since_obj_change_cnt = min(OMASA_FEC_MAX, hOMasa->since_obj_change_cnt)
[0142] The OMASA_FEC_MAX value is an example threshold value. It can be any suitable value, for example 20.
[0143] The method can then check whether the current object selected is one which has been selected for more than the minimum or threshold value. if ( hOMasa->since_obj_change_cnt < OMASA_FEC_MAX )
[0144] If the selected object is one which has been selected for more than the minimum or threshold value then the non-differential or absolute coding mode is forced on (*force_absolute_coding = 1 ).
[0145] The above logic can in some embodiments be applied to the encoding of both the azimuth and the elevation encoding or for one or other of the azimuth and the elevation encoding.
[0146] Additionally in some embodiments the threshold can differ for the azimuth and the elevation coding.
[0147] With respect to Figure 9 is shown an example flow diagram of the example method as shown above and represents the encoding 501 step as shown in Figure 5 in further detail. As shown in Figure 9 by 901 , is the operation of selecting one audio object for current frame (select the separated object for the frame).
[0148] Then there is the operation, as shown by 903 in Figure 9, of comparing the current frame selected audio object against the previous frame selected audio object and determining if the current frame selected audio object is the same as previous frame selected audio object?
[0149] Where the current frame selected audio object is not the same as the previous frame selected audio object then, as shown by 909 in Figure 9 is the operation of encoding the current frame selected audio object (metadata) using non-differential (or absolute) coding.
[0150] Where the current frame selected audio object is the same as the previous frame selected audio object then, as shown by 905 in Figure 9 is the operation of checking if the current frame selected audio object has been the same for more than a threshold number of frames?
[0151] If the current frame selected audio object has not been the same for the threshold values then, as shown by 909 in Figure 9, perform the operation of encoding the current frame selected audio object (metadata) using non-differential (or absolute) coding.
[0152] If the current frame selected audio object has been the same for more than a threshold number of frames, then as shown by 907 in Figure 9, perform the operation of encoding the current frame selected audio object (metadata) using a suitable differential coding.
[0153] In some embodiments the differential coding mechanism is configured to send the absolute values for key frames, and then differential values for some successive frames. Additionally in some embodiments the selection between differential and nondifferential (absolute) encoding can furthermore be implemented based further on the encoding efficiency. For example a differential encoding mechanism can be selected for encoding the parameter difference if the parameter difference is zero (in other words the parameter values for a first and second instance or frame are the same) and also if the parameter difference (between the current frame and previous frame) is not zero. In other words, differential encoding can be employed not only when the parameter is a constant value.
[0154] Furthermore in some embodiments the differential encoding mechanism can limit the number of consecutive differential frames, or define a maximum number of consecutive frames used for differential encoding, the number can be labelled as a maximum differential encoding frame value. These encoding efficiency based encoding mode selections are implemented independently of the object selection mechanism.
[0155] Then, as shown by 911 , is the encoding of the current frame selected audio object audio signal using any suitable audio signal encoding method.
[0156] The impact of the example embodiments as discussed herein can be seen for instance with respect to Figures 10a and 10b with respect to an example simulated encoded azimuth 1087 and decoded azimuth 1089 values. Figure 10a shows an example situation where there are three simulated packet loss situations 1095, 1097, and 1099, where the conventional encoding methods show errors (differences) between the encoded and decoded azimuth values. Figure 10b shows the same example situation where there are three simulated packet loss situations 1075, 1077, and 1079, where the encoding methods according to some embodiments show smaller errors (differences) between the encoded 1087 and decoded 1089 azimuth values.
[0157] The highlighted regions in the figures thus demonstrate the impact of the proposed improvement in recovering after frame erasures according to some embodiments. This improvement ensures recovery of object position quickly after a lost frame. Otherwise, the errors can produce errors which are very objectionable for the listener when a switching object. For example, the decoder continues in a spatial position of a previous object and later abruptly jumps to a new location.
[0158] With respect to Figure 11 an example electronic device which may be used as any of the apparatus parts of the system as described above. The device may be any suitable electronics device or apparatus. For example, in some embodiments the device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc. The device may for example be configured to implement the encoder / analyser part and / or the decoder part as shown in Figures 1 , 2 or 8 or any functional block as described above.
[0159] In some embodiments the device 1700 comprises at least one processor or central processing unit 1707. The processor 1707 can be configured to execute various program codes such as the methods such as described herein.
[0160] In some embodiments the device 1700 comprises at least one memory 1711. In some embodiments the at least one processor 1707 is coupled to the memory 1711. The memory 1711 can be any suitable storage means. In some embodiments the memory 1711 comprises a program code section for storing program codes implementable upon the processor 1707. Furthermore, in some embodiments the memory 1711 can further comprise a stored data section for storing data, for example data that has been processed or to be processed in accordance with the embodiments as described herein. The implemented program code stored within the program code section and the data stored within the stored data section can be retrieved by the processor 1707 whenever needed via the memory-processor coupling.
[0161] In some embodiments the device 1700 comprises a user interface 1705. The user interface 1705 can be coupled in some embodiments to the processor 1707. In some embodiments the processor 1707 can control the operation of the user interface 1705 and receive inputs from the user interface 1705. In some embodiments the user interface 1705 can enable a user to input commands to the device 1700, for example via a keypad. In some embodiments the user interface 1705 can enable the user to obtain information from the device 1700. For example, the user interface 1705 may comprise a display configured to display information from the device 1700 to the user. The user interface 1705 can in some embodiments comprise a touch screen or touch interface capable of both enabling information to be entered to the device 1700 and further displaying information to the user of the device 1700. In some embodiments the user interface 1705 may be the user interface for communicating.
[0162] In some embodiments the device 1700 comprises an input / output port 1709. The input / output port 1709 in some embodiments comprises a transceiver. The transceiver in such embodiments can be coupled to the processor 1707 and configured to enable a communication with other apparatus or electronic devices, for example via a wireless communications network. The transceiver or any suitable transceiver or transmitter and / or receiver means can in some embodiments be configured to communicate with other electronic devices or apparatus via a wire or wired coupling.
[0163] The transceiver can communicate with further apparatus by any suitable known communications protocol. For example in some embodiments the transceiver can use a suitable radio access architecture based on long term evolution advanced (LTE Advanced, LTE-A) or new radio (NR) (or can be referred to as 5G), universal mobile telecommunications system (UMTS) radio access network (UTRAN or E- UTRAN), long term evolution (LTE, the same as E-UTRA), 2G networks (legacy network technology), wireless local area network (WLAN or Wi-Fi), worldwide interoperability for microwave access (WiMAX), Bluetooth®, personal communications services (PCS), ZigBee®, wideband code division multiple access (WCDMA), systems using ultra-wideband (UWB) technology, sensor networks, mobile ad-hoc networks (MANETs), cellular internet of things (loT) RAN and Internet Protocol multimedia subsystems (IMS), any other suitable option and / or any combination thereof.
[0164] The transceiver input / output port 1409 may be configured to receive the signals. In some embodiments the device 1400 may be employed as at least part of the synthesis device. The input / output port 1409 may be coupled to headphones (which may be a headtracked or a non-tracked headphones) or similar and loudspeakers.
[0165] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the invention is not limited thereto. While various aspects of the invention may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0166] The embodiments of this invention may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard it should be noted that any blocks of the logic flow as in the Figures may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.
[0167] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.
[0168] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0169] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0170] As used in this application, the term “circuitry” may refer to one or more or all of the following:
[0171] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and
[0172] (b) combinations of hardware circuits and software, such as (as applicable):
[0173] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and
[0174] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0175] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0176] The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
[0177] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements
[0178] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention as defined in the appended claims.
Claims
CLAIMS:
1. An apparatus for encoding a combined input format spatial audio signal, wherein the apparatus comprises means configured to: obtain a first input spatial audio signal, the first input spatial audio signal comprises one or more audio objects; obtain a second input spatial audio signal; select one audio object from the one or more audio objects for a current frame, to be encoded separately from other objects of the first input format spatial audio signal; compare the selected current frame one audio object to at least one selected previous frame one audio object; and encode at least one parameter associated with the selected one audio object for the current frame based on the comparing.
2. The apparatus as claimed in claim 1 , wherein the means configured to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing is configured to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object.
3. The apparatus as claimed in any of claims 1 or 2, wherein the means configured to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing is configured to encode the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from at least one of the selected previous frame audio objects for a defined threshold number of frames.
4. The apparatus as claimed in any of claims 1 to 3, wherein the means configured to encode the at least one parameter associated with the selected one audio object for the current frame based on the comparing is configured to encodeeither the at least one parameter or a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object when the selected current frame one audio object is the same as the selected previous frames one audio object for a defined threshold number of frames.
5. The apparatus as claimed in claim 4, wherein the means configured to encode either the at least one parameter or the difference based on a coding efficiency.
6. The apparatus as claimed in any of claims 2 to 5, wherein the means configured to encode the at least one parameter associated with the selected current frame one audio object is configured to implement a non-differential encoding or absolute encoding with respect to the at least one parameter associated with the selected current frame.
7. The apparatus as claimed in any of claims 4 or 5, wherein the means configured to encode a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object is configured to implement a differential encoding with respect to the difference.
8. The apparatus as claimed in any of claims 1 to 7, wherein the at least one parameter associated with the selected current frame one audio object is at least one of: an azimuth parameter associated with the identified one audio object for the current frame; an elevation parameter associated with the identified one audio object for the current frame; a direction parameter associated with the identified one audio object for the current frame; a distance parameter associated with the identified one audio object for the current frame; anda location parameter associated with the identified one audio object for the current frame.
9. The apparatus as claimed in any of claims 1 to 8, wherein the means is further configured to: determine at least one audio object audio signal, the at least one object audio signal is associated with the selected current frame one audio object; and encode the determined at least one audio object audio signal.
10. The apparatus as claimed in any of claims 1 to 9, wherein the means is further configured to: determine at least one remaining audio signal based on a removal of the at least one object audio signal associated with the selected current frame one audio object from the first input format spatial audio signal; and encode a combined audio signal based on a combination of the at least one remaining audio signal and the second input format spatial audio signal.11 . The apparatus as claimed in any of claims 1 to 10, wherein the second input format spatial audio signal is a metadata assisted spatial audio (MASA) format.
12. A method for encoding a combined input format spatial audio signal, wherein the method comprises: obtaining a first input spatial audio signal, the first input spatial audio signal comprises one or more audio objects; obtaining a second input spatial audio signal; selecting one audio object from the one or more audio objects for a current frame, to be encoded separately from other objects of the first input format spatial audio signal; comparing the selected current frame one audio object to at least one selected previous frame one audio object; and encoding at least one parameter associated with the selected one audio object for the current frame based on the comparing.
13. The method as claimed in claim 12, wherein encoding the at least one parameter associated with the selected one audio object for the current frame based on the comparing comprises encoding the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from the selected immediately previous frame audio object.
14. The method as claimed in any of claims 12 or 13, wherein encoding the at least one parameter associated with the selected one audio object for the current frame based on the comparing comprises encoding the at least one parameter associated with the selected current frame one audio object when the selected current frame one audio object is different from at least one of the selected previous frame audio objects for a defined threshold number of frames.
15. The method as claimed in any of claims 12 to 14, wherein encoding the at least one parameter associated with the selected one audio object for the current frame based on the comparing comprises encoding either the at least one parameter or a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object when the selected current frame one audio object is the same as the selected previous frames one audio object for a defined threshold number of frames.
16. The method as claimed in claim 15, wherein encoding either the at least one parameter or the difference comprises encoding either the at least one parameter or the difference based on a coding efficiency.
17. The method as claimed in any of claims 13 to 16, wherein encoding the at least one parameter associated with the selected current frame one audio object is comprises implementing a non-differential encoding or absolute encoding with respect to the at least one parameter associated with the selected current frame.
18. The method as claimed in any of claims 15 or 16, wherein encoding a difference between at least one parameter associated with the selected current frame one audio object and at least one parameter associated with the selected previous frame one audio object comprises implementing a differential encoding with respect to the difference.
19. The method as claimed in any of claims 12 to 18, wherein the at least one parameter associated with the selected current frame one audio object is at least one of: an azimuth parameter associated with the identified one audio object for the current frame; an elevation parameter associated with the identified one audio object for the current frame; a direction parameter associated with the identified one audio object for the current frame; a distance parameter associated with the identified one audio object for the current frame; and a location parameter associated with the identified one audio object for the current frame.
20. The method as claimed in any of claims 12 to 19, further comprising: determining at least one audio object audio signal, the at least one object audio signal is associated with the selected current frame one audio object; and encoding the determined at least one audio object audio signal.21 . The method as claimed in any of claims 12 to 20, further comprising: determining at least one remaining audio signal based on a removal of the at least one object audio signal associated with the selected current frame one audio object from the first input format spatial audio signal; and encoding a combined audio signal based on a combination of the at least one remaining audio signal and the second input format spatial audio signal.
22. The method as claimed in any of claims 12 to 21 , wherein the second input format spatial audio signal is a metadata assisted spatial audio (MASA) format.
Citation Information
Patent Citations
Spatial audio parameter merging
GB2574238A
Determination of targeted spatial audio parameters and associated spatial audio playback
WO2019086757A1
Quantization of spatial audio direction parameters
US20220335956A1
Separating spatial audio objects
WO2022214730A1
Low coding rate parametric spatial audio encoding
WO2024199801A1