Apparatus, method and computer program for encoding spatial metadata
By obtaining the source format configuration parameters of spatial audio content and selecting appropriate compression methods and codebooks, the problem of low efficiency in spatial metadata encoding and decoding is solved, thereby improving the efficiency of spatial characteristic reconstruction in immersive audio applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-28
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to effectively compress and decode spatial metadata of spatial audio content, resulting in inefficient reconstruction of spatial characteristics in immersive audio applications.
By obtaining the source format configuration parameters of the spatial audio content, and selecting appropriate compression methods and codebooks, the encoding and decoding of spatial metadata can be achieved.
It improves the efficiency of spatial characteristic reconstruction of spatial audio content and enhances the quality of immersive audio experiences.
Smart Images

Figure CN113228169B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Examples of the present disclosure relate to apparatuses, methods, and computer programs for encoding spatial metadata. Some examples relate to apparatuses, methods, and computer programs for encoding spatial metadata associated with spatial audio content. BACKGROUND
[0002] Spatial audio content can be used for immersive audio applications such as mediated reality content applications, which can be virtual reality, augmented reality, mixed reality, extended reality, or any other suitable type of application. Spatial metadata can be associated with the spatial audio content. The spatial metadata can contain information that enables spatial characteristics of the spatial audio content to be recreated. SUMMARY
[0003] According to various, but not necessarily all, examples of the present disclosure, an apparatus can be provided comprising means for obtaining spatial metadata associated with spatial audio content; obtaining configuration parameters indicative of a source format of the spatial audio content; and using the configuration parameters to select a compression method for the spatial metadata associated with the spatial audio content.
[0004] The configuration parameters can be used to select a codebook to compress the spatial metadata associated with the spatial audio content.
[0005] The configuration parameters can be used to enable creation of a codebook for compressing the spatial metadata.
[0006] The codebook can be used to encode and decode the spatial metadata.
[0007] The source format indicated by the configuration parameters can indicate a format of the spatial audio content from which the spatial metadata was obtained.
[0008] The spatial metadata can comprise data indicative of spatial parameters of the spatial audio content.
[0009] The compression method can be selected independently of content of the obtained spatial audio content.
[0010] The means can be configured to obtain the spatial audio content.
[0011] The source configuration parameters can be obtained together with the spatial audio content.
[0012] The source configuration parameters can be obtained separately from the spatial audio content.
[0013] According to various but not all examples of the present disclosure, there can be provided an apparatus comprising processing circuitry; and memory circuitry comprising computer program code, the memory circuitry and the computer program code configured to, with the processing circuitry, cause the apparatus to obtain spatial metadata associated with a spatial audio content; obtain a configuration parameter indicative of a source format of the spatial audio content; and use the configuration parameter to select a compression method of the spatial metadata associated with the spatial audio content.
[0014] According to various but not all examples of the present disclosure, there can be provided an encoding device comprising the apparatus of any of the preceding claims and one or more transceivers configured to at least transmit the spatial metadata to a decoding device.
[0015] According to various but not all examples of the present disclosure, there can be provided a method comprising obtaining spatial metadata associated with a spatial audio content; obtaining a configuration parameter indicative of a source format of the spatial audio content; and using the configuration parameter to select a compression method of the spatial metadata associated with the spatial audio content.
[0016] The configuration parameter can be used to select a codebook to compress the spatial metadata associated with the spatial audio content.
[0017] According to various but not all examples of the present disclosure, there can be provided a computer program comprising computer program instructions which, when executed by processing circuitry, cause the following operations to be performed: obtaining spatial metadata associated with a spatial audio content; obtaining a configuration parameter indicative of a source format of the spatial audio content; and using the configuration parameter to select a compression method of the spatial metadata associated with the spatial audio content.
[0018] The configuration parameter can be used to select a codebook to compress the spatial metadata associated with the spatial audio content.
[0019] According to various but not all examples of the present disclosure, there can be provided a physical entity embodying the computer program as described above.
[0020] According to various but not all examples of the present disclosure, there can be provided an electromagnetic carrier signal carrying the computer program as described above.
[0021] According to various but not all examples of the present disclosure, there can be provided an apparatus comprising means for performing the following operations: receiving a spatial audio content; receiving spatial metadata associated with the spatial audio content; and receiving information indicative of a method used to compress the spatial metadata associated with the spatial audio content, wherein the method used to compress the spatial metadata is selected based on a source format of the spatial audio content.
[0022] The information indicating the method for compressing the spatial metadata can comprise a source configuration parameter.
[0023] The information indicating the method for compressing the spatial metadata can comprise a codebook selected using a source configuration parameter.
[0024] According to various but not all examples of the present disclosure, an apparatus can be provided, comprising: processing circuitry; and memory circuitry comprising computer program code, the memory circuitry and the computer program code configured to, with the processing circuitry, cause the apparatus to: receive spatial audio content; receive spatial metadata associated with the spatial audio content; and receive information indicating a method for compressing the spatial metadata associated with the spatial audio content, wherein the method for compressing the spatial metadata is selected based on a source format of the spatial audio content.
[0025] According to various but not all examples of the present disclosure, a decoding device can be provided, comprising an apparatus as described above and one or more transceivers configured to receive spatial audio content and spatial metadata from the decoding device.
[0026] According to various but not all examples of the present disclosure, a method can be provided, comprising: receiving spatial audio content; receiving spatial metadata associated with the spatial audio content; and receiving information indicating a method for compressing the spatial metadata associated with the spatial audio content, wherein the method for compressing the spatial metadata is selected based on a source format of the spatial audio content.
[0027] The information indicating the method for compressing the spatial metadata can comprise a source configuration parameter.
[0028] According to various but not all examples of the present disclosure, a computer program can be provided, the computer program comprising computer program instructions which, when executed by processing circuitry, cause the following: receiving spatial audio content; receiving spatial metadata associated with the spatial audio content; and receiving information indicating a method for compressing the spatial metadata associated with the spatial audio content, wherein the method for compressing the spatial metadata is selected based on a source format of the spatial audio content.
[0029] The information indicating the method for compressing the spatial metadata can comprise a source configuration parameter.
[0030] According to various but not all examples of the present disclosure, a physical entity embodying the above computer program can be provided.
[0031] According to various but not all examples of the present disclosure, an electromagnetic carrier signal carrying the above computer program can be provided. BRIEF DESCRIPTION OF DRAWINGS
[0032] Some exemplary embodiments will now be described with reference to the accompanying drawings, in which:
[0033] Figure 1 An exemplary device is shown;
[0034] Figure 2 An exemplary method is shown;
[0035] Figure 3 An exemplary system is shown;
[0036] Figure 4 An exemplary encoding device is shown;
[0037] Figure 5 An exemplary decoding device is shown;
[0038] Figure 6 Another exemplary method is shown;
[0039] Figure 7 An exemplary encoding method is shown;
[0040] Figure 8 Another exemplary encoding method is shown;
[0041] Figure 9 An exemplary decoding method is shown. Detailed Implementation
[0042] The accompanying drawings illustrate an apparatus 101, which includes components for obtaining spatial metadata associated with spatial audio content. The spatial audio content may represent immersive audio content or any other suitable type of content. The components may also be configured to obtain configuration parameters indicating the source format of the spatial audio content; and to use these configuration parameters to select a compression method for the spatial metadata associated with the spatial audio content.
[0043] The device 101 can be used to record and / or process the captured audio signal.
[0044] Figure 1 An example of an apparatus 101 according to this disclosure is shown schematically. Figure 1 The device 101 shown may be a chip or chipset. In some examples, device 101 may be provided within a device such as a processing device. In some examples, device 101 may be provided within an audio capture device or an audio rendering device.
[0045] exist Figure 1 In the example, device 101 includes controller 103. Figure 1In examples, the implementation of the controller 103 can be as a controller circuit. In some examples, the controller 103 can be implemented in hardware only, have certain aspects in software only including firmware, or can be a combination of hardware and software (including firmware).
[0046] As Figure 1 The controller 103 can be implemented using instructions that enable hardware functionality, for example, by using executable instructions of a computer program 109 (which can be stored on a computer-readable storage medium (disk, memory etc.) to be executed by such a processor 105) in a general purpose or special purpose processor 105, as shown in
[0047] The processor 105 is configured to read from and write to memory 107. The processor 105 can also include an output interface and an input interface by which the processor 105 outputs data and / or commands and inputs data and / or commands, respectively.
[0048] The memory 107 is configured to store the computer program 109 including computer program instructions (computer program code 111) which, when loaded into the processor 105, control the operation of the apparatus 101. The computer program instructions of the computer program 109 provide the logic and routines Figure 2 and 6 The methods shown in Figures 1 to 9. By reading the memory 107, the processor 502 is able to load and execute the computer program 109.
[0049] Thus, the apparatus 101 comprises: at least one processor 105; at least one memory 107 including computer program code 111, the at least one memory 107 and the computer program code 111 configured to, with the at least one processor 105, cause the apparatus 101 at least to: obtain spatial metadata associated with a spatial audio content; obtain 203 configuration parameters indicative of a source format of the spatial audio content; and use 205 the configuration parameters to select a compression method of the spatial metadata associated with the spatial audio content.
[0050] As Figure 1As shown in the figure, the computer program 109 can arrive at the apparatus 101 via any suitable delivery mechanism 113, such as a machine readable medium, a computer readable medium, a non-transitory computer readable storage medium, a computer program product, a memory device, a record medium such as a CD-ROM or a DVD, or a storage device such as a flash memory, a hard disk, or a solid state drive. The delivery mechanism 113 can be a signal configured to reliably transfer the computer program 109. The apparatus 101 can propagate or transmit the computer program 109 as a computer data signal on a carrier that is modulated to have the computer program 109. In some examples, the computer program 109 can be transmitted from a first block of data to a second block of data, which can be contiguous with the first block of data or non-contiguous with the first block of data. The computer program 109 can be transmitted using a wireless protocol such as Bluetooth, Bluetooth Low Energy, Bluetooth Smart, 6L0WPan (IPv6-based Low-Power Personal Area Network), ZigBee, ANT+, Near Field Communication (NFC), Radio Frequency Identification, Wireless Local Area Network (Wireless LAN), or any other suitable protocol.
[0051] The computer program 109 comprises computer program instructions for causing the apparatus 101 to perform at least the following: obtaining 201 spatial metadata associated with a spatial audio content; obtaining 203 configuration parameters indicative of a source format of the spatial audio content; and using 205 the configuration parameters to select a compression method for the spatial metadata associated with the spatial audio content.
[0052] The computer program instructions can be included in the computer program 109, a non-transitory computer-readable medium, a computer program product, a machine-readable medium. In some but not all examples, the computer program instructions can be distributed over more than one computer program 109.
[0053] Although the memory 107 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cache storage.
[0054] Although the processor 105 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which can be integrated / removable. The processor 105 can be a single core or multicore processor.
[0055] References to "computer-readable storage medium", "computer program product", "tangibly embodied computer program" etc., or a "controller", "computer", "processor" etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application- specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the firmware a programmable logic controller is to be programmed with.
[0056] As used in this application, the term "circuitry" refers to all of the following:
[0057] (a) hardware-only circuitry such as comprises only analog and / or digital circuitry;
[0058] (b) combinations of hardware circuits and software, such as (as applicable):
[0059] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware;
[0060] (ii) any part of hardware processor with software (including digital signal processors); and
[0061] (c) hardware circuit(s) whether or not combined with software such as (as applicable):
[0062] This definition of "circuitry" applies to all uses of this term in this application, including in any claims. As a further example, as used in this application the term "circuitry" also covers an implementation that has a hardware component and a software component. The term "circuitry" also covers (for example and if applicable to the particular claim element), for a software implementation, a machine-readable medium having instructions. The term "circuitry" also covers (for example and if applicable to the particular claim element) for a hardware implementation, a hardware-only implementation of only hardware circuitry or only hardware processors.
[0063] Figure 2 An example method is shown. The method can be implemented using the apparatus 101 as shown in Figure 1 .
[0064] The method comprises obtaining spatial metadata associated with spatial audio content at block 201. In some examples, the spatial metadata can be obtained with the spatial audio content. In other examples, the spatial metadata can be obtained separately from the spatial audio content. For example, the apparatus 101 can obtain the spatial audio content and can separately process the spatial audio content to obtain the spatial metadata.
[0065] The spatial audio content comprises content that can be rendered to enable a user to perceive spatial characteristics of the audio content. For example, the spatial audio content can be rendered to enable a user to perceive a direction of origin and a distance to an audio source. The spatial audio can enable an immersive audio experience to be provided to a user. The immersive audio experience can comprise a virtual reality, augmented reality, mixed reality or extended reality experience or any other suitable experience.
[0066] The spatial metadata associated with the spatial audio content comprises information relating to spatial characteristics of a sound space represented by the spatial audio content. The spatial metadata can comprise information such as a direction of audio arrival, a distance to an audio source, a direct to total energy ratio, a diffuse to total energy ratio or any other suitable information. The spatial metadata can be provided in a frequency band.
[0067] At block 203, the method comprises obtaining a configuration parameter indicative of a source format of the spatial audio content. The configuration parameter can be indicative of a format of the spatial audio that has been used to obtain the spatial metadata. In some examples, the source format can be indicative of a configuration of microphones that have been used to capture the spatial audio content, where the spatial audio content has subsequently been used to obtain the spatial metadata.
[0068] The source format can be any suitable format type. Examples of different source formats include a three-dimensional spatial microphone configuration, a two-dimensional spatial microphone configuration, a mobile phone with four or more microphones configured for three-dimensional audio capture, a mobile phone with three or more microphones configured for two-dimensional audio capture, a mobile phone with two microphones, a surround sound such as a 5.1 mix or a 7.1 mix or any other suitable type of source format. Different source formats will result in spatial audio content with associated spatial metadata. Different spatial metadata associated with different source formats can have different characteristics.
[0069] The configuration parameter can comprise data bits indicative of the source format. For example, in some examples, the configuration parameter can comprise eight data bits which enable 256 different combinations for indicating the source format. In other examples of the disclosure, other numbers of bits can be used.
[0070] In this example, the data bits can be configured in a predefined format. For example, if the configuration parameter comprises eight bits, the first two bits can define the overall source type. This overall source type can indicate that the source is a microphone array, a channel-based source, a mobile device, or a mix. A mix source can comprise audio captured by a microphone array mixed with a channel-based source. For example, a microphone array can be used to capture spatial audio, with a channel-based music track added as background audio. The channel-based audio track can be provided from an audio file selected via a user interface or by any other suitable control means. It will be appreciated that other mix sources can be used in other examples of the disclosure.
[0071] The third bit can indicate whether the source contains an elevation angle. For example, the third bit can indicate "true" or "false" depending on whether the source contains an elevation angle.
[0072] The remaining five bits can comprise more detailed information about the source format. The more detailed information about the source format can be the type of microphone array, which can indicate the number of microphones and the relative positions of the microphones or any other suitable type of format. In some examples, the more detailed information about the source format can define a channel configuration, such as 5.1, 7.1, 7.1+4, 22.2, 2.0 or any other suitable type of channel configuration. In some examples, the more detailed information about the source format can indicate the type of mobile device that has been used to capture the spatial audio. For example, it can indicate that the device is a specific six-microphone mobile device, a generic four-microphone device, a generic three-microphone device or any other suitable type of device. In some examples, the more detailed information about the source type can define a combination of different source types. For example, it can comprise a 5.1 channel-based format and one or more mobile devices or any other type of combination.
[0073] It will be appreciated that other arrangements of bits can be used in other examples of the disclosure. For example, in some examples, it can be determined from the indication of the source format whether the source contains an elevation angle, and therefore in this case the third bit indicating whether the source contains an elevation angle can not be required. For example, if the source format is indicated as 5.1, it will inherently be a source format without an elevation angle, whereas if the source format is indicated as 7.1+4, it will inherently be a source format with an elevation angle.
[0074] In some examples, a list of source formats can be used, and the source configuration parameter can indicate a source format from this list.
[0075] At block 205, the method comprises using the configuration parameter to select a compression method for spatial metadata associated with the spatial audio content. For example, a plurality of compression methods can be available, and the configuration parameter can be used to select one of these available parameters.
[0076] In some examples, the configuration parameters can be used to select a codebook to compress the spatial metadata associated with the spatial audio content. The codebook can be any suitable spatial metadata compression codebook that can be used to both encode and decode the spatial metadata. The codebook can comprise a lookup table that can be used to compress and then reconstruct the values of the spatial metadata. In some examples, the codebook can comprise a combination of a lookup table and an algorithm and any other suitable method. In some examples, a switching system can be used that enables switching between different types of codebook.
[0077] In some examples, the configuration parameters can be used to select one or more algorithms. The algorithms in turn can be used to generate a codebook or other compression method. For example, in some examples, the configuration parameters can enable selection of an algorithm that enables calculation of values based on transmitted index values.
[0078] If the configuration parameters enable selection of a codebook, the codebook can be prepared in advance based on statistical information of a set of input samples representing a class of source formats. In turn, the correct codebook can be selected from the prepared codebooks based at least in part on the source configuration parameters.
[0079] In some examples, the configuration parameters can be used to enable creation of a codebook for compressing spatial metadata. The source configuration parameters can provide some information about the statistics of the parameters, and this information can be used to create a new codebook and / or modify an existing codebook.
[0080] Information indicating that a codebook has been selected can be transmitted from the encoding device to the decoding device. The information indicating that a codebook has been selected can be transmitted as a dynamic value in the metadata stream. In other examples, the information indicating that a codebook has been selected can be transmitted through a separate channel at the beginning of the transmission or at a particular point in time during the transmission.
[0081] Figure 3 An example system 301 that can be used in implementations of the present disclosure is shown. The system 301 includes an encoding device 303 and a decoding device 305. It should be understood that in other examples, the system 301 can include additional components not shown in the system 301 of FIG. 3, for example, the system can include one or more intermediate devices such as storage devices. Figure 1 It should be understood that in other examples, the system 301 can include additional components not shown in the system 301 of FIG. 3, for example, the system can include one or more intermediate devices such as storage devices.
[0082] The encoding device 303 can be any device configured to obtain spatial metadata associated with spatial audio content. In some examples, the encoding device 303 can be configured to encode the spatial audio content and the spatial metadata.
[0083] In some examples, the encoding device 303 can be configured to obtain the spatial metadata from a separate device. In other examples, the encoding device 303 can be configured to obtain the spatial metadata from a storage device. Figure 3In the example, encoding device 303 includes analysis processor 105A. Analysis processor 105A is configured to receive input audio signal 311. The input audio signal may represent a captured spatial audio signal. The input audio signal may be received from a microphone array, a multi-channel loudspeaker, or any other suitable source. In some examples, the input audio signal 311 may include an Ambisonics signal or a variant of an Ambisonics signal. In some examples, the audio signal may include a first-order Ambisonics (FOA) signal or a higher-order Ambisonics (HOA) signal or any other suitable type of spherical harmonic signal.
[0084] In some examples, the analysis processor 105A can be configured to analyze the input audio signal 311 to obtain spatial audio content and spatial metadata. It should be understood that in other examples, the analysis processor 105A can receive both spatial audio content and spatial metadata. In such examples, the analysis processor 105A does not need to analyze the spatial audio content to obtain spatial metadata.
[0085] The analysis processor 105A is configured to create a transmission signal 313 for spatial audio content and spatial metadata. The analysis processor 105A can be configured to encode both the spatial audio content and spatial metadata to provide the transmission signal 313.
[0086] exist Figure 3 In the exemplary system 301 shown, a transmission signal 313 is sent to a decoding device 305. In some examples, the transmission signal 313 may be sent to a storage device, from which it can then be retrieved by one or more decoding devices. In other examples, the transmission signal 313 may be stored in the memory of an encoding device 303. The transmission signal 313 can then be retrieved from the memory for decoding and rendering at a subsequent point in time.
[0087] exist Figure 3 In the example, the decoding device 305 includes a synthesis processor 105B. The synthesis processor 105B is configured to receive a transmitted signal 313 and synthesize a spatial audio output signal 315 based on the received transmitted signal 313. The synthesis processor 105B decodes the received transmitted signal to synthesize the spatial audio output signal 315.
[0088] The synthesis processor 105B uses spatial metadata to create spatial characteristics of the spatial audio content, providing the listener with spatial audio content that represents the spatial characteristics of the captured sound scene. Spatial audio enables the provision of immersive audio to the user. The spatial audio output signal 315 can be a multi-channel speaker signal, a binaural signal, a spherical harmonic signal, or any other suitable type of signal.
[0089] The spatial audio output signal 315 can be provided to any suitable rendering device, such as one or more loudspeakers, headphones, or any other suitable rendering device.
[0090] Figure 4 Features of the example encoding device 303 are shown in more detail. The example encoding device 303 includes a transport audio signal generator 401, a spatial analyzer 403, and a multiplexer 405. In some examples, the transport audio signal generator 401, the spatial analyzer 403, and the multiplexer 405 can comprise modules within the analysis processor 105A.
[0091] The transport audio signal generator 401 receives the input audio signal 311, which includes spatial audio content. The transport audio signal generator 401 is configured to generate a transport audio signal 411 from the received input audio signal 311. The source format of the spatial audio content can be used to generate the transport audio signal. For example, to generate a stereo transport audio signal, if the spatial audio content is captured by a microphone array such as a spherical microphone grid, then two opposing microphones can be selected as the transport signals. Equalization or other suitable processing can be applied to the transport signals.
[0092] The transport audio signal 411 can comprise a mono signal, a stereo signal, a binaural stereo signal, or any other suitable signal, e.g., a FOA signal.
[0093] The spatial analyzer 403 also receives the input audio signal 311, which includes spatial audio content. The spatial analyzer 403 is configured to analyze the spatial audio content to provide spatial parameters that form spatial metadata. The spatial parameters represent spatial characteristics of the sound space represented by the spatial audio content. The spatial parameters can include information such as audio direction of arrival, distance to audio source, direct-to-total energy ratio, diffuse-to-total energy ratio, or any other suitable parameter. The spatial analyzer 403 can analyze different frequency bands of the spatial audio content, so that spatial metadata can be provided in the frequency bands. For example, a suitable set of frequency bands can be 24 frequency bands following the Bark scale. Other sets of frequency bands can be used in other examples of the disclosure.
[0094] The spatial analyzer 403 provides one or more output signals that include spatial metadata. In the example shown in FIG. 4, the spatial analyzer 403 provides a first output 415 that indicates direction parameters and a second output 417 that indicates direct-to-total energy ratios for different frequency bands. It will be appreciated that other outputs and parameters can be provided in other examples of the disclosure. These other parameters can be provided in place of, or in addition to, the direction parameters and energy ratios. Figure 4 In the example shown in FIG. 4, the spatial analyzer 403 provides a first output 415 that indicates direction parameters and a second output 417 that indicates direct-to-total energy ratios for different frequency bands. It will be appreciated that other outputs and parameters can be provided in other examples of the disclosure. These other parameters can be provided in place of, or in addition to, the direction parameters and energy ratios.
[0095] The multiplexer 405 is configured to receive the transport audio signal 411 and the spatial metadata output 415, 417 and combine them to generate the transport signal 313.
[0096] In Figure 4 examples, the multiplexer also receives an additional input 419 comprising source configuration parameters. The source configuration parameters are indicative of a source format of the spatial audio content.
[0097] In Figure 4 examples, the source configuration parameters are received separately from the spatial audio content. For example, information about the source format can be stored in a memory and can be retrieved by the multiplexer. In other examples, information about the source format can be received together with the spatial audio content. In some examples, the transport audio signal generator 401 and / or the spatial analyzer 403 can also use the source configuration parameters.
[0098] The multiplexer 405 is configured to encode the spatial audio content as well as the spatial metadata. The source configuration parameters are used to select a compression method for the spatial metadata. For example, the source configuration parameters can be configured to select a codebook for encoding the spatial metadata.
[0099] In Figure 4 examples, the multiplexer 405 comprises a transport audio signal encoding module 421 and a spatial metadata encoding module 423. The transport audio signal encoding module 421 is configured to encode and / or compress the transport audio signal 411. The spatial metadata encoding module 423 is configured to encode and / or compress the spatial metadata available from the spatial analyzer 403. Different encoding and / or compression methods can be used to encode the audio content and the spatial metadata.
[0100] The multiplexer further comprises a data stream generator / combiner module 425. The data stream generator / combiner module 425 is configured to combine the compressed transport audio signal and the compressed spatial metadata into the transport signal 313, which is provided as an output of the encoding device 303.
[0101] In Figure 4 the example shown in Fig. 3, the transport audio signal generator 401, the spatial analyzer 403 and the multiplexer 405 are all shown as part of the same encoding device 303. It will be appreciated that other configurations can be used in other examples of the present disclosure. In some examples, the transport audio signal generator 401 and the spatial analyzer 403 can be provided in a device or system separate from the multiplexer 405. For example, if MASA (metadata assisted spatial audio) is used, the spatial analysis is performed before the content is provided to the encoding device 303. In such examples, the encoding device 303 obtains a file or stream comprising the spatial metadata and the transport audio signal 411.
[0102] Figure 5 Features of the exemplary decoding device 305 are shown in more detail below. The exemplary decoding device 305 includes a demultiplexer 501, a prototype signal generator module 503, a direct stream generator module 505, a diffuse stream generator module 507, and a stream combiner module 509. The demultiplexer 501, prototype signal generator module 503, direct stream generator module 505, diffuse stream generator module 507, and stream combiner module 509 may include modules within the synthesis processor 105B.
[0103] Demultiplexer 501 receives a transmission signal 313, comprising encoded spatial audio content and encoded spatial metadata, as input. The transmission signal may include configuration parameters. Demultiplexer 501 is configured to receive the transmission signal 313 and separate it into two or more individual components. Figure 5 In the example, demultiplexer 501 is configured to separate the transmitted signal 313 into a separate decoded transmitted audio signal 511 and one or more outputs 513, 515 including decoded spatial metadata.
[0104] exist Figure 5 In the example, demultiplexer 501 includes a data stream receiver / splitter module 521. The data stream receiver / splitter module 521 is configured to receive the transmitted signal 313 and split it into a first component that includes at least spatial audio content and a second component that includes spatial metadata.
[0105] The demultiplexer 501 also includes a transmit audio signal decompressor / decoder module 523. The transmit audio signal decompressor / decoder module 523 is configured to receive components containing audio content from the data stream receiver / splitter module 521 and decompress that audio content. Furthermore, the transmit audio signal decompressor / decoder module 523 provides a decoded transmit audio signal 511 as output.
[0106] exist Figure 5 In the example shown, demultiplexer 501 also includes a metadata decompressor / decoder module 525. The metadata decompressor / decoder module 525 is configured to receive components including metadata from data stream receiver / splitter module 521. The metadata decoder module 525 decompresses the spatial metadata using a decompression method indicated by source configuration parameters. This can be a different decompression method than that used for spatial audio content. Once the spatial metadata has been decompressed, the metadata decompressor / decoder module 525 provides one or more outputs 513, 515 including the decoded spatial metadata. Figure 5In the example shown in FIG. 5, the metadata decompressor / decoder module 525 provides a first output 513 and a second output 515, where the first output 513 includes spatial metadata related to the direction of the spatial audio content and the second output 515 includes spatial metadata related to the energy ratio of the spatial audio content. It will be appreciated that other outputs providing data related to other spatial parameters can be provided in other examples of the present disclosure.
[0107] In Figure 5 In the example shown in FIG. 5, the decoded transport audio signal 511 is provided to a prototype signal generator module 531. The prototype signal generator module 531 is configured to create appropriate prototype signals 541 for the output device being used to render the spatial audio content. For example, if the output device includes a speaker setup in a 5.1 configuration and the transport audio signal 511 is a stereo signal, the left channel would receive the left signal, the right channel would receive the right signal, and the center channel would receive a mix of the left and right signals. It will be appreciated that other types of output devices can be used in other examples of the present disclosure. For example, the output device can be a different arrangement of speakers, or can be headphones, or can be any other suitable type of output device.
[0108] The prototype signals 541 from the prototype signal generator module 531 are provided to both the direct stream generator module 505 and the diffuse stream generator module 507. In Figure 5 In the example shown in FIG. 5, the direct stream generator module 505 and the diffuse stream generator module 507 also receive the outputs 513, 515 including spatial metadata. In other embodiments, different and / or other types of spatial metadata can be used. In some examples, different spatial metadata can be provided to the direct stream generator module 505 and the diffuse stream generator module 507.
[0109] In Figure 5 In the example shown in FIG. 5, the direct stream generator module 505 and the diffuse stream generator module 507 use the spatial metadata to create direct streams 543 and diffuse streams 545, respectively. For example, the spatial metadata related to the direction parameter can be used to create the direct streams 543 by panning the sound to the direction indicated by the metadata. The diffuse streams 545 can be created from the decorrelated signals of all or substantially all of the available channels.
[0110] The diffuse streams 545 and the direct streams 543 are provided to a stream combiner module 509. The stream combiner module 509 is configured to combine the direct streams 543 and the diffuse streams 545 to provide the spatial audio output signal 315. The spatial metadata related to the energy ratio can be used to combine the direct streams 543 and the diffuse streams 545.
[0111] The spatial audio output signal 315 can be provided to a rendering device, such as one or more loudspeakers, headphones, or any other suitable device configured to convert the electronic spatial audio output signal 315 into an audible signal.
[0112] In the example shown in Figure 5 In the example shown in FIG. 5, the demultiplexer 501, the prototype signal generator module 503, the direct stream generator module 505, the diffuse stream generator module 507, and the stream combiner module 509 are all shown as part of the same decoding device 305. It will be appreciated that other configurations can be used in other examples of the disclosure. For example, in some examples, the output of the demultiplexer 501 can be stored as a file in memory. It can then be provided to a separate device or system for processing to obtain the spatial audio output signal 315.
[0113] Figure 6 A method that can be used in some examples of the disclosure to create a codebook for compressing spatial metadata is shown. Figure 6 The method shown in FIG. 6 can be performed by an encoding device 303 such as the encoding device 303 shown in Figure 4 The method shown in FIG. 6 can be performed by an encoding device 303 such as the encoding device 303 shown in
[0114] At block 601, a source configuration is selected. The source configuration is a format that is used to capture audio signals. The selection of the source configuration can include selecting a microphone arrangement to be used to capture audio signals, selecting a device to be used to capture audio signals, selecting a premixed channel format, or any other selection.
[0115] At block 603, spatial audio content is obtained. The obtained spatial audio content is captured using the source configuration selected at block 601. The spatial audio content can include a representative set of audio samples. The representative set of samples can include a standard set of acoustic signals that can be used for the purpose of creating a codebook for compressing spatial metadata. The representative set of samples can include one or more acoustic samples having different spatial characteristics.
[0116] At block 605, a spatial analysis is performed on the obtained spatial audio content. The spatial analysis determines one or more spatial parameters of the spatial audio content. The spatial parameters can be direction parameters, energy ratio parameters, coherence parameters, or any other suitable parameters. The spatial analysis performed can be the same spatial analysis process performed by the spatial analyzer 403 of the encoding device 303 to obtain the spatial metadata. If the obtained spatial audio content includes a representative set of samples, the same spatial analysis can be performed on each sample within the set.
[0117] At block 607, the statistical information of the spatial parameters obtained at block 605 is analysed. This analysis enables the determination of the probability of occurrence for each parameter value. The analysis can comprise counting each occurrence of a parameter value from the obtained spatial audio. A histogram or any other suitable means can be used to count the occurrences.
[0118] At block 609, the method comprises designing a codebook using the statistical information obtained at block 607. For example, the codebook can be designed such that the most likely parameters have the shortest code values, while the least likely parameters are assigned longer code values. This can be achieved by ordering the parameter values from the highest occurrence rate to the lowest occurrence rate, and in turn assigning code values to the ordered parameter values, starting with the parameter value having the highest occurrence rate which is assigned the shortest available code value. This ensures that the spatial metadata will use fewer bits for each value after it has been compressed. The codebook created thereby can comprise a lookup table, or any other suitable information. In some examples, one or more algorithms can be used to generate the codebook.
[0119] At block 611, the codebook is stored. The codebook can be stored in the memory of the encoding device 303 or any other suitable storage location. The codebook is stored so that it can be accessed during the compression and decompression of the spatial metadata.
[0120] Figure 6 The method of Figure 6 shows an example of creating a codebook. In other examples, an existing codebook can be modified by applying known restrictions to it. For example, a codebook for a three-dimensional microphone can be available, but the source format can be a two-dimensional microphone array. In this example, the codebook for the three-dimensional array can be modified so that all horizontal direction parameter values receive shorter code values in the codebook. As another example, a codebook can be available for 5.1 speaker input, but the source format can be 2.0 speaker input. In this example, the codebook for 5.1 speaker input can be modified so that direction parameter values between -30° and 30° receive shorter code values.
[0121] Figure 6 The example method of Figure 6 shows an example of creating a codebook. This method can be performed by a vendor such as a mobile device manufacturer as part of the product specification. Once the codebook has been created, it can be used to encode and decode spatial metadata. The codebook can be used by a device such as an immersive audio capture device. Configuration parameters can be associated with the codebook so that the correct codebook can be selected for the encoding and decoding of spatial metadata.
[0122] Figure 7 The example method of Figure 7 shows an example of encoding spatial audio and spatial metadata. Figure 7The exemplary method shown can be derived from Figure 4 The encoding device 303 shown in the diagram, or the multiplexer 405, or any other suitable device, shall perform this operation. Figure 7 In the example shown, the input signal is provided in a parameterized spatial audio format with separate spatial audio content and spatial metadata, and the source configuration parameters are provided as part of this format.
[0123] At box 701, multiplexer 405 obtains the audio content. The audio content can be obtained from the transmitted audio signal 411. For example... Figure 4 As shown, the transmitted audio signal 411 can be obtained from the transmitted audio signal generator 401. The audio content has been captured using a source format. The source format may have been pre-selected before capturing the audio content, or it may be defined by the device used to capture spatial audio.
[0124] At box 703, multiplexer 405 obtains spatial metadata. The spatial metadata may include outputs 415, 417 from spatial analyzer 403. The spatial metadata may be provided in a parameterized format, including values of one or more spatial parameters of the spatial audio content provided within the transmitted signal 411. For example... Figure 4 As shown, metadata can be obtained from spatial analyzer 403.
[0125] At box 705, multiplexer 405 obtains source configuration parameters. Input source configuration parameters indicate an equivalent description of the source format or source configuration used to capture spatial audio. Source configuration parameters can be received as input from the capture device, or they can be received in response to user input via a user interface or any other suitable means. Source configuration parameters can be obtained as part of a spatial metadata packet. In this example, obtaining source configuration parameters may include reading parameters from a spatial metadata packet.
[0126] At box 707, compress the spatial audio content. Any suitable technique can be used to compress the spatial audio content. Figure 7 In the example shown, the source configuration parameters are not used to compress the audio transmission signal 411, which includes spatial audio content. The audio transmission signal 411 can be compressed using any suitable process such as AAC (Advanced Audio Coding), EVS (Enhanced Voice Services), or any other suitable process.
[0127] At block 709, a compression method for the spatial metadata is selected. The obtained source configuration parameters are used to select the compression method for the spatial metadata. Selecting the compression method can comprise selecting a pre-formed codebook corresponding to the source format for the captured spatial audio. The pre-formed codebook can be stored in a memory of the encoding device 303, or any memory accessible to the encoding device 303. In some examples, selecting the compression method can comprise selecting a computable or algebraic codebook, where the codebook is based on an algorithm.
[0128] Once the pre-formed codebook has been obtained from the memory, it can be passed to the spatial metadata encoding module 423, so that at block 711, the codebook can be used to compress the spatial metadata. The method of compressing the spatial metadata can be any compression method that uses the codebook. For example, the method can comprise Huffman coding or any other suitable process.
[0129] In some examples, a quantization process can be performed prior to compressing the spatial metadata. The quantization process can comprise quantizing the parameter values of the parametric spatial metadata, so that each parameter value has a corresponding code value. In some examples, the source configuration parameters can also be used for the quantization process, as the optimal quantization can also depend on the source format. For example, when there is an elevation angle in the source format, a spherical uniform quantization can be applied to the direction parameters, in order to obtain a more uniform and perceptually better quantized direction distribution than what is achievable with other quantization processes.
[0130] In some examples, the source configuration parameters can be used to determine the quantization process used. In this case, a separate indication of the source configuration parameters can not have to be provided to the decoder device 305, as the correct source configuration and / or method compression can be inherent to the quantization process.
[0131] At block 713, the compressed spatial audio content and the compressed spatial metadata are encoded together to form the encoded transmission signal 313. The combination of the compressed spatial audio content and the compressed spatial metadata can be performed by the data stream generator / combiner module 425 or any other suitable module. In some examples, the combination of the compressed spatial audio content and the compressed spatial metadata can also comprise further compression, such as run-length encoding or any other lossless encoding.
[0132] Figure 8 An example method of encoding spatial audio and spatial metadata is shown. Figure 8 The example method shown in Figure 8 can be performed by the encoding device 303 of an audio capture device, or any other suitable device. In Figure 8 In the example shown in Figure 8 , the input signal is provided to the encoding device 303 in a parametric spatial audio format, as shown in Figure 7 In contrast, in the example shown in Figure 8 , the input signal is provided to the encoding device 303 in a non-parametric spatial audio format, as shown in Figure 8In the example of FIG. 3, spatial audio is analyzed within the encoding device 303 to determine spatial metadata.
[0133] At block 801, spatial audio is captured. The spatial audio is captured using a source format.
[0134] At block 805, the captured spatial audio is processed to form an audio transport signal 411. The audio transport signal 411 includes audio content. Processing the captured spatial audio to form the audio transport signal 411 can be performed by the transport audio signal generator 401 or any other suitable component.
[0135] At block 807, a spatial analysis is performed on the spatial audio content to obtain spatial metadata. The spatial analysis can be performed by the spatial analyzer 403 as shown in Figure 4 FIG. 4 or by any other suitable component. The spatial metadata can be provided in a parametric format. That is, the spatial metadata can include one or more spatial parameters and can include values for one or more spatial parameters of the spatial audio.
[0136] At block 803, source configuration parameters are obtained. The input source configuration parameters indicate the source format used to capture the spatial audio. The source configuration parameters can be stored in a memory of the audio capture device or can be received in response to user input via a user interface or by any other suitable means.
[0137] At block 809, the audio transport signal 411 including the spatial audio content is compressed. The audio transport signal 411 can be compressed using any suitable technique. In the example shown in Figure 8 FIG. 4, the source configuration parameters are not used to compress the audio transport signal 411 including the spatial audio content. The audio transport signal 411 can be compressed using any suitable process such as AAC (Advanced Audio Coding), EVS (Enhanced Voice Services), or any other suitable process.
[0138] At block 811, a compression method is selected for the spatial metadata. The obtained source configuration parameters are used to select the compression method for the spatial metadata. As shown in the method of Figure 7 FIG. 4, selecting the compression method can include selecting a pre-formed codebook corresponding to the source format used for the captured spatial audio. The pre-formed codebook can be stored in a memory of the encoding device 303 or any memory accessible to the encoding device 303.
[0139] Once the pre-formed codebook has been retrieved from memory, it can be passed to the spatial metadata encoding module 423, such that at block 813, the codebook can be used to compress the spatial metadata. The method of compressing the spatial metadata can be any compression method that uses the codebook. For example, the method can comprise Huffman coding or any other suitable process. A quantization process can be applied to the spatial metadata prior to compressing the spatial metadata.
[0140] At block 815, the compressed spatial audio content and the compressed spatial metadata are encoded together to form the encoded transmission signal 313. The combination of the compressed spatial audio content and the compressed spatial metadata can be performed by the data stream generator / combiner module 425 or any other suitable module. In some examples, the combination of the compressed spatial audio content and the compressed spatial metadata can also comprise further compression, such as run-length encoding or any other lossless encoding.
[0141] Figure 9 An example decoding method is shown. Figure 9 The example method shown in Figure 5 may be performed by the decoding device 305 shown in
[0142] At block 901, the received encoded transmission signal 313 is decoded into separate transmission audio streams and spatial metadata streams. The transmission audio streams comprise audio content, while the spatial metadata streams comprise parametric values relating to the spatial characteristics of the transmission audio streams.
[0143] At block 903, the spatial audio content is decompressed from the transmission audio streams. Any suitable process can be used to decompress the spatial audio content. At block 905, a prototype signal 541 is formed. The prototype signal 541 can be formed by the prototype signal generator module 531 shown in Figure 5 or any other suitable component.
[0144] At block 907, source configuration parameters are obtained. In some examples, the source configuration parameters can be received with the encoded transmission signal 313. For example, the source configuration parameters can be encoded into the spatial metadata streams. In such examples, the source configuration parameters can be provided as the first value in the spatial metadata streams or any other defined value in the spatial metadata streams. Providing the source configuration parameters with the spatial metadata streams can allow the source configuration to be updated for different signal frames, which can help to improve compression efficiency.
[0145] In other examples, the source configuration parameters can be received separately from the encoded transmission signal 313. This can be provided through a separate signalling channel for the spatial metadata or the spatial audio content. For example, the source configuration parameters can be provided separately to the bitstream that transmits the audio content and the spatial metadata.
[0146] At box 909, the source configuration parameters are used to select the decompression method for spatial metadata. Selecting the decompression method may include selecting the codebook based on the source configuration parameters.
[0147] At box 911, the selected decompression method is used to decompress the spatial metadata and provide the spatial metadata parameters to the synthesizer. The decompression of spatial metadata can be the reverse process of a process already used to compress spatial metadata. For example, decompressing spatial metadata may include reading code values from the spatial metadata stream and obtaining the corresponding parameter values from a selected codebook. In other examples, the code values from the spatial metadata stream can be used in algorithms that provide the corresponding parameter values via computation. In some examples, algorithms can be used instead of lookup tables. In other examples, algorithms can be used in addition to lookup tables.
[0148] In box 913, spatial metadata and prototype signal 541 are synthesized into a spatial audio output signal.
[0149] exist Figure 9 In the exemplary method shown, source configuration parameters are provided to decoding device 305. In other examples, a codebook, which has been selected by encoding device 303 based on source configuration parameters, can be passed between encoding device 303 and decoding device 305.
[0150] Therefore, this disclosure provides examples of apparatus, methods, and computer programs for efficiently encoding spatial metadata by enabling suitable compression methods to be used. This can be implemented as a separate process for encoding audio content.
[0151] The examples described above demonstrate applications that implement the following components:
[0152] Automotive systems; telecommunications systems; electronic systems including consumer electronics; distributed computing systems; media systems for generating or rendering media content including audio, visual, and audiovisual content, as well as mixed reality, mediated reality, virtual reality, and / or augmented reality; personal systems including personal health systems or personal fitness systems; navigation systems; user interfaces also known as human-machine interfaces; networks including cellular networks, non-cellular networks, and optical networks; self-organizing networks; the Internet; the Internet of Things; virtual networks; and related software and services.
[0153] The term “includes” as used herein has an inclusive rather than exclusive meaning. That is, any statement “X includes Y” means that X may include only one Y or may include more than one Y. If the intention is to use “includes” with an exclusive meaning, it will be made clear in the context by referring to “includes only one…” or by using “consisting of…”.
[0154] Various examples have been described in this specification. Descriptions of features or aspects within an example indicate that such features or aspects are present in that example. Whether explicitly stated or not, the use of the term "example" or "for example" or "may" or "can" in the text indicates that such features or aspects are present in at least the described example, whether explicitly described or not, and that such features or aspects can be present in some or all examples. Thus, "example," "for example," or "may" or "can" refers to a particular instance in a class of examples. The nature of an instance can be only that of the instance, or of the class of instances, or of a subclass of the class of instances that includes some but not all of the class of instances. Thus, features described for one example but not for another example are implicitly disclosed as part of a working combination for the other example, but are not necessarily part of the other example.
[0155] While embodiments have been described in the preceding paragraphs in terms of various examples, it is to be understood that modifications can be made to the examples given without departing from the scope of the claims.
[0156] Features described in the preceding description can be used in combinations other than the combinations explicitly described above.
[0157] "Explicitly" indicates that features from different embodiments (e.g., different methods with different flowcharts) can be combined.
[0158] Although functions have been described with reference to certain features, those functions can be performed by other features whether described or not.
[0159] Although features have been described with reference to certain embodiments, those features can also be present in other embodiments whether described or not.
[0160] The term "a" or "the" has the inclusive rather than the exclusive meaning. That is, any reference to "X including a / the Y" indicates that "X can include only one Y" or "X can include more than one Y", unless the context clearly indicates otherwise. If it is intended that "a" or "the" be used in an exclusive sense, then the context will so indicate. In some environments, "at least one" or "one or more" can be used to emphasize the inclusive meaning, but the absence of these terms does not imply an exclusive meaning.
[0161] The presence of a feature (or combination of features) in a claim is a reference to that feature (or combination of features) itself and also to equivalent features (equivalent features) that implement the same technical effect in essentially the same way. Equivalent features include, for example, features that are variants and that achieve essentially the same result in essentially the same way. Equivalent features include, for example, features that perform essentially the same function in essentially the same way to achieve essentially the same result.
[0162] In this specification, reference has been made to various examples of the use of adjectives or adjective phrases to describe the features of examples. Such description of the features of examples indicates that the feature is exactly as described in some examples and essentially the same in other examples.
[0163] While in the foregoing specification this application has been described in terms of certain embodiments, it will be understood that applicant can seek to claim, and an appropriate apparatus may be constructed to perform, the full breadth of equivalent features or combinations of features anywhere in this disclosure whether or not they are explicitly mentioned in the foregoing description.
Claims
1. An apparatus comprising processing circuitry; and memory circuitry including computer program code, the memory circuitry and the computer program code being configured to, together with the processing circuitry, cause the apparatus to: Obtain spatial metadata associated with spatial audio content; Obtain configuration parameters indicating the input source format of the spatial audio content; as well as Use the configuration parameters to select a compression method for the spatial metadata associated with the spatial audio content.
2. The apparatus according to claim 1, wherein, The configuration parameters are used to select a codebook to compress the spatial metadata associated with the spatial audio content.
3. The apparatus according to claim 1, wherein, The configuration parameters are used to enable the creation of a codebook for compressing the spatial metadata.
4. The apparatus according to claim 2, wherein, The codebook is used to encode and decode the spatial metadata.
5. The apparatus according to claim 1, wherein, The indicated input source format includes the format of the spatial audio content used to obtain the spatial metadata, wherein the indicated input source format may optionally include at least one of the following: At least one configuration of a microphone used to capture the spatial audio content; At least one configuration for a channel used to represent the spatial audio content; At least one type of mobile device used to capture the spatial audio content; or Does the source of the spatial audio content include an elevation angle? 6. The apparatus according to claim 1, wherein, The spatial metadata includes data indicating spatial parameters of the spatial audio content.
7. The apparatus according to claim 1, wherein, The compression method is selected independently of the content of the obtained spatial audio content.
8. The apparatus of claim 1 is further configured to obtain the spatial audio content.
9. The apparatus according to claim 8, wherein, The configuration parameters are obtained together with the spatial audio content.
10. The apparatus according to claim 8, wherein, The configuration parameters are obtained separately from the spatial audio content.
11. The apparatus of claim 1 is further configured to send compressed space metadata to the decoding device.
12. A method comprising: Obtain spatial metadata associated with spatial audio content; Obtain configuration parameters indicating the input source format of the spatial audio content; as well as Use the configuration parameters to select a compression method for the spatial metadata associated with the spatial audio content.
13. The method according to claim 12, wherein, The configuration parameters are used to select a codebook to compress the spatial metadata associated with the spatial audio content.
14. An apparatus comprising processing circuitry; and memory circuitry including computer program code, the memory circuitry and the computer program code being configured to, together with the processing circuitry, cause the apparatus to: Receive spatial audio content; Receive compressed spatial metadata associated with the spatial audio content; as well as Receive information indicating a method for compressing the spatial metadata associated with the spatial audio content, wherein the method is selected based on configuration parameters indicating the input source format of the spatial audio content.
15. The apparatus according to claim 14, wherein, The information indicating the method for compressing the spatial metadata includes the configuration parameters.
16. The apparatus according to claim 14, wherein, The information indicating the method for compressing the spatial metadata includes the codebook that has been selected using the configuration parameters.
17. The apparatus of claim 14, further comprising one or more transceivers configured to receive the spatial audio content and the compressed spatial metadata from the encoding device.
18. A method comprising: Receive spatial audio content; Receive compressed spatial metadata associated with the spatial audio content; as well as Receive information indicating a method for compressing the spatial metadata associated with the spatial audio content, wherein the method for compressing the spatial metadata is selected based on configuration parameters indicating the input source format of the spatial audio content.
19. The method according to claim 18, wherein, Information indicating the method used to compress the spatial metadata includes the configuration parameters.
20. The method according to claim 18, wherein, Information indicating the method used to compress the spatial metadata includes the codebook that has been selected using the configuration parameters.
Citation Information
Patent Citations
Selecting codebooks for coding vectors decomposed from higher-order ambisonic audio signals
CN106463129A
Quantization of spatial vectors
CN108140389A