Method, apparatus, and medium for encoding and decoding audio bitstreams with parametric flexible rendering configuration data
Patent Information
- Application Number
- JP2025519774
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2023-09-15
- Publication Date
- 2026-09-07
AI Technical Summary
Existing audio streaming technologies face challenges in delivering high-quality, low-latency audio signals over wireless links, especially in complex user setups with multiple speakers, where wireless quality is often inadequate, affecting both audio and embedded information.
A method and apparatus for encoding and decoding audio bitstreams using joint coding tools, flexible rendering, and metadata to adjust delay, gain, and equalization based on playback device positions, with support for various channel configurations and echo management, enabling efficient decoding and rendering on multiple devices.
Enables high-quality, low-latency audio delivery across multiple devices with adaptable rendering and efficient decoding, optimizing wireless transmission and reducing computational complexity.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates generally to audio signals, and more particularly to audio source encoding and decoding for low-latency exchange of audio signals of immersive audio programs between devices. [Background technology]
[0002] Streaming audio is commonplace in today's society. Audio streaming is becoming more demanding as user expectations for quality increase, but the number and variety of speakers is also increasing, and user setups are becoming more complex. Streaming is usually done at least in part over a wireless link, which places requirements on the wireless link to be of good quality, but as many of us have experienced, this is not always the case.
[0003] Therefore, there is a need to define an exchange format for use cases where a particular format is streamed from a cloud / server and then transcoded on the device to a more suitable low-latency format for delivery over a wireless (or possibly wired) link. Example use cases are in-home connectivity or phone-to-car connectivity, but this format could be beneficial in any scenario where low-latency delivery of an audio signal from a single device to one or more connected devices is desired.
[0004] In addition to audio information being transmitted wirelessly and streamed, there may be other types of information embedded in the stream that will then also be affected by the quality of the wireless link and may have similar drawbacks as audio.
[0005] It would therefore be advantageous to solve problems associated with wireless streaming for various types of streaming audio combined with other types of information or signals. Summary of the Invention
[0006] It is an object of the present disclosure to at least partially solve the above problems by wirelessly streaming audio and other types of information.
[0007] According to a first aspect of the present disclosure, there is provided a method for generating an encoded bitstream from an audio program including a plurality of audio signals, the method comprising: receiving, for each of a plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; receiving, for each playback device, information indicating at least one of a delay, a gain, and an equalization curve associated with each playback device; determining a group of two or more related audio signals from the plurality of audio signals; applying one or more joint coding tools to two or more related audio signals of the group to obtain a jointly coded audio signal; combining the jointly coded audio signal, an indication of a playback device with which the jointly coded audio signal is associated, and information indicating at least one of a delay, a gain, and an equalization curve associated with each playback device with which the jointly coded audio signal is associated into a separate block of the coded bitstream; It has.
[0008] According to a second aspect of the present disclosure, there is provided a method of decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising one or more independent blocks of encoded data, the method comprising: identifying, from the encoded bitstream, discrete blocks of encoded data corresponding to one or more audio signals associated with a playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; determining that the retrieved independent blocks of encoded data include two or more jointly coded audio signals; applying one or more joint decoding tools to the two or more jointly coded audio signals to obtain one or more audio signals associated with a playback device; determining at least one of a delay, a gain, and an equalization curve associated with the playback device from the retrieved individual blocks of encoded data; applying a delay, gain, and / or equalization curve associated with the playback device to one or more audio signals associated with the playback device; It has.
[0009] According to a third aspect of the present disclosure, there is an apparatus configured to carry out the method according to the first and / or second aspect.
[0010] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium comprising a sequence of instructions that, when executed, cause one or more devices to perform any one of the methods according to the first and / or second aspects.
[0011] Further examples of the present disclosure are defined in the dependent claims.
[0012] In some examples, the delay, gain, and / or equalization curve associated with each playback device depends on the position of each playback device relative to the position of the listener and / or the position of each playback device relative to the positions of other playback devices. In some embodiments, the delay, gain, and / or equalization curve change dynamically.
[0013] In some examples, the delay, gain, and / or equalization curves are adjusted in response to changes in the position of the listener, changes in the position of the playback device, and / or changes in the position of one or more of the other playback devices.
[0014] In some examples, the method further comprises determining an audio signal from the plurality of audio signals that is not part of a group of two or more related audio signals.
[0015] In some examples, the method further includes, for an audio signal that is not part of the group of two or more related audio signals, applying a delay, gain, and / or equalization curve associated with a playback device with which the audio signal is associated.
[0016] In some examples, the method further includes independently coding audio signals that are not part of a group of two or more related audio signals, and combining the independently coded audio signals and an indication of a playback device with which the independently coded audio signals are associated into a separate, independently decodable subset of the encoded bitstream.
[0017] In this disclosure, a frame represents an entire time slice of all signals. A block stream represents a collection of signals for the duration of a session. A block represents one frame of a block stream. For digital audio with a given sampling frequency, the frame size corresponds to the number of audio samples in a frame of any audio signal. The frame size typically remains constant for the duration of a session.
[0018] The wake word may consist of a single word, or a frame containing two or more words in a fixed order.
[0019] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X outputs in response to multiple inputs, where the subsystem generates M of the inputs and the remaining XM inputs are received from external sources) may also be referred to as a decoder system.
[0020] Examples of the present disclosure will be described in more detail with reference to the accompanying drawings. In the following drawings, like numbers are used to refer to like elements. The following figures depict various examples, although one or more implementations are not limited to the examples depicted in the figures. [Brief explanation of the drawings]
[0021] [Figure 1] 1 shows an example of low-latency transcoding for a home connection. [Figure 2] 1 illustrates an example of automotive audio streaming. [Figure 3] 1 illustrates an example of enhanced TV audio streaming using wireless devices. [Figure 4] 1 shows an example of audio streaming through a simple wireless speaker. [Figure 5] 10 illustrates an example of bitstream element to device metadata mapping. [Figure 6] This shows an example of how to deploy a simple flexible rendering of audio streaming. [Figure 7] 10 illustrates an example of a mapping of flexible rendering data to bitstream elements. [Figure 8] 10 illustrates an example of echo reference signaling when playing audio and listening commands with multiple devices. [Figure 9] 1 shows an example of how frames, blocks, and packets relate to each other. [Figure 10] This shows an example of additional codec information in the current MPEG-4 structure. [Figure 11] This represents an example of further codec information in the current MPEG-4 structure. [Figure 12] This shows an example of integrating listening ability and speech recognition. [Figure 13] 1 illustrates an example of a frame containing multiple blocks. [Figure 14] 1 illustrates an example of a bitstream containing multiple blocks. [Figure 15] 1 shows an example of blocks with different priorities. [Figure 16] 1 illustrates an example of a bitstream containing multiple blocks with different priorities. [Figure 17] 1 shows an example of frames with different priorities. [Figure 18] 1 shows an example of a bitstream containing frames with different priorities. DETAILED DESCRIPTION OF THE INVENTION
[0022] The principles of the present disclosure will now be described with reference to various examples illustrated in the drawings. Of course, the description of these examples is merely to enable those skilled in the art to better understand and further practice the present invention, and is not intended to limit the scope of the present disclosure in any way.
[0023] 1, an immersive audio stream is streamed from a cloud or server 10 and decoded at a TV or hub device 20. The immersive audio stream may be coded in any existing format, including, for example, Dolby Digital Plus, AC-4, etc. The output is then transcoded into a low-latency interchange format for further transmission to a connected device 30, preferably connected via a local wireless connection, for example, a WiFi soft access point or a Bluetooth connection. Low latency typically depends on various factors, such as, for example, frame size, sampling rate, hardware and / or software computational resources, etc., but low latency is typically less than 40 ms, 20 ms, or 10 ms.
[0024] In another typical use case, a phone fetches an immersive audio stream from a cloud or server, transcodes it to an interchange format, and then transmits it to a connected car. In the example depicted in Figure 2, a mobile device (e.g., a phone or tablet) 20 connects to a server 10, receives an immersive audio stream, performs low-latency transcoding to a low-latency interchange format, and sends the transcoded signal to a car 30 that supports immersive audio playback. An example of an immersive audio stream is a stream containing audio in Dolby Atmos format, and an example of a car 30 that supports immersive audio playback is a car configured to play the Dolby Atmos immersive format.
[0025] In general, the interchange format preferably has low delay, low encoding and decoding complexity, the ability to scale to high quality, and reasonable coding efficiency. The format also preferably supports configurable delay so that delay can be traded off for efficiency and even error resilience so as to be operable under a variety of connection conditions.
[0026] [Hub or extended display with wireless speakers (with or without listening capabilities)] 3 shows an example where a hub 20 driving a set of wireless speakers 30, or possibly a display (e.g., a television or TV) with built-in speakers 30, is extended by several wireless speakers. Extended implies that in examples where the display 20 includes speakers 30, the display 20 is also part of the audio playback. The wireless speakers / devices 30 may receive the same complete signal (broadcast mode, shown on the left side of FIG. 3) or individual streams tailored to specific devices (unicast-multipoint mode, shown on the right side of FIG. 3).
[0027] Each speaker may include multiple drivers covering different frequency ranges of the same channel or corresponding to different channels in a canonical mapping. For example, a speaker may have two drivers, one of which may be an upward-firing driver that outputs a signal corresponding to a height channel to emulate a speaker installed at a high altitude. Wireless devices may have listening capabilities (e.g., "smart speakers") and therefore may require echo management, which may require the speaker to receive one or more echo references. The echo references may be local (e.g., for the same speaker / device) or, alternatively, may represent relevant signals from other nearby speakers / devices.
[0028] The speakers may be located anywhere, in which case so-called "flexible rendering" may be performed, whereby the rendering takes into account the actual location of the speakers (e.g., as opposed to the standard assumption that speakers are located in fixed, predetermined locations). Flexible rendering may be performed at the hub or TV, and the rendered signals are then sent to the speakers / devices either in a broadcast mode, where each of the rendered signals is sent to each device and each device extracts and outputs the appropriate signal, or as individual streams to each device. Alternatively, flexible rendering may be performed locally at each device, whereby each device receives a representation of the complete immersive program, e.g., a 7.1.4 channel-based immersive representation, and renders from that representation an output signal appropriate for each device.
[0029] Wireless devices may be connected to the hub or TV through a soft access point provided by the device or through a local access point in the home, which may impose different requirements on bit rate and latency.
[0030] [Mobile device projection in automotive applications] In another use case, a mobile device (e.g., a phone or tablet) 20 fetches an immersive audio stream from the cloud or server 10, transcodes the immersive audio stream into an interchange format, and then sends the transcoded stream to a connected device. In the example of FIG. 2, a Dolby Atmos-enabled phone 20 connects to the server 10 and receives a Dolby Atmos stream for low-latency transcoding and transmission to a Dolby Atmos-enabled car 30. In this use case, higher bit rates may be available than in the living room use case, and different wireless channel characteristics may also be present. The higher bit rate in the car use case may be due to the less noisy wireless environment, as the car acts as a kind of Faraday cage, protecting the interior environment from external wireless interference, compared to the living room, where any wireless devices from neighbors or other rooms add noise to the wireless environment. At the same time, in a car, there are typically very few wireless devices competing for the available radio bandwidth, whereas in a living room use case, all wireless devices in the home, including the living room wireless devices, are typically competing for the available bandwidth.
[0031] For such use cases, it is envisioned that the signal transcoded by the mobile device and sent to the car may be a channel-based immersive representation, an object-based representation, a scene-based representation (e.g., and Ambisonics representation), or a combination of different representations. In this example, different rendering architectures and broadcast vs. multipoint may be irrelevant, since the complete presentation is generally transferred from the mobile device to a single endpoint (e.g., the car).
[0032] [Exchange format description] The immersive interchange format is built on the Modified Discrete Cosine Transform (MDCT) with perceptually based quantization and coding. The format has configurable delay (e.g., support for different transform sizes at a given sampling rate). Exemplary frame sizes are 128, 256, 512, 1024, 120, 240, 480, 960, 192, 384, and 768 samples at sampling rates of 48 kHz and 44.1 kHz.
[0033] The format may support other channel configurations, including mono, stereo, 5.1, and immersive channel configurations (e.g., including, but not limited to, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, and 22.1, or any other channel configuration consisting of unique channels defined in ISO / IEC 23091-3:2018, Table-2). The format may also support object-based audio and scene-based representations, such as Ambisonics (first order or higher). The format may also use a signaling scheme that is suitable for integration with existing formats (e.g., those described in ISO / IEC 14496-3, also known as the MPEG-4 Audio Standard, ISO 14496-1, which may also be referred to as the MPEG-4 Systems Standard, ISO 14496-12, which may also be referred to as the ISO Base Media File Format Standard, and / or ISO 14496-14, which may also be referred to as the MP4 File Format Standard).
[0034] Additionally, the system may include support for skipping some syntax elements and decoding only the relevant portions for a particular speaker, support for flex rendering aspects of metadata control such as delay adjustment, level adjustment, equalization, etc., support for using smart speakers with listening capabilities (including scenarios where a set of signals is transmitted that allow each driver / speaker in a smart speaker to be fed independently) and associated support for echo management through echo reference signaling, support for quantization and coding in the MDCT domain with both low overlap windows and 50% overlap windows, and / or support for filtering along the frequency axis in the MDCT domain for temporal shaping of quantization noise, e.g., TNS (Temporal Noise Shaping). Thus, in various examples, the overlap windows may be symmetric or asymmetric.
[0035] The disclosed interchange format may, in some examples, provide improved joint channel coding by utilizing increased correlation between signals in flexible rendering use cases, improved coding efficiency due to scale factors shared across channel elements, incorporation of high frequency reconstruction and noise addition techniques in the MDCT domain, and various improvements relative to conventional coding structures to improve efficiency and suitability for the use cases at hand.
[0036] In some instances of controlling individual playback devices in broadcast mode, skippable blocks and metadata may be used. In a broadcast mode configuration, each wireless device needs to play the relevant portions of the complete audio for the specific driver / channel corresponding to the specific device. This implies that the device needs to know which portions of the stream are relevant to that device and extract those portions from the complete audio stream that was broadcast. To enable low-complexity decoding operations, it is desirable if the stream is constructed in such a way that a decoder for a specific device can efficiently skip over elements that are irrelevant to decoding to elements that are relevant to that device and a given driver on that device.
[0037] However, it should be noted that it may still be advantageous (from a compression perspective) to allow joint coding between signals that are defined for different devices (e.g., in the simplest scenario, the left and right channels of a stereo representation).
[0038] Thus, the format can enable "skippable blocks" in the bitstream to allow efficient decoding of portions that are only relevant to a particular device, and can include metadata that allows flexible mapping of one or more skippable blocks to specific devices while preserving the ability to apply joint coding techniques between signals corresponding to different speakers / devices.
[0039] 4 shows an example of an arbitrary setup. There are three connected devices 31, 32, 33, two of which are single-channel speakers 32, 33 (i.e., one audio channel) and one speaker 31 is a more advanced speaker 31 that has three different drivers and thereby operates on three different signals. The first speaker 31 operates on three individual signals in this example, while the second speaker 32 and the third speaker 33 operate on a stereo representation (e.g., left and right). It may therefore be advantageous to perform joint coding of the signals of the stereo representation.
[0040] Taking the above scenario into account, the format specifies a bitstream with skippable blocks, allowing a first device to extract only the relevant part of the stream to decode the signal for that speaker, while in the case of a stereo pair, it has the advantage of being able to do joint coding, building a signal where each speaker needs to decode two signals to output one signal, thus trading off decoder complexity and efficiency.
[0041] The format specifies a metadata format that enables a general and flexible representation of the mapping of a particular block of a plurality of skippable blocks to one or more devices. This is depicted in Figure 5, where the mapping may be expressed as a matrix that associates each device 31, 32, 33 with one or more bitstream elements, so that a decoder for a given device knows which bitstream elements to output and decode.
[0042] 5, the first block or skip block Blk1 contains three single-channel elements (one for each driver in Device 1 31), while the second block or skip block Blk2 contains channel pair elements, which may include jointly coded versions of the signals output by Device 2 32 and Device 3 33. Device 1 31 retrieves the mapping metadata and determines that the signal it needs is Skip Block 1 Blk1. Accordingly, it retrieves Skip Block 1 Blk1, decodes the three single-channel elements therein, and provides them to Drivers 1, 2a, and 2b.
[0043] Further, Device 1 31 ignores Skip Block 2 Blk2. Similarly, Device 2 32 retrieves the mapping metadata and determines that the signal it needs is Skip Block 2 Blk2. Therefore, Device 2 32 skips Skip Block 1 Blk1 and retrieves Skip Block 2 Blk2. Device 2 32 decodes the channel pair element and provides the CPE's left channel output to its driver. Similarly, Device 3 33 retrieves the mapping metadata and determines that the signal it needs is in Skip Block 2 Blk2. Therefore, Device 3 33 skips Skip Block 1 Blk1 and also retrieves Skip Block 2 Blk2. Device 3 33 decodes the channel pair element and provides the CPE's right channel output to its driver.
[0044] In some examples, device 2 32 may determine that it requires only a subset of the signals from skip block 2 Blk2. In such examples, device 2 32 may perform only a subset of the operations necessary to fully decode the signals in skip block 2 Blk2, if possible. Specifically, in the example of FIG. 5, device 2 32 may reduce computational complexity because it only needs to perform the processing operations necessary to extract the left channel of the CPE. Similarly, device 3 33 may only need to perform the processing operations necessary to extract the right channel of the CPE.
[0045] For example, in the case of Figure 5, there may be cases where the CPEs are coded using joint channel coding and other cases where they are coded using independent channel coding. When the CPEs are coded using independent channel coding, device 2 32 only needs to extract the first (e.g., left) channel of the CPE, while device 3 33 only needs to extract the second (e.g., right) channel of the CPE.
[0046] In another example, the CPE's channels may be coded using joint channel coding, in which case Device 2 32 and Device 3 33 should extract the CPE's two intermediate channels. However, Device 2 32 can still operate with reduced computational complexity by performing only the operations necessary to extract the left channel from the intermediate channels. Similarly, Device 3 33 can still operate with reduced computational complexity by performing only the operations necessary to extract the right channel from the intermediate decoded channels.
[0047] Depending on the specifications of the joint channel coding, other optimizations may be possible. The identities of each of the different devices / decoders may be defined during a system initialization or setup phase. Such setup is common and typically involves measuring the room acoustics, speaker distance to sweet spot, etc.
[0048] [Flexible rendering (pre-encoding and post-decoding) distribution and post-decoder application such as delay] For use cases in which flexible rendering is applied at the hub / TV before sending the associated signal to a particular device / speaker, the rendering may produce a signal that is more difficult to code from a coding perspective (e.g., when considering joint coding of such a signal). One reason is that flexible rendering may apply different delay, equalization, and / or gain adjustments to different devices (e.g., depending on the placement of a speaker relative to other speakers and the listener). It is also possible to preset, for example, gain and delay information at initial setup and flexibly render only equalization. Other variations of preset information and flexible rendering information may also be possible in other examples. It should be noted that, in this specification, the term "gain" should be interpreted to mean any level adjustment (e.g., attenuation, amplification, or pass-through) rather than being limited to only a specific level adjustment (e.g., amplification).
[0049] 6, right channel speaker 33 and left channel speaker 32 are given different delays to reflect their different placement relative to the listener (e.g., because speakers 32, 33 may not be equidistant to the listener, different delays may be applied to the signals output from speakers 32, 33 so that coherent sounds from the different speakers 32, 33 arrive at the listener simultaneously). The introduction of such different delays to coherent signals intended for playback by different speakers 31, 32, 33 makes joint coding of such signals difficult.
[0050] To address these challenges, aspects of the flexible rendering process can be parameterized and applied at the endpoint device after decoding of the signal.
[0051] In the example of Figure 7, delay and gain values for each device are parameterized and included in the decoded signal sent to each speaker 31, 32, 33. Each signal is decoded by each device 31, 32, 33, and then the parameterized gain and delay values may be introduced into each decoded signal.
[0052] 7, if coded signals for different devices 31, 32, 33 are sent in separable blocks (e.g., in skippable blocks), parameters (e.g., delays and gains) may also be transmitted in separable blocks, so that a device 31, 32, 33 can extract only a subset of parameters that are necessary for that device 31, 32, 33 and ignore (skip) parameters that are not necessary for that device 31, 32, 33. In such a case, mapping metadata may be provided to each device indicating which parameters are included in which blocks.
[0053] It should also be noted that while Figure 7 only shows delay and gain parameters, other parameters may also be included, such as equalization parameters, which may include, for example, multiple gains applied to different frequency regions, an indication of a predetermined equalization curve to be applied by the playback devices 31, 32, 33, one or more sets of infinite impulse response (IIR) or finite impulse response (FIR) filter coefficients, a set of biquad filter coefficients, parameters specifying the characteristics of a parametric equalizer, and other parameters specifying equalization known to those skilled in the art.
[0054] Furthermore, the parameterization of aspects of flexible rendering need not be static but can be dynamic (e.g., if a listener moves during playback of an audio program). Thus, it may be desirable to allow parameters to change dynamically. If one or more parameters change during an audio program, the device may interpolate between the previous delay and / or gain parameters and the updated delay and / or gain parameters to provide a smooth transition. This can be particularly useful in situations where the system dynamically tracks the listener's position and updates the sweet spot for dynamic rendering accordingly.
[0055] As already mentioned and as will be discussed below, applying flexible rendering can lead to higher levels of correlation between channels, which can be exploited by more flexible joint coding.
[0056] [Echo Reference Coding and Signaling] For the use case illustrated in FIG. 8, the need for echo management arises when there are multiple devices / speakers 30 receiving signals and operating together in either broadcast mode, shown on the left of FIG. 8, or unicast-multipoint mode, shown on the right of FIG. 8, and at the same time, the devices 30 have microphones 40 that enable the "listen" function.
[0057] When performing echo management with multiple speakers / devices, it may be beneficial to use echo references from other devices in addition to the local speaker device. For example, if one device is placed close to another, the signals from the nearby device will affect the echo management of the device with the active microphone. In broadcast mode use cases, each device receives signals from all devices. If a device has signals for other devices, it may be beneficial to use those signals as echo references. To do this, a specific device must be informed of which signals it can use as echo references for which other devices.
[0058] As an example, this can be done by providing metadata that not only maps the channel / signal to be played by a particular speaker / device (from the entire set), but also the channel / signal to be used as an echo reference (from the entire set) for a particular speaker / device. Such metadata or signaling can be dynamic, allowing the indication of a preferred echo reference to change over time.
[0059] For use cases where each device / speaker receives only the specific signal it is to play, it may be necessary to send an additional signal (e.g., an echo reference signal) to each device / speaker to provide the appropriate echo reference. Again, to do so, it is necessary to provide device-specific signaling so that each device can select its appropriate signal for playback and the appropriate signal for echo management.
[0060] Because the signal for echo management is used only by the device performing the echo management and is not reproduced by the device for a listener, the echo reference signal may be coded or represented differently from the signal intended for reproduction by the device. In particular, echo management can be successful with signals coded at a lower rate than would normally be used for reproduction to a listener, so that additional compression tools, such as a parametric representation of the signal that may not be suitable for reproduction to a listener at all, but that captures the necessary characteristics of the audio signal, can provide good echo management while significantly reducing transmission costs.
[0061] [Related syntax elements] In some examples, blocks are used to optimize audio transmission. In the described format, each frame may be divided into blocks, as described above with respect to skippable blocks. A block may be identified by the frame number to which it belongs, a block ID that can be used to associate consecutive blocks from different frames with the same ID into a block stream, and a retransmission priority. An example of the above is a stream including multiple frames N-2, N-1, and N, where frame N includes multiple blocks identified as ID1, ID2, and ID3, as shown in FIG. 13. An example of this stream, albeit in bitstream format, is shown in FIG. 14. In some examples, for audio signals of an immersive audio program, the block ID of a block may indicate which signal set of the entire immersive audio program is carried by that block.
[0062] The use case for the described format is the reliable transmission of audio over wireless networks such as Wi-Fi with low latency. For example, Wi-Fi uses a packet-based network protocol. Packet sizes are usually limited; a typical maximum packet size in IP networks is 1500 bytes. The stream's block-based architecture allows for greater flexibility in assembling packets for transmission. For example, packets containing smaller frames can be filled with retransmitted blocks from other frames. Larger frames can be split at block boundaries before packetizing them to reduce inter-packet dependencies at the network protocol layer.
[0063] Figure 9 illustrates the relationship between frames, blocks, and packets. A frame carries audio data, preferably all audio data, representing a contiguous segment of an audio signal with a start time, an end time, and a duration that is the difference between the end time and the start time. Contiguous segments may have durations in accordance with ISO / IEC 14496-3, subpart 4, section 4.5.2.1.1. Section 4.5.2.1.1 describes the contents of raw_data_block(). A frame may also carry a redundant representation of that segment, for example, encoded at a lower data rate. After encoding, the frame may be divided into blocks. The blocks may be combined into packets for transmission over a packet-based network. Blocks from different frames may be combined into a single packet and / or transmitted out of order.
[0064] As an example, blocks are used to address individual devices. Packets of data are received by individual devices or related device groups. The concept of skippable blocks may also be used to address individual devices or related device groups. Even if the network is operating in broadcast mode when sending packets to different devices, the audio processing (decoding, rendering, etc.) may be reduced to the block addressed to that device. All other blocks may simply be skipped, even if received within the same packet. In some examples, blocks may be ordered correctly based on their decoding or presentation time. Retransmitted blocks with lower priority may be dropped if the same block with higher priority is also received. The stream of blocks may then be provided to the decoder.
[0065] In some examples, stream and device configurations are transmitted out-of-band. Codecs allow for setting up connections over which audio streams are transmitted at a relatively high rate but with low latency. The configuration for such a connection can remain stable for the duration of such a connection. In that case, the configuration can be transmitted out-of-band rather than being part of the audio stream. Such out-of-band transmissions can also use different networks or network protocols. For example, the audio stream may use User Datagram Protocol (UDP) for low-latency transmission, while the configuration can use Transmission Control Protocol (TCP) to ensure reliable transmission of the configuration.
[0066] [Enabling codecs in the context of MPEG-4 audio] One particular application of this technology is within MPEG-4 Audio, where different Audio Object Types (AOTs) are defined for different codec technologies. The format described herein defines a new AOT, which allows for format-specific signaling and data. Furthermore, MPEG-4 decoder configuration is performed in the DecoderSpecificInfo() payload, which carries the AudioSpecificConfig() payload. The latter defines certain general signaling that is independent of a particular format, such as sampling rate, channel configuration, and specific information for a particular AOT. This may make sense in traditional formats where the entire stream is decoded by a single device. However, in broadcast mode, where a single stream is sent to multiple decoders, each of which decodes only a portion of the stream, advance channel configuration signaling (as a means of setting the device's output capabilities) may not be optimal.
[0067] Figure 10 shows the traditional MPEG-4 high-level structure (in black), with modifications highlighted in grey. To support broadcast use cases, codecSpecificConfig() (where "codec" is a generic placeholder name) is defined. This function redefines the signaling to suit a specific use case, allowing for mapping specific channel elements to specific devices and including other relevant static parameters. The MPEG-4 element channelConfiguration with value "0" is defined as the channel configuration defined in codecSpecificConfig. This value can therefore be used to enable modification of the signaling of the channel configuration within the codec-specific configuration.
[0068] Furthermore, in keeping with the spirit of MPEG-4, the raw payload is specified for the particular decoder at hand, assuming that codecSpecificConfig() is capable of decoding it. However, the format ensures that dynamic metadata is part of the raw payload and ensures that length information is available for all raw payloads, so that decoders can easily skip over material that is not relevant to a particular device.
[0069] In Figure 11, a portion of an example raw_data_block as defined in MPEG-4 is given on the right. The raw block data contains channel elements (single channel elements (SCE) or channel pair elements (CPE)) in a given order. However, a decoder wishing to skip over portions of these channel elements that may not be relevant to the output device at hand would, in the conventional MPEG-4 Audio syntax, be required to parse (and to some extent decode) all channel elements to be able to extract the relevant portions. In the new raw_data_block shown on the left of Figure 11, the content is organized into skippable blocks so that the decoder can skip over the irrelevant portions and decode only those channel elements indicated by the metadata as being relevant to the device at hand. As an example, the skippable block contains the raw_data_block and related information.
[0070] It is also possible to use blocks for retransmission. Each frame can be divided into blocks, as described in the skippable blocks section above. A block can be identified by the frame number to which it belongs, a block ID that can be used to associate consecutive blocks from different frames with the same ID into a block stream, and a retransmission priority (see Figures 15 and 16). For example, a high retransmission priority (e.g., shown as priority 0 in Figures 15 and 16) indicates that the block will be prioritized at the receiver over another block with the same block ID and frame counter but a lower retransmission priority (e.g., shown as priority 1 in Figures 15 and 16). Note that, as in the examples of Figures 15 and 16, increasing the priority index may decrease the priority (e.g., priority 1 is lower than priority 0), while in other examples, increasing the priority index may increase the priority (e.g., priority 0 is lower than priority 1). Further examples will be apparent to those skilled in the art.
[0071] The syntax may support retransmission of audio elements. For retransmission, different quality levels may be supported. Therefore, retransmitted blocks may carry a "priority" flag to indicate which block with the same block ID should be prioritized to the decoder. This is because blocks received with the same frame counter and block ID are redundant and therefore mutually exclusive to the decoder (see Figures 17 and 18).
[0072] Retransmission of the blocks may be performed at a reduced data rate. Such data rate reduction may be achieved by reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, reducing the number of channels in the audio signal (e.g., as described in U.S. Pat. No. 1,289,103, which is incorporated by reference in its entirety), or any combination thereof. In order for the decoder to select audio blocks that provide the best possible quality, the block providing the highest quality signal may be given the highest decoding priority, the block providing the second highest quality signal may be given the second highest priority, and so on.
[0073] Blocks may also be retransmitted at the same quality level, in which case the priority may reflect the delay of the retransmitted block.
[0074] [Tools to improve core coding efficiency] In a further example, MDCT quantization scale factors can be shared between channel elements. In some cases, it may be possible to share MDCT scale factors between certain channels to reduce the associated side bit rate. Furthermore, scale factor sharing can be extended to different channel elements, as in the case of 7.1.4 input, for example. The use of shared scale factors can be indicated by two signaling bits that allow three active configurations. One possible configuration is to share scale factors in the left horizontal channel, the top-left channel, the right horizontal channel, and the top-right channel. The specific syntax can be tailored to the concept of skippable blocks, so that scale factors are only shared within skip blocks.
[0075] Joint coding of more than two channels is also possible. In the example shown in Figures 4, 5, 6, and 7, block Blk1 carries three SCEs, one for each driver of the smart speaker considered in this scenario. In such a situation, it can be beneficial to enable joint coding of more than two channels, not just two channels (provided by the CPE and using stereo prediction, e.g., the MDCT-based complex predictive stereo tool described in ISO / IEC 23003-3, which can be extended to include the so-called MPEG-D USAC), as described, for example, in the Stereo Audio Processing (SAP) tool introduced in ETSI TS 103 190.
[0076] Similarly, for flexible rendering, it should be noted that there may be high correlation between signals between devices. For such signals, it may be beneficial to configure skip blocks that cover multiple devices and allow joint channel coding to be applied to more than two channels using the tools described above.
[0077] Another joint channel coding tool that can be used for a particular transform length is channel coupling, in which composite channel and scale factor information is transmitted at medium and high frequencies, allowing for bitrate reduction for playback in the good quality range. For example, for low-delay coding, it may be beneficial to use the channel coupling tool for frames with a transform length of 256, which corresponds to a frame length of approximately 256 samples.
[0078] The present disclosure also enables efficient coding of band-limited signals. In scenarios where separate signals are fed to different drivers of a smart speaker, some of these signals may be band-limited, as in the case of a three-way driver configuration with a woofer, midrange, and tweeter. Efficient encoding of such band-limited signals is therefore desirable, which may result in specially tailored psychoacoustic models and bit allocation strategies, as well as potential modifications to the syntax to handle such scenarios, resulting in improved coding efficiency and / or reduced computational complexity (e.g., allowing the use of a band-limited IMDCT for the woofer feed).
[0079] In some cases, a combination of high-frequency reconstruction and noise addition in the MDCT domain is also possible. Modern audio codecs are typically designed to support parametric coding techniques, and similarly, here we consider incorporating high-frequency reconstruction methods in the MDCT domain to maintain low latency, or noise addition schemes in the MDCT domain to manage the tone-to-noise ratio of the reconstructed high-bandwidth. Such parametric coding techniques can be particularly useful at low operating points, especially in scenarios where retransmissions occur as part of a forward error correction (FEC) scheme, where a main signal is typically transmitted followed by a delayed retransmission of the same signal at a lower bit rate.
[0080] These examples allow further reduction of the peak complexity of the encoder. This is applicable when there is no buffer model limitation for a constant bit rate transport channel. In buffered constant bit rate mode, after quantization and counting, the resulting number of bits used for the frame being encoded may exceed the limit of the number of allowed bits. To meet the bit requirement, at least one new, coarser quantization and bit counting step must be performed. If this limit is relaxed, the encoder can retain the initial quantization results and run at a slightly higher instantaneous bit rate. Nevertheless, by using the same bit reservoir control mechanism and a virtual buffer fullness that is updated according to the encoder's compliance with the buffer requirement, encoder operation similar to that of a buffered constant bit rate model encoder can be achieved. This saves additional quantization and bit counting steps, and the resulting audio quality is the same or better than that of the constant bit rate case, at the expense of a slight increase in overall bit rate.
[0081] [Return channel concept] For play and listen use cases, the smart speaker 30 may have one or more microphones 40 (e.g., microphones or microphone arrays) configured to capture a mono or spatial sound field (e.g., in mono, stereo, A-format, B-format, or other isotropic or anisotropic channel format), and the codec must be capable of efficient coding of such formats in a low-latency manner. Thus, the same codec used for the return channel is used for broadcast / transmission to the smart speaker device 30, as shown, for example, in FIG. 12.
[0082] It should also be noted that when wake word detection is performed for a smart speaker device, speech recognition is typically performed in the cloud using an appropriate segment of recorded audio triggered by the wake word detection.
[0083] In the present context, it may be interesting to reduce the overall system complexity by creating a speech analysis system based on an intermediate format / representation, for example within a codec. In the simplest use case, where no human-to-human conversation is ongoing, one can define a suitable representation specific to the speech recognition task, since this is not what another human would hear. This could be band energies, Mel-Frequency Cepstral Coefficients (MFCCs), etc., or a low-bitrate version of the MDCT-coded spectrum.
[0084] In a use case where a human conversation is ongoing and the wake-up word alternates with the required speech recognition, one would want to be able to extract the relevant part of the stream for this. Such extraction may require a layered coding structure with just additional data sent in parallel to the main audio signal. We also envision a predefined transcoding of the existing decoded MDCT spectrum to the relevant representation, which would simplify the structure. Essentially, the decoder decodes and outputs human-audible audio until it receives a signal that it also needs to decode (in parallel) the speech recognition representation. This representation may be "peeled" from the layered stream, or it may just be an alternative decoding of the same complete data, or just the decoding and output of an additional representation in the stream. In this use case, the signaling is an enabling part that indicates that the receiving decoder should output the representation relevant for speech recognition.
[0085] [Enumerated examples] Below, seven sets of enumerated embodiments (EEE-A, EEE-B, EEE-C, EEE-D, EEE-E, EEE-F, and EEE-G), which are not claims, describe aspects of the embodiments disclosed herein.
[0086] EEE-A1. A method for decoding an audio signal, comprising: receiving a bitstream including at least one frame, each frame of the at least one frame including a plurality of blocks; determining from the signaling data information identifying portions of one or more blocks of the plurality of blocks to be skipped during decoding based on device information of an output device; decoding the bitstream while skipping the identified portion of the one or more blocks; A method having the following.
[0087] EEE-A2. The method of EEE-A1, the information identifying portions of one or more blocks among the plurality of blocks to be skipped during decoding includes a matrix associating each output device among a plurality of output devices with one or more bitstream elements. method.
[0088] EEE-A3. The method of EEE-A2, comprising: the one or more bitstream elements are required for decoding the bitstream for a corresponding associated output device; method.
[0089] EEE-A4. Any of the methods EEE-A1 to EEE-A3, The output device may include at least one of a wireless device, a mobile device, a tablet, a single-channel speaker, and / or a multi-channel speaker; method.
[0090] EEE-A5. Any of the methods EEE-A1 to EEE-A4, the identified portion includes at least one block; method.
[0091] EEE-A6. Any of the methods EEE-A1 to EEE-A5, The output device is a first output device, and the method further comprises applying a joint coding technique between one or more signals of the bitstream to a second output device and a third output device.
[0092] EEE-A7. Any of the methods EEE-A1 to EEE-A6, The ID of each output device and / or decoder is defined during the system initialization phase, method.
[0093] EEE-A8. Any of the methods EEE-A1 to EEE-A7, the signaling data is determined from metadata of the bitstream; method.
[0094] EEE-A9. An apparatus configured to perform any one of the methods of EEE-A1 to EEE-A8.
[0095] EEE-A10. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform any one of the methods EEE-A1 through EEE-A8.
[0096] EEE-B1. A method for generating an encoded bitstream from an audio program including a plurality of audio signals, comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; receiving, for each playback device, information indicating at least one of a delay, a gain, and an equalization curve associated with each playback device; determining a group of two or more related audio signals from the plurality of audio signals; applying one or more joint coding tools to the two or more related audio signals of the group to obtain a jointly coded audio signal; combining the jointly coded audio signal, an indication of the playback devices with which the jointly coded audio signal is associated, and an indication of delays and gains associated with each playback device with which the jointly coded audio signal is associated into separate blocks of an encoded bitstream; A method having the following.
[0097] EEE-B2. The method of EEE-B1, the delay, gain, and / or equalization curve associated with each reproduction device depends on the position of said each reproduction device relative to the position of the listener; method.
[0098] EEE-B3. The method of EEE-B1 or EEE-B2, the delay, gain, and / or equalization curve associated with each reproduction device depends on the position of each reproduction device relative to the positions of the other reproduction devices; method.
[0099] Any one of the methods EEE-B4, EEE-B1 to EEE-B3, The delay, gain, and / or equalization curves are dynamically changed; method.
[0100] EEE-B5. A method according to EEE-B4, comprising: The delay, gain, and / or equalization curves are adjusted in response to changes in the listener's position; method.
[0101] The method of EEE-B6, EEE-B4 or EEE-B5, The delay, gain, and / or equalization curves are adjusted in response to changes in the position of the playback device; method.
[0102] Any one of the methods EEE-B7, EEE-B4 to EEE-B6, The delay, gain, and / or equalization curves are adjusted in response to changes in the position of one or more of the other playback devices; method.
[0103] Any one of the methods EEE-B8, EEE-B1 to EEE-B7, The method further comprising determining from the plurality of audio signals audio signals that are not part of the group of two or more related audio signals.
[0104] EEE-B9. A method according to EEE-B8, comprising: The method further comprising applying, for an audio signal that is not part of the group of two or more related audio signals, a delay, gain, and / or equalization curve associated with a playback device with which the audio signal is associated.
[0105] A method according to EEE-B10.EEE-B9, comprising: The method further comprising independently coding the audio signals that are not part of the group of two or more related audio signals, and combining the independently coded audio signals and an indication of a playback device with which the independently coded audio signals are associated into a separate, independently decodable subset of the coded bitstream.
[0106] A method according to EEE-B11.EEE-B8, comprising: independently coding the audio signals that are not part of the group of two or more related audio signals, and combining the independently coded audio signals, an indication of a playback device with which the independently coded audio signals are associated, and an indication of a delay, gain, and / or equalization curve associated with a playback device with which the independently coded audio signals are associated, into a separate, independently decodable subset of the coded bitstream.
[0107] EEE-B12. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, wherein the frames comprise one or more independent blocks of encoded data, comprising: identifying from the encoded bitstream discrete blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; determining that the retrieved independent blocks of encoded data comprise two or more jointly coded audio signals; applying one or more joint decoding tools to the two or more jointly coded audio signals to obtain the one or more audio signals associated with the playback device; determining at least one of a delay, a gain, and an equalization curve associated with the playback device from the retrieved discrete blocks of encoded data; applying the delay, the gain, and / or the equalization curve associated with the playback device to the one or more audio signals associated with the playback device; A method having the following.
[0108] EEE-B13. The method of EEE-B12, comprising: the determined delay, gain, and / or equalization curve associated with the playback device is dependent on the position of the playback device relative to the position of a listener; method.
[0109] EEE-B14. The method of EEE-B12 or EEE-B13, the determined delay, gain, and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices; method.
[0110] Any one of the methods EEE-B15, EEE-B12 to EEE-B14, the determined delay, gain, and / or equalization curve of the playback device is dynamically changed; method.
[0111] EEE-B16. The method of EEE-B15, the determined delay, gain, and / or equalization curve associated with the playback device is different from a previously determined delay, gain, and / or equalization curve associated with the playback device, the method further comprising interpolating between the previously determined delay, gain, and / or equalization curve associated with the playback device and the determined delay, gain, and / or equalization curve associated with the playback device. method.
[0112] EEE-B17. The method of EEE-B16, comprising: the determined delay, gain, and / or equalization curve differs from the previously determined delay, gain, and / or equalization curve due to a change in listener position; method.
[0113] EEE-B18. The method of EEE-B16 or EEE-B17, the determined delay, gain, and / or equalization curve differs from the previously determined delay, gain, and / or equalization curve due to a change in the position of the playback device; method.
[0114] Any one of the methods EEE-B19, EEE-B16 to EEE-B18, the determined delay, gain, and / or equalization curve differs from the previously determined delay, gain, and / or equalization curve due to a change in the position of one or more of the other playback devices; method.
[0115] Any one of the methods EEE-B20.EEE-B16 to EEE-B18, The frame of the encoded bitstream comprises two or more independent blocks of encoded data, and the method comprises: determining that one or more of the independent blocks includes an audio signal not associated with the playback device; ignoring the one or more independent blocks containing audio signals not associated with the playback device; and Further comprising: method.
[0116] Any one of the methods EEE-B21, EEE-B12 to EEE-B20, applying the one or more joint decoding tools includes identifying a subset of the jointly coded audio signals associated with the playback device and reconstructing only the subset of the jointly coded audio signals to obtain the one or more audio signals associated with the playback device. method.
[0117] Any one of the methods EEE-B22, EEE-B12 to EEE-B20, applying the one or more joint decoding tools includes reconstructing each of the jointly coded audio signals, identifying a subset of the reconstructed jointly coded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of the reconstructed jointly coded audio signals associated with the playback device. method.
[0118] EEE-B23. An apparatus configured to perform any one of the methods EEE-B1 to EEE-B22.
[0119] EEE-B24. A non-transitory computer readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform any one of the methods EEE-B1 through EEE-B22.
[0120] EEE-C1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, said frames including two or more independent blocks of encoded data, comprising: receiving, for one or more of the plurality of audio signals, information indicating a playback device with which the one or more audio signals are associated; receiving, for the indicated playback device, information indicating one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding one or more audio signals associated with the playback device; encoding one or more audio signals associated with the indicated one or more additional associated playback devices; and combining one or more encoded audio signals associated with the playback device and signaling information indicating the one or more additional associated playback devices into a first independent block; combining one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method having the following.
[0121] EEE-C2. The method of EEE-C1, the plurality of audio signals includes one or more audio signal groups that are not associated with the playback device or the one or more additional associated playback devices; encoding each of the one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices into respective independent blocks; combining each independent block of each of the one or more audio signals into the frame of the encoded bitstream; A method having the following.
[0122] EEE-C3. The method of EEE-C1 or EEE-C2, the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended for use as echo references for performing echo management of the playback devices; method.
[0123] EEE-C4. The method of EEE-C3, comprising: the one or more audio signals intended for use as echo references are transmitted using less data than the one or more audio signals associated with the playback device; method.
[0124] The method of EEE-C5, EEE-C3 or EEE-C4, the one or more audio signals intended for use as echo references are coded using a parametric coding tool; method.
[0125] EEE-C6. The method of EEE-C1 or EEE-C2, one or more audio signals associated with one or more other playback devices are suitable for playback from said one or more other playback devices; method.
[0126] EEE-C7. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising two or more independent blocks of the encoded bitstream, and the playback device having one or more microphones, identifying, from the encoded bitstream, discrete blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; extracting the one or more audio signals associated with the playback device from the identified discrete blocks of encoded data; identifying from the encoded bitstream one or more other independent blocks of encoded data corresponding to one or more audio signals associated with one or more other playback devices; extracting one or more audio signals associated with the one or more other playback devices from the one or more other independent blocks of encoded data; capturing one or more audio signals with the one or more microphones of the playback device; using, in response to the one or more captured audio signals, one or more retrieved audio signals associated with the one or more other playback devices as an echo reference for performing echo management for the playback device; A method having the following.
[0127] EEE-C8. The method of EEE-C7, comprising: determining that the encoded bitstream includes one or more additional independent blocks of encoded data; ignoring said one or more additional independent blocks of encoded data; and The method further comprises:
[0128] EEE-C9. The method of EEE-C8, comprising: ignoring the one or more additional independent blocks of encoded data includes skipping the one or more additional independent blocks of encoded data without retrieving the one or more additional independent blocks of encoded data. method.
[0129] Any one of the methods EEE-C10, EEE-C7 to EEE-C9, one or more audio signals associated with the one or more other playback devices are specifically intended for use as echo references for performing echo management for the playback devices; method.
[0130] EEE-C11. The method of EEE-C10, comprising: the one or more audio signals specifically intended for use as the echo reference are transmitted using less data than the one or more audio signals associated with the playback device; method.
[0131] EEE-C12. The method of EEE-C10 or EEE-C11, the one or more audio signals specifically intended for use as the echo reference are reconstructed from a parametric representation of the one or more audio signals. method.
[0132] Any one of the methods EEE-C13, EEE-C7 to EEE-C9, one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices; method.
[0133] The method of EEE-C14.EEE-C7, comprising: the encoded signal includes signaling information instructing the one or more other playback devices to use as an echo reference for the playback device. method.
[0134] EEE-C15. The method of EEE-C14, comprising: the one or more other playback devices indicated by the signaling information for the current frame are different from the one or more other playback devices used as echo references for the previous frame; method.
[0135] EEE-C16. An apparatus configured to perform any one of the methods EEE-C1 to EEE-C15.
[0136] EEE-C17. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform any one of the methods EEE-C1 to EEE-C15.
[0137] EEE-D1. A method for transmitting an audio signal, comprising: generating a packet of data comprising a portion of a bitstream, the bitstream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks, the generating including assembling the packet of data with one or more of the plurality of blocks, wherein blocks from different frames are combined into a single packet and / or transmitted out of order; transmitting the packets of data over a packet-based network; A method having the following.
[0138] EEE-D2. The method of EEE-D1, each block of the plurality of blocks includes identification information; method.
[0139] EEE-D3. The method of EEE-D2, the identification information includes at least one of a block ID, a corresponding frame number associated with the block, and / or a retransmission priority; method.
[0140] EEE-D4. Any of the methods EEE-D1 to EEE-D3, each frame of the plurality of frames carries all audio data representing a contiguous segment of the audio signal having a start time, an end time, and a duration; method.
[0141] EEE-D5. A method for decoding an audio signal, comprising: receiving a packet of data comprising a portion of a bitstream, the bitstream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks; determining a set of blocks of said plurality of blocks addressed to a device; decoding the set of blocks addressed to the device and skipping decoding of blocks of the plurality of blocks not addressed to the device; A method having the following.
[0142] EEE-D6. A method for transmitting an audio stream, comprising: transmitting the audio stream, the audio stream including a plurality of frames, each frame of the plurality of frames including a plurality of blocks, and the transmitting includes transmitting configuration information of the audio stream out-of-band. method.
[0143] EEE-D7. The method of EEE-D6, comprising: transmitting the configuration information of the audio stream out-of-band, transmitting the audio stream over a first network and / or a first network protocol; transmitting the configuration information over a second network and / or a second network protocol; having method.
[0144] EEE-D8. The method of EEE-D7, comprising: the first network protocol is User Datagram Protocol (UDP) and the second network protocol is Transmission Control Protocol (TCP); method.
[0145] EEE-D9. A method for decoding an audio signal, comprising: receiving a bitstream containing information corresponding to signaling of aspects of a static configuration and static metadata; mapping one or more channel elements to one or more devices based on said information and / or said static metadata; A method having the following.
[0146] EEE-D10. A method according to EEE-D9, comprising: the bitstream is received by a plurality of decoders configured to decode the bitstream, each decoder of the plurality of decoders configured to decode a portion of the bitstream; method.
[0147] The method of EEE-D11, EEE-D9 or EEE-D10, the bitstream further includes dynamic metadata; method.
[0148] Any of the methods EEE-D12.EEE-D9 to EEE-D11, The bitstream includes a plurality of blocks, each block of the plurality of blocks comprising: information enabling parts of said block not required by the device to be skipped during decoding; Dynamic Metadata and A method comprising:
[0149] EEE-D13. A method for retransmitting blocks of an audio signal, comprising: transmitting one or more blocks of a bitstream, the bitstream including a plurality of blocks, each of the one or more blocks of the bitstream having been previously transmitted; each of the one or more blocks includes a decoding priority indicator; method.
[0150] EEE-D14. The method of EEE-D13, comprising: the decoding priority indicator indicates to a decoder an order of priority for decoding the one or more blocks of the bitstream. method.
[0151] EEE-D15. The method of EEE-D13 or EEE-D14, each block of the one or more blocks has the same block ID; method.
[0152] EEE-D16. Any of the methods EEE-D13 to EEE-D15, wherein the transmission of the one or more blocks of the bitstream is transmitted at a reduced data rate compared to a previous transmission. method.
[0153] EEE-D17. The method of EEE-D16, comprising: reducing the data rate includes at least one of reducing a signal-to-noise ratio of the audio signal, reducing a bandwidth of the audio signal, and / or reducing a channel count of the audio signal. method.
[0154] EEE-E1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, said frames comprising one or more independent blocks of encoded data, comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; encoding one or more audio signals associated with each playback device to obtain one or more encoded audio signals; combining the one or more encoded audio signals associated with each of the playback devices into a first independent block of the frame; encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method having the following.
[0155] EEE-E2. A method according to EEE-E1, comprising: two or more audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and each of the band-limited signals employing a different encoding technique; method.
[0156] The method of EEE-E3.EEE-E2, a different psychoacoustic model and / or a different bit allocation technique is used for each of the bandlimited signals; method.
[0157] Any one of the methods EEE-E4, EEE-E1 to EEE-E3, the instantaneous frame rate of the encoded signal is variable and is constrained by a buffer fullness model; method.
[0158] A method according to any one of claims 1 to 5, further comprising: encoding the one or more audio signals associated with each playback device includes jointly encoding the one or more audio signals associated with each playback device and one or more additional audio signals associated with one or more additional playback devices into the first independent block of the frame. method.
[0159] The method of EEE-E6.EEE-E5, and jointly encoding the one or more audio signals and the one or more additional audio signals includes sharing one or more scale factors between two or more audio signals. method.
[0160] EEE-E7. The method of EEE-E6, the two or more audio signals are spatially related; method.
[0161] EEE-E8. The method of EEE-E7, the two or more spatially related audio signals include a left horizontal channel, an upper left channel, a right horizontal channel, or an upper right channel; method.
[0162] The method of EEE-E9.EEE-E5, Jointly encoding the one or more audio signals and the one or more additional audio signals includes: combining two or more audio signals into a composite signal above a specified frequency; determining, for each of the two or more audio signals, a scale factor related to the energy of the composite signal and the energy of each signal; applying a coupling tool having method.
[0163] The method of EEE-E10.EEE-E5, and jointly encoding the one or more audio signals and the one or more additional audio signals includes applying a joint coding tool to more than two signals. method.
[0164] EEE-E11. A method of decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising one or more independent blocks of encoded data, the method comprising: identifying, from the encoded bitstream, discrete blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; decoding one or more audio signals associated with the playback device from the discrete blocks of encoded data to obtain one or more decoded audio signals; identifying, from the encoded bitstream, one or more additional independent blocks of encoded data corresponding to one or more additional audio signals; decoding or skipping said one or more additional independent blocks of encoded data; A method having the following.
[0165] The method of EEE-E12.EEE-E11, two or more audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and different decoding techniques are used to decode the two or more audio signals; method.
[0166] EEE-E13. The method of EEE-E12, different psychoacoustic models and / or different bit allocation techniques are used to encode each of the bandlimited signals; method.
[0167] The method of EEE-E14, EEE-E11 or EEE-E13, the instantaneous frame rate of the encoded bitstream is variable and is constrained by a buffer fullness model; method.
[0168] The method of EEE-E15.EEE-E11, comprising: decoding the one or more audio signals associated with the playback devices includes jointly decoding the one or more audio signals associated with each playback device and one or more additional audio signals associated with one or more additional playback devices from the independent blocks of encoded data; method.
[0169] EEE-E16. The method of EEE-E15, and jointly decoding the one or more audio signals and the one or more additional audio signals includes deriving a scale factor shared between two or more audio signals. method.
[0170] EEE-E17. The method of EEE-E16, comprising: the two or more audio signals are spatially related; method.
[0171] EEE-E18. The method of EEE-E17, comprising: the two or more spatially related audio signals include a left horizontal channel, an upper left channel, a right horizontal channel, or an upper right channel; method.
[0172] The method of EEE-E19.EEE-E15, comprising: and jointly decoding the one or more audio signals and the one or more additional audio signals includes applying a decoupling tool. method.
[0173] A method according to EEE-E20.EEE-E19, comprising: The decoupling tool comprises: extracting independently decoded signals below a specified frequency; extracting a composite signal above the specified frequency; determining each decoupled signal above the specified frequency from the composite signal and a scale factor related to the energy of the composite signal and the energy of each signal; A method comprising:
[0174] A method according to EEE-E21.EEE-E15, comprising: and jointly decoding the one or more audio signals and the one or more additional audio signals includes applying a joint decoding tool to derive more than two audio signals. method.
[0175] 1. A method according to claim 8, further comprising: and decoding the one or more audio signals associated with the playback device includes applying a bandwidth extension to the audio signals in the same domain in which the audio signals were coded. method.
[0176] EEE-E23. The method of EEE-E22, the domain is a modified discrete cosine transform (MDCT) domain; method.
[0177] The method of EEE-E24, EEE-E22 or EEE-E23, said bandwidth extension including adaptive noise addition; method.
[0178] EEE-E25. An apparatus configured to perform any one of the methods EEE-E1 to EEE-E24.
[0179] EEE-E26. A non-transitory computer readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform any one of the methods EEE-E1 to EEE-E24.
[0180] EEE-F1. A method performed by a device equipped with one or more microphones to generate an encoded bitstream, comprising: capturing one or more audio signals with the one or more microphones; analyzing the captured audio signal to determine the presence of a wake word; Upon detecting the presence of the wake word, setting a flag to indicate that a speech recognition task should be performed on the captured audio signal; encoding the captured audio signal; assembling the encoded audio signal and the flag into the encoded bitstream; A method having the following.
[0181] EEE-F2. The method of EEE-F1, the one or more microphones are configured to capture a mono or spatial sound field; method.
[0182] The method of EEE-F3.EEE-F2, The spatial sound field is in A format or B format, method.
[0183] Any one of the methods EEE-F4, EEE-F1 to EEE-F3, the captured audio signal is intended only for use in performing the speech recognition task; method.
[0184] EEE-F5. The method of EEE-F4, the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient to perform the speech recognition task but insufficient for human listening; method.
[0185] Method EEE-F6, EEE-F4 or EEE-F5, the captured audio signal is converted into a representation comprising one or more of band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spatial coefficients before encoding the captured audio signal; method.
[0186] Any one of methods EEE-F7, EEE-F1 to EEE-F3, the captured audio signal is intended both for use in performing the speech recognition task and for human listening; method.
[0187] EEE-F8. The method of EEE-F7, comprising: the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening; method.
[0188] EEE-F9. The method of EEE-F7, comprising: encoding the captured audio signal includes generating a first encoded representation of the captured audio signal and a second encoded representation of the captured audio signal; the first coded representation is generated such that when the captured audio signal is decoded from the first coded representation, the quality of the decoded audio signal is sufficient for human hearing; the second coded representation is generated such that when the captured audio signal is decoded from the second coded representation, the quality of the decoded audio signal is sufficient for performing a speech recognition task but insufficient for human hearing. method.
[0189] The method of EEE-F10.EEE-F9, generating the second coded representation of the captured audio signal includes converting the captured audio signal into one or more of a parametric representation, a coarse waveform representation, or a representation including one or more of band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spatial coefficients before encoding the captured audio signal; method.
[0190] The method of EEE-F11, EEE-F9 or EEE-F10, assembling the encoded audio signal into a bitstream includes inserting the first coded representation into a first independent block of the encoded bitstream and inserting the second coded representation into a second independent block of the encoded bitstream. method.
[0191] The method of EEE-F12, EEE-F9 or EEE-F10, the first coded representation is included in a first layer of the coded bitstream, the second coded representation is included in a second layer of the coded bitstream, and the first layer and the second layer are included in a single block of the coded bitstream. method.
[0192] Any one of the methods EEE-F13.EEE-F1 to EEE-F12, If the presence of the wake word is not detected, setting a flag to indicate that a speech recognition task should not be performed on the captured audio signal; encoding the captured audio signal; assembling the encoded audio signal and the flag into the encoded bitstream; The method further comprises:
[0193] EEE-F14. A method for decoding an audio signal, comprising: receiving an encoded bitstream including an encoded audio signal and a flag indicating whether a speech recognition task should be performed; decoding the encoded audio signal to obtain a decoded audio signal; performing a speech recognition task on the decoded audio signal if the flag indicates that the speech recognition task should be performed. A method having the following.
[0194] EEE-F15. The method of EEE-F14, the decoded audio signal is intended only for use in performing the speech recognition task; method.
[0195] EEE-F16. The method of EEE-F15, the quality of the decoded audio signal is sufficient to perform the speech recognition task but insufficient for human listening; method.
[0196] The method of EEE-F17, EEE-F15 or EEE-F16, the decoded audio signal is in a representation comprising one or more of band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spatial coefficients. method.
[0197] EEE-F18. The method of EEE-F14, the encoded audio signal is encoded such that, when the encoded audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening; method.
[0198] EEE-F19. The method of EEE-F18, comprising: the encoded audio signals include a first encoded representation of one or more audio signals and a second encoded representation of the one or more audio signals; method.
[0199] The method of EEE-F20.EEE-F18, comprising: the quality of the audio signal decoded from said first coded representation is sufficient for human hearing; the quality of the audio signal decoded from said second coded representation is sufficient for performing a speech recognition task but insufficient for human hearing; method.
[0200] The method of EEE-F21, EEE-F19 or EEE-F20, the first coded representation is in a first independent block of the coded bitstream and the second coded representation is in a second independent block of the coded bitstream. method.
[0201] The method of EEE-F22, EEE-F19 or EEE-F20, the first coded representation is in a first layer of the coded bitstream, the second coded representation is in a second layer of the coded bitstream, and the first layer and the second layer are contained in a single block of the coded bitstream. method.
[0202] Any one of methods EEE-F23, EEE-F18 to EEE-F22, decoding the encoded audio signal includes decoding only the second coded representation and ignoring the first coded representation. method.
[0203] Any one of methods EEE-F24, EEE-F18 to EEE-F23, the audio signal decoded from the second coded representation is in a parametric representation, a waveform representation, or a representation comprising one or more of band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spatial coefficients. method.
[0204] EEE-F25. An apparatus configured to perform any one of the methods EEE-F1 to EEE-F24.
[0205] EEE-F26. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform any one of the methods EEE-F1 to EEE-F24.
[0206] EEE-G1. A method for encoding an audio signal of an immersive audio program for low-latency transmission to one or more playback devices, comprising: receiving a plurality of time-domain audio signals of the immersive audio program; Select the frame size, extracting frames of the time-domain audio signal according to the frame size, wherein the frames of the time-domain audio signal overlap with previous frames of the time-domain audio signal; segmenting the audio signal into overlapping frames; converting the frames of a time domain audio signal into a frequency domain signal; coding the frequency domain signal; quantizing the coded frequency domain signal using a perceptually based quantization tool; assembling the quantized coded frequency domain signals into one or more independent blocks within the frame; assembling said one or more independent blocks into an encoded frame; A method having the following.
[0207] EEE-G2. The method of EEE-G1, the plurality of audio signals includes channel-based signals having a defined channel configuration; method.
[0208] The method of EEE-G3.EEE-G2, the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2; method.
[0209] Any one of the methods EEE-G4, EEE-G1 to EEE-G3, the plurality of audio signals includes one or more object-based signals; method.
[0210] Any one of the methods EEE-G5, EEE-G1 to EEE-G4, the plurality of audio signals comprising a scene-based representation of the immersive audio program; method.
[0211] Any one of the methods EEE-G6, EEE-G1 to EEE-G5, the selected frame size is one of 128, 256, 512, 1024, 120, 240, 480, or 960 samples; method.
[0212] Any one of the methods EEE-G7, EEE-G1 to EEE-G6, an overlap between a frame of the time-domain audio signal and a previous frame of the time-domain audio signal is less than or equal to 50%; method.
[0213] Any one of the methods EEE-G8, EEE-G1 to EEE-G7, The transform is a modified discrete cosine transform (MDCT). method.
[0214] Any one of the methods EEE-G9, EEE-G1 to EEE-G8, two or more of the plurality of audio signals are jointly coded; method.
[0215] Any one of the methods EEE-G10.EEE-G1 to EEE-G9, each individual block containing encoded signals for one or more playback devices; method.
[0216] Any one of the methods EEE-G11.EEE-G1 to EEE-G10, At least one independent block includes coded signals for two or more playback devices, the coded signals including jointly coded audio signals. method.
[0217] Any one of the methods EEE-G12.EEE-G1 to EEE-G10, At least one independent block contains a plurality of coded signals covering different bandwidths intended for playback from different drivers of a playback device; method.
[0218] EEE-G13. Any one of the methods EEE-G1 to EEE-G12, At least one independent block contains an encoded echo reference signal used in echo management performed by the playback device; method.
[0219] EEE-G14. Any one of the methods EEE-G1 to EEE-G13, Coding the quantized frequency domain signal includes applying one or more of the following tools: Temporal Noise Shaping (TNS), Joint Channel Coding, Sharing of scale factors between signals, Determining control parameters for high frequency reconstruction, and Determining control parameters for noise substitution. method.
[0220] EEE-G15. Any one of the methods EEE-G1 to EEE-G14, one or more independent blocks containing parameters controlling one or more of delay, gain, and equalization of the playback device; method.
[0221] EEE-G16. A low-delay method for decoding an audio signal of an immersive audio program from an encoded signal, comprising: receiving an encoded frame including one or more independent blocks; extracting the quantized coded frequency domain signal from one or more independent blocks; dequantizing the quantized coded frequency domain signal; decoding the dequantized frequency domain signal; inverse transforming the decoded frequency domain signal to obtain a time domain signal; overlapping and adding the time domain signals with time domain signals from previous frames to provide multiple audio signals of the immersive audio program; A method having the following.
[0222] EEE-G17. The method of EEE-G16, comprising: the plurality of audio signals includes channel-based signals having a defined channel configuration; method.
[0223] EEE-G18. The method of EEE-G17, comprising: the channel configuration is one of mono, stereo, 5.1, 5.1.4, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2; method.
[0224] Any one of the methods EEE-G19.EEE-G16 to EEE-G18, the plurality of audio signals includes one or more object-based signals; method.
[0225] Any one of the methods EEE-G20. EEE-G16 to EEE-G19, the plurality of audio signals comprising a scene-based representation of the immersive audio program; method.
[0226] Any one of the methods EEE-G21.EEE-G16 to EEE-G20, a frame of time domain samples having one of 128, 256, 512, 1024, 120, 240, 480, or 960 samples; method.
[0227] Any one of the methods EEE-G22, EEE-G16 to EEE-G21, The overlap with the previous frame is 50% or less. method.
[0228] Any one of the methods EEE-G23.EEE-G16 to EEE-G22, The inverse transform is an inverse modified discrete cosine transform (IMDCT). method.
[0229] Any one of the methods EEE-G24. EEE-G16 to EEE-G23, each independent block containing a quantized coded frequency domain signal for one or more playback devices; method.
[0230] Any one of the methods EEE-G25, EEE-G16 to EEE-G24, At least one independent block includes quantized coded frequency domain signals for two or more playback devices, and the quantized coded frequency domain signals are jointly coded audio signals. method.
[0231] Any one of the methods EEE-G26, EEE-G16 to EEE-G25, at least one independent block includes a plurality of quantized coded frequency domain signals covering different bandwidths intended for reproduction from different drivers of a reproduction device; method.
[0232] Any one of the methods EEE-G27.EEE-G16 to EEE-G26, At least one independent block contains an encoded echo reference signal used in echo management performed by the playback device; method.
[0233] Any one of the methods EEE-G28.EEE-G16 to EEE-G27, Decoding the dequantized frequency domain includes applying one or more of the following decoding tools: temporal noise shaping (TNS), joint channel decoding, sharing scale factors between signals, high frequency reconstruction, and noise substitution. method.
[0234] Any one of the methods EEE-G29.EEE-G16 to EEE-G28, one or more independent blocks containing parameters controlling one or more of delay, gain, and equalization of the playback device; method.
[0235] Any one of the methods EEE-G30. EEE-G16 to EEE-G29, The method is performed by a playback device, extracting the quantized coded frequency domain signal from the one or more independent blocks includes selecting only blocks containing quantized coded frequency domain signals for reproduction by the reproduction device and ignoring independent blocks containing quantized coded frequency domain signals for reproduction by other reproduction devices. method.
[0236] EEE-G31. An apparatus configured to perform any one of the methods EEE-G1 to EEE-G30.
[0237] EEE-G32. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform any one of the methods EEE-G1 through EEE-G30.
[0238] In the specification and claims, terms such as "consisting of" and "including" are open-ended terms that mean at least the elements / features listed before it, and are not exclusive. Therefore, when used in a claim, the term "having" should not be interpreted as being limited to the means, elements, or steps listed before it. For example, the scope of an expression "a device having A and B" should not be limited to a device consisting only of elements A and B. The terms "comprise" and "comprising" used in this application are open-ended terms that mean at least the elements / features listed before it, and are not exclusive. Therefore, "comprising" is synonymous with "having" and means the same.
[0239] It should be understood that in the foregoing description of examples of the present invention, various features may be grouped together in a single example, figure, or description for the purpose of simplifying the disclosure and facilitating an understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that any features other than those expressly recited in each claim are required. Rather, as the following claims reflect, inventive aspects may lie in fewer than all features of a single foregoing disclosed example. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate example of the present invention.
[0240] Furthermore, although some examples described herein include some features but not others that are included in other examples, combinations of features from different examples are intended to be encompassed and form different examples, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed examples may be used in any combination.
[0241] Furthermore, some examples are described herein as methods or combinations of elements of methods that may be implemented by a processor of a computer system or other means for performing a function. Thus, a processor with the necessary instructions for performing such a method or element of a method forms a means for performing the method or element of a method. Furthermore, a described element of an apparatus herein is an example of a means for performing the function performed by that element.
[0242] Additionally, some examples described herein disclose connectivity solutions that should be construed as being implemented in distribution and / or transmission systems, such as wired and / or wireless systems, for example, by using any electrical, optical, and / or mobile systems, such as 3G, 4G, and 5G.
[0243] Thus, while specific examples of the invention have been described, those skilled in the art will recognize that other and further modifications may be made, and it is intended to claim all such modifications and improvements. For example, any formulas given above are merely representative of procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted from the methods described.
[0244] The above systems, devices, and methods may be implemented as software, firmware, hardware, or a combination thereof. For example, aspects of the present disclosure may be embodied at least in part in a device, a system including more than one device, a method, a computer program product, etc.
[0245] In a hardware implementation, the division of tasks among functional units mentioned in the above description does not necessarily correspond to a division into physical units; in contrast, one physical component may have multiple functions and one task may be performed cooperatively by multiple physical components.
[0246] Certain or all components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application specific integrated circuits. Such software may be distributed on computer-readable media, which may include computer storage media (i.e., non-transitory media) and communication media (i.e., transitory media).
[0247] As is well known to those skilled in the art, the term computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data, including, but not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer.
[0248] Additionally, those skilled in the art will appreciate that communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0249] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 378,497, filed October 5, 2022, and U.S. Provisional Patent Application No. 63 / 578,534, filed August 24, 2023, which applications are incorporated herein by reference in their entireties.
Claims
1. A method for generating an encoded bitstream from an audio program containing multiple audio signals, For each of the aforementioned multiple audio signals, information is received that indicates the playback device to which each audio signal is associated. For each playback device, information is received that indicates the delay associated with that playback device, To determine two or more related groups of audio signals from the aforementioned plurality of audio signals, The process involves applying one or more co-coding tools to two or more related audio signals of the group to obtain a co-coded audio signal, The co-coded audio signal, the instructions for the playback device associated with the co-coded audio signal, and the information indicating the delay associated with each playback device associated with the co-coded audio signal are combined to form an encoded bitstream. A method of having.
2. The delay associated with each of the playback devices depends on the position of each playback device relative to the listener's position, and / or the position of each playback device relative to the position of other playback devices. The method according to claim 1.
3. The aforementioned delay is adjusted according to changes in the listener's position, the position of the playback device, and / or the position of one or more other playback devices. The method according to claim 1 or 2.
4. The further comprises determining from the plurality of audio signals an audio signal that is not part of the group of two or more related audio signals, The method according to any one of claims 1 to 3.
5. The method further includes applying a delay associated with the playback device to which the audio signal is associated with the audio signal, to the audio signal that is not part of the aforementioned group of two or more related audio signals. The method according to claim 4.
6. The method further comprises independently coding an audio signal that is not part of the group of two or more related audio signals, and combining the independently coded audio signal with instructions for a playback device to which the independently coded audio signal is associated into the coded bitstream. The method according to claim 5.
7. The method further comprises independently coding an audio signal that is not part of the group of two or more associated audio signals, and combining the independently coded audio signal, an instruction for a playback device to which the independently coded audio signal is associated, and an instruction for a delay associated with the playback device to which the independently coded audio signal is associated, into the coded bitstream. The method according to claim 4.
8. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, wherein the frames include one or more independent blocks of encoded data, Identifying independent blocks of encoded data corresponding to one or more audio signals associated with the playback device from the encoded bitstream, Extracting identified independent blocks of the encoded data from the encoded bitstream, It is determined that the extracted independent blocks of the encoded data include two or more jointly coded audio signals, Applying one or more co-decoding tools to the two or more co-coded audio signals to obtain the one or more audio signals associated with the playback device, From the encoded bitstream, determine the delay associated with the playback device, The delay associated with the playback device is applied to one or more audio signals associated with the playback device. A method of having.
9. The determined delay associated with the playback device depends on the listener's position and / or the position of the playback device relative to other playback devices. The method according to claim 8.
10. The determined delay associated with the playback device is different from a previously determined delay associated with the playback device, and the method further comprises interpolating between the previously determined delay associated with the playback device and the determined delay associated with the playback device. The method according to claim 8 or 9.
11. The determined delay differs from the previously determined delay, gain, and / or equalization curve due to changes in the listener's position, the position of the playback device, and / or the position of one or more other playback devices. The method according to claim 10.
12. The frame of the encoded bitstream comprises two or more independent blocks of encoded data, and the method is Determining that one or more of the aforementioned independent blocks contain audio signals not associated with the playback device, Ignoring one or more independent blocks containing audio signals not associated with the aforementioned playback device. It further has, The method according to any one of claims 8 to 11.
13. Applying one or more of the aforementioned co-decryption tools means Identifying a subset of the co-coded audio signals associated with the playback device, reconstructing only that subset of the co-coded audio signals to obtain one or more audio signals associated with the playback device, or Reconstructing each of the jointly coded audio signals, identifying a subset of the reconstructed jointly coded audio signals associated with the playback device, and obtaining one or more audio signals associated with the playback device from the subset of the reconstructed jointly coded audio signals associated with the playback device. including, The method according to any one of claims 8 to 12.
14. An apparatus configured to perform the method described in any one of claims 1 to 13.
15. A non-temporary computer-readable storage medium comprising a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of claims 1 to 13.