Method, apparatus and medium for encoding and decoding an audio bitstream and associated echo reference signal

The method addresses low-latency audio streaming challenges by encoding audio signals into independent blocks for flexible rendering and echo management, enhancing data efficiency and echo handling across multiple devices.

JP2025535060APending Publication Date: 2025-10-22DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025519783
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-24
Filing Date
2023-09-15
Publication Date
2025-10-22

Smart Images

  • Figure 2025535060000001_ABST
    Figure 2025535060000001_ABST
Patent Text Reader

Abstract

1. A method for generating a frame of an encoded bitstream of an audio program including a plurality of audio signals, the frame comprising two or more independent blocks of encoded data, the method comprising: receiving, for one or more of the plurality of audio signals, information indicative of a playback device with which the one or more audio signals are associated; receiving, for the indicated playback device, information indicative of one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Related Applications) This application claims priority to U.S. Provisional Patent Application No. 63 / 378,498, filed October 5, 2022, and U.S. Provisional Patent Application No. 63 / 578,537, filed August 24, 2023, both of which are incorporated by reference in their entireties.

[0002] The present disclosure relates generally to processing audio signals, and more particularly to encoding and decoding audio sources for low latency interchange of immersive audio programs between devices. [Background technology]

[0003] Streaming audio is common in today's society. Audio streaming is becoming increasingly demanding as user expectations for quality rise, but user setups are becoming more complex with the number of speakers and even different types of speakers, possibly even within the same setup. Streaming is usually done, at least in part, over a wireless link, which places requirements on the wireless link to be of good quality, but as many people have likely experienced, this is not always the case.

[0004] Therefore, for use cases where a particular format is streamed from a cloud / server, then transcoded on a device to a more suitable low-latency format and delivered over a wireless link (or in some cases, a wired link), an interchange format needs to be defined. Example use cases are home connectivity and phone-to-car connectivity. However, such formats can be beneficial in scenarios where low-latency delivery of audio signals from a single device to one or more connected devices is desired.

[0005] In addition to audio information being transmitted wirelessly and streamed, there may also be other types of information incorporated into the stream, which may then also be subject to the quality of the wireless link and have the same drawbacks as audio.

[0006] Therefore, it would be advantageous to overcome the problems associated with wirelessly streaming various types of streaming audio in combination with other types of information or signals. Summary of the Invention [Problem to be solved by the invention]

[0007] The present disclosure is directed to overcoming the above-mentioned problems associated at least in part with wirelessly streaming a combination of audio and other types of information. [Means for solving the problem]

[0008] A first aspect of the present disclosure is directed to a method for generating a frame of an encoded bitstream of an audio program including a plurality of audio signals, the frame including two or more independent blocks of encoded data, the method including: receiving, for one or more of the plurality of audio signals, information indicative of a playback device with which the one or more audio signals are associated; receiving, for the indicated playback device, information indicative of one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and combining the first independent block and the one or more additional independent blocks into a frame of the encoded bitstream.

[0009] A second aspect of the present disclosure is directed to a method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, the frame including two or more independent blocks of encoded data, the playback device including one or more microphones, the method including: identifying independent blocks of encoded data from the encoded bitstream that correspond to one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; extracting one or more audio signals associated with the playback device from the identified independent blocks of encoded data; identifying one or more other independent blocks of encoded data from the encoded bitstream that correspond to one or more audio signals associated with one or more other playback devices; extracting one or more audio signals associated with the one or more other playback devices from the one or more other independent blocks of encoded data; capturing the one or more audio signals using the one or more microphones of the playback device; and using, in response to the one or more captured audio signals, the one or more extracted audio signals associated with the one or more other playback devices as echo references for echo management of the playback device.

[0010] A third aspect of the present disclosure is directed to an apparatus configured to perform the method of any of the first and / or second aspects.

[0011] A fourth aspect of the present disclosure is directed to a non-transitory computer-readable storage medium including a sequence of instructions that, when executed, cause one or more devices to perform the methods of any of the first and / or second aspects.

[0012] Further embodiments of the present disclosure are set forth in the dependent claims.

[0013] In some embodiments, one or more audio signals associated with a playback device are specifically intended for use as an echo reference for echo management of the playback device.

[0014] In some embodiments, one or more audio signals intended for use as an echo reference are transmitted using less data than one or more audio signals associated with a playback device.

[0015] In this disclosure, a frame represents a time slice across all signals. A block stream represents a collection of signals in a session. A block represents one frame of a block stream. For digital audio at a given sampling frequency, the frame size is equivalent to the number of audio samples in a frame of any audio signal. The frame size is typically constant during a session.

[0016] A wake word may include a single word or a phrase containing two or more words in a fixed order.

[0017] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that acts as a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and the other (XM) inputs are received from external sources) may also be referred to as a decoder system.

[0018] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings, in which like elements are designated by like reference numerals, and in which various embodiments are shown, although one or more implementations are not limited to the embodiments shown in the drawings. [Brief explanation of the drawings]

[0019] [Figure 1]FIG. 1 illustrates an example of low latency transcoding in a home connection.

[0020] [Figure 2] FIG. 1 illustrates an example of audio streaming in a car connection.

[0021] [Figure 3] FIG. 1 illustrates an example of enhanced television audio streaming using a wireless device.

[0022] [Figure 4] FIG. 1 illustrates an example of audio streaming over a simple wireless speaker.

[0023] [Figure 5] FIG. 10 illustrates an example of bitstream element to device metadata mapping.

[0024] [Figure 6] FIG. 1 shows an example of how simple flexible rendering of audio streaming can be deployed.

[0025] [Figure 7] FIG. 10 illustrates an example of mapping flexible rendering data to bitstream elements.

[0026] [Figure 8] FIG. 10 illustrates an example of signaling echo references for commands when playing and listening to audio on multiple devices.

[0027] [Figure 9] FIG. 2 illustrates an example of how frames, blocks, and packets relate to each other.

[0028] [Figure 10]FIG. 10 shows an example of additional codec information in the current MPEG-4 structure.

[0029] [Figure 11] FIG. 10 shows an example of additional codec information in the current MPEG-4 structure.

[0030] [Figure 12] FIG. 1 illustrates an example of integrating listening capabilities and speech recognition.

[0031] [Figure 13] FIG. 1 illustrates an example of a frame including multiple blocks.

[0032] [Figure 14] FIG. 1 illustrates an example of a bitstream including multiple blocks.

[0033] [Figure 15] FIG. 1 illustrates an example of blocks with different priorities.

[0034] [Figure 16] FIG. 1 illustrates an example of a bitstream containing multiple blocks with different priorities.

[0035] [Figure 17] FIG. 2 illustrates an example of frames with different priorities.

[0036] [Figure 18] FIG. 2 illustrates an example of a bitstream containing frames with different priorities. DETAILED DESCRIPTION OF THE INVENTION

[0037] The principles of the present invention will now be described with reference to various examples shown in the drawings. It should be understood that the descriptions of these examples are merely intended to enable those skilled in the art to better understand and further practice the present invention, and are not intended to limit the scope of the present invention.

[0038] In Figure 1, an immersive audio stream is streamed from a cloud or server 10 and decoded at a television or hub device 20. The immersive audio stream may be encoded in any existing format, including, for example, Dolby Digital Plus, AC-4, etc. The output is then transcoded into a low-latency interchange format and further transmitted to connected devices 30. These devices are preferably connected via a local wireless connection, such as a WiFi soft access point, or a Bluetooth connection. Low latency typically depends on various factors, such as frame size, sampling rate, and hardware and / or software computational resources, but is typically less than 40 ms, 20 ms, or 10 ms.

[0039] In another typical use case, a phone retrieves an immersive audio stream from a cloud or server, transcodes it into an interchange format, and then transmits it to a connected car. In the example shown in FIG. 2, a mobile device (e.g., a phone or tablet) 20 is connected to a server 10, receives the immersive audio stream, performs low-latency transcoding into a low-latency interchange format, and transmits the transcoded signal to a car 30. The car 30 supports immersive audio playback. An example of an immersive audio stream is a stream containing audio in the Dolby Atmos immersive format. An example of a car 30 that supports immersive audio playback is a car configured to play the Dolby Atmos immersive format.

[0040] In general, an interchange format preferably has low latency, low encoding and decoding complexity, the ability to scale to high quality, and reasonable coding efficiency. Additionally, the format preferably supports latency configurability so that latency can be traded off against efficiency, and error resilience so that it can operate under varying connection conditions.

[0041] (Hub or extended display with wireless speakers with or without listening capabilities) In the example shown in Figure 3, a hub 20 driving a set of wireless speakers 30, or a display (e.g., television or TV) 20 that may have built-in speakers 30, is augmented with several wireless speakers 30. The augmentation suggests that in examples where the display 20 includes speakers 30, the display 20 is also part of the audio playback. The wireless speakers / devices 30 may be receiving the same complete signal (broadcast mode, shown on the left side of Figure 3) or individual streams dedicated to specific devices (unicast multipoint mode, shown on the right side of Figure 3).

[0042] Each speaker may contain multiple drivers covering different frequency ranges of the same channel or corresponding to different channels within a standard mapping. For example, a speaker may have two drivers, one of which may be an upward-firing driver that emulates a high-positioned speaker by outputting a signal corresponding to a height channel. Wireless devices may have listening capabilities (e.g., are "smart speakers") and therefore may require echo management. Echo management may require a speaker to receive one or more echo references. The echo references may be local (e.g., for the same speaker / device) or may represent related signals from other nearby speakers / devices.

[0043] The speakers can be placed in any position. In this case, so-called "flexible rendering" can be performed, whereby the rendering takes into account the actual position of the speakers (e.g., as opposed to the standard assumption that loudspeakers are placed in fixed, predetermined locations). Flexible rendering can be performed at the hub or TV, where the rendered signals are then sent to the speakers / devices in broadcast mode or as individual streams to the individual devices. In broadcast mode, each of the rendered signals is sent to each device, and each corresponding device extracts and outputs the appropriate signal. Alternatively, flexible rendering can be performed locally at each device, whereby each device receives a representation of the full immersive program, e.g., a 7.1.4-channel-based immersive representation, and from that representation, renders an output signal appropriate for that device.

[0044] Wireless devices may connect to a hub or TV through a soft access point provided by the device or through a local access point in the home, which may impose different requirements on bit rate and latency.

[0045] (Mobile Device Projection in Car Applications) In another use case, a mobile device (e.g., a phone or tablet) 20 retrieves an immersive audio stream from the cloud or server 10, transcodes the immersive audio stream into an interchange format, and then transmits the transcoded stream to a connected car 30. In the example shown in FIG. 2, a Dolby Atmos-capable phone 20 is connected to the server 10, receives the Dolby Atmos stream, performs low-latency transcoding, and transmits it to a Dolby Atmos-capable car 3. This use case may allow for higher bit rates than the living room use case, and may have different wireless channel characteristics. The higher bit rate in the car use case is likely due to the lower noise in the wireless environment. This is because the car acts as a kind of Faraday cage, insulating the interior environment from external wireless interference. In contrast, in a living room use case, all wireless devices, including neighbors' wireless devices and those in other rooms, add noise to the wireless environment. At the same time, the number of wireless devices competing for available wireless bandwidth in a car is typically very low. In comparison, in a living room scenario, all wireless devices in the home, including those in the living room, are typically competing for the available wireless spectrum.

[0046] For this use case, it is envisioned that the signal to be transcoded by the mobile device and sent to the car may be a channel-based immersive representation, an object-based representation, a scene-based representation (e.g., an Ambisonics representation), or even a combination of various representations. In this example, various rendering architectures and broadcast versus multipoint may be irrelevant, since typically the complete presentation is transmitted from the mobile device to a single endpoint (e.g., the car).

[0047] (Interchange format explanation) The Immersive Interchange Format is based on the Modified Discrete Cosine Transform (MDCT) with perceptually motivated quantization and coding. It has latency configurability, e.g., support for various transform sizes at a given sampling rate. Example frame sizes are 128, 256, 512, 1024 samples, 120, 240, 480, 960 samples, and 192, 384, 768 samples at sampling rates of 48 kHz and 44.1 kHz.

[0048] The format supports mono, stereo, 5.1, and other channel configurations including immersive channel configurations (e.g., including but not limited to, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, and 22.2, or any other channel configuration consisting of unique channels identified in ISO / IEC 23091-3:2018, Table 2). The format may also support object-based audio and scene-based representations such as Ambisonics (e.g., first-order or higher-order). The format may further use a signaling scheme suitable for integration with existing formats (e.g., formats such as those described in ISO / IEC 14496-3 (which may also be referred to as the MPEG-4 Audio Standard), ISO 14496-1 (which may also be referred to as the MPEG-4 Systems Standard), ISO 14496-12 (which may also be referred to as the ISO Base Media File Format Standard), and / or ISO 14496-14 (which may also be referred to as the MP4 File Format Standard)).

[0049] Further, the system may have support for the ability to skip some syntax elements and decode only the relevant parts for a given speaker, support for metadata-controlled flex rendering aspects such as delay alignment, level adjustment and equalization, support for the use of smart speakers with listening capabilities (including scenarios where one set of signals is sent that allows each of the drivers / loudspeakers in a smart speaker to be fed independently) and associated support for echo management through echo reference signaling, support for quantization and encoding in the MDCT domain with both low and 50% overlap windows, and / or support for temporal shaping of the quantization noise by filtering along the frequency axis in the MDCT domain, e.g., TNS (Temporal Noise Shaping), so that in various embodiments the overlapping windows are symmetric or asymmetric.

[0050] The interchange format disclosed herein, in some embodiments, allows for greater efficiency and suitability for the use cases at hand by providing improved joint channel coding that takes advantage of increased correlation between signals for flexible rendering use cases, improved coding efficiency due to shared scale factors between channel elements, high frequency reconstruction and noise addition techniques being included in the MDCT domain, and various improvements over older coding structures.

[0051] In some embodiments, skippable blocks and metadata may be used to control individual playback devices in broadcast mode. In a broadcast mode configuration, each wireless device needs to play the relevant portions of the complete audio for the specific driver / channel corresponding to that specific device. This implies that the device needs to know which portions of the stream are relevant to that device and extract these portions from the complete audio stream being broadcast. To enable low-complexity decoding operations, the stream is preferably constructed in such a way that a decoder for a specific device can efficiently skip over elements not relevant for decoding to get to the elements relevant to the device and a given driver on the device.

[0052] However, it should be noted that there may still be advantages (from a compression perspective) to allowing joint coding between signals sent to various devices (e.g., in the most simplified scenario, the left and right channels of a stereo presentation).

[0053] As such, the format may enable "skippable blocks" in the bitstream to allow efficient decoding of portions that are only relevant to specific devices, and may include metadata that allows for flexible mapping of one or more skippable blocks to specific devices, while maintaining the ability to apply joint coding techniques between signals corresponding to various speakers / devices.

[0054] 4 shows an example of an arbitrary setup. There are three connected wireless devices 31, 32, 33. Two of these are single-channel speakers 32, 33 (meaning one channel of audio), and one speaker 31 is a more advanced speaker 31 that has three different drivers and operates with three different signals. In this example, the first speaker 31 operates with three individual signals, while the second and third speakers 32, 33 operate with a stereo representation (e.g., left and right). Therefore, it may be advantageous to jointly encode the signals of the stereo representation.

[0055] According to the above scenario, the format specifies a bitstream containing skippable blocks, allowing a first device to extract only the relevant part of the stream to decode the signal for its loudspeaker, while trading off decoder complexity and efficiency for a stereo pair. The tradeoff can be made by constructing the signal such that two signals need to be decoded in order for each loudspeaker to output a single signal. In this case, the advantage is that joint coding can be performed.

[0056] The format specifies a metadata format that allows a general and flexible representation of the mapping of a particular one of a plurality of skippable blocks to one or more devices, as shown in Figure 5. Here, the mapping may be represented as a matrix that associates each device 31, 32, 33 with one or more bitstream elements, so that a decoder for a given device knows which bitstream elements to output and decode.

[0057] 5, the first block or skip block Blk1 contains three single-channel elements (one for each driver of device 1 31), and the second block or skip block Blk2 contains a channel-pair element that may contain jointly encoded versions of the signals to be output by device 2 32 and device 3 33. Device 1 31 extracts the mapping metadata and determines that the signal it needs is within skip block 1 Blk1. Therefore, it extracts skip block 1 Blk1, decodes the three single-channel elements therein, and provides them to drivers 1, 2a, and 2b, respectively.

[0058] Furthermore, device 1 31 ignores skip block 2, Blk2. Similarly, device 2 32 extracts the mapping metadata and determines that the signal it requires is in skip block 2, Blk2. Therefore, device 2 32 skips skip block 1, Blk1, and extracts skip block 2, Blk2. Device 2 32 decodes the channel pair element and provides the left channel output of the CPE to its driver. Similarly, device 3 33 extracts the mapping metadata and determines that the signal it requires is in skip block 2, Blk2. Therefore, device 3 33 skips skip block 1, Blk1, and further extracts skip block 2, Blk2. Device 3 33 decodes the channel pair element and provides the right channel output of the CPE to its driver.

[0059] In some embodiments, device 2 32 may determine that it requires only a subset of the signals from skip block 2 Blk2. In such embodiments, device 2 32 may perform only a subset of the operations necessary to fully decode the signals in skip block 2 Blk2, if possible. In particular, in the example of Figure 5, device 2 32 may perform only the processing operations necessary to extract the left channel of the CPE, thereby reducing computational complexity. Similarly, device 3 33 may perform only the processing operations necessary to extract the right channel of the CPE.

[0060] 5, the CPEs may be coded at times using joint channel coding and at other times using independent channel coding. While the CPEs are coded using independent channel coding, device 2 32 may extract only the first (e.g., left) channel of the CPE, while device 3 33 may extract only the second (e.g., right) channel of the CPE.

[0061] In another embodiment, the CPE's channels may be coded using joint channel coding. In this case, device 2 32 and device 3 33 must extract the two middle channels of the CPE. However, device 2 32 may still be able to operate with reduced computational complexity by performing only the operations necessary to extract the left channel from the middle channels. Similarly, device 3 33 may still be able to operate with reduced computational complexity by performing only the operations necessary to extract the right channel from the middle decoded channels.

[0062] Depending on the specifics of the joint channel coding, other optimizations may be possible. The identity of each of the various devices / decoders may be defined during the initialization or setup phase of the system. Such configuration is general and usually involves measuring the acoustics of the room, the distance from the loudspeaker to the sweet spot, etc.

[0063] (Apply flexible rendering (pre-encode and post-decode) distribution and post-decode delay, etc.) In use cases where flexible rendering is applied to a hub / TV before sending relevant signals to specific devices / speakers, the rendering may create signals that are more difficult to encode from an encoding perspective, for example, when considering joint encoding of each signal. One reason is that flexible rendering may apply different delay, equalization, and / or gain adjustments to different devices (e.g., depending on where a speaker is located relative to other speakers and the listener). For example, it may be possible to preset gain and delay information in the initial setup and then flexibly render only equalization. In other embodiments, other types of preset information and flexible rendering information may also be possible. However, in this document, the term "gain" should be interpreted to mean any level adjustment (e.g., attenuation, amplification, or pass-through) and not limited to a specific level adjustment (e.g., amplification).

[0064] 6, the right channel 33 speaker and the left channel 32 speaker are given different latencies to reflect their different positions relative to the listener (e.g., because speakers 32, 33 may not be equidistant to the listener, different latencies may be applied to the signals output from speakers 32, 33, so that coherent sounds from the different speakers 32, 33 arrive at the listener simultaneously). Introducing different latencies in coherent signals intended to be reproduced by different speakers 31, 32, 33 in this way makes joint coding of such signals more difficult.

[0065] To address this challenge, several aspects of the flexible rendering process can be parameterized and applied within the device as an endpoint after signal decoding.

[0066] 7, the delay and gain of each device are parameterized and included in the encoded signal sent to each speaker 31, 32, 33. Each signal may be decoded by a corresponding device 31, 32, 33, which may then introduce the parameterized gain and delay values ​​into the corresponding decoded signal.

[0067] 7, the encoded signals for different devices 31, 32, 33 are sent in separate blocks (e.g., skippable blocks). In this case, the parameters (e.g., delays and gains) are also sent in separate blocks, such that devices 31, 32, 33 extract only the subset of parameters required for that device 31, 32, 33 and ignore (and skip) the parameters not required for that device 31, 32, 33. In such a case, mapping metadata may be provided to each device indicating which parameters are contained in which blocks.

[0068] 7 shows only delay and gain parameters, other parameters may also be included, such as equalization parameters, including, for example, multiple gains applied to various frequency regions, a representation of a predetermined equalization curve applied by the reproduction devices 31, 32, 33, one or more sets of infinite impulse response (IIR) or finite impulse response (FIR) filter coefficients, a set of biquad filter coefficients, parameters specifying the characteristics of a parametric equalizer, and other parameters specifying equalization known to those skilled in the art.

[0069] Furthermore, the parameterization of the flexible rendering aspect need not be static but can be dynamic (e.g., if the listener moves during playback of an audio program). As such, it may be preferable to allow parameters to change dynamically. When one or more parameters change during an audio program, the device may provide a smooth transition by interpolating between previous and updated delay and / or gain parameters. This may be particularly useful if the system dynamically tracks the listener's position and updates the sweet spot for dynamic rendering accordingly.

[0070] As mentioned above and will be discussed below, when flexible rendering is applied, correlation levels between channels may increase, which can be exploited by more flexible joint coding.

[0071] (Echo Reference Coding and Signaling) In the use case outlined in Figure 8, multiple devices / speakers 30 may cooperate to receive signals in either a broadcast mode, as shown on the left side of Figure 8, or a unicast multipoint mode, as shown on the right side of Figure 8, while simultaneously having a microphone 40 on the device present to enable "listening" capabilities, which creates the need for echo management.

[0072] When performing echo management with multiple speakers / devices, it may be beneficial to use an echo reference from more than just the local speaker device. As an example, one device may be placed in close proximity to another device. Thus, signals from nearby devices impact the echo management of the device with the active microphone. In the broadcast mode use case, each device receives signals for all devices. When a device has signals for other devices, it may be beneficial to use these signals as echo references. To do this, it is necessary to signal specific devices that correspond to signals that can be used as echo references for other devices.

[0073] In one embodiment, this can be done by providing metadata that not only maps which channels / signals (from the entire set) are played by a particular speaker / device, but also which channels / signals (from the entire set) are used as echo references for a particular speaker / device. Such metadata or signaling can be dynamic, allowing for the indication of different preferred echo references over time.

[0074] For use cases where each device / speaker receives only a specific signal to play, it may be necessary to send an additional signal (e.g., an echo reference signal) to each device / speaker to provide an appropriate echo reference. Again, this requires providing device-specific signaling to enable each device to select the appropriate signal to play and the appropriate signal for echo management.

[0075] The signal for echo management is used only by the device performing the echo management and is not reproduced by the device intended for the listener. Therefore, the echo reference signal may be coded or represented in a different way than the signal intended for reproduction by the device. In particular, echo management can be successful with signals coded at a lower rate than would normally be used for reproduction for the listener. Therefore, additional compression tools, such as signal parameterization, that capture the necessary characteristics of the audio signal but are not entirely suitable for reproduction to the listener, can provide good echo management at significantly lower transmission costs.

[0076] (various syntax elements) In some embodiments, blocks are used to optimize the transport of audio. In the above format, each frame may be divided into blocks, as described above for skippable blocks. A block may be identified by the frame number to which it belongs, a block ID that may be used to associate consecutive blocks from different frames with the same ID into a block stream, and a retransmission priority. The above example is a stream with multiple frames N-2, N-1, and N. Frame N contains multiple blocks identified as ID1, ID2, and ID3. This is shown in Figure 13. An example of this stream, but in bitstream format, is shown in Figure 14. In some embodiments, for an audio signal of an immersive audio program, the block ID of a block may indicate which set of signals from the entire immersive audio program is carried by that block.

[0077] The use of the above format is to transmit audio reliably and with low latency over wireless networks such as WiFi. WiFi, for example, uses a packet-based network protocol. Packet sizes are usually limited. A typical maximum packet size in IP networks is 1500 bytes. The packet-based architecture of streams allows flexibility in assembling packets for transmission. For example, packets with smaller frames can be filled with retransmitted blocks from other frames. Larger frames can be separated at inter-block boundaries before packetization, thereby reducing inter-packet dependencies at the network protocol layer.

[0078] Figure 9 illustrates the relationship between frames, blocks, and packets. A frame carries audio data, preferably the entire audio data, representing a contiguous segment of an audio signal, having a start time, an end time, and a duration that is the difference between the start and end times. Contiguous segments may include durations in accordance with ISO / IEC 14496-3, subpart 4, section 4.5.2.1.1. Section 4.5.2.1.1 describes the contents of raw_data_block(). A frame may also carry a redundant representation of that segment, e.g., a representation encoded at a lower data rate. After encoding, the frame can be divided into blocks. Blocks can be combined into packets for retransmission over a packet-based network. Blocks from various frames can be combined into a single packet and / or sent out of order.

[0079] In one embodiment, blocks are used to address individual devices. Packets of data are received by individual devices or groups of related devices. The concept of skippable blocks can be used to address individual devices or groups of related devices. Even if the network is operating in broadcast mode when sending packets to various devices, the processing of the audio (e.g., decoding, rendering, etc.) can be processing of the block addressed to that device. All other blocks can simply be skipped, even if they are received within the same packet. In some embodiments, blocks can be provided in the correct order based on their decoding or presentation time. A retransmitted block of lower priority can be removed if the same block of higher priority has already been received. The stream of blocks can then be provided to a decoder.

[0080] In some embodiments, stream and device configurations are sent out-of-band. Codecs allow for the setup of connections over which audio streams are transmitted at relatively high rates and low latency. The configuration of such connections may remain stable for the duration of such connections. In this case, the components of the audio stream may be transmitted out-of-band rather than created. For such out-of-band transmission, different networks or even network protocols may be used. For example, the audio stream may use User Datagram Protocol (UDP) for low-latency transmission, and the configuration may use Transmission Control Protocol (TCP) to ensure that the configuration is transmitted reliably.

[0081] (Enable the codec in the context of MPEG-4 Audio) One particular application of this technology is its use within MPEG-4 Audio, where different AOTs (Audio Object Types) are defined for different codec technologies. For the formats described in this document, new AOTs can be defined, allowing for format-specific signaling and data. Furthermore, decoder configuration within MPEG-4 is performed within the DecoderSpecificInfo() payload, which in turn carries the AudioSpecificConfig() payload. In the latter case, certain general signaling is defined that is agnostic to a particular format, such as sampling rate and channel configuration, as well as specific information about a particular AOT. This may make sense for traditional formats where the entire stream is decoded by a single device. However, in broadcast mode, a single stream is sent to several decoders, each of which decodes only a portion of the stream. In this case, the initial channel configuration signaling (as a means of configuring the device's output capabilities) may not be optimal.

[0082] Figure 10 shows the traditional MPEG-4 high-level structure (in black), with changes highlighted in gray. To support broadcast use cases, codecSpecificConfig() ("codec" can be a generic placeholder name) is defined. In this case, the signaling is redefined for the specific use case, allowing specific channel elements to be device specific and including other related static parameters. An MPEG-4 element channel configuration with value "0" is defined as the channel configuration defined in codecSpecificConfig. This value can therefore be used to allow revision of the channel configuration signaling within the codec specific config.

[0083] Furthermore, in the MPEG-4 sense, if codecSpecificConfig() is decodable, it specifies the raw payload for the particular decoder at hand. However, the format ensures that dynamic metadata is part of the raw payload, and that length information is available for all raw payloads. This allows decoders to easily skip over elements that are not relevant to a particular device.

[0084] In Figure 11, a portion of an example of a raw_data_block as defined in MPEG-4 is shown on the right. A raw data block contains channel elements (single channel elements (SCE) or channel pair elements (CPE)) in a given order. However, in conventional MPEG-4 audio syntax, a decoder wishing to skip some of the channel elements that may not be relevant to the output device at hand must analyze (and to some extent decode) all of the channel elements to be able to extract the relevant portions. In the novel raw_data_block shown on the left side of Figure 11, the contents are formed from skippable blocks, allowing the decoder to skip the irrelevant portions and decode only those channel elements that the metadata indicates are relevant to the device at hand. In one embodiment, the skippable block contains a raw_data_block and related information.

[0085] Blocks can also be used for retransmission. As described above in the section on skippable blocks, each frame can be divided into blocks. As shown in FIGS. 15 and 16, blocks are identified by the frame number to which they belong, a block ID that can be used to associate consecutive blocks from different frames with the same ID into a block stream, and a retransmission priority. For example, a high retransmission priority (e.g., shown as priority 0 in FIGS. 15 and 16) indicates that a block is preferred at the receiver side over another block with the same block ID and frame number but a lower retransmission priority (shown as priority 1 in FIGS. 15 and 16). However, as shown in the examples of FIGS. 15 and 16, the priority can decrease as the priority index increases (e.g., priority 1 can be lower than priority 0). On the other hand, in other embodiments, the priority can increase as the priority index increases (e.g., priority 0 is lower than priority 1). Still other embodiments will be apparent to those skilled in the art.

[0086] The syntax may support retransmission of audio elements, where different quality levels may be supported. Therefore, retransmitted blocks may carry a "priority" flag to indicate which block with the same block ID has priority to the decoder, because blocks received with the same frame counter and the same block ID are redundant and therefore mutually exclusive to the decoder, as shown in Figures 17 and 18.

[0087] Retransmission of the block may be at a lower data rate. Such a lower data rate may be achieved by reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, reducing the channel count of the audio signal (as described in U.S. Pat. No. 11,289,103, which is incorporated herein by reference in its entirety), or any combination thereof. In order for the decoder to select the audio block that provides the best possible quality, the block that provides the highest quality signal may have the highest decoding priority, the block with the second highest quality signal may have the second highest priority, and so on.

[0088] The block may also be retransmitted at the same quality level, in which case the priority may reflect the latency of the retransmitted block.

[0089] (Tools to improve core coding efficiency) In further embodiments, it may be possible to share MDCT quantization scale factors between channel elements. In some cases, sharing MDCT scale factors between certain channels may allow for a reduction in the bit rate of the associated side. Furthermore, scale factor sharing may be extended to various channel elements, e.g., 7.1.4 inputs. The use of shared scale factors may be indicated by two signaling bits, allowing for three active structures. One possible structure is to share scale factors among the left horizontal channel, the left top channel, the right horizontal channel, and the right top channel. Specific syntax may be aligned with the skippable block concept to ensure that scale factors are shared only within skip blocks.

[0090] It is also possible to have joint coding of more than two channels. In the examples shown in Figures 4, 5, 6, and 7, block Blk1 carries three SCEs, one for each driver in the smart speaker considered in this scenario. In such cases, it may be beneficial to enable joint coding of not only two channels (provided by the CPE, which can be extended to include stereo prediction, e.g., the MDCT-based complex prediction stereo tool as described in ISO / IEC 20330-3, which can be extended to include the tool called MPEG-D USAC), but also more than two channels, e.g., as described in the SAP (Stereo Audio Processing) tool introduced in ETSI TS 103 190.

[0091] Similarly, for flexible rendering, note that there may be a high amount of signal correlation between devices. For such signals, it may be beneficial to construct skip blocks that cover multiple devices, allowing joint channel coding to be applied to more than two channels using the tools outlined above.

[0092] Another joint channel coding tool that can be used for a particular transform length is channel coupling, in which composite channel and scale factor information is transmitted at mid- and high-frequency frequencies. This can reduce the bit rate for playback in the good quality range. For example, it can be beneficial to use the channel coupling tool for a transform length of 256 frames, which corresponds to a frame length of approximately 256 samples for low-latency coding.

[0093] The present disclosure also enables efficient encoding of band-limited signals. In scenarios where separate signals are fed to feed the various drivers of a smart speaker, for example, in a three-way driver configuration with a woofer, midrange, and tweeter, some of these signals may be band-limited. Therefore, efficient encoding of such band-limited signals is desirable. This may imply specifically tuned psychoacoustic models and bit allocation strategies, as well as potential modifications to the syntax to handle such scenarios with improved coding efficiency and / or reduced computational complexity (e.g., enabling the use of a band-limited IMDCT for the woofer feed).

[0094] In some embodiments, a combination of high-frequency reconstruction and noise addition within the MDCT domain is possible. Modern audio codecs are typically designed to support parametric coding techniques. Similarly, it is possible to maintain low latency by, for example, including a high-frequency reconstruction method within the MDCT domain, and to perform a noise addition scheme within the MDCT domain, thereby managing the tone-to-noise ratio in the reconstructed high band. Such parametric coding techniques may be particularly useful at low operating points, especially in scenarios where retransmissions occur as part of an FEC (forward error correction) scheme. An FEC scheme typically involves transmitting a primary signal, followed by a delayed retransmission of the same signal at a lower bit rate.

[0095] These embodiments also allow for reduced peak complexity in the encoder. This is applicable when there is no buffer model limitation for a constant bitrate transport channel. At constant bitrate in buffered mode, after quantization and counting, the resulting number of bits used for the encoded frame may be greater than the limited number of allowed bits. To meet the bit requirement, at least one new, coarser quantization and bit-counting step must be performed. If this restriction were less restrictive, the encoder could retain the initial quantization results and get by with a somewhat higher instantaneous bitrate. Nevertheless, using the same bit reservoir control mechanism, with a virtual buffer fullness that is updated as the encoder follows the buffer requirement, allows the encoder to behave similarly to the constant bitrate behavior of a buffer-model encoder. This saves additional quantization and bit-counting steps, and the resulting audio quality is the same as or better than the constant bitrate case, with the drawback of a slightly increased overall bitrate.

[0096] (Return Channel Concept) For playback and listening use cases, the smart speaker 30 may have one or more microphones 40 (e.g., microphones or microphone arrays) configured to capture a mono or spatial sound field (e.g., mono, stereo, A-format, B-format, or any other isotropic or anisotropic channel format). The codec must provide efficient encoding of the above formats with low latency. Thus, the same codec used for the return channel is used for broadcast / transmission to the smart speaker 30, as shown in the example of FIG. 12.

[0097] Furthermore, it should be noted that while wake word detection occurs on the smart speaker device, speech recognition typically occurs on the cloud on appropriate segments of pre-recorded audio sent by wake word detection.

[0098] In this context, it may be interesting to save overall system complexity by forming the speech analysis system on an intermediate format / representation within the codec, for example. For the simplest use cases where there is no person-to-person conversation, one can define a suitable representation specialized for the speech recognition task, since it is not something that will be heard by other people. This could be band energies, Mel-frequency ceptostrum coefficients (MFCCs), etc., or a low-bitrate version of the MDCT coded spectrum.

[0099] For use cases involving person-to-person conversations, where the wake word and required speech recognition are interleaved, it is sufficient to extract the relevant portion of the stream for that purpose. Such extraction may involve a layered coding structure, with simply additional data sent in parallel with the main speech signal. We further assume a defined transcoding of the existing decoded MDCT spectrum to the relevant representation, and in doing so, a simplified structure. The decoder essentially decodes and outputs the audio that a person can hear, until it is signaled that it should decode (in parallel) the speech recognition representation. This representation can be "peeled" from the layered stream, and may simply be another decoding of the same complete data or simply the decoding and output of an additional representation within the stream. In this use case, the signaling is an enabling piece indicating that the receiving decoder should output a representation relevant to speech recognition.

[0100] (Listing of Examples) The following seven sets of enumerated examples (EEE) (EEE-A, EEE-B, EEE-C, EEE-D, EEE-E, EEE-F, and EEE-G) describe various aspects of the embodiments disclosed herein, and are not claims.

[0101] EEE-A1. A method for decoding an audio signal, comprising: receiving a bitstream including at least one frame, each frame including a plurality of blocks; determining from the signaling data information for identifying portions of one or more of the plurality of blocks that should be skipped during decoding based on device information for an output device; decoding the bitstream while skipping the identified one or more blocks; A method comprising:

[0102] EEE-A2. The method of EEE-A1, wherein the information for identifying portions of one or more of the plurality of blocks to skip during decoding includes a matrix associating each output device of a plurality of output devices with one or more bitstream elements.

[0103] EEE-A3. The method of EEE-A2, wherein said one or more bitstream elements are required for said decoding of said bitstream for said associated output device corresponding thereto.

[0104] EEE-A4. The method of any one of EEE-A1 to EEE-A3, wherein the output device may include at least one of a wireless device, a mobile device, a tablet, a single channel speaker, and / or a multi-channel speaker.

[0105] EEE-A5. The method of any one of EEE-A1 to EEE-A4, wherein the identified portion comprises at least one block.

[0106] EEE-A6. The method of any one of EEE-A1 to EEE-A5, wherein the output device is a first output device, and the method further comprises applying a joint coding technique between one or more signals of the bitstream to a second output device and a third output device.

[0107] EEE-A7. The method of any one of EEE-A1 to EEE-A6, wherein the identity of each output device and / or decoder is defined during a system initialization phase.

[0108] EEE-A8. The method of any one of EEE-A1 to EEE-A7, wherein the signaling data is determined from metadata of the bitstream.

[0109] EEE-A9. An apparatus configured to perform the method described in any one of EEE-A1 to EEE-A8.

[0110] EEE-A10. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the method described in any one of EEE-A1 to EEE-A8.

[0111] EEE-B1. A method for generating an encoded bitstream from an audio program including a plurality of audio signals, comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; receiving, for each playback device, information indicative of at least one of a delay, a gain, and an equalization curve associated with said respective playback device; determining a group comprising two or more related audio signals from the plurality of audio signals; - applying one or more joint coding tools to two or more related audio signals of said set to obtain a jointly coded audio signal; combining the jointly coded audio signal, an indication of the playback device with which the jointly coded audio signal is associated, and the indication of the delay and gain associated with each playback device with which the jointly coded audio signal is associated into an independent block of an encoded bitstream; A method comprising:

[0112] EEE-B2. The method of EEE-B1, wherein the delay, gain and / or equalization curve associated with each playback device depends on the position of each playback device relative to the position of the listener.

[0113] EEE-B3. The method of EEE-B1 or EEE-B2, wherein the delay, gain and / or equalization curve associated with each playback device depends on the position of each playback device relative to the positions of other playback devices.

[0114] EEE-B4. The method of any one of EEE-B1 to EEE-B3, wherein the delay, gain and / or equalization curve is dynamically variable.

[0115] EEE-B5. The method of EEE-B4, wherein the delay, gain and / or equalization curves are adjusted in response to a change in the position of the listener.

[0116] EEE-B6. The method of EEE-B4 or EEE-B5, wherein the delay, gain and / or equalization curves are adjusted in response to a change in the position of the playback device.

[0117] EEE-B7. The method of any one of EEE-B4 to EEE-B6, wherein the delay, gain and / or equalization curves are adjusted in response to a change in the position of one or more of the other playback devices.

[0118] EEE-B8. The method of any one of EEE-B1 to EEE-B7, further comprising determining from said plurality of audio signals an audio signal that is not part of said group comprising two or more related audio signals.

[0119] The method of EEE-B8 further comprising, for the audio signal that is not part of the group containing related audio signals EEE-B9.2 or higher, applying the delay, gain and / or equalization curves associated with the playback device with which the audio signal is associated.

[0120] The method of EEE-B9, further comprising independently encoding the audio signals that are not part of the group that includes an associated audio signal of EEE-B10.2 or higher, and combining the independently encoded audio signals and the playback device indications with which they are associated into separate and independent decodable subsets of the encoded bitstream.

[0121] The method of EEE-B8, further comprising independently encoding the audio signals that are not part of the group comprising EEE-B11.2 or higher related audio signals, and combining the independently encoded audio signals, an indication of the playback device with which the independently encoded audio signals are associated, and an indication of the delay, gain and / or equalization curves associated with the playback device with which the independently encoded audio signals are associated into a separate and independent decodable subset of the encoded bitstream.

[0122] EEE-B12. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, said frames comprising one or more independent blocks of encoded data, said method comprising: identifying discrete blocks of encoded data from the encoded bitstream corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; determining that the extracted independent blocks of encoded data comprise two or more jointly encoded audio signals; applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device; determining at least one of a delay, a gain, and an equalization curve associated with the playback device from the extracted independent blocks of encoded data; applying the delay, gain and / or equalization curve associated with the playback device to the one or more audio signals associated with the playback device.

[0123] EEE-B13. The method of EEE-B12, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to the position of the listener.

[0124] EEE-B14. The method of EEE-B12 or EEE-B13, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices.

[0125] EEE-B15. The method of any one of EEE-B12 to EEE-B14, wherein the determined delay, gain and / or equalization curve of the playback device is dynamically variable.

[0126] EEE-B16. The method of EEE-B15, further comprising, when the determined delay, gain and / or equalization curve associated with the playback device differs from a previously determined delay, gain and / or equalization curve associated with the playback device, interpolating between the previously determined delay, gain and / or equalization curve associated with the playback device and the determined delay, gain and / or equalization curve associated with the playback device.

[0127] EEE-B17. The method of EEE-B16, wherein the determined delay, gain and / or equalization curve differs from the previously determined delay, gain and / or equalization curve due to a change in the position of the listener.

[0128] EEE-B18. The method of EEE-B16 or EEE-B17, wherein the determined delay, gain and / or equalization curve differs from the previously determined delay, gain and / or equalization curve because the position of the playback device has changed.

[0129] EEE-B19. The method of any one of EEE-B16 to EEE-B18, wherein the determined delay, gain and / or equalization curve differs from the previously determined delay, gain or equalization curve due to a change in the position of one or more of the other playback devices.

[0130] EEE-B20. The frame of the encoded bitstream comprises two or more independent blocks of encoded data, and the method comprises: determining that one or more independent blocks contain audio signals not associated with said playback device; ignoring the one or more independent blocks containing audio signals not associated with the playback device; 10. The method of any one of EEE-B12 to EEE-B19, comprising:

[0131] The method of any one of EEE-B12 to EEE-B20, wherein applying a joint decoding tool of EEE-B21.1 or above comprises identifying a subset of the jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device by reconstructing only the subset of the jointly encoded audio signals.

[0132] The method according to any one of EEE-B12 to EEE-B20, wherein applying a joint decoding tool of EEE-B22.1 or above comprises reconstructing each of the jointly encoded audio signals, identifying a subset of the reconstructed jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of the reconstructed jointly encoded audio signals associated with the playback device.

[0133] EEE-B23. Apparatus configured to perform a method according to any one of EEE-B1 to EEE-B22.

[0134] EEE-B24. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the methods described in any one of EEE-B1 to EEE-B22.

[0135] EEE-C1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, said frames including two or more independent blocks of encoded data, said method comprising: receiving, for one or more of the plurality of audio signals, information indicating a playback device with which the one or more audio signals are associated; receiving, for the indicated playback device, information indicating one or more additional associated playback devices; receiving one or more audio signals associated with said one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method comprising:

[0136] EEE-C2. The plurality of audio signals includes one or more groups of audio signals that are not associated with either the playback device or the one or more additional associated playback devices, and the method further comprises: encoding each of the one or more groups of audio signals that are not associated with the playback device or the one or more additional associated playback devices into individual blocks; combining independent blocks corresponding to each of the one or more groups into the frame of the encoded bitstream; The method of EEE-C1 further comprises:

[0137] EEE-C3. The method of EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended for use as echo references for echo management of the playback devices.

[0138] EEE-C4. The method of EEE-C3, wherein the one or more audio signals intended for use as echo references are transmitted using less data than the one or more audio signals associated with the playback device.

[0139] EEE-C5. The method of EEE-C3 or EEE-C4, wherein the one or more audio signals intended for use as echo references are encoded using a parametric coding tool.

[0140] EEE-C6. The method of EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

[0141] EEE-C7. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising two or more independent blocks of encoded data, the playback device comprising one or more microphones, the method comprising: identifying discrete blocks of encoded data from the encoded bitstream corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; extracting the one or more audio signals associated with the playback device from the identified independent blocks of encoded data; identifying one or more other independent blocks of encoded data from the encoded bitstream corresponding to one or more audio signals associated with one or more other playback devices; extracting the one or more audio signals associated with the one or more other playback devices from the one or more other independent blocks of encoded data; capturing one or more audio signals using the one or more microphones of the playback device; using the one or more extracted audio signals associated with the one or more other playback devices as echo references for echo management of the playback device in response to the one or more captured audio signals; A method comprising:

[0142] EEE-C8. Determining that the encoded bitstream includes one or more additional independent blocks of encoded data; ignoring the one or more additional independent blocks of encoded data; and The method of EEE-C7, further comprising:

[0143] EEE-C9. The method of EEE-C8, wherein ignoring the one or more additional independent blocks of encoded data comprises skipping the one or more additional independent blocks of encoded data without extracting the one or more additional independent blocks of encoded data.

[0144] EEE-C10. The method of any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are specifically intended for use as echo references for echo management of the playback devices.

[0145] EEE-C11. The method of EEE-C10, wherein the one or more audio signals specifically intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.

[0146] EEE-C12. The method of EEE-C10 or EEE-C11, wherein the one or more audio signals specifically intended for use as an echo reference are reconstructed from a parametric representation of the one or more audio signals.

[0147] EEE-C13. The method of any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

[0148] EEE-C14. The method of EEE-C7, wherein the encoded signal includes signaling information indicating which of the one or more other playback devices to use as an echo reference for the playback device.

[0149] EEE-C15. The method of EEE-C14, wherein the one or more other playback devices indicated by the signaling information for the current frame are different from the one or more other playback devices used as echo references for previous frames.

[0150] EEE-C16. An apparatus configured to perform the method described in any one of EEE-C1 to EEE-C15.

[0151] EEE-C17. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the methods described in any one of EEE-C1 to EEE-C15.

[0152] EEE-D1. A method for transmitting an audio signal, comprising: generating a packet of data comprising a portion of a bitstream, the bitstream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks; The generating step comprises: assembling a packet of data having one or more of the plurality of blocks, wherein blocks from different frames are combined into a single packet and / or transmitted out of order; and transmitting the packets of data over a packet-based network; A method comprising:

[0153] EEE-D2. The method of EEE-D1, wherein each block of the plurality of blocks includes information indicating its identity.

[0154] EEE-D3. The method of EEE-D2, wherein the identity indicative information includes at least one of a block ID, a corresponding frame number associated with the block, and / or a retransmission priority.

[0155] EEE-D4. The method of any one of EEE-D1 to EEE-D3, wherein each frame of the plurality of frames carries all audio data representing a continuous segment of an audio signal, having a start time, an end time and a duration.

[0156] EEE-D5. A method for decoding an audio signal, comprising: receiving a packet of data comprising a portion of a bitstream, the bitstream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks; determining a set of blocks addressed by the device from among the plurality of blocks; decrypting the set of blocks addressed by the device and skipping decoding of blocks of the plurality of blocks that are not addressed by the device; A method comprising:

[0157] EEE-D6. A method for transmitting an audio stream, comprising: 10. A method comprising: transmitting the audio stream, the audio stream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks, the transmitting comprising transmitting configuration information regarding the audio stream out-of-band.

[0158] EEE-D7. Transmitting configuration information about the audio stream out-of-band transmitting the audio stream over a first network and / or a first network protocol; transmitting the configuration information over a second network and / or a second network protocol; The method according to EEE-D6, comprising:

[0159] EEE-D8. The method of EEE-D7, wherein the first network protocol is User Datagram Protocol (UDP) and the second network protocol is Transmission Control Protocol (TCP).

[0160] EEE-D9. A method for decoding an audio signal, comprising: receiving a bitstream including information corresponding to signaling of static configuration aspects and static metadata; mapping one or more channel elements to one or more devices based on the information and / or static metadata; A method comprising:

[0161] EEE-D10. The method of EEE-D9, wherein the bitstream is received by a plurality of decoders configured to decode the bitstream, each decoder of the plurality of decoders configured to decode a portion of the bitstream.

[0162] EEE-D11. The method of EEE-D9 or EEE-D10, wherein the bitstream further includes dynamic metadata.

[0163] EEE-D12. The bitstream includes a plurality of blocks, each block of the plurality of blocks comprising: information enabling portions of said blocks to be skipped during decoding, which portions are not required by the device; Dynamic metadata; 10. The method of any one of EEE-D9 to EEE-D11, comprising:

[0164] EEE-D13. A method for retransmitting blocks of an audio signal, comprising: transmitting one or more blocks of a bitstream, the bitstream including a plurality of blocks, each of the one or more blocks of the bitstream having been previously transmitted; A method wherein each of the one or more blocks includes a decoding priority indicator.

[0165] EEE-D14. The method of EEE-D13, wherein the decoding priority indicator indicates to a decoder an order of priority for decoding the one or more blocks of the bitstream.

[0166] EEE-D15. The method of EEE-D13 or EEE-D14, wherein each block of the one or more blocks includes the same block ID.

[0167] EEE-D16. The method of any one of EEE-D13 to EEE-D15, wherein the transmission of the one or more blocks of the bitstream is performed by reducing the data rate of the previous transmission.

[0168] EEE-D17. The method of EEE-D16, wherein reducing the data rate includes at least one of reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, and / or reducing the channel count of the audio signal.

[0169] EEE-E1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, said frames including one or more independent blocks of encoded data, said method comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; encoding the one or more audio signals associated with each playback device to obtain one or more encoded audio signals; combining the one or more encoded audio signals associated with each of the playback devices into a first independent block of the frame; encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method comprising:

[0170] The method according to EEE-E1, wherein EEE-E2.2 or higher audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended to be played back by a respective driver of the playback device, and wherein a different encoding technique is used for each of the band-limited signals.

[0171] EEE-E3. The method of EEE-E2, wherein a different psychoacoustic model and / or a different bit allocation technique is used for each of said band-limited signals.

[0172] EEE-E4. The method of any one of EEE-E1 to EEE-E3, wherein the instantaneous frame rate of the encoded signal is variable and is constrained by a buffer fullness model.

[0173] EEE-E5. The method of EEE-E1, wherein encoding one or more audio signals associated with each playback device includes jointly encoding the one or more audio signals associated with each playback device and one or more additional audio signals associated with one or more additional playback devices into the first independent block of the frame.

[0174] EEE-E6. The method of EEE-E5, wherein jointly encoding the one or more audio signals with the one or more additional audio signals comprises sharing one or more scale factors between the two or more audio signals.

[0175] EEE-E7. The method of EEE-E6, wherein the two or more audio signals are spatially related.

[0176] EEE-E8. The method of EEE-E7, wherein the two or more spatially related audio signals comprise a left horizontal channel, a left top channel, a right horizontal channel or a right top channel.

[0177] EEE-E9. Jointly encoding the one or more audio signals and the one or more additional audio signals includes applying a coupling tool, and applying the coupling tool includes: The combination of two or more audio signals into a composite signal above a certain frequency; determining, for each of the two or more audio signals, a scale factor relating energy of the composite signal to energy of each respective corresponding signal; The method according to EEE-E5, comprising:

[0178] EEE-E10. The method of EEE-E5, wherein jointly encoding the one or more audio signals and one or more further audio signals comprises applying a joint encoding tool to more than two signals.

[0179] EEE-E11. A method of decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, said frames comprising one or more independent blocks of encoded data, said method comprising: identifying discrete blocks of encoded data from the encoded bitstream corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; decoding the one or more audio signals associated with the playback device from the discrete blocks of encoded data to obtain one or more decoded audio signals; identifying one or more additional independent blocks of encoded data from the encoded bitstream corresponding to one or more additional audio signals; decoding or skipping said one or more additional independent blocks of encoded data; A method comprising:

[0180] The method according to EEE-E11, wherein two or more audio signals EEE-E12.2 or higher are associated with the playback device, each of the two or more audio signals being a band-limited signal intended to be played back by a respective driver of the playback device, and wherein different decoding techniques are used to decode the two or more audio signals.

[0181] EEE-E13. The method of EEE-E12, wherein different psychoacoustic models and / or different bit allocation techniques are used to encode each of said band-limited signals.

[0182] EEE-E14. The method of EEE-E11 or EEE-E13, wherein the instantaneous frame rate of the encoded signal is variable and is constrained by a buffer fullness model.

[0183] EEE-E15. The method of EEE-11, wherein decoding the one or more audio signals associated with the playback device comprises jointly decoding the one or more audio signals associated with the respective playback device and one or more additional audio signals associated with one or more additional playback devices from the independent blocks of encoded data.

[0184] EEE-E16. The method of EEE-15, wherein jointly decoding the one or more audio signals and one or more additional audio signals comprises extracting scale factors shared between two or more audio signals.

[0185] EEE-E17. The method of EEE-E16, wherein the two or more audio signals are spatially related.

[0186] EEE-E18. The method of EEE-E17, wherein the two or more spatially related audio signals comprise a left horizontal channel, a left top channel, a right horizontal channel or a right top channel.

[0187] EEE-E19. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and one or more further audio signals comprises applying a decoupling tool.

[0188] EEE-E20. The decoupling tool comprises: Extracting the independent decoded signal below a certain frequency; extracting a composite signal above the particular frequency; determining corresponding decoupled signals above the particular frequency from the composite signal and a scale factor relating the energy of the composite signal to the energy of each corresponding signal; combining each independently decoded signal with its corresponding decoupled signal into said jointly decoded signal; The method according to EEE-E19, comprising:

[0189] EEE-E21. The method of EEE-E15, wherein jointly decoding the one or more audio signals and one or more further audio signals comprises applying a joint decoding tool to extract more than two audio signals.

[0190] EEE-E22. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device comprises applying bandwidth extension to the audio signals in the same domain in which the audio signals were encoded.

[0191] EEE-E23. The method of EEE-E22, wherein said domain is a modified discrete cosine transform (MDCT) domain.

[0192] EEE-E24. The method of EEE-E22 or EEE-E23, wherein said bandwidth extension comprises adaptive noise addition.

[0193] EEE-E25. An apparatus configured to perform the method described in any one of EEE-E1 to EEE-E24.

[0194] EEE-E26. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the methods described in any one of EEE-E1 to EEE-E24.

[0195] 1. A method for generating an encoded bitstream performed by a device having an EEE-F1.1 or higher microphone, comprising: capturing one or more audio signals with the one or more microphones; determining the presence of a wake word by analyzing the captured audio signal; When the presence of the wake word is detected, setting a flag indicating a speech recognition task to be performed on the captured audio signal; encoding the captured audio signal; assembling the encoded audio signal and the flag into the encoded bitstream; A method comprising:

[0196] EEE-F2. The method of EEE-F1, wherein the one or more microphones are configured to capture a mono or spatial sound field.

[0197] EEE-F3. The method of EEE-F2, wherein said spatial sound field is of A-format or B-format.

[0198] EEE-F4. The method of any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended for use only in performing the speech recognition task.

[0199] EEE-F5. The method of EEE-F4, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient to perform the speech recognition task but not sufficient for human listening.

[0200] EEE-F6. The method of EEE-F4 or EEE-F5, wherein the captured audio signal is converted into a representation comprising one or more of band energies, Mel-frequency ceptostral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients prior to encoding the captured audio signal.

[0201] EEE-F7. The method of any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended both for human listening and for use in the speech recognition task.

[0202] EEE-F8. The method of EEE-F7, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening.

[0203] EEE-F9. The method of EEE-F7, wherein encoding the captured audio signal includes generating a first encoded representation of the captured audio signal and a second encoded representation of the captured audio signal, wherein the first encoded representation is generated such that when the captured audio signal is decoded from the first encoded representation, the quality of the decoded audio signal is sufficient for human hearing, and the second encoded representation is generated such that when the captured audio signal is decoded from the second encoded representation, the quality of the decoded audio signal is sufficient for performing the speech recognition task but not sufficient for human hearing.

[0204] EEE-F10. The method of EEE-F9, wherein generating the second encoded representation of the captured audio signal comprises, prior to encoding the captured audio signal, converting the captured audio signal into one or more of a parametric representation; a coarse waveform representation; or a representation comprising one or more of band energy, Mel-frequency ceptostrum coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0205] EEE-F11. The method of EEE-F9 or EEE-F10, wherein assembling the encoded audio signal into a bitstream comprises inserting the first encoded representation into a first independent block of the encoded bitstream, and inserting the second encoded representation into a second independent block of the encoded bitstream.

[0206] EEE-F12. The method of EEE-F9 or EEE-F10, wherein the first encoded representation is included in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first and second layers are included in a single block of the encoded bitstream.

[0207] EEE-F13. If the presence of the wake word is not detected, raising a flag to indicate that no speech recognition tasks will be performed on the captured audio signal; encoding the captured audio signal; assembling the encoded audio signal and the flag into the encoded bitstream; The method of any one of EEE-F1 to EEE-F12, further comprising:

[0208] EEE-F14. A method for decoding an audio signal, comprising: receiving an encoded bitstream including an encoded audio signal and a flag indicating whether a speech recognition task is to be performed; decoding the encoded audio signal to obtain a decoded audio signal; performing the speech recognition task on the decoded audio signal when the flag indicates that the speech recognition task is to be performed; A method comprising:

[0209] EEE-F15. The method of EEE-F14, wherein the decoded audio signal is intended for use only in performing the speech recognition task.

[0210] EEE-F16. The method of EEE-F15, wherein the quality of the decoded audio signal is sufficient for performing the speech recognition task but not sufficient for human listening.

[0211] EEE-F17. The method of EEE-F15 or EEE-F16, wherein the decoded audio signal is a representation of the captured audio signal, prior to encoding, that includes one or more of band energies, Mel-frequency ceptostral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0212] EEE-F18. The method of EEE-F14, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening.

[0213] EEE-F19. The method of EEE-F18, wherein the encoded audio signal comprises a first encoded representation of one or more audio signals and a second encoded representation of the one or more audio signals.

[0214] EEE-F20. The method according to EEE-F18, wherein the quality of the audio signal decoded from the first representation is sufficient for human hearing and the quality of the audio signal decoded from the second representation is sufficient for performing the speech recognition task but not sufficient for human hearing.

[0215] EEE-F21. The method of EEE-F19 or EEE-F20, wherein the first representation is in a first independent block of the encoded bitstream and the second representation is in a second independent block of the encoded bitstream.

[0216] EEE-F22. The method of EEE-F19 or EEE-F20, wherein the first representation is in a first layer of the encoded bitstream and the second representation is included in a second layer of the encoded bitstream, and the first and second layers are included in a single block of the encoded bitstream.

[0217] EEE-F23. The method of any one of EEE-F18 to EEE-F22, wherein decoding the encoded audio signal comprises decoding only the second representation and ignoring the first representation.

[0218] EEE-F24. The method of any one of EEE-F18 to EEE-F23, wherein the audio signal decoded from the second encoded representation is a parameter representation; a waveform representation; or a representation including one or more of band energies, Mel-frequency keptostrum coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0219] EEE-F25. An apparatus configured to perform the method described in any one of EEE-F1 to EEE-F24.

[0220] EEE-F26. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the methods described in any one of EEE-F1 to EEE-F24.

[0221] EEE-G1. A method of encoding an audio signal of an immersive audio program for low latency transmission to one or more playback devices, comprising: receiving a plurality of time-domain audio signals of the immersive audio program; Select the frame size, extracting the frame of the time domain audio signal in response to the frame size, wherein the frame of the time domain audio signal overlaps with a previous frame of the time domain audio signal; splitting the audio signal into overlapping frames; transforming the frames of a time domain audio signal into a frequency domain signal; encoding the frequency domain signal; quantizing the encoded frequency-domain signal using a perceptually motivated quantization tool; assembling the quantized coded frequency domain signal into one or more independent blocks within the frame; assembling the one or more independent blocks into an encoded frame; A method comprising:

[0222] EEE-G2. The method of EEE-G1, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration.

[0223] EEE-G3. The method of EEE-G2, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2.

[0224] EEE-G4. The method of any one of EEE-G1 to EEE-G3, wherein the plurality of audio signals includes one or more object-based signals.

[0225] EEE-G5. The method of any one of EEE-G1 to EEE-G4, wherein the plurality of audio signals comprises a scene-based representation of the immersive audio program.

[0226] EEE-G6. The method of any one of EEE-G1 to EEE-G5, wherein the selected frame size is one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.

[0227] EEE-G7. The method of any one of EEE-G1 to EEE-G6, wherein the overlap between said frame of the time domain audio signal and said previous frame of the time domain audio signal is less than or equal to 50%.

[0228] EEE-G8. The method of any one of EEE-G1 to EEE-G7, wherein said transform is a modified discrete cosine transform (MDCT).

[0229] EEE-G9. The method of any one of EEE-G1 to EEE-G8, wherein two or more of the plurality of audio signals are jointly encoded.

[0230] EEE-G10. The method of any one of EEE-G1 to EEE-G9, wherein each independent block contains an encoded signal intended for one or more playback devices.

[0231] EEE-G11. The method of any one of EEE-G1 to EEE-G10, wherein at least one independent block comprises encoded signals for two or more playback devices, said encoded signals comprising jointly encoded audio signals.

[0232] EEE-G12. The method of any one of EEE-G1 to EEE-G11, wherein at least one independent block contains multiple encoded signals covering different bandwidths intended for playback from different drivers of a playback device.

[0233] EEE-G13. The method of any one of EEE-G1 to EEE-G12, wherein at least one independent block includes an encoded echo reference signal for use in echo management performed by the playback device.

[0234] EEE-G14. The method of any one of EEE-G1 to EEE-G13, wherein encoding the quantized frequency domain signal comprises applying one or more of the following tools: temporal noise shaping (TNS), joint channel coding, scale factor sharing between signals, deterministic control parameters for high frequency reconstruction, and deterministic control parameters for noise substitution.

[0235] The method of any one of EEE-G1 to EEE-G14, wherein the independent block EEE-G15.1 or higher includes parameters controlling one or more of delay, gain and equalization of the playback device.

[0236] EEE-G16. A low latency method for decoding an audio signal of an immersive audio program from an encoded signal, comprising: receiving an encoded frame including one or more independent blocks; Extracting the quantized coded frequency domain signal from one or more independent blocks; dequantizing the quantized encoded frequency domain signal; decoding the dequantized frequency domain signal; inverse transforming the decoded frequency domain signal to obtain a time domain signal; providing a plurality of audio signals of the immersive audio program by overlapping and adding the time domain signal with a time domain signal from a previous frame; A method comprising:

[0237] EEE-G17. The method of EEE-G16, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration.

[0238] EEE-G18. The method of EEE-G17, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2.

[0239] EEE-G19. The method of any one of EEE-G16 to EEE-G18, wherein the plurality of audio signals includes one or more object-based signals.

[0240] EEE-G20. The method of any one of EEE-G16 to EEE-G19, wherein the plurality of audio signals comprises a scene-based representation of the immersive audio program.

[0241] EEE-G21. The method of any one of EEE-G16 to EEE-G20, wherein the frame of time domain samples comprises one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.

[0242] EEE-G22. The method of any one of EEE-G16 to EEE-G21, wherein the overlap with the previous frame is 50% or less.

[0243] EEE-G23. The method of any one of EEE-G16 to EEE-G22, wherein the inverse transform is an inverse modified discrete cosine transform (IMDCT).

[0244] EEE-G24. The method of any one of EEE-G16 to EEE-G23, wherein each independent block contains a quantized encoded frequency domain signal intended for one or more playback devices.

[0245] EEE-G25. The method according to any one of EEE-G16 to EEE-G24, wherein at least one independent block comprises quantized coded frequency domain signals for two or more playback devices, said quantized coded frequency domain signals being jointly coded audio signals.

[0246] EEE-G26. The method according to any one of EEE-G16 to EEE-G25, wherein at least one independent block comprises a plurality of quantized coded frequency domain signals covering different bandwidths intended to be reproduced from different drivers of a reproduction device.

[0247] EEE-G27. The method of any one of EEE-G16 to EEE-G26, wherein at least one independent block includes an encoded echo reference signal for use in echo management performed by the playback device.

[0248] EEE-G28. The method of any one of EEE-G16 to EEE-G27, wherein decoding the quantized frequency domain signal comprises applying one or more of the following decoding tools: temporal noise shaping (TNS), joint channel decoding, scale factor sharing between signals, high frequency reconstruction, and noise substitution.

[0249] The method of any one of EEE-G16 to EEE-G28, wherein the independent blocks EEE-G29.1 or higher contain parameters controlling one or more of delay, gain and equalization of the playback device.

[0250] EEE-G30. The method according to any one of EEE-G16 to EEE-G29, wherein the method is performed by a playback device and wherein extracting the quantized encoded signal from one or more independent blocks comprises selecting only blocks containing quantized encoded frequency domain signals that are reproduced by the playback device, and ignoring independent blocks containing quantized encoded frequency domain signals that are reproduced by other playback devices.

[0251] EEE-G31. An apparatus configured to perform the method described in any one of EEE-G1 to EEE-G30.

[0252] EEE-G32. A non-transitory computer-readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the methods described in any one of EEE-G1 to EEE-G30.

[0253] In the following claims and in the description herein, the terms "comprising," "comprised of," or "which comprises" are all open terms, meaning the inclusion of at least the subsequent element(s) / feature(s) but not the exclusion of other elements / features. Therefore, when used in a claim, the term "comprising" should not be interpreted as being limited to the means, elements, or steps listed thereafter. For example, the scope of the expression "a device comprising A and B" should not be limited to a device consisting only of elements A and B. The terms "including," "which includes," or "that includes" are also all open terms, meaning the inclusion of at least the subsequent element(s) / feature(s) but not the exclusion of other elements / features. Therefore, "including" is synonymous with "comprising" and means the same thing.

[0254] In the foregoing description of embodiments of the present invention, it is understood that various features may be grouped together in a single embodiment, drawing, or description to streamline the disclosure and / or aid in understanding one or more of the various aspects of the present invention. The disclosed method, however, is not to be interpreted as reflecting an intention requiring more features than are expressly recited in each claim. Rather, as the following claims reflect, aspects of the present invention may lie in less than all features of a single above-disclosed embodiment. As such, the claims following the Detailed Description are expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of the present invention.

[0255] Furthermore, some embodiments described herein may include some features included in other embodiments but not other features. On the other hand, combinations of features from different embodiments are intended to be included in the present invention and form different embodiments. This will be understood by those skilled in the art. For example, in the following claims, any of the embodiments described in the claims can be used in any combination.

[0256] Furthermore, some of the embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or other means for performing that function. Thus, a processor with the necessary instructions for performing such a method or method elements forms a means for performing the method or method elements. Furthermore, apparatus elements described herein are examples of means for performing the functions performed by the elements.

[0257] Furthermore, some of the embodiments described herein disclose related solutions that should be construed as potentially being implemented in distribution and / or transmission systems, such as wired and / or wireless systems, e.g., any electrical, optical, and / or mobile systems (e.g., 3G, 4G, and 5G).

[0258] Thus, while particular embodiments of the present invention have been described, those skilled in the art will recognize that other variations and further modifications may be made, and that all of these variations and modifications are intended to be within the scope of the claims. For example, any formulas described above are merely representative of procedures that may be used. Functions may be added or deleted from block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted from methods described.

[0259] The systems, devices, and methods disclosed above may be implemented as software, firmware, hardware, or a combination thereof. For example, aspects of the present application may be embodied at least in part in a device, a system including more than one device, a method, a computer program product, etc.

[0260] In the case of a hardware implementation, the division of tasks between the functional units mentioned above does not necessarily correspond to a division between physical units: on the contrary, one physical component may have multiple functions, and one task may be performed by the cooperation of several physical components.

[0261] Certain or all components may be implemented as software executed by a single digital processor or microprocessor, or as hardware or application specific integrated circuits. Such software may be distributed on computer readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media).

[0262] As is well known to those skilled in the art, the term "computer storage media" includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data, including, but not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer.

[0263] Additionally, as known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

Claims

1. 1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, the frames including two or more independent blocks of encoded data, the method comprising: receiving, for one or more of the plurality of audio signals, information indicating a playback device with which the one or more audio signals are associated; receiving, for the indicated playback device, information indicating one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method comprising:

2. the plurality of audio signals includes one or more groups of audio signals that are not associated with the playback device or the one or more additional associated playback devices, the method comprising: encoding each of the one or more groups of audio signals that are not associated with the playback device or the one or more additional associated playback devices into individual blocks; combining independent blocks corresponding to each of the one or more groups into the frame of the encoded bitstream; The method of claim 1 further comprising:

3. 3. The method of claim 1, wherein the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended for use as echo references for echo management of the playback devices.

4. The method of claim 3 , wherein the one or more audio signals intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.

5. 5. The method of claim 3 or 4, wherein the one or more audio signals intended to be used as echo references are encoded using a parametric coding tool.

6. The method of claim 1 or 2, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

7. 1. A method of decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising two or more independent blocks of encoded data, the playback device comprising one or more microphones, the method comprising: identifying discrete blocks of encoded data from the encoded bitstream corresponding to the one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; extracting the one or more audio signals associated with the playback device from the identified independent blocks of encoded data; identifying one or more other independent blocks of encoded data from the encoded bitstream corresponding to one or more audio signals associated with one or more other playback devices; extracting the one or more audio signals associated with the one or more other playback devices from the one or more other independent blocks of encoded data; capturing one or more audio signals using the one or more microphones of the playback device; using the one or more extracted audio signals associated with the one or more other playback devices as echo references for echo management of the playback device in response to the one or more captured audio signals; A method comprising:

8. determining that the encoded bitstream includes one or more additional independent blocks of encoded data; ignoring the one or more additional independent blocks of encoded data; and The method of claim 7 further comprising:

9. 9. The method of claim 8, wherein ignoring the one or more additional independent blocks of encoded data comprises skipping the one or more additional independent blocks of encoded data without extracting the one or more additional independent blocks of encoded data.

10. 10. The method of claim 7, wherein the one or more audio signals associated with the one or more other playback devices are specifically intended to be used as echo references for echo management of the playback device.

11. The method of claim 10 , wherein the one or more audio signals specifically intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.

12. 12. The method of claim 10 or 11, wherein the one or more audio signals specifically intended for use as echo references are reconstructed from a parametric representation of the one or more audio signals.

13. 10. The method of claim 7, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

14. 8. The method of claim 7, wherein the encoded signal includes signaling information indicating which of the one or more other playback devices to use as an echo reference for the playback device.

15. The method of claim 14 , wherein the one or more other playback devices indicated by the signaling information for a current frame are different from the one or more other playback devices used as echo references for previous frames.

16. Apparatus configured to perform the method of any one of claims 1 to 15.

17. 16. A non-transitory computer readable storage medium comprising sequences of instructions that, when executed, cause one or more devices to perform the method of any one of claims 1 to 15.