Method, apparatus, and medium for efficient encoding and decoding of audio bitstreams
The method addresses low-latency audio streaming challenges by using device-specific encoding and decoding techniques with independent blocks and metadata, ensuring efficient decoding and flexible rendering across complex setups.
Patent Information
- Application Number
- JP2025519519
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2023-09-15
- Publication Date
- 2025-11-06
AI Technical Summary
Existing audio streaming technologies face challenges in delivering low-latency, high-quality audio signals over wireless links, especially in complex user setups with multiple speakers, where wireless quality is often inconsistent, affecting both audio and embedded information.
A method for generating and decoding audio bitstreams that includes independent blocks for each playback device, using different encoding techniques for band-limited signals and applying joint encoding/decoding tools, with variable frame rates and metadata for flexible rendering and echo management, allowing efficient decoding of relevant parts.
Enables low-latency, high-quality audio distribution across multiple devices with flexible rendering and echo management, optimizing decoding complexity and efficiency by allowing devices to skip irrelevant data blocks.
Smart Images

Figure 2025536466000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 378,500, filed October 5, 2022, and U.S. Provisional Application No. 63 / 534,458, filed August 24, 2023, each of which is incorporated by reference in its entirety.
[0002] Technical Field The present disclosure relates generally to audio signal processing, and more particularly to audio source encoding and decoding for low latency exchange of audio signals of immersive audio programs between devices. [Background technology]
[0003] Streaming audio is common in today's society. It is becoming increasingly demanding as user expectations for quality rise, but user setups are also becoming more complex, not only in terms of the number of speakers but also the type of speakers. Streaming is usually done at least in part over a wireless link, which requires the wireless link to be of good quality, but as many of you have probably experienced, this is not always the case.
[0004] Therefore, there is a need to define an exchange format for use cases where a format is streamed from a cloud / server and then transcoded on the device to a lower latency format more suitable for delivery over a wireless (or possibly wired) link. Example use cases are in-home connectivity and phone-to-car connectivity, but this format can be beneficial in any scenario where low-latency distribution of an audio signal from a single device to one or more connected devices is desired.
[0005] In addition to audio information being transmitted wirelessly and streamed, there may be other types of information embedded in the stream that are also affected by the quality of the wireless link and may have similar drawbacks as audio.
[0006] Therefore, it would be advantageous to overcome the problems associated with wireless streaming for various types of streamed audio combined with other types of information or signals. Summary of the Invention [Problem to be solved by the invention]
[0007] An object of the present disclosure is to at least partially overcome the above problems by using wireless streaming of audio combined with other types of information. [Means for solving the problem]
[0008] According to a first aspect of the present disclosure, there is provided a method for generating a frame of an encoded bitstream of an audio program including a plurality of audio signals, the frame including one or more independent blocks of encoded data, the method comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; encoding the one or more audio signals associated with each playback device to obtain one or more encoded audio signals; combining the one or more encoded audio signals associated with each playback device into a first independent block of the frame; encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; and combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.
[0009] According to a second aspect of the present disclosure, there is provided a method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, the frame including one or more independent blocks of encoded data, the method including: identifying from the encoded bitstream independent blocks of encoded data corresponding to one or more audio signals associated with the playback device; extracting the identified independent blocks of encoded data from the encoded bitstream; decoding the one or more audio signals associated with the playback device from the independent blocks of encoded data to obtain one or more decoded audio signals; identifying from the encoded bitstream one or more additional independent blocks of encoded data corresponding to one or more additional audio signals; and decoding or skipping the one or more additional independent blocks of encoded data.
[0010] According to a third aspect of the present disclosure, there is provided an apparatus configured to perform the method of the first and / or second aspects.
[0011] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium comprising a sequence of instructions that, when executed, cause one or more devices to perform the methods of the first and / or second aspects.
[0012] Further examples of the present disclosure are defined in the dependent claims.
[0013] In some examples, two or more audio signals are associated with a playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and a different encoding technique is used for each of the band-limited signals. In some examples, a different psychoacoustic model and / or a different bit allocation technique is used for each of the band-limited signals. In one example, the instantaneous frame rate of the encoded signals is variable and constrained by a buffer fullness model.
[0014] In some examples, encoding one or more audio signals associated with each playback device includes jointly encoding one or more audio signals associated with each playback device and one or more additional audio signals associated with one or more additional playback devices into a first independent block of a frame. In one example, jointly encoding the one or more audio signals and the one or more additional audio signals includes sharing one or more scale factors across two or more audio signals. In some examples, the two or more audio signals are spatially related. In one example, the two or more spatially related audio signals include a left horizontal channel, an upper-left channel, a right horizontal channel, or an upper-right channel.
[0015] In some examples, jointly encoding the one or more audio signals and one or more additional audio signals includes applying a coupling tool. In some examples, the coupling tool includes combining two or more audio signals into a composite signal above a specified frequency and determining, for each of the two or more audio signals, a scale factor relating the energy of the composite signal to the energy of each respective signal. In some examples, jointly encoding the one or more audio signals and one or more additional audio signals includes applying a joint encoding tool to more than two signals.
[0016] In some examples, two or more audio signals are associated with a playback device, each of the two or more audio signals being band-limited signals intended for playback by a respective driver of the playback device, and different decoding techniques are used to decode the two or more audio signals. In one example, different psychoacoustic models and / or different bit allocation techniques were used to encode each of the band-limited signals.
[0017] For some examples, the instantaneous frame rate of the encoded bitstream is variable and is constrained by a buffer fullness model. 15. The method of claim 11, wherein decoding the one or more audio signals associated with the playback device includes jointly decoding the one or more audio signals associated with the respective playback device and one or more additional audio signals associated with the one or more additional playback devices from the independent blocks of encoded data.
[0018] In one example, jointly decoding the one or more audio signals and one or more additional audio signals includes applying a decoupling tool, which in one example includes extracting the independently decoded signals below a specified frequency, extracting a composite signal above a specified frequency, determining each separated signal above the specified frequency from the composite signal and a scale factor relating the energy of the composite signal and the energy of each signal, and combining each independently decoded signal with each separated signal to obtain a jointly decoded signal.
[0019] In this disclosure, a frame represents an entire time slice of all signals. A block stream represents a collection of signals over the duration of a session. A block represents one frame of a block stream. For digital audio with a given sampling frequency, the frame size is equal to the number of audio samples in a frame for a given audio signal. The frame size typically remains constant over the duration of a session.
[0020] The wake word can include a single word or a phrase containing two or more words in a fixed order.
[0021] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system. [Brief explanation of the drawings]
[0022] Examples of the present disclosure will now be described in more detail with reference to the accompanying drawings, in which like reference numerals are used to refer to like elements. The following figures illustrate various examples, however, one or more implementations are not limited to the examples shown in the figures.
[0023] [Figure 1] An example of low latency transcoding for home connectivity.
[0024] [Figure 2] An example of automotive connectivity audio streaming is shown.
[0025] [Figure 3]Demonstrates enhanced TV audio streaming using wireless devices.
[0026] [Figure 4] This example demonstrates audio streaming using a simple wireless speaker.
[0027] [Figure 5] An example of bitstream element to device metadata mapping is shown.
[0028] [Figure 6] This example shows how to deploy a simple flexible rendering of streaming audio.
[0029] [Figure 7] An example of mapping flexible rendering data to bitstream elements is shown.
[0030] [Figure 8] 10 shows an example of echo reference signaling when using multiple devices to play audio and listen for commands.
[0031] [Figure 9] Here is an example of how frames, blocks, and packets relate to each other:
[0032] [Figure 10] An example of additional codec information in the current MPEG-4 structure is shown below.
[0033] [Figure 11] Further examples of codec information in the current MPEG-4 structure are shown below.
[0034] [Figure 12] An example of integrating listening skills and speech recognition is given.
[0035] [Figure 13]1 shows an example of a frame containing multiple blocks.
[0036] [Figure 14] 1 shows an example of a bitstream containing multiple blocks.
[0037] [Figure 15] An example of blocks with different priorities is shown below.
[0038] [Figure 16] 1 shows an example of a bitstream containing multiple blocks with different priorities.
[0039] [Figure 17] Examples of frames with different priorities are shown below.
[0040] [Figure 18] 1 shows an example of a bitstream containing frames with different priorities. DETAILED DESCRIPTION OF THE INVENTION
[0041] The principles of the present invention will now be described with reference to various examples shown in the drawings. It should be understood that the depictions of these embodiments are merely intended to enable those skilled in the art to better understand and further practice the present invention, and are not intended to limit the scope of the invention in any way.
[0042] In FIG. 1, an immersive audio stream is streamed from a cloud or server 10 and decoded on a TV or hub device 20. The immersive audio stream may be encoded in any existing format, including, for example, Dolby Digital Plus, AC-4, etc. The output is then transcoded into a low-latency interchange format for further transmission to a connected device 30. The connected device is preferably connected through a local wireless connection, such as a WiFi soft access point, or a Bluetooth connection. Low latency typically depends on various factors, such as frame size, sampling rate, and hardware and / or software computational resources, but low latency is typically less than 40 ms, 20 ms, or 10 ms.
[0043] In another typical use case, a phone fetches an immersive audio stream from a cloud or server, transcodes it into an interchange format, and then transmits it to a connected car. In the illustrated example of Figure 2, a mobile device (e.g., a phone or tablet) 20 connects to server 10, receives the immersive audio stream, performs low-latency transcoding into a low-latency interchange format, and transmits the transcoded signal to a car 30 that supports immersive audio playback. One example of an immersive audio stream is a stream containing audio in Dolby Atmos format, and one example of a car 30 that supports immersive audio playback is a car configured to play the Dolby Atmos immersive format.
[0044] In general, the interchange format preferably has low latency, low encoding and decoding complexity, the ability to scale to high quality, and reasonable coding efficiency. The format also preferably supports configurable latency, allowing latency to be traded for efficiency and also for error resilience, so as to be operable under varying connection conditions.
[0045] Hub or extended display with wireless speakers with or without listening capabilities 3 shows an example of a hub 20 driving a set of wireless speakers 30, or alternatively, a display (e.g., a television or TV) 20, possibly with built-in speakers 30, is augmented with several wireless speakers 30. Augmentation implies that the display 20 is also part of the audio playback in examples where the display 20 includes speakers 30. The wireless speakers / devices 30 may be receiving the same complete signal (broadcast mode, shown on the left side of FIG. 3) or individual streams tailored for specific devices 30 (unicast multipoint mode, shown on the right side of FIG. 3).
[0046] Each speaker may include multiple drivers covering different frequency ranges of the same channel or corresponding to different channels in a canonical mapping. For example, a speaker may have two drivers, one of which may be an upward-firing driver that outputs a signal corresponding to a height channel to emulate a high-positioned speaker. Wireless devices may have listening capabilities (e.g., "smart speakers") and thus may require echo management, which may require the speaker to receive one or more echo-references. The echo-reference may be local (e.g., for the same speaker / device) or alternatively, may represent relevant signals from other nearby speakers / devices.
[0047] Speakers can be located in arbitrary positions, in which case so-called "flexible rendering" can be performed, whereby rendering takes into account the actual positions of the speakers (as opposed to, for example, the canonical assumption that loudspeakers are located in fixed, predefined positions). Flexible rendering can be performed at the hub or TV, and the rendered signals are then sent to the speakers / devices in broadcast mode (where each of the rendered signals is sent to each device, and each respective device extracts and outputs the appropriate signal) or as individual streams to individual devices. Alternatively, flexible rendering can be performed locally on each device, whereby each device receives a full immersive program representation, e.g., a 7.1.4-channel-based immersive representation, and renders from that representation an appropriate output signal for each device.
[0048] Wireless devices may connect to a hub or TV through a soft access point provided by the device or through a local access point in the home, which may impose different requirements on bitrate and latency.
[0049] Mobile Device Projection in Automotive Applications In another use case, a mobile device (e.g., a phone or tablet) 20 fetches an immersive audio stream from the cloud or server 10, transcodes the immersive audio stream into an interchange format, and then transmits the transcoded stream to a connected car 30. In the example of Figure 2, a Dolby Atmos-enabled phone 20 connects to the server 10 and receives the Dolby Atmos stream for low-latency transcoding and transmission to a Dolby Atmos-enabled car 3. In this use case, a higher bit rate may be available than in the living room use case, and the wireless channel may have different characteristics. A possible reason for the higher bit rate in the car use case is the quieter wireless environment, as the car acts as a kind of Faraday cage, shielding the interior environment from external wireless disturbances. In contrast, in a living room, all wireless devices from neighbors, other rooms, etc. add noise to the wireless environment. At the same time, there are typically very few wireless devices in the car competing for the available wireless bandwidth. In contrast, in the living room use case, all wireless devices in the home, including the living room wireless devices, compete for the available bandwidth.
[0050] For this use case, it is envisioned that the signal transcoded by the mobile device and sent to the car could be a channel-based immersive representation, an object-based representation, a scene-based representation (e.g., an Ambisonics representation), or even a combination of different representations. For this example, different rendering architectures, broadcast or multipoint, may not be important, since typically the complete presentation is transferred from the mobile device to a single endpoint (e.g., the car).
[0051] Interchange format description The Immersive Interchange Format is built on the Modified Discrete Cosine Transform (MDCT) with perceptually motivated quantization and coding. It has configurable latency, e.g., support for different transform sizes at a given sampling rate. Exemplary frame sizes are 128, 256, 512, 1024, 120, 240, 480, 960, and 192, 384, and 768 samples at sampling rates of 48 kHz and 44.1 kHz.
[0052] The format may support other channel configurations, including mono, stereo, 5.1, and immersive channel configurations (e.g., but not limited to, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, and 22.2, or any other channel configuration consisting of unique channels as specified in ISO / IEC 23091-3:2018, Table 2). The format may also support object-based audio and scene-based representations, such as Ambisonics (e.g., first-order or higher-order). The format may also use signaling schemes suitable for integration with existing formats (e.g., those described in ISO / IEC 14496-3, sometimes referred to as the MPEG-4 Audio Standard; ISO 14496-1, sometimes referred to as the MPEG-4 Systems Standard; ISO 14496-12, sometimes referred to as the ISO Base Media File Format Standard; and / or ISO 14496-14, sometimes referred to as the MP4 File Format Standard).
[0053] Additionally, the system may have support for the ability to skip some syntax elements and decode only the relevant portions for a given speaker, support for metadata-controlled Flex-rendering aspects such as delay alignment, level adjustment, and equalization, support for the use of smart speakers with listening capabilities (including scenarios where a collection of signals is sent that allow each of the drivers / loudspeakers in a smart speaker to be fed independently) and associated support for echo management through echo reference signaling, support for quantization and encoding in the MDCT domain with both low-overlap and 50%-overlap windows, and / or support for filtering along the frequency axis in the MDCT domain to perform temporal shaping of the quantization noise, e.g., TNS (Temporal Noise Shaping). Thus, in various examples, the overlapping windows may be symmetric or asymmetric.
[0054] The disclosed interchange format may, in some examples, provide improved joint channel coding to take advantage of increased correlation between signals in flexible rendering use cases, improved coding efficiency due to shared scale factors across channel elements, inclusion of high frequency reconstruction and noise addition techniques in the MDCT domain, and various improvements to legacy coding structures to allow better efficiency and suitability for the use cases at hand.
[0055] In some instances of controlling individual playback devices in broadcast mode, skippable blocks and metadata may be used. In a broadcast mode configuration, each wireless device needs to play the relevant portions of the complete audio for the specific driver / channel corresponding to that particular device. This implies that the device needs to know which portions of the stream are relevant to that device and extract those portions from the complete audio stream being broadcast. To enable low-complexity decoding operations, it is preferable that the stream be constructed in such a way that a decoder for a particular device can efficiently skip over irrelevant elements for decoding until the elements that are relevant for that device and a given driver on that device.
[0056] However, it should be noted that there may still be benefits (from a compression perspective) to allowing joint coding between signals destined for different devices (e.g., in the simplest scenario, the left and right channels of a stereo presentation).
[0057] Thus, the format may allow "skippable blocks" in the bitstream to enable efficient decoding of only those parts that are relevant to a particular device, and may include metadata to allow flexible mapping of one or more skippable blocks to particular devices, while retaining the ability to apply joint coding techniques between signals corresponding to different speakers / devices.
[0058] An example of an arbitrary setup is shown in Figure 4. There are three connected wireless devices 31, 32, 33, two of which are single-channel speakers 32, 33 (meaning one audio channel) and one speaker 31 is a more advanced speaker 31 that has three different drivers and therefore operates on three different signals. The first speaker 31 operates on three individual signals in this example, while the second and third speakers 32, 33 operate on a stereo representation (e.g., left and right). Therefore, it may be beneficial to perform joint coding of the signals of the stereo representation.
[0059] Given the above scenario, the format may specify a bitstream that includes skippable blocks so that the first device can extract only the relevant part of the stream to decode the signal for that speaker, while for a stereo pair, decoder complexity and efficiency may be traded off by constructing a signal where each speaker needs to decode two signals to output a single signal (which has the advantage that joint coding can be performed).
[0060] The format specifies a metadata format to allow a general and flexible representation of the mapping of particular ones of a plurality of skippable blocks to one or more devices. This is shown in Figure 5, where the mapping may be represented as a matrix that associates each device 31, 32, 33 with one or more bitstream elements so that a decoder for a given device knows which bitstream elements to output and decode.
[0061] 5, the first block or skip block Blk1 contains three single-channel elements (one for each driver of Device 1 31), and the second block or skip block Blk2 contains a channel pair element that may contain a jointly encoded version of the signal output by Devices 2 32 and 3 33. Device 1 31 extracts the mapping metadata and determines that the signal it needs is within Skip Block 1 Blk1. Therefore, it extracts Skip Block 1 Blk1, decodes the three single-channel elements within it, and provides them to Drivers 1, 2a, and 2b, respectively.
[0062] Further, Device 1 31 ignores Skip Block 2, Blk2. Similarly, Device 2 32 extracts the mapping metadata and determines that the signal it requires is in Skip Block 2, Blk2. Therefore, Device 2 32 skips Skip Block 1, Blk1, and extracts Skip Block 2, Blk2. Device 2 32 decodes the channel pair element and provides the CPE's left channel output to its driver. Similarly, Device 3 33 extracts the mapping metadata and determines that the signal it requires is in Skip Block 2, Blk2. Therefore, Device 3 33 skips Skip Block 1, Blk1, and also extracts Skip Block 2, Blk2. Device 3 33 decodes the channel pair element and provides the CPE's right channel output to its driver.
[0063] In some examples, Device 2 32 may determine that it requires only a subset of the signals from Skip Block 2 Blk2. In such examples, when possible, Device 2 32 may perform only a subset of the operations required to fully decode the signals in Skip Block 2 Blk2. Specifically, in the example of FIG. 5, Device 2 32 may perform only the processing operations required to extract the left channel of the CPE, thus allowing for reduced computational complexity. Similarly, Device 3 33 may perform only the processing operations required to extract the right channel of the CPE.
[0064] 5, the CPEs may be coded using joint channel coding or may be coded using independent channel coding. When the CPEs are coded using independent channel coding, device 2 32 may extract only the first (e.g., left) channel of the CPE, and device 3 33 may extract only the second (e.g., right) channel of the CPE.
[0065] In another example, the CPE's channels may be coded using joint channel coding, in which case Devices 2 32 and 3 33 must extract the CPE's two intermediate channels. However, Device 2 32 may still be able to operate with reduced computational complexity by performing only the operations required to extract the left channel from the intermediate channels. Similarly, Device 3 33 may still be able to operate with reduced computational complexity by performing only the operations required to extract the right channel from the intermediate decoded channels.
[0066] Depending on the details of the joint channel coding, other optimizations may be possible. The identities of each of the different devices / decoders may be defined during a system initialization or setup phase. Such setup is common and typically involves measuring the room acoustics, speaker distance to the sweet spot, etc.
[0067] Flexible rendering (pre-encoding and post-decoding) distribution and application of delays (post-decoder) In use cases where flexible rendering is applied at the hub / TV before sending the associated signal to a particular device / speaker, the rendering may create a signal that is more difficult to code from a coding perspective, for example, when considering joint encoding of such signals. One reason is that flexible rendering may apply different delay, equalization, and / or gain adjustments for different devices (e.g., depending on the placement of the speaker relative to other speakers and the listener). It is also possible to preset information such as gain and delay during initial setup and only perform equalization flexibly. Other variations of preset information and flexible rendering information may be possible in other examples. It should be noted that, in this specification, the term “gain” should be interpreted to mean any level adjustment (e.g., attenuation, amplification, or pass-through) rather than being limited to only certain types of level adjustment (e.g., amplification).
[0068] 6, the right channel 33 and left channel 32 speakers are given different latencies to reflect their different placement relative to the listener (e.g., because speakers 32, 33 may not be equidistant to the listener, different latencies may be applied to the signals output from speakers 32, 33 so that coherent sounds from the different speakers 32, 33 arrive at the listener simultaneously). The introduction of such different latencies to coherent signals intended for playback through different speakers 31, 32, 33 makes joint encoding of such signals difficult.
[0069] To address this challenge, aspects of the flexible rendering process may be parameterized and applied at the endpoint device after the signal is decoded.
[0070] 7, delay and gain values for each device are parameterized and included in the encoded signal sent to each speaker 31, 32, 33. Each signal may be decoded by each device 31, 32, 33, and each device may then introduce the parameterized gain and delay values into each decoded signal.
[0071] 7, in which encoded signals for different devices 31, 32, 33 are sent in separable blocks (e.g., skippable blocks), parameters (e.g., delays and gains) may also be sent in separable blocks, so that devices 31, 32, 33 can extract only the subset of parameters needed for that device 31, 32, 33 and ignore (and skip) those parameters not needed for that device 31, 32, 33. In such a case, each device may be provided with mapping metadata indicating which parameters are included in which blocks.
[0072] Also, note that while Figure 7 only shows delay and gain parameters, other parameters may also be included, such as equalization parameters, which may include, for example, multiple gains applied to different frequency regions, an indication of a predetermined equalization curve applied by the playback devices 31, 32, and 33, one or more sets of infinite impulse response (IIR) or finite impulse response (FIR) filter coefficients, a set of biquad filter coefficients, parameters specifying the characteristics of a parametric equalizer, and other parameters for specifying equalization known to those skilled in the art.
[0073] Furthermore, the parameterization of flexible rendering aspects need not be static but may be dynamic (e.g., as the listener moves while the audio program is playing). Thus, it may be preferable to allow parameters to change dynamically. When one or more parameters change during the audio program, the device may interpolate between the previous and updated delay and / or gain parameters to provide a smooth transition. This may be particularly useful in situations where the system dynamically tracks the listener's position and correspondingly updates the sweet spot for dynamic rendering.
[0074] Also, as mentioned above and will be discussed below, when flexible rendering is applied, an increased level of correlation between channels may occur, which can be exploited by more flexible joint coding.
[0075] Echo reference coding and signaling For use cases such as that outlined in Figure 8, the need for echo management arises when there are multiple devices / speakers 30 operating together receiving signals in either broadcast mode, as shown on the left side of Figure 8, or unicast multipoint mode, as shown on the right side of Figure 8, and at the same time, there are microphones 40 on the devices 30 to enable "listen in" capabilities.
[0076] When performing echo management involving multiple speakers / devices, it can be beneficial to use an echo reference from more than just the local speaker device. As an example, one device may be located close to another, so that signals from the nearby device will affect the echo management of the device with the active microphone. In a broadcast mode use case, each device receives signals for all devices. When a device has signals for other devices, it can be beneficial to use those signals as echo references. To do so, it is necessary to signal to specific devices which signals can be used as echo references for which other devices.
[0077] In one example, this may be done by providing metadata that not only maps the channel / signal to be played by a particular speaker / device (from the entire set), but also the channel / signal to be used as an echo reference (from the entire set) for a particular speaker / device. Such metadata or signaling may be dynamic, allowing the indication of a preferred echo reference to change over time.
[0078] For use cases where each device / speaker receives only the specific signal that it should play, it may be necessary to send an additional signal (e.g., an echo reference signal) to each device / speaker to provide the appropriate echo reference. Again, to do so, it is necessary to provide device-specific signaling so that each device can select the appropriate signal for playback and the appropriate signal for echo management.
[0079] Because the signal for echo management is only used by the device to perform echo management and is not reproduced by the device for a listener, the echo reference signal may be coded or represented differently from the signal intended for reproduction by the device. In particular, because successful echo management can be achieved with signals coded at lower rates than those normally used for reproduction to a listener, additional compression tools, such as parametric representations of the signal, which are not entirely suitable for reproduction to a listener but which capture the necessary characteristics of the audio signal, may provide good echo management at significantly reduced transmission costs.
[0080] Various syntax elements In some examples, blocks are used to optimize audio transport. In the described formats, each frame may be divided into blocks, as described above with respect to skippable blocks. A block may be identified by the frame number to which it belongs, a block ID that may be used to associate consecutive blocks from different frames with the same ID into a block stream, and a priority for retransmission. An example of the above is a stream having multiple frames N-2, N-1, and N, where frame N includes multiple blocks identified as ID1, ID2, and ID3, as shown in FIG. 13. An example of this stream, but in bitstream format, is shown in FIG. 14. In some examples, for an audio signal of an immersive audio program, the block ID of a block may indicate which set of signals within the entire immersive audio program is carried by that block.
[0081] One use case for the described format is the reliable transmission of audio with low latency over wireless networks such as Wi-Fi. For example, Wi-Fi uses a packet-based network protocol. Packet sizes are usually limited; a typical maximum packet size in IP networks is 1500 bytes. The stream's block-based architecture allows flexibility when assembling packets for transmission. For example, packets with smaller frames may be filled with retransmitted blocks from other frames. Large frames can be split on block boundaries before packetization to reduce dependencies between packets on the network protocol layer.
[0082] Figure 9 illustrates the relationship between frames, blocks, and packets. A frame carries audio data, preferably all audio data, representing a contiguous segment of an audio signal having a start time, an end time, and a duration that is the difference between the end time and the start time. Contiguous segments may contain time periods in accordance with ISO / IEC 14496-3, Subpart 4, Section 4.5.2.1.1. Section 4.5.2.1.1 describes the contents of raw_data_block(). A frame may also carry a redundant representation of that segment, for example, encoded at a lower data rate. After encoding, the frame can be divided into blocks. The blocks can be combined into packets for transmission over packet-based networks. Blocks from different frames can be combined into a single packet and / or transmitted out of order.
[0083] In one example, blocks are used to address individual devices. Packets of data are received by individual devices or groups of related devices. The concept of skippable blocks may be used to address individual devices or groups of related devices. Even if the network operates in broadcast mode when transmitting packets to different devices, audio processing (e.g., decoding, rendering, etc.) may be reduced to blocks addressed to that device. All other blocks can simply be skipped, even if received within the same packet. In some examples, blocks may be brought in the correct order based on their decode or presentation time. Retransmitted blocks with a lower priority may be removed if the same block with a higher priority has also been received. The stream of blocks may then be provided to a decoder.
[0084] In some examples, stream and device configurations are sent out-of-band. A codec allows for setting up a connection over which an audio stream is transmitted at a relatively high rate but with low latency. The configuration of such a connection may remain stable for the duration of such a connection. In that case, instead of making the audio stream a component, it can be transmitted out-of-band. For such out-of-band transmission, a different network or even a network protocol may be used. For example, the audio stream may use User Datagram Profile (UDP) for low-latency transmission, while the configuration may use Transmission Control Protocol (TCP) to ensure reliable transmission of the configuration.
[0085] Codec enablement in the context of MPEG-4 audio One particular application of this technology is within MPEG-4 Audio. In MPEG-4 Audio, different AOTs (Audio Object Types) are defined for different codec technologies. New AOTs may be defined for the formats described herein, allowing for format-specific signaling and data. Furthermore, decoder configuration in MPEG-4 is performed in the DecoderSpecificInfo() payload, which carries the AudioSpecificConfig() payload. The latter defines some general signaling that is agnostic to a particular format, such as sampling rate and channel configuration, as well as specific information for a particular AOT. This may make sense for legacy formats where the entire stream is decoded by a single device. However, in broadcast mode, where a single stream is sent to several decoders, each of which decodes only a portion of the stream, advance channel configuration signaling (as a means of setting up the device's output capabilities) may not be optimal.
[0086] Figure 10 shows the traditional MPEG-4 high-level structure (in black), with modifications in gray. To support broadcast use cases, codecSepcificConfig() (where "codec" can be a generic placeholder name) is defined, and signaling can be redefined for specific use cases, thereby including mapping specific channel elements to specific devices, as well as other related static parameters. The MPEG-4 element channelConfiguration having the value "0" is defined as the channel configuration being defined in codecSpecificConfig. Thus, this value can be used to enable revision of the signaling of the channel configuration within the codec-specific configuration.
[0087] Furthermore, in the MPEG-4 sense, the raw payload is specified for the particular decoder at hand, assuming that codecSpecificConfig() is decodable. However, the present format ensures that dynamic metadata is part of the raw payload, and that length information is available for all of the raw payload, allowing decoders to easily skip over elements that are not relevant to a particular device.
[0088] In Figure 11, a portion of an example raw_data_block defined in MPEG-4 is given on the right. The raw data block contains channel elements (single channel elements (SCE) or channel pair elements (CPE)) in a given order. However, a decoder wishing to skip some of these channel elements that may not be relevant to the output device at hand would, in the conventional MPEG-4 audio syntax, have to parse (and to some extent decode) all channel elements to be able to extract the relevant parts. In the new raw_data_block shown on the left of Figure 11, the content consists of skippable blocks, allowing the decoder to skip the irrelevant parts and decode only the channel elements indicated by the metadata as relevant for the device at hand. In one example, the skippable block contains the raw_data_block and related information.
[0089] It is also possible to use blocks for retransmission. Each frame can be divided into blocks, as described in the skippable blocks section above. A block can be identified by the frame number to which it belongs, a block ID that can be used to associate consecutive blocks from different frames with the same ID into a block stream, and a priority for retransmission. This is shown in Figures 15 and 16. For example, a high priority for retransmission (e.g., indicated by priority 0 in Figures 15 and 16) signals that the block is to be prioritized at the receiver over another block with the same block ID and frame counter but a lower priority for retransmission (e.g., indicated by priority 1 in Figures 15 and 16). Note that, as in the examples of Figures 15 and 16, the priority may decrease as the priority index increases (e.g., priority 1 may be a lower priority than priority 0), while in other examples, the priority may increase as the priority index increases (e.g., priority 0 is a lower priority than priority 1). Still other examples will be apparent to those skilled in the art.
[0090] The syntax may support retransmission of audio elements. For retransmission, different quality levels may be supported. Therefore, retransmitted blocks may carry a "priority" flag to indicate which block with the same block ID should be prioritized for the decoder. This is because received blocks with the same frame counter and block ID are redundant and therefore mutually exclusive for the decoder, as shown in Figures 17 and 18.
[0091] Retransmission of a block may be performed at a reduced data rate. Such a reduced data rate may be achieved by reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, reducing the number of channels of the audio signal (e.g., as described in U.S. Pat. No. 11,289,103, which is incorporated by reference in its entirety), or any combination thereof. In order for a decoder to select an audio block that provides the best possible quality, the block that provides the highest quality signal may have the highest decoding priority, the block with the second highest quality signal may have the second highest priority, and so on.
[0092] A block may also be retransmitted at the same quality level, in which case the priority may reflect the latency of the retransmitted block.
[0093] Tools to improve core coding efficiency A further example may allow sharing of MDCT quantization scale factors across channel elements. In some cases, it may be possible to share MDCT scale factors across certain channels to reduce the associated side bit rate. Furthermore, scale factor sharing may be extended to different channel elements, such as 7.1.4 inputs. The use of shared scale factors may be indicated by two signaling bits that allow three active configurations. One possible configuration is to share scale factors in the left horizontal channel, the upper-left channel, the right horizontal channel, and the upper-right channel. To ensure that scale factors are shared only within skip blocks, the specific syntax may be aligned with the concept of skippable blocks.
[0094] It is also possible to have joint coding of more than two channels. In the example shown in Figures 4, 5, 6, and 7, block Blk 1 carries three SCEs, one for each driver in the smart speaker considered in this scenario. In such a situation, it may be beneficial to enable joint coding of more than two channels, not just two channels (as provided by the CPE) (which can be extended to include stereo prediction, e.g., the MDCT-based complex predictive stereo tool described in ISO / IEC 23003-3 called MPEG-D USAC), as described, for example, in the SAP (Stereo Audio Processing) tool introduced in ETSI TS 103 190.
[0095] Similarly, for flexible rendering, note that there may be a high amount of correlation between signals across devices. For such signals, it may be beneficial to construct skip blocks that cover multiple devices and allow joint channel coding to be applied across more than two channels using the tools outlined above.
[0096] Another joint channel coding tool that can be used for certain transform lengths is channel combining, in which combined channel and scale factor information is transmitted for mid- and high-frequency signals. This can provide bitrate reduction for playback in the good quality range. For example, it can be beneficial to use a channel combining tool for frames with a transform length of 256, which corresponds to a frame length of approximately 256 samples for low-latency coding.
[0097] The present disclosure also allows for efficient encoding of band-limited signals. In scenarios where separate signals are transmitted to feed different drivers of a smart speaker, some of these signals may be band-limited, such as in the case of a three-way driver configuration with a woofer, midrange, and tweeter. Thus, efficient encoding of such band-limited signals is desirable, which can translate into specifically tuned psychoacoustic models and bit allocation strategies, as well as potential modifications of syntax, to handle such scenarios with improved coding efficiency and / or reduced computational complexity (e.g., enabling the use of a band-limited IMDCT for the woofer feed).
[0098] In some cases, a combination of high-frequency reconstruction in the MDCT domain and noise addition is possible. Modern audio codecs are typically designed to support parametric coding techniques, and similarly, it is possible to envision including high-frequency reconstruction methods in the MDCT domain to maintain low latency and noise addition methods in the MDCT domain to manage the tone-to-noise ratio in the reconstructed high band. Such parametric coding techniques can be particularly useful for lower operating points, especially in scenarios where retransmissions are performed as part of an FEC (forward error correction) scheme. Here, typically, a main signal is transmitted, followed by a delayed retransmission of the same signal at a lower bit rate.
[0099] These examples further allow for a reduction in peak complexity in the encoder. This would be applicable in the absence of the buffer model limitations for a constant bitrate transport channel. In a constant bitrate with buffer mode, after quantization and counting, it is possible that the resulting number of bits used for the encoded frame exceeds the allowed bit limit. To meet the bit requirement, at least one new, coarser quantization and bit counting stage must be performed. If this limit is more relaxed, the encoder can retain the first quantization result, performed at a slightly higher instantaneous bitrate. Nevertheless, by using the same bit reservoir control mechanism, along with a virtual buffer filling level that is updated according to the encoder's buffer requirements, it is possible to achieve encoder behavior similar to that of a constant bitrate with buffer model. This saves additional quantization and bit counting stages, and the resulting audio quality is the same or better than the constant bitrate case, at the expense of a slight increase in overall bitrate.
[0100] Return Channel Concept For playback and listening use cases, the smart speaker 30 may have one or more microphones 40 (e.g., microphones or microphone arrays) configured to capture a mono or spatial sound field (e.g., mono, stereo, A-format, B-format, or any other isotropic or anisotropic channel format), and the codec needs to be capable of efficient coding of such formats with low latency, so that the same codec is used as is used for the return channel for broadcast / transmission to the smart speaker device 30, e.g., as shown in FIG.
[0101] It should also be noted that while wake word detection is performed on the smart speaker device, speech recognition is typically performed in the cloud, with appropriate segments of recorded speech triggered by wake word detection.
[0102] In this context, it can be advantageous to economize on the overall system complexity by building a speech analysis system on an intermediate format / representation within the codec, for example. For the simplest use cases, where no human-to-human conversation is ongoing and the utterance should not be heard by another human, a suitable representation specific to the speech recognition task can be defined. This could be band energy, Mel-frequency Cepstral Coefficients (MFCCs), etc., or a low-bitrate version of the MDCT-encoded spectrum.
[0103] In use cases where there is ongoing human-to-human conversation and the wake word and required speech recognition are interleaved, it is desirable to be able to extract the relevant portion of the stream for this purpose. Such extraction requires a layered coding structure of simply additional data sent in parallel with the main speech signal. It also assumes a defined transcoding of the existing decoded MDCT spectrum to the relevant representation, thereby simplifying the structure. Essentially, the decoder decodes and outputs human-audible audio until it is signaled (in parallel) that it should also decode a speech recognition representation. This representation can be "peeled" from the layered stream and can simply be an alternative decoding of the same complete data, or simply a decoding and output of an additional representation within the stream. In this use case, the signaling is an enabling portion indicating that the receiving decoder should output the speech recognition-related representation. [Example]
[0104] Itemized implementation form Below, seven sets of non-claimed itemized examples (EEE-A, EEE-B, EEE-C, EEE-D, EEE-E, EEE-F, EEE-G) illustrate aspects of the examples disclosed herein.
[0105] EEE-A1. A method for decoding an audio signal, the method comprising: receiving a bitstream including at least one frame, each frame of the at least one frame including a plurality of blocks; determining, based on device information for an output device, information from the signaling data to identify portions of one or more blocks of the plurality of blocks to be skipped when decoding; decoding the bitstream while skipping the identified portion of the one or more blocks; A method comprising:
[0106] EEE-A2. The method of EEE-A1, wherein the information for identifying portions of one or more blocks of the plurality of blocks to be skipped when decoding includes a matrix associating each output device of a plurality of output devices with one or more bitstream elements.
[0107] EEE-A3. The method of EEE-A2, wherein the one or more bitstream elements are required for decoding of the bitstream for a corresponding associated output device.
[0108] EEE-A4. The method of any one of EEE-A1 to EEE-A3, wherein the output device may include at least one of a wireless device, a mobile device, a tablet, a single channel speaker, and / or a multi-channel speaker.
[0109] EEE-A5. The method of any one of EEE-A1 to EEE-A4, wherein the identified portion comprises at least one block.
[0110] EEE-A6. The method of any one of EEE-A1 to EEE-A5, wherein the output device is a first output device, and further comprising applying a joint coding technique between one or more signals of the bitstream to a second output device and a third output device.
[0111] EEE-A7. The method of any one of EEE-A1 to EEE-A6, wherein the identity of each output device and / or decoder is defined during a system initialization phase.
[0112] EEE-A8. The method of any one of EEE-A1 to EEE-A7, wherein the signaling data is determined from metadata of the bitstream.
[0113] EEE-A9. An apparatus configured to perform the method of any one of EEE-A1 to EEE-A8.
[0114] EEE-A10. A non-transitory computer readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the method of any one of EEE-A1 to EEE-A8.
[0115] EEE-B1. A method for generating an encoded bitstream from an audio program including a plurality of audio signals, the method comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; receiving, for each playback device, information indicative of at least one of a delay, a gain, and an equalization curve associated with said respective playback device; determining a group of two or more related audio signals from the plurality of audio signals; applying one or more joint encoding tools to the two or more related audio signals of the group to obtain a jointly encoded audio signal; combining the jointly coded audio signals, an indication of the playback devices with which the jointly coded audio signals are associated, and an indication of the delays and gains associated with each of the playback devices with which the jointly coded audio signals are associated into a separate block of an encoded bitstream; A method comprising:
[0116] EEE-B2. The method of EEE-B1, wherein the delay, gain, and / or equalization curve associated with each playback device depends on the position of the each playback device relative to the position of a listener.
[0117] EEE-B3. The method of EEE-B1 or EEE-B2, wherein the delay, gain, and / or equalization curve associated with each playback device depends on the position of the each playback device relative to the positions of other playback devices.
[0118] EEE-B4. The method of any one of EEE-B1 to EEE-B3, wherein the delay, gain and / or equalization curve is dynamically variable.
[0119] EEE-B5. The method of EEE-B4, wherein the delay, gain and / or equalization curves are adjusted in response to changes in the position of the listener.
[0120] EEE-B6. The method of EEE-B4 or EEE-B5, wherein the delay, gain, and / or equalization curve is adjusted in response to changes in the position of the playback device.
[0121] EEE-B7. The method of any one of EEE-B4 to EEE-B6, wherein the delay, gain, and / or equalization curve is adjusted in response to a change in the position of one or more of the other playback devices.
[0122] EEE-B8. The method of any one of EEE-B1 to EEE-B7, further comprising determining from said plurality of audio signals audio signals that are not part of said group of two or more related audio signals.
[0123] EEE-B9. The method of EEE-B8, further comprising, for an audio signal that is not part of the group of two or more related audio signals, applying the delay, gain and / or equalization curve associated with the playback device with which the audio signal is associated.
[0124] EEE-B10. The method of EEE-B9, further comprising independently encoding the audio signals that are not part of the group of two or more related audio signals, and combining the independently encoded audio signals and an indication of the playback device with which the independently encoded audio signals are associated into a separate, independently decodable subset of the encoded bitstream.
[0125] EEE-B11. The method of EEE-B8, further comprising independently encoding the audio signals that are not part of the group of two or more related audio signals, and combining the independently coded audio signals, an indication of the playback device with which the independently coded audio signals are associated, and an indication of the delay, gain and / or equalization curves associated with the playback device with which the independently coded audio signals are associated into a separate, independently decodable subset of the encoded bitstream.
[0126] EEE-B12. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, said frames comprising one or more independent blocks of encoded data, the method comprising: identifying, from the encoded bitstream, discrete blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting identified independent blocks of encoded data from the encoded bitstream; determining that the extracted independent blocks of encoded data comprise two or more jointly encoded audio signals; applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device; determining at least one of a delay, a gain, and an equalization curve associated with the playback device from the extracted independent blocks of encoded data; applying the delay, gain, and / or equalization curve associated with the playback device to the one or more audio signals associated with the playback device; A method comprising:
[0127] EEE-B13. The method of EEE-B12, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to the position of a listener.
[0128] EEE-B14. The method of EEE-B12 or EEE-B13, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices.
[0129] EEE-B15. The method of any one of EEE-B12 to EEE-B14, wherein the determined delay, gain and / or equalization curve of the playback device is dynamically variable.
[0130] EEE-B16. The method of EEE-B15, wherein when the determined delay, gain and / or equalization curve associated with the playback device differs from a previously determined delay, gain and / or equalization curve associated with the playback device, the method further comprises interpolating between the previously determined delay, gain and / or equalization curve associated with the playback device and the determined delay, gain and / or equalization curve associated with the playback device.
[0131] EEE-B17. The method of EEE-B16, wherein the determined delay, gain and / or equalization curves differ from the previously determined delay, gain and / or equalization curves due to changes in listener position.
[0132] EEE-B18. The method of EEE-B16 or EEE-B17, wherein the determined delay, gain, and / or equalization curve differs from the previously determined delay, gain, and / or equalization curve due to a change in the position of the playback device.
[0133] EEE-B19. The method of any one of EEE-B16 to EEE-B18, wherein the determined delay, gain, and / or equalization curve differs from the previously determined delay, gain, or equalization curve due to a change in the position of one or more of the other reproduction devices.
[0134] EEE-B20. The frame of the encoded bitstream includes two or more independent blocks of encoded data, the method further comprising: determining that one or more of the independent blocks include audio signals not associated with the playback device; Ignoring the one or more independent blocks containing audio signals not relevant to the playback device. 10. The method of any one of EEE-B12 to EEE-B19, comprising:
[0135] EEE-B21. The method of any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools comprises identifying a subset of the jointly encoded audio signals associated with the playback device and reconstructing only that subset of the jointly encoded audio signals to obtain the one or more audio signals associated with the playback device.
[0136] EEE-B22. The method of any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools comprises reconstructing each of the jointly encoded audio signals, identifying a subset of reconstructed jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of reconstructed jointly encoded audio signals associated with the playback device.
[0137] EEE-B23. An apparatus configured to perform the method of any one of EEE-B1 to EEE-B22.
[0138] EEE-B24. A non-transitory computer readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the method described in any one of EEE-B1 to EEE-B22.
[0139] EEE-C1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, said frames including two or more independent blocks of encoded data, the method comprising: receiving, for one or more of the plurality of audio signals, information indicating a playback device with which the one or more audio signals are associated; receiving, for the indicated playback device, information indicating one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method comprising:
[0140] EEE-C2. The plurality of audio signals includes one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices; encoding each of the one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices into a respective independent block; combining the respective independent blocks for each of the one or more groups into the frame of the encoded bitstream; The method of EEE-C1 further comprises:
[0141] EEE-C3. The method according to EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended for use as an echo reference for performing echo management for the playback devices.
[0142] EEE-C4. The method of EEE-C3, wherein the one or more audio signals intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.
[0143] EEE-C5. The method according to EEE-C3 or EEE-C4, wherein the one or more audio signals intended for use as an echo reference are encoded using parametric coding tools.
[0144] EEE-C6. The method of EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.
[0145] EEE-C7. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising two or more independent blocks of encoded data, the playback device having one or more microphones, the method comprising: identifying, from the encoded bitstream, discrete blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting identified independent blocks of encoded data from the encoded bitstream; extracting the one or more audio signals associated with the playback device from the identified independent blocks of encoded data; identifying, from the encoded bitstream, one or more other independent blocks of encoded data corresponding to one or more audio signals associated with one or more other playback devices; extracting the one or more audio signals associated with the one or more other playback devices from the one or more other independent blocks of encoded data; capturing one or more audio signals using the one or more microphones of the playback device; in response to the one or more captured audio signals, using the one or more extracted audio signals associated with the one or more other playback devices as echo references for performing echo management on the playback device; A method comprising:
[0146] EEE-C8. The method: determining that the encoded bitstream includes one or more additional independent blocks of encoded data; Ignoring the one or more additional independent blocks of encoded data. The method of EEE-C7, further comprising:
[0147] EEE-C9. The method of EEE-C8, wherein ignoring the one or more additional independent blocks of encoded data comprises skipping the one or more additional independent blocks of encoded data without extracting the one or more additional independent blocks of encoded data.
[0148] EEE-C10. A method according to any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are specifically intended for use as an echo reference for performing echo management for the playback devices.
[0149] EEE-C11. The method of EEE-C10, wherein the one or more audio signals specifically intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.
[0150] EEE-C12. The method of EEE-C10 or EEE-C11, wherein the one or more audio signals specifically intended for use as an echo reference are reconstructed from a parametric representation of the one or more audio signals.
[0151] EEE-C13. The method of any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.
[0152] EEE-C14. The method of EEE-C7, wherein the encoded signal includes signaling information indicating the one or more other playback devices to use as an echo reference for the playback device.
[0153] EEE-C15. The method of EEE-C14, wherein the one or more other playback devices indicated by the signaling information for the current frame are different from the one or more other playback devices used as echo references for the previous frame.
[0154] EEE-C16. An apparatus configured to perform the method of any one of EEE-C1 to EEE-C15.
[0155] EEE-C17. A non-transitory computer readable storage medium containing a sequence of instructions that, when executed, cause one or more devices to perform the method of any one of EEE-C1 to EEE-C15.
[0156] EEE-D1. A method for transmitting an audio signal, the method comprising: generating packets of data comprising portions of a bitstream, the bitstream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks, said generating comprising: assembling a packet of data using one or more of the plurality of blocks, wherein blocks from different frames are combined into a single packet and / or transmitted out of order; transmitting said packets of data over a packet-based network; A method comprising:
[0157] EEE-D2. The method of EEE-D1, wherein each block of the plurality of blocks includes identification information.
[0158] EEE-D3. The method of EEE-D2, wherein the identification information includes at least one of a block ID, a corresponding frame number associated with the block, and / or a priority for retransmission.
[0159] EEE-D4. A method according to any one of EEE-D1 to EEE-D3, wherein each frame of the plurality of frames carries all audio data representing a continuous segment of an audio signal having a start time, an end time and a duration.
[0160] EEE-D5. A method for decoding an audio signal, the method comprising: receiving packets of data comprising portions of a bitstream, the bitstream comprising a plurality of frames, each frame of the plurality of frames comprising a plurality of blocks; determining a set of blocks of the plurality of blocks that are destined for a device; decoding the set of blocks destined for the device and skipping decoding the blocks of the plurality of blocks not destined for the device; A method comprising:
[0161] EEE-D6. A method for transmitting an audio stream, the method comprising: transmitting the audio stream, the audio stream including a plurality of frames, each frame of the plurality of frames including a plurality of blocks, and the transmitting includes transmitting configuration information for the audio stream out-of-band. method.
[0162] EEE-D7. Sending out-of-band configuration information about said audio stream comprises: transmitting the audio stream over a first network and / or a first network protocol; transmitting the configuration information over a second network and / or a second network protocol; Method described in EEE-D6.
[0163] EEE-D8. The method of EEE-D7, wherein the first network protocol is User Datagram Protocol (UDP) and the second network protocol is Transmission Control Protocol (TCP).
[0164] EEE-D9. A method for decoding an audio signal, the method comprising: receiving a bitstream, the bitstream comprising: Information corresponding to signaling of static configuration aspects; including static metadata, stages and; mapping one or more channel elements to one or more devices based on the information and / or static metadata; A method comprising:
[0165] EEE-D10. The method of EEE-D9, wherein the bitstream is received by a plurality of decoders configured to decode the bitstream, each decoder of the plurality of decoders configured to decode a portion of the bitstream.
[0166] EEE-D11. The method of EEE-D9 or EEE-D10, wherein the bitstream further comprises dynamic metadata.
[0167] EEE-D12. The bitstream includes a plurality of blocks, each block of the plurality of blocks comprising: information that allows portions of the block that are not needed for a device to be skipped during decoding; Dynamic Metadata and 10. The method of any one of EEE-D9 to EEE-D11, comprising:
[0168] EEE-D13. A method for retransmitting a block of an audio signal, the method comprising: transmitting one or more blocks of a bitstream, the bitstream including a plurality of blocks, each of the one or more blocks of the bitstream having been previously transmitted; each of the one or more blocks includes a decoding priority indicator; method.
[0169] EEE-D14. The method of EEE-D13, wherein the decoding priority indicator indicates to a decoder a priority for decoding the one or more blocks of the bitstream.
[0170] EEE-D15. The method of EEE-D13 or EEE-D14, wherein each block of the one or more blocks includes the same block ID.
[0171] EEE-D16. The method of any one of EEE-D13 to EEE-D15, wherein the transmission of the one or more blocks of the bitstream is transmitted by reducing the data rate compared to a previous transmission.
[0172] EEE-D17. The method of EEE-D16, wherein reducing the data rate includes at least one of reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, and / or reducing the number of channels of the audio signal.
[0173] EEE-E1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, said frames including one or more independent blocks of encoded data, the method comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; encoding one or more audio signals associated with each playback device to obtain one or more encoded audio signals; combining the one or more encoded audio signals associated with each playback device into a first independent block of the frame; encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method comprising:
[0174] EEE-E2. The method of EEE-E1, wherein two or more audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and wherein a different encoding technique is used for each of the band-limited signals.
[0175] EEE-E3. The method of EEE-E2, wherein a different psychoacoustic model and / or a different bit allocation technique is used for each of said band-limited signals.
[0176] EEE-E4. The method of any one of EEE-E1 to EEE-E3, wherein the instantaneous frame rate of the encoded signal is variable and is constrained by a buffer fullness model.
[0177] EEE-E5. The method of EEE-E1, wherein encoding the one or more audio signals associated with each playback device includes jointly encoding the one or more audio signals associated with each playback device and one or more additional audio signals associated with one or more additional playback devices into the first independent block of the frame.
[0178] EEE-E6. The method of EEE-E5, wherein jointly encoding the one or more audio signals and one or more additional audio signals includes sharing one or more scale factors across two or more audio signals.
[0179] EEE-E7. The method of EEE-E6, wherein the two or more audio signals are spatially related.
[0180] EEE-E8. The method of EEE-E7, wherein the two or more spatially related audio signals comprise a left horizontal channel, an upper left channel, a right horizontal channel or an upper right channel.
[0181] EEE-E9. Jointly encoding the one or more audio signals and one or more additional audio signals Combine two or more audio signals into a composite signal above a specified frequency, determining, for each of the two or more audio signals, a scale factor relating the energy of the composite signal to the energy of each respective signal; Method described in EEE-E5.
[0182] EEE-E10. The method of EEE-E5, wherein jointly encoding the one or more audio signals and one or more additional audio signals comprises applying a joint encoding tool to more than two signals.
[0183] EEE-E11. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising one or more independent blocks of encoded data, the method comprising: identifying, from the encoded bitstream, discrete blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting identified independent blocks of encoded data from the encoded bitstream; decoding the one or more audio signals associated with the playback device from the discrete blocks of encoded data to obtain one or more decoded audio signals; identifying, from the encoded bitstream, one or more additional independent blocks of encoded data corresponding to one or more additional audio signals; decoding or skipping the one or more additional independent blocks of encoded data; A method comprising:
[0184] EEE-E12. The method of EEE-E11, wherein two or more audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and wherein different decoding techniques are used to decode the two or more audio signals.
[0185] EEE-E13. The method of EEE-E12, wherein different psychoacoustic models and / or different bit allocation techniques are used to encode each of said band-limited signals.
[0186] EEE-E14. The method of EEE-E11 or EEE-E13, wherein the instantaneous frame rate of the encoded bitstream is variable and constrained by a buffer fullness model.
[0187] EEE-E15. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device includes jointly decoding the one or more audio signals associated with the respective playback device and one or more additional audio signals associated with one or more additional playback devices from the independent blocks of encoded data.
[0188] EEE-E16. The method of EEE-E15, wherein jointly decoding the one or more audio signals and one or more further audio signals comprises extracting a scale factor shared across two or more audio signals.
[0189] EEE-E17. The method of EEE-E16, wherein the two or more audio signals are spatially related.
[0190] EEE-E18. The method of EEE-E17, wherein the two or more spatially related audio signals comprise a left horizontal channel, a top left channel, a right horizontal channel or a top right channel.
[0191] EEE-E19. The method of EEE-E15, wherein jointly decoding the one or more audio signals and one or more further audio signals comprises applying a decoupling tool.
[0192] EEE-E20. The separation tool: Extracting independently decoded signals below a specified frequency; extracting the composite signal above the specified frequency; determining each separated signal above the specified frequency from the composite signal and a scale factor relating the energy of the composite signal and the energy of each signal; combining each independently decoded signal with a respective separated signal to obtain said jointly decoded signal. Method as described in EEE-E19.
[0193] EEE-E21. The method of EEE-E15, wherein jointly decoding the one or more audio signals and one or more additional audio signals comprises applying a joint decoding tool to extract more than two audio signals.
[0194] EEE-E22. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device comprises applying bandwidth extension to the audio signals in the same domain in which they were encoded.
[0195] EEE-E23. The method of EEE-E22, wherein said domain is a modified discrete cosine transform (MDCT) domain.
[0196] EEE-E24. The method of EEE-E22 or EEE-E23, wherein said bandwidth extension comprises adaptive noise addition.
[0197] EEE-E25. An apparatus configured to perform the method according to any one of EEE-E1 to EEE-E24.
[0198] EEE-E26. A non-transitory computer readable storage medium containing sequences of instructions that, when executed, cause one or more devices to perform the method of any one of EEE-E1 to EEE-E24.
[0199] EEE-F1. A method performed by a device having one or more microphones for generating an encoded bitstream, the method comprising: capturing one or more audio signals with the one or more microphones; analyzing the captured audio signal to determine the presence of a wake word; Upon detecting the presence of the wake word: setting a flag to indicate that a speech recognition task should be performed on the captured audio signal; encoding the captured audio signal; assembling the encoded audio signal and the flag into the encoded bitstream. A method comprising:
[0200] EEE-F2. The method of EEE-F1, wherein the one or more microphones are configured to capture a mono or spatial sound field.
[0201] EEE-F3. The method of EEE-F2, wherein said spatial sound field is in A-format or B-format.
[0202] EEE-F4. The method of any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended for use only in performing the speech recognition task.
[0203] EEE-F5. The method of EEE-F4, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient to perform a speech recognition task, but not sufficient for human hearing.
[0204] EEE-F6. The method of EEE-F4 or EEE-F5, wherein the captured audio signal is converted into a representation comprising one or more of band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients before encoding the captured audio signal.
[0205] EEE-F7. The method of any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended for both human hearing and for use in performing speech recognition tasks.
[0206] EEE-F8. The method of EEE-F7, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human hearing.
[0207] EEE-F9. The method of EEE-F7, wherein encoding the captured audio signal includes generating a first encoded representation of the captured audio signal and a second encoded representation of the captured audio signal, the first encoded representation being generated such that when the captured audio signal is decoded from the first encoded representation, the quality of the decoded audio signal is sufficient for human hearing, and the second encoded representation being generated such that when the captured audio signal is decoded from the second encoded representation, the quality of the decoded audio signal is sufficient to perform a speech recognition task but not sufficient for human hearing.
[0208] EEE-F10. The method of EEE-F9, wherein generating a second encoded representation of the captured audio signal comprises converting the captured audio signal to one or more of a parametric representation, a coarse waveform representation, or a representation including one or more of band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients before encoding the captured audio signal.
[0209] EEE-F11. The method of EEE-F9 or EEE-F10, wherein assembling the encoded audio signal into a bitstream includes inserting the first encoded representation into a first independent block of the encoded bitstream and inserting the second encoded representation into a second independent block of the encoded bitstream.
[0210] EEE-F12. The method of EEE-F9 or EEE-F10, wherein the first encoded representation is included in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first and second layers are included in a single block of the encoded bitstream.
[0211] EEE-F13. When the presence of the wake word is not detected: setting the flag to indicate that a speech recognition task should not be performed on the captured audio signal; encoding the captured audio signal; assembling the encoded audio signal and the flag into the encoded bitstream. 10. The method according to any one of claims EEE-F1 to EEE-F12.
[0212] EEE-F14. A method for decoding an audio signal, comprising: receiving an encoded bitstream including an encoded audio signal and a flag indicating whether a speech recognition task should be performed; decoding the encoded audio signal to obtain a decoded audio signal; performing the speech recognition task on the decoded audio signal when the flag indicates that the speech recognition task should be performed; A method comprising:
[0213] EEE-F15. The method of EEE-F14, wherein the decoded audio signal is intended for use only in performing the speech recognition task.
[0214] EEE-F16. The method of EEE-F15, wherein the quality of the decoded audio signal is sufficient to perform a speech recognition task but not sufficient for human hearing.
[0215] EEE-F17. The method of EEE-F15 or EEE-F16, wherein the decoded audio signal is in a representation including one or more of band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients before encoding the captured audio signal.
[0216] EEE-F18. The method of EEE-F14, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human hearing.
[0217] EEE-F19. The method of EEE-F18, wherein the encoded audio signal comprises a first encoded representation of one or more audio signals and a second encoded representation of the one or more audio signals.
[0218] EEE-F20. The method according to EEE-F18, wherein the quality of the audio signal decoded from the first representation is sufficient for human hearing and the quality of the audio signal decoded from the second representation is sufficient to perform a speech recognition task but not sufficient for human hearing.
[0219] EEE-F21. The method of EEE-F19 or EEE-F20, wherein the first representation is in a first independent block of the encoded bitstream and the second representation is in a second independent block of the encoded bitstream.
[0220] EEE-F22. The method of EEE-F19 or EEE-F20, wherein the first representation is in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first and second layers are included in a single block of the encoded bitstream.
[0221] EEE-F23. The method of any one of EEE-F18 to EEE-F22, wherein decoding the encoded audio signal comprises decoding only the second representation and ignoring the first representation.
[0222] EEE-F24. A method according to any one of EEE-F18 to EEE-F23, wherein the audio signal decoded from the second encoded representation is a parametric representation, a waveform representation, or a representation including one or more of band energies, Mel-frequency cepstral coefficients or modified discrete cosine transform (MDCT) spectral coefficients.
[0223] EEE-F25. An apparatus configured to carry out the method according to any one of EEE-F1 to EEE-F24.
[0224] EEE-F26. A non-transitory computer readable storage medium containing a sequence of instructions that, when executed, cause one or more devices to perform the method of any one of EEE-F1 to EEE-F24.
[0225] EEE-G1. A method for encoding an audio signal of an immersive audio program for low latency transmission to one or more playback devices, the method comprising: receiving a plurality of time-domain audio signals of the immersive audio program; Step 1: Select the frame size; extracting a frame of the time-domain audio signal in response to the frame size, wherein the frame of the time-domain audio signal overlaps with a previous frame of the time-domain audio signal; segmenting the audio signal into overlapping frames; converting the frames of time domain audio signals into frequency domain signals; encoding the frequency domain signal; quantizing the encoded frequency domain signal using a perceptually motivated quantization tool; assembling the quantized coded frequency domain signals into one or more independent blocks within the frame; assembling the one or more independent blocks into an encoded frame; A method comprising:
[0226] EEE-G2. The method of EEE-G1, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration.
[0227] EEE-G3. The method of EEE-G2, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2.
[0228] EEE-G4. The method of any one of EEE-G1 to EEE-G3, wherein the plurality of audio signals includes one or more object-based signals.
[0229] EEE-G5. The method of any one of EEE-G1 to EEE-G4, wherein the plurality of audio signals comprises a scene-based representation of the immersive audio program.
[0230] EEE-G6. The method of any one of EEE-G1 to EEE-G5, wherein the selected frame size is one of 128, 256, 512, 1024, 120, 240, 480, or 960 samples.
[0231] EEE-G7. The method of any one of EEE-G1 to EEE-G6, wherein the overlap between said frame of the time-domain audio signal and said previous frame of the time-domain audio signal is less than or equal to 50%.
[0232] EEE-G8. The method of any one of EEE-G1 to EEE-G7, wherein said transform is a modified discrete cosine transform (MDCT).
[0233] EEE-G9. The method of any one of EEE-G1 to EEE-G8, wherein two or more of the plurality of audio signals are jointly encoded.
[0234] EEE-G10. The method of any one of EEE-G1 to EEE-G9, wherein each independent block contains encoded signals for one or more playback devices.
[0235] EEE-G11. The method of any one of EEE-G1 to EEE-G10, wherein at least one independent block comprises encoded signals for two or more playback devices, said encoded signals comprising jointly encoded audio signals.
[0236] EEE-G12. The method of any one of EEE-G1 to EEE-G11, wherein at least one independent block contains multiple encoded signals covering different bandwidths intended for playback from different drivers of a playback device.
[0237] EEE-G13. The method of any one of EEE-G1 to EEE-G12, wherein at least one independent block includes an encoded echo reference signal for use in echo management performed by the playback device.
[0238] EEE-G14. The method of any one of EEE-G1 to EEE-G13, wherein encoding the quantized frequency domain signal comprises applying one or more of the following tools: temporal noise shaping (TNS), joint channel coding, sharing scale factors across signals, determining control parameters for high frequency reconstruction, and determining control parameters for noise substitution.
[0239] EEE-G15. The method of any one of EEE-G1 to EEE-G14, wherein the one or more independent blocks include parameters for controlling one or more of delay, gain and equalization of the playback device.
[0240] EEE-G16. A low latency method for decoding an audio signal of an immersive audio program from an encoded signal, comprising: receiving an encoded frame including one or more independent blocks; extracting the quantized coded frequency domain signal from one or more independent blocks; dequantizing the quantized encoded frequency domain signal; decoding the dequantized frequency domain signal; inverse transforming the decoded frequency domain signal to obtain a time domain signal; overlapping and adding the time domain signal with a time domain signal from a previous frame to provide a plurality of audio signals of the immersive audio program; A method comprising:
[0241] EEE-G17. The method of EEE-G16, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration.
[0242] EEE-G18. The method of EEE-G17, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2.
[0243] EEE-G19. The method of any one of EEE-G16 to EEE-G18, wherein the plurality of audio signals includes one or more object-based signals.
[0244] EEE-G20. The method of any one of EEE-G16 to EEE-G19, wherein the plurality of audio signals comprises a scene-based representation of the immersive audio program.
[0245] EEE-G21. The method of any one of EEE-G16 to EEE-G20, wherein the frame of time domain samples comprises one of 128, 256, 512, 1024, 120, 240, 480, or 960 samples.
[0246] EEE-G22. The method of any one of EEE-G16 to EEE-G21, wherein the overlap with the immediately preceding frame is 50% or less.
[0247] EEE-G23. The method of any one of EEE-G16 to EEE-G22, wherein the inverse transform is an inverse modified discrete cosine transform (IMDCT).
[0248] EEE-G24. The method of any one of EEE-G16 to EEE-G23, wherein each independent block contains a quantized encoded frequency domain signal for one or more playback devices.
[0249] EEE-G25. The method of any one of EEE-G16 to EEE-G24, wherein at least one independent block contains quantized coded frequency domain signals for two or more playback devices, and wherein the quantized coded frequency domain signals are jointly coded audio signals.
[0250] EEE-G26. The method of any one of EEE-G16 to EEE-G25, wherein at least one independent block comprises a plurality of quantized coded frequency domain signals covering different bandwidths intended for playback from different drivers of a playback device.
[0251] EEE-G27. The method of any one of EEE-G16 to EEE-G26, wherein at least one independent block includes an encoded echo reference signal for use in echo management performed by the playback device.
[0252] EEE-G28. The method of any one of EEE-G16 to EEE-G27, wherein decoding the dequantized frequency domain signal comprises applying one or more of the following decoding tools: temporal noise shaping (TNS), joint channel decoding, sharing scale factors across signals, high frequency reconstruction and noise substitution.
[0253] EEE-G29. The method of any one of EEE-G16 to EEE-G28, wherein the one or more independent blocks contain parameters for controlling one or more of delay, gain and equalization of the playback device.
[0254] EEE-G30. A method according to any one of EEE-G16 to EEE-G29, performed by a playback device and wherein extracting the quantized coded signal from one or more independent blocks comprises selecting only blocks containing quantized coded frequency domain signals for playback by said playback device, and ignoring independent blocks containing quantized coded frequency domain signals for playback by other playback devices.
[0255] EEE-G31. An apparatus configured to perform the method of any one of EEE-G1 to EEE-G30.
[0256] EEE-G32. A non-transitory computer readable storage medium containing a sequence of instructions that, when executed, cause one or more devices to perform the method of any one of EEE-G1 to EEE-G30.
[0257] In the following claims and in the description of this specification, the terms "comprise," "consist of," or "have" are all open terms meaning to include at least the element / feature that it follows, but not to exclude others. Thus, when used in the claims, the term "comprise" should not be interpreted as being limited to the means or elements or steps listed. For example, the scope of the expression "device including A and B" should not be limited to a device consisting only of elements A and B. As used herein, the terms "comprise," "include," or "comprising" are all open terms meaning to include at least the element / feature that it follows, but not to exclude others. Thus, "comprise" is synonymous with "have" and means "have."
[0258] In the foregoing description of examples of the present invention, it should be understood that various features may be grouped together in a single example, figure, or description thereof to facilitate flow of the disclosure and to aid in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that more features are required than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in fewer than all features of a single, foregoing disclosed example. Thus, the claims following the detailed description are expressly incorporated into this detailed description, with each claim standing on its own as a separate example of the present invention.
[0259] Furthermore, some examples described herein include some features that are included in other examples but not others. However, as will be understood by one of ordinary skill in the art, combinations of features from different examples are intended to be encompassed and form different examples. For example, in the following claims, any of the claimed examples may be used in any combination.
[0260] Furthermore, some of the examples are described herein as methods or combinations of elements of methods that may be implemented by a processor of a computer system or by other means for performing a function. Thus, a processor with the necessary instructions for performing such a method or element of a method forms a means for performing the method or element of a method. Furthermore, a described element of an apparatus herein is an example of a means for performing the function performed by that element.
[0261] Furthermore, some of the examples described herein disclose connected solutions that should be construed as potentially being implemented in distribution and / or transmission systems, such as wired and / or wireless systems, for example, by using any electrical, optical, and / or mobile system, such as 3G, 4G, and 5G.
[0262] Thus, while particular examples of the present invention have been described, those skilled in the art will recognize that other and further modifications may be made, and it is intended to claim all such changes and modifications. For example, any formulas above are merely representative of procedures that may be used. Functions may be added to or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted in the methods described.
[0263] The systems, devices, and methods disclosed above may be implemented as software, firmware, hardware, or a combination thereof. For example, aspects of the present application may be embodied, at least in part, in a device, a system including two or more devices, a method, a computer program product, etc.
[0264] In hardware implementation, the division of tasks among functional units mentioned in the above description does not necessarily correspond to a division into physical units; conversely, one physical component may have multiple functions and one task may be performed by several physical components working together.
[0265] Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application specific integrated circuits. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transitory media).
[0266] As known to those skilled in the art, the term computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer.
[0267] Additionally, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and is known to those skilled in the art to include any information delivery media.
Claims
1. 1. A method for generating frames of an encoded bitstream of an audio program including a plurality of audio signals, the frames including one or more independent blocks of encoded data, the method comprising: receiving, for each of the plurality of audio signals, information indicating a playback device with which the respective audio signal is associated; encoding one or more audio signals associated with each playback device to obtain one or more encoded audio signals; combining the one or more encoded audio signals associated with each playback device into a first independent block of the frame; encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream; A method comprising:
2. 10. The method of claim 1, wherein two or more audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and wherein a different encoding technique is used for each of the band-limited signals.
3. The method of claim 2 , wherein a different psychoacoustic model and / or a different bit allocation technique is used for each of the band-limited signals.
4. 4. The method of claim 1, wherein the instantaneous frame rate of the encoded signal is variable and is constrained by a buffer fullness model.
5. 2. The method of claim 1, wherein encoding the one or more audio signals associated with each playback device comprises jointly encoding the one or more audio signals associated with each playback device and one or more additional audio signals associated with one or more additional playback devices into the first independent block of the frame.
6. 6. The method of claim 5, wherein jointly encoding the one or more audio signals and the one or more additional audio signals comprises sharing one or more scale factors across two or more audio signals.
7. The method of claim 6 , wherein the two or more audio signals are spatially related.
8. The method of claim 7 , wherein the two or more spatially related audio signals comprise a left horizontal channel, an upper left channel, a right horizontal channel, or an upper right channel.
9. Jointly encoding the one or more audio signals and one or more additional audio signals comprises: Combine two or more audio signals into a composite signal above a specified frequency, determining, for each of the two or more audio signals, a scale factor relating the energy of the composite signal to the energy of each respective signal; The method of claim 5.
10. The method of claim 5 , wherein jointly encoding the one or more audio signals and one or more additional audio signals comprises applying a joint encoding tool to more than two signals.
11. 1. A method for decoding one or more audio signals associated with a playback device from frames of an encoded bitstream, the frames comprising one or more independent blocks of encoded data, the method comprising: identifying, from the encoded bitstream, independent blocks of encoded data corresponding to the one or more audio signals associated with the playback device; extracting identified independent blocks of encoded data from the encoded bitstream; decoding the one or more audio signals associated with the playback device from the discrete blocks of encoded data to obtain one or more decoded audio signals; identifying, from the encoded bitstream, one or more additional independent blocks of encoded data corresponding to one or more additional audio signals; decoding or skipping the one or more additional independent blocks of encoded data; A method comprising:
12. 12. The method of claim 11, wherein two or more audio signals are associated with the playback device, each of the two or more audio signals being a band-limited signal intended for playback by a respective driver of the playback device, and wherein different decoding techniques are used to decode the two or more audio signals.
13. The method of claim 12 , wherein different psychoacoustic models and / or different bit allocation techniques are used to encode each of the band-limited signals.
14. 14. The method of claim 12 or 13, wherein the instantaneous frame rate of the encoded bitstream is variable and is constrained by a buffer fullness model.
15. 12. The method of claim 11 , wherein decoding the one or more audio signals associated with the playback device comprises jointly decoding, from the independent blocks of encoded data, the one or more audio signals associated with the respective playback device and one or more additional audio signals associated with one or more additional playback devices.
16. 16. The method of claim 15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises extracting a scale factor that is shared across two or more audio signals.
17. The method of claim 16 , wherein the two or more audio signals are spatially related.
18. 20. The method of claim 17, wherein the two or more spatially related audio signals comprise a left horizontal channel, an upper left channel, a right horizontal channel, or an upper right channel.
19. The method of claim 15 , wherein jointly decoding the one or more audio signals and the one or more further audio signals comprises applying a separation tool.
20. The separation tool: Extracting independently decoded signals below a specified frequency; extracting the composite signal above the specified frequency; determining each separated signal above the specified frequency from the composite signal and a scale factor relating the energy of the composite signal and the energy of each signal; combining each independently decoded signal with a respective separated signal to obtain the jointly decoded signal.
20. The method of claim 19.
21. 16. The method of claim 15, wherein jointly decoding the one or more audio signals and one or more additional audio signals comprises applying a joint decoding tool to extract more than two audio signals.
22. 12. The method of claim 11, wherein decoding the one or more audio signals associated with the playback device includes applying bandwidth extension to the audio signals in the same domain in which the audio signals were encoded.
23. 23. The method of claim 22, wherein the domain is a modified discrete cosine transform (MDCT) domain.
24. The method of claim 22 or 23, wherein the bandwidth extension comprises adaptive noise addition.
25. Apparatus configured to carry out the method of any one of claims 1 to 24.
26. 25. A non-transitory computer readable storage medium comprising sequences of instructions that, when executed, cause one or more devices to perform the method of any one of claims 1 to 24.