Method, apparatus and medium for encoding and decoding of audio bitstreams and associated echo reference signals

By generating frames of encoded bitstreams containing multiple audio signals, the delay and quality problems when wireless streaming audio signals are solved, and low-latency audio signal switching and high-quality wireless transmission are realized.

CN120077434APending Publication Date: 2025-05-30DOLBY LABORATORIES LICENSING CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380071307.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-24
Filing Date
2023-09-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, when wireless streaming audio signals, it is difficult to effectively solve the problem of delay and quality when combining audio signals with other types of information.

Method used

Low-latency audio signal exchange is achieved by generating a frame of an encoded bitstream containing a plurality of audio signals, including two or more independent encoded data blocks, respectively for the playback device and the additional associated playback device.

Benefits of technology

It realizes low-latency audio signal exchange, improves the audio signal quality of wireless streaming, and is suitable for scenarios such as in-home connections and automobile connections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077434A_ABST
    Figure CN120077434A_ABST
Patent Text Reader

Abstract

A method for generating a frame of an encoded bitstream of an audio program comprising a plurality of audio signals, where the frame comprises two or more independent blocks of encoded data, the method comprising for one or more audio signals of the plurality of audio signals, receiving information indicative of a playback device associated with the one or more audio signals; for the indicated playback device, receiving information indicating one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 378,498, filed Oct. 5, 2022, and U.S. Provisional Application No. 63 / 578,537, filed Aug. 24, 2023, which are hereby incorporated by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to audio signal processing, and more particularly to audio source encoding and decoding for low - latency exchange of audio signals for immersive audio programs between devices. Background Art

[0004] Audio streaming is common in today's society. As users' expectations for quality increase, the requirements for audio streaming become higher, and user setups have become more complex, with a large number of speakers and different speaker types, possibly within the same setup. Streaming is typically at least partially performed over a wireless link, which then requires the wireless link to have good quality. However, as many people may have experienced, this is not always the case.

[0005] Therefore, there is a need to define an exchange format for use cases where a certain format is streamed from the cloud / server and then transcoded on the device to a more suitable low - latency format for distribution over a wireless (or, in some cases, wired) link. Exemplary use cases are in - home connections and phone - to - car connections, however the format may be beneficial in any scenario where low - latency distribution of an audio signal from a single device to one or more connected devices is desired.

[0006] In addition to audio information being sent and streamed wirelessly, there can be other types of information incorporated into the stream. Such other types of information are also affected by the quality of the wireless link and may have similar drawbacks to audio.

[0007] Therefore, it would be advantageous to overcome the problems associated with wireless streaming of different types of streamed audio combined with other types of information or signals. Summary of the Invention

[0008] An object of the present disclosure is to overcome at least in part the above - mentioned problems regarding wireless streaming of audio combined with other types of information.

[0009] According to a first aspect of the present disclosure, there is provided a method for generating a frame of an encoded bitstream for an audio program comprising a plurality of audio signals, wherein the frame comprises two or more independent encoded data blocks, the method comprising, for one or more of the plurality of audio signals, receiving information indicative of a playback device associated with the one or more audio signals; for the indicated playback device, receiving information indicative of one or more additional associated playback devices; receiving one or more audio signals associated with the indicated one or more additional associated playback devices; encoding the one or more audio signals associated with the playback device; encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block; combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.

[0010] According to a second aspect of the present disclosure, there is provided a method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises two or more independent encoded data blocks, wherein the playback device comprises one or more microphones, the method comprising: identifying from the encoded bitstream an independent encoded data block corresponding to the one or more audio signals associated with the playback device; extracting the identified independent encoded data block from the encoded bitstream; extracting from the identified independent encoded data block the one or more audio signals associated with the playback device; identifying from the encoded bitstream one or more other independent encoded data blocks corresponding to one or more audio signals associated with one or more other playback devices; extracting from the one or more other independent encoded data blocks the one or more audio signals associated with the one or more other playback devices; capturing one or more audio signals using the one or more microphones of the playback device; and using the extracted one or more audio signals associated with the one or more other playback devices as an echo reference for performing echo management on the playback device in response to the captured one or more audio signals.

[0011] According to a third aspect of the present disclosure, there is provided an apparatus configured to perform any one of the first and / or second aspects.

[0012] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium including a sequence of instructions which, when executed, cause one or more devices to perform the method of any one of the first and / or second aspects.

[0013] Further examples of the present disclosure are defined in the dependent claims.

[0014] In some examples, one or more audio signals associated with an associated playback device are specifically intended to be used as an echo reference for performing echo management on the playback device.

[0015] In some examples, less data is used to transmit one or more audio signals intended to be used as an echo reference compared to one or more audio signals associated with the playback device.

[0016] In the present disclosure, a frame represents a time slice of the whole of all signals. A block stream represents a set of signals during a session duration. A block represents a frame of the block stream. For digital audio with a given sampling frequency, the frame size is equivalent to the number of audio samples in a frame of any audio signal. The frame size generally remains constant during the duration of a session.

[0017] A wake word may include a single word, or a phrase containing two or more words in a fixed order.

[0018] Throughout the present disclosure included in the claims, the expression "system" is used in a broad sense to denote a device, a system, or a subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Examples of the present disclosure will be described in more detail with reference to the drawings. In the following drawings, the same reference numerals are used to refer to the same elements. Although the following drawings depict various examples, one or more embodiments are not limited to the examples depicted in the drawings.

[0020] Figure 1 An example of in-home connected low-latency transcoding is shown.

[0021] Figure 2 An example of automotive connected audio streaming is shown.

[0022] Figure 3 An example of enhanced TV audio streaming using a wireless device is shown.

[0023] Figure 4Shows an example of audio streaming with a simple wireless speaker.

[0024] Figure 5 Shows an example of metadata mapping from bitstream elements to a device.

[0025] Figure 6 Shows an example of how to deploy simple and flexible rendering of audio streaming.

[0026] Figure 7 Shows an example of mapping flexible rendering data to bitstream elements.

[0027] Figure 8 Shows an example of signaling of echo reference when playing audio and listening commands with multiple devices.

[0028] Figure 9 Shows an example of how frames, blocks, and packets are related to each other.

[0029] Figure 10 Shows an example of additional codec information in the current MPEG-4 structure.

[0030] Figure 11 Shows an example of still additional codec information in the current MPEG-4 structure.

[0031] Figure 12 Shows an example of integrating listening capabilities and speech recognition.

[0032] Figure 13 Shows an example of a frame including multiple blocks.

[0033] Figure 14 Shows an example of a bitstream including multiple blocks.

[0034] Figure 15 Shows an example of blocks with different priorities.

[0035] Figure 16 Shows an example of a bitstream including multiple blocks with different priorities.

[0036] Figure 17 Shows an example of frames with different priorities.

[0037] Figure 18 Shows an example of a bitstream including frames with different priorities. Detailed Description

[0038] The principles of the present invention will now be described with reference to various examples shown in the accompanying drawings. It should be recognized that the depiction of these embodiments is merely to enable those skilled in the art to better understand and further implement the present invention, and is not intended to limit the scope of the present invention in any way.

[0039] In Figure 1 , an immersive audio stream is streamed from a cloud or server 10 and decoded on a TV or hub device 20. The immersive audio stream can be encoded in any existing format, including, for example, Dolby Digital Plus, AC-4, etc. The output is then transcoded into a low-latency exchange format for further transmission to a connected device 30, which is preferably connected via a local wireless connection (such as a WiFi soft access point or a Bluetooth connection). Low latency generally depends on various factors such as frame size, sampling rate, hardware and / or software computing resources, etc., but the low latency is typically less than 40 ms, 20 ms, or 10 ms.

[0040] In another typical use case, a phone obtains an immersive audio stream from a cloud or server, transcodes it into an exchange format, and then transmits it to a connected car. In Figure 2 the example shown, a mobile device (e.g., a phone or a tablet) 20 connects to a server 10 and receives an immersive audio stream, performs a low-latency transcoding into a low-latency exchange format, and transmits the transcoded signal to a car 30 that supports immersive audio playback. An example of an immersive audio stream is a stream that includes audio in the Dolby Atmos format, and an example of a car 30 that supports immersive audio playback is a car that is configured to playback the Dolby Atmos immersive format.

[0041] Generally, the exchange format preferably has low latency, low encoding and decoding complexity, the ability to scale to high quality, and reasonable encoding efficiency. The format preferably also supports configurable latency such that the latency can be traded off against efficiency and error tolerance in order to be operable under varying connection conditions.

[0042] A hub or enhanced display with wireless speakers with or without listening capabilities

[0043] Figure 3 The following example is shown in Figure 3shown on the left side in) or a separate stream customized for a specific device (unicast-multicast mode, Figure 3 shown on the right side in).

[0044] Each speaker may include multiple drivers, covering different frequency ranges of the same channel or corresponding to different channels in the specification mapping. For example, a speaker may have two drivers, where one driver may be an upward-firing driver whose output corresponds to the signal of the height channel to simulate an overhead speaker. The wireless device may have listening capabilities (e.g., "smart speaker"), and thus may require echo management, which may require the speaker to receive one or more echo references. The echo reference may be local (e.g., for the same speaker / device), or alternatively may represent a relevant signal from other nearby speakers / devices.

[0045] The speakers can be placed in any position, in which case so-called "flexible rendering" can be performed, whereby the rendering takes into account the actual position of the speakers (e.g., contrary to the specification assumption where the speakers are assumed to be located in fixed, predefined positions). Flexible rendering can occur in the hub or TV, and subsequently, the rendered signal is either transmitted to the speakers / devices in broadcast mode, where each rendered signal is transmitted to each device and each corresponding device extracts and outputs the appropriate signal, or is transmitted as a separate stream to a separate device. Alternatively, flexible rendering can occur locally on each device, whereby each device receives a representation of the full immersive program, e.g., an immersive representation based on 7.1.4 channels, and renders the output signal suitable for the corresponding device based on this representation.

[0046] The wireless device can be connected to the hub or TV via a soft access point provided by the device or via a local access point in the home. This may impose different requirements on bitrate and latency.

[0047] Mobile Device Projection in Automotive Applications

[0048] In another use case, a mobile device (e.g., a phone or tablet) 20 obtains an immersive audio stream from the cloud or server 10, transcodes the immersive audio stream into an interchange format, and then transmits the transcoded stream to the connected car 30. In Figure 2In the example, the Dolby Atmos-capable phone 20 connects to the server 10 and receives a Dolby Atmos stream for low-latency transcoding and transmission to the Dolby Atmos-capable car 3. In this use case, a higher bitrate can be utilized compared to the living room use case, and the wireless channel can also have different characteristics. The higher bitrate in the car use case may be attributed to a less noisy wireless environment, as the car acts as a kind of Faraday cage and shields the in-vehicle environment from external wireless interference compared to the living room, where all wireless devices such as those from neighbors, other rooms, etc. add noise to the wireless environment. Also, compared to the living room use case where all wireless devices in the home (including living room wireless devices) typically compete for the available bandwidth, there are usually very few wireless devices in the car competing for the available wireless bandwidth.

[0049] For this use case, the signal to be transcoded by the mobile device and transmitted to the car can be a channel-based immersive representation, an object-based representation, a scene-based representation (e.g., Ambisonics representation), or even a combination of different representations. For this example, different rendering architectures and broadcast vs. multicast may not be relevant as the complete rendering will typically be delivered from the mobile device to a single endpoint (e.g., the car).

[0050] Exchange format description

[0051] The immersive exchange format is based on the modified discrete cosine transform (MDCT) with perceptual excitation quantization and coding. It has a configurable latency, e.g., supporting different transform sizes at a given sampling rate. Exemplary frame sizes are 128, 256, 512, 1024 samples at sampling rates of 48 kHz and 44.1 kHz and 120, 240, 480, 960 and 192, 384, 768 samples.

[0052] This format can support mono, stereo, 5.1, and other channel configurations, including immersive channel configurations (e.g., including but not limited to 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, and 22.2, or any other channel configuration with unique channel compositions as specified in Table - 2 of ISO / IEC 23091 - 3:2018). This format can also support object - based audio and scene - based representations, such as Ambisonics (e.g., first - order or higher - order). This format can also use a signaling scheme suitable for integration with existing formats (e.g., those described in ISO / IEC 14496 - 3, which can also be referred to as the MPEG - 4 audio standard, ISO14496 - 1, which can also be referred to as the MPEG - 4 system standard, ISO14496 - 12, which can also be referred to as the ISO base media file format standard, and / or ISO14496 - 14, which can also be referred to as the MP4 file format standard).

[0053] In addition, the system can support the ability to skip a portion of the syntax elements and only decode the relevant parts for a given speaker, support flexible rendering aspects of metadata control such as delay alignment, level adjustment, and equalization, support the use of smart speakers with listening capabilities (including scenarios where a set of signals is transmitted, allowing each driver / speaker in the smart speaker to be fed independently), and, associated therewith, support echo management through signal transmission via echo reference, support quantization and coding in the MDCT domain with both low - overlap windows and 50% overlap windows, and / or support filtering along the frequency axis in the MDCT domain for temporal shaping of quantization noise, e.g., TNS (temporal noise shaping). Thus, in various examples, the overlap window is symmetric or asymmetric.

[0054] In some examples, the disclosed interchange format can provide improved joint channel coding, leveraging increased correlation between signals in flexible rendering use cases, improved coding efficiency through scale factors shared across channel elements, inclusion of high - frequency reconstruction and noise addition techniques in the MDCT domain, and various improvements over traditional coding structures to allow for better efficiency and applicability to the use cases at hand.

[0055] In some examples of controlling individual playback devices in broadcast mode, skippable blocks and metadata can be used. In a broadcast mode setting, each wireless device will need to play the relevant parts of the complete audio for a specific driver / channel corresponding to that device. This means that the device needs to know which parts of the stream are relevant to it and extract those parts from the complete audio stream being broadcast. To enable low - complexity decoding operations, preferably the stream is structured such that the decoder for a specific device can efficiently skip over elements that are not relevant to decoding and jump to the elements relevant to that device and a given driver on that device.

[0056] However, it should be noted that it may be beneficial (from a compression perspective) to still allow joint encoding between signals destined for different devices (e.g., the left and right channels in a stereo presentation in the simplest case).

[0057] Thus, the format can enable "skippable blocks" in the bitstream to enable efficient decoding of portions related only to specific devices, including metadata that enables flexible mapping of one or more skippable blocks to specific devices while retaining the ability to apply joint encoding techniques between signals corresponding to different speakers / devices.

[0058] In Figure 4 an example of an arbitrary setup is shown. There are three connected wireless devices 31, 32, 33, where two are single-channel speakers 32, 33 (meaning one channel of audio), and one speaker 31 is a more advanced speaker 31 with three different drivers and thus operates on three different signals. In this example, the first speaker 31 operates on 3 separate signals, while the second and third speakers 32, 33 operate on a stereo representation (e.g., left and right). As such, it may be beneficial to jointly encode the signals of the stereo representation.

[0059] Given the above scenario, the format stipulates that the bitstream contains skippable blocks such that the first device can extract only the relevant portion of the stream for decoding the signals of that speaker, while the decoder complexity versus efficiency can be traded off for the stereo pair by constructing signals where each speaker needs to decode two signals to output a single signal, with the advantage that joint encoding can be performed.

[0060] The format stipulates a metadata format that enables a general and flexible representation, which maps a specific one of the multiple skippable blocks to one or more devices. This is shown in Figure 5 where the mapping can be represented as a matrix associating each device 31, 32, 33 with one or more bitstream elements such that the decoder for a given device will know what bitstream elements to output and decode.

[0061] For example, in the example of Figure 5 , the first block or skip block Blk1 includes 3 single-channel elements (one for each driver of device 1 31), while the second block or skip block Blk2 contains channel pair elements, which may include a jointly encoded version of the signals to be output by devices 2 32 and 3 33. Device 1 31 extracts the mapping metadata and determines that its required signals are in skip block 1 Blk1. Thus, it extracts skip block 1 Blk1, decodes the three single-channel elements therein, and supplies them to drivers 1, 2a, and 2b respectively.

[0062] In addition, device 1 31 ignores skip block 2, Blk2. Similarly, device 2 32 extracts mapping metadata and determines that its required signals are in skip block 2 Blk2. Thus, device 2 32 skips skip block 1 Blk1 and extracts skip block 2 Blk2. Device 2 32 decodes the channel pair elements and provides the left channel output of the CPE to its driver. Similarly, device 3 33 extracts mapping metadata and determines that its required signals are in skip block 2 Blk2, so device 3 33 skips skip block 1 Blk1 and also extracts skip block 2 Blk2. Device 3 33 decodes the channel pair elements and provides the right channel output of the CPE to its driver.

[0063] In some examples, device 2 32 may determine that it only needs a subset of the signals from skip block 2 Blk2. In such examples, when possible, device 2 32 may only perform a subset of the operations required to fully decode the signals in skip block 2 Blk2. Specifically, in Figure 5 the example of, device 2 32 may only perform those processing operations required to extract the left channel of the CPE, thus enabling a reduction in computational complexity. Similarly, device 3 33 may only perform those processing operations required to extract the right channel of the CPE.

[0064] For example, in Figure 5 the case of, there may be times when the CPE is encoded using joint channel coding, and other times when it is encoded using independent channel coding. During the time when independent channel coding is used to encode the CPE, device 2 32 may only extract the first (e.g., left) channel of the CPE, while device 3 33 may only extract the second (e.g., right) channel of the CPE.

[0065] In another example, joint channel coding may be used to encode the channels of the CPE, in which case devices 2 32 and 3 33 must extract both intermediate channels of the CPE. However, device 2 32 is still able to operate with reduced computational complexity by only performing those operations required to extract the left channel from the intermediate channels. Similarly, device 3 33 is still able to operate with reduced computational complexity by only performing those operations required to extract the right channel from the intermediate decoded channels.

[0066] Depending on the specific circumstances of the joint channel coding, other optimizations may be possible. The identity of each of the different devices / decoders can be defined during the system initialization or setup phase. Such setup is common and typically involves measuring the acoustics of the room, the distance from the speakers to the optimal listening point (sweet-spot), etc.

[0067] Applications of the post - decoder such as distribution and latency of flexible rendering (pre - encoding and post - decoding)

[0068] For use cases where flexible rendering is applied in a hub / TV before sending relevant signals to a specific device / speaker, the rendering may create signals that are more difficult to encode from an encoding perspective, such as signals that are more difficult to encode when considering joint encoding of such signals. One reason is that flexible rendering can apply different delays, equalization, and / or gain adjustments to different devices (e.g., depending on the placement of the speaker relative to other speakers and the listener). Information such as gain and delay can also be pre - set during initial setup, and only equalization can be flexibly rendered. In other examples, other variants of pre - setting information and flexible rendering information are also possible. It should be noted that in this document, the term "gain" should be interpreted to mean any level adjustment (e.g., attenuation, amplification, or pass - through), rather than being limited to certain level adjustments (e.g., amplification).

[0069] In Figure 6 the example, the right - channel 33 and left - channel 32 speakers are given different delays, reflecting different placements relative to the listener (e.g., since speakers 32, 33 may not be equidistant from the listener, different delays can be applied to the signals output from speakers 32, 33 so that coherent sounds from different speakers 32, 33 reach the listener simultaneously). Introducing such different delays to coherent signals intended for playback on different speakers 31, 32, 33 makes the joint encoding of such signals challenging.

[0070] To address this challenge, aspects of the flexible rendering process can be parameterized and applied to the endpoint device after decoding of the signal.

[0071] In Figure 7 the example, the delay and gain values for each device are parameterized and included in the encoded signal sent to the corresponding speakers 31, 32, 33. The individual signals can be decoded by the corresponding devices 31, 32, 33, which can then introduce the parameterized gain and delay values into the corresponding decoded signals.

[0072] In one example, such as Figure 7 the example shown, the encoded signals for different devices 31, 32, 33 can be transmitted in separable blocks (e.g., skip - able blocks), and parameters (e.g., delay and gain) can also be transmitted in separable blocks, such that devices 31, 32, 33 can extract only the subset of parameters required by that device 31, 32, 33 and ignore (and skip) those parameters not needed by that device 31, 32, 33. In this case, mapping metadata indicating which parameters are included in which blocks can be provided to each device.

[0073] It should also be noted that although Figure 7 only indicates delay and gain parameters, other parameters such as equalization parameters may also be included. Equalization parameters may include, for example, multiple gains to be applied to different frequency regions, an indication of a predetermined equalization curve to be applied by playback devices 31, 32, 33, one or more sets of infinite impulse response (IIR) or finite impulse response (FIR) filter coefficients, a set of biquadratic filter coefficients, parameters specifying the characteristics of a parametric equalizer, and other parameters for specifying equalization known to those skilled in the art.

[0074] In addition, the parameterization in flexible rendering need not be static, but may be dynamic (e.g., in the case where a listener moves during the playback of an audio program). Therefore, it may be preferable to allow the parameters to change dynamically. In the case where one or more parameters change during an audio program, the device may interpolate between the previous delay and / or gain parameters and the updated delay and / or gain parameters to provide a smooth transition. This may be particularly useful in situations where the system dynamically tracks the position of the listener and updates the optimal listening point for dynamic rendering accordingly.

[0075] It should also be noted and will be described subsequently that when flexible rendering is applied, an increase in the level of correlation between channels may occur, and this can be exploited by more flexible joint coding.

[0076] Echo reference coding and signaling

[0077] For the use case as Figure 8 outlined, in which there are multiple devices / speakers 30 operating together to receive signals in the broadcast mode shown on the left or in the unicast-multipoint mode shown on the right, and there are microphones 40 on the devices 30 to enable the "listening" capability, there is a need for echo management. Figure 8 Figure 8

[0078]

[0079] When performing echo management for multiple speakers / devices, it may be beneficial to use echo references from not only the local speaker device. As an example, one device may be placed close to another device, so the signal from the nearby device will affect the echo management of the device with an active microphone. In the broadcast mode use case, each device receives the signals of all devices. When a device has the signals of other devices, it may be beneficial to use those signals as echo references. For this purpose, it is necessary to signal to a particular device which signals can be used as echo references for other devices.

[0079] In one example, this can be done by providing metadata that not only maps the channels / signals (from the entire set) to be played by a particular speaker / device, but also maps the channels / signals (from the entire set) to be used as an echo reference for a particular speaker / device. Such metadata or signaling can be dynamic, such that the indication of the preferred echo reference can vary over time.

[0080] For use cases where each device / speaker only receives the specific signals it is to play, in order to provide a suitable echo reference, it may be necessary to transmit additional signals (e.g., echo reference signals) to each device / speaker. Additionally, in order to do so, device-specific signaling needs to be provided so that each device can select the appropriate signals for playback and the appropriate signals for echo management.

[0081] Since the signals for echo management are only used by the device for performing echo management and not played by the device to the listener, the echo reference signals can be encoded or represented differently from the signals intended to be played back by the device. Specifically, since successful echo management can be achieved with signals encoded at a lower rate than is typically used for playing back to the listener, additional compression tools (e.g., parametric representations of the signal, which may not be suitable at all for playing back to the listener but capture the necessary characteristics of the audio signal) can provide good echo management at a significantly reduced transmission cost.

[0082] Classification syntax elements

[0083] In some examples, blocks are used to optimize the transmission of audio. In the described format, each frame can be divided into blocks, as described above for skippable blocks. A block can be identified by the frame number to which it belongs, a block ID that can be used to associate consecutive blocks of different frames with the same ID into a block stream, and a retransmission priority. An example of the above is a stream with multiple frames N-2, N-1, and N, where frame N includes multiple blocks identified as ID1, ID2, ID3, which are shown in Figure 13 In Figure 14 An example of this stream is shown, which is in bitstream format. In some examples, for the audio signals of an immersive audio program, the block ID of a block can indicate which set of signals of the entire immersive audio program is carried by that block.

[0084] One use case of the described format is the reliable transmission of audio with low latency over a wireless network such as Wifi. For example, Wifi uses a packet-based network protocol. The packet size is typically limited. A typical maximum packet size in an IP network is 1500 bytes. The block-based architecture of the stream allows flexibility in assembling packets for transmission. For example, packets with smaller frames can be filled with retransmission blocks from other frames. Large frames can be split at block boundaries before packing to reduce the dependencies between packets at the network protocol layer.

[0085] Figure 9 The relationship between frames, blocks, and packets is shown. A frame carries audio data, preferably all of the audio data, which represents a continuous segment of an audio signal having a start time, an end time, and a duration that is the difference between the end time and the start time. The continuous segment can include a time period according to ISO / IEC 14496-3, Subpart 4, Section 4.5.2.1.1. Section 4.5.2.1.1 describes the content of raw_data_block(). The frame can also carry a redundant representation of the segment, for example, encoded at a lower data rate. After encoding, the frame can be split into blocks. These blocks can be combined into packets for transmission over a packet-based network. Blocks from different frames can be combined in a single packet, and / or can be transmitted out of order.

[0086] In an example, blocks are used to address individual devices. Data packets are received by individual devices or a related group of devices. The concept of skip blocks can be used to address individual devices or a related group of devices. Even if the network operates in a broadcast mode when sending packets to different devices, the processing of the audio (e.g., decoding, rendering, etc.) can be reduced to the blocks addressed to that device. All other blocks, even if received within the same packet, can be simply skipped. In some examples, blocks can be introduced in the correct order based on their decoding or presentation time. If the same block with a higher priority has also been received, the retransmission block with a lower priority can be removed. The block stream can then be fed into a decoder.

[0087] In some examples, the configuration of the stream and the device is transmitted out-of-band. The codec allows the establishment of a connection where the audio stream is transmitted at a fairly high rate but with low latency. The configuration of such a connection can remain stable for the duration of the connection. In that case, instead of making the configuration part of the audio stream, it can be transmitted out-of-band. For such out-of-band transmission, a different network or network protocol can even be used. For example, the audio stream can use the User Datagram Profile (UDP) for low-latency transmission, while the configuration can use the Transmission Control Protocol (TCP) to ensure reliable transmission of the configuration.

[0088] Enable the codec in the context of MPEG-4 audio

[0089] One specific application of this technology is for use within MPEG-4 audio. In MPEG-4 audio, different AOTs (Audio Object Types) are defined for different codec technologies. For the format described in this document, a new AOT can be defined which allows signaling and data specific to that format. Additionally, the configuration of the decoder in MPEG-4 is done within the DecoderSpecificInfo() payload, which in turn carries the AudioSpecificConfig() payload. In the latter, some general signaling that is unaware of the specific format is defined, such as sample rate and channel configuration, as well as specific information for a particular AOT. For traditional formats where the entire stream is decoded by a single device, this may make sense. However, in a broadcast mode where a single stream is transmitted to several decoders, where each corresponding decoder only decodes a part of the stream, the pre-channel configuration signaling (as a means to set the output function of the device) may be sub-optimal.

[0090] Figure 10 A traditional MPEG-4 high-level structure is shown (in black), which is modified to gray. To support the broadcast use case, codecSpecificConfig() is defined (where "codec" can be a general placeholder name), where the signaling is redefined for the specific use case such that specific channel elements can be mapped to specific devices, as well as including other relevant static parameters. The MPEG-4 element channelConfiguration with the value "0" is defined as the channel configuration defined in codecSpecificConfig. Thus, this value can be used to enable the revision of the signaling for the channel configuration within the codec-specific configuration.

[0091] Furthermore, in the spirit of MPEG-4, assuming that codecSpecificConfig() is decodable, the raw payload is specified for the specific decoder at hand. However, the format ensures that the dynamic metadata is part of the raw payload and that length information is available for all raw payloads such that the decoder can easily skip elements that are not relevant to a particular device.

[0092] In Figure 11In the figure on the right, a part of an example raw_data_block as defined in MPEG-4 is given. The raw_data_block contains channel elements (single-channel elements (SCE) or channel pair elements (CPE)) in a given order. However, in the conventional MPEG-4 audio syntax, a decoder that wishes to skip parts of these channel elements that may be irrelevant to the output device at hand will have to parse (and to some extent also decode) all channel elements in order to be able to extract the relevant parts. In Figure 11 In the new raw_data_block shown on the left, the content will consist of skippable blocks such that the decoder can skip the irrelevant parts and only decode the channel elements indicated by the metadata as being relevant to the device at hand. In one example, the skippable blocks include the raw_data_block and associated information.

[0093] Retransmission using blocks will also be possible. Each frame can be divided into blocks as described in the skippable block section above. The blocks can be identified by the frame number to which they belong, a block ID that can be used to associate consecutive blocks of different frames with the same ID to a block stream, and a retransmission priority, as Figure 15 and 16 illustrated. For example, a high retransmission priority (e.g., indicated by priority 0 in Figure 15 and Figure 16 signals that a block will be preferred over another block at the receiver that has the same block ID and frame count but a lower retransmission priority (e.g., indicated by priority 1 in Figure 15 and Figure 16 ). It should be noted that, as in the examples of Figure 15 and Figure 16 , the priority can decrease as the priority index increases (e.g., priority 1 can be a lower priority than priority 0), while in other examples, the priority can increase as the priority index increases (e.g., priority 0 is a lower priority than priority 1). Other examples will be obvious to those skilled in the art.

[0094] The syntax can support retransmission of audio elements. For retransmission, various quality levels can be supported. Thus, the retransmitted blocks can carry a 'priority' flag to indicate which blocks with the same block ID should be preferred by the decoder, since the received blocks with the same frame count and block ID are redundant and thus mutually exclusive for the decoder, as Figure 17 and 18 shown.

[0095] Retransmission of blocks can be performed at a reduced data rate. This reduced data rate can be achieved by reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, reducing the channel count of the audio signal (e.g., as described in U.S. Patent 11,289,103, which is incorporated herein by reference in its entirety), or any combination thereof. To enable the decoder to select the audio block that provides the best possible quality, the block providing the highest quality signal can have the highest decoding priority, the block with the second highest quality signal can have the second highest priority, and so on.

[0096] Blocks can also be retransmitted at the same quality level. In this case, the priority can reflect the latency of the retransmitted block.

[0097] Tools for improving core coding efficiency

[0098] Additional examples can allow sharing of MDCT quantization scale factors across channel elements. In some cases, MDCT scale factors can be shared across certain channels to reduce the associated side bitrate. Additionally, scale factor sharing can be extended to different channel elements, e.g., for 7.1.4 inputs. The use of the shared scale factor can be indicated by two signaling bits that allow three active configurations. One possible configuration would be to share the scale factor among the left horizontal channel, left top channel, right horizontal channel, and right top channel. Specific syntax can be combined with the skip block concept to ensure that the scale factor is shared only within skip blocks.

[0099] There may also be joint coding of more than two channels. In the examples illustrated in Figure 4 , 5 , 6, and 7, block Blk 1 carries three SCEs, one for each of the drivers in the smart speaker considered in this scenario. In such a case, it may be beneficial to enable joint coding of more than two channels (such as provided by CPE, which can be extended to include stereo prediction, e.g., the MDCT-based complex prediction stereo tool described in ISO / IEC 23003-3, known as MPEG-D USAC), rather than just two channels, as described, for example, in the SAP (Stereo Audio Processing) tool introduced in ETSI TS103 190.

[0100] Similarly, it should be noted that in the case of flexible rendering, there may be a high correlation between signals across devices. For such signals, it may be beneficial to construct skip blocks that cover multiple devices and allow the application of joint channel coding on more than two channels using the tools outlined above.

[0101] For certain transform lengths, another joint channel coding tool that can be used is channel coupling, where composite channel and scale factor information is transmitted for medium and high frequencies. This can provide a bit rate reduction for playback in the good quality range. For example, for a frame with a transform length of 256, corresponding to a frame length of approximately 256 samples for low-latency coding, using the channel coupling tool may be beneficial.

[0102] The present disclosure will also allow for efficient coding of band-limited signals. In scenarios where separate signals are sent to different drivers of a smart speaker, some of these signals may be band-limited, such as in the case of a three-way driver configuration with a woofer, midrange, and tweeter, for example. Thus, it is desirable to efficiently code such band-limited signals, which can translate to potentially modified psychoacoustic models and bit allocation strategies and grammars that are specifically tuned to handle such cases with improved coding efficiency and / or reduced computational complexity (e.g., enabling the use of band-limited IMDCT for the woofer feed).

[0103] In some examples, a combination of high-frequency reconstruction and noise addition in the MDCT domain will be possible. Modern audio codecs are typically designed to support parametric coding techniques, and similarly, it is conceivable to include, for example, high-frequency reconstruction methods in the MDCT domain to retain low latency in the MDCT domain and perform a noise addition scheme to manage the tonal noise ratio in the reconstructed high-frequency band. Such parametric coding techniques are particularly useful for lower operating points and especially in scenarios where retransmission occurs as part of an FEC (forward error correction) scheme, in which the main signal is typically transmitted and then the same signal is retransmitted at a lower bit rate in a delayed manner.

[0104] These examples will further allow for a reduction in peak complexity in the encoder. This will be applicable in cases where the buffer model limitations for a constant bit rate transmission channel are not in place. With a constant bit rate in buffer mode, after quantization and counting, the resulting number of bits for the frame to be encoded may be higher than the limit of the allowed bits. At least one new, coarser quantization and bit counting step is required to meet the bit requirement. If the limit is more relaxed, the encoder can retain the first quantization result and will complete with a slightly higher instantaneous bit rate. Similar encoder behavior to a constant bit rate encoder with a buffer model can still be achieved by using the same bit storage control mechanism and the virtual buffer fullness updated according to the encoder that adheres to the buffer requirements. This will save additional quantization and bit counting steps, and compared to the constant bit rate case with the drawback of a slightly increased total bit rate, the resulting audio quality will be the same or better.

[0105] Backhaul channel concept

[0106] For playback and listening use cases, the smart speaker 30 may have one or more microphones 40 (e.g., a microphone or a microphone array) configured to capture a mono or spatial sound field (e.g., mono, stereo, A-format, B-format, or any other isotropic or anisotropic channel format), and a codec is required to be able to efficiently encode such formats in a low-latency manner. Thus, the same codec is used for broadcast / transmission to the smart speaker device 30 as for the backhaul channel, e.g., Figure 12 as shown.

[0107] It should also be noted that while wake word detection is performed on the smart speaker device, speech recognition is typically done in the cloud, where appropriate segments of the recorded speech are triggered by wake word detection.

[0108] In this context, it may be interesting to generally save system complexity by creating a speech analysis system on an intermediate format / representation, e.g., within a codec. For the simplest use cases, where there is no person-to-person conversation, a suitable representation specific to the speech recognition task can be defined since it will not be listened to by another person. This can be band energy, Mel Frequency Cepstral Coefficients (MFCC), etc., or a low-bitrate version of the MDCT-encoded spectrum.

[0109] In use cases where there is an ongoing person-to-person conversation and there are intertwined wake words and required speech recognition, one may also want to be able to extract relevant portions of the stream for this. Such extraction may require a hierarchical coding structure with simple additional data transmitted in parallel with the main speech signal. A code conversion from an existing decoded MDCT spectrum to a defined representation is also envisioned, and doing so simplifies the structure. Substantially, the decoder will decode and output human-audible audio until signaled that it should also (in parallel) decode the speech recognition representation. The representation can be a "stripped" hierarchical stream, simply an alternative decoding of the same complete data, or simply the decoding and output of an additional representation in the stream. In this use case, the signaling is an enabling segment that indicates to the decoder on the receiving side that it should output the speech recognition-related representation.

[0110] Enumerating examples

[0111] In the following, seven groups of enumerated exemplary examples (EEE-A, EEE-B, EEE-C, EEE-D, EEE-E, EEE-F, and EEE-G) of the claims describe aspects of the examples disclosed herein.

[0112] EEE-A1. An audio signal decoding method, the method comprising:

[0113] receiving a bitstream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks;

[0114] Determine information for identifying one or more portions of the plurality of blocks to be skipped during decoding from signaling data based on device information of an output device; and

[0115] Decode the bitstream with the identified portion of the one or more blocks skipped.

[0116] EEE - A2. The method according to EEE - A1, wherein the information for identifying one or more portions of the plurality of blocks to be skipped during decoding includes a matrix associating each of the plurality of output devices with one or more bitstream elements.

[0117] EEE - A3. The method according to EEE - A2, wherein the one or more bitstream elements are required for decoding the bitstream for the corresponding associated output device.

[0118] EEE - A4. The method according to any one of EEE - A1 to EEE - A3, wherein the output device may include at least one of a wireless device, a mobile device, a tablet computer, a mono speaker, and / or a multi - channel speaker.

[0119] EEE - A5. The method according to any one of EEE - A1 to EEE - A4, wherein the identified portion includes at least one block.

[0120] EEE - A6. The method according to any one of EEE - A1 to EEE - A5, wherein the output device is a first output device, and further includes applying joint coding techniques between one or more signals of the bitstream to a second output device and a third output device.

[0121] EEE - A7. The method according to any one of EEE - A1 to EEE - A6, wherein the identity of each output device and / or decoder is defined during a system initialization phase.

[0122] EEE - A8. The method according to any one of EEE - A1 to EEE - A7, wherein the signaling data is determined from metadata of the bitstream.

[0123] EEE - A9. An apparatus configured to perform the method according to any one of EEE - A1 to EEE - A8.

[0124] EEE - A10. A non - transitory computer - readable storage medium including a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of EEE - A1 to EEE - A8.

[0125] EEE-B1. A method for generating an encoded bitstream from an audio program comprising a plurality of audio signals, the method comprising:

[0126] For each of the plurality of audio signals, receiving information indicating a playback device associated with the corresponding audio signal;

[0127] For each playback device, receiving information indicating at least one of a delay, a gain, and an equalization curve associated with the corresponding playback device;

[0128] Determining a group of two or more correlated audio signals from the plurality of audio signals;

[0129] Applying one or more joint encoding tools to the two or more correlated audio signals in the group to obtain a jointly encoded audio signal;

[0130] Combining the jointly encoded audio signal, an indication of the playback device associated with the jointly encoded audio signal, and indications of a delay and a gain associated with the corresponding playback device associated with the jointly encoded audio signal into an independent block of the encoded bitstream.

[0131] EEE-B2. The method according to EEE-B1, wherein the delay, the gain, and / or the equalization curve associated with the corresponding playback device depends on the position of the corresponding playback device relative to the listener's position.

[0132] EEE-B3. The method according to EEE-B1 or EEE-B2, wherein the delay, the gain, and / or the equalization curve associated with the corresponding playback device depends on the position of the corresponding playback device relative to the positions of other playback devices.

[0133] EEE-B4. The method according to any one of EEE-B1 to EEE-B3, wherein the delay, the gain, and / or the equalization curve is dynamically variable.

[0134] EEE-B5. The method according to EEE-B4, wherein the delay, the gain, and / or the equalization curve is adjusted in response to a change in the position of the listener.

[0135] EEE-B6. The method according to EEE-B4 or EEE-B5, wherein the delay, the gain, and / or the equalization curve is adjusted in response to a change in the position of the playback device.

[0136] EEE-B7. The method according to any one of EEE-B4 to EEE-B6, wherein the delay, the gain, and / or the equalization curve is adjusted in response to a change in the position of one or more of the other playback devices.

[0137] EEE-B8. The method according to any one of EEE-B1 to EEE-B7 further includes determining, from the plurality of audio signals, an audio signal that is not part of a group of the two or more relevant audio signals.

[0138] EEE-B9. The method according to EEE-B8 further includes, for an audio signal that is not part of a group of the two or more relevant audio signals, applying a delay, gain, and / or equalization curve associated with a playback device associated with the audio signal.

[0139] EEE-B10. The method according to EEE-B9 further includes independently encoding an audio signal that is not part of a group of the two or more relevant audio signals, and combining the independently encoded audio signal and an indication of a playback device associated with the independently encoded audio signal into a separately independently decodable subset of an encoded bitstream.

[0140] EEE-B11. The method according to EEE-B8 further includes independently encoding an audio signal that is not part of a group of the two or more relevant audio signals, and combining the independently encoded audio signal, an indication of a playback device associated with the independently encoded audio signal, and an indication of a delay, gain, and / or equalization curve associated with a playback device associated with the independently encoded audio signal into a separately independently decodable subset of an encoded bitstream.

[0141] EEE-B12. A method for decoding, from a frame of an encoded bitstream, one or more audio signals associated with a playback device, wherein the frame includes one or more independently encoded data blocks, the method includes:

[0142] Identifying, from the encoded bitstream, independently encoded data blocks corresponding to the one or more audio signals associated with the playback device;

[0143] Extracting the identified independently encoded data blocks from the encoded bitstream;

[0144] Determining that the extracted independently encoded data blocks include two or more jointly encoded audio signals;

[0145] Applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device;

[0146] Determining at least one of a delay, a gain, and an equalization curve associated with the playback device from the extracted independently encoded data blocks;

[0147] Apply the latency, gain, and / or equalization curve associated with the playback device to the one or more audio signals associated with the playback device.

[0148] EEE-B13. The method according to EEE-B12, wherein the determined latency, gain, and / or equalization curve associated with the playback device depends on the position of the playback device relative to the position of the listener.

[0149] EEE-B14. The method according to EEE-B12 or EEE-B13, wherein the determined latency, gain, and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices.

[0150] EEE-B15. The method according to any one of EEE-B12 to EEE-B14, wherein the determined latency, gain, and / or equalization curve of the playback device is dynamically variable.

[0151] EEE-B16. The method according to EEE-B15, wherein when the determined latency, gain, and / or equalization curve associated with the playback device is different from the previously determined latency, gain, and / or equalization curve associated with the playback device, the method further includes interpolating between the previously determined latency, gain, and / or equalization curve associated with the playback device and the determined latency, gain, and / or equalization curve associated with the playback device.

[0152] EEE-B17. The method according to EEE-B16, wherein the determined latency, gain, and / or equalization curve is different from the previously determined latency, gain, and / or equalization curve due to a change in the position of the listener.

[0153] EEE-B18. The method according to EEE-B16 or EEE-B17, wherein the determined latency, gain, and / or equalization curve is different from the previously determined latency, gain, and / or equalization curve due to a change in the position of the playback device.

[0154] EEE-B19. The method according to any one of EEE-B16 to EEE-B18, wherein the determined latency, gain, and / or equalization curve is different from the previously determined latency, gain, or equalization curve due to a change in the position of one or more other playback devices.

[0155] EEE-B20. The method according to any one of EEE-B12 to EEE-B19, wherein a frame of the encoded bitstream includes two or more independent encoded data blocks, and the method further includes:

[0156] Determine that one or more of the independent blocks contain audio signals not associated with the playback device; and

[0157] Ignore the one or more independent blocks that contain audio signals not associated with the playback device.

[0158] EEE-B21. The method according to any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools includes: identifying a subset of the jointly encoded audio signals associated with the playback device, and reconstructing only the subset of the jointly encoded audio signals to obtain the one or more audio signals associated with the playback device.

[0159] EEE-B22. The method according to any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools includes: reconstructing each of the jointly encoded audio signals, identifying a subset of the reconstructed jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of the reconstructed jointly encoded audio signals associated with the playback device.

[0160] EEE-B23. An apparatus configured to perform the method according to any one of EEE-B1 to EEE-B22.

[0161] EEE-B24. A non-transitory computer-readable storage medium including a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of EEE-B1 to EEE-B22.

[0162] EEE-C1. A method for generating a frame of an encoded bitstream of an audio program including a plurality of audio signals, wherein the frame includes two or more independent encoded data blocks, the method comprising:

[0163] For one or more of the plurality of audio signals, receive information indicating a playback device associated with the one or more audio signals;

[0164] For the indicated playback device, receive information indicating one or more additional associated playback devices;

[0165] Receive one or more audio signals associated with the indicated one or more additional associated playback devices;

[0166] Encode the one or more audio signals associated with the playback device;

[0167] Encode the one or more audio signals associated with the indicated one or more additional associated playback devices;

[0168] Combine the one or more encoded audio signals associated with the playback device and signaling information indicating the one or more additional associated playback devices into a first independent block;

[0169] Combine the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and

[0170] Combine the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.

[0171] EEE-C2. The method according to EEE-C1, wherein the plurality of audio signals includes one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices, and the method further includes:

[0172] Encode each of the one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices into a corresponding independent block; and

[0173] Combine the corresponding independent blocks of each of the one or more groups into the frame of the encoded bitstream.

[0174] EEE-C3. The method according to EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended to be used as an echo reference for performing echo management on the playback device.

[0175] EEE-C4. The method according to EEE-C3, wherein the one or more audio signals intended to be used as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.

[0176] EEE-C5. The method according to EEE-C3 or EEE-C4, wherein the one or more audio signals intended to be used as an echo reference are encoded using parametric coding tools.

[0177] EEE-C6. The method according to EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

[0178] EEE-C7. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame includes two or more independent encoded data blocks, and wherein the playback device includes one or more microphones, the method comprising:

[0179] Identifying, from the encoded bitstream, independent encoded data blocks corresponding to the one or more audio signals associated with the playback device;

[0180] Extracting the identified independent encoded data blocks from the encoded bitstream;

[0181] Extracting the one or more audio signals associated with the playback device from the identified independent encoded data blocks;

[0182] Identifying, from the encoded bitstream, one or more other independent encoded data blocks corresponding to one or more audio signals associated with one or more other playback devices;

[0183] Extracting the one or more audio signals associated with the one or more other playback devices from the one or more other independent encoded data blocks;

[0184] Capturing, using the one or more microphones of the playback device, one or more audio signals; and

[0185] Using the extracted one or more audio signals associated with the one or more other playback devices as an echo reference for performing echo management on the playback device in response to the captured one or more audio signals.

[0186] EEE-C8. The method according to EEE-C7, the method further comprising:

[0187] Determining that the encoded bitstream includes one or more additional independent encoded data blocks; and

[0188] Ignoring the one or more additional independent encoded data blocks.

[0189] EEE-C9. The method according to EEE-C8, wherein ignoring the one or more additional independent encoded data blocks includes skipping the one or more additional independent encoded data blocks without extracting the one or more additional independent encoded data blocks.

[0190] EEE-C10. The method according to any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are specifically intended to be used as an echo reference for performing echo management on the playback device.

[0191] EEE-C11. The method according to EEE-C10, wherein the one or more audio signals specifically intended to be used as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device.

[0192] EEE-C12. The method according to EEE-C10 or EEE-C11, wherein the one or more audio signals specifically intended to be used as an echo reference are reconstructed based on a parametric representation of the one or more audio signals.

[0193] EEE-C13. The method according to any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

[0194] EEE-C14. The method according to EEE-C7, wherein the encoded signal includes signaling information indicating that the one or more other playback devices are used as an echo reference for the playback device.

[0195] EEE-C15. The method according to EEE-C14, wherein the one or more other playback devices indicated by the signaling information for the current frame are different from the one or more other playback devices used as an echo reference for the previous frame.

[0196] EEE-C16. An apparatus configured to perform the method according to any one of EEE-C1 to EEE-C15.

[0197] EEE-C17. A non-transitory computer-readable storage medium including a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of EEE-C1 to EEE-C15.

[0198] EEE-D1. A method for transmitting an audio signal, the method comprising:

[0199] Generating a data packet including a portion of a bitstream, wherein the bitstream includes a plurality of frames, wherein each frame of the plurality of frames includes a plurality of blocks, wherein the generating includes:

[0200] Integrating the data packet with one or more of the plurality of blocks, wherein blocks from different frames are combined into a single packet and / or transmitted out of order; and

[0201] Transmitting the data packet via a packet-based network.

[0202] EEE-D2. The method according to EEE-D1, wherein each of the plurality of blocks includes identification information.

[0203] EEE-D3. The method according to EEE-D2, wherein the identification information includes at least one of a block ID, a corresponding frame number associated with the block, and / or a retransmission priority.

[0204] EEE-D4. The method according to any one of EEE-D1 to EEE-D3, wherein each of the plurality of frames carries all audio data representing a continuous segment of an audio signal having a start time, an end time, and a duration.

[0205] EEE-D5. An audio signal decoding method, the method comprising:

[0206] Receiving a data packet including a portion of a bitstream, wherein the bitstream includes a plurality of frames, and wherein each of the plurality of frames includes a plurality of blocks;

[0207] Determining a set of blocks addressed to a device among the plurality of blocks; and

[0208] Decoding the set of blocks addressed to the device, and skipping decoding of blocks in the plurality of blocks not addressed to the device.

[0209] EEE-D6. A method for transmitting an audio stream, the method comprising:

[0210] Transmitting the audio stream, wherein the audio stream includes a plurality of frames, and wherein each of the plurality of frames includes a plurality of blocks, and wherein the transmitting includes transmitting configuration information of the audio stream out-of-band.

[0211] EEE-D7. The method according to EEE-D6, wherein transmitting the configuration information of the audio stream out-of-band includes:

[0212] Transmitting the audio stream via a first network and / or a first network protocol; and

[0213] Transmitting the configuration information via a second network and / or a second network protocol.

[0214] EEE-D8. The method according to EEE-D7, wherein the first network protocol is the User Datagram Protocol (UDP), and the second network protocol is the Transmission Control Protocol (TCP).

[0215] EEE-D9. An audio signal decoding method, the method comprising:

[0216] Receive a bitstream, the bitstream including: information corresponding to signaling of static configuration aspects; static metadata; and

[0217] Map one or more channel elements to one or more devices based on the information and / or the static metadata.

[0218] EEE-D10. The method according to EEE-D9, wherein the bitstream is received by a plurality of decoders configured to decode the bitstream, and each decoder of the plurality of decoders is configured to decode a part of the bitstream.

[0219] EEE-D11. The method according to EEE-D9 or EEE-D10, wherein the bitstream further includes dynamic metadata.

[0220] EEE-D12. The method according to any one of EEE-D9 to EEE-D11, wherein the bitstream includes a plurality of blocks, and each block of the plurality of blocks includes: information enabling a part of the block to be skipped during decoding, wherein the part is not required by the device; and dynamic metadata.

[0221] EEE-D13. A method for retransmitting blocks of an audio signal, the method including:

[0222] Transmit one or more blocks of a bitstream, wherein the bitstream includes a plurality of blocks, and each of the one or more blocks of the bitstream has been transmitted previously; and wherein each of the one or more blocks includes a decoding priority indicator.

[0223] EEE-D14. The method according to EEE-D13, wherein the decoding priority indicator indicates to a decoder a priority order for decoding the one or more blocks of the bitstream.

[0224] EEE-D15. The method according to EEE-D13 or EEE-D14, wherein each block of the one or more blocks includes the same block ID.

[0225] EEE-D16. The method according to any one of EEE-D13 to EEE-D15, wherein the transmission of the one or more blocks of the bitstream is transmitted at a reduced data rate compared to a previous transmission.

[0226] EEE-D17. The method according to EEE-D16, wherein reducing the data rate includes reducing at least one of a signal-to-noise ratio of an audio signal, a bandwidth of the audio signal, and a channel count of the audio signal.

[0227] EEE-E1. A method for generating a frame of an encoded bitstream for an audio program comprising a plurality of audio signals, wherein the frame comprises one or more independent encoded data blocks, the method comprising:

[0228] For each of the plurality of audio signals, receiving information indicative of a playback device associated with the respective audio signal;

[0229] Encoding one or more audio signals associated with the respective playback device to obtain one or more encoded audio signals;

[0230] Combining the one or more encoded audio signals associated with the respective playback device into a first independent block of the frame;

[0231] Encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; and

[0232] Combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.

[0233] EEE-E2. The method according to EEE-E1, wherein two or more audio signals are associated with the playback device, and each of the two or more audio signals is a band-limited signal intended to be played back by a respective driver of the playback device, and wherein a different encoding technique is used for each band-limited signal.

[0234] EEE-E3. The method according to EEE-E2, wherein a different psychoacoustic model and / or a different bit allocation technique is used for each band-limited signal.

[0235] EEE-E4. The method according to any one of EEE-E1 to EEE-E3, wherein an instantaneous frame rate of the encoded signal is variable and is constrained by a buffer fullness model.

[0236] EEE-E5. The method according to EEE-E1, wherein encoding one or more audio signals associated with the respective playback device comprises jointly encoding one or more audio signals associated with the respective playback device and one or more additional audio signals associated with one or more additional playback devices into a first independent block of the frame.

[0237] EEE-E6. The method according to EEE-E5, wherein jointly encoding the one or more audio signals and one or more additional audio signals comprises sharing one or more scale factors across two or more audio signals.

[0238] EEE-E7. The method according to EEE-E6, wherein the two or more audio signals are spatially correlated.

[0239] EEE-E8. The method according to EEE-E7, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left top channel, a right horizontal channel, or a right top channel.

[0240] EEE-E9. The method according to EEE-E5, wherein jointly encoding the one or more audio signals and one or more additional audio signals includes applying a coupling tool, including:

[0241] Combining two or more audio signals into a composite signal above a specified frequency; and

[0242] For each of the two or more audio signals, determining a scaling factor that relates the energy of the composite signal to the energy of each respective signal.

[0243] EEE-E10. The method according to EEE-E5, wherein jointly encoding the one or more audio signals and one or more additional audio signals includes applying a joint encoding tool to more than two signals.

[0244] EEE-E11. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame includes one or more independent encoded data blocks, the method comprising:

[0245] Identifying from the encoded bitstream the independent encoded data blocks corresponding to the one or more audio signals associated with the playback device;

[0246] Extracting the identified independent encoded data blocks from the encoded bitstream;

[0247] Decoding the one or more audio signals associated with the playback device from the independent encoded data blocks to obtain one or more decoded audio signals;

[0248] Identifying from the encoded bitstream one or more additional independent encoded data blocks corresponding to one or more additional audio signals; and

[0249] Decoding or skipping the one or more additional independent encoded data blocks.

[0250] EEE-E12. The method according to EEE-E11, wherein two or more audio signals are associated with the playback device, and each of the two or more audio signals is a band-limited signal intended to be played back by a corresponding driver of the playback device, and wherein different decoding techniques are used to decode the two or more audio signals.

[0251] EEE-E13. The method according to EEE-E12, wherein different psychoacoustic models and / or different bit allocation techniques are used for encoding each band-limited signal.

[0252] EEE-E14. The method according to EEE-E11 or EEE-E13, wherein the instantaneous frame rate of the encoded bitstream is variable and is constrained by a buffer fullness model.

[0253] EEE-E15. The method according to EEE-E11, wherein decoding one or more audio signals associated with a playback device includes jointly decoding one or more audio signals associated with a corresponding playback device and one or more additional audio signals associated with one or more additional playback devices from independent encoded data blocks.

[0254] EEE-E16. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and one or more additional audio signals includes extracting one or more scale factors shared across two or more audio signals.

[0255] EEE-E17. The method according to EEE-E16, wherein the two or more audio signals are spatially correlated.

[0256] EEE-E18. The method according to EEE-E17, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left top channel, a right horizontal channel, or a right top channel.

[0257] EEE-E19. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and one or more additional audio signals includes applying a decoupling tool.

[0258] EEE-E20. The method according to EEE-E19, wherein the decoupling tool includes:

[0259] extracting an independently decoded signal below a specified frequency;

[0260] extracting a composite signal above the specified frequency;

[0261] Determine each decoupled signal above the specified frequency from the composite signal and a scaling factor that correlates the energy of the composite signal with the energy of each signal; and

[0262] Combine each independently decoded signal with the corresponding decoupled signal to obtain a jointly decoded signal.

[0263] EEE-E21. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and one or more additional audio signals includes applying a joint decoding tool to extract more than two audio signals.

[0264] EEE-E22. The method according to EEE-E11, wherein decoding the one or more audio signals associated with the playback device includes applying bandwidth extension to the audio signals in the same domain in which the audio signals are encoded.

[0265] EEE-E23. The method according to EEE-E22, wherein the domain is the modified discrete cosine transform (MDCT) domain.

[0266] EEE-E24. The method according to EEE-E22 or EEE-E23, wherein bandwidth extension includes adaptive noise addition.

[0267] EEE-E25. An apparatus configured to perform the method according to any one of EEE-E1 to EEE-E24.

[0268] EEE-E26. A non-transitory computer-readable storage medium including a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of EEE-E1 to EEE-E24.

[0269] EEE-F1. A method for generating an encoded bitstream performed by a device having one or more microphones, the method comprising:

[0270] Capture one or more audio signals by the one or more microphones;

[0271] Analyze the captured audio signals to determine the presence of a wake word;

[0272] When the presence of the wake word is detected:

[0273] Set a flag to indicate that a speech recognition task will be performed on the captured audio signals;

[0274] Encode the captured audio signals;

[0275] Integrate the encoded audio signals and the flag into the encoded bitstream.

[0276] EEE-F2. The method according to EEE-F1, wherein the one or more microphones are configured to capture a mono or spatial sound field.

[0277] EEE-F3. The method according to EEE-F2, wherein the spatial sound field is in A format or B format.

[0278] EEE-F4. The method according to any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended for performing only the speech recognition task.

[0279] EEE-F5. The method according to EEE-F4, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for performing the speech recognition task but not sufficient for human listening.

[0280] EEE-F6. The method according to EEE-F4 or EEE-F5, wherein before encoding the captured audio signal, the captured audio signal is converted into a representation including one or more of band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0281] EEE-F7. The method according to any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended for human listening and for performing the speech recognition task.

[0282] EEE-F8. The method according to EEE-F7, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening.

[0283] EEE-F9. The method according to EEE-F7, wherein encoding the captured audio signal includes generating a first encoded representation of the captured audio signal and a second encoded representation of the captured audio signal, wherein the first encoded representation is generated such that when the captured audio signal is decoded from the first encoded representation, the quality of the decoded audio signal is sufficient for human listening, and wherein the second encoded representation is generated such that when the captured audio signal is decoded from the second encoded representation, the quality of the decoded audio signal is sufficient for performing the speech recognition task but not sufficient for human listening.

[0284] EEE-F10. The method according to EEE-F9, wherein generating the second encoded representation of the captured audio signal includes: before encoding the captured audio signal, converting the captured audio signal into a parametric representation, a rough waveform representation, or one or more of a representation including one or more of band energy, Mel frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0285] EEE-F11. The method according to EEE-F9 or EEE-F10, wherein integrating the encoded audio signal into a bitstream includes inserting the first encoded representation into a first independent block of the encoded bitstream and inserting the second encoded representation into a second independent block of the encoded bitstream.

[0286] EEE-F12. The method according to EEE-F9 or EEE-F10, wherein the first encoded representation is included in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first layer and the second layer are included in a single block of the encoded bitstream.

[0287] EEE-F13. The method according to any one of EEE-F1 to EEE-F12, further comprising: when the wake word is not detected:

[0288] Setting the flag to indicate that the speech recognition task is not to be performed on the captured audio signal;

[0289] Encoding the captured audio signal;

[0290] Integrating the encoded audio signal and the flag into the encoded bitstream.

[0291] EEE-F14. An audio signal decoding method, comprising:

[0292] Receiving an encoded bitstream, the encoded bitstream including an encoded audio signal and a flag indicating whether the speech recognition task is to be performed;

[0293] Decoding the encoded audio signal to obtain a decoded audio signal; and

[0294] When the flag indicates that the speech recognition task is to be performed, performing the speech recognition task on the decoded audio signal.

[0295] EEE-F15. The method according to EEE-F14, wherein the decoded audio signal is intended to be used only for performing the speech recognition task.

[0296] EEE-F16. The method according to EEE-F15, wherein the quality of the decoded audio signal is sufficient for performing the speech recognition task but not sufficient for human listening.

[0297] EEE-F17. The method according to EEE-F15 or EEE-F16, wherein, before encoding the captured audio signal, the representation of the decoded audio signal includes one or more of band energy, Mel frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0298] EEE-F18. The method according to EEE-F14, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening.

[0299] EEE-F19. The method according to EEE-F18, wherein the encoded audio signal includes a first encoded representation of one or more audio signals and a second encoded representation of the one or more audio signals.

[0300] EEE-F20. The method according to EEE-F18, wherein the quality of the audio signal decoded from the first representation is sufficient for human listening, and wherein the quality of the audio signal decoded from the second representation is sufficient for performing a speech recognition task but not sufficient for human listening.

[0301] EEE-F21. The method according to EEE-F19 or EEE-F20, wherein the first representation is in a first independent block of the encoded bitstream, and the second representation is in a second independent block of the encoded bitstream.

[0302] EEE-F22. The method according to EEE-F19 or EEE-F20, wherein the first representation is in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first layer and the second layer are included in a single block of the encoded bitstream.

[0303] EEE-F23. The method according to any one of EEE-F18 to EEE-F22, wherein decoding the encoded audio signal includes decoding only the second representation while ignoring the first representation.

[0304] EEE-F24. The method according to any one of EEE-F18 to EEE-F23, wherein the audio signal decoded from the second encoded representation is a parametric representation, a waveform representation, or a representation including one or more of band energy, Mel frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.

[0305] EEE-F25. An apparatus configured to perform the method according to any one of EEE-F1 to EEE-F24.

[0306] EEE-F26. A non-transitory computer-readable storage medium includes a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of EEE-F1 to EEE-F24.

[0307] EEE-G1. An encoding method for an audio signal of an immersive audio program for low-latency transmission to one or more playback devices, the method comprising:

[0308] Receiving a plurality of time-domain audio signals of the immersive audio program;

[0309] Selecting a frame size;

[0310] Extracting frames of the time-domain audio signals in response to the frame size, wherein the frames of the time-domain audio signals overlap with the previous frame of the time-domain audio signals;

[0311] Dividing the audio signal into overlapping frames;

[0312] Transforming the frames of the time-domain audio signals into frequency-domain signals;

[0313] Encoding the frequency-domain signals;

[0314] Quantizing the encoded frequency-domain signals using a perceptual excitation quantization tool;

[0315] Integrating the quantized and encoded frequency-domain signals into one or more independent blocks within a frame; and

[0316] Integrating the one or more independent blocks into an encoded frame.

[0317] EEE-G2. The method according to EEE-G1, wherein the plurality of audio signals includes channel-based signals having a defined channel configuration.

[0318] EEE-G3. The method according to EEE-G2, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, or 22.2.

[0319] EEE-G4. The method according to any one of EEE-G1 to EEE-G3, wherein the plurality of audio signals includes one or more object-based signals.

[0320] EEE-G5. The method according to any one of EEE-G1 to EEE-G4, wherein the plurality of audio signals includes a scene-based representation of the immersive audio program.

[0321] EEE-G6. The method according to any one of EEE-G1 to EEE-G5, wherein the selected frame size is one of 128, 256, 512, 1024, 120, 240, 480, or 960 samples.

[0322] EEE-G7. The method according to any one of EEE-G1 to EEE-G6, wherein the overlap between the frame of the time-domain audio signal and the previous frame of the time-domain audio signal is 50% or less.

[0323] EEE-G8. The method according to any one of EEE-G1 to EEE-G7, wherein the transform is a modified discrete cosine transform (MDCT).

[0324] EEE-G9. The method according to any one of EEE-G1 to EEE-G8, wherein two or more of the plurality of audio signals are jointly encoded.

[0325] EEE-G10. The method according to any one of EEE-G1 to EEE-G9, wherein each independent block contains an encoded signal for one or more playback devices.

[0326] EEE-G11. The method according to any one of EEE-G1 to EEE-G10, wherein at least one independent block contains an encoded signal for two or more playback devices, and the encoded signal includes a jointly encoded audio signal.

[0327] EEE-G12. The method according to any one of EEE-G1 to EEE-G11, wherein at least one independent block contains a plurality of encoded signals covering different bandwidths and intended to be played back from different drivers of a playback device.

[0328] EEE-G13. The method according to any one of EEE-G1 to EEE-G12, wherein at least one independent block contains an encoded echo reference signal for echo management performed by a playback device.

[0329] EEE-G14. The method according to any one of EEE-G1 to EEE-G13, wherein encoding the quantized frequency-domain signal includes applying one or more of the following tools: time noise shaping (TNS), joint channel coding, cross-signal sharing of scale factors, determining control parameters for high-frequency reconstruction, and determining control parameters for noise substitution.

[0330] EEE-G15. The method according to any one of EEE-G1 to EEE-G14, wherein one or more independent blocks include parameters for controlling one or more of the delay, gain, and equalization of a playback device.

[0331] EEE-G16. A low-latency method for an audio signal to decode an immersive audio program from an encoded signal, the method comprising:

[0332] Receiving an encoded frame comprising one or more independent blocks;

[0333] Extracting a quantized and encoded frequency-domain signal from the one or more independent blocks;

[0334] Dequantizing the quantized and encoded frequency-domain signal;

[0335] Decoding the dequantized frequency-domain signal;

[0336] Performing an inverse transform on the decoded frequency-domain signal to obtain a time-domain signal; and

[0337] Overlapping and adding the time-domain signal with the time-domain signal from a previous frame to provide a plurality of audio signals of the immersive audio program.

[0338] EEE-G17. The method according to EEE-G16, wherein the plurality of audio signals includes channel-based signals having a defined channel configuration.

[0339] EEE-G18. The method according to EEE-G17, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6 or 22.2.

[0340] EEE-G19. The method according to any one of EEE-G16 to EEE-G18, wherein the plurality of audio signals includes one or more object-based signals.

[0341] EEE-G20. The method according to any one of EEE-G16 to EEE-G19, wherein the plurality of audio signals includes a scene-based representation of the immersive audio program.

[0342] EEE-G21. The method according to any one of EEE-G16 to EEE-G20, wherein a frame of time-domain samples includes one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.

[0343] EEE-G22. The method according to any one of EEE-G16 to EEE-G21, wherein the overlap with the previous frame is 50% or less.

[0344] EEE-G23. The method according to any one of EEE-G16 to EEE-G22, wherein the inverse transform is an inverse modified discrete cosine transform (IMDCT).

[0345] EEE-G24. The method according to any one of EEE-G16 to EEE-G23, wherein each independent block contains a frequency-domain signal that is quantized and encoded for one or more playback devices.

[0346] EEE-G25. The method according to any one of EEE-G16 to EEE-G24, wherein at least one independent block contains a frequency-domain signal that is quantized and encoded for two or more playback devices, and the quantized and encoded frequency-domain signal is a jointly encoded audio signal.

[0347] EEE-G26. The method according to any one of EEE-G16 to EEE-G25, wherein at least one independent block contains a plurality of frequency-domain signals that are quantized and encoded and cover different bandwidths and are intended to be played back from different drivers of a playback device.

[0348] EEE-G27. The method according to any one of EEE-G16 to EEE-G26, wherein at least one independent block contains an encoded echo reference signal for echo management performed by a playback device.

[0349] EEE-G28. The method according to any one of EEE-G16 to EEE-G27, wherein decoding the dequantized frequency-domain signal includes applying one or more of the following decoding tools: temporal noise shaping (TNS), joint channel decoding, cross-signal shared scale factors, high-frequency reconstruction, and noise substitution.

[0350] EEE-G29. The method according to any one of EEE-G16 to EEE-G28, wherein one or more independent blocks include parameters for controlling one or more of the delay, gain, and equalization of a playback device.

[0351] EEE-G30. The method according to any one of EEE-G16 to EEE-G29, wherein the method is performed by a playback device, and wherein extracting the quantized and encoded signal from one or more independent blocks includes only selecting those blocks that contain the quantized and encoded frequency-domain signal for playback by the playback device and ignoring the independent blocks that contain the quantized and encoded frequency-domain signal for playback by other playback devices.

[0352] EEE-G31. An apparatus configured to perform the method according to any one of EEE-G1 to EEE-G30.

[0353] EEE-G32. A non-transitory computer-readable storage medium including a sequence of instructions that, when executed, cause one or more devices to perform the method according to any one of EEE-G1 to EEE-G30.

[0354] In the following claims and the description herein, any of the terms "comprising", "consisting of" or "which includes" is an open term, which means at least including the elements / features that follow, but not excluding other elements / features. Therefore, when used in the claims, the term "comprising" should not be interpreted as being limited to the devices or elements or steps listed thereafter. For example, the scope of the expression "the device comprises A and B" should not be limited to the device consisting of only elements A and B. As used herein, any of the terms "comprising" or "which includes" is also an open term, which also means at least including the elements / features that follow the term, but not excluding other elements / features. Therefore, including is synonymous with including and means including.

[0355] It should be appreciated that in the above description of examples of the present invention, various features are sometimes grouped together in a single example, figure, or description thereof for the purpose of streamlining the present disclosure and aiding in the understanding of one or more of the various inventive aspects. However, the disclosed method should not be interpreted as reflecting the intention of requiring more features than those explicitly recited in each claim. On the contrary, as reflected in the appended claims, the inventive aspects lie in less than all the features of a single aforementioned disclosed example. Therefore, the claims following the specific embodiments are hereby expressly incorporated into the present specific embodiments, with each claim independently serving as a separate example of the present invention.

[0356] In addition, although some examples described herein include some features included in other examples, but do not include other features, as will be understood by those skilled in the art, combinations of features of different examples are intended to be covered and form different examples. For example, in the following claims, any of the claimed examples may be used in any combination.

[0357] In addition, some of the examples are described herein as methods or combinations of elements of methods that can be implemented by a processor of a computer system or by other means of performing functions. Therefore, a processor with the necessary instructions for performing such a method or elements of a method forms a means for performing such a method or elements of a method. In addition, the elements described herein of a device are examples of means for carrying out the functions performed by the element.

[0358] In addition, some examples described herein disclose connection solutions that should be interpreted as being possible to implement in distribution and / or transmission systems such as wired and / or wireless systems, for example by using any electrical, optical and / or mobile systems, such as 3G, 4G and 5G.

[0359] Accordingly, while specific examples of the invention have been described, those skilled in the art will recognize that other and further modifications can be made and are intended to claim all such changes and modifications. For example, any of the formulas given above are merely representative of the processes that can be used. Functions can be added or removed from the block diagrams, and operations can be interchanged between functional blocks. Steps can be added or removed from the described methods.

[0360] The systems, devices, and methods disclosed above can be implemented as software, firmware, hardware, or a combination thereof. For example, aspects of the present application can be embodied, at least in part, in a device, a system including more than one device, a method, a computer program product, etc.

[0361] In a hardware implementation, the partitioning of tasks between the functional units mentioned in the above description does not necessarily correspond to the partitioning of physical units; rather, one physical component can have multiple functions, and one task can be performed by several physical components working together.

[0362] Certain components or all components can be implemented as software executed by a digital signal processor or a microprocessor or as hardware or an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium).

[0363] As is well known to those skilled in the art, the term "computer storage medium" includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer.

[0364] Furthermore, as is well known to those skilled in the art, a communication medium typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism and includes any information delivery medium.

Claims

1. A method for generating a frame of an encoded bitstream for an audio program comprising a plurality of audio signals, wherein the frame comprises two or more independent encoded data blocks, the method comprises: For one or more of the plurality of audio signals, receiving information indicating a playback device associated with the one or more audio signals; For the indicated playback device, receiving information indicating one or more additional associated playback devices; Receiving one or more audio signals associated with the indicated one or more additional associated playback devices; Encoding the one or more audio signals associated with the playback device; Encoding the one or more audio signals associated with the indicated one or more additional associated playback devices; Combining the one or more encoded audio signals associated with the playback device and signaling information indicating the one or more additional associated playback devices into a first independent block; Combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and Combining the first independent block and the one or more additional independent blocks into the frame of the encoded bitstream.

2. The method according to claim 1, wherein, The plurality of audio signals includes one or more sets of audio signals not associated with the playback device or the one or more additional associated playback devices, and the method further comprises: Encoding each set of audio signals in the one or more sets of audio signals not associated with the playback device or the one or more additional associated playback devices into a corresponding independent block; and Combining the corresponding independent blocks of each set in the one or more sets into the frame of the encoded bitstream.

3. The method according to claim 1 or 2, wherein the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended to be used as an echo reference for performing echo management on the playback device.

4. The method according to claim 3, wherein, Compared with the one or more audio signals associated with the playback device, less data is used to transmit the one or more audio signals intended to be used as an echo reference.

5. The method according to claim 3 or 4, wherein the one or more audio signals intended to be used as an echo reference are encoded using parametric coding tools.

6. The method according to claim 1 or 2, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

7. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises two or more independent encoded data blocks, wherein the playback device comprises one or more microphones, the method comprises: Identifying from the encoded bitstream an independent encoded data block corresponding to the one or more audio signals associated with the playback device; Extracting the identified independent encoded data block from the encoded bitstream; Extract the one or more audio signals associated with the playback device from the identified independent encoded data blocks; Identify, from the encoded bitstream, one or more other independent encoded data blocks corresponding to one or more audio signals associated with one or more other playback devices; Extract the one or more audio signals associated with the one or more other playback devices from the one or more other independent encoded data blocks; Capture one or more audio signals using the one or more microphones of the playback device; And Use the extracted one or more audio signals associated with the one or more other playback devices as an echo reference for performing echo management on the playback device in response to the captured one or more audio signals.

8. The method according to claim 7, the method further comprises: Determine that the encoded bitstream includes one or more additional independent encoded data blocks; And Ignore the one or more additional independent encoded data blocks.

9. The method according to claim 8, wherein ignoring the one or more additional independent encoded data blocks comprises skipping the one or more additional independent encoded data blocks without extracting the one or more additional independent encoded data blocks.

10. The method according to any one of claims 7 to 9, wherein, The one or more audio signals associated with the one or more other playback devices are specifically intended to be used as an echo reference for performing echo management on the playback device.

11. The method according to claim 10, wherein, Compared with the one or more audio signals associated with the playback device, less data is used to transmit the one or more audio signals specifically intended to be used as an echo reference.

12. The method according to claim 10 or 11, wherein, Reconstruct the one or more audio signals specifically intended to be used as an echo reference according to the parameter representation of the one or more audio signals.

13. The method according to any one of claims 7 to 9, wherein, The one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.

14. The method according to claim 7, wherein, The encoded signal includes signaling information indicating that the one or more other playback devices are used as an echo reference for the playback device.

15. The method according to claim 14, wherein, For the one or more other playback devices indicated by the signaling information for the current frame, they are different from the one or more other playback devices used as an echo reference for the previous frame.

16. An apparatus configured to perform the method according to any one of claims 1 - 15.

17. A non - transitory computer - readable storage medium comprising a sequence of instructions which, when executed, cause one or more devices to perform the method according to any one of claims 1 - 15.

Citation Information

Patent Citations

  • Selective forward error correction for spatial audio codecs

    US11289103B2