Method, apparatus and medium for decoding audio signal with skippable blocks
By receiving multiple bit streams and determining a subset of decoding blocks based on output device information, the problem of signal delay and complexity in wireless streaming is solved, and low-latency and efficient audio signal exchange and decoding are achieved.
Patent Information
- Application Number
- CN202380071308.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2023-09-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively solve the problem of signal delay and complexity when wirelessly streaming audio signals, especially in multi-device scenarios.
By receiving a bit stream including a plurality of blocks, based on device information of the output device, signaling data for identifying a subset of blocks to be decoded, and decodes the identified subset of blocks.
It realizes low-latency audio signal switching and decoding, which is suitable for multi-device scenarios, and improves the efficiency and flexibility of audio signal processing.
Smart Images

Figure CN119998875A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 378,496, filed on October 5, 2022, and U.S. Provisional Application No. 63 / 578,606, filed on August 24, 2023, both of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates generally to audio signal processing, and more particularly to audio source encoding and decoding for low-latency exchange of audio signals for immersive audio programs between devices. Background Art
[0004] Audio streaming is common in today's society. Audio streaming is becoming more and more demanding as user expectations of quality increase, and user setups are becoming more and more complex, with a large number of speakers and speaker types. Streaming is usually done at least partially over a wireless link, which then requires that the wireless link is of good quality, however, as many people may have experienced, this is not always the case.
[0005] Therefore, there is a need to define an exchange format for use cases where a certain format is streamed from the cloud / server and then transcoded on the device to a more suitable low-latency format for distribution over a wireless (or, in some cases, wired) link. Exemplary use cases are in-home connections, and phone-to-car connections, however the format may be beneficial in any scenario where low-latency distribution of audio signals from a single device to one or more connected devices is desired.
[0006] In addition to audio information being wirelessly sent and streamed, there may be other types of information incorporated into the stream. Such other types of information may also be affected by the quality of the wireless link and may have similar disadvantages as audio.
[0007] Therefore, it would be advantageous to overcome the problems associated with wireless streaming of different types of streaming audio for combination with other types of information or signals. Summary of the invention
[0008] One object of the present disclosure is to overcome at least in part the above-mentioned problems with wireless streaming of audio combined with other types of information.
[0009] According to a first aspect of the present disclosure, a method for decoding an audio signal is provided, the method comprising: receiving a bit stream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks; determining, based on device information of an output device, information for identifying a subset of one or more blocks to be decoded among the plurality of blocks for each frame according to signaling data; and decoding the identified subset of one or more blocks among the plurality of blocks.
[0010] According to a second aspect of the present disclosure, a method for generating an encoded bit stream of an audio signal is provided, the method comprising: generating a bit stream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks; for each frame, determining, based on device information of an output device, signaling data for identifying a subset of one or more blocks to be decoded among the plurality of blocks; and encoding the bit stream using the signaling data, comprising the subset of one or more blocks to be decoded among the plurality of blocks when decoding.
[0011] According to a third aspect of the present disclosure, there is provided an apparatus configured to execute the method of the first and / or second aspect.
[0012] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium comprising an instruction sequence, which, when executed, causes one or more devices to perform the method of the first and / or second aspect.
[0013] Further examples of the disclosure are defined in the dependent claims.
[0014] In some examples, the information for identifying a subset of one or more blocks in the plurality of blocks includes a matrix associating each of a plurality of output devices with a subset of the one or more blocks. In some examples, one or more frames are required to decode the bitstream for the corresponding associated output device. In some examples, the output device may include at least one of a wireless device, a mobile device, a tablet computer, a mono speaker, a multi-channel speaker, and / or a sound reproduction system.
[0015] In some examples, one or more blocks partially and / or completely include at least one audio signal for playback based on device information at an output device. In some examples, the output device is a first output device, and also includes applying a joint coding technique between one or more signals to at least a second output device in the bitstream. In some examples, the identity of each output device and / or decoder is defined during a system initialization phase. In some examples, the signaling data is determined by metadata of the bitstream.
[0016] In this disclosure, a frame refers to a time slice of the whole of all signals. A block stream refers to a collection of signals over the duration of a session. A block refers to one frame of a block stream. For digital audio with a given sampling frequency, the frame size is equivalent to the number of audio samples in a frame of any audio signal. The frame size is usually kept constant over the duration of a session.
[0017] The wake-up word may include a single word, or a phrase containing two or more words in a fixed order.
[0018] Throughout the present disclosure including in the claims, the expression "system" is used in a broad sense to refer to a device, system or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and other XM inputs are received from external sources) may also be referred to as a decoder system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Examples of the present disclosure will be described in more detail with reference to the accompanying drawings. In the following drawings, the same reference numerals are used to refer to the same elements. Although the following drawings depict various examples, one or more implementations are not limited to the examples depicted in the drawings.
[0020] Figure 1 An example of low-latency transcoding for an in-home connection is shown.
[0021] Figure 2 An example of car-connected audio streaming is shown.
[0022] Figure 3 An example of enhanced TV audio streaming using wireless devices is shown.
[0023] Figure 4 An example of audio streaming with a simple wireless speaker is shown.
[0024] Figure 5 An example of metadata mapping from bitstream elements to devices is shown.
[0025] Figure 6 An example showing how to deploy a simple and flexible rendering of audio streaming.
[0026] Figure 7 An example of mapping of flexible rendering data to bitstream elements is shown.
[0027] Figure 8 An example of echo-referenced signaling when playing audio and listening for commands with multiple devices is shown.
[0028] Fig. 9An example of how frames, blocks, and packets relate to each other is shown.
[0029] Fig.10 An example of additional codec information in the current MPEG-4 structure is shown.
[0030] Fig.11 An example of still further codec information in the current MPEG-4 structure is shown.
[0031] Fig.12 An example of integrated listening capabilities and speech recognition is shown.
[0032] Fig.13 An example of a frame including a plurality of blocks is shown.
[0033] Fig.14 An example of a bitstream including a plurality of blocks is shown.
[0034] Fig.15 An example of blocks with different priorities is shown.
[0035] Fig.16 An example of a bitstream including multiple blocks with different priorities is shown.
[0036] Fig.17 An example of frames with different priorities is shown.
[0037] Fig.18 An example of a bitstream including frames with different priorities is shown. DETAILED DESCRIPTION
[0038] The principle of the present invention will now be described with reference to various examples shown in the accompanying drawings. It should be appreciated that the description of these embodiments is only to enable those skilled in the art to better understand and further implement the present invention, and is not intended to limit the scope of the present invention in any way.
[0039] exist Figure 1 In the immersive audio stream, the immersive audio stream is streamed from the cloud or server 10 and decoded on the TV or hub device 20. The immersive audio stream can be encoded in any existing format, including, for example, Dolby Digital Plus, AC-4, etc. The output is then transcoded into a low-latency exchange format for further transmission to a connected device 30, which is preferably connected via a local wireless connection (e.g., a WiFi soft access point or a Bluetooth connection). Low latency typically depends on various factors, such as frame size, sampling rate, hardware and / or software computing resources, etc., but low latency is typically less than 40ms, 20ms, or 10ms.
[0040] In another typical use case, the phone obtains the immersive audio stream from the cloud or server, transcodes it into an exchange format, and then transmits it to the connected car. Figure 2 In the example shown, a mobile device (e.g., a phone or tablet) 20 connects to the server 10 and receives an immersive audio stream, performs low-latency transcoding to a low-latency exchange format, and transmits the transcoded signal to an immersive audio playback-enabled car 30. An example of an immersive audio stream is a stream that includes audio in the Dolby Atmos format, and an example of an immersive audio playback-enabled car 30 is a car configured to play back the Dolby Atmos immersive format.
[0041] In general, the exchange format preferably has low latency, low encoding and decoding complexity, the ability to scale to high quality, and reasonable coding efficiency. The format preferably also supports configurable delays so that delays can be traded off with efficiency and error tolerance to operate under varying connection conditions.
[0042] A hub or enhanced display with wireless speakers with or without listening capabilities
[0043] Figure 3 , a hub 20 driving a set of wireless speakers 30, or a display (e.g., a television or TV) 20 that may have built-in speakers 30, which may be augmented with several wireless speakers 30. The augmentation indicates that in the example where the display 20 includes speakers 30, the display 20 is also part of the audio reproduction. The wireless speakers / devices 30 can receive the same complete signal (broadcast mode, Figure 3 ) or a separate stream tailored for a specific device (Unicast-Multipoint mode, Figure 3 ).
[0044] Each speaker may include multiple drivers, covering different frequency ranges of the same channel, or corresponding to different channels in the canonical mapping. For example, a speaker may have two drivers, one of which may be an upward-firing driver that outputs a signal corresponding to a height channel to emulate an overhead speaker. Wireless devices may have listening capabilities (e.g., "smart speakers"), and therefore may require echo management, which may require the speaker to receive one or more echo references. The echo reference may be local (e.g., to the same speaker / device), or alternatively may represent relevant signals from other nearby speakers / devices.
[0045] The speakers may be placed at arbitrary locations, in which case so-called "flexible rendering" may be performed, whereby the rendering takes into account the actual locations of the speakers (e.g., as opposed to the specification assumptions, whereby the speakers are assumed to be at fixed, predefined locations). Flexible rendering may occur in the hub or TV, with the rendered signals then being transmitted to the speakers / devices either in broadcast mode, whereby each rendered signal is transmitted to each device and each respective device extracts and outputs the appropriate signal, or as separate streams to separate devices. Alternatively, flexible rendering may occur locally on each device, whereby each device receives a representation of the complete immersive program, e.g., a 7.1.4 channel based immersive representation, and renders output signals appropriate for the respective device based on that representation.
[0046] Wireless devices can connect to a hub or TV through a soft access point provided by the device or through a local access point in the home. This may impose different requirements on bit rate and latency.
[0047] Mobile Device Projection in Automotive Applications
[0048] In another use case, a mobile device (e.g., a phone or tablet) 20 obtains an immersive audio stream from a cloud or server 10, transcodes the immersive audio stream into an exchange format, and then transmits the transcoded stream to a connected car 30. Figure 2 In the example of , a Dolby Atmos enabled phone 20 connects to a server 10 and receives a Dolby Atmos stream for low latency transcoding and transmission to a Dolby Atmos enabled car 3. In this use case, higher bit rates can be utilized than in the living room use case, and there may also be different characteristics of the wireless channel. The higher bit rate in the car use case may be attributed to a less noisy wireless environment, because the car acts as a kind of Faraday cage and shields the interior environment from external wireless interference compared to the living room, where all wireless devices, such as from neighbors, other rooms, etc., add noise to the wireless environment. At the same time, there are typically very few wireless devices in the car to compete for the available wireless bandwidth, compared to the living room use case where all wireless devices in the home (including the living room wireless devices) typically compete for the available bandwidth.
[0049] For this use case, it is envisioned that the signal to be transcoded by the mobile device and transmitted to the car may be a channel-based immersive representation, an object-based representation, a scene-based representation (e.g., an Ambisonics representation), or even a combination of different representations. For this example, different rendering architectures and broadcast vs. multipoint may not be relevant because the complete presentation will typically be delivered from the mobile device to a single endpoint (e.g., the car).
[0050] Interchange Format Description
[0051] The Immersive Interchange Format is built on a Modified Discrete Cosine Transform (MDCT) with perceptually excited quantization and coding. It has configurable delays, for example to support different transform sizes at a given sampling rate. Exemplary frame sizes are 128, 256, 512, 1024 and 120, 240, 480, 960 and 192, 384, 768 samples at sampling rates of 48kHz and 44.1kHz.
[0052] The format may support mono, stereo, 5.1, and other channel configurations, including immersive channel configurations (e.g., including but not limited to 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, and 22.2, or any other channel configuration consisting of unique channels as specified in ISO / IEC 23091-3:2018, Table-2). The format may also support object-based audio and scene-based representations such as Ambisonics (e.g., first order or higher order). The format may also use a signaling scheme suitable for integration with existing formats (e.g., such as those described in ISO / IEC 14496-3, which may also be referred to as the MPEG-4 Audio standard, ISO 14496-1, which may also be referred to as the MPEG-4 Systems standard, ISO 14496-12, which may also be referred to as the ISO Base Media File Format standard, and / or ISO 14496-14, which may also be referred to as the MP4 file format standard).
[0053] In addition, the system can support the ability to skip a portion of syntax elements and only decode the relevant portion for a given speaker, support flexible rendering aspects such as metadata control of delay alignment, level adjustment and equalization, support the use of smart speakers with listening capabilities (including scenarios where a set of signals is transmitted, allowing the set of signals to independently feed each driver / speaker in the smart speaker), and associated support for echo management through signal delivery of echo references, support quantization and encoding in the MDCT domain with both low overlap windows and 50% overlap windows, and / or support filtering in the MDCT domain along the frequency axis for time shaping of quantization noise, such as TNS (temporal noise shaping). Therefore, in various examples, the overlapping windows are symmetric or asymmetric.
[0054] In some examples, the disclosed exchange format can provide improved joint channel coding, taking advantage of increased correlation between signals in flexible rendering use cases, improved coding efficiency through scale factors shared across channel elements, inclusion of high frequency reconstruction and noise addition techniques in the MDCT domain, various improvements over traditional coding structures to allow better efficiency and applicability to the use case at hand.
[0055] In some examples of controlling individual playback devices in broadcast mode, skippable blocks and metadata may be used. In a broadcast mode setting, each wireless device will need to play the relevant portion of the complete audio for a specific driver / channel corresponding to a specific device. This means that the device needs to know which parts of the stream are relevant to that device and extract those parts from the complete audio stream being broadcast. In order to achieve a low-complexity decoding operation, it is preferred that the stream is constructed so that the decoder of a specific device can effectively skip over elements that are not relevant to decoding and jump to elements that are relevant to that device and a given driver on that device.
[0056] However, it should be noted that it may still be beneficial (from a compression point of view) to allow joint encoding between signals destined for different devices (eg, the left and right channels of a stereo presentation in the simplest case).
[0057] Thus, the format may enable "skippable blocks" in the bitstream to enable efficient decoding of portions that are only relevant to a specific device, including metadata that enables flexible mapping of one or more skippable blocks to a specific device while retaining the ability to apply joint coding techniques between signals corresponding to different speakers / devices.
[0058] exist Figure 4 In FIG. 1 , an example of an arbitrary setup is shown. There are three connected wireless devices 31, 32, 33, two of which are mono speakers 32, 33 (meaning one channel of audio), and one speaker 31 is a more advanced speaker 31 having three different drivers and thus operating on three different signals. In this example, the first speaker 31 operates on 3 separate signals, while the second and third speakers 32, 33 operate on a stereo representation (e.g. left and right). As such, it may be beneficial to jointly encode the signals for the stereo representation.
[0059] Given the above scenario, the format specifies that the bitstream contains skippable blocks so that the first device can extract only the relevant part of the stream for decoding the signal of the speaker, and the decoder complexity versus efficiency can be compromised for a stereo pair by constructing a signal where each speaker needs to decode two signals to output a single signal, with the good news that joint encoding can be performed.
[0060] The format specifies a metadata format that enables a generic and flexible representation of a mapping of a particular one of multiple skippable blocks to one or more devices. Figure 5 , where the mapping can be represented as a matrix associating each device 31, 32, 33 to one or more bitstream elements, so that the decoder for a given device will know what bitstream elements to output and decode.
[0061] For example, in Figure 5 In the example of , the first block or skip block Blk1 includes three single channel elements (one for each driver of device 1 31), while the second block or skip block Blk2 contains channel pair elements, which may include jointly encoded versions of the signals to be output by devices 2 32 and 3 33. Device 1 31 extracts the mapping metadata and determines that the signal it needs is in skip block 1Blk1. Therefore, it extracts skip block 1Blk1, decodes the three single channel elements therein, and provides them to drivers 1, 2a, and 2b, respectively.
[0062] In addition, device 1 31 ignores skip block 2, Blk2. Similarly, device 2 32 extracts mapping metadata and determines that the signal it needs is in skip block 2Blk2. Therefore, device 2 32 skips skip block 1Blk1 and extracts skip block 2Blk2. Device 2 32 decodes the channel pair element and provides the left channel output of the CPE to its driver. Similarly, device 3 33 extracts mapping metadata and determines that the signal it needs is in skip block 2Blk2, so device 3 33 skips skip block 1Blk 1 and also extracts skip block 2Blk2. Device 3 33 decodes the channel pair element and provides the right channel output of the CPE to its driver.
[0063] In some examples, device 2 32 may determine that it only needs a subset of the signal from skip block 2Blk2. In such examples, when possible, device 2 32 may only perform a subset of the operations required to fully decode the signal in skip block 2Blk2. Specifically, in Figure 5 In the example of , device 2 32 can only perform those processing operations required to extract the left channel of the CPE, thereby enabling reduced computational complexity. Similarly, device 3 33 can only perform those processing operations required to extract the right channel of the CPE.
[0064] For example, in Figure 5 In the case of , there may be times when the CPE is encoded using joint channel coding and other times when it is encoded using independent channel coding. During the time when the CPE is encoded using independent channel coding, device 2 32 may extract only the first (e.g., left) channel of the CPE, while device 3 33 may extract only the second (e.g., right) channel of the CPE.
[0065] In another example, joint channel coding may be used to encode the channels of the CPE, in which case devices 2 32 and 3 33 must extract the two middle channels of the CPE. However, device 2 32 is still able to operate with reduced computational complexity by performing only those operations required to extract the left channel from the middle channel. Similarly, device 3 33 is still able to operate with reduced computational complexity by performing only those operations required to extract the right channel from the middle decoded channel.
[0066] Other optimizations may be possible depending on the specifics of the joint channel coding. The identity of each of the different devices / decoders may be defined during a system initialization or setup phase. Such setup is common and typically involves measuring the acoustics of the room, the distance of the speakers to the sweet-spot, etc.
[0067] Flexible rendering (pre-encoding and post-decoding) distribution and delay, etc. Post-decoder applications
[0068] For use cases where flexible rendering is applied in the hub / TV before sending the relevant signal to a specific device / speaker, rendering may create a signal that is more difficult to encode from an encoding perspective, such as a signal that is more difficult to encode when considering joint encoding of such signals. One reason is that flexible rendering can apply different delays, equalization, and / or gain adjustments to different devices (e.g., depending on the placement of the speaker relative to other speakers and the listener). It is also possible to pre-set information such as gain and delay during initial setup and flexibly render only the equalization. In other examples, other variations of pre-set information and flexible rendering information are also possible. It should be noted that in this article, the term "gain" should be interpreted as meaning any level adjustment (e.g., attenuation, amplification, or pass-through), rather than being limited to certain level adjustments (e.g., amplification).
[0069] exist Figure 6 In the example of , the right channel 33 and left channel 32 speakers are given different delays, reflecting the different placement relative to the listener (e.g., since the speakers 32, 33 may not be equidistant from the listener, different delays may be applied to the signals output from the speakers 32, 33 so that coherent sounds from different speakers 32, 33 arrive at the listener at the same time). Introducing such different delays to coherent signals intended for playback on different speakers 31, 32, 33 makes joint encoding of such signals challenging.
[0070] To address this challenge, aspects of a flexible rendering process can be parameterized and applied in the endpoint device after decoding of the signal.
[0071] exist Figure 7In the example of , the delay and gain values of each device are parameterized and included in the encoded signal transmitted to the respective loudspeaker 31, 32, 33. The respective signals may be decoded by the respective device 31, 32, 33, which may then introduce the parameterized gain and delay values into the respective decoded signal.
[0072] In an example, for example Figure 7 In the example shown, the encoded signals for different devices 31, 32, 33 may be transmitted in separable blocks (e.g., skippable blocks), and parameters (e.g., delays and gains) may also be transmitted in separable blocks, so that the devices 31, 32, 33 may extract only a subset of the parameters required by the devices 31, 32, 33, and ignore (and skip) those parameters not required by the devices 31, 32, 33. In this case, mapping metadata indicating which parameters are included in which blocks may be provided to each device.
[0073] It should also be noted that although Figure 7 Only delay and gain parameters are indicated, but other parameters such as equalization parameters may also be included. The equalization parameters may, for example, include a plurality of gains to be applied to different frequency regions, an indication of a predetermined equalization curve to be applied by the playback device 31, 32, 33, one or more sets of infinite impulse response (IIR) or finite impulse response (FIR) filter coefficients, a set of biquad filter coefficients, parameters specifying the characteristics of a parametric equalizer, and other parameters for specifying equalization known to those skilled in the art.
[0074] Furthermore, parameterization of aspects of flexible rendering need not be static, but can be dynamic (e.g., in the case of a listener moving during playback of an audio program). Thus, it may be preferable to allow parameters to change dynamically. In the case of one or more parameters changing during an audio program, the device may interpolate between previous delay and / or gain parameters and updated delay and / or gain parameters to provide a smooth transition. This may be particularly useful in scenarios where the system dynamically tracks the position of the listener and updates the dynamically rendered sweet spot accordingly.
[0075] It should also be noted and will be described subsequently that when flexible rendering is applied, an increased level of correlation between channels may occur and this may be exploited by a more flexible joint coding.
[0076] Echo Reference Coding and Signaling
[0077] For Figure 8 The use case outlined in this use case is that multiple devices / speakers 30 operate together to Figure 8 The left side shows the broadcast mode or Figure 8Receiving signals in the unicast-multipoint mode shown on the right, while having a microphone 40 on the device 30 to enable "listening" capability, the need for echo management arises.
[0078] When performing echo management for multiple speakers / devices, it may be beneficial to use echo references from more than just the local speaker device. As an example, a device may be placed close to another device, so the signal from the nearby device will affect the echo management of the device with an active microphone. In a broadcast mode use case, each device receives the signal of all devices. When a device has the signals of other devices, it may be beneficial to use those signals as echo references. To do this, it is necessary to signal to a specific device which signals can be used as echo references for other devices.
[0079] In one example, this can be accomplished by providing metadata that maps not only the channels / signals (from the entire set) to be played by a particular speaker / device, but also the channels / signals (from the entire set) to be used as echo references for the particular speaker / device. Such metadata or signaling can be dynamic, so that the indication of a preferred echo reference can change over time.
[0080] For use cases where each device / speaker receives only the specific signal it is to play, additional signals (e.g., echo reference signals) may need to be transmitted to each device / speaker in order to provide a suitable echo reference. Furthermore, in order to do so, device-specific signaling needs to be provided so that each device can select the appropriate signal for playback and the appropriate signal for echo management.
[0081] Since the signal used for echo management is only used by the device for performing echo management and not played by the device to a listener, the echo reference signal may be encoded or represented differently than the signal intended to be played back by the device. Specifically, because successful echo management may be achieved with a signal encoded at a lower rate than would normally be used for playback to a listener, additional compression tools (e.g., a parametric representation of the signal that may not be suitable at all for playback to a listener but captures the necessary characteristics of the audio signal) may provide good echo management at a significantly reduced transmission cost.
[0082] Classification Syntax Elements
[0083] In some examples, chunks are used to optimize the transmission of audio. In the described format, each frame may be divided into chunks, as described above with respect to skippable chunks. Chunks may be identified by the frame number to which they belong, a chunk ID that may be used to associate consecutive chunks of different frames with the same ID into a chunk stream, and a retransmission priority. An example of the above is a stream with multiple frames N-2, N-1, and N, frame N comprising multiple chunks identified as ID1, ID2, ID3, which are in Fig.13 In Fig.14 An example of this stream is shown in , which is in a bitstream format. In some examples, for an audio signal of an immersive audio program, a block ID of a block may indicate which group of signals of the entire immersive audio program is carried by the block.
[0084] One use case for the described format is low-latency reliable transmission of audio over wireless networks such as Wifi. For example, Wifi uses a packet-based network protocol. Packet sizes are usually limited. A typical maximum packet size in IP networks is 1500 bytes. The block-based architecture of the stream allows flexibility in assembling packets for transmission. For example, packets with smaller frames can be filled with retransmitted blocks from other frames. Large frames can be split on block boundaries before packing to reduce dependencies between packets at the network protocol layer.
[0085] Fig. 9 The relationship between frames, blocks and packets is shown. A frame carries audio data, preferably all audio data, representing a continuous segment of an audio signal having a start time, an end time, and a duration that is the difference between the end time and the start time. A continuous segment may include a time period according to ISO / IEC 14496-3, subpart 4, section 4.5.2.1.1. Section 4.5.2.1.1 describes the contents of raw_data_block(). A frame may also carry a redundant representation of the segment, for example, encoded at a lower data rate. After encoding, the frame may be segmented into blocks. These blocks may be combined into packets for transmission over a packet-based network. Blocks from different frames may be combined in a single packet and / or may be transmitted out of order.
[0086] In the example, blocks are used to address individual devices. Data packets are received by individual devices or groups of related devices. The concept of skippable blocks can be used to address individual devices or groups of related devices. Even if the network operates in broadcast mode when sending packets to different devices, the processing of the audio (e.g., decoding, rendering, etc.) can be reduced to the blocks addressed to the device. All other blocks, even if received within the same packet, can be simply skipped. In some examples, blocks can be introduced in the correct order based on their decoding or presentation time. If the same block with a higher priority has also been received, the retransmitted block with a lower priority can be removed. The block stream can then be fed to the decoder.
[0087] In some examples, the configuration of the stream and the device is transmitted out-of-band. The codec allows a connection to be established where the audio stream is transmitted at a fairly high rate but with low latency. The configuration of such a connection can remain stable for the duration of such a connection. In that case, instead of making the configuration part of the audio stream, it can be transmitted out-of-band. For such out-of-band transmission, even different networks or network protocols can be used. For example, the audio stream can use the User Datagram Profile (UDP) for low-latency transmission, while the configuration can use the Transmission Control Protocol (TCP) to ensure reliable transmission of the configuration.
[0088] Enable codecs in MPEG-4 audio contexts
[0089] One specific application of the present technology is for use within MPEG-4 Audio. In MPEG-4 Audio, different AOTs (Audio Object Types) are defined for different codec technologies. For the formats described in this document, new AOTs can be defined that allow signaling and data specific to that format. Furthermore, the configuration of a decoder in MPEG-4 is done in a DecoderSpecificInfo() payload, which in turn carries an AudioSpecificConfig() payload. In the latter, certain generic signaling that is agnostic to that specific format, such as sampling rate and channel configuration, is defined, as well as specific information for a specific AOT. For traditional formats, where the entire stream is decoded by a single device, this may make sense. However, in broadcast mode, where a single stream is transmitted to several decoders, where each respective decoder only decodes a portion of the stream, pre-channel configuration signaling (as a means of setting the output capabilities of the device) may be suboptimal.
[0090] Fig.10 The traditional MPEG-4 high-level structure is shown (in black), with modifications in grey. To support broadcast use cases, codecSpecificConfig() is defined (where "codec" may be a generic placeholder name), where the signalling is redefined for the specific use case, such that specific channel elements may be mapped to specific devices, as well as including other relevant static parameters. The MPEG-4 element channelConfiguration with value "0" is defined as the channel configuration defined in codecSpecificConfig. Thus, this value may be used to enable revisions to the signalling of the channel configuration within a codec specific configuration.
[0091] Further, in the spirit of MPEG-4, the raw payload is specified for the specific decoder at hand, assuming codecSpecificConfig() is decodable. However, the format ensures that dynamic metadata is part of the raw payload and that length information is available for all raw payloads so that the decoder can easily skip elements that are not relevant to a specific device.
[0092] exist Fig.11 In , a part of an example raw_data_block as defined in MPEG-4 is given on the right. A raw data block contains channel elements (Single Channel Elements (SCE) or Channel Pair Elements (CPE)) in a given order. However, in conventional MPEG-4 audio syntax, a decoder wishing to skip parts of these channel elements that may not be relevant to the output device at hand will have to parse (and to some extent also decode) all channel elements to be able to extract the relevant parts. Fig.11 In the new raw data block shown on the left, the content will consist of skippable blocks, so that the decoder can skip irrelevant parts and only decode the channel elements indicated by the metadata as relevant to the device at hand. In one example, the skippable block includes raw_data_block and related information.
[0093] It will also be possible to use blocks for retransmission. Each frame can be divided into blocks, as described in the skippable blocks section above. A block can be identified by the frame number it belongs to, a block ID that can be used to associate consecutive blocks of different frames with the same ID into a block stream, and a retransmission priority, as described in the skippable blocks section above. Fig.15 and 16 For example, a high retransmission priority (e.g., Fig.15 and Fig.16 0 in the , signals that a block will take precedence at the receiver over another block with the same block ID and frame count but with a lower retransmission priority (e.g., Fig.15 and Fig.16 It should be noted that, if Fig.15 and Fig.16 In some examples, the priority may decrease as the priority index increases (e.g., priority 1 may be a lower priority than priority 0), while in other examples, the priority may increase as the priority index increases (e.g., priority 0 is a lower priority than priority 1). Still other examples will be apparent to those skilled in the art.
[0094] The syntax may support retransmission of audio elements. For retransmissions, various quality levels may be supported. Thus, retransmitted blocks may carry a 'priority' flag to indicate which blocks with the same block ID should be given priority to the decoder, since received blocks with the same frame count and block ID are redundant and therefore mutually exclusive to the decoder, as Fig.17 and 18 as shown in .
[0095] The retransmission of the block may be performed at a reduced data rate. Such a reduced data rate may be achieved by reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, reducing the channel count of the audio signal (e.g., as described in U.S. Pat. No. 11,289,103, which is incorporated herein by reference in its entirety), or any combination thereof. In order for the decoder to select the audio block that provides the best possible quality, the block that provides the highest quality signal may have the highest decoding priority, the block with the next highest quality signal may have the next highest priority, and so on.
[0096] The block may also be retransmitted at the same quality level. In this case, the priority may reflect the delay of the retransmitted block.
[0097] Tools to improve core coding efficiency
[0098] Another example may allow sharing of MDCT quantization scale factors across channel elements. In some cases, MDCT scale factors may be shared across certain channels to reduce the associated side bit rate. In addition, scale factor sharing may be extended to different channel elements, for example for 7.1.4 inputs. The use of shared scale factors may be indicated by 2 signaling bits that allow 3 active configurations. One possible configuration would be to share scale factors across the left horizontal channel, the left top channel, the right horizontal channel, and the right top channel. Specific syntax may be combined with the skippable block concept to ensure that scale factors are only shared within a skip block.
[0099] It is also possible to have joint coding of more than two channels. Figure 4 , 5 , 6 and 7, block Blk 1 carries three SCEs, one for each of the drivers in the smart speaker considered in this scenario. In such a case, it may be beneficial to enable joint coding of not just two channels (as provided by CPE, which can be extended to include stereo prediction, for example, as described in ISO / IEC 23003-3, the MDCT-based complex prediction stereo tool, known as MPEG-D USAC), but more than two channels, for example as described in the SAP (Stereo Audio Processing) tool introduced in ETSI TS103 190.
[0100] Similarly, it should be noted that in the case of flexible rendering, there may be high correlation between signals across devices. For such signals, it may be beneficial to construct skip blocks that cover multiple devices and allow joint channel coding to be applied on more than two channels using the tools outlined above.
[0101] For certain transform lengths, another joint channel coding tool that can be used is channel coupling, where composite channel and scale factor information is transmitted for medium and high frequencies. This can provide bit rate reduction for playback in the good quality range. For example, for a transform length of 256 frames, corresponding to a frame length of about 256 samples for low-delay coding, it may be beneficial to use the channel coupling tool.
[0102] The present disclosure will also allow for efficient encoding of band-limited signals. In scenarios where separate signals are sent to feed different drivers of a smart speaker, some of these signals may be band-limited, such as in the case of a three-way driver configuration with a woofer, midrange, and tweeter, for example. It is therefore desirable to efficiently encode such band-limited signals, which can translate into specially tuned psychoacoustic models and bit allocation strategies as well as potential modifications of the syntax to handle this case with improved coding efficiency and / or reduced computational complexity (e.g., enabling the use of band-limited IMDCT for woofer feeds).
[0103] In some examples, a combination of high frequency reconstruction and noise addition in the MDCT domain will be possible. Modern audio codecs are typically designed to support parametric coding techniques, and similarly, it is conceivable to include, for example, a high frequency reconstruction method in the MDCT domain to retain low latency in the MDCT domain and to perform a noise addition scheme to manage the tone-to-noise ratio in the reconstructed high frequency band. Such parametric coding techniques are particularly useful for lower operating points, and particularly in scenarios where retransmission occurs as part of an FEC (forward error correction) scheme, where a main signal is typically transmitted and then the same signal is retransmitted at a lower bit rate in a delayed manner.
[0104] These examples will further allow peak complexity reduction in the encoder. This will be applicable in the case where the buffer model restrictions for constant bit rate transmission channels are not in place. In the case of constant bit rate with buffer mode, it may happen that the number of bits obtained for the frame to be encoded is higher than the limit of the allowed bits after quantization and counting. At least one new, coarser quantization and bit counting step is required to meet the bit requirements. If the restriction is looser, the encoder can keep the first quantization result and will be completed at a slightly higher instantaneous bit rate. The encoder behavior similar to the constant bit rate encoder with buffer model can still be achieved by using the same bit storage control mechanism and the virtual buffer fullness updated according to the encoder that complies with the buffer requirements. This will save additional quantization and bit counting steps, and the resulting audio quality will be the same or better than the constant bit rate case with the disadvantage of a slight increase in the total bit rate.
[0105] Return Channel Concept
[0106] For playback and listening use cases, the smart speaker 30 may have one or more microphones 40 (e.g., a microphone or microphone array) configured to capture monophonic or spatial sound fields (e.g., mono, stereo, A-format, B-format, or any other isotropic or anisotropic channel format), and require a codec capable of efficient encoding of such formats in a low-latency manner, so that the same codec is used for broadcasting / transmission to the smart speaker device 30 as for the return channel, e.g. Fig.12 shown.
[0107] It should also be noted that while wake word detection is performed on the smart speaker device, speech recognition is typically done in the cloud, where the appropriate segment of recorded speech is triggered by the wake word detection.
[0108] In this context, it may be interesting to save system complexity overall by creating a speech analysis system on an intermediate format / representation, e.g. within a codec. For the simplest use case, where no human-to-human conversation is taking place, a suitable representation specific to the speech recognition task can be defined, since it will not be listened to by another person. This could be band energies, Mel-Frequency Cepstral Coefficients (MFCC), etc., or a low-bitrate version of the MDCT-coded spectrum.
[0109] In the use case where there is an ongoing human-to-human conversation and there is an intertwined wake-up word and required speech recognition, one might want to be able to extract the relevant portion of the stream for this as well. Such an extraction might require a layered coding structure where simple additional data is transmitted in parallel with the main speech signal. Transcoding from the existing decoded MDCT spectrum to the definition of the relevant representation is also envisioned, and doing so simplifies the structure. Essentially, the decoder will decode and output human audible audio until it is signaled that it should also decode the speech recognition representation (in parallel). This representation can be a "stripped" layered stream, simply an alternative decoding of the same complete data, or simply the decoding and output of an additional representation in the stream. In this use case, the signaling is an enabling segment that indicates that the decoder on the receiving side should output the speech recognition related representation.
[0110] List examples
[0111] In the following, seven sets of non-claimed exemplary examples (EEE-A, EEE-B, EEE-C, EEE-D, EEE-E, EEE-F, and EEE-G) describe aspects of the examples disclosed herein.
[0112] EEE-A1. A method for decoding an audio signal, the method comprising:
[0113] receiving a bitstream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks;
[0114] determining, from the signaling data, information for identifying a portion of one or more blocks of the plurality of blocks to be skipped in decoding, based on device information of an output device; and
[0115] The bitstream is decoded while skipping the portion of the identified one or more blocks.
[0116] EEE-A2. A method according to EEE-A1, wherein the information for identifying portions of one or more of the plurality of blocks to be skipped when decoding comprises a matrix associating each of a plurality of output devices with one or more bitstream elements.
[0117] EEE-A3. A method according to EEE-A2, wherein the one or more bitstream elements are required for decoding the bitstream for a corresponding associated output device.
[0118] EEE-A4. A method according to any one of EEE-A1 to EEE-A3, wherein the output device may include at least one of a wireless device, a mobile device, a tablet computer, a mono speaker and / or a multi-channel speaker.
[0119] EEE-A5. A method according to any one of EEE-A1 to EEE-A4, wherein the identified portion includes at least one block.
[0120] EEE-A6. A method according to any one of EEE-A1 to EEE-A5, wherein the output device is a first output device, and further comprising applying a joint coding technique between one or more signals of the bit streams to the second output device and the third output device.
[0121] EEE-A7. A method according to any one of EEE-A1 to EEE-A6, wherein the identity of each output device and / or decoder is defined during a system initialization phase.
[0122] EEE-A8. A method according to any one of EEE-A1 to EEE-A7, wherein the signaling data is determined by metadata of the bitstream.
[0123] EEE-A9. An apparatus configured to perform a method according to any one of EEE-A1 to EEE-A8.
[0124] EEE-A10. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, cause one or more devices to perform a method according to any one of EEE-A1 to EEE-A8.
[0125] EEE-B1. A method for generating an encoded bitstream from an audio program comprising a plurality of audio signals, the method comprising:
[0126] for each of the plurality of audio signals, receiving information indicating a playback device associated with the respective audio signal;
[0127] for each playback device, receiving information indicative of at least one of a delay, a gain, and an equalization curve associated with the respective playback device;
[0128] determining a group of two or more related audio signals from the plurality of audio signals;
[0129] applying one or more joint coding tools to the two or more related audio signals in the group to obtain a jointly encoded audio signal;
[0130] The jointly encoded audio signal, an indication of a playback device associated with the jointly encoded audio signal, and an indication of delays and gains associated with the respective playback devices associated with the jointly encoded audio signal are combined into separate blocks of an encoded bitstream.
[0131] EEE-B2. A method according to EEE-B1, wherein the delay, gain and / or equalization curve associated with the corresponding playback device depends on the position of the corresponding playback device relative to the listener position.
[0132] EEE-B3. A method according to EEE-B1 or EEE-B2, wherein the delay, gain and / or equalization curve associated with the corresponding playback device depends on the position of the corresponding playback device relative to the position of other playback devices.
[0133] EEE-B4. A method according to any one of EEE-B1 to EEE-B3, wherein the delay, gain and / or equalization curve is dynamically variable.
[0134] EEE-B5. A method according to EEE-B4, wherein the delay, gain and / or equalization curves are adjusted in response to changes in the position of the listener.
[0135] EEE-B6. A method according to EEE-B4 or EEE-B5, wherein the delay, gain and / or equalization curve is adjusted in response to a change in the position of the playback device.
[0136] EEE-B7. A method according to any one of EEE-B4 to EEE-B6, wherein the delay, gain and / or equalization curve is adjusted in response to a change in the position of one or more of the other playback devices.
[0137] EEE-B8. The method according to any one of EEE-B1 to EEE-B7, further comprising determining an audio signal from the plurality of audio signals that is not part of the group of two or more related audio signals.
[0138] EEE-B9. The method according to EEE-B8 also includes, for an audio signal that is not part of the group of two or more related audio signals, applying a delay, gain and / or equalization curve associated with a playback device associated with the audio signal.
[0139] EEE-B10. The method according to EEE-B9 also includes independently encoding audio signals that are not part of the group of two or more related audio signals, and combining the independently encoded audio signals and indications of playback devices associated with the independently encoded audio signals into separate independently decodable subsets of the encoded bitstream.
[0140] EEE-B11. The method according to EEE-B8 also includes independently encoding audio signals that are not part of the group of two or more related audio signals, and combining the independently encoded audio signals, indications of playback devices associated with the independently encoded audio signals, and indications of delay, gain and / or equalization curves associated with the playback devices associated with the independently encoded audio signals into separate independently decodable subsets of the encoded bitstream.
[0141] EEE-B12. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises one or more independent encoded data blocks, the method comprising:
[0142] identifying, from the encoded bitstream, independent blocks of encoded data corresponding to the one or more audio signals associated with the playback device;
[0143] extracting the identified independent coded data blocks from the coded bitstream;
[0144] determining that the extracted independent coded data block comprises two or more jointly coded audio signals;
[0145] applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device;
[0146] determining at least one of a delay, gain, and equalization curve associated with the playback device from the extracted independent encoded data blocks;
[0147] The delay, gain, and / or equalization curve associated with the playback device is applied to the one or more audio signals associated with the playback device.
[0148] EEE-B13. A method according to EEE-B12, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to the position of the listener.
[0149] EEE-B14. A method according to EEE-B12 or EEE-B13, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices.
[0150] EEE-B15. A method according to any one of EEE-B12 to EEE-B14, wherein the determined delay, gain and / or equalization curve of the playback device is dynamically variable.
[0151] EEE-B16. A method according to EEE-B15, wherein, when the determined delay, gain and / or equalization curve associated with the playback device is different from the previously determined delay, gain and / or equalization curve associated with the playback device, the method further includes interpolating between the previously determined delay, gain and / or equalization curve associated with the playback device and the determined delay, gain and / or equalization curve associated with the playback device.
[0152] EEE-B17. A method according to EEE-B16, wherein the determined delay, gain and / or equalization curve differs from a previously determined delay, gain and / or equalization curve due to a change in the position of the listener.
[0153] EEE-B18. A method according to EEE-B16 or EEE-B17, wherein the determined delay, gain and / or equalization curve is different from a previously determined delay, gain and / or equalization curve due to a change in the position of the playback device.
[0154] EEE-B19. A method according to any one of EEE-B16 to EEE-B18, wherein the determined delay, gain and / or equalization curve is different from a previously determined delay, gain or equalization curve due to a change in the position of one or more other playback devices.
[0155] EEE-B20. A method according to any one of EEE-B12 to EEE-B19, wherein a frame of the coded bitstream comprises two or more independent coded data blocks, the method further comprising:
[0156] determining that one or more of the independent blocks contain audio signals not associated with the playback device; and
[0157] The one or more independent blocks containing audio signals not associated with the playback device are ignored.
[0158] EEE-B21. A method according to any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools includes: identifying a subset of the jointly encoded audio signals associated with the playback device, and reconstructing only the subset of the jointly encoded audio signals to obtain the one or more audio signals associated with the playback device.
[0159] EEE-B22. A method according to any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools includes: reconstructing each of the jointly encoded audio signals, identifying a subset of the reconstructed jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of the reconstructed jointly encoded audio signals associated with the playback device.
[0160] EEE-B23. An apparatus configured to perform a method according to any one of EEE-B1 to EEE-B22.
[0161] EEE-B24. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, causes one or more devices to perform a method according to any one of EEE-B1 to EEE-B22.
[0162] EEE-C1. A method for generating a frame of an encoded bitstream of an audio program comprising a plurality of audio signals, wherein the frame comprises two or more independent encoded data blocks, the method comprising:
[0163] for one or more audio signals of the plurality of audio signals, receiving information indicating a playback device associated with the one or more audio signals;
[0164] For the indicated playback device, receiving information indicating one or more additional associated playback devices;
[0165] receiving one or more audio signals associated with the indicated one or more additional associated playback devices;
[0166] encoding the one or more audio signals associated with the playback device;
[0167] encoding the one or more audio signals associated with the indicated one or more additional associated playback devices;
[0168] combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block;
[0169] combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and
[0170] The first independent block and the one or more additional independent blocks are combined into the frame of the encoded bitstream.
[0171] EEE-C2. A method according to EEE-C1, wherein the plurality of audio signals comprises one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices, the method further comprising:
[0172] encoding each of one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices into respective independent blocks; and
[0173] The respective independent blocks of each of the one or more groups are combined into the frame of the encoded bitstream.
[0174] EEE-C3. A method according to EEE-C1 or EEE-C2, wherein one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended to be used as an echo reference for performing echo management on the playback device.
[0175] EEE-C4. A method according to EEE-C3, wherein the one or more audio signals intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device. EEE-C4.
[0176] EEE-C5. A method according to EEE-C3 or EEE-C4, wherein the one or more audio signals intended for use as echo references are encoded using a parametric coding tool.
[0177] EEE-C6. A method according to EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.
[0178] EEE-C7. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises two or more independent encoded data blocks, wherein the playback device comprises one or more microphones, the method comprising:
[0179] identifying, from the encoded bitstream, independent blocks of encoded data corresponding to the one or more audio signals associated with the playback device;
[0180] extracting the identified independent coded data blocks from the coded bitstream;
[0181] extracting the one or more audio signals associated with the playback device from the identified independent encoded data blocks;
[0182] identifying from the encoded bitstream one or more other independent encoded data blocks corresponding to one or more audio signals associated with one or more other playback devices;
[0183] extracting the one or more audio signals associated with the one or more other playback devices from the one or more other independent encoded data blocks;
[0184] capturing one or more audio signals using the one or more microphones of the playback device; and
[0185] The extracted one or more audio signals associated with the one or more other playback devices are used as an echo reference for performing echo management on the playback device in response to the captured one or more audio signals.
[0186] EEE-C8. The method according to EEE-C7, further comprising:
[0187] determining that the coded bitstream includes one or more additional independent coded data blocks; and
[0188] The one or more additional independent coded data blocks are ignored.
[0189] EEE-C9. A method according to EEE-C8, wherein ignoring the one or more additional independent encoded data blocks includes skipping the one or more additional independent encoded data blocks without extracting the one or more additional independent encoded data blocks.
[0190] EEE-C10. A method according to any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are specifically intended to be used as an echo reference for performing echo management on the playback device.
[0191] EEE-C11. A method according to EEE-C10, wherein the one or more audio signals specifically intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device. EEE-C12.
[0192] EEE-C12. A method according to EEE-C10 or EEE-C11, wherein the one or more audio signals specifically intended for use as echo references are reconstructed based on parameter representations of the one or more audio signals.
[0193] EEE-C13. A method according to any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.
[0194] EEE-C14. A method according to EEE-C7, wherein the encoded signal includes signaling information indicating that the one or more other playback devices are used as an echo reference for the playback device.
[0195] EEE-C15. A method according to EEE-C14, wherein the one or more other playback devices indicated by the signaling information for the current frame are different from the one or more other playback devices used as echo references for the previous frame.
[0196] EEE-C16. An apparatus configured to perform a method according to any one of EEE-C1 to EEE-C15.
[0197] EEE-C17. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, causes one or more devices to perform a method according to any one of EEE-C1 to EEE-C15.
[0198] EEE-D1. A method for transmitting an audio signal, the method comprising:
[0199] generating a data packet comprising a portion of a bitstream, wherein the bitstream comprises a plurality of frames, wherein each frame of the plurality of frames comprises a plurality of blocks, wherein the generating comprises:
[0200] integrating a data packet with one or more of the plurality of blocks, wherein blocks from different frames are combined into a single packet and / or transmitted out of order; and
[0201] The data packets are transmitted via a packet-based network.
[0202] EEE-D2. A method according to EEE-D1, wherein each of the plurality of blocks includes identification information.
[0203] EEE-D3. A method according to EEE-D2, wherein the identification information includes at least one of a block ID, a corresponding frame number associated with the block, and / or a retransmission priority.
[0204] EEE-D4. A method according to any one of EEE-D1 to EEE-D3, wherein each frame of the plurality of frames carries all audio data representing a continuous segment of an audio signal having a start time, an end time and a duration.
[0205] EEE-D5. A method for decoding an audio signal, the method comprising:
[0206] receiving a data packet comprising a portion of a bitstream, wherein the bitstream comprises a plurality of frames, wherein each frame of the plurality of frames comprises a plurality of blocks;
[0207] determining a set of blocks from the plurality of blocks that are addressed to a device; and
[0208] The set of blocks addressed to the device are decoded, and decoding of blocks in the plurality of blocks not addressed to the device is skipped.
[0209] EEE-D6. A method for transmitting an audio stream, the method comprising:
[0210] The audio stream is transmitted, wherein the audio stream comprises a plurality of frames, wherein each frame of the plurality of frames comprises a plurality of blocks, wherein the transmitting comprises transmitting configuration information of the audio stream out-of-band.
[0211] EEE-D7. The method according to EEE-D6, wherein transmitting the configuration information of the audio stream out-of-band comprises:
[0212] transmitting the audio stream via a first network and / or a first network protocol; and
[0213] The configuration information is transmitted via a second network and / or a second network protocol.
[0214] EEE-D8. A method according to EEE-D7, wherein the first network protocol is a user datagram protocol (UDP) and the second network protocol is a transmission control protocol (TCP).
[0215] EEE-D9. A method for decoding an audio signal, the method comprising:
[0216] receiving a bitstream, the bitstream comprising: information corresponding to signaling of static configuration aspects; static metadata; and
[0217] One or more channel elements are mapped to one or more devices based on the information and / or the static metadata.
[0218] EEE-D10. A method according to EEE-D9, wherein the bitstream is received by a plurality of decoders configured to decode the bitstream, wherein each of the plurality of decoders is configured to decode a portion of the bitstream.
[0219] EEE-D11. A method according to EEE-D9 or EEE-D10, wherein the bitstream also includes dynamic metadata.
[0220] EEE-D12. A method according to any one of EEE-D9 to EEE-D11, wherein the bitstream comprises a plurality of blocks, wherein each of the plurality of blocks comprises: information enabling a portion of the block to be skipped during decoding, wherein the portion is not required by the device; and dynamic metadata.
[0221] EEE-D13. A method for retransmitting a block of an audio signal, the method comprising:
[0222] transmitting one or more blocks of a bitstream, wherein the bitstream comprises a plurality of blocks, wherein each of the one or more blocks of the bitstream has been previously transmitted; and
[0223] Each of the one or more blocks includes a decoding priority indicator.
[0224] EEE-D14. A method according to EEE-D13, wherein the decoding priority indicator indicates to a decoder a priority order for decoding the one or more blocks of the bitstream.
[0225] EEE-D15. A method according to EEE-D13 or EEE-D14, wherein each of the one or more blocks includes the same block ID.
[0226] EEE-D16. A method according to any one of EEE-D13 to EEE-D15, wherein the transmission of the one or more blocks of the bitstream is transmitted by reducing the data rate compared to the previous transmission.
[0227] EEE-D17. A method according to EEE-D16, wherein reducing the data rate includes at least one of reducing a signal-to-noise ratio of the audio signal, reducing a bandwidth of the audio signal, and / or reducing a channel count of the audio signal. EEE-D18. The method of claim 15, wherein reducing the data rate comprises reducing a signal-to-noise ratio of the audio signal, reducing a bandwidth of the audio signal, and / or reducing a channel count of the audio signal.
[0228] EEE-E1. A method for generating a frame of an encoded bitstream of an audio program comprising a plurality of audio signals, wherein the frame comprises one or more independent encoded data blocks, the method comprising:
[0229] for each of the plurality of audio signals, receiving information indicating a playback device associated with the respective audio signal;
[0230] encoding one or more audio signals associated with a respective playback device to obtain one or more encoded audio signals;
[0231] combining the one or more encoded audio signals associated with the respective playback devices into a first independent block of the frame;
[0232] encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; and
[0233] The first independent block and the one or more additional independent blocks are combined into the frame of the encoded bitstream.
[0234] EEE-E2. A method according to EEE-E1, wherein two or more audio signals are associated with the playback device, and each of the two or more audio signals is a band-limited signal intended to be played back by a corresponding driver of the playback device, and wherein a different encoding technique is used for each band-limited signal.
[0235] EEE-E3. A method according to EEE-E2, wherein a different psychoacoustic model and / or a different bit allocation technique is used for each band-limited signal.
[0236] EEE-E4. A method according to any one of EEE-E1 to EEE-E3, wherein the instantaneous frame rate of the encoded signal is variable and constrained by a buffer fullness model.
[0237] EEE-E5. A method according to EEE-E1, wherein encoding one or more audio signals associated with the corresponding playback device includes jointly encoding one or more audio signals associated with the corresponding playback device and one or more additional audio signals associated with one or more additional playback devices into a first independent block of the frame.
[0238] EEE-E6. The method of EEE-E5, wherein jointly encoding the one or more audio signals and the one or more additional audio signals comprises sharing one or more scale factors across two or more audio signals. ...
[0239] EEE-E7. A method according to EEE-E6, wherein the two or more audio signals are spatially correlated.
[0240] EEE-E8. The method of EEE-E7, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left top channel, a right horizontal channel, or a right top channel. EEE-E8.
[0241] EEE-E9. A method according to EEE-E5, wherein jointly encoding the one or more audio signals and the one or more additional audio signals comprises applying a coupling tool, comprising:
[0242] combining two or more audio signals into a composite signal above a specified frequency; and
[0243] For each of the two or more audio signals, a scaling factor is determined that relates the energy of the composite signal to the energy of each respective signal.
[0244] EEE-E10. The method according to EEE-E5, wherein jointly encoding the one or more audio signals and the one or more additional audio signals comprises applying a joint coding tool to more than two signals. EEE-E11.
[0245] EEE-E11. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises one or more independent encoded data blocks, the method comprising:
[0246] identifying, from the encoded bitstream, independent blocks of encoded data corresponding to the one or more audio signals associated with the playback device;
[0247] extracting the identified independent coded data blocks from the coded bitstream;
[0248] decoding the one or more audio signals associated with the playback device from the independent encoded data blocks to obtain one or more decoded audio signals;
[0249] identifying from the coded bitstream one or more additional independent coded data blocks corresponding to one or more additional audio signals; and
[0250] The one or more additional independent coded data blocks are decoded or skipped.
[0251] EEE-E12. A method according to EEE-E11, wherein two or more audio signals are associated with the playback device, and each of the two or more audio signals is a band-limited signal intended to be played back by a corresponding driver of the playback device, and wherein different decoding techniques are used to decode the two or more audio signals.
[0252] EEE-E13. A method according to EEE-E12, wherein each band-limited signal is encoded using a different psychoacoustic model and / or a different bit allocation technique.
[0253] EEE-E14. A method according to EEE-E11 or EEE-E13, wherein the instantaneous frame rate of the encoded bit stream is variable and constrained by a buffer fullness model.
[0254] EEE-E15. A method according to EEE-E11, wherein decoding one or more audio signals associated with a playback device comprises jointly decoding one or more audio signals associated with the corresponding playback device and one or more additional audio signals associated with one or more additional playback devices from independent encoded data blocks.
[0255] EEE-E16. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises extracting one or more scale factors shared across two or more audio signals. EEE-E17.
[0256] EEE-E17. A method according to EEE-E16, wherein the two or more audio signals are spatially correlated.
[0257] EEE-E18. The method of EEE-E17, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left top channel, a right horizontal channel, or a right top channel. EEE-E19. The method of claim 11, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left top channel, a right horizontal channel, or a right top channel.
[0258] EEE-E19. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises applying a decoupling tool.
[0259] EEE-E20. The method according to EEE-E19, wherein the decoupling tool comprises:
[0260] Extract independent decoding signals below the specified frequency;
[0261] Extracting the composite signal above the specified frequency;
[0262] determining individual decoupled signals above the specified frequency from the composite signal and a scaling factor relating the energy of the composite signal to the energy of the individual signals; and
[0263] Each independently decoded signal is combined with the corresponding decoupled signal to obtain a jointly decoded signal.
[0264] EEE-E21. The method according to EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises applying a joint decoding tool to extract more than two audio signals. EEE-E22.
[0265] EEE-E22. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device comprises applying bandwidth extension to the audio signals in the same domain as the audio signals were encoded. EEE-E23. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device comprises applying bandwidth extension to the audio signals in the same domain as the audio signals were encoded.
[0266] EEE-E23. A method according to EEE-E22, wherein the domain is a modified discrete cosine transform (MDCT) domain.
[0267] EEE-E24. A method according to EEE-E22 or EEE-E23, wherein bandwidth extension includes adaptive noise addition.
[0268] EEE-E25. An apparatus configured to perform a method according to any one of EEE-E1 to EEE-E24.
[0269] EEE-E26. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, causes one or more devices to perform a method according to any one of EEE-E1 to EEE-E24.
[0270] EEE-F1. A method for generating a coded bit stream performed by a device having one or more microphones, the method comprising:
[0271] capturing one or more audio signals by the one or more microphones;
[0272] analyzing the captured audio signal to determine the presence of a wake word;
[0273] When the wake word is detected:
[0274] Setting a flag to indicate that a speech recognition task is to be performed on the captured audio signal;
[0275] encoding the captured audio signal;
[0276] The encoded audio signal and the flag are integrated into the encoded bit stream.
[0277] EEE-F2. A method according to EEE-F1, wherein the one or more microphones are configured to capture a monophonic or spatial sound field.
[0278] EEE-F3. A method according to EEE-F2, wherein the spatial sound field is in A format or B format.
[0279] EEE-F4. A method according to any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended only for performing the speech recognition task.
[0280] EEE-F5. A method according to EEE-F4, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for performing a speech recognition task but is insufficient for human listening.
[0281] EEE-F6. A method according to EEE-F4 or EEE-F5, wherein before encoding the captured audio signal, the captured audio signal is converted into a representation comprising one or more of frequency band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0282] EEE-F7. A method according to any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended for human listening and for performing the speech recognition task.
[0283] EEE-F8. A method according to EEE-F7, wherein the captured audio signal is encoded so that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening.
[0284] EEE-F9. A method according to EEE-F7, wherein encoding the captured audio signal includes generating a first encoded representation of the captured audio signal and a second encoded representation of the captured audio signal, wherein the first encoded representation is generated so that when the captured audio signal is decoded from the first encoded representation, the quality of the decoded audio signal is sufficient for human listening, and wherein the second encoded representation is generated so that when the captured audio signal is decoded from the second encoded representation, the quality of the decoded audio signal is sufficient for performing a speech recognition task but not sufficient for human listening.
[0285] EEE-F10. A method according to EEE-F9, wherein generating a second encoded representation of the captured audio signal includes: before encoding the captured audio signal, converting the captured audio signal into one or more of a parametric representation, a coarse waveform representation, or a representation including one or more of frequency band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0286] EEE-F11. A method according to EEE-F9 or EEE-F10, wherein integrating the encoded audio signal into a bitstream comprises inserting a first encoded representation into a first independent block of the encoded bitstream, and inserting a second encoded representation into a second independent block of the encoded bitstream.
[0287] EEE-F12. A method according to EEE-F9 or EEE-F10, wherein the first encoded representation is included in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first layer and the second layer are included in a single block of the encoded bitstream.
[0288] EEE-F13. The method according to any one of EEE-F1 to EEE-F12, further comprising: when the presence of the wake-up word is not detected:
[0289] Setting the flag to indicate that a speech recognition task is not to be performed on the captured audio signal;
[0290] encoding the captured audio signal;
[0291] The encoded audio signal and the flag are integrated into the encoded bit stream.
[0292] EEE-F14. An audio signal decoding method, comprising:
[0293] receiving a coded bit stream, the coded bit stream comprising a coded audio signal and a flag indicating whether a speech recognition task is to be performed;
[0294] decoding the encoded audio signal to obtain a decoded audio signal; and
[0295] When the flag indicates that the speech recognition task is to be performed, the speech recognition task is performed on the decoded audio signal.
[0296] EEE-F15. A method according to EEE-F14, wherein the decoded audio signal is intended only for performing the speech recognition task.
[0297] EEE-F16. A method according to EEE-F15, wherein the quality of the decoded audio signal is sufficient for performing the speech recognition task but is insufficient for human listening.
[0298] EEE-F17. A method according to EEE-F15 or EEE-F16, wherein, before encoding the captured audio signal, the representation of the decoded audio signal includes one or more of frequency band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0299] EEE-F18. The method of EEE-F14, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening. EEE-F18.
[0300] EEE-F19. A method according to EEE-F18, wherein the encoded audio signal comprises a first encoded representation of one or more audio signals and a second encoded representation of the one or more audio signals.
[0301] EEE-F20. A method according to EEE-F18, wherein the quality of the audio signal decoded from the first representation is sufficient for human listening, and wherein the quality of the audio signal decoded from the second representation is sufficient for performing a speech recognition task but is insufficient for human listening.
[0302] EEE-F21. A method according to EEE-F19 or EEE-F20, wherein the first representation is in a first independent block of the encoded bitstream and the second representation is in a second independent block of the encoded bitstream.
[0303] EEE-F22. A method according to EEE-F19 or EEE-F20, wherein the first representation is in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first layer and the second layer are included in a single block of the encoded bitstream.
[0304] EEE-F23. A method according to any one of EEE-F18 to EEE-F22, wherein decoding the encoded audio signal comprises decoding only the second representation and ignoring the first representation.
[0305] EEE-F24. A method according to any one of EEE-F18 to EEE-F23, wherein the audio signal decoded from the second encoded representation is a parametric representation, a waveform representation, or a representation comprising one or more of frequency band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0306] EEE-F25. An apparatus configured to perform a method according to any one of EEE-F1 to EEE-F24.
[0307] EEE-F26. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, causes one or more devices to perform a method according to any one of EEE-F1 to EEE-F24.
[0308] EEE-G1. A method for encoding an audio signal of an immersive audio program for low-delay transmission to one or more playback devices, the method comprising:
[0309] receiving a plurality of time-domain audio signals of the immersive audio program;
[0310] Select the frame size;
[0311] extracting a frame of the time-domain audio signal in response to the frame size, wherein the frame of the time-domain audio signal overlaps with a previous frame of the time-domain audio signal;
[0312] Splitting the audio signal into overlapping frames;
[0313] transforming frames of the time-domain audio signal into frequency-domain signals;
[0314] Encoding the frequency domain signal;
[0315] quantize the coded frequency domain signal using a perceptually motivated quantization tool;
[0316] integrating the quantized and encoded frequency domain signal into one or more independent blocks within a frame; and
[0317] The one or more independent blocks are integrated into a coded frame.
[0318] EEE-G2. A method according to EEE-G1, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration. EEE-G2.
[0319] EEE-G3. A method according to EEE-G2, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6 or 22.2.
[0320] EEE-G4. A method according to any one of EEE-G1 to EEE-G3, wherein the plurality of audio signals comprises one or more object-based signals. EEE-G4.
[0321] EEE-G5. A method according to any one of EEE-G1 to EEE-G4, wherein the plurality of audio signals comprises a scene-based representation of the immersive audio program.
[0322] EEE-G6. A method according to any one of EEE-G1 to EEE-G5, wherein the selected frame size is one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.
[0323] EEE-G7. A method according to any one of EEE-G1 to EEE-G6, wherein the overlap between the frame of the time-domain audio signal and the previous frame of the time-domain audio signal is 50% or less.
[0324] EEE-G8. A method according to any one of EEE-G1 to EEE-G7, wherein the transform is a modified discrete cosine transform (MDCT).
[0325] EEE-G9. A method according to any one of EEE-G1 to EEE-G8, wherein two or more of the multiple audio signals are jointly encoded.
[0326] EEE-G10. A method according to any one of EEE-G1 to EEE-G9, wherein each independent block contains an encoded signal for one or more playback devices.
[0327] EEE-G11. A method according to any one of EEE-G1 to EEE-G10, wherein at least one independent block contains an encoded signal for two or more playback devices, and the encoded signal includes a jointly encoded audio signal.
[0328] EEE-G12. A method according to any one of EEE-G1 to EEE-G11, wherein at least one independent block contains multiple encoded signals covering different bandwidths intended for playback from different drives of a playback device.
[0329] EEE-G13. A method according to any one of EEE-G1 to EEE-G12, wherein at least one independent block contains a coded echo reference signal for echo management performed by a playback device.
[0330] EEE-G14. A method according to any one of EEE-G1 to EEE-G13, wherein encoding the quantized frequency domain signal includes applying one or more of the following tools: temporal noise shaping (TNS), joint channel coding, sharing scaling factors across signals, determining control parameters for high frequency reconstruction, and determining control parameters for noise replacement.
[0331] EEE-G15. A method according to any one of EEE-G1 to EEE-G14, wherein one or more independent blocks include parameters for controlling one or more of delay, gain, and equalization of the playback device.
[0332] EEE-G16. A low-latency method for decoding an audio signal of an immersive audio program from an encoded signal, the method comprising:
[0333] receiving a coded frame comprising one or more independent blocks;
[0334] extracting a quantized and encoded frequency domain signal from the one or more independent blocks;
[0335] Dequantizing the quantized and coded frequency domain signal;
[0336] Decoding the dequantized frequency domain signal;
[0337] Performing an inverse transformation on the decoded frequency domain signal to obtain a time domain signal; and
[0338] The time domain signal is overlapped and added with a time domain signal from a previous frame to provide a plurality of audio signals of the immersive audio program.
[0339] EEE-G17. A method according to EEE-G16, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration. EEE-G17.
[0340] EEE-G18. A method according to EEE-G17, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6 or 22.2.
[0341] EEE-G19. A method according to any one of EEE-G16 to EEE-G18, wherein the plurality of audio signals include one or more object-based signals. EEE-G19.
[0342] EEE-G20. A method according to any one of EEE-G16 to EEE-G19, wherein the plurality of audio signals comprises a scene-based representation of the immersive audio program.
[0343] EEE-G21. A method according to any one of EEE-G16 to EEE-G20, wherein the frame of time domain samples includes one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.
[0344] EEE-G22. A method according to any one of EEE-G16 to EEE-G21, wherein the overlap with the previous frame is 50% or less.
[0345] EEE-G23. A method according to any one of EEE-G16 to EEE-G22, wherein the inverse transform is an inverse modified discrete cosine transform (IMDCT).
[0346] EEE-G24. A method according to any one of EEE-G16 to EEE-G23, wherein each independent block contains a quantized and encoded frequency domain signal for one or more playback devices.
[0347] EEE-G25. A method according to any one of EEE-G16 to EEE-G24, wherein at least one independent block contains quantized and encoded frequency domain signals for two or more playback devices, and the quantized and encoded frequency domain signals are jointly encoded audio signals.
[0348] EEE-G26. A method according to any one of EEE-G16 to EEE-G25, wherein at least one independent block contains multiple quantized and encoded frequency domain signals covering different bandwidths intended for playback from different drives of a playback device.
[0349] EEE-G27. A method according to any one of EEE-G16 to EEE-G26, wherein at least one independent block contains an encoded echo reference signal for echo management performed by a playback device.
[0350] EEE-G28. A method according to any one of EEE-G16 to EEE-G27, wherein decoding the dequantized frequency domain signal includes applying one or more of the following decoding tools: temporal noise shaping (TNS), joint channel decoding, cross-signal sharing of scaling factors, high frequency reconstruction, and noise replacement.
[0351] EEE-G29. A method according to any one of EEE-G16 to EEE-G28, wherein one or more independent blocks include parameters for controlling one or more of delay, gain, and equalization of the playback device.
[0352] EEE-G30. A method according to any one of EEE-G16 to EEE-G29, wherein the method is performed by a playback device, and wherein extracting the quantized and encoded signal from one or more independent blocks includes selecting only those blocks containing quantized and encoded frequency domain signals for playback by the playback device, while ignoring independent blocks containing quantized and encoded frequency domain signals for playback by other playback devices.
[0353] EEE-G31. An apparatus configured to perform a method according to any one of EEE-G1 to EEE-G30.
[0354] EEE-G32. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, causes one or more devices to perform a method according to any one of EEE-G1 to EEE-G30.
[0355] In the following claims and the description herein, any of the terms "comprising", "consisting of" or "which includes" is an open term, which means at least including the elements / features that follow, but not excluding other elements / features. Therefore, when used in the claims, the term "comprising" should not be interpreted as being limited to the devices or elements or steps listed thereafter. For example, the scope of the expression "the device comprises A and B" should not be limited to the device consisting of only elements A and B. As used herein, any of the terms "comprising" or "which includes" is also an open term, which also means at least including the elements / features that follow the term, but not excluding other elements / features. Therefore, including is synonymous with including and means including.
[0356] It should be appreciated that in the above description of examples of the present invention, various features are sometimes grouped together in a single example, figure, or description thereof for the purpose of streamlining the present disclosure and aiding in the understanding of one or more of the various inventive aspects. However, the disclosed method should not be interpreted as reflecting the intention of requiring more features than those explicitly recited in each claim. On the contrary, as reflected in the appended claims, the inventive aspects lie in less than all the features of a single aforementioned disclosed example. Therefore, the claims following the specific embodiments are hereby expressly incorporated into the present specific embodiments, with each claim independently serving as a separate example of the present invention.
[0357] In addition, although some examples described herein include some features included in other examples, but do not include other features, as will be understood by those skilled in the art, combinations of features of different examples are intended to be covered and form different examples. For example, in the following claims, any of the claimed examples may be used in any combination.
[0358] In addition, some of the examples are described herein as methods or combinations of elements of methods that can be implemented by a processor of a computer system or by other means of performing functions. Therefore, a processor with the necessary instructions for performing such a method or elements of a method forms a means for performing such a method or elements of a method. In addition, the elements described herein of a device are examples of means for carrying out the functions performed by the element.
[0359] In addition, some examples described herein disclose connection solutions that should be interpreted as being possible to implement in distribution and / or transmission systems such as wired and / or wireless systems. For example, by using any electrical, optical and / or mobile system, such as 3G, 4G and 5G.
[0360] Thus, while specific examples of the present invention have been described, those skilled in the art will recognize that other and further modifications may be made, and it is intended that all such changes and modifications be claimed. For example, any formula given above is merely representative of a process that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted from the described methods.
[0361] The systems, devices and methods disclosed above may be implemented as software, firmware, hardware or a combination thereof. For example, aspects of the present application may be at least partially embodied in a device, a system including more than one device, a method, a computer program product, etc.
[0362] In hardware implementation, the division of tasks between functional units mentioned in the above description does not necessarily correspond to the division of physical units; on the contrary, one physical component may have multiple functions, and one task may be performed by several physical components in collaboration.
[0363] Some or all components may be implemented as software executed by a digital signal processor or microprocessor or as hardware or an application specific integrated circuit. Such software may be distributed on a computer readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium).
[0364] As is well known to those skilled in the art, the term "computer storage media" includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0365] Further, as is well known to those skilled in the art, communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
Claims
1. A method for decoding an audio signal, the method comprising: receiving a bitstream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks; determining, based on device information of the output device, information for identifying, for each frame, a subset of one or more blocks of the plurality of blocks to be decoded according to the signaling data; as well as A subset of the identified one or more blocks of the plurality of blocks is decoded. 2 . The method of claim 1 , wherein the information identifying the subset of one or more of the plurality of blocks comprises a matrix associating each of a plurality of output devices with the subset of one or more blocks.
3. The method of claim 2, wherein: The one or more frames are required for decoding of the bitstream for the corresponding associated output device.
4. The method according to any one of claims 1 to 3, wherein: The output device can include at least one of a wireless device, a mobile device, a tablet computer, a mono speaker, a multi-channel speaker, and / or a sound reproduction system.
5. The method according to any one of claims 1 to 4, wherein: The one or more blocks partially and / or completely include at least one audio signal for playback at the output device based on the device information.
6. The method of any one of claims 1-5, wherein the output device is a first output device, further comprising applying a joint coding technique between one or more signals of the bitstream to at least a second output device.
7. The method according to any one of claims 1 to 6, wherein: The identity of each output device and / or decoder is defined during the system initialization phase.
8. The method according to any one of claims 1 to 7, wherein: The signalling data is determined by metadata of the bitstream.
9. A method for generating a coded bit stream of an audio signal, the method comprising: generating a bitstream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks; For each frame, determining, based on device information of an output device, signaling data for identifying a subset of one or more blocks of the plurality of blocks to be decoded; as well as The bitstream is encoded with signaling data including the subset of one or more blocks of the plurality of blocks to be decoded upon decoding.
10. The method of claim 9, wherein the information identifying the one or more of the plurality of blocks to be decoded upon decoding comprises a matrix associating each of a plurality of output devices with the subset of at least one audio signal.
11. The method of claim 10, wherein: The one or more frames are required for decoding of the bitstream for the corresponding associated output device.
12. The method of any one of claims 9-11, wherein the output device can include at least one of a wireless device, a mobile device, a tablet computer, a mono speaker, a multi-channel speaker, and / or a sound reproduction system.
13. The method according to any one of claims 9 to 12, wherein: The one or more blocks partially and / or completely include at least one audio signal for playback at the output device based on the device information.
14. An apparatus configured to perform the method according to any one of claims 1 to 13.
15. A non-transitory computer-readable storage medium comprising a sequence of instructions which, when executed, causes one or more devices to perform the method of any one of claims 1-13.
Citation Information
Patent Citations
Selective forward error correction for spatial audio codecs
US11289103B2