Method, apparatus and medium for encoding and decoding audio bitstream with parameter-flexible rendering configuration data
By jointly encoding multiple audio signals and generating encoded bitstreams, the delay and quality problems in wireless streaming are solved, and the low-latency interchange and efficient encoding and decoding of audio signals between devices are realized.
Patent Information
- Application Number
- CN202380070780.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2023-09-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively solve the delay and quality problems when wirelessly streaming audio signals, resulting in low latency interchange and encoding and decoding efficiency of audio signals between devices.
By generating encoded bitstreams from multiple audio signals, information from the playback device is received for each audio signal, and a joint encoding tool is applied to encode the related audio signals to generate independent blocks to achieve transmission in a low-latency format.
Low-latency interchange and efficient encoding and decoding of audio signals between devices are realized, improving the quality and efficiency of wireless streaming.
Smart Images

Figure CN119998871A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Application Serial No. 63 / 378,497 filed on October 5, 2022 and U.S. Provisional Application Serial No. 63 / 578,534 filed on August 24, 2023, both of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates generally to audio signal processing, and more particularly to audio source encoding and decoding for low-latency interchange of audio signals between devices for immersive audio programs. Background Art
[0004] Streaming of audio is commonplace in today's world. It is becoming more demanding as users' expectations of quality continue to rise, and the user's setup is becoming more complex as the number of speakers and the types of speakers increase. Streaming is usually done at least partially over a wireless link, which requires that the wireless link be of good quality, which, as many people have experienced, is not always the case.
[0005] Therefore, there is a need to define an interchange format for use cases where streaming occurs from the cloud / server in a certain format and is subsequently transcoded on the device into a low latency format more suitable for distribution over a wireless (or in some cases wired) link. Exemplary use cases are home connectivity and phone to car connectivity, but such a format could be beneficial in any scenario where it is desirable to distribute audio signals from a single device to one or more connected devices in a low latency manner.
[0006] In addition to sending and streaming audio information wirelessly, other types of information can be incorporated into the stream. This other type of information is then also affected by the quality of the wireless link and may have similar drawbacks as audio.
[0007] Therefore, it would be advantageous to overcome the problems associated with wireless streaming for different types of streamed audio and combinations of other types of information or signals. Summary of the invention
[0008] It is an object of the present disclosure to at least partially overcome the above problems by utilizing wireless streaming of audio as well as other types of information.
[0009] According to a first aspect of the present disclosure, a method for generating an encoded bitstream from an audio program including multiple audio signals, the method comprising: receiving, for each of the multiple audio signals, information indicating a playback device associated with the corresponding audio signal; receiving, for each playback device, information indicating at least one of a delay, a gain, and an equalization curve associated with the corresponding playback device; determining a group of two or more related audio signals from the multiple audio signals; applying one or more joint encoding tools to the two or more related audio signals in the group to obtain a jointly encoded audio signal; and combining the jointly encoded audio signal, an indication of the playback device associated with the jointly encoded audio signal, and information indicating at least one of a delay, a gain, and an equalization curve associated with the corresponding playback device associated with the jointly encoded audio signal into independent blocks of the encoded bitstream.
[0010] According to a second aspect of the present disclosure, a method for decoding one or more audio signals associated with a playback device from a frame of an encoded bit stream, wherein the frame includes one or more independent encoded data blocks, and the method includes: identifying independent encoded data blocks corresponding to the one or more audio signals associated with the playback device from the encoded bit stream; extracting the identified independent encoded data blocks from the encoded bit stream; determining that the extracted independent encoded data blocks include two or more jointly encoded audio signals; applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device; determining at least one of a delay, a gain, and an equalization curve associated with the playback device from the extracted independent encoded data blocks; and applying the delay, the gain, and / or the equalization curve associated with the playback device to the one or more audio signals associated with the playback device.
[0011] According to a third aspect of the present disclosure, a device is provided, wherein the device is configured to execute the method according to the first aspect and / or the second aspect.
[0012] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, which includes a series of instructions, which, when executed, cause one or more devices to perform a method according to any one of the methods of the first aspect and / or the second aspect.
[0013] Further examples of the disclosure are defined in the dependent claims.
[0014] In some examples, the delay, gain, and / or equalization curve associated with the respective playback device depends on the position of the respective playback device relative to the position of the listener and / or the position of the respective playback device relative to the position of other playback devices. In some examples, the delay, gain, and / or equalization curve are dynamically variable.
[0015] In some examples, delays, gains, and / or equalization curves are adjusted in response to changes in the listener's position, changes in the position of the playback device, and / or changes in the position of one or more of the other playback devices.
[0016] In some examples is a method that includes determining an audio signal from a plurality of audio signals that is not part of a group of two or more related audio signals.
[0017] In some examples is a method that includes applying, for an audio signal that is not part of a set of two or more related audio signals, a delay, gain, and / or equalization curve associated with a playback device associated with the audio signal.
[0018] In some examples is a method that includes independently encoding an audio signal that is not part of a group of two or more related audio signals, and combining the independently encoded audio signal and an indication of a playback device associated with the independently encoded audio signal into separate independently decodable subsets of an encoded bitstream.
[0019] In this disclosure, a frame refers to a time slice of the whole of all signals. A block stream refers to a collection of signals over the duration of a session. A block refers to one frame of a block stream. For digital audio with a given sampling frequency, the frame size is equal to the number of audio samples in a frame of any audio signal. The frame size is typically kept constant over the duration of a session.
[0020] The wake-up word may include a single word, or a phrase including two or more words in a fixed order.
[0021] Throughout this disclosure, including in the claims, the expression "system" is used in a broad sense to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Examples of the present disclosure will be described in more detail with reference to the accompanying drawings. In the following drawings, the same reference numerals are used to refer to the same elements. Although the following drawings depict various examples, one or more implementations are not limited to the examples depicted in the drawings.
[0023] Figure 1 Diagram showing an example of low-latency transcoding for a home connection
[0024] Figure 2 Illustration of an example of connected car audio streaming
[0025] Figure 3 Illustrated example of enhanced TV audio streaming using wireless devices
[0026] Figure 4 Illustration of an example of audio streaming using a simple wireless speaker
[0027] Figure 5 Illustration of an example of metadata mapping from bitstream elements to devices
[0028] Figure 6 Example showing how to deploy simple flexible rendering for audio streaming
[0029] Figure 7 Illustration of an example of mapping flexible rendering data to bitstream elements
[0030] Figure 8 Illustration of an example of echo-referenced signaling when playing audio and listening for commands with multiple devices
[0031] Fig. 9 An example illustrating how frames, blocks, and packets relate to each other
[0032] Fig.10 Illustrated example of further codec information in the current MPEG-4 structure
[0033] Fig.11 An example of further codec information in the current MPEG-4 structure is shown.
[0034] Fig.12 Illustration of an example of integrating listening capabilities and speech recognition
[0035] Fig.13 An example of a frame including a plurality of blocks is shown
[0036] Fig.14 An example of a bitstream including multiple blocks is shown
[0037] Fig.15 Illustration of an example of blocks with different priorities
[0038] Fig.16 An example of a bitstream including multiple blocks with different priorities is illustrated
[0039] Fig.17 Illustration of an example of frames with different priorities
[0040] Fig.18 illustrates an example of a bitstream including frames with different priorities DETAILED DESCRIPTION
[0041] The principle of the present invention will now be described with reference to the various examples illustrated in the accompanying drawings. It should be understood that the description of these examples is only to enable those skilled in the art to better understand and further implement the present invention; it is not intended to limit the scope of the present invention in any way.
[0042] exist Figure 1 In the immersive audio stream, the immersive audio stream is streamed from the cloud or server 10 and decoded on the TV or hub device 20. The immersive audio stream can be encoded in any existing format, including, for example, Dolby Digital Plus, AC-4, etc. The output is then transcoded into a low-latency interchange format for further transmission to a connected device 30, which is preferably connected via a local wireless connection (e.g., a WiFi soft access point) or a Bluetooth connection. Low latency typically depends on various factors, such as, for example, frame size, sampling rate, hardware and / or software computing resources, etc., and low latency is typically less than 40ms, 20ms, or 10ms.
[0043] In another typical use case, the phone fetches the immersive audio stream from the cloud or server, transcodes it into an interchangeable format, and then transmits it to the connected car. Figure 2 In the illustrated example of , a mobile device (e.g., a phone or tablet) 20 connects to a server 10 and receives an immersive audio stream, performs low-latency transcoding to a low-latency interchange format, and transmits the transcoded signal to an immersive audio playback-enabled car 30. An example of an immersive audio stream is a stream that includes audio in the Dolby Atmos format, and an example of an immersive audio playback-enabled car 30 is a car configured to play back the Dolby Atmos immersive format.
[0044] Typically, the interchange format preferably has low latency, low encoding and decoding complexity, the ability to scale to high quality, and reasonable coding efficiency. The format preferably also supports configurable latency so that latency can be weighed against efficiency and error resilience to operate under varying connection conditions.
[0045] A hub or enhanced display with wireless speakers with or without listening capabilities
[0046] Figure 3 , an example of a hub 20 driving a set of wireless speakers 30 is illustrated, or a display (e.g., a television or TV) 20 that may have built-in speakers 30 is augmented with several wireless speakers 30. The augmentation indicates that in the example where the display 20 includes speakers 30, the display 20 is also part of the audio reproduction. The wireless speakers / devices 30 may be receiving the same complete signal (broadcast mode, in Figure 3 ) or individual streams customized for specific devices (Unicast Multipoint mode, in Figure 3 ).
[0047] Each speaker may include multiple drivers that cover different frequency ranges of the same channel, or correspond to different channels in a canonical mapping. For example, a speaker may have two drivers, one of which may be an upward-firing driver that outputs a signal corresponding to a height channel to simulate a height speaker. Wireless devices may have listening capabilities (e.g., a "smart speaker"), and therefore may require echo management, which may require the speaker to receive one or more echo references. The echo reference may be local (e.g., to the same speaker / device), or may instead represent relevant signals from other nearby speakers / devices.
[0048] The speakers may be placed at arbitrary locations, in which case so-called "flexible rendering" may be performed, whereby the rendering takes into account the actual locations of the speakers (e.g. in contrast to the specification assumptions, whereby it is assumed that the loudspeakers are located at fixed predefined locations). Flexible rendering may be done in the hub or TV, whereby the rendered signals are then sent to the speakers / devices in a broadcast mode, whereby each rendered signal is sent to each device and each respective device extracts and outputs the appropriate signal, or as separate streams sent to separate devices. Alternatively, flexible rendering may be done locally on each device, whereby each device receives a representation of the complete immersive program, e.g. a 7.1.4 channel based immersive representation, and renders from that representation an output signal suitable for the respective device.
[0049] Wireless devices can connect to a hub or TV through a soft access point provided by the device or through a local access point in the home. This may impose different requirements on bit rate and latency.
[0050] Mobile Device Projection in Automotive Applications
[0051] In another use case, a mobile device (e.g., a phone or tablet) 20 retrieves an immersive audio stream from a cloud or server 10, transcodes the immersive audio stream into an interchange format, and then transmits the transcoded stream to a connected car 30. Figure 2In the example of , a Dolby Atmos-enabled phone 20 connects to a server 10 and receives a Dolby Atmos stream for low-latency transcoding and transmission to a Dolby Atmos-enabled car 3. In this use case, a higher bitrate can be obtained than the bitstream in the living room use case, and the wireless channel may have different characteristics. The reason for the higher bitrate in the car use case may be due to the less noisy wireless environment, because the car acts as a kind of Faraday cage and protects the interior car environment from external wireless interference compared to the living room where all wireless devices, such as from neighbors, other rooms, etc., add noise to the wireless environment. At the same time, there are typically few wireless devices in the car competing for the available wireless bandwidth compared to the living room use case where all wireless devices in a typical home (including living room wireless devices) compete for the available bandwidth.
[0052] For this use case, it is envisioned that the signal to be transcoded by the mobile device and transmitted to the car may be a channel-based immersive representation, an object-based representation, a scene-based representation (e.g., as well as a high-fidelity stereo representation), or even a combination of different representations. For this example, different rendering architectures and broadcast vs. multipoint may not be relevant, since typically the complete presentation will be transmitted from the mobile device to a single endpoint (e.g., the car).
[0053] Description of the interchange format
[0054] The Immersive Interchange Format is built on a modified discrete cosine transform (MDCT) with perceptually excited quantization and coding. It has configurable latency, for example, to support different transform sizes at a given sampling rate. At sampling rates of 48kHz and 44.1kHz, exemplary frame sizes are 128, 256, 512, 1024 and 120, 240, 480, 960 and 192, 384, 768 samples.
[0055] The format may support mono, stereo, 5.1, and other channel configurations, including immersive channel configurations (e.g., including but not limited to 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, and 22.2, or any other channel configuration consisting of unique channels as specified in Table-2 of ISO / IEC 23091-3:2018). The format may also support object-based audio and scene-based representations, such as high-fidelity stereo (e.g., first order or higher order). The format may also use a signaling scheme suitable for integration with existing formats (e.g., such as those described in ISO / IEC 14496-3 (which may also be referred to as the MPEG-4 audio standard), ISO 14496-1 (which may also be referred to as the MPEG-4 systems standard), ISO 14496-12 (which may also be referred to as the ISO base media file format standard), and / or ISO 14496-14 (which may also be referred to as the MP4 file format standard)).
[0056] In addition, the system can support the ability to skip portions of syntax elements and decode only the relevant portions for a given speaker, support metadata-controlled flexible rendering aspects such as delay alignment, level adjustment and equalization, support the use of smart speakers with listening capabilities (including scenarios where a set of signals are sent to allow each driver / speaker in the smart speaker to be fed independently), and relatedly support echo management through signaling of echo references, support quantization and encoding in the MDCT domain with low overlap windows as well as 50% overlap windows, and / or support filtering along the frequency axis in the MDCT domain for time shaping of quantization noise, such as TNS (temporal noise shaping). Therefore, in various examples, the overlapping windows are symmetric or asymmetric.
[0057] In some examples, the disclosed interchange format can provide the following: improved joint-channel coding to exploit increased correlation between signals in flexible rendering use cases; improved coding efficiency through scale factors shared across channel elements; inclusion of high frequency reconstruction and noise addition techniques in the MDCT domain; various improvements to traditional coding structures to allow for increased efficiency and make them suitable for the use case at hand.
[0058] In some examples of controlling individual playback devices in broadcast mode, skippable chunks and metadata may be used. In a broadcast mode setting, each wireless device will need to play the relevant portion of the complete audio for a specific driver / channel corresponding to a specific device. This means that the device needs to know which portions of the stream are relevant to that device and extract those portions from the complete audio stream being broadcast. In order to achieve a low complexity decoding operation, it is preferred that the stream is constructed in such a way that the decoder of a specific device can efficiently skip over elements that are not relevant to decoding to get to the elements that are relevant to the device and a given driver on the device.
[0059] However, it should be noted that there may still be benefits (from a compression perspective) in allowing joint encoding between signals destined for different devices (eg, in the simplest scenario, the left and right channels of a stereo presentation).
[0060] Thus, the format may: implement "skippable blocks" in the bitstream to enable efficient decoding of portions that are only relevant to a specific device, including enabling flexible mapping of one or more skippable blocks to metadata for a specific device while retaining the ability to apply joint coding techniques between signals corresponding to different speakers / devices.
[0061] exist Figure 4 In FIG. 1 , an example of an arbitrary setup is shown. There are three connected wireless devices 31, 32, 33, two of which are mono speakers 32, 33 (meaning one audio channel), and one speaker 31 is a more advanced speaker 31 having three different drivers and thus operating on three different signals. In this example, the first speaker 31 operates on 3 separate signals, while the second speaker 32 and the third speaker 33 operate on a stereo representation (e.g. left and right). Therefore, it may be beneficial to jointly encode the signals of the stereo representation.
[0062] Given the above scenario, the format specifies a bitstream containing skippable blocks so that the first device can extract only the relevant part of the stream to decode the signal for that speaker, and for a stereo pair, the decoder complexity and efficiency can be traded off by constructing a signal where each speaker needs to decode two signals to output a single signal, with the benefit that joint encoding can be performed.
[0063] The format specifies a metadata format to enable a generic and flexible representation of mapping a particular block of a plurality of skippable blocks to one or more devices. Figure 5 , where the mapping can be represented as a matrix associating each device 31, 32, 33 with one or more bitstream elements, so that the decoder of a given device will know what bitstream elements to output and decode.
[0064] For example, in Figure 5 In the example of , the first block or skip block Blk1 includes 3 mono elements (one for each driver of device 1 31), while the second block or skip block Blk2 contains channel pair elements, which may include jointly encoded versions of the signals to be output by devices 2 32 and 3 33. Device 1 31 extracts the mapping metadata and determines that the signal it needs is in skip block 1Blk1. Therefore, it extracts skip block 1Blk1, decodes the three mono elements therein, and provides the three mono elements to drivers 1, 2a, and 2b, respectively.
[0065] In addition, device 1 31 ignores skip block 2Blk2. Similarly, device 2 32 extracts mapping metadata and determines that the signal it needs is in skip block 2Blk2. Therefore, device 2 32 skips skip block 1Blk1 and extracts skip block 2Blk2. Device 2 32 decodes the channel pair element and provides the left channel output of the CPE to its driver. Similarly, device 3 33 extracts mapping metadata and determines that the signal it needs is in skip block 2Blk2, so device 3 33 skips skip block 1Blk1 and also extracts skip block 2Blk2. Device 3 33 decodes the channel pair element and provides the right channel output of the CPE to its driver.
[0066] In some examples, device 2 32 may determine that it only needs a subset of the signal from skip block 2Blk2. In such examples, device 2 32 may perform only a subset of the operations required to fully decode the signal in skip block 2Blk2, if possible. Specifically, Figure 5 In the example of , device 2 32 may only perform those processing operations required to extract the left channel of the CPE, thereby enabling reduced computational complexity. Similarly, device 3 33 may only perform those processing operations required to extract the right channel of the CPE.
[0067] For example, in Figure 5 In the case of , there may be times when the CPE is encoded using joint channel coding and other times when it is encoded using independent channel coding. During the time when the CPE is encoded using independent channel coding, device 2 32 may extract only the first (e.g., left) channel of the CPE, while device 3 33 may extract only the second (e.g., right) channel of the CPE.
[0068] In another example, the channels of the CPE may be encoded using joint channel coding, in which case the two middle channels of the CPE must be extracted by devices 2 32 and 3 33. However, by performing only those operations required to extract the left channel from the middle channel, device 2 32 may still be able to operate with reduced computational complexity. Similarly, by performing only those operations required to extract the right channel from the middle decoded channel, device 3 33 may still be able to operate with reduced computational complexity.
[0069] Depending on the specifics of joint channel encoding, other optimizations may be possible. The identity of each of the different devices / decoders may be defined during a system initialization or setup phase. Such setup is common and typically involves measuring the room acoustics, the distance of the speakers to the sweet spot, etc.
[0070] Flexible rendering (pre-encoding and post-decoding) distribution and delayed application of post-decoder
[0071] For the use case of applying flexible rendering in the hub / TV before sending the relevant signal to a specific device / speaker, rendering can create a signal that is more difficult to encode from an encoding perspective, for example when considering joint encoding of such a signal. One reason is that flexible rendering can apply different delays, equalization, and / or gain adjustments to different devices (e.g., depending on the placement of the speaker relative to other speakers and the listener). It is also possible to preset information (e.g., gain and delay) at initial setup and flexibly render only the equalization. In other examples, other variations of preset information and flexible rendering information are also possible. Note that in this document, the term "gain" should be interpreted as meaning any level adjustment (e.g., attenuation, amplification, or pass-through), and not limited to only certain level adjustments (e.g., amplification).
[0072] exist Figure 6 In the example of , the right channel 33 and left channel 32 speakers are given different delays, which reflects the different placement relative to the listener (e.g., because the distances of the speakers 32, 33 from the listener may not be equal, different delays can be applied to the signals output from the speakers 32, 33 so that coherent sounds from different speakers 32, 33 arrive at the listener at the same time). Introducing such different delays to coherent signals intended for playback on different speakers 31, 32, 33 makes joint encoding of such signals challenging.
[0073] To address this challenge, aspects of a flexible rendering process can be parameterized and applied in the endpoint device after the signal is decoded.
[0074] exist Figure 7In the example of , the delay and gain values of each device are parameterized and included in the encoded signal sent to the respective loudspeaker 31, 32, 33. The respective signals may be decoded by the respective devices 31, 32, 33, which may then introduce the parameterized gain and delay values to the respective decoded signals.
[0075] In the example (such as Figure 7 In the example shown in ), in the case where the encoded signals for different devices 31, 32, 33 are sent in separable blocks (e.g., skippable blocks), parameters (e.g., delays and gains) may also be sent in separable blocks, so that the devices 31, 32, 33 can extract only a subset of the parameters required by the devices 31, 32, 33 and ignore (and skip) those parameters not required by the devices 31, 32, 33. In this case, mapping metadata indicating which parameters are included in which blocks may be provided to each device.
[0076] Also note that although Figure 7 Only delay and gain parameters are indicated, but other parameters such as equalization parameters may also be included. For example, the equalization parameters may include a plurality of gains to be applied to different frequency regions, an indication of a predetermined equalization curve to be applied by the playback device 31, 32, 33, one or more sets of infinite impulse response (IIR) or finite impulse response (FIR) filter coefficients, a set of biquad filter coefficients, parameters specifying parametric equalizer characteristics, and other parameters known to those skilled in the art for specifying equalization.
[0077] Furthermore, parameterization of aspects of flexible rendering need not be static, and may be dynamic (e.g., where a listener moves during playback of an audio program). Thus, it may be preferable to allow parameters to change dynamically. Where one or more parameters change during an audio program, the device may interpolate between previous delay and / or gain parameters and updated delay and / or gain parameters to provide a smooth transition. This may be particularly useful where the system dynamically tracks the position of the listener and correspondingly updates the sweet spot for dynamic rendering.
[0078] As has also been noted and will be described later, when flexible rendering is applied the level of correlation between channels may increase and this may be exploited by a more flexible joint coding.
[0079] Echo Reference Coding and Signaling
[0080] For Figure 8 In the use case outlined in , there are multiple devices / speakers 30 operating together in broadcast mode (in Figure 8 ) or in unicast multipoint mode (on the left side of Figure 8The need for echo management arises when the device 30 receives signals and a microphone 40 is present on the device 30 to enable "listening" capabilities.
[0081] When performing echo management with multiple speakers / devices, it may be beneficial to use echo references that are not just from the local speaker device. As an example, a device may be placed close to another device, and thus the signal from the nearby device will affect the echo management of the device with the active microphone. In the broadcast mode use case, each device receives the signal for all devices. When a device has signals for other devices, it may be beneficial to use those signals as echo references. To do this, it is necessary to signal to a specific device which signals can be used as echo references for which other devices.
[0082] In one example, this can be done by providing metadata that maps not only the channels / signals to be played by a particular speaker / device (from the entire set), but also the channels / signals to be used as echo references (from the entire set) for the particular speaker / device. Such metadata or signaling can be dynamic, allowing the indication of a preferred echo reference to change over time.
[0083] For use cases where each device / speaker receives only the specific signal it is to play, it may be necessary to transmit an additional signal (e.g., an echo reference signal) to each device / speaker in order to provide an appropriate echo reference. Furthermore, for this purpose, it may be necessary to provide device-specific signaling so that each device can select the appropriate signal for playback and the appropriate signal for echo management.
[0084] Since the signal used for echo management is only used by the device to perform echo management and is not played back by the device to a listener, the echo reference signal may be encoded or represented in a different manner than the signal intended for playback by the device. In particular, because successful echo management can be achieved with a signal encoded at a lower rate than would normally be used for playback to a listener, additional compression tools (such as parametric representation of the signal) that may not be suitable for playback to a listener at all but captures the necessary characteristics of the audio signal can provide good echo management at significantly reduced transmission costs.
[0085] Various grammatical elements
[0086] In some examples, blocks are used to optimize the transmission of audio. In the described format, each frame can be split into multiple blocks, as described above with respect to skippable blocks. A block can be identified by the frame number to which it belongs, a block ID that can be used to associate consecutive blocks with the same ID from different frames with a block stream, and a retransmission priority. An example of the above is a stream with multiple frames N-2, N-1, and N, frame N including multiple blocks identified as ID1, ID2, ID3, and the multiple blocks in Fig.13 Medium picture. Fig.14 An example of this stream is illustrated in , but in a bitstream format. In some examples, for an audio signal of an immersive audio program, a block ID of a block may indicate which group of signals of the entire immersive audio program is carried by the block.
[0087] A use case for the described format is to reliably transmit audio with low latency over wireless networks such as Wifi. For example, Wifi uses a packet-based network protocol. Packet sizes are usually limited. A typical maximum packet size in IP networks is 1500 bytes. The block-based architecture of the stream allows for flexibility in assembling packets for transmission. For example, packets with smaller frames can be filled with retransmitted blocks from other frames. Large frames can be split on block boundaries before packetization to reduce dependencies between packets at the network protocol layer.
[0088] Fig. 9 The relationship between frames, blocks and packets is shown. A frame carries audio data, preferably all audio data, representing a continuous segment of an audio signal having a start time, an end time and a duration that is the difference between the end time and the start time. The continuous segment may include a time period according to Section 4.5.2.1.1 of Subpart 4 of ISO / IEC 14496-3. Section 4.5.2.1.1 describes the contents of raw_data_block(). A frame may also carry a redundant representation of the segment, for example encoded at a lower data rate. After encoding, the frame may be segmented into multiple blocks. The blocks may be combined into packets for transmission over a packet-based network. Blocks from different frames may be combined in a single packet and / or may be sent out of order.
[0089] In the examples, blocks are used to address individual devices. Data packets are received by individual devices or groups of related devices. The concept of skippable blocks can be used to address individual devices or groups of related devices. Even if the network operates in broadcast mode when sending packets to different devices, the processing of the audio (e.g., decoding, rendering, etc.) can be simplified to blocks addressed to that device. All other blocks, even if received within the same packet, can simply be skipped. In some examples, blocks can be introduced in the correct order based on the decoding or presentation time of the blocks. If the same block with a higher priority has also been received, the retransmitted block with a lower priority can be removed. The block stream can then be fed to the decoder.
[0090] In some examples, the configuration of the stream and the device is sent out-of-band. The codec allows a connection to be established where the audio stream is transmitted at a relatively high rate, but with low latency. The configuration of such a connection can remain stable for the duration of such a connection. In this case, instead of making the configuration part of the audio stream, the configuration can be transmitted out-of-band. For such out-of-band transmission, even different networks or network protocols can be used. For example, the audio stream can use the User Datagram Profile (UDP) for low-latency transmission, while the configuration can use the Transmission Control Protocol (TCP) to ensure reliable transmission of the configuration.
[0091] Enabling codecs in the context of MPEG-4 Audio
[0092] One specific application of the present technique is for use within MPEG-4 Audio. In MPEG-4 Audio, different AOTs (Audio Object Types) are defined for different codec technologies. For the formats described in this document, new AOTs may be defined, which allow specific signaling and data for that format. Furthermore, the configuration of a decoder in MPEG-4 is done in a DecoderSpecificInfo() payload, which in turn carries an AudioSpecificConfig() payload. In the latter, certain generic signaling that is agnostic to the specific format, such as sampling rate and channel configuration, is defined, as well as specific information for a specific AOT. This may make sense for traditional formats where the entire stream is decoded by a single device. However, in broadcast mode, where a single stream is transmitted to several decoders, each respective decoder only decoding part of the stream, pre-channel configuration signaling (as a means of setting the output capabilities of the device) may be suboptimal.
[0093] Fig.10The traditional MPEG-4 high-level structure is illustrated (in black), with modifications in grey. To support broadcast use cases, codecSepcificConfig() is defined (where "codec" may be a generic placeholder name), where the signalling is redefined for a specific use case such that a specific channel element may be mapped to a specific device, and other relevant static parameters may be included. An MPEG-4 element channelConfig with a value of "0" is defined as the channelConfig defined in codecSpecificConfig. Thus, this value may be used for signalling that enables modification of the channel configuration inside a codec specific configuration.
[0094] Further, in the spirit of MPEG-4, codecSpecificConfig() is assumed to be decodable, and the raw payload is specified for the specific decoder at hand. However, the format ensures that dynamic metadata is part of the raw payload, and that length information is available for all raw payloads, so that the decoder can easily skip elements that are not relevant to a specific device.
[0095] exist Fig.11 A portion of an example raw_data_block as defined in MPEG-4 is given on the right. A raw data block contains channel elements (single channel elements (SCE) or channel pair elements (CPE)) in a given order. However, in the conventional MPEG-4 audio syntax, a decoder wishing to skip parts of these channel elements that may not be relevant for the output device at hand would have to parse (and to some extent also decode) all channel elements in order to be able to extract the relevant parts. Fig.11 In the new raw_data_block illustrated on the left side of , the content will consist of skippable blocks so that the decoder can skip irrelevant parts and only decode the channel elements indicated by the metadata as being relevant to the device at hand. In the example, the skippable block includes the raw_data_block and the relevant information.
[0096] It is also possible to use blocks for retransmission. Each frame can be split into multiple blocks, as described above in the section on skippable blocks. A block can be identified by the frame number to which it belongs, a block ID that can be used to associate consecutive blocks with the same ID from different frames with a block stream, and a retransmission priority, such as Fig.15 and Fig.16 For example, a high retransmission priority (e.g., Fig.15 and Fig.16 indicates that a block should be given priority at the receiver over another block with the same block ID and frame counter but with a lower retransmission priority (e.g., Fig.15 and Fig.16Note that, as in Fig.15 and Fig.16 In some examples, the priority may decrease as the priority index increases (e.g., priority 1 may be a lower priority than priority 0), while in other examples, the priority may increase as the priority index increases (e.g., priority 0 is a lower priority than priority 1). Still other examples will be apparent to the skilled person.
[0097] The syntax may support retransmission of audio elements. For retransmissions, various quality levels may be supported. Therefore, retransmitted blocks may carry a 'priority' flag to indicate which blocks with the same block ID should be given priority for the decoder, since received blocks with the same frame counter and block ID are redundant and therefore mutually exclusive for the decoder, as Fig.17 and Fig.18 As shown in the figure.
[0098] The retransmission of the block may be done at a reduced data rate. Such a reduced data rate may be achieved by reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, reducing the channel count of the audio signal (e.g., as described in U.S. Pat. No. 11,289,103, which is incorporated herein by reference in its entirety), or any combination thereof. In order for the decoder to select the audio block that provides the best possible quality, the block that provides the highest quality signal may have the highest decoding priority, the block with the second highest quality signal may have the second highest priority, and so on.
[0099] These blocks can also be retransmitted at the same quality level. In this case, the priority can reflect the latency of the retransmitted blocks.
[0100] Tools for improving core coding efficiency
[0101] Further examples may allow sharing of MDCT quantization scale factors across channel elements. In some cases, MDCT scale factors may be shared across certain channels to reduce the associated edge bit rate. In addition, scale factor sharing may be extended to different channel elements, for example for 7.1.4 input. The use of shared scale factors may be indicated by 2 signaling bits that allow 3 active configurations. One possible configuration is to share scale factors in the left horizontal channel, the left upper channel, the right horizontal channel, and the right upper channel. Specific syntax may be aligned with the skippable block concept to ensure that scale factors are only shared within a skipped block.
[0102] It is also possible to jointly encode more than two channels. Figure 4 , Figure 5 , Figure 6 and Figure 7In the example illustrated in , block Blk 1 carries three SCEs, one for each driver in the smart speaker considered in this scenario. In this case, it may be beneficial to be able to jointly code not only two channels (as provided by CPE, which can be extended to include stereo prediction, for example, as described in ISO / IEC 23003-3, the MDCT-based complex prediction stereo tool, called MPEG-D USAC), but also more than two channels, for example, as described in the SAP (Stereo Audio Processing) tool introduced in ETSI TS103 190.
[0103] Similarly, it should be noted that in the case of flexible rendering, there may be a large amount of correlation between signals across devices. For such signals, it may be beneficial to build skip blocks that cover multiple devices and allow joint channel coding to be applied on more than two channels using the tools outlined above.
[0104] For certain transform lengths, another joint channel coding tool that can be used is channel coupling, where composite channel and scale factor information is transmitted for both intermediate and high frequencies. This can provide bit rate reduction for playback within a good quality range. For example, it may be beneficial to use the channel coupling tool for frames with a transform length of 256, which corresponds to a frame length of about 256 samples for low-latency coding.
[0105] The present disclosure will also allow for efficient encoding of bandwidth limited signals. In scenarios where separate signals are sent to feed different drivers of a smart speaker, some of these signals may be band limited, such as, for example, in the case of a three-way driver configuration with a woofer, midrange speaker, and tweeter. Therefore, it is desirable to efficiently encode such band limited signals, which can translate into specially tuned psychoacoustic models and bit allocation strategies as well as potential modifications to the syntax to handle such scenarios with improved coding efficiency and / or reduced computational complexity (e.g., enabling the use of band limited IMDCT for the woofer feed).
[0106] In some examples, high frequency reconstruction and noise addition may be combined in the MDCT domain. Modern audio codecs are typically designed to support parametric coding techniques, and similarly, it is contemplated here to include, for example, a high frequency reconstruction method in the MDCT domain to keep latency low, and to perform a noise addition scheme in the MDCT domain to manage the tone-to-noise ratio in the reconstructed high frequency band. Such parametric coding techniques may be particularly useful for lower operating points, and in particular in scenarios where retransmission occurs as part of an FEC (forward error correction) scheme, where a main signal is typically transmitted, and then the same signal is retransmitted in a delayed manner at a lower bit rate.
[0107] These examples will further allow to reduce the peak complexity in the encoder. If there is no buffer model restriction of the constant bit rate transmission channel, this will be applicable. Under the constant bit rate with buffer mode, after quantization and counting, the final number of bits to be used for the frame to be encoded may exceed the limit of the allowed bits. At least one new, rougher quantization and bit counting step is required to meet the bit requirements. If this restriction is more relaxed, the encoder can keep the first quantization result and will be completed with a slightly higher instantaneous bit rate. By using the same bit repository control mechanism and the virtual buffer fullness updated according to the encoder following the buffer requirements, the encoder behavior similar to the constant bit rate of the buffer model encoder can still be achieved. This will save additional quantization and bit counting steps, and the resulting audio quality will be the same or better than the case of a constant bit rate with the disadvantage of a slightly increased overall bit rate.
[0108] Return channel concept
[0109] For the playback and listening use case, the smart speaker 30 may have one or more microphones 40 (e.g., a microphone or microphone array) configured to capture a monophonic or spatial sound field (e.g., in mono, stereo, A-format, B-format, or any other isotropic or anisotropic channel format), and a codec is required that can efficiently encode such a format in a low-latency manner so that the same codec as the codec used for the return channel is used for the broadcast / transmission to the smart speaker device 30, e.g. Fig.12 As shown in the figure.
[0110] It should also be noted that while wake word detection is performed on the smart speaker device, speech recognition is typically done in the cloud, where the wake word detection triggers the appropriate recorded speech segment.
[0111] In this context, it may be interesting to save system complexity overall by creating a speech analysis system on an intermediate format / representation, such as within a codec. For the simplest use case (where no human-to-human conversation is ongoing), a suitable representation specific to the speech recognition task can be defined, since the other person will not hear it. This could be band energies, Mel-Frequency Cepstral Coefficients (MFCC), etc., or a low-bitrate version of the MDCT-coded spectrum.
[0112] In the use case where a human-to-human conversation is ongoing and the wake-up word and the required speech recognition are intertwined, one would want to be able to extract the relevant portion of the stream for this. Such an extraction may require a layered coding structure of simple additional data sent in parallel with the main speech signal. We also envision a defined transcoding from an existing decoded MDCT spectrum to a related representation, and doing so simplifies the structure. Essentially, the decoder will decode and output human-audible audio until it is indicated that the decoder should also decode (in parallel) a speech recognition representation. This representation can be "peeled off" from the layered stream, simply as an alternative decoding of the same complete data, or simply as a decoding and output of an additional representation in the stream. In this use case, signaling is the enabler that indicates that the decoder on the receiving side should output a speech recognition related representation.
[0113] Examples of examples
[0114] In the following, seven sets of enumerated exemplary examples (EEE-A, EEE-B, EEE-C, EEE-D, EEE-E, EEE-F, and EEE-G) describe various aspects of the examples disclosed herein, which examples are not claims.
[0115] EEE-A1. A method for decoding an audio signal, the method comprising:
[0116] Receiving a bitstream comprising at least one frame, wherein each frame of the at least one frame comprises a plurality of blocks;
[0117] determining, based on device information of an output device, information for identifying a portion to be skipped when decoding one or more blocks of the plurality of blocks according to the signaling data; and
[0118] The bitstream is decoded while skipping the identified portion of the one or more blocks.
[0119] EEE-A2. A method as described in EEE-A1, wherein the information for identifying the portion to be skipped when decoding one or more blocks of the plurality of blocks includes a matrix associating each output device of a plurality of output devices with one or more bitstream elements.
[0120] EEE-A3. A method as described in EEE-A2, wherein the one or more bitstream elements are required for a corresponding associated output device to decode the bitstream.
[0121] EEE-A4. A method as described in any one of EEE-A1 to EEE-A3, wherein the output device may include at least one of a wireless device, a mobile device, a tablet computer, a mono speaker and / or a multi-channel speaker.
[0122] EEE-A5. A method as described in any one of EEE-A1 to EEE-A4, wherein the identified portion includes at least one block.
[0123] EEE-A6. A method as described in any one of EEE-A1 to EEE-A5, wherein the output device is a first output device, and the method further includes applying a joint encoding technology between one or more signals of the bit stream to a second output device and a third output device.
[0124] EEE-A7. A method as described in any of EEE-A1 to EEE-A6, wherein the identity of each output device and / or decoder is defined during a system initialization phase.
[0125] EEE-A8. A method as described in any one of EEE-A1 to EEE-A7, wherein the signaling data is determined based on metadata of the bit stream.
[0126] EEE-A9. An apparatus configured to perform a method as described in any one of EEE-A1 to EEE-A8.
[0127] EEE-A10. A non-transitory computer-readable storage medium comprising a series of instructions, which, when executed, cause one or more devices to perform a method as described in any one of EEE-A1 to EEE-A8.
[0128] EEE-B1. A method for generating an encoded bitstream from an audio program comprising a plurality of audio signals, the method comprising:
[0129] receiving, for each audio signal of the plurality of audio signals, information indicating a playback device associated with the respective audio signal;
[0130] receiving, for each playback device, information indicative of at least one of a delay, a gain, and an equalization curve associated with the respective playback device;
[0131] determining a set of two or more related audio signals from the plurality of audio signals;
[0132] applying one or more joint coding tools to the two or more related audio signals in the group to obtain a jointly encoded audio signal;
[0133] The jointly encoded audio signal, the indication of the playback device associated with the jointly encoded audio signal, and the indication of the delay and the gain associated with the respective playback device associated with the jointly encoded audio signal are combined into separate blocks of an encoded bitstream.
[0134] EEE-B2. The method of EEE-B1, wherein the delay, the gain, and / or the equalization curve associated with the respective playback device depends on a position of the respective playback device relative to a position of a listener.
[0135] EEE-B3. A method as described in EEE-B1 or EEE-B2, wherein the delay, the gain and / or the equalization curve associated with the corresponding playback device depends on the position of the corresponding playback device relative to the position of other playback devices.
[0136] EEE-B4. A method as described in any one of EEE-B1 to EEE-B3, wherein the delay, the gain and / or the equalization curve are dynamically variable.
[0137] EEE-B5. A method as described in EEE-B4, wherein the delay, the gain and / or the equalization curve are adjusted in response to a change in the position of the listener.
[0138] EEE-B6. A method as described in EEE-B4 or EEE-B5, wherein the delay, the gain and / or the equalization curve are adjusted in response to a change in the position of the playback device.
[0139] EEE-B7. A method as described in any one of EEE-B4 to EEE-B6, wherein the delay, the gain and / or the equalization curve are adjusted in response to a change in the position of one or more of the other playback devices.
[0140] EEE-B8. The method of any one of EEE-B1 to EEE-B7, further comprising determining an audio signal from the plurality of audio signals that is not part of the set of two or more related audio signals.
[0141] EEE-B9. The method as described in EEE-B8 further includes applying the delay, the gain and / or the equalization curve associated with the playback device associated with the audio signal to the audio signal that is not part of the group of two or more related audio signals.
[0142] EEE-B10. The method as described in EEE-B9 further includes independently encoding the audio signal that is not part of the group of two or more related audio signals, and combining the independently encoded audio signal and an indication of the playback device associated with the independently encoded audio signal into a separate independently decodable subset of the encoded bitstream.
[0143] EEE-B11. The method as described in EEE-B8 further includes independently encoding the audio signal that is not part of the group of two or more related audio signals, and combining the independently encoded audio signal, an indication of the playback device associated with the independently encoded audio signal, and an indication of the delay, the gain and / or the equalization curve associated with the playback device associated with the independently encoded audio signal into a separate independently decodable subset of the encoded bitstream.
[0144] EEE-B12. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises one or more independent encoded data blocks, the method comprising:
[0145] identifying, from the encoded bitstream, independent encoded data blocks corresponding to the one or more audio signals associated with the playback device;
[0146] extracting the identified independent coded data blocks from the coded bitstream;
[0147] determining that the extracted independent coded data block comprises two or more jointly coded audio signals;
[0148] applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device;
[0149] determining at least one of a delay, gain, and equalization curve associated with the playback device from the extracted independent encoded data blocks;
[0150] The delay, the gain, and / or the equalization curve associated with the playback device are applied to the one or more audio signals associated with the playback device.
[0151] EEE-B13. A method as described in EEE-B12, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to the position of the listener.
[0152] EEE-B14. A method as described in EEE-B12 or EEE-B13, wherein the determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices.
[0153] EEE-B15. A method as described in any one of EEE-B12 to EEE-B14, wherein the determined delay, gain and / or equalization curve of the playback device is dynamically variable.
[0154] EEE-B16. A method as described in EEE-B15, wherein, when the determined delay, gain and / or equalization curve associated with the playback device is different from the previously determined delay, gain and / or equalization curve associated with the playback device, the method further includes interpolating between the previously determined delay, gain and / or equalization curve associated with the playback device and the determined delay, gain and / or equalization curve associated with the playback device.
[0155] EEE-B17. A method as described in EEE-B16, wherein the determined delay, gain and / or equalization curve is different from the previously determined delay, gain and / or equalization curve due to a change in the position of the listener.
[0156] EEE-B18. A method as described in EEE-B16 or EEE-B17, wherein the determined delay, gain and / or equalization curve is different from the previously determined delay, gain and / or equalization curve due to a change in the position of the playback device.
[0157] EEE-B19. A method as described in any one of EEE-B16 to EEE-B18, wherein the determined delay, gain and / or equalization curve is different from the previously determined delay, gain or equalization curve due to a change in the position of one or more of the other playback devices.
[0158] EEE-B20. The method of any one of EEE-B12 to EEE-B19, wherein the frame of the encoded bitstream comprises two or more independent encoded data blocks, the method further comprising:
[0159] determining that one or more of the independent blocks contain audio signals not associated with the playback device; and
[0160] The one or more independent blocks containing audio signals not associated with the playback device are ignored.
[0161] EEE-B21. A method as described in any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools includes identifying a subset of the jointly encoded audio signal associated with the playback device, and reconstructing only the subset of the jointly encoded audio signal to obtain the one or more audio signals associated with the playback device.
[0162] EEE-B22. A method as described in any one of EEE-B12 to EEE-B20, wherein applying one or more joint decoding tools includes reconstructing each of the jointly encoded audio signals, identifying a subset of the reconstructed jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of the reconstructed jointly encoded audio signals associated with the playback device.
[0163] EEE-B23. An apparatus configured to perform the method as described in any one of EEE-B1 to EEE-B22.
[0164] EEE-B24. A non-transitory computer-readable storage medium comprising a series of instructions, which, when executed, cause one or more devices to perform a method as described in any one of EEE-B1 to EEE-B22.
[0165] EEE-C1. A method for generating a frame of an encoded bitstream of an audio program comprising a plurality of audio signals, wherein the frame comprises two or more independent encoded data blocks, the method comprising:
[0166] receiving, for one or more audio signals of the plurality of audio signals, information indicating a playback device associated with the one or more audio signals;
[0167] receiving, for the indicated playback device, information indicating one or more additional associated playback devices;
[0168] receiving one or more audio signals associated with the indicated one or more additional associated playback devices;
[0169] encoding the one or more audio signals associated with the playback device;
[0170] encoding the one or more audio signals associated with the indicated one or more additional associated playback devices;
[0171] combining the one or more encoded audio signals associated with the playback device and signaling information indicative of the one or more additional associated playback devices into a first independent block;
[0172] combining the one or more encoded audio signals associated with the one or more additional associated playback devices into one or more additional independent blocks; and
[0173] The first independent block and the one or more additional independent blocks are combined into a frame of the encoded bitstream.
[0174] EEE-C2. The method of EEE-C1, wherein the plurality of audio signals comprises one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices, the method further comprising:
[0175] encoding each of the one or more groups of audio signals not associated with the playback device or the one or more additional associated playback devices into a respective independent block; and
[0176] The respective independent blocks of each of the one or more groups are combined into a frame of the encoded bitstream.
[0177] EEE-C3. A method as described in EEE-C1 or EEE-C2, wherein the one or more audio signals associated with the indicated one or more additional associated playback devices are specifically intended to be used as echo references for the playback devices to perform echo management.
[0178] EEE-C4. The method of EEE-C3, wherein the one or more audio signals intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device. EEE-C4.
[0179] EEE-C5. A method as described in EEE-C3 or EEE-C4, wherein the one or more audio signals intended for use as echo references are encoded using a parametric coding tool.
[0180] EEE-C6. A method as described in EEE-C1 or EEE-C2, wherein one or more audio signals associated with one or more other playback devices are suitable for playback from the one or more other playback devices.
[0181] EEE-C7. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises two or more independent encoded data blocks, wherein the playback device comprises one or more microphones, the method comprising:
[0182] identifying, from the encoded bitstream, independent encoded data blocks corresponding to the one or more audio signals associated with the playback device;
[0183] extracting the identified independent coded data blocks from the coded bitstream;
[0184] extracting the one or more audio signals associated with the playback device from the identified independent encoded data blocks;
[0185] identifying from the encoded bitstream one or more other independent encoded data blocks corresponding to one or more audio signals associated with one or more other playback devices;
[0186] extracting the one or more audio signals associated with the one or more other playback devices from the one or more other independent encoded data blocks;
[0187] capturing one or more audio signals using the one or more microphones of the playback device; and
[0188] One or more extracted audio signals associated with the one or more other playback devices are used as echo references for the playback devices to perform echo management in response to the one or more captured audio signals.
[0189] EEE-C8. The method as described in EEE-C7, further comprising:
[0190] determining that the encoded bitstream includes one or more additional independent encoded data blocks; and
[0191] The one or more additional independent encoded data blocks are ignored.
[0192] EEE-C9. A method as described in EEE-C8, wherein ignoring the one or more additional independent encoded data blocks includes skipping the one or more additional independent encoded data blocks without extracting the one or more additional independent encoded data blocks.
[0193] EEE-C10. A method as described in any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are specifically intended to be used as echo references for the playback device to perform echo management.
[0194] EEE-C11. The method of EEE-C10, wherein the one or more audio signals specifically intended for use as an echo reference are transmitted using less data than the one or more audio signals associated with the playback device. EEE-C12.
[0195] EEE-C12. A method as described in EEE-C10 or EEE-C11, wherein the one or more audio signals specifically intended for use as echo references are reconstructed based on parameter representations of the one or more audio signals.
[0196] EEE-C13. A method as described in any one of EEE-C7 to EEE-C9, wherein the one or more audio signals associated with the one or more other playback devices are suitable for playback from the one or more other playback devices.
[0197] EEE-C14. A method as described in EEE-C7, wherein the encoded signal includes signaling information indicating that the one or more other playback devices are used as echo references for the playback device.
[0198] EEE-C15. A method as described in EEE-C14, wherein the one or more other playback devices indicated by the signaling information of the current frame are different from the one or more other playback devices used as the echo reference of the previous frame.
[0199] EEE-C16. An apparatus configured to perform the method as described in any one of EEE-C1 to EEE-C15.
[0200] EEE-C17. A non-transitory computer-readable storage medium comprising a series of instructions, which, when executed, cause one or more devices to perform a method as described in any one of EEE-C1 to EEE-C15.
[0201] EEE-D1. A method for transmitting an audio signal, the method comprising:
[0202] Generating a data packet comprising a portion of a bitstream, wherein the bitstream comprises a plurality of frames, wherein each frame of the plurality of frames comprises a plurality of blocks, wherein the generating comprises:
[0203] assembling a data packet having one or more of the plurality of blocks, wherein blocks from different frames are combined into a single packet and / or transmitted out of order; and
[0204] The data packets are transmitted via a packet-based network.
[0205] EEE-D2. A method as described in EEE-D1, wherein each block of the plurality of blocks includes identification information.
[0206] EEE-D3. A method as described in EEE-D2, wherein the identification information includes at least one of the following: a block ID, a corresponding frame number associated with the block, and / or a retransmission priority.
[0207] EEE-D4. A method as described in any one of EEE-D1 to EEE-D3, wherein each frame of the multiple frames carries all audio data representing a continuous segment of an audio signal with a start time, an end time and a duration.
[0208] EEE-D5. A method for decoding an audio signal, the method comprising:
[0209] receiving a data packet comprising a portion of a bitstream, wherein the bitstream comprises a plurality of frames, wherein each frame of the plurality of frames comprises a plurality of blocks;
[0210] determining a set of blocks in the plurality of blocks that are addressed to a device; and
[0211] The set of blocks addressed to the device are decoded and decoding of blocks of the plurality of blocks not addressed to the device is skipped.
[0212] EEE-D6. A method for transmitting an audio stream, the method comprising:
[0213] The audio stream is transmitted, wherein the audio stream comprises a plurality of frames, wherein each frame of the plurality of frames comprises a plurality of blocks, wherein the transmitting comprises transmitting configuration information of the audio stream out-of-band.
[0214] EEE-D7. A method as described in EEE-D6, wherein the configuration information of transmitting the audio stream out-of-band includes:
[0215] transmitting the audio stream via a first network and / or a first network protocol; and
[0216] The configuration information is transmitted via a second network and / or a second network protocol.
[0217] EEE-D8. A method as described in EEE-D7, wherein the first network protocol is the User Datagram Protocol (UDP) and the second network protocol is the Transmission Control Protocol (TCP).
[0218] EEE-D9. A method for decoding an audio signal, the method comprising:
[0219] Receive a bit stream, the bit stream comprising:
[0220] Information corresponding to signalling in terms of static configuration;
[0221] Static metadata; and
[0222] One or more channel elements are mapped to one or more devices based on the information and / or the static metadata.
[0223] EEE-D10. A method as described in EEE-D9, wherein the bitstream is received by a plurality of decoders configured to decode the bitstream, wherein each of the plurality of decoders is configured to decode a portion of the bitstream.
[0224] EEE-D11. A method as described in EEE-D9 or EEE-D10, wherein the bitstream further includes dynamic metadata.
[0225] EEE-D12. A method as described in any one of EEE-D9 to EEE-D11, wherein the bitstream includes a plurality of blocks, wherein each block of the plurality of blocks includes:
[0226] information enabling skipping of a portion of the block during decoding, wherein the device does not need the portion; and
[0227] Dynamic metadata.
[0228] EEE-D13. A method for retransmitting a block of an audio signal, the method comprising:
[0229] transmitting one or more blocks of a bitstream, wherein the bitstream comprises a plurality of blocks, wherein each of the one or more blocks of the bitstream has been previously transmitted; and
[0230] Wherein each of the one or more blocks comprises a decoding priority indicator.
[0231] EEE-D14. A method as described in EEE-D13, wherein the decoding priority indicator indicates to a decoder a priority order for decoding the one or more blocks of the bitstream.
[0232] EEE-D15. A method as described in EEE-D13 or EEE-D14, wherein each of the one or more blocks includes the same block ID.
[0233] EEE-D16. A method as described in any one of EEE-D13 to EEE-D15, wherein the transmission of the one or more blocks of the bit stream is transmitted by reducing the data rate compared to the previous transmission.
[0234] EEE-D17. A method as described in EEE-D16, wherein reducing the data rate includes at least one of the following: reducing the signal-to-noise ratio of the audio signal, reducing the bandwidth of the audio signal, and / or reducing the channel count of the audio signal.
[0235] EEE-E1. A method for generating a frame of an encoded bitstream of an audio program comprising a plurality of audio signals, wherein the frame comprises one or more independent encoded data blocks, the method comprising:
[0236] receiving, for each audio signal of the plurality of audio signals, information indicating a playback device associated with the respective audio signal;
[0237] encoding one or more audio signals associated with respective playback devices to obtain one or more encoded audio signals;
[0238] combining the one or more encoded audio signals associated with the respective playback devices into a first independent block of the frame;
[0239] encoding one or more other audio signals of the plurality of audio signals into one or more additional independent blocks; and
[0240] The first independent block and the one or more additional independent blocks are combined into a frame of the encoded bitstream.
[0241] EEE-E2. A method as described in EEE-E1, wherein two or more audio signals are associated with the playback device, and each of the two or more audio signals is a band-limited signal intended to be played back by a corresponding driver of the playback device, and wherein a different encoding technique is used for each of the band-limited signals.
[0242] EEE-E3. A method as described in EEE-E2, wherein a different psychoacoustic model and / or a different bit allocation technique is used for each of the band-limited signals.
[0243] EEE-E4. A method as described in any one of EEE-E1 to EEE-E3, wherein the instantaneous frame rate of the encoded signal is variable and constrained by a buffer fullness model.
[0244] EEE-E5. A method as described in EEE-E1, wherein encoding one or more audio signals associated with the corresponding playback device includes jointly encoding the one or more audio signals associated with the corresponding playback device and one or more additional audio signals associated with one or more additional playback devices into a first independent block of the frame.
[0245] EEE-E6. The method of EEE-E5, wherein jointly encoding the one or more audio signals and the one or more additional audio signals comprises sharing one or more scale factors across two or more audio signals. ...
[0246] EEE-E7. A method as described in EEE-E6, wherein the two or more audio signals are spatially correlated.
[0247] EEE-E8. The method of EEE-E7, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left upper channel, a right horizontal channel, or a right upper channel.
[0248] EEE-E9. A method as described in EEE-E5, wherein the joint encoding of the one or more audio signals and the one or more additional audio signals includes applying a coupling tool, the coupling tool comprising:
[0249] combining two or more audio signals into a composite signal above a specified frequency; and
[0250] A scaling factor related to the energy of the composite signal and the energy of each respective signal is determined for each of the two or more audio signals.
[0251] EEE-E10. The method of EEE-E5, wherein jointly encoding the one or more audio signals and the one or more additional audio signals comprises applying a joint coding tool to more than two signals. EEE-E11.
[0252] EEE-E11. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein the frame comprises one or more independent encoded data blocks, the method comprising:
[0253] identifying, from the encoded bitstream, independent encoded data blocks corresponding to the one or more audio signals associated with the playback device;
[0254] extracting the identified independent coded data blocks from the coded bitstream;
[0255] decoding the one or more audio signals associated with the playback device from the independent encoded data blocks to obtain one or more decoded audio signals;
[0256] identifying from the encoded bitstream one or more additional independent encoded data blocks corresponding to one or more additional audio signals; and
[0257] The one or more additional independent encoded data blocks are decoded or skipped.
[0258] EEE-E12. A method as described in EEE-E11, wherein two or more audio signals are associated with the playback device, and each of the two or more audio signals is a band-limited signal intended to be played back by a corresponding driver of the playback device, and wherein the two or more audio signals are decoded using different decoding techniques.
[0259] EEE-E13. A method as described in EEE-E12, wherein each of the band-limited signals is encoded using a different psychoacoustic model and / or a different bit allocation technique.
[0260] EEE-E14. A method as described in EEE-E11 or EEE-E13, wherein the instantaneous frame rate of the encoded bit stream is variable and constrained by a buffer fullness model.
[0261] EEE-E15. A method as described in EEE-E11, wherein decoding the one or more audio signals associated with the playback device includes jointly decoding the one or more audio signals associated with the corresponding playback device and one or more additional audio signals associated with one or more additional playback devices from the independent encoded data blocks.
[0262] EEE-E16. The method of EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises extracting a scale factor shared across two or more audio signals. EEE-E17.
[0263] EEE-E17. The method of EEE-E16, wherein the two or more audio signals are spatially correlated.
[0264] EEE-E18. The method of EEE-E17, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left upper channel, a right horizontal channel, or a right upper channel. EEE-E19. The method of EEE-E110, wherein the two or more spatially correlated audio signals include a left horizontal channel, a left upper channel, a right horizontal channel, or a right upper channel.
[0265] EEE-E19. The method of EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises applying a decoupling tool.
[0266] EEE-E20. The method of EEE-E19, wherein the decoupling tool comprises:
[0267] Extract independently decoded signals below a specified frequency;
[0268] extracting a composite signal having a frequency higher than the specified frequency;
[0269] determining a corresponding decoupled signal above the specified frequency based on the composite signal and a proportionality factor related to the energy of the composite signal and the energy of the corresponding signal; and
[0270] Each independently decoded signal is combined with the corresponding decoupled signal to obtain the jointly decoded signal.
[0271] EEE-E21. The method of EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises applying a joint decoding tool to extract more than two audio signals. EEE-E22. The method of EEE-E15, wherein jointly decoding the one or more audio signals and the one or more additional audio signals comprises applying a joint decoding tool to extract more than two audio signals.
[0272] EEE-E22. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device comprises applying bandwidth extension to the audio signals in a same domain as the audio signals were encoded. EEE-E23. The method of EEE-E11, wherein decoding the one or more audio signals associated with the playback device comprises applying bandwidth extension to the audio signals in a same domain as the audio signals were encoded.
[0273] EEE-E23. A method as described in EEE-E22, wherein the domain is a modified discrete cosine transform (MDCT) domain.
[0274] EEE-E24. A method as described in EEE-E22 or EEE-E23, wherein the bandwidth extension includes adaptive noise addition.
[0275] EEE-E25. An apparatus configured to perform the method as described in any one of EEE-E1 to EEE-E24.
[0276] EEE-E26. A non-transitory computer-readable storage medium comprising a series of instructions, which, when executed, cause one or more devices to perform a method as described in any one of EEE-E1 to EEE-E24.
[0277] EEE-F1. A method for generating an encoded bitstream performed by a device having one or more microphones, the method comprising:
[0278] capturing one or more audio signals by the one or more microphones;
[0279] analyzing the captured audio signal to determine the presence of a wake word;
[0280] When the presence of the wake word is detected:
[0281] Set a flag to indicate that a speech recognition task will be performed on the captured audio signal:
[0282] encoding the captured audio signal;
[0283] The encoded audio signal and the flag are assembled into the encoded bitstream.
[0284] EEE-F2. A method as described in EEE-F1, wherein the one or more microphones are configured to capture a monophonic or spatial sound field.
[0285] EEE-F3. A method as described in EEE-F2, wherein the spatial sound field adopts A format or B format.
[0286] EEE-F4. A method as described in any one of EEE-F1 to EEE-F3, wherein the captured audio signal is intended only for performing the speech recognition task.
[0287] EEE-F5. The method of EEE-F4, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient to perform the speech recognition task but insufficient for human listening.
[0288] EEE-F6. A method as described in EEE-F4 or EEE-F5, wherein, before encoding the captured audio signal, the captured audio signal is converted into a representation comprising one or more of the following: frequency band energy, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0289] EEE-F7. A method as described in any of EEE-F1 to EEE-F3, wherein the captured audio signal is intended both for human listening and for performing the speech recognition task.
[0290] EEE-F8. The method of EEE-F7, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening.
[0291] EEE-F9. A method as described in EEE-F7, wherein encoding the captured audio signal includes generating a first encoded representation of the captured audio signal and a second encoded representation of the captured audio signal, wherein the first encoded representation is generated so that when the captured audio signal is decoded from the first encoded representation, the quality of the decoded audio signal is sufficient for human listening, and wherein the second encoded representation is generated so that when the captured audio signal is decoded from the second encoded representation, the quality of the decoded audio signal is sufficient to perform the speech recognition task but not sufficient for human listening.
[0292] EEE-F10. A method as described in EEE-F9, wherein generating the second encoded representation of the captured audio signal includes converting the captured audio signal into one or more of the following before encoding the captured audio signal: a parametric representation, a coarse waveform representation, or a representation including one or more of frequency band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0293] EEE-F11. A method as described in EEE-F9 or EEE-F10, wherein assembling the encoded audio signal into a bitstream includes inserting the first encoded representation into a first independent block of the encoded bitstream, and inserting the second encoded representation into a second independent block of the encoded bitstream.
[0294] EEE-F12. A method as described in EEE-F9 or EEE-F10, wherein the first encoded representation is included in a first layer of the encoded bitstream, the second encoded representation is included in a second layer of the encoded bitstream, and the first layer and the second layer are included in a single block of the encoded bitstream.
[0295] EEE-F13. The method of any one of EEE-F1 to EEE-F12, further comprising, when the presence of the wake-up word is not detected:
[0296] Setting the flag to indicate that a speech recognition task is not to be performed on the captured audio signal;
[0297] encoding the captured audio signal;
[0298] The encoded audio signal and the flag are assembled into the encoded bitstream.
[0299] EEE-F14. A method for decoding an audio signal, the method comprising:
[0300] receiving an encoded bit stream, the encoded bit stream comprising an encoded audio signal and a flag indicating whether a speech recognition task is to be performed;
[0301] decoding the encoded audio signal to obtain a decoded audio signal; and
[0302] When the flag indicates that the speech recognition task is to be performed, the speech recognition task is performed on the decoded audio signal.
[0303] EEE-F15. A method as described in EEE-F14, wherein the decoded audio signal is intended to be used only for performing the speech recognition task.
[0304] EEE-F16. The method of EEE-F15, wherein the quality of the decoded audio signal is sufficient to perform the speech recognition task but is insufficient for human listening.
[0305] EEE-F17. A method as claimed in claim EEE-F15 or EEE-F16, wherein, before encoding the captured audio signal, the decoded audio signal is in a representation comprising one or more of frequency band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0306] EEE-F18. The method of EEE-F14, wherein the captured audio signal is encoded such that when the captured audio signal is decoded, the quality of the decoded audio signal is sufficient for human listening. EEE-F19.
[0307] EEE-F19. The method of EEE-F18, wherein the encoded audio signal comprises a first encoded representation of one or more audio signals and a second encoded representation of the one or more audio signals. ...
[0308] EEE-F20. A method as described in EEE-F18, wherein the quality of the audio signal decoded from the first representation is sufficient for human listening, and wherein the quality of the audio signal decoded from the second representation is sufficient to perform the speech recognition task but not sufficient for human listening.
[0309] EEE-F21. A method as described in EEE-F19 or EEE-F20, wherein the first representation is in a first independent block of the encoded bitstream and the second representation is in a second independent block of the encoded bitstream.
[0310] EEE-F22. A method as described in EEE-F19 or EEE-F20, wherein the first representation is in a first layer of the encoded bit stream, the second encoded representation is included in a second layer of the encoded bit stream, and the first layer and the second layer are included in a single block of the encoded bit stream.
[0311] EEE-F23. A method as described in any one of EEE-F18 to EEE-F22, wherein decoding the encoded audio signal includes decoding only the second representation and ignoring the first representation.
[0312] EEE-F24. A method as described in any one of EEE-F18 to EEE-F23, wherein the audio signal decoded from the second encoded representation is in a parametric representation, a waveform representation, or a representation comprising one or more of frequency band energies, Mel-frequency cepstral coefficients, or modified discrete cosine transform (MDCT) spectral coefficients.
[0313] EEE-F25. An apparatus configured to perform a method as described in any one of EEE-F1 to EEE-F24.
[0314] EEE-F26. A non-transitory computer-readable storage medium comprising a series of instructions, which, when executed, cause one or more devices to perform a method as described in any one of EEE-F1 to EEE-F24.
[0315] EEE-G1. A method for encoding an audio signal of an immersive audio program for low-latency transmission to one or more playback devices, the method comprising:
[0316] receiving a plurality of time-domain audio signals of the immersive audio program;
[0317] Select the frame size;
[0318] extracting a frame of the time-domain audio signal in response to the frame size, wherein the frame of the time-domain audio signal overlaps with a previous frame of the time-domain audio signal;
[0319] segmenting the audio signal into overlapping frames;
[0320] transforming frames of the time-domain audio signal into frequency-domain signals;
[0321] Encoding the frequency domain signal;
[0322] quantizing the encoded frequency domain signal using a perceptually motivated quantization tool;
[0323] assembling the quantized and encoded frequency domain signal into one or more independent blocks within the frame; and
[0324] The one or more independent blocks are assembled into an encoded frame.
[0325] EEE-G2. The method of EEE-G1, wherein the plurality of audio signals comprises channel-based signals having a defined channel configuration. ...
[0326] EEE-G3. A method as described in EEE-G2, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6 or 22.2.
[0327] EEE-G4. A method as described in any one of EEE-G1 to EEE-G3, wherein the multiple audio signals include one or more object-based signals.
[0328] EEE-G5. A method as described in any of EEE-G1 to EEE-G4, wherein the plurality of audio signals include a scene-based representation of the immersive audio program.
[0329] EEE-G6. A method as described in any of EEE-G1 to EEE-G5, wherein the selected frame size is one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.
[0330] EEE-G7. A method as described in any one of EEE-G1 to EEE-G6, wherein the overlap between a frame of the time domain audio signal and a previous frame of the time domain audio signal is 50% or less.
[0331] EEE-G8. A method as described in any one of EEE-G1 to EEE-G7, wherein the transform is a modified discrete cosine transform (MDCT).
[0332] EEE-G9. A method as described in any one of EEE-G1 to EEE-G8, wherein two or more audio signals of the multiple audio signals are jointly encoded.
[0333] EEE-G10. A method as described in any of EEE-G1 to EEE-G9, wherein each independent block contains an encoded signal for one or more playback devices.
[0334] EEE-G11. A method as described in any of EEE-G1 to EEE-G10, wherein at least one independent block contains an encoded signal for two or more playback devices, and the encoded signal includes a jointly encoded audio signal.
[0335] EEE-G12. A method as described in any of EEE-G1 to EEE-G11, wherein at least one independent block contains multiple encoded signals covering different bandwidths intended for playback from different drives of a playback device.
[0336] EEE-G13. A method as described in any of EEE-G1 to EEE-G12, wherein at least one independent block contains an encoded echo reference signal used in echo management performed by a playback device.
[0337] EEE-G14. A method as described in any one of EEE-G1 to EEE-G13, wherein encoding the quantized frequency domain signal includes applying one or more of the following tools: temporal noise shaping (TNS), joint channel coding, sharing scaling factors across signals, determining control parameters for high frequency reconstruction, and determining control parameters for noise replacement.
[0338] EEE-G15. A method as described in any of EEE-G1 to EEE-G14, wherein one or more independent blocks include parameters for controlling one or more of delay, gain, and equalization of the playback device.
[0339] EEE-G16. A low-latency method for decoding an audio signal of an immersive audio program from an encoded signal, the method comprising:
[0340] receiving an encoded frame comprising one or more independent blocks;
[0341] extracting a quantized and encoded frequency domain signal from one or more independent blocks;
[0342] Dequantizing the quantized and encoded frequency domain signal;
[0343] Decoding the dequantized frequency domain signal;
[0344] Inversely transforming the decoded frequency domain signal to obtain a time domain signal; and
[0345] The time domain signal is overlapped and added with a time domain signal from a previous frame to provide a plurality of audio signals of the immersive audio program.
[0346] EEE-G17. A method as described in EEE-G16, wherein the plurality of audio signals include channel-based signals having a defined channel configuration.
[0347] EEE-G18. A method as described in EEE-G17, wherein the channel configuration is one of mono, stereo, 5.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6 or 22.2.
[0348] EEE-G19. A method as described in any one of EEE-G16 to EEE-G18, wherein the multiple audio signals include one or more object-based signals.
[0349] EEE-G20. A method as described in any of EEE-G16 to EEE-G19, wherein the plurality of audio signals include a scene-based representation of the immersive audio program.
[0350] EEE-G21. A method as described in any of EEE-G16 to EEE-G20, wherein a frame of time domain samples includes one of 128, 256, 512, 1024, 120, 240, 480 or 960 samples.
[0351] EEE-G22. A method as described in any of EEE-G16 to EEE-G21, wherein the overlap with the previous frame is 50% or less.
[0352] EEE-G23. A method as described in any one of EEE-G16 to EEE-G22, wherein the inverse transform is an inverse improved discrete cosine transform (IMDCT).
[0353] EEE-G24. A method as described in any of EEE-G16 to EEE-G23, wherein each independent block contains a quantized and encoded frequency domain signal for one or more playback devices.
[0354] EEE-G25. A method as described in any one of EEE-G16 to EEE-G24, wherein at least one independent block contains quantized and encoded frequency domain signals for two or more playback devices, and the quantized and encoded frequency domain signals are jointly encoded audio signals.
[0355] EEE-G26. A method as described in any of EEE-G16 to EEE-G25, wherein at least one independent block contains multiple quantized and encoded frequency domain signals covering different bandwidths intended for playback from different drives of a playback device.
[0356] EEE-G27. A method as described in any of EEE-G16 to EEE-G26, wherein at least one independent block contains an encoded echo reference signal used in echo management performed by a playback device.
[0357] EEE-G28. A method as described in any one of EEE-G16 to EEE-G27, wherein decoding the dequantized frequency domain signal includes applying one or more of the following decoding tools: temporal noise shaping (TNS), joint channel decoding, cross-signal sharing of scaling factors, high frequency reconstruction, and noise replacement.
[0358] EEE-G29. A method as described in any of EEE-G16 to EEE-G28, wherein one or more independent blocks include parameters for controlling one or more of delay, gain, and equalization of a playback device.
[0359] EEE-G30. A method as described in any one of EEE-G16 to EEE-G29, wherein the method is performed by a playback device, and wherein extracting the quantized and encoded signals from one or more independent blocks includes selecting only those blocks containing quantized and encoded frequency domain signals for playback by the playback device, and ignoring independent blocks containing quantized and encoded frequency domain signals for playback by other playback devices.
[0360] EEE-G31. An apparatus configured to perform a method as described in any one of EEE-G1 to EEE-G30.
[0361] EEE-G32. A non-transitory computer-readable storage medium comprising a series of instructions, which, when executed, cause one or more devices to perform a method as described in any one of EEE-G1 to EEE-G30.
[0362] In the claims below and the description herein, any of the terms comprising, including or which comprises is an open term, which means at least including the elements / features that follow, but not excluding other elements / features. Therefore, when the term comprising is used in a claim, the term should not be interpreted as being limited to the devices or elements or steps listed thereafter. For example, the scope of the expression "a device comprising A and B" should not be limited to a device comprising only elements A and B. As used herein, the term including or any of which includes or that includes is also an open term, which also means at least including the elements / features that follow the term, but not excluding other elements / features. Therefore, including is synonymous with comprising and means comprising.
[0363] It should be understood that in the above description of examples of the invention, various features are sometimes grouped together in a single example, figure, or description thereof to simplify the disclosure and aid in understanding one or more of the various creative aspects. However, this method of disclosure should not be interpreted as reflecting an intention to require more features than those explicitly recited in each claim. On the contrary, as reflected in the following claims, the creative aspects lie in less than all the features of a single preceding disclosed example. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim stands on its own as a separate example of the present invention.
[0364] In addition, although some examples described herein include some features included in other examples and do not include other features included in other examples, as will be understood by those skilled in the art, combinations of features of different examples are intended to be covered and combinations of features of different examples form different examples. For example, in the following claims, any of the claimed examples may be used in any combination.
[0365] In addition, some examples are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices that perform functions. Therefore, a processor with instructions required to perform such methods or method elements forms a device for performing the methods or method elements. In addition, the elements of the devices described herein are examples of devices for performing the functions performed by the elements.
[0366] Furthermore, some of the examples described herein disclose connection solutions that should be interpreted as being possible to implement in distribution and / or transmission systems such as wired systems and / or wireless systems, for example, by using any electrical, optical and / or mobile systems such as 3G, 4G and 5G.
[0367] Thus, while specific examples of the present invention have been described, those skilled in the art will recognize that other and further modifications may be made, and it is intended that all such changes and modifications be claimed. For example, any formula given above merely represents a procedure that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged among functional blocks. The described methods may have steps added or deleted.
[0368] The systems, devices and methods disclosed above may be implemented as software, firmware, hardware or a combination thereof. For example, aspects of the present application may be at least partially embodied in a device, a system including more than one device, a method, a computer program product, etc.
[0369] In hardware implementations, the division of tasks between functional units mentioned in the above description does not necessarily correspond to the division of physical units; on the contrary, one physical component may have multiple functions, and one task may be performed by several physical components in collaboration.
[0370] Some or all components may be implemented as software executed by a digital signal processor or microprocessor or as hardware or an application specific integrated circuit. Such software may be distributed on a computer readable medium, which may include a computer storage medium (or non-transient medium) and a communication medium (or transient medium).
[0371] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to: RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0372] Furthermore, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules or other data in the form of modulated data signals such as carrier waves or other transmission mechanisms, and includes any information delivery media.
Claims
1. A method for generating an encoded bitstream from an audio program comprising a plurality of audio signals, the method comprising: receiving, for each audio signal of the plurality of audio signals, information indicating a playback device associated with the respective audio signal; receiving, for each playback device, information indicative of at least one of a delay, a gain, and an equalization curve associated with the respective playback device; determining a set of two or more related audio signals from the plurality of audio signals; applying one or more joint coding tools to the two or more related audio signals in the group to obtain a jointly encoded audio signal; The jointly encoded audio signal, an indication of the playback device associated with the jointly encoded audio signal, and information indicating at least one of a delay, gain, and equalization curve associated with the respective playback device associated with the jointly encoded audio signal are combined into separate blocks of an encoded bitstream.
2. The method according to claim 1, wherein: The information indicates the delay and the gain.
3. The method according to claim 1, wherein: The delay, the gain and / or the equalization curve associated with the respective playback device depends on the position of the respective playback device relative to the position of the listener.
4. The method according to any one of claims 1 to 3, wherein: The delay, the gain, and / or the equalization curve associated with the respective playback device depends on the position of the respective playback device relative to the positions of other playback devices.
5. The method according to any one of claims 1 to 4, wherein: The delay, the gain and / or the equalization curve are dynamically variable.
6. The method according to claim 5, wherein: The delay, the gain and / or the equalization curve are adjusted in response to a change in the position of the listener.
7. The method according to claim 5 or claim 6, wherein: The delay, the gain and / or the equalization curve are adjusted in response to a change in the position of the playback device.
8. The method according to any one of claims 5 to 7, wherein: The delay, the gain, and / or the equalization curve are adjusted in response to a change in position of one or more of the other playback devices.
9. The method according to any one of claims 1 to 8, further comprising determining from the plurality of audio signals an audio signal that is not part of the set of two or more related audio signals.
10. The method according to claim 9, further comprising: For the audio signal that is not part of the set of two or more related audio signals, the delay, the gain and / or the equalization curve associated with the playback device associated with the audio signal is applied.
11. The method of claim 10 , further comprising independently encoding the audio signal that is not part of the set of two or more related audio signals, and combining the independently encoded audio signal and the indication of the playback device associated with the independently encoded audio signal into a separate independently decodable subset of the encoded bitstream.
12. The method of claim 9, further comprising independently encoding the audio signal that is not part of the group of two or more related audio signals, and combining the independently encoded audio signal, an indication of the playback device associated with the independently encoded audio signal, and an indication of the delay, the gain and / or the equalization curve associated with the playback device associated with the independently encoded audio signal into a separate independently decodable subset of the encoded bitstream.
13. A method for decoding one or more audio signals associated with a playback device from a frame of an encoded bitstream, wherein: The frame comprises one or more independent encoded data blocks, and the method comprises: identifying, from the encoded bitstream, independent encoded data blocks corresponding to the one or more audio signals associated with the playback device; extracting the identified independent coded data blocks from the coded bitstream; determining that the extracted independent coded data block comprises two or more jointly coded audio signals; applying one or more joint decoding tools to the two or more jointly encoded audio signals to obtain the one or more audio signals associated with the playback device; determining at least one of a delay, gain, and equalization curve associated with the playback device from the extracted independent encoded data blocks; The delay, the gain, and / or the equalization curve associated with the playback device are applied to the one or more audio signals associated with the playback device.
14. The method of claim 13, wherein: The determined delay, gain and / or equalization curve associated with the playback device depends on the position of the playback device relative to the position of the listener.
15. The method according to claim 13 or claim 14, wherein: The determined delay, gain, and / or equalization curve associated with the playback device depends on the position of the playback device relative to other playback devices.
16. The method according to any one of claims 13 to 15, wherein: The determined delay, gain and / or equalization curve of the playback device is dynamically variable.
17. The method according to claim 16, wherein: When the determined delay, gain and / or equalization curve associated with the playback device is different from a previously determined delay, gain and / or equalization curve associated with the playback device, the method further includes interpolating between the previously determined delay, gain and / or equalization curve associated with the playback device and the determined delay, gain and / or equalization curve associated with the playback device.
18. The method according to claim 17, wherein: Due to the change in the position of the listener, the determined delay, gain and / or equalization curve differs from the previously determined delay, gain and / or equalization curve.
19. The method according to claim 17 or claim 18, wherein: Due to the change in position of the playback device, the determined delay, gain and / or equalization curve differs from the previously determined delay, gain and / or equalization curve.
20. The method according to any one of claims 17 to 19, wherein: Due to a change in position of one or more of the other playback devices, the determined delay, gain, and / or equalization curve differs from the previously determined delay, gain, or equalization curve.
21. The method according to any one of claims 13 to 20, wherein: The frame of the encoded bitstream comprises two or more independent encoded data blocks, the method further comprising: determining that one or more of the independent blocks contain audio signals not associated with the playback device; and The one or more independent blocks containing audio signals not associated with the playback device are ignored.
22. The method according to any one of claims 13 to 21, wherein: Applying one or more joint decoding tools includes identifying a subset of the jointly encoded audio signals associated with the playback device, and reconstructing only the subset of the jointly encoded audio signals to obtain the one or more audio signals associated with the playback device.
23. The method according to any one of claims 13 to 21, wherein: Applying one or more joint decoding tools includes reconstructing each of the jointly encoded audio signals, identifying a subset of the reconstructed jointly encoded audio signals associated with the playback device, and obtaining the one or more audio signals associated with the playback device from the subset of the reconstructed jointly encoded audio signals associated with the playback device.
24. An apparatus configured to perform the method according to any one of claims 1 to 23.
25. A non-transitory computer readable storage medium comprising a series of instructions which, when executed, cause one or more devices to perform the method of any one of claims 1 to 23.
Citation Information
Patent Citations
Selective forward error correction for spatial audio codecs
US11289103B2