Converting audio signals captured in different formats to fewer formats to simplify encoding and decoding operations.
By converting audio signals to a limited set of formats using a simplification unit and metadata, the complexity and cost of IVAS codecs are reduced, enabling deployment across diverse devices with maintained quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-11
AI Technical Summary
Existing audio codecs, such as IVAS, face complexity and cost issues due to the need to support a wide variety of audio capture and rendering formats, making it impractical to implement across diverse devices.
A simplification unit converts audio signals from various capture formats to a limited set of formats, such as mono, stereo, and spatial, before encoding, using metadata to represent spatial information, thereby reducing codec complexity.
This approach allows IVAS codecs to be deployed across various devices with reduced complexity and cost, while maintaining equivalent experiential quality.
Smart Images

Figure 2026076253000001_ABST
Abstract
Description
Technical Field
[0004]
[0001] Cross - Reference to Related Applications This application claims the benefit of priority from U.S. Provisional Patent Application No. 62 / 742,729, filed Oct. 8, 2018. The content of the application is incorporated herein by reference in its entirety.
[0002] Technique Embodiments of the present disclosure generally relate to audio signal processing, and more particularly to the distribution of captured audio signals.
Background Art
[0003] The development of audio and video encoder / decoder (“codec”) standards has recently focused on the development of codecs for Immersive Voice and Audio Services (IVAS). IVAS is expected to support a range of service functions such as operation from monaural to stereo and even fully immersive audio encoding, decoding, and rendering. A suitable IVAS codec is also expected to provide high error resilience against packet loss and delay jitter under different transmission conditions. IVAS is intended to be supported by a wide range of devices, endpoints, and network nodes including, but not limited to, mobile phones and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality and augmented reality devices, home theater devices, and other suitable devices. Since these devices, endpoints, and network nodes can have various acoustic interfaces for sound capture and rendering, it may not be practical for an IVAS codec to accommodate all the various ways in which audio signals are captured and rendered.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The disclosed embodiments enable the conversion of audio signals captured in various formats by various capture devices into a limited number of formats that can be processed by a codec, such as an IVAS codec. [Means for solving the problem]
[0005] In some embodiments, a simplification unit incorporated into the audio device receives an audio signal. This audio signal may be a signal captured by one or more audio capture devices coupled to the audio device. The audio signal may, for example, be the audio of a video conference between people in different locations. The simplification unit determines whether the audio signal is in a format not supported by the audio device's encoding unit, commonly referred to as an "encoder." For example, the simplification unit may determine whether the audio signal is mono, stereo, or in a standard or proprietary spatial format. Based on the determination that the audio signal is in a format not supported by the encoding unit, the simplification means converts the audio signal to a format supported by the encoding unit. For example, if the simplification unit determines that the audio signal is in a proprietary spatial format, the simplification unit may convert the audio signal to a spatial "mezzanine" format supported by the encoding unit. The simplification unit then transfers the converted audio signal to the encoding unit.
[0006] The advantage of the disclosed embodiment is that the complexity of the codec, such as the IVAS codec, can be reduced by decreasing the potentially large number of audio capture formats to a limited number of formats, such as mono, stereo, and spatial. As a result, the codec can be deployed in a variety of devices, regardless of the device's audio capture capabilities.
[0007] These and other aspects, features, and embodiments can be expressed as methods, apparatus, systems, components, program products, means or steps for performing functions, and in other ways.
[0008] In some implementations, the simplification unit of the audio device receives an audio signal in a first format. The first format is one of a set of multiple audio formats supported by the audio device. The simplification unit determines whether the first format is supported by the encoder of the audio device. Based on the finding that the first format is not supported by the encoder, the simplification unit converts the audio signal to a second format supported by the encoder. The second format is an alternative representation of the first format. The simplification unit transfers the audio signal in the second format to the encoder. The encoder encodes the audio signal. The audio device either stores the encoded audio signal or transmits the encoded audio signal to one or more other devices.
[0009] Converting an audio signal to a second format may include generating metadata about the audio signal. This metadata may include representations of parts of the audio signal. Encoding an audio signal may include encoding the audio signal in the second format into a transport format supported by a second device. The audio device may transmit the encoded audio signal by transmitting metadata that includes representations of parts of the audio signal not supported by the second format.
[0010] In some implementations, the simplification unit may determine whether an audio signal is in a first format by determining the number of audio capture devices and the corresponding location of each capture device used to capture the audio signal. Each of the one or more other devices may be configured to reproduce an audio signal from a second format. At least one of the one or more other devices may not be able to reproduce an audio signal from the first format.
[0011] The second format can represent an audio signal as several audio objects within an audio scene, each relying on several audio channels to carry spatial information. The second format may include metadata to carry further parts of the spatial information. Both the first and second formats can be spatial audio formats. The second format may be a spatial audio format, while the first format may be a mono format associated with metadata, or a stereo format associated with metadata. A set of multiple audio formats supported by an audio device may include multiple spatial audio formats. The second format may also be an alternative representation of the first format, further characterized by enabling an equivalent degree of experiential quality.
[0012] In some implementations, the rendering unit of the audio device receives the audio signal in a first format. The rendering unit determines whether the audio device can reproduce the audio signal in the first format. In response to determining that the audio device cannot reproduce the audio signal in the first format, the rendering unit adapts the audio signal to make it available in a second format. The rendering unit then transfers the audio signal in the second format for rendering.
[0013] In some implementations, the rendering unit may convert the audio signal to a second format by combining it with metadata containing a portion of the audio signal not supported by the fourth format used for encoding, in conjunction with the audio signal in the third format. Here, the third format corresponds to the term "first format" in the context of the simplification unit, which is one of a set of multiple audio formats supported by the encoder. The fourth format corresponds to the term "second format" in the context of the simplification unit, which is a format supported by the encoder and is an alternative representation of the third format. Hereafter and elsewhere in this specification, the terms first, second, third and fourth are used for identification purposes and do not necessarily indicate a particular order.
[0014] A decoding unit receives the audio signal in the transport format. The decoding unit decodes the audio signal in the transport format into a first format and transfers the audio signal in the first format to a rendering unit. In some implementations, adapting the audio signal to be available in a second format may include adapting the decoding to produce the received audio in the second format. In some implementations, each of a plurality of devices is configured to play the audio signal in the second format. One or more of the plurality of devices may not be able to play the audio signal in the first format.
[0015] In some implementations, a simplification unit receives audio signals in multiple formats from an audio preprocessing unit. The simplification unit receives device attributes from the device, which include indications of one or more audio formats supported by the device. These one or more audio formats include at least one of mono, stereo, or spatial formats. The simplification unit converts the audio signals into ingest formats, which are alternative representations of the one or more audio formats. The simplification unit provides the converted audio signals to an encoding unit for downstream processing. Each of the audio preprocessing unit, the simplification unit, and the encoding unit may include one or more computer processors.
[0016] In some implementations, the encoding system includes a capture unit configured to capture an audio signal, an acoustic preprocessing unit configured to perform operations including preprocessing the audio signal, an encoder, and a simplification unit. The simplification unit is configured to perform the following operations: The simplification unit receives an audio signal in a first format from the acoustic preprocessing unit. The first format is one of a set of multiple audio formats supported by the encoder. The simplification unit determines whether the first format is supported by the encoder. In response to determining that the first format is not supported by the encoder, the simplification unit converts the audio signal to a second format supported by the encoder. The simplification unit transfers the audio signal in the second format to the encoder. The encoder is configured to perform operations including encoding the audio signal, storing the encoded audio signal, or transmitting the encoded audio signal to another device, at least one of these.
[0017] In some implementations, converting an audio signal to a second format involves generating metadata for the audio signal. This metadata may include representations of parts of the audio signal that are not supported by the second format. The encoder's operation may further include transmitting the encoded audio signal by transmitting the metadata, which includes representations of parts of the audio signal that are not supported by the second format.
[0018] In some implementations, the second format represents the audio signal as several (a number of) channels to carry several (a number of) objects and spatial information in the audio scene. In some implementations, preprocessing of the audio signal may include one or more of the following: performing noise cancellation, performing echo cancellation, reducing the number of channels in the audio signal, increasing the number of audio channels in the audio signal, or generating acoustic metadata.
[0019] In some implementations, the decoding system includes a decoder, a rendering unit, and a playback unit. The decoder is configured to perform operations including, for example, decoding an audio signal from a transport format to a first format. The rendering unit is configured to perform the following operations: The rendering unit receives the audio signal in the first format. The rendering unit determines whether the audio device can play the audio signal in a second format. The second format allows the use of more output devices than the first format. In response to determining that the audio device can play the audio signal in the second format, the rendering unit converts the audio signal to the second format. The rendering unit renders the audio signal in the second format. The playback unit is configured to perform operations including initiating playback of the rendered audio signal on a speaker system.
[0020] In some implementations, converting an audio signal to a second format may involve using metadata that combines the audio signal of the third format with a representation of parts of the audio signal that are not supported by the fourth format used for encoding. Here, the third format corresponds to the term "first format" in the context of the simplification unit, which is one of a set of multiple audio formats supported by the encoder. The fourth format corresponds to the term "second format" in the context of the simplification unit, which is a format supported by the encoder and is an alternative representation of the third format.
[0021] In some implementations, the decoder's operation may further include receiving the audio signal in a transport format and transferring the audio signal in a first format to a rendering unit.
[0022] These and other aspects, features, and embodiments will become apparent from the following description, including the claims. [Brief explanation of the drawing]
[0023] Drawings often show specific arrangements or sequences of schematic elements, such as those representing devices, units, instruction blocks, and data elements, for the sake of clarity. However, those skilled in the art should understand that the specific ordering or arrangement of schematic elements in the drawings is not intended to imply that a particular order or sequence of operations or separation of processes is required. Furthermore, the inclusion of schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, or that features represented by such elements should not be included in or combined with other elements in some embodiments.
[0024] Furthermore, when connecting elements such as solid or dashed lines or arrows are used in drawings to illustrate the connection, relationship, or association of two or more other schematic elements, the absence of such connecting elements is not intended to imply that the connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. Moreover, for the sake of simplicity of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of signals, data, or instructions, a person skilled in the art should understand that such an element represents one or more signal paths as necessary to affect the communication. [Figure 1] This disclosure illustrates various devices that can be supported by the IVAS system according to several embodiments of this disclosure. [Figure 2]A is a block diagram of a system for converting a captured audio signal into a format ready for encoding, according to some embodiments of the present disclosure. B is a block diagram of a system for converting and restoring the captured audio into a suitable playback format, according to some embodiments of the present disclosure. [Figure 3] It is a flowchart of exemplary actions for converting an audio signal into a format supported by an encoding unit, according to some embodiments of the present disclosure. [Figure 4] It is a flowchart of exemplary actions for determining whether an audio signal is in a format supported by an encoding unit, according to some embodiments of the present disclosure. [Figure 5] It is a flowchart of exemplary actions for converting an audio signal into an available playback format, according to some embodiments of the present disclosure. [Figure 6] It is another flowchart of exemplary actions for converting an audio signal into an available playback format, according to some embodiments of the present disclosure. [Figure 7] It is a block diagram of a hardware architecture for implementing the features described with reference to FIGS. 1 - 6, according to some embodiments of the present disclosure.
Mode for Carrying Out the Invention
[0025] In the following description, for the purpose of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the present disclosure may be practiced without these specific details.
[0026] Herein, we refer in detail to embodiments illustrated in the accompanying drawings. The following detailed description includes numerous specific details to provide a full understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments can be carried out without these specific details. On the other hand, well-known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure aspects of the embodiments. Hereinafter, several features are described that can be used independently of each other or in any combination of other features.
[0027] Where used herein, the term "includes" and its variations should be read as an open-ended term meaning "including, but not limited to..." The term "or" should be read as "and / or" unless the context explicitly indicates otherwise. The term "based on" should be read as "based at least in part on..."
[0028] Figure 1 shows various devices that can be supported by the IVAS system. In some implementations, these devices communicate through a call server 102 that can receive audio signals from a public switched telephone network (PSTN) or public terrestrial mobile network equipment (PLMN), for example, by a PSTN / other PLMN device 104. This device can use the G.711 and / or G.722 standards for audio (speech) compression and decompression. Device 104 can generally only capture and render monaural audio. The IVAS system is also capable of supporting legacy user devices 106. These legacy devices may include enhanced voice services (EVS) devices, adaptive multi-rate wideband (AMR-WB) speech-to-audio coding standard supporting devices, adaptive multi-rate narrowband (AMR-NB) supporting devices, and other suitable devices. These devices typically render and capture audio in mono only.
[0029] The IVAS system is also capable of supporting user devices that capture and render audio signals in a variety of formats, including advanced audio formats. For example, the IVAS system is capable of supporting stereo capture and rendering devices (e.g., user device 108, laptop 114, and conference room system 118), mono capture and binaural rendering devices (e.g., user device 110 and computer device 112), immersive capture and rendering devices (e.g., conference room use device 116), stereo capture and immersive rendering devices (e.g., home theater 120), mono capture and immersive rendering (e.g., virtual reality (VR) gear 122), immersive content consumption 124, and other suitable devices. To directly support all these formats, the codecs for the IVAS system would need to be very complex and expensive to implement. Therefore, a system for simplifying the codecs prior to the encoding stage is desirable.
[0030] While the following description focuses on IVAS systems and codecs, the disclosed embodiments are applicable to any codecs for any audio system where reducing the number of audio capture formats to a smaller number would be beneficial, either to reduce the complexity of the audio codec or for any other desired reason.
[0031] Figure 2A is a block diagram of a system 200 for converting captured audio signals into a format ready for encoding, according to some embodiments of the present disclosure. The capture unit 210 receives audio signals from one or more capture devices, such as microphones. For example, the capture unit 210 may receive audio signals from one microphone (e.g., a mono signal), from two microphones (e.g., a stereo signal), from three microphones, or from a different number and configuration of audio capture devices. The capture unit 210 may include customizations by one or more third parties, and these customizations may be specific to the capture devices used.
[0032] In some implementations, a mono audio signal is captured by a single microphone. The mono signal can be captured, for example, by a PSTN / PLMN telephone 104, a legacy user device 106, a user device 110 with a hands-free headset, a computer device 112 with a connected headset, and a virtual reality gear 122, as shown in Figure 1.
[0033] In some implementations, the capture unit 210 receives stereo audio captured using various recording / microphone techniques. Stereo audio can be captured by, for example, user equipment 108, laptop 114, conference room system 118, and home theater 120. In one example, stereo audio is captured by two directional microphones located in the same position, positioned at a spread angle of approximately 90 degrees or more. The stereo effect is due to the level difference between channels. In another example, stereo audio is captured by two spatially displaced microphones. In some implementations, the spatially displaced microphones are omnidirectional microphones. The stereo effect in this configuration is due to the level and time difference between channels. The distance between microphones significantly affects the perceived stereo width. In yet another example, audio is captured by two directional microphones with a displacement of 17 cm and a spread angle of 110 degrees. This system is often referred to as the French television broadcasting office (Office de Radiodiffusion Television Francaise, "ORTF") stereo microphone system. Another stereo capture system involves two microphones with different characteristics, arranged so that the signal from one microphone is the mid-range signal and the other is the side-range signal. This arrangement is often called mid-side (M / S) recording. The stereo effect of the signal from M / S is typically formed based on the level difference between channels.
[0034] In some implementations, the capture unit 210 receives audio captured using a multi-microphone technique. In these implementations, audio capture involves the placement of three or more microphones. This placement is generally required to capture spatial audio and can also be effective in suppressing ambient noise. As the number of microphones increases, so does the number of details of the spatial scene that can be captured by the microphones. In some cases, increasing the number of microphones also improves the accuracy of the captured scene. For example, the various user equipment (UEs) in Figure 1, which can be operated in hands-free mode, can utilize multiple microphones to generate mono, stereo, or spatial audio signals. Furthermore, an open laptop computer 114 with multiple microphones can be used to generate stereo capture. Some manufacturers have released laptop computers with 2 to 4 micro-electromechanical system ("MEMS") microphones that allow for stereo capture. Immersive audio capture with multiple microphones can be implemented, for example, in a conference room user equipment 216.
[0035] The captured audio generally undergoes a pre-processing stage before being taken up by a voice or audio codec. Thus, the acoustic pre-processing unit 220 receives the audio signal from the capture unit 210. In some implementations, the acoustic pre-processing unit 220 performs noise and echo cancellation, channel downmixing and upmixing (e.g., reducing or increasing the number of audio channels), and / or any kind of spatial processing. The audio signal output of the acoustic pre-processing unit 220 is generally suitable for encoding and transmission to other devices. In some implementations, the specific design of the acoustic pre-processing unit 220 is carried out by the device manufacturer, as it depends on the details of audio capture using the specific device. However, requirements set by appropriate acoustic interface specifications can set limits on these designs and ensure that certain quality requirements are met. Acoustic pre-processing is performed with the aim of generating one or more different types of audio signals or audio input formats supported by the IVAS codec, enabling various IVAS target use cases or service levels. Depending on the specific IVAS service requirements associated with these use cases, IVAS codecs may be required to support mono, stereo, and spatial formats.
[0036] Generally, the mono format is used when it is the only available format, for example, based on the type of capture device, such as when the capture capability of the transmitting device is limited. For stereo audio signals, the acoustic pre-processing unit 220 converts the captured signal into a normalized representation that satisfies a specific convention (e.g., left / right channel ordering convention). For M / S stereo capture, this process may involve matrix operations, for example, so that the signal is represented using the left / right convention. After pre-processing, the stereo signal satisfies a certain convention (e.g., left / right convention), however, information about the specific stereo capture device (e.g., number and configuration of microphones) is removed.
[0037] With regard to spatial formats, the type of spatial input signal or specific spatial audio format obtained after acoustic preprocessing may depend on the type of transmitting device and its ability to capture audio. Simultaneously, spatial audio formats that may be required by IVAS service requirements include low-resolution spatial, high-resolution spatial, metadata-assisted spatial audio (MASA) formats, and Higher Order Ambisonics ("HOA") transport format (HTF), or yet another spatial audio format. Thus, the acoustic preprocessing unit 220 of a transmitting device with spatial audio capabilities must be prepared to provide spatial audio signals in a suitable format that meets these requirements.
[0038] Low-resolution spatial formats include spatial WXY, primary ambisonics ("FOA"), and other formats. The spatial WXY format relates to a three-channel primary planar B-format audio representation with the height component (Z) omitted. This format is useful for bitrate-efficient immersive phone and conference scenarios where spatial resolution requirements are not very high and the spatial height component is considered insignificant. This format is particularly useful for conference calls as it allows the receiving client to perform immersive rendering of a conference scene captured in a conference room with multiple participants. Similarly, this format is useful for conference servers that spatially position conference participants in a virtual conference room. In contrast, FOA includes the height component (Z) as a fourth component signal. FOA representations are significant for low-rate VR applications.
[0039] High-resolution spatial formats include channel, object, and scene-based spatial formats. Depending on the number of audio component signals involved, each of these formats allows for the representation of spatial audio with virtually unlimited resolution. However, for various reasons (e.g., bitrate limitations and complexity limitations), they are practically limited to a relatively small number of component signals (e.g., 12). Further spatial formats may include, or rely on, the MASA or HTF formats.
[0040] Requiring IVAS-supporting devices to support the numerous diverse audio input formats mentioned above could result in substantial costs in terms of complexity, memory footprint, implementation testing, and maintenance. However, not all devices support all audio formats, nor do they benefit from supporting all of them. For example, some IVAS-enabled devices may support only stereo but not spatial capture. Other devices may support only low-resolution spatial input, and yet another class of devices may support only HOA capture. Thus, various devices will likely only utilize certain subsets of audio formats. Therefore, if an IVAS codec had to support direct encoding of all audio formats, it would become unnecessarily complex and expensive.
[0041] To solve this problem, the system 200 in Figure 2A includes a simplification unit 230. The acoustic preprocessing unit 220 transfers the audio signal to the simplification unit 130. In some implementations, the acoustic preprocessing unit 220 generates acoustic metadata that is transferred to the simplification unit 230 along with the audio signal. The acoustic metadata may include data related to the audio signal (e.g., format metadata such as mono, stereo, spatial). The acoustic metadata may also include noise cancellation data and other preferred data, such as data related to the physical or geometric characteristics of the capture unit 210.
[0042] The simplification unit 230 converts the various input formats supported by the device into a reduced common set of codec intake formats. For example, the IVAS codec can support three intake formats (mono, stereo, and spatial). The mono and stereo formats are similar to or identical to the respective formats produced by the acoustic preprocessing unit, while the spatial format may be a “mezzanine” format. The mezzanine format is a format that can accurately represent any of the aforementioned spatial audio signals obtained from the acoustic preprocessing unit 220. This includes spatial audio represented in any channel, object, and scene-based format (or a combination thereof). In some implementations, the mezzanine format can represent the audio signal as several objects in an audio scene and several channels to carry spatial information about that audio scene. Furthermore, the mezzanine format can represent MASA, HTF, or other spatial audio formats. One preferred spatial mezzanine format can represent spatial audio as m objects and an nth-order HOA ("mObj+HOAn"). Here, m and n are small integers, including zero.
[0043] Process 300 in Figure 3 illustrates exemplary actions for converting audio data from a first format to a second format. In 302, the simplification unit 230 receives the audio signal from, for example, the acoustic pre-processing unit 220. As described above, the audio signal received from the acoustic pre-processing unit 220 may be a signal that has undergone noise and echo cancellation processing, and channel downmixing and upmixing processing to reduce or increase the number of audio channels, for example. In some implementations, the simplification unit 230 receives acoustic metadata along with the audio signal. The acoustic metadata may include format instructions and other information as described above.
[0044] In 304, the simplification unit 230 determines whether the audio signal is in a first format supported or unsupported by the audio device's encoding unit 240. For example, the audio format detection unit 232 can analyze the audio signal received from the acoustic preprocessing unit 220 and identify the format of the audio signal, as shown in Figure 2A. If the audio format detection unit 232 determines that the audio signal is in mono or stereo format, the simplification unit 230 passes the signal to the encoding unit 240. However, if the audio format detection unit 232 determines that the signal is in spatial format, the audio format detection unit 232 passes the audio signal to the conversion unit 234. In some implementations, the audio format detection unit 232 can use acoustic metadata to determine the format of the audio signal.
[0045] In some implementations, the simplification unit 230 determines whether an audio signal is in a first format by determining the number, configuration, or location of audio capture devices (e.g., microphones) used to capture the audio signal. For example, if the audio format detection unit 232 determines that the audio signal was captured by a single capture device (e.g., a single microphone), the audio format detection unit 232 can determine that it is a monaural signal. If the audio format detection unit 232 determines that the audio signal was captured by two capture devices positioned at a specific angle to each other, the audio format detection unit 232 can determine that the signal is a stereo signal.
[0046] Figure 4 is a flowchart illustrating exemplary actions for determining whether an audio signal is in a format supported by an encoding unit, according to some embodiments of the present disclosure. In 402, a simplification unit 230 accesses the audio signal. For example, an audio format detection unit 232 may receive the audio signal as input. In 404, the simplification unit 230 determines the acoustic capture configuration of the audio equipment used to capture the audio signal, e.g., the number and position configuration of microphones. For example, the audio format detection unit 232 may analyze the audio signal and determine that three microphones were positioned at different locations in space. In some implementations, the audio format detection unit 232 may use acoustic metadata to determine the acoustic capture configuration. That is, an acoustic preprocessing unit 220 may generate acoustic metadata indicating the position and number of each capture device. The metadata may also include a description of the detected audio characteristics, such as the direction and directivity of the sound source. In 406, the simplification unit 230 compares the acoustic capture configuration with one or more stored acoustic capture configurations. For example, a stored acoustic capture configuration may include the number and position of each microphone to identify a particular configuration (e.g., mono, stereo, or spatial). The simplification unit 230 compares each of those acoustic capture configurations with the acoustic capture configuration of the audio signal.
[0047] In step 408, the simplification unit 230 determines whether the acoustic capture configuration matches a stored acoustic capture configuration associated with the spatial format. For example, the simplification unit 230 may determine the number of microphones used to capture the audio signal and their positions in space. The simplification unit 230 can compare this data with a stored known configuration for the spatial format. If the simplification unit 230 determines that there is no match for the spatial format, this may indicate that the audio format is mono or stereo, and process 400 proceeds to 412, where the simplification unit 230 transfers the audio signal to the encoding unit 240. However, if the simplification unit 230 identifies the audio format as belonging to a set of spatial formats, process 400 proceeds to 410, where the simplification unit 230 converts the audio signal to a mezzanine format.
[0048] Referring again to Figure 3, in 306, the simplification unit 230 determines that the audio signal is in a format not supported by the encoding unit and converts the audio signal to a second format supported by the encoding unit. For example, the conversion unit 234 may convert the audio signal to a mezzanine format. A mezzanine format accurately represents a spatial audio signal originally represented in any channel, object, and scene-based format (or a combination thereof). Furthermore, a mezzanine format can represent MASA, HTF, or other preferred formats. For example, a format that can function as a spatial mezzanine format can represent audio as m objects and n-th order HOA ("mObj+HOAn"), where m and n are small integers including zero. Thus, a mezzanine format may also involve representing the audio in waveforms (signals) and metadata that can capture the explicit characteristics of the audio signal.
[0049] In some implementations, the conversion unit 234 generates metadata about the audio signal when converting it to a second format. The metadata may be associated with object metadata, which includes parts of the audio signal in the second format, such as the locations of one or more objects. Another example is when the audio is captured using its own set of capture devices, and the number and configuration of the devices are not supported or efficiently represented by the encoding unit and / or the mezzanine format. In such cases, the conversion unit 234 can generate metadata. The metadata may include at least one of conversion metadata or acoustic metadata. The conversion metadata may include a subset of metadata associated with parts of the format that are not supported by the encoding process and / or the mezzanine format. For example, when the audio signal is played back on a system configured to specifically output audio captured by its own configuration, the conversion metadata may include device settings for the capture (e.g., microphone) configuration and / or device settings for the output device (e.g., speaker) configuration. Metadata originating from either the acoustic preprocessing unit 220 and / or the conversion unit 234 may also include acoustic metadata, which describes certain audio signal characteristics such as the spatial direction from which the captured sound arrives, and the directivity or diffusion of the sound. In this example, the audio may be represented as a mono or stereo signal with additional metadata, but there may be a determination that it is spatial in a spatial format. In this case, the mono or stereo signal and metadata are propagated to the encoder 240.
[0050] In 308, the simplification unit 230 transfers the audio signal in the second format to the encoding unit. As shown in Figure 2A, if the audio format detection unit 232 determines that the audio is in a mono or stereo format, the audio format detection unit 232 transfers the audio signal to the encoding unit. However, if the audio format detection unit 232 determines that the audio signal is in a spatial format, the audio format detection unit 232 transfers the audio signal to the conversion unit 234. The conversion unit 234 converts the spatial audio to, for example, a mezzanine format, and then transfers the audio signal to the encoding unit 240. In some implementations, the conversion unit 234 also transfers conversion metadata and acoustic metadata to the encoding unit 240 in addition to the audio signal.
[0051] The encoding unit 240 receives an audio signal in a second format (e.g., mezzanine format) and encodes the audio signal in the second format into a transport format. The encoding unit 240 propagates the encoded audio signal to some transmitting entity that sends it to a second device. In some implementations, the encoding unit 240 or a subsequent entity stores the encoded audio signal for later transmission. The encoding unit 240 can receive audio signals in mono, stereo, or mezzanine format and encode those signals for audio transport. If the audio signal is in mezzanine format and the encoding unit receives conversion metadata and / or acoustic metadata from the simplification unit 230, the encoding unit transfers the conversion metadata and / or acoustic metadata to the second device. In some implementations, the encoding unit 240 encodes the conversion metadata and / or acoustic metadata into a specific signal that the second device can receive and decode. The encoding unit then outputs the encoded audio signal to an audio transport that carries it to one or more other devices. Thus, each device (for example, among the devices in Figure 1) can encode audio signals in a second format (for example, mezzanine format), but these devices generally cannot encode audio signals in a first format.
[0052] In one embodiment, the encoding unit 240 (for example, the IVAS codec described above) operates on a mono, stereo, or spatial audio signal provided by the simplification stage. Encoding is performed depending on a codec mode selection, which may be based on one or more of the negotiated IVAS service level, the capabilities of the transmitting and receiving devices, and the available bitrate.
[0053] The service level may include, for example, IVAS Stereo Telephony, IVAS Immersive Conferencing, IVAS User-Generated VR Streaming, or other suitable service levels. A certain audio format (mono, stereo, spatial) can be assigned to a specific IVAS service level in which a preferred mode of IVAS codec operation is selected.
[0054] Furthermore, the operating mode of the IVAS codec can be selected in response to the capabilities of the transmitting and receiving devices. For example, depending on the capabilities of the transmitting device, the encoding unit 240 may not be able to access spatially ingested signals, for example, because the encoding unit 240 is only providing mono or stereo signals. In addition, an end-to-end capability exchange or corresponding codec mode request may indicate that the receiving end has certain rendering limitations and does not need to encode and transmit spatial audio signals, or vice versa. In another example, another device may request spatial audio.
[0055] In some implementations, end-to-end capability exchange cannot fully resolve remote device capabilities. For example, the encoding point may not have information about whether the decoding unit (sometimes called a decoder) is for a single mono speaker, a stereo speaker, or whether it is rendering binaurally. The actual rendering scenario may change during a service session. For example, if the connected playback device changes, the rendering scenario may change. In one example, end-to-end capability exchange may not occur because the sink device is not connected during an IVAS encoding session. This may occur with voicemail services or (user-generated) virtual reality content streaming services. Another example where the capabilities of the receiving device are unknown or cannot be resolved due to ambiguity is a single encoder that needs to support multiple endpoints. For example, in IVAS conferencing or virtual reality content distribution, one endpoint may use a headset and another endpoint may render to stereo speakers.
[0056] One way to address this problem is to assume the lowest possible receiving device capability and select a corresponding IVAS codec operating mode, which in some cases may be monaural. Another way to address this problem is to require the IVAS decoder to derive a decoded audio signal that can be rendered by devices with the respective lower audio capabilities, even if the encoder is operating in a mode that supports spatial audio or stereo audio. That is, a signal encoded as spatial audio should be decodeable for both stereo and mono rendering. Similarly, a signal encoded as stereo should be decodeable for mono rendering.
[0057] For example, in an IVAS conference, the call server should only need to perform a single encoding and send the same encoding to multiple endpoints. Some of these endpoints may be binaural, while others may be stereo. Thus, a single two-channel encoding can support both, for example, rendering on a laptop 114 and conference room system 118 with stereo speakers, and immersive rendering with binaural presentation on user device 110 and virtual reality gear 122. Thus, a single encoding can support both outcomes simultaneously. As a result, one implication is that this two-channel encoding supports both stereo speaker playback and binaurally rendered playback with a single encoding.
[0058] Another example involves high-quality mono extraction. This system can support the extraction of high-quality mono signals from encoded spatial or stereo audio signals. In some implementations, it is possible to extract an Enhanced Audio Services ("EVS") codec bitstream for mono decoding, for example, using a standard EVS decoder.
[0059] In addition to service levels and device capabilities, the available bitrate is another parameter that can control codec mode selection. In some implementations, the bitrate needs to increase along with the quality of experience that can be provided at the receiving end and the number of relevant components of the audio signal. At the lowest bitrate, only mono audio rendering is possible. The EVS codec offers mono operation down to 5.9 kilobits / second. As the bitrate increases, higher quality service can be achieved. However, the Quality of Encoding (QoE) remains limited to mono-only operation and rendering. The next higher level of QoE is possible with (conventional) 2-channel stereo. However, since this system has two audio signal components to be transmitted, it requires a bitrate higher than the lowest mono bitrate to provide useful quality. Spatial sound experiences require a higher QoE than stereo. At the lower end of the bitrate range, this experience can be made possible with a binaural representation of a spatial signal, which may be called "Spatial Stereo". Spatial stereo is likely to be the most compact spatial representation because it relies on encoder-side binaural pre-rendering (with an appropriate head-related transfer function (HRTF)) to an encoder (e.g., encoding unit 240) of the spatial audio signal intake, and consists of only two audio component signals. Because spatial stereo carries more perceptual information, the bitrate required to achieve sufficient quality is likely to be higher than that required for conventional stereo signals. However, spatial stereo representations may have limitations in relation to the customization of rendering at the receiving end. These limitations may include constraints on headphone rendering, the use of a pre-selected set of HRTFs, or rendering without head tracking.Higher QoE at higher bitrates is enabled not by relying on binaural pre-rendering in the encoder, but rather by a codec mode for encoding audio signals in a spatial format that represents the ingested spatial mezzanine format. The number of audio component signals represented by that format can be adjusted depending on the bitrate. For example, this can result in somewhat powerful spatial representation, ranging from spatial WXY to high-resolution spatial audio formats, as described above. This allows for low to high spatial resolution depending on the available bitrate, providing flexibility to handle a wide range of rendering scenarios, including binaural with head tracking. This mode is referred to as the "Versatile Spatial" mode.
[0060] In some implementations, the IVAS codec operates within the bitrate range of the EVS codec, i.e., 5.9 to 128 kilobits per second. For low-rate stereo operation in bandwidth-constrained environments, bitrates as low as 13.2 kbps may be required. This requirement may depend on the technical feasibility of using a particular IVAS codec, potentially enabling attractive IVAS service operation. For low-rate spatial stereo operation in bandwidth-constrained environments, the minimum bitrate enabling spatial rendering and simultaneous stereo rendering is possible down to 24.4 kilobits per second. For operation in versatile spatial mode, low spatial resolution (spatial WXY, FOA) is probably possible down to 24.4 kilobits per second, but at this rate, audio quality can be achieved similarly to spatial stereo operation mode.
[0061] Referring to Figure 2B, the receiver receives an audio transport stream containing the encoded audio signal. The receiver's decoding unit 250 receives the encoded audio signal (for example, in the transport format encoded by the encoder) and decodes it. In some implementations, the decoding unit 250 receives the audio signal encoded in one of four modes: mono, (conventional) stereo, spatial stereo, or versatile spatial. The decoding unit 250 transfers the audio signal to the rendering unit 260. The rendering unit 260 receives the audio signal from the decoding unit 250 and renders the audio signal. It should be noted that, in general, there is no need to restore the original first spatial audio format that was captured by the simplification unit 230. This allows for significant savings in decoder complexity and / or memory footprint in IVAS decoder implementations.
[0062] Figure 5 is a flowchart illustrating exemplary actions for converting an audio signal to an available playback format according to some embodiments of the present disclosure. In 502, the rendering unit 260 receives an audio signal in a first format. For example, the rendering unit 260 may receive audio signals in mono, conventional stereo, spatial stereo, or versatile spatial formats. In some implementations, a mode selection unit 262 receives the audio signal. The mode selection unit 262 identifies the format of the audio signal. If the mode selection unit 262 determines that the format of the audio signal is supported by the playback configuration, the mode selection unit 262 transfers the audio signal to the renderer 264. However, if the mode selection unit determines that the audio signal is not supported, the mode selection unit performs further processing. In some implementations, the mode selection unit 262 selects a different decoding unit.
[0063] In 504, the rendering unit 260 determines whether the audio device can reproduce the audio signal in a second format supported by the playback configuration. For example, the rendering unit 260 may determine that the audio signal is in a spatial stereo format, but the audio device can only reproduce the received audio in mono (for example, based on the number and configuration of speakers and other output devices and / or metadata related to the decoded audio). In some implementations, not all devices in the system (for example, as shown in Figure 1) can reproduce the audio signal in the first format, but all devices can reproduce the audio signal in the second format.
[0064] In 506, the rendering unit 260 adapts the audio decoding to generate a signal in a second format, based on its determination that an output device can reproduce the audio signal in a second format. Alternatively, the rendering unit 260 (for example, a mode selection unit 262 or a renderer 264) may adapt the audio signal to a second format using metadata, such as acoustic metadata, conversion metadata, or a combination of acoustic and conversion metadata. In 508, the rendering unit 260 transfers the audio signal in either a supported first format or a supported second format for audio output (for example, to a driver interfacing with a speaker system).
[0065] In some implementations, the rendering unit 260 converts the audio signal to the second format by using metadata that includes representations of parts of the audio signal not supported by the second format, in combination with the audio signal in the first format. For example, if the audio signal is received in mono format and the metadata includes spatial format information, the rendering unit can use the metadata to convert the mono format audio signal to a spatial format.
[0066] Figure 6 is another block diagram of exemplary actions for converting an audio signal to an usable playback format according to some embodiments of the present disclosure. In 602, the rendering unit 260 receives an audio signal in a first format. For example, the rendering unit 260 may receive an audio signal in mono, conventional stereo, spatial stereo, or versatile spatial format. In some implementations, a mode selection unit 262 receives the audio signal. In 604, the rendering unit 260 acquires the audio output capability (e.g., audio playback capability) of an audio device. For example, the rendering unit 260 may acquire the speaker locations, their positional configurations, and / or the configurations of other playback devices available for playback. In some implementations, a mode selection unit 262 performs the acquisition operation.
[0067] In 606, the rendering unit 260 compares the audio characteristics of the first format with the output capability of the audio device. For example, the mode selection unit 262 can determine that the audio signal is in a spatial stereo format (for example, based on acoustic metadata, conversion metadata, or a combination of acoustic metadata and conversion metadata) and that the audio device can only reproduce the audio signal in a conventional stereo format on a stereo speaker system (for example, based on the speaker and other output device configuration). The rendering unit 260 can compare the audio characteristics of the first format with the output capability of the audio device. In 608, the rendering unit 260 determines whether the output capability of the audio device matches the audio output characteristics of the first format. If the output capability of the audio device does not match the audio characteristics of the first format, process 600 proceeds to 610, where the rendering unit 260 (for example, the mode selection unit 262) takes action to obtain the audio signal in a second format. For example, the rendering unit 260 may adapt the decoding unit 250 to decode the received audio in a second format, or the rendering unit may use acoustic metadata, conversion metadata, or a combination of acoustic and conversion metadata to convert the audio from a spatial stereo format to a supported second format, the second format being conventional stereo in the given example. If the output capability of the audio device matches the audio output characteristics of the first format, or after the conversion operation 610, process 600 proceeds to 612, in which the rendering unit 260 (for example, using renderer 264) transfers the audio signal, which is now guaranteed to be supported, to the output device.
[0068] Figure 7 shows a block diagram of an exemplary system 700 suitable for carrying out exemplary embodiments of the present disclosure. As shown, the system 700 includes a central processing unit (CPU) 701 capable of executing various processes according to, for example, a program stored in a read-only memory (ROM) 702, or a program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 also stores data as needed when the CPU 701 executes various processes. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output interface (I / O) 705 is also connected to the bus 704.
[0069] The following components are connected to the I / O interface 705: an input unit 706 which may include a keyboard, mouse, etc.; an output unit 707 which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 708 which includes a hard disk or another suitable storage device; and a communication unit 709 which includes a network interface card such as a network card (e.g., wired or wireless).
[0070] In some implementations, the input unit 706 includes one or more microphones at different locations (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other preferred formats).
[0071] In some implementations, the output unit 707 includes a system with varying numbers of speakers. As shown in Figure 1, the output unit 707 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other preferred formats).
[0072] The communication unit 709 is configured to communicate with other devices (for example, via a network). If necessary, the drive 710 is also connected to the I / O interface 705. A removable medium 711, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is mounted on the drive 710, and if necessary, a computer program read from there is installed in the storage unit 708. Those skilled in the art will understand that although the system 700 is described as including the components described above, in actual applications it is possible to add, remove, and / or replace some of these components, and all such modifications or changes are within the scope of this disclosure.
[0073] According to exemplary embodiments of the present disclosure, the processes described above may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the method. In such embodiments, the computer program may be downloaded from a network via a communication unit 709, mounted, and / or installed from a removable medium 711.
[0074] In general, various exemplary embodiments of this disclosure may be implemented in hardware or special-purpose circuits (e.g., control circuits), software, logic, or any combination thereof. For example, the simplification unit 230 and the other units described above may be performed by a control circuit (e.g., a CPU in combination with the other components of Figure 7), and thus the control circuit may perform the actions described in this disclosure. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be performed by a controller, microprocessor, or other computing device (e.g., a control circuit). Although various aspects of the exemplary embodiments of this disclosure are illustrated and described using block diagrams, flowcharts, or any other pictorial representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware, or controllers, or other computing devices, or any combination thereof, as an example without limitation.
[0075] Furthermore, the various blocks shown in the flowchart can be viewed as method steps and / or as operations resulting from the operation of computer program code and / or as a plurality of coupled logic circuit elements constructed to perform associated functions. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program comprising program code configured to perform the methods described above.
[0076] In the context of this disclosure, a machine-readable medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-temporary and may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections having one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0077] Computer program code for performing the methods of this disclosure can be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, a dedicated computer, or another programmable data processing device having a control circuit, and when executed by the computer or other programmable data processing device processor, the program code will implement the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer, partially on a remote computer, entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0078] Several aspects are described below. [Aspect 1] A step in which a simplified unit of an audio device receives an audio signal of a first format, wherein the first format is one of a set of multiple audio formats supported by the audio device; The simplification unit includes the step of determining whether the first format is supported by the encoder of the audio device; A step in which, based on the fact that the first format is not supported by the encoder, the simplification unit converts the audio signal to a second format supported by the encoder, wherein the second format is an alternative representation of the first format; The simplification unit transfers the audio signal in the second format to the encoder; The steps include: encoding the audio signal using the encoder; The process includes the steps of storing the encoded audio signal or transmitting the encoded audio signal to one or more other devices. method. [Aspect 2] The method according to embodiment 1, wherein converting the audio signal to a second format includes generating metadata about the audio signal, the metadata including a representation of a portion of the audio signal. [Aspect 3] The method according to embodiment 1, wherein encoding the audio signal includes encoding the audio signal in the second format into a transport format supported by the second device. [Aspect 4] The method according to embodiment 3, further comprising transmitting the encoded audio signal by transmitting the metadata, which includes a representation of the portion of the audio signal not supported by the second format. [Aspect 5] The method according to embodiment 1, wherein determining whether the audio signal is in the first format by the simplification unit includes determining the number of audio capture devices and the corresponding location of each capture device used to capture the audio signal. [Aspect 6] The method according to embodiment 1, wherein each of the one or more other devices is configured to reproduce the audio signal from the second format, and at least one of the one or more other devices is unable to reproduce the audio signal from the first format. [Aspect 7] The method according to embodiment 1, wherein the second format represents the audio signal as several audio objects in an audio scene, each of which relies on several audio channels to carry spatial information. [Aspect 8] The method according to embodiment 1, wherein the second format further includes metadata for carrying further portions of spatial information. [Aspect 9] The method according to embodiment 1, wherein both the first format and the second format are spatial audio formats. [Aspect 10] The method according to aspect 1, wherein the second format is a spatial audio format, and the first format is a mono format associated with metadata, or a stereo format associated with metadata. [Aspect 11] The method according to any one of embodiments 1 to 10, wherein the set of multiple audio formats supported by the audio device includes multiple spatial audio formats. [Aspect 12] The method according to any one of embodiments 1 to 11, wherein the second format is an alternative representation of the first format and further features enabling an equivalent degree of empirical quality. [Aspect 13] The rendering unit of the audio device receives an audio signal in a first format; The rendering unit determines whether the audio device can reproduce the audio signal in the first format; In response to the audio device determining that it cannot reproduce the audio signal in the first format, the rendering unit adapts the audio signal so that it can be used in the second format; The rendering unit includes the step of transferring the audio signal in the second format for rendering, method. [Aspect 14] The method according to embodiment 13, wherein the rendering unit converts the audio signal to a second format, and includes using metadata, which includes representations of portions of the audio signal not supported by the fourth format used for encoding, in combination with the audio signal in the third format. [Aspect 15] The decoding unit receives the audio signal in transport format; The steps include: decoding the audio signal of the transport format into the first format; The process further includes the step of transferring the audio signal in the first format to the rendering unit, The method described in Embodiment 13. [Aspect 16] The method according to aspect 15, wherein adapting the audio signal to be available in the second format includes adapting the decoding to produce the received audio in the second format. [Aspect 17] The method according to embodiment 13, wherein each of the multiple devices is configured to reproduce the audio signal in the second format, and one or more of the multiple devices are unable to reproduce the audio signal in the first format. [Aspect 18] The simplification unit receives various audio signals in multiple formats from the audio pre-processing unit; The simplification unit receives attributes of the device from the device, the attributes of which include an indication of one or more audio formats supported by the device, the one or more audio formats of which include at least one of mono, stereo, or spatial formats; The simplification unit converts the audio signals into an intake format which is an alternative representation of one or more audio formats; A method comprising the step of providing the audio signal converted by the simplification unit to an encoding unit for downstream processing, Each of the aforementioned sound preprocessing unit, simplification unit, and encoding unit has one or more computer processors. method. [Aspect 19] One or more computer processors; The system comprises one or more non-temporary storage media that store instructions causing the one or more computer processors to perform the operations described in any one of the embodiments 1 to 18 when executed by the one or more computer processors, Device. [Aspect 20] A capture unit configured to capture an audio signal; An audio preprocessing unit configured to perform operations including preprocessing the aforementioned audio signal; Encoder and; An encoding system having a simplification unit, The aforementioned simplified unit is: A step of receiving an audio signal of a first format from the sound preprocessing unit, wherein the first format is one of a set of multiple audio formats supported by the encoder; A step of determining whether the first format is supported by the encoder; The steps include: in response to determining that the first format is not supported by the encoder, converting the audio signal to a second format supported by the encoder; The system is configured to perform an operation that includes the step of transferring the audio signal in the second format to the encoder, The aforementioned encoder is: Encode the aforementioned audio signal; It is configured to perform operations including storing an encoded audio signal or transmitting an encoded audio signal to another device. Encoding system. [Aspect 21] The encoding system according to embodiment 20, wherein converting the audio signal to a second format includes generating metadata for the audio signal, the metadata including a representation of the portion of the audio signal not supported by the second format. [Aspect 22] The encoding system according to embodiment 20, wherein the operation of the encoder further includes transmitting the encoded audio signal by transmitting the metadata, which includes a representation of the portion of the audio signal not supported by the second format. [Aspect 23] The encoding system according to embodiment 20, wherein the second format represents the audio signal audio as several channels for carrying some objects and spatial information in an audio scene. [Aspect 24] Preprocessing the aforementioned audio signal: Perform noise cancellation; Perform echo cancellation; To reduce the number of channels in the aforementioned audio signal; Increasing the number of audio channels in the aforementioned audio signal; or Including one or more of the following: generating acoustic metadata, The encoding system described in embodiment 20. [Aspect 25] It is a decoding system: A decoder configured to perform an operation including decoding an audio signal from a transport format to a first format; A rendering unit, The steps include receiving the audio signal in the first format; A step of determining whether an audio device can reproduce the audio signal in a second format, wherein the second format allows the use of more output devices than the first format; The steps include: converting the audio signal to the second format in response to the audio device determining that it can reproduce the audio signal in the second format; A rendering unit configured to perform an operation including the step of rendering the audio signal in the second format; A playback unit configured to perform operations including initiating playback of a rendered audio signal on a speaker system, Decoding system. [Aspect 26] The decoding system according to embodiment 25, wherein converting the audio signal to a second format includes using metadata, which includes representations of portions of the audio signal not supported by the fourth format used for encoding, in combination with the audio signal in the third format. [Aspect 27] The operation of the decoder described above is further: The audio signal in transport format is received; This includes transferring the audio signal in the first format to the rendering unit, The decoding system described in aspect 25.
Claims
[Claim 1] The simplification stage is a step of receiving audio signals in multiple formats and metadata for those audio signals from the sound preprocessing stage, wherein the audio signals represent audio captured by at least one microphone; The simplification stage includes a step of receiving attributes of the device from the device, wherein the attributes include one or more audio formats supported by the device, and the one or more audio formats include spatial formats; The simplification stage includes the step of converting the audio signal into a spatial mezzanine format compatible with the one or more audio formats; The simplification stage includes the step of providing the converted audio signal to the encoding stage, the output of which is for downstream processing in the apparatus. method.