Rendering an M-channel input to S speakers (S <M)
The audio renderer dynamically adjusts multi-channel audio output on portable devices with fewer speakers by using primary and secondary matrices and mixing gains, addressing unbalanced loudness and spatial collapse issues.
Patent Information
- Application Number
- JP2024079078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-17
- Filing Date
- 2024-05-15
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2040-06-17
AI Technical Summary
Existing rendering techniques for multi-channel audio on portable devices with fewer speakers than channels do not account for the time-varying behavior of input audio channels, leading to unbalanced loudness and spatial collapse of sound images.
An audio renderer that applies primary and secondary rendering matrices, calculates mixing gains based on time-varying channel distributions, and mixes pre-rendered signals to dynamically adjust audio output, ignoring inactive channels when necessary.
Enhances audio rendering efficiency and clarity by balancing loudness and maintaining distinct sound images across speakers, even with varying channel activity.
Smart Images

Figure 0007800998000028 
Figure 0007800998000029 
Figure 0007800998000030
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to PCT Application No. PCT / CN2019 / 092021, filed June 20, 2019, and U.S. Provisional Application No. 62 / 875,160, filed July 17, 2019, each of which is incorporated by reference in its entirety.
[0002] Technical Field The present invention relates to rendering an M channel input on S speakers, where S is less than M. [Background technology]
[0003] Portable devices such as mobile phones and tablet computers have become increasingly popular and are now very common. Such devices are often used for media playback, including movies and music from YouTube® or similar sources. To achieve an immersive listening experience, portable devices often include multiple independent speakers. For example, a tablet may include two top-layer speakers and two bottom-layer speakers. Furthermore, these devices typically include multiple independent power amplifiers (PAs) for their speakers to provide the device with flexibility for playback control.
[0004] At the same time, multi-channel audio content, i.e., content with more than two channels, e.g., 5.1, 5.1.2, is becoming more common. Multi-channel audio can be produced as is or can be converted from other formats, e.g., object-based audio, or by various upmixing methods.
[0005] There are different approaches to rendering multi-channel audio to portable devices with fewer speakers than the number of channels. One approach to rendering a 5.1.2 audio signal (8 channels) to a 4-speaker tablet is to render the height channels of the input signal to two upper speakers. To balance the reproduced sound with respect to the upper and lower speakers, the direct channels (i.e., non-height channels) are rendered to the two lower speakers. An example of such a rendering approach is provided in U.S. Pat. No. 6,233,599.
[0006] However, prior art rendering techniques do not take into account the time-varying behavior of the input audio channels. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] WO2017 / 165837 Summary of the Invention [Problem to be solved by the invention]
[0008] It is an object of the present invention to provide a more dynamic rendering technique based on the input audio. [Means for solving the problem]
[0009] According to a first aspect of the present invention, this and other objects are achieved by an audio renderer that renders a multi-channel audio signal having M channels to a portable device having S independent speakers, where S < M, and the audio renderer includes: a first matrix application module that applies a primary rendering matrix to an input audio signal to provide a first pre-rendered signal suitable for reproduction by the plurality of independent speakers; a second matrix application module that applies a secondary rendering matrix to the input audio signal to provide a second pre-rendered signal suitable for reproduction by the plurality of independent speakers; a channel analysis module configured to calculate a mixing gain according to a time-varying channel distribution; and a mixing module configured to generate a rendered output signal by mixing the first and second pre-rendered signals based on the mixing gain.
[0010] According to a second aspect of the present invention, this and other objects are achieved by a method of rendering a multi-channel audio signal having M channels to a portable device having S independent speakers, where S < M, and the method includes: applying a primary rendering matrix to an input audio signal to provide a first pre-rendered signal suitable for reproduction by the plurality of independent speakers; applying a secondary rendering matrix to the input audio signal to provide a second pre-rendered signal suitable for reproduction by the plurality of independent speakers; calculating a mixing gain according to a time-varying channel distribution; and mixing the first and second pre-rendered signals based on the mixing gain to generate a rendered output signal.
[0011] The present invention is based on the recognition that a multi-channel audio input can have a varying number of active channels. By providing several (at least two) different rendering matrices and selecting an appropriate mix of rendering matrices based on an analysis of the input signal, more efficient rendering on the available speakers can be achieved.
[0012] In extreme cases, the rendered output corresponds to one of the pre-rendered signals, in other cases it is a mix of both.
[0013] The secondary rendering matrix can be configured to ignore at least one of the channels in the input audio format. This may be appropriate when one or more channels of the input signal are relatively weak and therefore no longer contribute significantly to the rendered output. One example of a channel that may be weak for periods of time is a height channel, i.e., a channel intended for reproduction on (height) speakers located above the listener, or at least higher than other (direct) speakers.
[0014] A specific example relates to 5.1.2 audio, i.e., audio having left, right, center, left rear, right rear, LFE, and left / right height channels. For some periods, for example, the height channel may be relatively weak, in which case the 5.1.2 signal degenerates into a 5.1 signal, i.e., six channels, rather than eight channels. In that situation, the original rendering matrix (adapted for 5.1.2) may lead to unbalanced loudness between the upper and lower level speakers. In accordance with the present invention, rendering may be dynamically adjusted to focus on the currently active channel. Thus, in the given example, the input audio may be rendered using a rendering matrix adapted for 5.1 rather than a rendering matrix adapted for 5.1.2. The detailed description below provides more detailed examples of rendering matrices. [Brief explanation of the drawings]
[0015] The present invention will now be described in more detail with reference to the accompanying drawings showing presently preferred embodiments of the invention. [Figure 1] FIG. 1 is a block diagram of an audio renderer according to one embodiment of the present invention. [Figure 2] 1 is a flowchart of an embodiment of the present invention. [Figure 3] 1 a-b show two examples of four-speaker layouts for landscape orientation of a portable device, corresponding to top / bottom firing (a) and left / right firing (b). DETAILED DESCRIPTION OF THE INVENTION
[0016] Detailed Description of the Presently Preferred Embodiments The systems and methods disclosed below may be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks does not necessarily correspond to the division into physical units; conversely, one physical component may have multiple functions, and one task may be performed by multiple cooperating physical components. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those skilled in the art, the term “computer storage media” includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVDs), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by a computer. Additionally, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and is well known to those skilled in the art to include any information delivery media.
[0017] An embodiment of the present invention will now be discussed with reference to the block diagram of FIG. 1 and the flow chart of FIG.
[0018] This method is executed in real time. First, multi-channel input audio is received (e.g., decoded) in step S1, and a set of rendering matrices is generated based on the number M of received channels and the number S of available speakers in step S2. Each rendering matrix is configured to render M received signals to S speaker feeds. Here, S < M. In the illustrated example, the set includes a primary (default) matrix and a secondary (alternate) matrix, although one or more additional alternate matrices are possible. In step S3, each matrix is applied to the input signal by matrix application modules 11, 12 to generate pre-rendered signals for further mixing. In parallel step S4, the input audio is analyzed by a channel analysis module 13. In step S5, gains are calculated by the analysis module 13, for example based on the energy distribution between channels. This gain is further smoothed by a smoothing module 14 in step S6 and then input to a mixing module 15, which also receives the outputs from the matrix application modules 11, 12. In step S7, the mixing module 15 mixes (weights) the pre-rendered signals based on the smoothed gain and outputs a rendered audio signal. Details of the rendering process are described below.
[0019] Rendering Matrix When an M-channel input signal and an S-speaker device are provided, the general rendering process is represented by the following equation: y = Rx (1) Here, x is an M-dimensional vector representing the input signal, y is an S-dimensional vector representing the rendered signal, and R is an S×M rendering matrix. For the rendering matrix R, the rows correspond to the speakers and the columns correspond to the channels of the input signal. The entries of the rendering matrix indicate the mapping from channels to speakers.
[0020] When a portable device with S independent speakers (S > 2) is provided, the primary rendering matrix Rprim and the secondary rendering matrix R sec is determined according to the number of input channels M. R prim and R sec both have S×M of the same size. Specifically, the matrix R prim and R sec can be written as follows.
Equation
Equation
[0021] Generally, multi-channel audio usually includes four categories of channels: 1) Front channels, that is, left, right, and center channels (L, R, C) 2) Surround channels on the listener's plane, such as left / right surround (Ls / Rs) for 5.1 / 5.1.2 / 5.1.4, etc. or left / right surround (Lrs / Rrs) for 7.1 / 7.1.2 / 7.1.4, etc. 3) Height channels, such as left / right up (Lt / Rt) for 5.1.2 / 7.1.2 / 9.1.2, etc., left / right front / back (Ltf / Rtf, Ltr / Rtr) for 5.1.4 / 7.1.4 / 9.1.4, etc. 4) LFE channel.
[0022] Given a target speaker layout, the linear matrix defined in equation (2) can be rewritten as a block matrix:
number
[0023] quadratic matrix R sec is a function of R using one or more zero columns. prim can be derived from
[0024] Some more specific examples of rendering matrices according to embodiments of the present invention are described below.
[0025] Figures 3a and 3b show two examples of portable devices, here tablets in landscape orientation, equipped with multiple independently controlled speakers. In both examples, the devices have four speakers a through d (S=4). In Figure 1a, the speakers are located on the upper and lower sides of the device, thus including two speakers a and b that emit sound upward and two speakers c and d that emit sound downward. In Figure 1b, the speakers are located on the left and right sides of the device, thus including two upper speakers a and b that emit sound to the sides and two lower speakers c and d that emit sound to the sides.
[0026] In this example, a 5.1.2 channel audio signal (M=8) is played on a portable device as shown in FIG. 3a or 3b.
[0027] In this case, the linear matrix R prim can be defined by the following equation:
number
[0028] During periods when the height channel of the original 5.1.2 signal is nearly silent, the audio signal degenerates to a 5.1 signal plus two channels that can be ignored. Thus, the quadratic rendering matrix R sec1 can be defined by the following equation:
number
[0029] Multiple secondary rendering matrices R for a given device and input signal secx Note that in the above example of rendering 5.1.2 audio to four speakers, if the surround channels Ls and Rs are also mostly silent in addition to the height channel, the signal degenerates to a 3.1 signal containing only the C, L, R and LFE channels, plus a set of channels that can be ignored. In that case, the corresponding quadratic matrix R sec2 becomes:
number
[0030] In practice, when multiple quadratic matrices exist, the appropriate quadratic matrix is dynamically selected based on channel analysis as described below.
[0031] In addition to ensuring efficient rendering of the input signal, there is also the challenge of ensuring that all input channels (e.g., height channels) are clearly distinguishable after rendering. This is due to the small distance between speaker positions in portable devices. For example, height channels are likely to be rendered to speakers that are relatively close to the speakers for non-height channels. This leads to spatial collapse of the height sound image.
[0032] To mitigate the spatial collapse and make the height channel distinguishable after rendering, we use the rendering matrix R prim It is crucial to generate the correct entries for the height channels. Specifically, it is desirable to render most of the height channels to the top speakers, while rendering the front channels to the bottom speakers. This mitigates the "sinking" of the height channels into the front channels.
[0033] For the above example, see R prim The entries can be set as follows:
number
[0034] Alternatively, R prim The entries can be set as follows:
number
[0035] In both examples above, the columns (from left to right) correspond to channels L, R, C, LFE, Ls, Rs, Lt and Rt, respectively.
[0036] A first quadratic matrix R configured to ignore the two height channels Lt and Rt (columns 7 and 8) sec1 The entry can be set as follows:
number
[0037] The entries of the second quadratic matrix configured to ignore the two height channels Lt and Rt (columns 7 and 8) and the two surround channels Ls and Rs (columns 5 and 6) can be set as follows:
number
[0038] In another example, a 7.1.2 channel (M=10) input signal is reproduced by the device of Figure 3a or Figure 3b (S=4). In this case, R prim The entries can be set as follows:
number
[0039] quadratic matrix R sec1 and R sec2 The entries can be set as follows:
number
[0040] Rendering matrix R prim and R secx Note that the entries of can be real constants or frequency-dependent complex vectors. For example, R in Eq. prim The entries of can be expanded to a B-dimensional complex vector, where B is the number of frequency bands. In the use case mentioned above, to improve the height channel, R in equation (2) primSpecific frequency bands can be modified for the entries in the last two columns. An example of a specific frequency band can be 7 kHz to 9 kHz.
[0041] Also, as illustrated by the above example, R prim and R secx Note that at least some of the entries in the matrix can be set to be the same.
[0042] Channel Analysis The channel analysis module 23 is intended to determine whether the input signal is degenerate, so that a proper pre-rendered signal or a suitable mix of them can be used. The module 23 is executed for each frame.
[0043] One approach is based on the energy distribution among the input channels.
[0044] The aforementioned use case with only two different rendering matrices is taken as an example. For a 4-speaker portable device and a 5.1.2 input signal, the gain g raw is calculated by the following formula:
number
[0045] In addition to energy, diffuseness can be an alternative or additional criterion for analyzing the input channels. High diffuseness tends to result in disproportionate weighting of the L / R channels between the top and bottom speakers.
[0046] Adaptive Smoothing and Blending gain g rawcan be further smoothed by the smoothing module 14 according to the history of the input signal. For the current frame n (n>1), the smoothed gain can be calculated as follows:
number
[0047] The final rendering signal y is obtained by the blending process as follows:
number
[0048] If there are more than two different rendering matrices, the rendered output will contain a mix of three or more pre-rendered signals, depending on the channel analysis.
[0049] Conclusion As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object merely indicates that different instances of similar objects are being referred to and is not intended to imply that the objects so described must be in a given sequence, temporally, spatially, in ranking, or in any other way.
[0050] In the claims and the description herein, the terms "having," "comprising," and "including" are all open terms meaning the inclusion of at least the subsequent element / feature but not the exclusion of other elements. Thus, when used in the claims, the terms "having" and "comprising" should not be interpreted as being limited to the means, elements, or steps listed thereafter. For example, the scope of an expression "device including A and B" should not be limited to a device consisting only of elements A and B. As used herein, the terms "comprise," "comprising," and "including" are also all open terms meaning the inclusion of at least the subsequent element / feature but not the exclusion of other elements. Thus, when used in the claims, the terms "having" and "comprising" should not be interpreted as being limited to the means, elements, or steps listed thereafter. For example, the scope of an expression "device including A and B" should not be limited to a device consisting only of elements A and B. Thus, "comprising" is synonymous with "having" and means "having."
[0051] As used herein, the term "exemplary" is used in the sense of providing an example rather than indicating a characteristic; that is, an "exemplary embodiment" is an embodiment served as an example, not necessarily an embodiment of exemplary character.
[0052] In the foregoing description of exemplary embodiments of the invention, it should be understood that various features of the invention may be grouped together in a single embodiment, drawing, or description thereof for the purpose of improving the flow of the disclosure and aiding in understanding one or more of the various inventive aspects. This method of disclosure, however, should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0053] Furthermore, although some embodiments described herein include some features but not others included in other embodiments, it will be understood by those skilled in the art that combinations of features from different embodiments are intended to form different embodiments within the scope of the present invention. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0054] Furthermore, some of the embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or other means for performing functions. Thus, a processor with the necessary instructions for performing such a method or method element constitutes a means for performing the method or method element. Furthermore, the elements described herein of apparatus embodiments are examples of means for performing the functions performed by the elements for purposes of implementing the invention.
[0055] In the description provided herein, numerous specific details are set forth. However, it will be understood that embodiments of the present invention may be practiced without such specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this document.
[0056] Similarly, it should be noted that the term coupled, when used in the claims, should not be interpreted as being limited only to a direct connection. The terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms are not intended as synonyms for each other. Thus, the scope of the expression device A coupled to device B should not be limited to devices or systems in which the output of device A is directly connected to the input of device B, but rather means that a path exists between the output of device A and the input of device B, and that the path may include other devices or means. "Coupled" may mean that two or more elements are in direct physical or electrical contact, or that two or more elements are not in direct contact with each other but still cooperate or interact with each other.
[0057] Thus, while particular embodiments of the present invention have been described, those skilled in the art will appreciate that other and further modifications may be made thereto without departing from the spirit of the invention, and that all such changes and modifications are intended to be claimed as falling within the scope of the present invention. For example, the formulas above merely represent procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.
[0058] Thus, while specific embodiments of the present invention have been described, those skilled in the art will appreciate that other and further modifications may be made thereto without departing from the spirit of the invention, and that all such changes and modifications are intended to be claimed as falling within the scope of the present invention. For example, the formulas above merely represent procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the present invention. For example, in the illustrated embodiment, the portable device has four speakers (S = 4). Of course, it is possible to have more (or fewer) than four speakers, resulting in different matrix sizes.
[0059] Some aspects will be described. [Aspect 1] An audio renderer that renders a multi-channel audio signal having M channels to a portable device having S independent speakers, where S < M, and the audio renderer: A first matrix application module that applies a primary rendering matrix to the input audio signal to provide a first pre-rendered signal suitable for reproduction by the plurality of independent speakers; A second matrix application module that applies a secondary rendering matrix to the input audio signal to provide a second pre-rendered signal suitable for reproduction by the plurality of independent speakers; A channel analysis module configured to calculate a mixing gain according to a time-varying channel distribution; A mixing module configured to generate a rendered output signal by mixing the first and second pre-rendered signals based on the mixing gain. Audio renderer. [Aspect 2] The audio renderer according to Aspect 1, wherein the secondary rendering matrix is configured to ignore at least one of the channels in the input audio signal. [Aspect 3] The audio renderer according to Aspect 2, wherein the input audio signal includes two height channels, and the secondary rendering matrix is configured to ignore the height channels. [Aspect 4] The input audio signal is a 5.1.2 audio signal having seven channels (M = 7), the number of independent speakers is four (S = 4), and the primary rendering matrix is
number
number
number
Number
number
number
Claims
1. 1. An audio renderer for rendering a multi-channel input audio signal having M channels to a portable device having S independent speakers, where S<M, the audio renderer comprising: a first matrix application module that applies a primary rendering matrix to the input audio signal to provide a first pre-rendered signal suitable for playback on the S independent speakers; a second matrix application module that applies a secondary rendering matrix to the input audio signal to provide a second pre-rendered signal suitable for playback on the S independent speakers; a channel analysis module configured to calculate a mixing gain according to a time-varying channel distribution; a mixing module configured to generate a rendered output signal by mixing the first and second pre-rendered signals based on the mixing gain; the channel analysis module determines the mix gain based on an energy distribution among channels of the input audio signal and / or a diffuseness of the input audio signal. Audio renderer.
2. The audio renderer of claim 1 , wherein the secondary rendering matrix is configured to ignore at least one channel in the input audio signal.
3. 3. The audio renderer of claim 2, wherein the input audio signal includes two height channels, and the secondary rendering matrix is configured to ignore the height channels.
4. The input audio signal is a 5.1.2 audio signal with eight channels (M=8), the number of independent speakers is four (S=4), and the primary rendering matrix is: [0016] 4. An audio renderer according to claim 1, wherein the audio renderer is set to:
5. The input audio signal is a 5.1.2 audio signal with eight channels (M=8), the number of independent speakers is four (S=4), and the primary rendering matrix is: [Equation 17] 4. An audio renderer according to claim 1, wherein the audio renderer is set to:
6. The input audio signal is a 5.1.2 audio signal with eight channels (M=8), the number of independent speakers is four (S=4), and the secondary rendering matrix is [Equation 18] 6. An audio renderer according to claim 1, wherein the audio renderer is set to:
7. 7. An audio renderer according to any one of claims 1 to 6, further comprising a smoothing module for smoothing blend gains for a current frame based on blending gains for a set of previous frames.
8. 8. An audio renderer according to any one of claims 1 to 7, wherein the entries of the primary and secondary rendering matrices are real constants or frequency-dependent complex vectors.
9. 9. An audio renderer according to any one of claims 1 to 8, wherein at least some entries of the primary rendering matrix are complex vectors of dimension B, where B is the number of frequency bands.
10. 10. The audio renderer of claim 9, wherein the frequency band ranges from 7 kHz to 9 kHz.
11. 11. An audio renderer according to any one of claims 1 to 10, wherein at least some entries of the primary rendering matrix and the secondary rendering matrix are equal.
12. 1. A method for rendering a multi-channel input audio signal having M channels for a portable device having S independent speakers, where S<M, the method comprising: applying a primary rendering matrix to the input audio signal to provide a first pre-rendered signal suitable for playback on the S independent speakers; applying a secondary rendering matrix to the input audio signal to provide a second pre-rendered signal suitable for playback on the S independent speakers; calculating a mixing gain according to a time-varying channel distribution; mixing the first and second pre-rendered signals based on the mixing gain to generate a rendered output signal; the mixing gain is calculated based on an energy distribution among the channels of the input audio signal, a diffuseness of the input audio signal, or both; method.
13. The method of claim 12 , wherein the secondary rendering matrix is configured to ignore at least one of the channels of the input audio signal.
14. The method of claim 13 , wherein the input audio signal includes two height channels, and the secondary rendering matrix is configured to ignore the two height channels.
15. The input audio signal is a 5.1.2 audio signal with eight channels (M=8), the number of independent speakers is four (S=4), and the primary rendering matrix is: [Equation 19] 15. The method according to claim 12, wherein the parameter is set to:
16. The input audio signal is a 5.1.2 audio signal with eight channels (M=8), the number of independent speakers is four (S=4), and the primary rendering matrix is: [Equation 20] 15. The method according to claim 12, wherein the parameter is set to:
17. The input audio signal is a 5.1.2 audio signal with eight channels (M=8), the number of independent speakers is four (S=4), and the primary rendering matrix is: [0000] 15. The method according to claim 12, wherein the parameter is set to:
18. 15. The method of any one of claims 12 to 14, further comprising smoothing the blended gains for a current frame based on blending gains for a set of previous frames.
19. A non-transitory computer readable medium comprising computer program code portions configured to perform the steps of any one of claims 12 to 18 when executed on a processor.
Citation Information
Patent Citations
Apparatus and method for generating an audio output signal using object-based metadata.
JP2011528200A
Three-dimensional acoustic reproduction device and program
JP2016100877A
Near-field rendering of immersive audio content in portable computers and devices
WO2017165837A1