Render an audio object with an apparent size

By using a thicker virtual sound source grid to render audio objects, the problem of high computational complexity of large audio objects processing in the prior art is solved, and the resource consumption is reduced and the audio playback quality is maintained.

CN110603821BActive Publication Date: 2025-06-24DOLBY INTERNATIONAL AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201880029053.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-07-05
Filing Date
2018-05-01
Publication Date
2025-06-24
Estimated Expiration
2038-05-01

AI Technical Summary

Technical Problem

The existing audio rendering technology has high computational complexity when processing large audio objects, especially in low-power embedded systems, resulting in excessive resource consumption.

Method used

Reduce the number of virtual sound sources by using a thicker, lower density virtual sound source grid to render audio objects, reducing the number of virtual sound sources, thereby reducing computational complexity and storage requirements.

Benefits of technology

The same audio playback effect as the traditional high-density virtual sound source grid is achieved, but the computing complexity and storage requirements are greatly reduced, reducing system cost and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110603821B_ABST
    Figure CN110603821B_ABST
Patent Text Reader

Abstract

Methods, systems, and computer program products are disclosed for rendering an audio object having an apparent size. An audio processing system receives audio panning data that includes a first grid that maps a first virtual sound source and speaker positions in space to speaker gains. The first grid specifies a first speaker gain for the first virtual sound source in the space. The audio processing system determines a second grid for a second virtual sound source in the space, including the second virtual sound source that maps the first virtual sound source to the second virtual source. The audio processing system selects at least one of the first grid or the second grid for rendering the audio object based on the apparent size of the audio object. The audio processing system renders the audio object based on the selected one or more grids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to audio playback systems.

[0002] Cross - Reference to Related Applications

[0003] This application claims priority to the following priority applications: Spanish application P201730658 filed on May 4, 2017 (our reference number: D16134ES), US Provisional Application 62 / 528,798 filed on July 5, 2017 (reference number: D16134USP1#), and EP application 17179710.3 filed on July 5, 2017 (reference number: D16134EP), which are incorporated herein by reference. Background Art

[0004] Modern audio processing systems can be configured to render one or more audio objects. An audio object can include a stream of audio signals associated with metadata. The metadata can indicate the position and apparent size of the audio object. Apparent size refers to the spatial size of the sound that a listener should perceive when the audio object is rendered in a reproduction environment. The rendering can include calculating a set of audio object gain values for each channel in a set of output channels. Each output channel can correspond to a playback device, e.g., a speaker.

[0005] Audio objects can be generated without reference to any particular reproduction environment. The audio processing system can render audio objects in a reproduction environment with a multi - step process that includes a setup process and a runtime process. During the setup process, the audio processing system can define a plurality of virtual sound sources in a space: the audio objects are located within the space and can move within the space. The virtual sound sources correspond to the positions of static point sources. The setup process receives speaker layout data. Speaker layout data refers to the positions of some or all of the speakers in the reproduction environment. The setup process calculates speaker gain values for each virtual sound source of each speaker based on the speaker positions and the virtual source positions. At runtime when rendering an audio object, the runtime process calculates the contribution of one or more virtual sound sources located within the area or volume defined by the audio object position and the audio object apparent size. Thus, the runtime process represents the audio object with the one or more virtual sound sources and outputs speaker gains for the audio object. Summary of the Invention

[0006] Techniques for rendering an audio object having an apparent size are described. An audio processing system receives audio panning data that includes a first grid that maps a first virtual sound source and speaker positions in space to speaker gains. The first grid specifies a first speaker gain for the first virtual sound source in the space. The audio processing system determines a second grid for a second virtual sound source in the space, including mapping the first speaker gain to a second speaker gain for the second virtual source. The first grid is denser than the second grid in terms of the number of virtual sound sources. The audio processing system selects at least one of the first grid or the second grid to render the audio object, the selection being based on the apparent size of the audio object. The audio processing system renders the audio object based on the selected grid, including representing the audio object using one or more virtual sound sources in the selected grid that are enclosed within a volume or area having the apparent size.

[0007] Compared with conventional audio rendering techniques for reproducing three-dimensional sound effects, the features described in this specification can obtain one or more advantages. For example, the disclosed techniques reduce the computational complexity of audio rendering. Conventional systems use many virtual sound sources to represent large audio objects. When processing the size of large audio objects, conventional systems need to consider these many virtual sound sources simultaneously. Simultaneous computation can be challenging, especially in low-power embedded systems. For example, a grid can have a size of 11×11×11 virtual sound sources. For an audio object whose size spans the entire listening area (which is not common), a conventional rendering system needs to consider 1331 virtual sound sources simultaneously and add them together. By generating a coarser, lower-density grid of virtual sources, the disclosed techniques can obtain results that are roughly the same as those produced by a conventional higher-density grid of virtual sound sources, but with much lower computational complexity. For example, by using a coarse grid with a size of 7×7×7 virtual sound sources, an audio rendering system using the disclosed techniques needs at most 343 virtual sound sources and uses approximately 26% of the memory of a conventional system using an 11×11×11 grid. An audio rendering system using a 5×5×5 coarse grid uses approximately 9% of the memory. An audio rendering system using a 3×3×3 coarse grid uses only approximately 2% of the memory. The reduced memory requirements can reduce system cost and power consumption without sacrificing playback quality.

[0008] Details of one or more embodiments of the disclosed subject matter are set forth in the following figures and description. Based on the specification, figures, and claims, other features, aspects, and advantages of the disclosed subject matter will become apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1is a block diagram showing an example audio processing system that implements coarse grid rendering.

[0010] Figure 2 is a schematic diagram showing an example audio object associated with a corresponding apparent size.

[0011] Figure 3 is a schematic diagram showing an example technique of a unit that creates fine virtual sound sources.

[0012] Figure 4 is a schematic diagram showing an example technique of reducing the number of virtual sound sources.

[0013] Figure 5 is a schematic diagram showing an example technique of a unit that creates coarse virtual sound sources.

[0014] Figure 6 is a schematic diagram showing an example technique of mapping fine virtual sound sources to coarse virtual sound sources when determining speaker gain.

[0015] Figure 7 is a schematic diagram showing an example technique of reducing the number of virtual sound sources for a large audio object.

[0016] Figure 8 is a flowchart of an example process of rendering an audio object with an apparent size.

[0017] Figure 9 is to implement the reference Figures 1 to 8 is a block diagram of an example system architecture of an audio rendering system that performs the features and operations described.

[0018] Similar reference numerals in the various figures indicate similar elements. Detailed Description

[0019] Rendering an Audio Object Using Coarse Grid

[0020] Figure 1is a block diagram showing an example audio processing system 100 that implements coarse grid rendering. The audio processing system 100 includes a grid mapper 102. The grid mapper 102 is a component of the audio processing system 100 that includes hardware components and software components configured to perform a setup process. The grid mapper 102 may receive translation data 104. The translation data 104 may include a pre-computed original grid (e.g., a first grid). Example techniques for determining the original grid are described in U.S. Publication No. 2016 / 0007133. The received original grid includes a two-dimensional or three-dimensional grid of virtual sound sources (e.g., a first virtual sound source) distributed across a cell space (e.g., a listening room). The received original grid has a first density, which is measured by the number of virtual sound sources in the space (e.g., 11×11×11 virtual sound sources), corresponding to eleven virtual sound sources across the width of the space, eleven virtual sound sources along the length of the space, and eleven virtual sound sources in the height of the space. For convenience, the examples in this specification have equal width, length, and height in terms of the number of virtual sound sources. In various embodiments, the width, length, and height may be different. For example, the grid may have 11×11×9 virtual sound sources. Each virtual sound source is a point source. In the example shown, the virtual sound sources are evenly distributed in the space, where the distance between two adjacent virtual sound sources along the length dimension, width dimension, and optionally the height dimension is equal. In some embodiments, the virtual sound sources may be unevenly distributed, e.g., such that the distribution is denser where the expected sound energy is higher or the required spatial resolution is higher. The received original grid maps the speaker gain (e.g., a first speaker gain) of the virtual sound sources to one or more speakers according to the speaker layout in the listening environment. The received original grid specifies the corresponding amount of speaker gain contributed by each virtual sound source to each speaker.

[0021] By performing the setup process, the grid mapper 102 maps the received original fine grid to one or more coarser grids. The terms "fine" and "coarse" used in this specification are relative terms. If grid A is denser than grid B, e.g., if grid A has more virtual sound sources than grid B, then grid A is a fine grid relative to grid B, and grid B is a coarse grid relative to grid A. The virtual sound sources in grid A may be referred to as fine virtual sound sources. The virtual sounds in grid B are referred to as coarse virtual sound sources.

[0022] The grid mapper 102 may determine a second grid 106 filled with fewer virtual sound sources (e.g., 5×5×5) than the virtual sound sources in the received original grid. Relatively speaking, the second grid 106 is a coarse grid and the original grid is a fine grid. The grid mapper 102 may determine a third grid 108 filled with even fewer virtual sound sources (e.g., 3×3×3 virtual sound sources). The third grid 108 is an even coarser grid. Each of the second grid 106 and the third grid 108 maps the speaker gains of the virtual sound sources in the corresponding virtual grid to the speaker gains according to the same speaker layout in the listening environment. Each of the second grid 106 and the third grid 108 specifies the amount of the speaker gain contributed by each coarse virtual sound source to each speaker. Then, the grid mapper 102 stores the second grid 106, the third grid 108, and the original grid 110 in the storage device 112. The storage device 112 may be a non-transitory storage device, such as a disk or a memory of the audio processing system 100.

[0023] After setting the speaker positions, the renderer 114 may render one or more audio objects at runtime. The runtime may be the playback time when the audio signal is played on the speakers. The renderer 114 (e.g., an audio panner) includes one or more hardware and software components configured to perform a panning operation that maps the audio object to the speakers. The renderer 114 receives the audio object 116. The audio object 116 may include a position parameter and a size parameter. The position parameter may specify the apparent position of the audio object in space. The size parameter may specify the apparent size that the spatial sound field of the audio object 116 should exhibit during playback. Based on the size parameter, the renderer 114 may select one or more of the original grid 110, the second grid 106, or the third grid 108 to render the audio object. Generally, the renderer 114 may select a finer grid for a smaller apparent size. The renderer 114 may map the audio object 116 to one or more audio channels, each corresponding to a speaker. The renderer 114 may output the mapping as one or more speaker gains 118. The renderer 114 may submit the speaker gains to one or more amplifiers, or directly to one or more speakers. The renderer 114 may dynamically select the grid, using a fine grid for smaller audio objects and a coarse grid for larger audio objects.

[0024] Figure 2 is a schematic diagram showing example audio objects associated with corresponding apparent sizes. An audio encoding system may encode a particular audio scene (e.g., a band playing at a venue) as one or more audio objects. In the example shown, the audio processing system (e.g., Figure 1The audio processing system 100) renders audio objects 202 and 204. Each of the audio objects 202 and 204 includes a position parameter and a size parameter. The position parameter may include position coordinates indicating the respective positions of the corresponding audio object in a unit space. The space may be a three-dimensional volume having any geometric shape. In the example shown, a two-dimensional projection of the space is shown. In the example shown, the positions of the audio objects 202 and 204 are represented as black circles at the centers of the audio objects 202 and 204, respectively.

[0025] A grid 206 of virtual sound sources represents positions in the space. The virtual sound sources include, for example, virtual sound source 208, virtual sound source 210, and virtual sound source 212. Each virtual sound source is represented as Figure 2 a white circle in. The grid 206 spatially coincides with the space. For convenience, a 7×7 projection is shown. Virtual sound sources located on the outer boundary of the grid 206 (e.g., virtual sound sources 208 and 212) are designated as external virtual sound sources. Virtual sound sources located within the grid 206 (e.g., virtual sound source 210) are designated as internal virtual sound sources. External virtual sound sources that do not lie at the corners of the grid 206 (e.g., virtual sound source 208) are designated as non-corner sound sources. External virtual sound sources located at the corners of the grid 206 (e.g., virtual sound source 212) are designated as corner sound sources.

[0026] The shapes of the audio objects 202 and 204 can be zero-dimensional, one-dimensional, two-dimensional, three-dimensional, spherical, cubic, or have any other regular or irregular form. The size parameter of each of the audio objects 202 and 204 can specify the respective apparent sizes of each audio object. The renderer can activate all virtual sound sources that simultaneously fall within the size shape with an activation factor that depends on the exact number of virtual sound sources and optionally a window factor. During playback, the contributions of all virtual sound sources to the available speakers are added together. The addition of the sound sources is not necessarily linear. A quadratic addition law for maintaining the RMS value may be implemented. Other addition laws may be used. For an audio object at the boundary, e.g., audio object 204, the renderer can add together only the external virtual sound sources located on that boundary. In this example, if the audio object 204 extends across the entire boundary, seven virtual sound sources (49 in three-dimensional space) would be required to represent the audio object 204. Similarly, in this example, if the audio object 202 fills the entire space, 49 virtual sound sources (343 in three-dimensional space) would be required to represent the audio object 202. The audio processing system (e.g., Figure 1An audio processing system 100) can use a coarser grid than grid 206 to reduce the number of virtual sound sources required to represent audio objects 202 and 204. The audio processing system can use a cell assignment technique to create the coarser grid, which is described in more detail below.

[0027] The audio processing system can determine which virtual sound source or sources represent an audio object based on the position parameters and size parameters associated with the object. In the example shown, audio object 202 is represented by six virtual sound sources, including four internal virtual sound sources and two external audio sources. Audio object 204 is represented by four external virtual sound sources. The audio processing system should perform a separation operation and a mapping operation to represent audio objects 202 and 204 with fewer virtual sound sources in the coarser grid. For example, the audio processing system can use one or more coarse virtual sound sources (e.g., coarse virtual sound source 214) to represent audio objects 202 and 204 in the coarser grid. Figure 2 The coarse virtual sound source is shown as a white triangle in the figure.

[0028] Figure 3 is a schematic diagram showing an example technique for creating cells of fine virtual sound sources. Assigning virtual sound sources to cells is a stage in generating the coarser grid. After the grid mapper (e.g., Figure 1 grid mapper 102 of ) receives the original fine grid 206 of fine virtual sound sources in space, it assigns a corresponding cell to each virtual sound source in the grid. The original fine grid 206 can include an original number (e.g., K×L×M) of fine virtual sound sources evenly distributed in three-dimensional space. The positive integers K, L, and M can correspond to the number of virtual sound sources along the length, width, and height of the space, respectively. For convenience, Figure 3 shows a two-dimensional projection with dimensions of 7×7.

[0029] Assigning cells to virtual sound sources can include determining boundaries, e.g., boundaries 302 and 304, to divide the space into cells called fine cells. The boundaries 302 and 304 that divide the virtual sound sources in the fine grid 206 are designated as fine boundaries, as indicated by the dashed lines in the figure. The fine boundaries 302 and 304 can be the midlines or midplanes between virtual sound sources. The midline or midplane can be a line or plane on which the points are equidistant from two adjacent virtual sound sources. The grid mapper can assign each corresponding area or volume surrounded by the corresponding boundary around a virtual sound source to the cell corresponding to that virtual sound source. For example, the grid mapper can assign the area or volume around virtual sound source 210 to cell 306 corresponding to virtual sound source 210. The grid mapper creates the corresponding cells for each virtual sound source in the fine grid 206.

[0030] Figure 4is a schematic diagram showing an example technique for reducing the number of virtual sound sources. Reducing the number of virtual sound sources is another stage in generating a coarse grid. A grid mapper (e.g., Figure 1 's grid mapper 102) creates a set of virtual sound sources in the same space represented by the Figure 3 's fine grid 206. The grid mapper designates a set of positions in the space as a set of coarse virtual sound sources. The coarse virtual sound sources are fewer than the fine virtual sound sources represented in the original fine grid 206. For example, the grid mapper may specify that the coarse grid 402 has P×Q×R virtual sound sources, where at least one of P, Q, and R is less than K, L, and M respectively. For convenience, Figure 4 shows a two-dimensional projection of a coarse grid of 5×5 virtual sound sources. Each coarse virtual sound source in the grid 402 is represented as a triangle. The coarse virtual sound sources may have a uniform distribution in space. After creating the coarse grid 402, the grid mapper moves to the next processing stage: calculating the speaker gain for each coarse virtual sound source.

[0031] Figure 5 is a schematic diagram showing an example technique for creating units for coarse virtual sound sources. Assigning units to the reduced virtual sound sources is another stage in generating a coarse grid. A grid mapper (e.g., Figure 1 's grid mapper 102) assigns corresponding coarse units to each coarse virtual sound source in the coarse grid 402. Assigning coarse units to coarse virtual sound sources may include determining boundaries for dividing the space into coarse units, e.g., boundaries 502 and 504. The boundaries 502 and 504 that divide the coarse virtual sound sources in the coarse grid 402 are designated as coarse boundaries, as represented by the dashed lines in the accompanying drawings. The coarse boundaries 502 and 504 may be midlines or mid-planes between internal virtual sound sources (e.g., internal virtual sound sources 506 and 508) and between external virtual sound sources that are not corner sources (e.g., external virtual sound sources 510 and 512). In some first embodiments, between the external virtual sound source 510 and the internal virtual sound source 506 or between the non-corner source 510 and the corner source 514, the grid mapper may determine the midline. In some second embodiments, the grid mapper may designate the fine boundaries of the fine grid 206 that are between internal and external virtual sound sources or between non-corner and corner sound sources as coarse boundaries. For example, in the second embodiment, the grid mapper may use Figure 3 's boundary 304 to divide the internal virtual sound source 506 and the external sound source 510, and also use Figure 3 's boundary 302 to divide the non-corner source 510 and the corner source 514.

[0032] The grid mapper designates each corresponding area or volume surrounded by a corresponding boundary around a corresponding coarse virtual sound source as a coarse cell corresponding to that coarse virtual sound source. For example, the grid mapper may designate the space around virtual sound source 508 as coarse cell 516 corresponding to coarse virtual sound source 508. Then, the grid mapper can proceed to the next processing stage.

[0033] Figure 6 is a schematic diagram showing an example technique for mapping fine virtual sound sources to coarse virtual sound sources when determining speaker gains. The grid mapper (e.g., Figure 1 grid mapper 102) creates coarse virtual sound sources, including a specific virtual sound source 602, for which there is currently no information on the corresponding speaker gain. The grid mapper can determine the speaker gain corresponding to a coarse virtual sound source based on the overlap between fine cells and coarse cells.

[0034] For example, the grid mapper determines that coarse virtual sound source 602 is associated with coarse cell 603. The grid mapper determines that coarse cell 603 overlaps with four fine cells, which are respectively associated with fine virtual sound sources 604, 606, 608, and 610. The grid mapper can calculate the corresponding overlap ratio, which refers to the corresponding amount of overlap. The overlap ratio can be the ratio between the area (or volume) of the corresponding fine cell that overlaps with the coarse cell and the total area (or volume) of that corresponding fine cell.

[0035] For example, as Figure 6 shown, the grid mapper can determine that the entire fine cell corresponding to fine virtual sound source 604 is located within coarse cell 603. In response, the grid mapper can determine that the overlap ratio of the fine cell corresponding to the original virtual sound source 604 is 1.00 or 100%. Similarly, the grid mapper can determine that the corresponding overlap ratios of the fine cells corresponding to fine virtual sound sources 606 and 608 are approximately 0.83 or 83%, and the overlap ratio of the fine cell corresponding to fine virtual sound source 610 is approximately 0.69 or 69%.

[0036] Therefore, the grid mapper can determine the speaker gain contribution of virtual sound source 602 by summing the contributions of virtual sound sources 604, 606, 608, 610 weighted by the overlap ratios. The summation can be implemented using various techniques. For example, the same technique used to add the contributions from all virtual sound sources to available speakers during playback can be used to implement the summation.

[0037] More generally, the grid mapper can use Equation 1 below to determine the speaker gain contribution.

[0038] G ui =[∑ v w uv (h vv g vi) p 1 / p (1)

[0039] In Equation 1, G ui represents the contribution of the coarse virtual sound source u to the loudspeaker i; p = 1, 2, 3...; h uv is a height correction term that can assign equal or different weights to different sound sources. For example, in some embodiments, h uv can give more weight to the fine virtual sound sources that are closer to the bottom (e.g., the floor of the listening room) relative to the position of the coarse virtual sound source, g vi represents the gain contribution of the original fine virtual sound source v to the loudspeaker i. In some other embodiments, if no distinction between sound sources at different heights is desired, h uv can be set to 1 for all fine virtual sound sources. Additionally, w uv is the weight of the fine virtual sound source v with respect to the coarse virtual sound source u, where for fine cells that completely fall within the coarse cell, w uv = 1; for fine cells that partially fall within the coarse cell corresponding to u, 0 < w uv < 1; and for fine cells that do not overlap with the coarse cell, w uv = 0. For example, the weight can correspond to the overlap rate.

[0040] The grid mapper can perform an additional coarsening stage from the original grid or from the coarse grid. During rendering, the renderer can use the coarse grid to determine the contribution of the coarse virtual sound sources to audio objects with non-zero apparent sizes. The renderer can use the fine grid in translations of zero size (where the apparent size of the audio object is zero).

[0041] In the example shown, the audio object 202 is initially represented by six fine virtual sound sources (including four internal virtual sound sources and two external audio sources). The audio object 204 is initially represented by four fine external virtual sound sources. The renderer can use the coarse grid to represent the audio object 202 and the audio object 204. In the coarse grid, the audio object 202 is represented by two coarse virtual sound sources (one internal and one external). The audio object 204 is represented by three coarse virtual sound sources (all external). The reduction in the number of represented sound sources reduces the demand for computational resources without sacrificing playback quality.

[0042] Figure 7 ​FIG. is a schematic diagram showing an example technique for reducing the number of virtual sound sources for a large audio object. For a large audio object having an apparent size approaching the entire space (e.g., an entire room), the grid mapper may create a coarse grid 702 having only one internal coarse virtual sound source 704. The other coarse virtual sound sources in the coarse grid 702 are external coarse virtual sound sources. All coarse virtual sound sources may be evenly distributed in the coarse grid 702. The coarse grid 702 may be a grid having 3×3×3 virtual sound sources. Figure 7 A two-dimensional projection is shown in.

[0043] At runtime, the renderer may select the fine grid 206, the coarse grid 402, or the coarsest grid 702 based on the size of the audio object and one or more size thresholds. For example, the grid mapper may generate a series of grids Grid0, Grid1, Grid2...GridN, where Grid0 is the original fine grid, e.g., Figure 2 the grid 206, and Grid1 to GridN are a series of successively coarser grids including Figure 4 the coarse grid 402, and the coarse grid 702. The renderer may define a series of successively larger size thresholds s1, s2...sN. The renderer may determine the output speaker gain as follows.

[0044] · If the size of the audio object s satisfies the condition s < s1, the renderer interpolates the gain calculated from Grid0 with the gain calculated from Grid1;

[0045] · If s(i - 1) <= s < si, the renderer interpolates the gain from Grid(i - 1) with the gain calculated from Grid(i);

[0046] · If s > sN, the renderer calculates the speaker gain based on GridN.

[0047] For example, at runtime, the renderer may interpolate the gain from the grid 206 and the gain from the grid 402 when determining that the size of the audio object is less than 0.2, interpolate the gain from the grid 402 and the gain from the grid 702 when determining that the size of the audio object is between 0.2 and 0.5, and use the grid 702 to determine the gain when determining that the size of the audio object is greater than 0.5, where the size of the space is 1.

[0048] Figure 8 FIG. is a flowchart of an example process 800 for rendering an audio object having an apparent size. The process 800 may be performed by a system including one or more computer processors (e.g., Figure 1 the audio processing system 100).

[0049] The system receives (802) audio panning data. The audio panning data includes a first grid that assigns a first speaker gain of a first virtual sound source in space to a speaker gain. The panning data can be data provided by a conventional panner with full resolution. For example, the first grid can be a fine grid with K×L×M fine virtual sound sources. The conventional panner has determined the first speaker gain of the fine virtual sound sources.

[0050] The system determines (804) a second grid of a second virtual sound source in space. The second grid is a coarse grid relative to the first grid and is less dense than the first grid. Determining the second grid includes mapping the first speaker gain of the first virtual sound source to the second speaker gain of the second virtual sound source. Determining the second grid can include the following operations. The system divides the space of the first grid into first cells. Each first cell is a fine cell corresponding to a respective first virtual sound source in the first grid. The system divides the space into second cells, which are fewer and coarser than the first cells. Each second cell corresponds to a respective second virtual sound source created by the system. The system maps the respective first speaker gain from each first virtual sound source to one or more second speaker gains of one or more second virtual sound sources based on the amount of overlap between the corresponding first cell and one or more corresponding second cells.

[0051] Mapping the respective first contribution (e.g., first speaker gain) from each first virtual sound source to one or more second contributions (e.g., second speaker gain) can include the following operations. The system determines the respective amount of overlap of the corresponding first cell in each of the one or more corresponding second cells. The system determines the respective speaker gain weight in each of the second speaker gains based on the respective amount of overlap. The system assigns the first speaker gain to each of the one or more second contributions based on the respective weights.

[0052] The space can be a two-dimensional space or a three-dimensional space. The first virtual sound source can include an external first sound source located on the outer boundary of the space and an internal first sound source located within the space. The second virtual sound source can include an external second sound source located on the outer boundary of the space and an internal second sound source located within the space. The external second sound source can include corner sound sources and non-corner sources. Dividing the space into second cells includes the following steps. Between each external sound source and the corresponding internal sound source, or between each corner sound source and the corresponding non-corner source, the system divides the corresponding second cell according to the fine cell boundary of the corresponding first cell as a fine cell. Between each pair of internal second sound sources or between each pair of non-corner sound sources, the system divides the corresponding second cell by the midline between the two sound sources of the pair.

[0053] The system selects at least one of a first grid or a second grid (806) to render an audio object based on a size parameter of the audio object. In some embodiments, selecting at least one of the first grid or the second grid may include the following operations. The system receives the audio object. The system determines an apparent size of the sound space based on the size parameter in the audio object. The system selects the first grid when determining that the apparent size is not greater than a threshold, or selects the second grid when determining that the apparent size is greater than the threshold.

[0054] The system renders (808) the audio object based on the selected one or more grids, including representing the audio object using one or more virtual sound sources within each selected grid that are enclosed within the sound space defined by the size parameter. Rendering the audio object includes: providing a signal representing the audio object to one or more speakers according to the output speaker gain determined in stage 806.

[0055] In some embodiments, the system renders the audio object using two or more grids. In this case, the system determines a third grid for a third virtual sound source in the space. The first grid is a fine grid; the second grid is a coarse grid; the third grid is intermediate, coarser than the first grid but not as coarse as the second grid. The third grid has fewer third virtual sound sources than the first virtual sound sources and more third virtual sound sources than the second virtual sound sources. Determining the third grid includes mapping a first contribution (e.g., a first speaker gain) to a third contribution (e.g., a third speaker gain) corresponding to the third virtual sound source. Selecting a grid among the three grids may include the following operations. The system selects the first grid and the third grid when determining that the apparent size is less than a first threshold (e.g., 0.2), where the space is a unit space of 1.

[0056] When the system uses two or more grids, the system determines the output speaker gain by interpolating the speaker gains. For example, when the first grid and the third grid are selected, the system may determine the output speaker gain by interpolating the speaker gains calculated based on the first grid and the third grid. The system selects the third grid and the second grid when determining that the apparent size is between the first threshold and a second threshold (e.g., 0.5 greater than the first threshold). The system determines the output speaker gain by interpolating the speaker gains determined based on the third grid and the second grid. The system selects the second grid when determining that the apparent size is greater than the second threshold. The system designates the speaker gain determined based on the second grid as the output speaker gain.

[0057] Example system architecture

[0058] Figure 9 is the implementation reference Figures 1 to 8Block diagram of an example system architecture of an audio rendering system with the described features and operations. Other architectures with more or fewer components are also possible. In some embodiments, architecture 900 includes one or more processors 902 (e.g., dual-core processor), one or more output devices 904 (e.g., LCD), one or more network interfaces 906, one or more input devices 908 (e.g., mouse, keyboard, touch-sensitive display), and one or more computer-readable media 912 (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components may exchange communications and data via one or more communication channels 910 (e.g., bus), which may utilize various hardware and software to facilitate the transfer of data and control signals between components.

[0059] The term "computer-readable media" refers to a medium that participates in providing instructions to processor 902 for execution, including but not limited to non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media includes but is not limited to coaxial cables, copper wire, and fiber optics.

[0060] Computer-readable media 912 may further include an operating system 914 (e.g., operating system), a network communication module 916, speaker layout mapping instructions 920, grid mapping instructions 930, and rendering instructions 940. Operating system 914 may be multi-user, multi-processor, multi-tasking, multi-threaded, real-time, etc. Operating system 914 performs basic tasks including but not limited to: identifying input from network interface 906 and / or device 908 and providing output to the network interface and / or device 908; keeping track of and managing files and directories on computer-readable media 912 (e.g., memory or storage device); controlling peripheral devices; managing traffic on one or more communication channels 910. Network communication module 916 includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.).

[0061] Speaker layout mapping instructions 920 may include computer instructions that, when executed, cause processor 902 to perform the following operations: receive speaker layout information specifying where in space each speaker is located, receive configuration information specifying a grid size (e.g., 11×11×11), and determine a grid of speaker gains that map the location of virtual sound sources to each speaker. Grid mapping instructions 930 may include instructions that, when executed, cause processor 902 to perform Figure 1Computer instructions for the operation of the grid mapper 102, the operations including mapping a grid generated by the speaker layout mapping instruction 920 to one or more coarse grids. The rendering instruction 940 may include causing the processor 902 to perform, when executed, Figure 1 Computer instructions for the operation of the renderer 114 of Figure 1 , the operations including selecting one or more grids to render an audio object.

[0062] The architecture 900 may be implemented in a parallel processing infrastructure or a peer-to-peer infrastructure or on a single device having one or more processors. The software may include multiple software components or may be a single body of code.

[0063] The described features may advantageously be implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to send data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform some activity or bring about some result. The computer program may be written in any form of programming language including compiled or interpreted languages (e.g., Objective-C, Java), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0064] By way of example, suitable processors for executing instruction programs include both general and special purpose microprocessors, as well as one or more processors or cores of any type of computer's sole processor. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing the instructions and one or more memories for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data files, or be operatively coupled to communicate with them; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, which by way of example includes semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may both be supplemented by, or incorporated in, an ASIC (application specific integrated circuit).

[0065] To provide interaction with a user, features may be implemented on a computer having a display device such as a CRT (cathode ray tube) monitor, an LCD (liquid crystal display) monitor, or a retinal display device for presenting information to a user. The computer may have a touch surface input device (e.g., a touch screen), a keyboard, and a pointing device such as a mouse or a trackball by which the user can provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0066] Features may be implemented in a computer system that includes backend components such as a data server, or includes middleware components such as an application server or an Internet server, or includes frontend components such as a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of the system may be connected by any form or medium of digital data communication such as a communication network. Examples of communication networks include, for example, LANs, WANs, and the computers and networks that form the Internet.

[0067] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. The relationship of the client to the server is created by computer programs that run on the respective computers and have a client - server relationship to each other. In some embodiments, the server transmits data (e.g., an HTML page) to a client device (e.g., for the purpose of presenting data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., the result of a user interaction) may be received at the server from the client device.

[0068] A system of one or more computers may be configured to perform particular actions by software, firmware, hardware, or a combination thereof installed on the system that, when operating, causes the performance of the actions or causes the system to perform the actions. One or more computer programs may be configured to perform particular actions by instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.

[0069] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of any claimed subject matter, but rather as descriptions of features specific to particular embodiments of a specific invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be removed from the combination, and the claimed combination may refer to a sub-combination or a variant of a sub-combination.

[0070] Similarly, although the operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in an sequential order, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0071] Accordingly, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order shown or sequential order to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0072] Multiple embodiments of the invention have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention.

Claims

1. A method for rendering an audio object, comprising: Receiving, by one or more processors, audio panning data, the audio panning data including a first grid that specifies a first speaker gain for a first virtual sound source in a space, the first speaker gain corresponding to one or more speakers in the space; Determining, based on the first grid, a second grid for a second virtual sound source in the space, including mapping the first speaker gain to a second speaker gain for the second virtual sound source, wherein the second virtual sound source is fewer than the first virtual sound source; Selecting at least one of the first grid or the second grid for rendering the audio object based on a size parameter of the audio object; and Rendering the audio object based on the selected one or more grids, including: representing the audio object using one or more virtual sound sources in each selected grid that are enclosed in a sound space at least partially defined by the size parameter.

2. The method according to claim 1, wherein Determining the second grid includes: Dividing the space into first cells, each first cell corresponding to a respective first virtual sound source in the first grid; Dividing the space into second cells, the second cells being fewer than the first cells, each second cell corresponding to a respective second virtual sound source; and Mapping the respective first speaker gain of each first virtual sound source to a respective second speaker gain of one or more second virtual sound sources based on an overlap amount between a corresponding first cell and one or more corresponding second cells.

3. The method according to claim 2, wherein Mapping the first speaker gain to the second speaker gain includes: Determining a respective overlap amount between each first cell and each second cell; Determining a respective weight for the contribution of the first speaker gain of each first virtual sound source to each second virtual sound source based on the corresponding overlap amount; and Allocating the first speaker gain to each of the second speaker gains according to the respective weights.

4. The method according to claim 2, wherein: The space is a two-dimensional space or a three-dimensional space, The first virtual sound sources include external first sound sources located on an outer boundary of the space and internal first sound sources located within the space, and The second virtual sound sources include external second sound sources located on the outer boundary of the space and internal second sound sources located within the space, the external second sound sources including corner sound sources and non-corner sources.

5. The method according to claim 4, wherein, Dividing the space into the second cells includes: Between each external sound source and a corresponding internal sound source, or between each corner sound source and a corresponding non-corner source, dividing the corresponding second cell according to a cell boundary of the corresponding first cell; and Between each pair of internal second sound sources or between each pair of non-corner sources, dividing the corresponding second cell by a midline between the two sound sources of the pair.

6. The method according to claim 1, wherein Selecting at least one of the first grid or the second grid includes: Receiving the audio object; Determining an apparent size of the sound space based on the size parameter in the audio object; and When it is determined that the apparent size is not greater than the threshold, the first grid is selected, or when it is determined that the apparent size is greater than the threshold, the second grid is selected.

7. The method according to claim 1, wherein: Selecting at least one of the first grid or the second grid includes: selecting the first grid and the second grid, and Rendering the audio object includes: determining an output speaker gain by interpolating the first speaker gain and the second speaker gain based on the apparent size of the sound space, the apparent size being determined based on sound parameters in the audio object.

8. The method according to claim 1, comprising: Determining a third grid for a third virtual sound source in the space includes mapping the first speaker gain to a third speaker gain corresponding to the third virtual source, wherein the third grid has fewer third virtual sound sources than the first virtual sound source and more third virtual sound sources than the second virtual sound source.

9. The method according to claim 8, wherein Selecting at least one of the first grid or the second grid for rendering the audio object includes: When it is determined that the apparent size of the sound space is less than a first threshold, selecting the first grid and the third grid, wherein rendering the audio object includes: determining an output speaker gain by interpolating the first speaker gain and the third speaker gain; When it is determined that the apparent size is between the first threshold and a second threshold greater than the first threshold, selecting the third grid and the second grid, wherein rendering the audio object includes: determining an output speaker gain by interpolating the third speaker gain and the second speaker gain; and When it is determined that the apparent size is greater than the second threshold, selecting the second grid, wherein rendering the audio object includes: determining that the second speaker gain is selected as the output speaker gain.

10. The method according to claim 9, wherein, Rendering the audio object includes: Providing a signal representing the audio object to one or more speakers according to the output speaker gain.

11. A system for rendering an audio object, comprising: One or more processors; And A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 10.

12. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Rendering of audio objects with apparent size to arbitrary loudspeaker layouts

    US20160007133A1

  • Sound reproducing system sound reproducing method

    CN101150890A

  • Rendering of audio objects with apparent size to arbitrary loudspeaker layouts

    CN105075292A