An audio object renderer, a method and a computer program for determining speaker gain, using panned object speaker gain and spread object speaker gain

The audio object renderer addresses the challenge of balancing computational complexity and auditory impression by combining panned and object feature information-based speaker gains, effectively simulating the expansion and diffusion of audio objects with extended sound sources.

JP7693664B2Active Publication Date: 2025-06-17FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022529355
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-20
Filing Date
2020-11-20
Publication Date
2025-06-17
Estimated Expiration
2040-11-20

AI Technical Summary

Technical Problem

Existing audio object rendering technologies face challenges in achieving a good trade-off between auditory impression and computational complexity, particularly when dealing with audio objects that require extended sound source sizes and complex psychoacoustic effects.

Method used

An audio object renderer that determines speaker gains by combining panned object speaker gains and object feature information speaker gains. The renderer uses point source panning for panned object speaker gains and considers object feature information, such as spread and distance, to calculate speaker gains that effectively simulate the expansion and diffusion of audio objects.

Benefits of technology

The proposed solution achieves a good balance between computational efficiency and auditory realism, enabling effective localization and perception of audio objects with extended sound sources, even for listeners not in the optimal listening arrangement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693664000003
    Figure 0007693664000003
  • Figure 0007693664000004
    Figure 0007693664000004
  • Figure 0007693664000005
    Figure 0007693664000005
Patent Text Reader

Abstract

The audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a-1214c) describing gains for including one or more audio object signals (1260) in a plurality of speaker signals (1262a-1262c) based on the object position information (210, 1210, azi, ele) and the object feature information or spread information (1212) is configured to obtain panned object speaker gains (202a, 1232, g) using point source panning of the audio object (202, 1230). The audio object renderer is configured to obtain spread object speaker gains (206a, 1242, gOS) taking into account the object position information (210, 1210, azi, ele) and the object feature information or spread information (1212). The audio object renderer is configured to combine the panned object speaker gains (202a, 1232, g) and the spread object speaker gains (206a, 1242, gOS) such that there is always a contribution of the panned object speaker gain to obtain a combined speaker gain (214, 1214, 1214a-1214c). A method and a computer program are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments according to the present invention relate to an audio object renderer.

[0002] Further embodiments according to the present invention relate to a method for determining speaker gain.

[0003] Further embodiments according to the present invention relate to a computer program.

[0004] Embodiments according to the present invention generally relate to panning for audio objects having an extended sound source size.

Background Art

[0005] The following provides some background on the present invention. However, note that the features, functions, and applications mentioned below are optionally and can also be used in combination with embodiments according to the present invention.

[0006] In the field of surround sound reproduction, speakers are typically placed at some specific locations in a room. The commonly used surround reproduction system "5.1" includes three speakers within the front hemisphere and two speakers within the rear hemisphere. When a signal (e.g., a monaural audio signal) is intended to be reproduced in the space between two speakers, the signal is evenly distributed to these two adjacent speakers. This procedure also works for a 3D speaker setup where speakers are added above and / or below the horizontal plane. A well-known panning algorithm is the so-called "Vector Base Amplitude Panning" (VBAP). After calculating the panning gain, the monaural signal is reproduced from the associated speakers with corresponding weighting.

[0007] Most panning techniques are known to reproduce a point-like sounding signal (object) within a space. It is also known, although frequent, that it may be desirable to change the size of the object, increase the diffusion of the sound, change the perceived distance, or bring about other psychoacoustic effects. Thus, the object should (or sometimes must) emit sound not only as a point, but also from a wider reproduction angle.

[0008] Figure 1 shows a graphic representation of the spread configurations of different objects. In the upper part, at reference numerals 100, 101, 102, objects with three different spread values are shown. In the lower part, at reference numerals 104 and 105, the objects are spread non-uniformly on the reproduction sphere.

[0009] In other words, Figure 1 shows different object spread configurations that are independent of the setup of the reproduction speakers. At reference number 100, a point-like sounding object is depicted. At reference numbers 101 and 102, the objects are uniformly spread over a wider / higher reproduction angle. At reference number 104, the object spreads vertically, and at reference number 105, it spreads horizontally. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0010] In view of such a situation, it is necessary to create a concept that provides an improved trade-off relationship between the auditory impression and the computational complexity. MEANS FOR SOLVING THE PROBLEM

[0011] One embodiment according to the present invention creates an audio object renderer for determining a speaker gain that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information and object feature information. The audio object renderer is configured to obtain a panned object speaker gain using point source panning of the audio object. The audio object renderer is configured to obtain an object feature information speaker gain (e.g., a speaker gain considering, for example, the expansion and / or perceived expansion and / or perceived angular expansion and / or diffusion and / or blur of one or more audio objects under consideration) in consideration of the object feature information (1212). For example, the object feature information may describe a distribution to a plurality of points that may correspond to the divergence from the sound source (or the audio object, or the sound derived from the audio object), for example, the broadening of the sound source. For example, the object, or the perception of the object, may be expanded according to the object feature information. Generally, for example, the object feature information may represent the spread (broadening) and / or range and / or diffusivity of the audio object, and the object feature information speaker gain may consider this spread and / or range and / or diffusivity of the audio object. Alternatively, or in addition, the object feature information may describe, for example, the distance of the audio object, and this distance may be converted to a spread, for example, in a preparation stage, and then this spread may be considered in providing the object feature information speaker gain. However, as another option, the object feature information speaker gain may also be directly derived from the distance.

[0012] The audio object renderer is further configured to combine the panned object speaker gains (202a, 1232, g) and the object feature information speaker gains (206a, 1242, gOS) such that there is always a contribution from the panned object speaker gain to obtain a combined speaker gain (214, 1214, 1214a - 1214c).

[0013] This embodiment according to the present invention is based on the discovery that a good configuration between the computational complexity and the achievable auditory impression can be obtained by obtaining an object feature information speaker gain (which may correspond to a spread object speaker gain) that describes the intensity of the object signal from an audio object in different speaker signals associated with different speakers, based on both the panned object speaker gain and the spread object speaker gain. In particular, by using the panned object speaker gain, which typically provides a "point sound source" auditory impression, the localization of the audio object by the user can be smoothed. For example, in deriving the panned object speaker gain, point sound source panning of the audio object may be used, which may, for example, select a single speaker for playback of the audio object, or, for example, distribute the audio object to a plurality of speakers closest to the audio object (while leaving, for example, a speaker that is close but not the closest to the audio object unused). Thus, the "point sound source" panning of the audio object typically provides a panned object speaker gain where only the speaker gains of a small number of speakers closest to the object position are non - zero.

[0014] In addition, the audio object renderer also obtains an object feature information object speaker gain, and the object extends over an extended area, for example, an extended range of azimuth angle and / or an extended range of elevation angle. Therefore, in determining the object feature information object speaker gain, the extension of the audio object is considered, which can be derived, for example, from the object feature information. In determining the object feature information object speaker gain, typically, the audio object will be spread over a larger number of speakers than in the determination of the panned object speaker gain, because in the determination of the object feature information object speaker gain, the extension of the audio object is considered and typically a relatively wide (e.g., wider than the extension of the audio object under consideration) stationary (e.g., stationary attenuation) distribution characteristic is used.

[0015] Therefore, by combining the panned object speaker gain based on the point source panning of the audio object and the object feature information object speaker gain considering the extension of the audio object, good localization of the audio object can be achieved (even for an object with a relatively large extension), and it is also possible to perceive the extension of the audio object. This is particularly applicable in situations where there are multiple uses and / or where the user is not located in the "sweet spot" of the listening arrangement configuration. By always introducing, for example, according to the object feature information, the contribution obtained by the panned object speaker, it can be achieved that the localization of the audio object is always ensured and is also possible for a listener not in the sweet spot arrangement.

[0016] Furthermore, it should be noted that object feature information typically enables the determination or estimation of the extension of an audio object. For example, the object feature information may indicate the type of the object, and this type of object may be meant as a parameter (e.g., a spread parameter) for the determination of the spread object speaker gain. For example, the object feature information may make it possible to distinguish a relatively small object from a relatively large object. Alternatively, or in addition thereto, the object feature information enables the distinction between a near object and a far object, which may mean one or more parameters for the determination of the spread object speaker gain. The object feature information may optionally describe the "blurred" or "diffused" extension of the audio object, or the distribution of the audio object to multiple localization positions.

[0017] In conclusion, the audio object renderer may derive one or more parameters for the determination of the object feature information object speaker gain from the object feature information. Thus, the object feature information makes it possible to appropriately adjust the derivation of the spread object speaker gain, whereby the object can be localized by providing the panned object speaker gain, and at the same time, can be perceived with an appropriate extension by considering the object feature information in the provision of the object feature information speaker gain.

[0018] In a preferred embodiment, the audio object renderer is configured to obtain the object feature information speaker gain by additionally considering the object position information. This allows both the extension and the position of the audio object to be taken into account.

[0019] In a preferred embodiment, the object feature information is audio object spread information. This enables particularly efficient calculations as there is no need to map "abstract" object feature information onto the object spread information in this case.

[0020] One embodiment according to the present invention creates an audio object renderer for determining a speaker gain (e.g., a combined speaker gain or the resulting speaker gain) that describes the gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele)), which can be given, for example, in spherical coordinates, and object feature information, using, for example, an azimuth value azi and an elevation value ele. The object feature information can be, for example, information indicating whether the object is small or expanded, e.g., an object size value, or the object feature information can be, for example, object distance information that can be mapped onto a spread value (e.g., spread angle information describing the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, e.g., spreadAngleEle). However, other types of object feature information are also possible.

[0021] The audio object renderer is configured to obtain a panned object speaker gain (e.g., also designated as "object speaker gain" or represented by the vector g) using point source panning of the audio object. In point source panning, the audio object can be regarded as a point source, for example, the spread information can be ignored, and the signal of the audio object is associated with two or more speakers within the environment of the object position of the audio object by an appropriate selection of the panned object speaker gain.

[0022] The audio object renderer is configured to obtain a spread object speaker gain (also designated, for example, as a spread speaker gain or represented as a vector gOS) in consideration of object position information and object characteristic information.

[0023] The audio object renderer is configured to combine a panned object speaker gain (e.g., g) and a spread object speaker gain (e.g., gOS) such that the contribution of the panned object speaker gain is always present (e.g., independent of the object characteristic information) to obtain a combined speaker gain.

[0024] This embodiment according to the present invention is based on the discovery that a good configuration between computational complexity and achievable auditory impression can be obtained by obtaining a spread object speaker gain, which describes the intensity of an object signal from an audio object in different speaker signals associated with different speakers, based on both the panned object speaker gain and the spread object speaker gain. In particular, the use of the panned object speaker gain, which typically provides a "point sound source" auditory impression, can smooth the localization of the audio object by the user. For example, in deriving the panned object speaker gain, point sound source panning of the audio object may be used, which may, for example, select a single speaker for playback of the audio object, or distribute the audio object to a plurality of speakers closest to the audio object (while leaving, for example, speakers that are not the closest to the audio object unused). Thus, the "point sound source" panning of the audio object typically provides a panned object speaker gain where only the speaker gains of a few speakers closest to the object position are non-zero.

[0025] In addition, the audio object renderer also obtains a spread of object speaker gains, where the object spreads over an extended region, e.g., an extended range of azimuth and / or an extended range of elevation. Thus, in determining the spread of object speaker gains, the spread of the audio object is taken into account, which can be derived, for example, from object characteristic information. In determining the spread of object speaker gains, typically, the audio object is spread over a larger number of speakers than in the determination of the panned object speaker gain, because in the determination of the spread of object speaker gains, the spread of the audio object is taken into account and typically a relatively wide (e.g., wider than the spread of the audio object under consideration) stationary (e.g., stationary attenuation) distribution characteristic is used.

[0026] Therefore, by combining the panned object speaker gain based on the point-source panning of the audio object and the spread of object speaker gain taking into account the spread of the audio object, good localization of the audio object is achieved (even for objects with a relatively large spread), and it is also possible to perceive the spread of the audio object. This is particularly applicable in situations where there are multiple uses and / or where the user is not located in the "sweet spot" of the listening arrangement configuration. By always introducing the contribution obtained by the panned object speaker, for example, depending on the spread information or regardless of the object characteristic information, it can be achieved that the localization of the audio object is always ensured and is possible for listeners not in the sweet spot arrangement.

[0027] Furthermore, it should be noted that object feature information typically enables the determination or estimation of the spread of an audio object. For example, the object feature information may indicate the type of the object, and this type of object may be meant as a parameter (e.g., a spread parameter) for the determination of the spread object speaker gain. For example, the object feature information may make it possible to distinguish a relatively small object from a relatively large object. Alternatively, or in addition thereto, the object feature information may enable the distinction between a near object and a far object, which may mean one or more parameters for the determination of the spread object speaker gain.

[0028] In conclusion, the audio object renderer may derive one or more parameters for the determination of the spread object speaker gain from the object feature information. Accordingly, the object feature information makes it possible to appropriately adjust the derivation of the spread object speaker gain, whereby the object can be localized by providing the panned object speaker gain, and at the same time, can be perceived with an appropriate spread by taking into account the object feature information in the provision of the spread object speaker gain.

[0029] In conclusion, the above-described audio object renderer enables the determination of speaker gain that provides a good auditory impression while moderately reducing the computational complexity.

[0030] As a further conclusion, the present invention generally creates an object renderer that pans an object using VBAP, then determines an object feature gain for the object, and always combines these in relation to the VBAP-panned object.

[0031] Another embodiment according to the present invention creates an audio object renderer for determining a speaker gain (e.g., a combined speaker gain or a resulting speaker gain) that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi) and / or elevation (ele)), which can be given in spherical coordinates (e.g., using azimuth value azi and elevation value ele), and spread information (e.g., spread angle information describing the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, e.g., spreadAngleEle).

[0032] The audio object renderer is configured to obtain a panned object speaker gain (also designated as "object speaker gain" or represented by vector g) using point source panning of the audio object (e.g., the audio object is considered a point source, e.g., the spread information is ignored, e.g., the signal of the audio object is associated with two or more speakers within the environment of the object position of the audio object by appropriate selection of the panned object speaker gain).

[0033] The audio object renderer is configured to obtain a spread object speaker gain (also, e.g., designated as spread speaker gain or represented, e.g., as vector gOS) taking into account the object position information and the spread information.

[0034] The audio object renderer is configured to combine panned object speaker gain (e.g., g) and spread object speaker gain (e.g., gOS) such that the contribution of the panned object speaker gain is always present (e.g., independent of spread information) to obtain a combined speaker gain.

[0035] This audio object renderer is based on the same considerations as the above-described audio object renderer. However, instead of object characteristic information, spread information that directly describes how the object should be spread is evaluated. For example, the spread information may be spread angle information that describes the spread in the azimuth direction and / or the spread in the elevation direction. Alternatively, the spread information may also be solid angle information or may specify the size of the object in other forms (e.g., using absolute size information and / or distance information, or the like). Thus, it is possible to obtain a combined speaker gain, whereby the combined speaker gain enables the representation of an object signal with a good auditory impression, allowing the listener to localize the object and perceive the object with appropriate expansion. For example, this can be achieved by always using panning even when the spread of the object is wide enough.

[0036] In the preferred embodiment of the foregoing audio object renderer, the audio object renderer maps the difference between the position of a support point (e.g., an artificially generated support point, e.g., specified by SSP) and the object position to one or more spread gain value contributions (e.g., aziGain(naz) or eleGain(nel)), evaluates one or more gain functions (e.g., one or more polynomial functions or one or more parabolic functions, e.g., a spread weighting curve), and is configured to determine a spread object speaker gain (e.g., gOS) based on the one or more spread gain value contributions.

[0037] It is possible to have a uniform and computationally efficient computational scheme for determining spread gain values by using the position of the support point and evaluating one or more gain functions that can map the difference between the position of the support point and the object position to one or more spread gain value contributions. The gain function can, for example, weight the angular difference (e.g., with respect to both azimuth and elevation) between the object position and the position of the support point, thereby enabling a simple yet accurate determination of the spread gain value contribution. Further, however, by using the support point position rather than the speaker position, a high degree of uniformity is ensured, which typically helps to reduce the complexity of the algorithm. For example, the mapping from the signal gain associated with the support point to the signal gain associated with the actual speaker can be pre-computed and need not be recomputed for each audio object. However, in the determination of gain values typically associated with geometrically regular support points, it is possible by using an algorithm with a relatively low complexity. Thus, the determination of the combined speaker gain is possible with a high degree of efficiency.

[0038] In the preferred embodiment of the foregoing audio object renderer, the audio object renderer has a spread in a first direction (e.g., spread azi) according to, and also according to the spread in the second direction (e.g., spread ele ) is configured to determine the weight of the spread object speaker gain in combination with the panned object speaker gain.

[0039] Thus, it can be determined what importance (or weight) the determination of the spread object speaker gain has when compared with the determination of the panned object speaker gain. For example, for a large spread angle, the relative importance of the spread object speaker gain in the combination can be higher when compared with a situation where the spread is relatively small. Further, by adjusting the weighting of the spread object speaker gain, for a relatively wide spread, the (relative) contribution of the panned object speaker gain is reduced, thereby ensuring that a bad auditory impression can be avoided. In contrast, when the spread is relatively small, the (relative) contribution of the panned object speaker gain is increased, thereby being able to reflect the localization of the audio object. Thus, for example, when the spread of the audio object may increase over time, it is also possible to have a smooth transition, which can be the case, for example, when the audio object under consideration is a moving audio object. In conclusion, the determination of the weighting of the spread object speaker gain in combination with the panned object speaker gain makes it possible to achieve a good auditory impression even in the case of a widely spread object.

[0040] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer has a spread angle (or a normalized spread angle or a normalized spread range or a weighted spread range) in a first direction (e.g., spread azi ) and a second direction (spread ele) the product with the spread angle (or normalized spread angle or normalized spread range or weighted spread range) in (at least a predetermined threshold value, for example g res (ranging over a lower spread angle range)) is configured to determine the weighting of the spread object speaker gain in combination with the panned object speaker gain.

[0041] It has been found that the product of the spread angle in the first direction and the spread angle in the second direction well reflects the spread characteristics of the audio object. In particular, when the audio object has a very small spread angle in one of these directions, the product becomes relatively small and the object is considered to be relatively localized. However, it has been found that the product of the spread angle in the first direction and the spread angle in the second direction (which may be perpendicular to the first direction, for example) works well in adjusting the weighting of the spread object speaker gain in combination with the panned object speaker gain and can be determined with very high computational efficiency.

[0042] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer adds a panned object speaker gain (for example, the vector g of the panned object speaker gain) weighted with a fixed weight (for example, 1) and a spread object speaker gain (for example, the vector gOS) (for example, attenGain or g atten ) weighted with a variable weight, which depends on the spread angle (for example, spread azi ) in the first direction and the spread angle (for example, spread ele ) in the second direction, and optionally, for example, it can be limited not to be greater than the fixed weight as shown in the formula for g atten .

[0043] By using such an approach, it can be achieved that the panned object speaker gain has sufficient weight in the combination regardless of the spread angle, but it is still possible to vary the effective (relative) weight between the panned object speaker gain and the spread object speaker gain. However, by ensuring that the contribution of the panned object speaker gain does not fall below a minimum weight, the object can always be appropriately localized regardless of the position of the listener within the speaker environment. This makes it possible to avoid a strong reduction in the auditory impression while maintaining the possibility of adjusting the perceived expansion of the audio object.

[0044] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer is configured to normalize (e.g., by dividing the result of the addition by the norm of the result of the addition) the result of the addition of the panned object speaker gain weighted with a fixed weight and the spread object speaker gain weighted with a variable weight.

[0045] Using such normalization, it can be achieved that the total energy, or total perceived loudness, is substantially independent of the spread of one or more audio objects. Thus, the adjustment of loudness can be performed separately from the adjustment of spread.

[0046] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer combines the panned object speaker gain with the spread object speaker gain weight attenuationGain attenGain = 0.89f * min(c1, max(spread azi , spread ele ) / g res1 ) + 0.11f * min(c2, min(spread azi , spread ele ) / gres2 )、 configured to be determined according to c1 is a predetermined value (for example, 1), and may be used to limit, for example, the first coefficient contributing to attenGain. c2 is a predetermined value (for example, 1), and may be used to limit, for example, the second coefficient contributing to attenGain. g res1 is a predetermined value (for example, the azimuth angle interval of the support point position, for example, 45 degrees, for example, SSP grid resolution), and g res2 is a predetermined value (for example, the elevation angle interval of the support point position, for example, 45 degrees, for example, SSP grid resolution), spread azi is the spread angle of the audio object in the azimuth direction, spread ele is the spread angle of the audio object in the elevation direction, min(.) is the minimum operator, and max(.) is the maximum operator.

[0047] By using such a determination of the spread speaker gain weight in combination with the panned object speaker gain, a particularly good auditory impression can be achieved. In this calculation, the weight depends on both the smaller and the larger spread angles, and the larger spread angle is given a larger weight. Such an evaluation of the two-dimensional spread (i.e., such a determination of the spread object speaker gain weight) results in a relative weighting of the spread object speaker gain and the panned object speaker gain that provides a good auditory impression.

[0048] However, it should be noted that the coefficients 0.89f and 0.11f may be changed, and the first coefficient applied to min(c1, max(spread azi , spread ele ) / g res1 ) should be larger (preferably, at least 50% larger) than the second coefficient.

[0049] In a preferred embodiment of the foregoing audio object renderer, the audio object renderer is configured such that the spread angle of the audio object increases (e.g., the product of the spread angle in the azimuth direction and the spread angle in the elevation direction increases) until the spread angles in the azimuth direction and the elevation direction reach predetermined values, and the relative contribution of the spread object speaker gain compared to the panned object speaker gain increases.

[0050] As a result of such a concept, for example, both the spread angle in the first direction (e.g., the azimuth direction) and the spread angle in the second direction (e.g., the elevation direction) are considered, and contribute to the weight of the contribution of the spread object speaker gain in the above-described combination of the panned object speaker gain and the spread object speaker gain. It has been found that a particularly good auditory impression can be obtained. However, by applying some limitations (which can be defined by "predetermined values"), it is still possible to avoid using excessive weighting for the spread object speaker gain that would reduce the possibility of localizing the sound source.

[0051] In a preferred embodiment of the foregoing audio object renderer, the audio object renderer is configured to obtain a spread object speaker gain (either designated as the spread speaker gain or represented as the vector gOS) using a representation of the support point position in polar coordinates (e.g., aziSSP and / or eleSSP) in consideration of the object position information and the spread information, and the audio object renderer is configured to provide a speaker gain based on the spread object speaker gain.

[0052] The representation of the support point position in polar coordinates is typically known to strongly smooth the calculations, because it is very useful for representing the spread in terms of the spread in the azimuthal direction and the spread in the elevation direction. Further, when using polar coordinates, it is often not necessary to consider the radius component, since it is often appropriate to assume a predetermined radius with respect to the support point position. Thus, for example, it is possible to represent the fulcrum using only two polar coordinates, the azimuth angle and the elevation angle. Thus, the computational complexity is typically reduced.

[0053] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer - evaluates one or more angular differences between the azimuthal position of the audio object and one or more support points (e.g., diffCLKDir, diffAntiCLKDir, and optionally diffCLKDir+(n-1)*Spread.openAngle and diffAntiCLKDir+(n-1)Spread.openAngle), - evaluates one or more angular differences between the elevation position of the audio object and the elevation positions of one or more support points (e.g., diffCLKDir, diffAntiCLKDir), thereby being configured to obtain the spread speaker gain.

[0054] It has been found that evaluating such an angular difference can be performed with a very small amount of computational effort. Also, the angular difference can be mapped, with a medium amount of computational effort, to a value or gain value associated with a support point, for example, describing how an audio signal of an audio object should be rendered at the support point. Further, in a computationally efficient implementation, the support points can be arranged in a very regular configuration such that, for example, they are the same for different elevation angles for which the azimuth difference to be evaluated is different, and also, for example, they are the same for different azimuth angles for which the elevation difference to be evaluated is different. By using such a concept, the azimuth difference only needs to be evaluated for one elevation angle, and the elevation difference only needs to be evaluated for one azimuth angle, so particularly high computational efficiency can be achieved. Therefore, the concept described here is useful for reducing the amount of computation, which facilitates the implementation of the concept in devices with small computational resources.

[0055] In a preferred embodiment of the aforementioned audio object renderer, the positions of the support points are arranged on a sphere within an allowable range of ±10% or ±20% of the radius of the sphere.

[0056] By using support points arranged on a sphere, typically it is not necessary to explicitly consider the radius. Also, it has been found that such support point positions are particularly useful when the extent of the audio object is expressed in terms of angular values (for example, elevation and azimuth values). In contrast, using Cartesian coordinates would result in a significant increase in the amount of computation because it is necessary to consider three spatial coordinates and typically it is necessary to apply computationally expensive trigonometric functions in the calculation to transition between Cartesian coordinates and angular representation.

[0057] In a preferred embodiment of the aforementioned audio object renderer, the support point positions include a uniform azimuth interval along a circle having a constant elevation angle (e.g., elevation angle -135 degrees or -90 degrees or -45 degrees or 0 degrees or 45 degrees or 90 degrees or 135 degrees) and a constant radius (e.g., a normalized radius of 1), or further along a plurality of circles having different constant elevation angles and constant radii (e.g., at 45 degrees).

[0058] Alternatively, or in addition thereto, the support point positions include a uniform elevation interval along a circle having a constant azimuth angle (e.g., azimuth angle -135 degrees or -90 degrees or -45 degrees or 0 degrees or 45 degrees or 90 degrees or 135 degrees) and a constant radius (e.g., a normalized radius of 1), or further along a plurality of circles having different constant azimuth angles and constant radii (e.g., at 45 degrees).

[0059] By using a uniform azimuth interval along a circle having a constant elevation angle and a uniform elevation interval along a circle having a constant azimuth angle, the computational amount can be reduced. Further, when there is a uniform azimuth interval along a plurality of circles having different constant elevation angles and constant radii, for the plurality of circles having different elevation angles and constant radii, since it is not necessary to repeat a significant part of the calculation when all have a uniform azimuth interval, the computational amount can be further reduced while maintaining a good coverage of the entire surface of the sphere (spherical surface) as it is. For example, the calculation of the azimuth difference only needs to be calculated for one of the circles, and the result can also be applied to other circles having the same uniform azimuth interval such as the first circle under consideration. Also, this applies to a plurality of circles having different constant azimuth angles and constant radii and preferably having the same uniform elevation interval. The elevation difference only needs to be calculated once, and the result can be carried over to the support point positions on another circle having the same uniform elevation interval. As a result, calculations can be performed for a large number of support points with a very small computational amount.

[0060] In a preferred embodiment of the aforementioned audio object renderer, the object renderer extends within a first hemisphere in which the audio object is disposed and also extends over a region that extends within a second hemisphere on the opposite side of the azimuth position from the first hemisphere, and is configured to obtain a spread object speaker gain such that the audio object is spread. For example, the first hemisphere may be a hemisphere in front of the listener's position having an azimuth angle between, for example, -90 degrees and +90 degrees, where 0 degrees is the direction of the listener's line of sight. For example, the second hemisphere may be a hemisphere behind the listener's position having an azimuth angle between, for example, -180 degrees and -90 degrees, or 90 degrees and 180 degrees, or vice versa.

[0061] By spreading the audio object over a region that extends within the first hemisphere and also extends within the second hemisphere, the object can be effectively spread above or below the user's head. Accordingly, an auditory impression can be achieved that gives the listener the impression that an enlarged object that is as large in front of the head as behind the head is above the head. Accordingly, a particularly realistic auditory impression can be provided.

[0062] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer is configured to use an extended elevation angle range between -180 degrees and +180 degrees.

[0063] By using an extended elevation range, discontinuities can be avoided and calculations can be simplified, thus reducing the need for many case distinctions and angle corrections. For example, it is much simpler to state (and represent mathematically or algorithmically) that an object should be rendered for a given azimuth and elevation angles between 80 and 100 degrees, compared to showing that the object should be rendered for an elevation angle of 80 degrees at two opposite azimuths, e.g., 0 and 180 degrees. Thus, by doubling the elevation range, which usually extends from -90 to +90 degrees, the calculations can be simplified and the representation of the calculation results can also be simplified.

[0064] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer, for a given object position (e.g., defined by an azimuth value azi and an elevation value ele) and a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle), - A first set of azimuth gain values (e.g., aziGain) that describe the contribution to a spread gain (e.g., g_spd) for a plurality of azimuth values associated with a support point position or a support point azimuth index (e.g., naz) that is associated with an elevation value in the original elevation value range that does not show the intersection of the poles of the spherical coordinate system (e.g., from -90 to +90 degrees), (e.g., using a polynomial or parabolic function whose width is adapted to the object spread width in the azimuth direction), and - For example, an extended elevation value range (e.g., from -180 degrees to -90 degrees and from +90 degrees to +180 degrees) that indicates an intersection of one of the poles of the spherical coordinate system (e.g., an intersection of the pole at an elevation angle of -90 degrees or an intersection of the pole at an elevation angle of +90 degrees) corresponding to the spread of the audio object over one of the poles of the spherical coordinate system, associated with elevation values within this range, and associated with a support point position or a support point azimuth index (e.g., naz) (e.g., using a polynomial function or a parabolic function whose width is adapted to the object spread width in the azimuth direction), describing the contribution to the spread gain for a plurality of azimuth values configured to calculate.

[0065] The audio object renderer is also configured to derive a spread gain (e.g., used in the determination of the spread object speaker gain gOS) using a first set of azimuth gain values (e.g., aziGain(naz)) and also using a second set of azimuth gain values (e.g., aziGainExtd).

[0066] By calculating two sets of azimuth gain values, namely the azimuth gain values for the original (or basic) elevation value range (or associated therewith) and the azimuth gain values for the extended elevation value range (or associated therewith), the spread of the audio object above the listener's head (or below the listener) can be calculated particularly efficiently. In particular, it has been found that the sets of azimuth gain values can be used in an efficient manner to derive the spread gain, for example, using a pre-defined combinatorial mapping. On the other hand, the elevation values within the extended elevation value range can be easily derived from the elevation values within the original elevation value range using addition or subtraction (e.g.) without multiple case distinctions and without changing the azimuth values, thus smoothing the calculation and the representation of the results.

[0067] For example, when spreading an audio object above the user's head, the spread above the user's head can be easily calculated by using an elevation angle value greater than 90 degrees. Thus, an object having a given azimuth angle (e.g., an azimuth angle within the range from -90 degrees to +90 degrees) and a positive elevation angle between 0 degrees and 90 degrees can be easily extended above the user's head without changing the azimuth angle (within the range from -90 degrees to +90 degrees) by using an elevation angle greater than 90 degrees. Thus, the elevation gain values associated with the elevation angle values in the extended elevation angle value range (in this example, between +90 degrees and +180 degrees) can be obtained as medium quality and then, for example, mapped back towards the support point or towards the coordinate system using only the elevation angles within the original elevation angle range in combination with, for example, the azimuth gain values of a second set of azimuth gain values. In conclusion, the concepts described significantly improve the computational efficiency when spreading an audio object above (or below) the user's head.

[0068] In a preferred embodiment of the aforementioned audio object renderer, for a given object position (e.g., defined by an azimuth value azi and an elevation value ele) and a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle), - A first set of elevation gain values (e.g., eleGain) that describe the contribution to the spread gain for a plurality of elevation values associated with a support point position or a speaker azimuth index or a support point elevation index (e.g., nel) that is associated with the elevation values in the original elevation angle range (e.g., from -90 degrees to +90 degrees) that does not show the intersection of the poles of the spherical coordinate system, and that uses, for example, a parabola function whose width is adapted to the object spread width in the azimuth direction, and - For example, an extended elevation value range (e.g., from -180 degrees to -90 degrees and from +90 degrees to +180 degrees) indicating an intersection of one of the poles of the spherical coordinate system (e.g., an intersection of the pole at an elevation angle of -90 degrees or an intersection of the pole at an elevation angle of +90 degrees) that can correspond to the spread of the audio object over one of the poles of the spherical coordinate system, associated with elevation values within this range, a support point position or a speaker elevation indicator or a support point elevation indicator (e.g., nel) (e.g., using a parabolic function whose width is adapted to the object spread width in the azimuth direction) describing the contribution to the spread gain for a plurality of elevation values associated with it, a second set of elevation gain values (e.g., eleGainExtd) is configured to calculate.

[0069] The audio object renderer is also configured to derive the spread gain using a first set of azimuth gain values (aziGain(naz)), using a second set of azimuth gain values (aziGainExtd), using a first set of elevation gain values (eleGain(nel)), and using a second set of elevation gain values (eleGainExtd(nel)).

[0070] This embodiment is based on considerations similar to those of the embodiment for calculating the azimuth gain value.

[0071] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer is also configured to (e.g., multiplicatively) combine the values of the first set of azimuth gain values and the first set of elevation gain values (e.g., corresponding values, e.g., aziGain(naz), eleGain(nel)), and to combine the (e.g., corresponding) values of the second set of azimuth gain values and the second set of elevation gain values (e.g., aziGainExtd(naz), eleGainExtd(nel)).

[0072] Using such combinations, values associated with the same location can be combined (e.g., summed). For example, each support point location can be referenced by a combination of a first azimuth value and a first elevation value within the original elevation value range, and also by a combination of a second azimuth value and a second elevation value within the extended elevation value range. For example, a point (e.g., a support point location) can be specified at an azimuth of +10 degrees at an elevation of +80 degrees and can also be specified by an azimuth of -170 degrees and an elevation of +100 degrees. In other words, it is possible to combine values, such as azimuth gain values and elevation gain values, that are associated with the same point but are referenced by different combinations of azimuth values and elevation values.

[0073] For example, the second set of azimuth gain values can include the same azimuth gain values that are included in the first set of azimuth gain values but are associated with the opposite direction (or azimuth value). For example, the second set of azimuth gain values can exhibit the same gain at an azimuth of -170 degrees (or, generally, x - 180 degrees or x + 180 degrees) as the first set of azimuth gain values exhibits at an azimuth of +10 degrees (or, generally, x degrees). However, the actual location designed by an azimuth of +10 degrees and an elevation of +80 degrees (or, generally, y degrees) is actually the same as the location specified by an azimuth of -170 degrees and an elevation value within the extended elevation value range of +100 degrees (or, generally, 180 degrees - y). As a result, the elevation gain values associated with elevation values within the extended elevation value range should be combined with the azimuth gain values associated with the "opposite" azimuth. Thus, the second set of azimuth gain values actually describes the contribution to the spread gain for the "opposite" azimuth value.

[0074] As a result, by combining the values of the azimuth gain values of the first set and the values of the elevation gain values of the first set, and also by combining the values of the azimuth gain values of the second set and the values of the elevation gain values of the second set, two pairs of values associated with the same position (described by different azimuth values and elevation values) are combined, thereby obtaining a meaningful gain value associated with the support point.

[0075] In a preferred embodiment of the aforementioned audio object renderer, the azimuth gain values of the second set represent a progressive change in the gain value over an azimuth that is shifted by 180 degrees when compared to the progressive change in the gain value over the azimuth represented by the azimuth gain values of the first set.

[0076] Using such an expression of the azimuth gain value and using two sets of azimuth gain values, it is conceivable that when considering the modified azimuth, the elevation value within the extended elevation value range specifies the same point as the elevation value within the original elevation value range. Therefore, by using two sets of azimuth gain values that represent the progressive change in the gain value over two shifted azimuth ranges, the elevation within the extended azimuth range can be easily and computationally efficiently handled because the "azimuth gain values of the second set" can be easily combined with the elevation gain value associated with the extended azimuth range.

[0077] In a preferred embodiment of the aforementioned audio object renderer, the azimuth gain values of the first set represent a progressive change in the gain value over a range of 360 degrees, taking into account the azimuth object position and the azimuth spread angle having an angular accuracy determined by the number of speakers or the number of support points.

[0078] Alternatively, or in addition, the second set of azimuth gain values represents a gradual change in gain values over a 360-degree range, taking into account the azimuth object position rotated by 180 degrees and the azimuth spread angle having an angular accuracy determined by the number of speakers or the number of support points.

[0079] By using a set of azimuth gain values that represents a gradual change in gain values over a 360-degree range, the listener's complete environment is efficiently considered, and both the audio object in front of the listener and the audio object behind the listener can be taken into account. By using such a set of azimuth gain values that represents a gradual change in gain values over a 360-degree range, it is also possible to handle the spread of audio objects above or below the user's head.

[0080] In a preferred embodiment of the aforementioned audio object renderer, the first set of elevation gain values represents a gradual change in gain values over an elevation range from -90 degrees to +90 degrees, taking into account the elevation object position (in the range between -90 degrees and +90 degrees) and the elevation spread angle (for example, when the audio object does not spread over the pole of the spherical coordinate system).

[0081] Alternatively, or in addition, the second set of elevation gain values represents a gradual change in gain values over an elevation range from -180 degrees to -90 degrees and from +90 degrees to +180 degrees, taking into account the elevation object position (in the range between -90 degrees and +90 degrees) and the elevation spread angle (for example, when the audio object spreads over the pole of the spherical coordinate system).

[0082] By using such a set of elevation gain values, the angular values can be simply added or subtracted without exceeding the elevation range covered by the first set of elevation gain values and the second set of elevation gain values, so that the spread of objects above the user's head can be easily handled.

[0083] One embodiment according to the present invention creates a method for determining a speaker gain (e.g., a combined speaker gain or the resulting speaker gain) that describes the gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele)), which can be given, for example, in spherical coordinates (e.g., using the azimuth value azi and the elevation value ele), and object feature information (e.g., spread angle information that describes the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information that describes the spread in the elevation direction, e.g., spreadAngleEle).

[0084] This method includes obtaining a panned object speaker gain (also designated as "object speaker gain" or represented by the vector g) using point source panning of the audio object (where the audio object is considered a point source, spread information is ignored, and the signal of the audio object is associated with two or more speakers within the environment of the object position of the audio object by appropriate selection of the panned object speaker gain).

[0085] This method also includes obtaining a spread object speaker gain (also designated as spread speaker gain or represented as the vector gOS) considering the object position information and the object feature information.

[0086] This method includes combining the panned object speaker gain (e.g., g) and the spread object speaker gain (e.g., gOS) such that the contribution of the panned object speaker gain is always present (independent of the spread information) to obtain a combined speaker gain.

[0087] This method is based on the same considerations as the corresponding device described above.

[0088] Furthermore, this method can optionally be complemented by any of the features, functions, and details described, either individually or in combination, with respect to the corresponding apparatus described above.

[0089] One embodiment according to the present invention creates a method for determining a speaker gain (e.g., a combined speaker gain or the resulting speaker gain) that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele)), which can be given, for example, in spherical coordinates, using, for example, azimuth value azi and elevation value ele, and spread information (e.g., spread angle information describing the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, e.g., spreadAngleEle).

[0090] This method includes obtaining a panned object speaker gain (also designated as "object speaker gain" or represented by vector g), which uses point-source panning of the audio object (e.g., the audio object is considered a point source, the spread information is ignored, and the signal of the audio object is associated with two or more speakers within the environment of the object position of the audio object by appropriate selection of the panned object speaker gain).

[0091] This method includes obtaining a spread object speaker gain (also designated as spread speaker gain or represented as vector gOS), taking into account the object position information and the spread information.

[0092] This method involves combining a panned object speaker gain (e.g., g) and a spread object speaker gain (e.g., gOS) such that the contribution of the panned object speaker gain is always present (independent of the spread information) to obtain a combined speaker gain.

[0093] This method is based on the same considerations as the corresponding apparatus described above.

[0094] Furthermore, this method can optionally be complemented by any of the features, functions, and details described, either individually or in combination, with respect to the corresponding apparatus described above.

[0095] In a preferred embodiment of the method described above, the method includes evaluating one or more gain functions (e.g., one or more polynomial functions or one or more parabolic functions, e.g., a spread weighting curve) that map the difference between the position of a support point (e.g., an artificially generated support point, SSP) and the object position to one or more spread gain value contributions (e.g., aziGain(naz) or eleGain(nel)), and determining a spread object speaker gain (e.g., gOS) based on the one or more spread gain value contributions.

[0096] One embodiment according to the present invention creates a computer program for performing the method described above when the computer program is executed on a computer.

[0097] The computer program can be complemented by any of the features, functions, and details described herein, either individually or in combination.

[0098] Next, further embodiments according to the present application are described. These embodiments can be used individually or in combination with any of the other embodiments disclosed herein. In other words, optionally, any of the features, functions, and details of the embodiments described next can be introduced into any other embodiment disclosed herein, optionally, both individually and in combination.

[0099] One embodiment according to the present invention creates an audio object renderer (200, 1300) for determining speaker gains (214, 1214, 1214a - 1214c) that describe gains for including one or more audio object signals (1260) into a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1310, azi, ele) and object feature information (1312). The object renderer is configured to obtain object feature information gains (314a, g_spd) using one or more polynomial functions having a degree of 3 or less.

[0100] In a preferred embodiment, the object renderer is configured to obtain object feature information speaker gains (206a, 1242, gOS) using object feature gains (314a, g_spd) based on object feature gain contributions (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd).

[0101] In a preferred embodiment, the object feature information is spread information (212, 1312, spreadAngleAzi, spreadAngleEle).

[0102] One embodiment according to the present invention creates an audio object renderer for determining a speaker gain (e.g., a combined speaker gain) that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele)), which can be given, for example, in spherical coordinates, using, for example, an azimuth value azi and an elevation value ele, and object feature information. The object feature information can be, for example, information indicating whether the object is small or extended, e.g., an object size value, or the object feature information can be, for example, object distance information that can be mapped onto a spread value (e.g., spread angle information describing the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, e.g., spreadAngleEle). However, other types of object feature information are also possible.

[0103] The object renderer is configured to obtain a spread object speaker gain (also designated as a spread speaker gain or represented as a vector gOS) considering the object position information and the object feature information.

[0104] The object renderer is configured to obtain a spread gain (e.g., object-to-anchor spread gain, e.g., g_spd), such as an element of the vector g_spd, for example, by using one or more polynomial functions of degree 3 or less, such as one or more parabolic functions or polynomial functions of degree 3 (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1), that map an angular difference (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) between an object position and an anchor position to a spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)) to describe the contribution of an audio object signal to a plurality of speaker signals or a plurality of anchor signals.

[0105] The object renderer is configured to obtain a spread object speaker gain using a spread gain (g_spd) based on the spread gain contribution.

[0106] This embodiment is based on the finding that a polynomial function having a degree of 2 or less is particularly suitable for obtaining an object spread gain based on the angular difference between the object position and the support point position. It has been found that a polynomial function having a degree of 3 or less can be evaluated with a moderate amount of calculation and can be suitably used in a situation where calculation resources are limited. However, it has been found that such a polynomial function still approximates well the characteristics necessary for obtaining an object spread gain that provides a good auditory impression of a spread audio object. In particular, it has been found that a polynomial function having a degree of 3 or less can be evaluated much more easily than other functions that require a very large amount of calculation or a large look-up table, such as an exponential function. Therefore, the audio object renderer can be implemented with particularly little calculation amount.

[0107] Furthermore, it should be noted that the object feature information can be used to adjust the determination of the spread object speaker gain. For example, the object feature information can determine the spread width. For example, the object feature information can be information indicating whether the object is small or expanded, and can include, for example, an object size value, or the object feature information can be mapped, for example, on a spread value, such as spread angle information describing the spread in the azimuth direction, for example, on spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, for example, on spreadAngleEle, and can include object distance information. In other words, it should be noted that the object feature information typically enables the determination or estimation of the expansion of the audio object. For example, the object feature information may indicate the type of the object, and this type of object may be meant as a parameter (e.g., a spread parameter) for the determination of the spread object speaker gain. For example, the object feature information can enable the distinction between a relatively small object and a relatively large object. Alternatively, or in addition, the object feature information enables the distinction between a near object and a far object, which can mean one or more parameters for the determination of the spread object speaker gain.

[0108] In conclusion, the audio object renderer can derive one or more parameters for the determination of the spread object speaker gain from the object feature information. Therefore, the object feature information enables the appropriate adjustment of the derivation of the spread object speaker gain so that a good auditory impression can be achieved.

[0109] As a conclusion, by using one or more polynomial functions having a degree of 3 or less and whose parameters may depend on, for example, object feature information, it is possible to efficiently determine the spread object speaker gain while moderately suppressing the amount of calculation.

[0110] However, it should be noted that the above approach can be applied in an optionally more general form. In particular, it may not be necessary to perform object spreading. For example, when support points are pre-rendered in order to render object characteristics (or object characteristics or object features) by applying a weighting curve to the support points, this weighting curve could be implemented, for example, by a polynomial of degree 3 for reasons of efficiency.

[0111] One embodiment according to the present invention creates an audio object renderer for determining a speaker gain (e.g., a combined speaker gain) that describes the gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele)), which can be given, for example, in spherical coordinates using, for example, an azimuth value azi and an elevation value ele, and spread information (e.g., spread angle information describing the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, e.g., spreadAngleEle).

[0112] The object renderer is configured to obtain a spread object speaker gain (also designated as a spread speaker gain or represented as a vector gOS) in consideration of the object position information and the spread information.

[0113] The object renderer is configured to obtain a spread gain (e.g., object-to-anchor spread gain, e.g., g_spd), e.g., a spread gain value (e.g., an element of the vector g_spd), using one or more polynomial functions of degree 3 or less, e.g., a parabola function or a polynomial function of degree 3 (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1), which map the angular difference between the object position and the anchor position (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) to the spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)) to describe the contribution of the audio object signal to a plurality of speaker signals or a plurality of anchor signals.

[0114] The object renderer is configured to obtain a spread object speaker gain using a spread gain [g_spd] based on the spread gain contribution.

[0115] Generally speaking, the function used (e.g., the polynomial function) preferably (but not necessarily) has the curve shape of point source panning (VBAP in the embodiments of the present invention) (at least approximately). This function does not necessarily have to be a parabola. This is the idea that the spread component is panned in the same way between the anchors (such that an object is usually panned between two speakers, e.g., using VBAP).

[0116] This audio object renderer is based on the same considerations as the above-described audio object renderer. However, instead of object feature information, spread information that directly describes how the object should be spread is evaluated. For example, the spread information may be spread angle information that describes the spread in the azimuth direction and / or the spread in the elevation direction. Alternatively, the spread information may also be solid angle information or may specify the size of the object in other forms (e.g., using absolute size information and / or distance information, or the like). Therefore, it is possible to obtain a spread object speaker gain in a computationally efficient manner, and the spread information may be used, for example, to adjust the parameters of one or more polynomial functions (e.g., the width of the parabola used to obtain the spread gain). Therefore, as indicated by the spread information, the calculation can be easily adjusted according to the actual spread, and the calculation can be performed in an efficient manner.

[0117] In a preferred embodiment of the above-described audio object renderer, the width of one or more polynomial functions (e.g., a parabola function that can be determined by the scaling value aziParable or the scaling value eleParable) is determined by the spread information or by the object feature information (the audio object renderer may be configured, for example, to adapt the width of one or more polynomial functions or parabola functions to the spread widths associated with different audio objects).

[0118] Since polynomial functions can be easily parameterized, it has been found that the width of a polynomial function (e.g., a parabolic function) can be easily adjusted. However, the evaluation of a parabolic function (e.g., a polynomial function) with one or more parameters is typically possible without excessive computational effort. Further, by adjusting one or more parameters of the parabolic function according to spread information or according to object feature information, the spread width can be adjusted very smoothly, and as a result, the perceptual impression can be very good.

[0119] In a preferred embodiment of the audio object renderer described above, the object renderer is configured to obtain a spread gain value (e.g., the value of the vector g_spd) using a first polynomial function (e.g., a polynomial function having a degree of 3 or less, e.g., a parabolic function) that maps the azimuth difference between the object position and the support point position to a first spread gain value contribution (e.g., aziGain(naz)), and using a second polynomial function (e.g., a polynomial function having a degree of 3 or less, e.g., a parabolic function) that maps the elevation difference between the object position and the support point position to a second spread gain value contribution (e.g., eleGain(nel)).

[0120] This concept is based on the idea that a two-dimensional spread function, which can depend on both, for example, the elevation difference between the audio object position and the spread support point position, and the azimuth difference between the audio object position and the spread support point position, can be efficiently determined using a combination (e.g., multiplication) of two spread functions, one applied to the azimuth difference and one applied to the elevation difference. In other words, it has been found that when two separate parabolic spread functions are applied to the azimuth difference and the elevation difference and the results are multiplied, a good spread result can be achieved. In particular, it has been found that such a type of two-dimensional spread result provides a moderately good auditory impression while keeping the computational amount small. In particular, the separate evaluation of the two spread functions typically requires a significantly smaller computational amount compared to the evaluation of the combined two-dimensional function while providing a good auditory impression.

[0121] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer is configured to obtain a spread gain value (e.g., the value of vector g_spd) by combining (e.g., multiplicatively combining) a first spread gain contribution (e.g., aziGain(naz)) and a second spread gain contribution (e.g., eleGain(nel)).

[0122] By combining the spread gain contributions, which can be performed using, for example, multiplication, a two-dimensional spread can be achieved that well approximates the two-dimensional spread by the spread function shown in FIG. 3. In other words, it has been found that multiplying the two spread gain contributions obtained based on two polynomial functions results in two-dimensional spread characteristics that provide a moderately good auditory impression.

[0123] In a preferred embodiment of the aforementioned audio object renderer, for a given object position (e.g., defined by an azimuth value azi and an elevation value ele) and a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle), the object renderer - A set of azimuth gain values (e.g., aziGain) that describe the contribution to the spread gain for a plurality of azimuth values associated with a support point position or a speaker azimuth index or a support point azimuth index (e.g., naz) (using, for example, a polynomial function of degree 3 or less or a parabola function whose width is adapted to the object spread width in the azimuth direction), and / or - A set of elevation gain values (e.g., eleGain) that describe the contribution to the spread gain for a plurality of elevation values associated with a support point position or a speaker elevation index or a support point elevation index (e.g., nel or naz) (using, for example, a polynomial function of degree 3 or less or a parabola function whose width is adapted to the object spread width in the elevation direction), configured to calculate and derive the spread gain using a set of azimuth gain values (aziGain(naz)) and / or using elevation gain values (eleGain(nel)).

[0124] By determining a set of azimuth gain values associated with a plurality of support point positions and / or a plurality of elevation gain values associated with a support point position, the spreading characteristics in one or two planes are determined, and the spread gain can be derived from this spreading characteristic in one or two planes. In a preferred case where both a set of azimuth gain values associated with a support point position and a set of elevation gain values associated with a support point position are determined, the spread gain can be easily obtained, for example, using the multiplication of pairs of elements (of these sets) associated with the respective azimuth and respective elevation of the support points under consideration. Thus, the polynomial function only needs to be evaluated for one set of support point positions having the same azimuth and different elevations, and one set of support point positions having the same elevation and different azimuths, and then the spread gain can be derived using, for example, a computationally simple multiplication of appropriate elements of the set of azimuth gain values and the set of elevation gain values (for all support points). Thus, high computational efficiency can be achieved.

[0125] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer combines an element (e.g., aziGain(naz)) of a set of azimuth gain values (e.g., aziGain) associated with the currently considered speaker or the currently considered support point (e.g., specified by objNo and having associated values naz and nel) with an element (e.g., eleGain(nel)) of a set of elevation gain values (e.g., eleGain) associated with the currently considered speaker or the currently considered support point, to obtain a spread gain value (e.g., represented by the vector g_spd, g_spd(objNo)) associated with a plurality of different speakers or a plurality of different support points (e.g., specified by different values of objNo).

[0126] Accordingly, by determining the spread gain values associated with different support points using the multiplication of each element of the set of azimuth gain values and each element of the set of elevation gain values (e.g., the elements associated with the azimuth and elevation of each support point), the spread values associated with different support points can be obtained in a computationally efficient manner. Particularly high efficiency can be achieved when a plurality of support points include the same azimuth value, and when a plurality of support points include the same elevation value. In this case, the number of elements of the set of azimuth gain values and the number of elements of the set of elevation gain values can be kept moderately small, and the elements of the set of azimuth gain values and the elements of the set of elevation gain values can be reused in the determination of the spread values associated with a plurality of support points. In other words, such a concept is particularly efficient when combined with a uniform spacing (with respect to azimuth and elevation) of the support points.

[0127] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer is configured for a given object position (e.g., defined by an azimuth value azi and an elevation value ele) and a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle) to - a first set of azimuth gain values (e.g., aziGain) that describe the contribution to the spread gain for a plurality of azimuth values associated with a support point position or a speaker azimuth index or a support point azimuth index (e.g., naz) that is associated with an elevation value within the original elevation value range that does not show the intersection of the poles of the spherical coordinate system (e.g., from -90 degrees to +90 degrees), (e.g., using a polynomial function having a degree of 3 or less, or using a parabolic function whose width is adapted to the object spread width in the azimuth direction), and - a second set of azimuth gain values (e.g., aziGainExtd) that describe the contribution to the spread gain for a plurality of azimuth values associated with a support point position or a speaker azimuth index or a support point azimuth index (e.g., naz) that is associated with an elevation value within the extended elevation value range that shows the intersection of the poles of the spherical coordinate system (e.g., from -180 degrees to -90 degrees and from +90 degrees to +180 degrees), (e.g., using a polynomial function having a degree of 3 or less, or using a parabolic function whose width is adapted to the object spread width in the azimuth direction), calculate and is configured to derive the spread gain using the set of azimuth gain values (aziGain(naz)) and / or using the elevation gain values (eleGain(nel)) (or using the second set of azimuth gain values).

[0128] By calculating two sets of azimuth gain values, namely the azimuth gain value for the original (or basic) elevation value range and the azimuth gain value for the extended elevation value range, the spread gain for the spread of the audio object above (or below) the listener's head can be calculated particularly efficiently. In particular, it has been found that the sets of azimuth gain values can be used in an efficient manner to derive the spread gain, for example, using a predefined combinatorial mapping. On the other hand, the elevation values within the extended elevation value range can be easily derived from the elevation values within the original elevation value range using addition or subtraction (for example) without making a plurality of case distinctions and without changing the azimuth value, so that separate sets of elevation gain values can also be calculated efficiently. For example, when spreading an audio object above the user's head, the spread above the user's head can be easily calculated by using elevation values greater than 90 degrees. Thus, an object having a given azimuth (e.g., an azimuth within the range from -90 degrees to +90 degrees) and a positive elevation between 0 degrees and 90 degrees can be easily spread above the user's head while maintaining the azimuth (within the range from -90 degrees to +90 degrees) using an elevation greater than 90 degrees. Thus, the azimuth gain value associated with the elevation value in the extended elevation value range (in this example, between +90 degrees and +180 degrees) is obtained as an intermediate quantity and can then be inverse mapped, for example, towards the support point or towards the coordinate system using only the elevations within the original elevation value range. The existence of the first set of azimuth gain values and the second set of azimuth gain values allows the second set of azimuth gain values to be adapted to be combined with the elevation gain values within the extended elevation value range, so that the spread value can be efficiently derived.

[0129] In a preferred embodiment of the aforementioned audio object renderer, for a given object position (defined, for example, by an azimuth value azi and an elevation value ele) and a given spread (defined, for example, by spreadAngleAzi or spreadAngleEle), the audio object renderer - A first set of elevation gain values (e.g., eleGain) that describe the contribution to the spread gain for a plurality of elevation values associated with an elevation value within the original elevation value range that does not show the intersection of the poles of the spherical coordinate system (e.g., from -90 degrees to +90 degrees), (e.g., using a polynomial function having a degree of 3 or less, or using a parabola function whose width is adapted to the object spread width in the azimuth direction) and is associated with a support point position or a speaker azimuth index or a support point elevation index (e.g., nel), and - A second set of elevation gain values (e.g., eleGainExtd) that describe the contribution to the spread gain for a plurality of elevation values associated with an elevation value within an extended elevation value range that shows the intersection of the poles of the spherical coordinate system (e.g., from -180 degrees to -90 degrees and from +90 degrees to +180 degrees), (e.g., using a polynomial function having a degree of 3 or less, or using a parabola function whose width is adapted to the object spread width in the azimuth direction) and is associated with a support point position or a speaker elevation index or a support point elevation index (e.g., nel), configured to calculate and use a set of azimuth gain values (aziGain(naz)) and use the elevation gain values (eleGain(nel)) (or use the first set of azimuth gain values, the second set of azimuth gain values, the first set of elevation gain values, and the second set of elevation gain values) to derive the spread gain.

[0130] This embodiment is based on considerations similar to those of the embodiment for calculating the azimuth gain value. In particular, the existence of the first set of elevation gain values and the second set of elevation gain values can make the handling when the object spreads over the listener's head simple and computationally efficient. In particular, it has been found that it is much easier to use such extended elevation values (e.g., greater than +90 degrees) in such cases compared to the immediate correction of the azimuth value and the "conversion" of the elevation value.

[0131] Although just an example, when the object position includes an elevation angle of +80 degrees, for the purpose of further calculations and for the purpose of evaluating the parabolic function, it is quite easy to assume that the support point is at an elevation angle of, for example, 135 degrees, because it can be said that the elevation angle difference between the audio object and the support point is 55 degrees. In contrast, when the support point at 135 degrees is referred to as the support point at an elevation angle of 45 degrees, it is not easily possible to calculate the correct elevation angle difference between the position of the audio object and the support point position.

[0132] In summary, it has been found that the use of such an extended elevation angle range is very efficient as it avoids many case divisions in the calculations and smoothes the evaluation of the polynomial function. Further, it has been found that the spread gain can be derived, for example, using a first set of azimuth gain values, a second set of azimuth gain values, a first set of elevation gain values, and a second set of elevation gain values. In this case, for example, the entry of the first set of azimuth gain values and the entry of the first set of elevation gain values are combined, and the entry of the second set of azimuth gain values and the entry of the second set of elevation gain values are efficiently combined, thereby deriving the spread value.

[0133] In a preferred embodiment of the aforementioned audio object renderer, the audio object renderer is configured to pre-calculate a support point panning gain (e.g., Spread.gainsSSP) for panning an audio signal associated with a plurality of support points to a plurality of speakers during initialization using panning (e.g., vector-based amplitude panning) (based on knowing the positions of the support points and the speakers).

[0134] The audio object renderer is configured to obtain an object-to-anchor-point spread gain (e.g., g_spd or a spread gain value, e.g., an element of the vector g_spd) that describes the contribution of the audio object signal to a plurality of anchor-point signals using a polynomial function having a degree of 3 or less (e.g., using a parabolic function).

[0135] The audio object renderer is configured to combine (e.g., multiply) the object-to-anchor-point spread gain and the anchor-point panning gain to obtain a spread object speaker gain.

[0136] It has been found that calculating the anchor-point panning gain and the object-to-anchor-point spread gain separately and then combining the object-to-anchor-point spread gain and the anchor-point panning gain provides particularly high computational efficiency, especially when there are multiple audio objects. Since the anchor points typically do not change for the handling of multiple audio objects, the anchor-point panning gain only needs to be calculated once, which can be done in a preparation stage. In contrast, the object-to-anchor-point spread gain typically depends on the object position and thus needs to be calculated separately for each audio object.

[0137] Accordingly, by reusing the support point panning gain for a plurality of objects, the computational efficiency can be improved without degrading the achievable audio quality. Further, the use of support points is particularly efficient because the spatial arrangement of the support points can be freely adjusted with an emphasis on computational efficiency without being restricted by the actual speaker positions or the actual speaker setup. Thus, for example, the support points can be selected to be uniformly distributed (e.g., having a uniform azimuth interval and a uniform elevation interval), which significantly smoothens the determination of the object-to-support point spread gain. Accordingly, the first spreading step can be made independent of the actual speaker arrangement configuration, and the second spreading step (from the spread support points to the speaker signals) only needs to be calculated once even when there are multiple audio objects, so it can be said that the use of the spread support points as an intermediate spreading target is actually useful for improving efficiency. Thus, this process is efficient.

[0138] In a preferred embodiment of the aforementioned audio object renderer, one or more polynomial functions having a degree of 3 or less are of the form p = max(0, c1 * anglediff 2 + c2) which is a parabola function that provides a return value p, where c1 is a parameter that determines the width of the parabola function, c2 is a predetermined value, angeldiff is the angle difference at which the parabola function is evaluated, and max(.,.) is a maximum value operator that returns the maximum value of its operands.

[0139] It has been found that such polynomial functions (e.g., using a max-operation) that are restricted to non-negative values can well approximate the desired spread characteristics and can be evaluated with a small amount of computation. Thus, it has been found that such polynomial functions are excellent for determining spread values.

[0140] One embodiment according to the present invention creates a method for determining a speaker gain (e.g., a combined speaker gain) that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele), which can be given, for example, in spherical coordinates, using, for example, an azimuth value azi and an elevation value ele) and object feature information (e.g., spread angle information that describes the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information that describes the spread in the elevation direction, e.g., spreadAngleEle).

[0141] This method includes obtaining a spread object speaker gain (also designated as a spread speaker gain or represented as vector gOS) in consideration of the object position information and the object feature information.

[0142] This method uses one or more polynomial functions having a degree of 3 or less, such as a parabola function or a polynomial function of degree 3 (e.g., parable*((diffCLKDir+(n - 1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n - 1)*Spread.openAngle).^2)+1), to map the angle difference between the object position and the support point position (e.g., diffCLKDir+(n - 1)*Spread.openAngle or diffAntiCLKDir+(n - 1)*Spread.openAngle) to a spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)), and includes obtaining a spread gain (e.g., an object-to-support point spread gain, e.g., g_spd), such as a spread gain value, e.g., an element of vector g_spd, that describes the contribution of the audio object signal to the plurality of speaker signals or the plurality of support point signals.

[0143] This method includes obtaining a spread object speaker gain based on a spread gain contribution, using a spread gain (e.g., g_spd), or using the spread gain (e.g., g_spd) as a spread object speaker gain for a spread object speaker gain.

[0144] This method is based on the same considerations as the corresponding apparatus described above. Further, this method may optionally be complemented by any of the features, functions, and details described for the corresponding apparatus above, individually or in combination.

[0145] One embodiment creates a method for determining a speaker gain (e.g., a combined speaker gain) that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information (e.g., azimuth (azi), elevation (ele)), which may be given in spherical coordinates, for example, using azimuth value azi and elevation value ele, and spread information (e.g., spread angle information describing the spread in the azimuth direction, e.g., spreadAngleAzi, and / or spread angle information describing the spread in the elevation direction, e.g., spreadAngleEle).

[0146] This method includes obtaining a spread object speaker gain (also designated as a spread speaker gain or represented as vector gOS) in consideration of object position information and spread information.

[0147] This method uses one or more polynomial functions of degree 3 or less, such as a parabola function or polynomial function of degree 3 (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1), which map the angular difference between the object position and the support point position (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) to the spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)), to describe the contribution of the audio object signal to a plurality of speaker signals or a plurality of support point signals, and obtain a spread gain (e.g., object-to-support point spread gain, e.g., g_spd), e.g., a spread gain value, e.g., an element of the vector g_spd.

[0148] This method includes obtaining a spread object speaker gain using a spread gain (e.g., g_spd) based on the spread gain contribution or using the spread gain (e.g., g_spd) as the spread object speaker gain.

[0149] This method is based on the same considerations as the corresponding apparatus described above. Further, this method can optionally be complemented by any of the features, functions, and details described for the corresponding apparatus above, either individually or in combination.

[0150] One embodiment according to the present invention creates a computer program for executing one of the methods described above when the computer program is executed on a computer.

[0151] The computer program can be complemented by any of the features, functions, and details described herein, either individually or in combination.

[0152] Next, embodiments according to the present invention will be described with reference to the accompanying drawings.

Brief Description of the Drawings

[0153]

Figure 1

Figure 2

Figure 3-1

Figure 3-2

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9-1

Figure 9-2

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Mode for Carrying Out the Invention

[0154] 1. Description of Some Embodiments 1.A. Audio Object Renderer According to FIG. 12 FIG. 12 shows a schematic block diagram of an audio object renderer 1200 according to an embodiment of the present invention.

[0155] The audio object renderer 1200 is configured to receive object position information 1210 and object characteristic information 1212. Further, the audio object renderer 1200 is configured to provide a speaker gain 1214 (e.g., a combined speaker gain or a resulting speaker gain) that describes the gain for including one or more audio object signals in a plurality of speaker signals.

[0156] The object position information 1210 may include, for example, object azimuth angle information (e.g., azi) and object elevation angle information (e.g., ele). For example, the object position information may be provided in spherical coordinates using, for example, the azimuth value azi and the elevation value ele. Further, the object feature information or spread information 1212 may, for example, describe the characteristics of the audio object and provide an indication of how the object should be spread. The object feature information may be, for example, information indicating whether the object is small or expanded, such as an object size value. The object feature information may, for example, include object distance information, which can be mapped to the spread value, for example, the spread angle information describing the spread in the azimuth direction, such as spreadAngleAzi, and / or the spread angle information describing the spread in the elevation direction, such as spreadAngleEle. However, different types of object feature information are also possible. Alternatively, the audio object renderer may directly receive, for example, spread information describing the spread of the audio object in the azimuth direction and / or the elevation direction.

[0157] The audio object renderer 1200 includes a panned object speaker gain determination 1230, which may be configured to obtain a panned object speaker gain 1232 (also designated, for example, as "object speaker gain" or represented by the vector g) using point source panning of the audio object. In point source panning, the audio object may be considered, for example, as a point source, spread information or object characteristic information 1212 may be ignored, for example, and the signal of the audio object is associated with two or more speakers within the environment of the object position of the audio object by an appropriate selection of the panned object speaker gain 1232. In other words, the panned object speaker gain determination 1230 may perform, for example, point source panning of the audio object considering the audio object as a point source and distribute the audio object signal to the speaker (only) closest to the audio object. However, with point source panning, no contribution of the audio object signal can be allocated to speakers that are further away from the audio object. However, the panned object speaker gain determination 1230 may use any point source panning concept.

[0158] Furthermore, the audio object renderer 1200 may include a spread object speaker gain determination 1240 that provides a spread object speaker gain 1242 (also designated, for example, as a spread speaker gain and represented, for example, as a vector gOS) in consideration of the object position information 1210 and the object characteristic information or spread information 1212. For example, the spread object speaker gain determination 1240 may determine a speaker gain that takes into account, for example, the spread of the audio object in the azimuth and elevation directions. Thus, the audio object under consideration is assumed to have a significant spread, and typically, in the spread, a smooth and steady attenuation of the spread object speaker gain that occurs with an increase in the distance from the center position of the object is assumed (and implemented), so the spread object speaker gain can be non-zero over a wide range.

[0159] Furthermore, the audio object renderer 1200 includes a combiner or combiner 1250 that combines the panned object speaker gain 1232 (e.g., g) and the spread object speaker gain (e.g., gOS) such that there is always a contribution from the panned object speaker gain to obtain a combined speaker gain. (For example, independently of the object characteristic information or spread information 1212).

[0160] Thus, it can be said that the audio object renderer 1200 provides a combined speaker gain 1214 based on both the point-source panning of the audio object signal and the spread of the audio object signal. For example, the combined speaker gain 1214 always has a contribution from the panned object speaker gain, which ensures that the object can be appropriately localized even when the audio object spreads over a relatively wide range.

[0161] Furthermore, it should be noted that the concepts described herein can be implemented with particularly high computational efficiency, as described below.

[0162] Furthermore, it should be noted that FIG. 12 also shows how the combined speaker gain 1214 can be used in further processing. For example, speaker gains 1214a, 1214b, 1214c associated with different speakers in a speaker setup can be provided. The audio object signal 1260, which is an audio signal associated with the audio object under consideration, may be scaled by the speaker gain 1214a associated with the first speaker to obtain a first speaker signal 1262a, the audio object signal 1260 may be scaled by the speaker gain 1214b associated with the second speaker to obtain a second speaker signal 1262b, and the audio object signal 1260 may be scaled by the third speaker gain 1214c to obtain a third speaker signal 1262c, and so on. The speaker signals 1262a, 1262b, 1262c can of course be combined with speaker signals associated with other audio objects to obtain actual speaker signals.

[0163] Accordingly, the speaker signals can be obtained based on the combined speaker gain 1214, whereby the audio object under consideration is represented in both a point-source panned form and a spread form that has been found to give a particularly good auditory impression.

[0164] Furthermore, it should be noted that the audio object renderer 1200 can be optionally complemented by any of the features, functions, and details described herein, either individually or in combination.

[0165] 1.B. Audio Object Renderer According to FIG. 13 FIG. 13 shows a schematic block diagram of an audio object renderer 1300 according to an embodiment of the present invention. The audio object renderer 1300 is configured to receive object position information 1310, which may correspond to, for example, the object position information 1210. Further, the audio object renderer 1300 is configured to receive object feature information or spread information 1312, which may correspond to the object feature information or spread information 1212. Further, the audio object renderer 1300 provides a (combined) speaker gain 1314, which may correspond to the speaker gain 1214 and may be applied to the audio object signal in the same manner as the speaker gains 1214a to 1214c. The audio object renderer is configured to obtain a spread object speaker gain (also designated as a spread speaker gain or represented as a vector gOS) in consideration of the object position information 1310 and the object feature information or spread information 1312.

[0166] The audio object renderer includes a spread gain determination 1330 configured to obtain a spread gain 1332, which may be, for example, an object-to-support-point spread gain (e.g., g_spd). For example, the spread gain determination may be configured to determine the elements of a vector g_spd that describe the contribution of the audio object signal to a plurality of speaker signals or to a plurality of support point signals. In particular, the spread gain determination 1330 may use one or more polynomial functions (e.g., one or more parabolic functions or polynomial functions of degree 3) having a degree of 3 or less to map one or more angular differences between the object position and one or more support point positions to one or more spread gain value contributions (e.g., to a "layer gain"), which may be represented, for example, by aziGain, aziGainExtd, eleGain, and eleGainExtd.

[0167] In other words, the spread gain value contribution 1336 is obtained using a mapping 1334 that uses a polynomial function having a degree of 3 or less, and the mapping 1334 maps one or more angular differences between an object position and one or more support point positions to the spread gain value contribution 1336. Further, the spread gain determination 1330 also includes a spread gain contribution process 1338 that provides a spread gain 1332 (e.g., spd) based on the spread gain value contribution 1336. Further, the audio object renderer 1300 includes a spread object speaker gain determination 1340 that obtains a spread object speaker gain 1340 based on the spread gain 1332, the latter being based on the spread gain contribution 1336.

[0168] In other words, the spread gain value contribution is derived using a polynomial function having a degree of 3 or less, and the spread gain value contribution 1336 is then mapped onto the spread gain 1332, said spread gain 1332 being able to describe, for example, how an audio object signal should be distributed among a plurality of support points. The spread object speaker gain determination 1340 may be regarded as mapping a spread gain (related to spread support points) to a spread object speaker gain, which corresponds to the mapping from the support point positions to the actual speakers (and actual speaker positions).

[0169] As a further conclusion, it should be noted that the derivation of the spread object speaker gain is evaluated using a polynomial function, which can be evaluated using a moderate amount of computational effort, rather than an exponential or trigonometric function that requires a relatively high amount of computational effort.

[0170] Accordingly, the spread gain value contribution 1336 enables the spread gain 1332 to be obtained very efficiently. Also, it has been found that using a polynomial function having a degree of 3 or less results in no significant loss in the auditory representation.

[0171] However, it should be noted that the audio object renderer 1300 can be optionally complemented by any of the features, functions, and details disclosed herein, either individually or in combination.

[0172] 1.C. Speaker Gain Calculation According to FIG. 2 FIG. 2 is a signal flow chart for object spread realization for an asymmetric and / or 2D speaker set-up.

[0173] The signal flow shown in FIG. 2 can be implemented, for example, in the audio object renderers 1200, 1300 according to FIGS. 12 and 13.

[0174] Furthermore, the concepts described with reference to FIG. 2 can optionally be included, either individually or in combination, in the concepts of FIG. 5.

[0175] The speaker gain determination 200 according to FIG. 2 receives object position information 210, which can correspond to, for example, the object position information 1210, 1310. The object position information can describe the object position, for example, with respect to azimuth information (e.g., azi) and elevation information (e.g., ele). The speaker gain determination 200 also receives spread angle information 212, which can correspond to, for example, the object feature information or spread information 1212, 1312. The spread angle information 212 can describe the spread angle, for example, with respect to azimuth information and / or elevation information, or with respect to width information and / or height information. Furthermore, the speaker gain determination 200 may provide a speaker gain 214 obtained as a result corresponding to, for example, the speaker gain 1214, 1314.

[0176] Speaker gain determination 200 includes grid creation 204 for the spread support position, which provides information 204a describing the spread support point positions. Further, speaker gain determination 200 includes panning 205 of the spread support points, which may provide, for example, a spread support point panning gain 205a.

[0177] The support point position information 204a may describe the positions of the spread support points, which may be uniformly distributed, for example, on a sphere (spherical surface). For example, the spread support point positions described by the information 204a may be selected independently of the actual speaker positions and may form a grid defined by, for example, uniform azimuth and elevation intervals.

[0178] The spread support point panning gain described by the information 205a may define, for example, how the audio object signals associated with the spread support point positions are distributed to the actual speakers or speaker signals. Thus, the spread support point panning gain described by the information 205a may describe, for example, for all spread support points, the distribution of the audio object signals associated with the spread support points to the actual speakers or speaker signals.

[0179] Note that grid creation 204 and panning 205 may be calculated, for example, only once and can be reused for the spreads of multiple audio objects, as the spread support points typically do not change for the spreads of multiple audio objects.

[0180] Speaker gain determination 200 also includes spread gain calculation 201, which takes into account the spread support point positions 204a, object position information 210, and spread angle information 212. Based on this, spread gain calculation 201 provides a spread gain 201a that may describe, for example, the spread of the audio object signals to a plurality of spread support points.

[0181] Speaker gain determination 200 also includes a combination 206 that combines a spread gain 201a and a spread support position panning gain 205a to thereby obtain a spread speaker gain 206a (e.g., gOS). For example, combination 206 may use the multiplication of the spread gain 201a and the spread support point panning gain 205a. For example, combination 206 combines the mapping of the audio object signal to the spread support point described by the spread gain 201a and the mapping of the signal at (or associated with) the spread support point to the actual speaker signal described by the spread support point panning gain 205a, such that the spread speaker gain 206a provides, in the form of a weighting value to be applied to the audio object signal, the contribution of the audio object signal of the currently considered audio object in the actual speaker signal (thereby deriving the actual speaker signal). However, it should be noted that the spread gain calculation 201 and the combination 206 are typically performed separately for each audio object to be spread (while the spread support point position 204a and the spread support point panning gain 205a can be reused without change).

[0182] Speaker gain determination 200 also includes object panning 202, which uses object position information 210 to derive an object speaker gain or panned object speaker gain 202a. Object panning 202, for example, performs point-source panning of the audio object and can determine which speaker the actual speaker signal contribution of the audio object signal of the audio object under consideration is passed on to. Typically, when using point-source panning, only the object speaker gain (or panned object speaker gain) for the speaker directly adjacent to the object position is non-zero.

[0183] Speaker gain determination 200 also includes a combination 203 that combines the spread speaker gain 206a and the panned object speaker gain 202a to obtain the resulting speaker gain 214 (which can be designated, for example, as g). This combination can include, for example, the sum or weighted combination of the spread speaker gain 206a and the panned object speaker gain 202a.

[0184] This concept enables the provision of the resulting speaker gain 214 that includes both the panned object speaker gain and the spread speaker gain 206a. Such a concept has been found to be calculable with high computational efficiency and to provide good quality audio perception.

[0185] Note that the concept for speaker gain determination described with respect to FIG. 2 can optionally be complemented by any of the features, functions, and details described herein, either individually or in combination.

[0186] 1.D. Spread Function Some details regarding possible spread functions that can be used in any of the audio object renderers described herein and in any of the audio object rendering concepts described herein are provided below.

[0187] For example, FIG. 3 shows a graphical representation of different spread functions. Note that the spread function is typically defined with respect to the object position. Thus, in the graphical representations 310, 320, 330, the first axis 312a, 322a, 332a describes the azimuth difference between the currently considered spread support position and the object azimuth angle position. The second axis 312b, 322b, 332b describes the elevation difference between the currently considered spread support point position and the object elevation angle position.

[0188] Graphic representation 310 illustrates a horizontal spread (the spread in the azimuth direction is greater than the spread in the elevation direction). Graphic representation 320 illustrates a uniform spread where the spread in the azimuth direction is equal to the spread in the elevation direction, and graphic representation 330 illustrates a vertical spread where the spread in the azimuth direction is less than the spread in the elevation direction.

[0189] The graphic representation depends on the azimuth difference between the spread support point position and the azimuth object position, and illustrates the (relative) gain by which the audio object signal should be scaled (e.g., to obtain the signal at the spread support point) according to the elevation difference between the spread support point position and the object elevation position. For example, the gain can be obtained by multiplying the respective azimuth spread functions 314a, 324a, 334a and the respective elevation spread functions 314b, 324b, 334b. Note that a computationally efficient concept for implementing a spread function (or spread gain function) that closely resembles the spread gain function of FIG. 3 is described in detail. Further, note that the spread gain function of FIG. 3 can be used or approximated in, for example, the audio object renderers 1200, 1300 described herein, or the audio object rendering concepts of FIGS. 2 and 5.

[0190] Further, note that additional (optional) details regarding the spread gain function are also described herein.

[0191] Furthermore, FIG. 4 shows another graphical representation of the one-dimensional gain curve for various spread angles. It can be seen that the smaller the spread angle, the steeper the graph becomes. In FIG. 4, note that the horizontal axis 412 describes the difference angle (e.g., difference azimuth or difference elevation angle) between the spread support point and the object. The vertical axis 414 describes the gain to be applied to derive the spread support point signal from the audio object signal considering the difference angle (e.g., assuming a one-dimensional spread). Curves 416a, 416b, 416c, 416d representing the gain as a function of the difference angle are associated with different spread angles. For example, curve 416a is associated with a relatively narrow spread angle, and curve 416d is associated with a relatively wide spread angle.

[0192] As an example, for a 50-degree difference angle (azimuth or elevation) between the spread support point (SSP) and the object, a gain value of 0.37 is determined from the gain curve 416c representing a fairly strong spread value.

[0193] In other words, the gain curve shown in FIG. 4, or an approximation thereof, can be used, for example, in the derivation of the spread gain 201a, or in the derivation of the spread object speaker gain 1242, or in the derivation of the spread gain 1332. Further details regarding the gain curve and its possible approximations are also described herein. Additionally, the use of the gain curve according to FIG. 4 is described in more detail herein.

[0194] 1.E. Spread Gain Determination According to FIG. 5 FIG. 5 shows a signal flow chart for rendering the spread gain according to an embodiment of the present invention. The azimuth processing path is covered by a blue straight line or a straight line having a first hatching type and a second hatching type, and the elevation angle (more precisely, the elevation angle processing path) is covered by an orange straight line or a straight line having a third hatching type and a fourth hatching type.

[0195] Spread gain determination 500 receives object position information including azimuth object position information 510a and elevation object position information 510b as an input. The azimuth object position information 510a and the elevation object position information 510b may correspond to, for example, the object position information 210 described above. Further, the spread gain determination 500 also receives spread support point position information including azimuth spread support point position information 513a and elevation spread support point position information 513b. The azimuth SSP position information 513a and the elevation SSP position information 513b may correspond to, for example, the SSP position information 204a described above.

[0196] This spread gain determination 500 includes an azimuth gain determination 530, which may include, for example, an azimuth difference angle calculation 301 and an azimuth gain function application 302. Further, the spread gain determination 500 also includes an elevation gain determination 540, which may include, for example, an elevation difference angle calculation 304 and an elevation gain function application 305. For example, in the azimuth gain determination 530, one or more azimuth gain values (e.g., aziGain) may be determined based on the azimuth object position information 510a and the azimuth SSP position information 513a. Further, in the azimuth gain determination 530, the azimuth spread information 512a, or its preprocessed version may also be considered. For example, in the azimuth difference angle calculation 301, the difference between the azimuth object position and one or more azimuth SSP positions may be calculated, thereby obtaining one or more angle differences, and in the azimuth gain function application 302, gain values may be determined for these one or more angle differences. For example, the azimuth gain function application may evaluate a gain function (e.g., a gain function as illustrated in FIGS. 3 and 4, or a polynomial or parabolic gain function as disclosed herein) for one or more angle differences. Thus, one or more azimuth gain values are obtained, which are determined by the value of the azimuth gain function for each angle difference between each spread support point position (azimuth position) associated with the spread support point position and each object position (azimuth position), and the azimuth spread information 502 is considered to adjust the width of the azimuth gain function.

[0197] A similar operation is performed by the elevation gain determination 540. The elevation gain determination 540 receives the elevation spread support point position information 513b and the elevation object position information 510b, and based thereon, provides one or more elevation gain values 305a. For example, one or more differences between the object position elevation and one or more spread support point position elevations can be calculated by the difference angle calculation 304. To obtain one or more elevation gain values 305a, an elevation gain function whose width can be determined by the elevation spread information 512b or a pre-processed version thereof can be applied (e.g., evaluated for one or more of the difference angles determined by the difference angle calculation 304). For example, while referring to FIG. 3 or while referring to FIG. 4, the elevation gain function described, or an approximation thereof (e.g., a polynomial gain function as disclosed herein), can be used in the elevation gain function application 305. The gain function is evaluated, for example, for one or more difference angles determined in the difference angle calculation 304 to obtain one or more elevation gain values 305a.

[0198] For example, the azimuth gain value 302a can be obtained, for example, for a given elevation and for a plurality of spread support point azimuths, but can be combined, for example, multiplicatively, with a plurality of elevation gain values 305a that can be obtained, for example, for a given azimuth value and a plurality of spread support point elevation values. This combination is specified at 313 and can be a multiplication in which different pairs of the azimuth gain value 302a and the elevation gain value 305a associated with different spread support points are multiplied, thereby obtaining contributions to the gain values associated with different spread support points.

[0199] The spread gain determination 500 optionally also includes handling of an extended elevation range. For example, the spread gain determination can optionally include an elevation range extension 307, which can adapt (or pre-process) the azimuth object position information 510a and / or the azimuth SSP position information 513a and / or the elevation SSP position information 513b and / or the elevation object position information 510b for use in, for example, extended elevation range calculations.

[0200] The extended elevation angle range calculation includes, for example, providing one or more extended azimuth angle gain values 309a. The extended azimuth angle gain values 309a may correspond to, for example, the azimuth angle gain values 302a, but may have a modified association with the azimuth angle. In other words, the extended azimuth angle gain values 309a (or a set of extended azimuth angle gain values) may be, for example, an angular shift version of the azimuth angle gain values 302a (or more precisely, a set of azimuth angle gain values). However, the extended azimuth angle gain values 309a can be derived, for example, from the azimuth angle gain values 302a, or can be obtained using the azimuth angle difference calculation 308 and the azimuth angle gain function application 309.

[0201] Similarly, the elevation angle range extension process includes providing an extended elevation angle gain value 311a based on the elevation object position information 510b and the elevation SSP position information 513b. For example, the elevation angle difference calculation 310 can calculate the angle difference between the elevation object position and the elevation spread support point position within an extended elevation angle range, for example, between +90 degrees and +180 degrees, or between -90 degrees and -180 degrees. Thus, the elevation angle gain function application 311 can apply a gain function to the angle difference determined by the elevation angle difference calculation 310 to obtain, for example, an extended elevation angle gain value 311a (or more precisely, a set of extended elevation angle gain values) associated with the elevation SSP position within the extended elevation angle range. However, it should be noted that the extended elevation angle gain values 311a can optionally be determined based on the elevation angle gain values 305a, for example, using an appropriate mapping (or resorting).

[0202] Furthermore, one or more extended azimuth gain values 309a may be multiplicatively combined with, for example, one or more corresponding extended elevation gain values 311, thereby obtaining a contribution 312a and obtaining a value 314a associated with the spread support point. For example, contributions 313a and 312a associated with the same spread support point may be summed at sum 314 to obtain a gain value 314a associated with the spread support point. Additionally, normalization 315 may optionally be applied to the gain value 314 to obtain a spread gain value 514. The spread gain value 514 may correspond to, for example, the spread gain 1332 described above. For example, normalization 315 may account for the spread width and may help avoid changes in signal energy due to spreading.

[0203] Regarding the overall function of the spread gain determination 500, it should be noted that the azimuth gain value and the elevation gain value can be determined independently for multiple spread support point positions having different azimuth values and for multiple spread support point positions having different elevation values. Then, gain values for more spread support points are obtained by combination with the azimuth gain value and the elevation gain value. The standard azimuth gain value and the extended azimuth gain value can be used to reflect the spread above or below the user's head. The azimuth gain value, elevation gain value, extended azimuth gain value, and extended elevation gain value associated with the same spread support point are combined, thereby efficiently obtaining the gain value 214 associated with each spread support point. Such combinations are performed for different spread support points (or for all spread support points, or for all spread support points except those arranged at the poles).

[0204] Thus, by using the extended elevation range processing (e.g., including blocks 307, 308, 309, 310, and 311), the spread above the head, or the spread below the listener, can be easily implemented by allowing elevation angles greater than 90 degrees and / or elevation angles less than -90 degrees in the extended elevation range. By providing the pair of extended azimuth gain values and extended elevation gain values associated with the spread support points, the computational efficiency can be improved because the pair of extended azimuth gain values and extended elevation gain values can be combined (e.g., multiplied) to obtain the spread gain 314a (or the contribution 312a to the spread gain 314a).

[0205] In conclusion, the spread gain determination 500 is highly efficient and enables the determination of the spread gain even when the audio object spreads above the listener's head or below the listener.

[0206] Furthermore, it should be noted that the spread gain determination 500 can be optionally complemented using any of the features, functions, and details disclosed herein, either individually or in combination.

[0207] 1.F. Spread support point width according to FIG. 6 Figure 6 shows an exemplary spread support point lattice having a resolution of 45 degrees. As can be seen in Figure 6, the spread support points 610a, 612a, 612a, 612b, 612c, 612d, 612e, 612f, 612g, 614b, 614c, 614d, 614e, 614f, 616c, 616d, 616e are arranged on a sphere (spherical surface). In particular, the spread support points 612a to 612g, 614b to 614f, and 616c to 616e are arranged and configured on a circle having a certain elevation angle. The spread support point 610a is at the pole of the spherical coordinate system (for example, an elevation angle of +90 degrees). Further, it should be noted that the spread support points 612b, 614b are on a semi-circle having a certain azimuth angle. Similarly, the spread support points 612c, 614c, 616c are on a semi-circle having a certain azimuth angle.

[0208] Generally speaking, the spread support points are preferably arranged and configured on a lattice defined by a circle of a certain elevation angle and a semi-circle or circle having a certain azimuth angle. Therefore, typically, there are a plurality of spread support points having the same elevation angle, and typically, there are a plurality of spread support points having the same azimuth angle.

[0209] For further (optional) details, reference is also made to the additional description regarding the positions of the spread support points provided herein.

[0210] However, it should be noted that the 45-degree resolution should be considered only as an example, and different resolutions can also be selected. For example, the resolution in the azimuth angle direction and the resolution in the elevation angle direction can naturally be different.

[0211] 1.G. Polynomial gain function As described above, the use of the polynomial gain function is advantageous because the polynomial gain function (for example, having a degree of 2 or 3) can be evaluated with a small amount of computation.

[0212] Figure 7 shows a comparison of the VBAP panning curve and the shape of an extended and flipped parabola. For example, the horizontal axis 712 describes the difference angle (e.g., elevation difference angle or azimuth difference angle) between the currently considered spread support point and the object. The vertical axis 714 describes the gain or the normalized gain. The first curve 720 describes the gain that would be obtained using VBAP. The second curve 730 describes the gain that can be obtained by a (extended and flipped) parabola, and its value is limited to being non - negative. As can be seen, the parabolic - based gain function is a very good approximation of the VBAP gain function.

[0213] Therefore, the parabolic gain function can be used in any of the apparatuses and methods described herein to map the angle difference to a gain value. In other words, the parabolic - based gain function can be used, for example, in the spread gain calculation 201, or in the spread object speaker gain determination 1240, or in the mapping 1334. For example, the parabolic - based gain function can also replace (or approximate) the spread gain function shown in FIG. 3 and the gain curve shown in FIG. 4. Further, the parabolic - based gain function shown in FIG. 7 can also be used in blocks 302, 309, 311, 305 of the spread gain determination 500.

[0214] However, it should be noted that the parabola can, of course, be adapted according to the spread (e.g., according to the azimuth spread or the elevation spread). Further, the parabola can, of course, also be scaled according to the specific requirements of the application, the central value of the parabola can be changed, and / or the width of the parabola can be changed.

[0215] 1.H. Implementations according to FIGS. 8 to 11 FIGS. 8 to 11 are MATLAB (registered trademark) code examples of concepts and methods for determining speaker gain that describes the gain for including one or more audio object signals in a plurality of speaker signals.

[0216] Note that the concepts, or parts or details thereof, described in general terms with reference to FIGS. 8 to 11 can be optionally used in any of the embodiments described herein.

[0217] The method includes an initialization executed by an initialization function "spread_pannSSP". This initialization function uses, as input information, a configuration structure including VBAP parameters. The initialization function also receives information regarding the number of speakers. However, note that the initialization function does not necessarily have to use these input parameters.

[0218] However, the initialization typically includes the selection (or definition) of spread support points. For example, the spread support points can be defined by a grid of azimuth and elevation angles in a spherical coordinate system (e.g., all spread support points can have equal radii). For example, the azimuth angles of the spread support points may be defined in an array aziSSP, and the elevation angles of the spread support points may be defined in an array eleSSP. The definition of the spread support points is indicated by reference numeral 810. The definition of the spread support points indicated by reference numeral 810 may correspond to, for example, grid creation 204.

[0219] The initialization also includes the panning of the spread support points, indicated by reference numeral 820. The panning of the spread support points indicated by reference numeral 820 may correspond to, for example, the panning of the spread support points indicated by reference numeral 205 in FIG. 2. For example, for each of the spread support points, the panning of the audio signal to be rendered at the position of each spread support point to the actual speaker signal is determined. In other words, for each spread support point, a scaling value is determined that describes the panning of the signal to be rendered at the position of each spread support point to the actual speaker signal. These gain values are stored in a data structure named "Spread.gainsSSP".

[0220] For spread support points arranged at the poles of the spherical coordinate system (e.g., at an elevation angle of ±90 degrees), special treatment may be applied. This is indicated by reference numeral 830. However, it should be noted that the specific details of these panning gain values (for panning an audio object at the spread support point position to a speaker signal) have no particular relevance to the present invention. In a given example, the function vbap is used, but other functions (e.g., other panning functions) may equally well be used.

[0221] Initialization 800 also includes the initialization of several variables (or constants) for use in further processing. This initialization is indicated by reference numeral 840.

[0222] However, it should be noted that the details of initialization 800 should be considered optional.

[0223] The following describes function calls that are typically executed multiple times for different objects. The main function is called "spread_calculateGains". This main function receives the gain g provided, for example, by object panning (e.g., by object panning 202 or by panned object speaker gain determination 1230), and based on this, provides a spread gain (also designated by g), which may correspond to "the resulting speaker gain" 214 or speaker gains 1214, 1314. In addition, the main function receives a data structure including object azimuth information azi, object elevation information ele, object spread width spdAzi (or spreadAngleAzi), object spread height spdEle (or spreadAngleEle), and spread parameters that may be provided, for example, by the above-described initialization.

[0224] The main function 900 may include, for example, the determination of an attenuation gain, indicated by reference numeral 910. For example, the attenuation gain attenGain may be determined according to the object spread width and the object spread height, and further according to the spread grid resolution. For example, the calculation rule indicated by reference numeral 910 may be used. However, generally speaking, the attenuation gain may increase with the increase of the maximum object spread and also with the increase of the minimum object spread. Therefore, when the spread of the object is relatively large, the spread object speaker gain is relatively strongly weighted (in relation to the panned object speaker gain), while when the spread is relatively small, the spread object speaker gain is relatively weakly weighted.

[0225] In a further preprocessing step 920, the object spread width and / or the object spread height are adjusted to ensure that the minimum object spread width and / or the minimum object spread height are used in further processing. In particular, the smaller of the object spread width and the object spread height is adjusted to take the minimum value if it is smaller than the respective minimum value.

[0226] In a further preprocessing step, indicated by reference numeral 930, the parabola parameters used in the determination of the gain value are calculated, for example, based on the respective spread angles.

[0227] Furthermore, in a preparation step 940, loop limit values aziLoopLim and eleLoopLim are calculated, which determine the number of calculation steps executed in the calculation of the layer gain. By an (optional) limitation of the calculation steps executed in the calculation of the layer gain, computational complexity can be achieved in some cases.

[0228] Furthermore, the main function 900 also includes the determination of the azimuth layer gain, indicated by reference numeral 950. In a first sub-step 951, the array of azimuth layer gains is determined by calling the function "calculateLayerGains". The azimuth layer gains calculated in step 951 are stored in the array aziGain. In a further sub-step 952, an angular shift version of the azimuth layer gains is obtained and stored in the array aziGainExtd. In other words, the entries of the array aziGain are copied into the array aziGainExtd in a modified order.

[0229] Note that step 951 may correspond to functions 301, 302 as shown in, for example, FIG. 3. Further note that step 952 may correspond to functions as shown by reference numerals 308, 309 in FIG. 3. In other words, instead of performing functions 301, 302 as shown in FIG. 3, the function indicated by reference numeral 951 may be used, or vice versa. Further, instead of performing the functions shown by reference numerals 308, 309 in FIG. 3, the function indicated by reference numeral 952 may be used, or vice versa. For example, the arrays aziGain and aziGainExtd may represent layer gains associated with different azimuths that are shifted by 180 degrees relative to each other. For example, the first element of the array aziGain may correspond to the azimuth φ1, and the first element of the array aziGainExtd may correspond to the azimuth φ1 + 180 degrees.

[0230] The main function 900 may also include the determination of the elevation layer gain, indicated by reference number 960. For this purpose, the function "calculateLayerGains" may be (re)used, which returns an intermediate array of values, as indicated by reference number 961. From this intermediate array of values eleGainTMP, an array of elevation layer gains eleGain may be determined by an appropriate selection of the entries of the intermediate array of values, as indicated by reference number 962. Similarly, an extended array of elevation layer gains may also be determined using an appropriate selection and order of the entries of the intermediate array, as indicated by reference number 963.

[0231] The functions as indicated by reference numbers 961 and 962 may correspond to the functions of blocks 304 and 305, for example, and the functions as indicated by reference numbers 961 and 963 may correspond to the functions as indicated by blocks 310 and 311, for example.

[0232] In other words, the functions as indicated by reference numbers 961 and 962 may replace the functions of functional blocks 304, 305, and the functions as indicated by reference numbers 961, 963 may replace the functions as indicated by blocks 310, 311. However, alternatively, the functions of blocks 304, 305 and the functions of blocks 310, 311 may be executed instead of function 960.

[0233] In a further step, indicated by reference number 970, the main function 900 calculates the spread support point spread gain based on the azimuth layer gain and based on the previously calculated elevation layer gain. For this purpose, the function "calculateSSPGains", described later, is called.

[0234] Accordingly, a spread support point spread gain, specified by the array g_spd, is obtained, which describes which scaling to use for rendering the audio object at the spread support point. However, since it is desired to know which scaling to use for rendering the audio object signal into the speaker signal, at step 980, the spread support point spread gain is mapped to a speaker gain, which can be understood as the panning from the audio signal to be rendered at the position of the spread support point to the actual speaker signal (typically associated with the speaker at an actual speaker position different from the spread support point).

[0235] For this purpose, the result of the previously performed panning of the spread support point (executed at initialization 800) is utilized. The product of the set of the spread support point panning gain and the spread support point spread gain is, for example, summed over all spread support points. In other words, the spread support point panning gain (or the set of spread support point panning gains) is associated with each spread support point (referenced by the running variable obj), and the spread support point spread gain is also associated with each spread support point.

[0236] Note that step 980 may, for example, correspond to the functionality indicated by reference numeral 206. Accordingly, it is also possible that the functionality of block 206 is replaced by the functionality indicated by reference numeral 980, and vice versa.

[0237] At step 990, the obtained (spread object) speaker gain gOS is combined with an input gain value g, which may be, for example, a panned object gain value. The scaling of the spread gain value gOS is determined, for example, by the attenuation gain attenGain described above. Further, step 990 optionally includes normalizing the result of the combination of the panned gain value and the spread gain value.

[0238] For example, step 990 may correspond to the function of block 203.

[0239] Regarding the overall function of the main function 900, note that the main function is at steps 950, 960, 970, and 980. At step 950, an array of "layered gain values" is calculated, which describes the spread of the audio object in the azimuth direction, more precisely, the gain values associated with the azimuth values of the spread support points. In this step, the object spread in the azimuth direction is considered. Also, the extended azimuth gain values that are cyclically shifted with respect to the azimuth gain values in the aziGain array help to form an "overhead" spread.

[0240] At step 960, spread values associated with different elevation values associated with a given azimuth value and spread support points are calculated. Here, the elevation position of the audio object and the object spread in the elevation direction are considered. Also, an extended array of elevation layer gains is obtained to support the overhead spread of the audio object.

[0241] At step 970, the values of the azimuth layer gain and the elevation layer gain are combined to calculate the gain values associated with all support point positions.

[0242] As a result, at step 980, the gain values associated with the support point positions are effectively mapped to the gain values associated with the speaker signals.

[0243] Below, some details of the functions "calculateLayerGains (calculate layer gains)" and "calculateSSPGains (calculate SSP gains)" will be described with reference to FIGS. 10 and 11.

[0244] Figure 10 shows the MATLAB (registered trademark) code for the function "calculateLayerGains (calculation of layer gains)". Note that the return value of the said function specified by "gains" is an array, and the index of the array elements is associated with the azimuth angle or elevation angle. Generally speaking, the entries of the said array include that they decay almost parabolically as the angle difference between the angular position (azimuth position or elevation position) of the audio object and the angle (for example, SSP azimuth position or SSP elevation position) associated with each entry of the array increases.

[0245] Function 1000 includes the optional determination of the signed value plumin, indicated by reference numeral 1010.

[0246] Furthermore, function 1000 includes the determination of the array index (a "plurality") associated with the object position (for example, object elevation or object azimuth). This determination is indicated by reference numeral 1020.

[0247] Function 1000 also includes the calculation of the deviation of the object position from the position (angle) of the adjacent spread support points, indicated by reference numeral 1030.

[0248] Furthermore, the function includes the calculation of the gain values, indicated by reference numeral 1040. These gain values are stored in the array "gains", and the array index is associated with the angle (azimuth or elevation) of the spread support points. The gain values themselves are determined using the evaluation of the parabola for each angle difference between the position of the audio object under consideration and the position of each spread support point. The gain values are provided by the evaluation of the limited parabola so that the values remain non - negative. Therefore, the array "gains" is filled with gain values (with the restriction to non - negative values applied) based on the evaluation of the parabola centered on the angle (azimuth or elevation) of the position of the audio object under consideration.

[0249] Thus, the function "calculate layer gain", denoted by reference numeral 1000, enables providing an array of gain values, more precisely spread gain values associated with a constant azimuth angle or alternatively a constant elevation angle.

[0250] Next, the details of the function "calculateSSPGains (calculate SSP gain)" will be described.

[0251] The function denoted by reference numeral 1100 includes the calculation of the spread gain denoted by reference numeral 1110. In particular, one spread gain value is calculated for each spread support point SSP, and the specific handling for the spread support point at the pole of the polar coordinate system is denoted by reference numeral 1120.

[0252] However, for each spread support point specified by the elevation index nel and the azimuth index naz, the azimuth gain value aziGain(naz) is multiplicatively combined with the elevation gain value eleGain(nel). This combination may correspond to, for example, the operation shown in block 313. In addition, the associated extended azimuth gain value aziGainExtd(naz) is also multiplicatively combined with the associated extended elevation gain value eleGainExtd(nel), which may correspond to the operation shown in block 312.

[0253] Furthermore, the results of the multiplicative combinations are then added, which may correspond to the operations shown in block 314. For example, the azimuth gain value and the extended azimuth gain value specified by the same index naz may correspond to angles that differ by only 180 degrees. For example, the azimuth gain value specified by a given index naz may be associated with an azimuth of +45 degrees, while the extended azimuth gain value associated with the same index naz may be associated with an azimuth of -135 degrees. Furthermore, the angles associated with the same index nel in the elevation gain array and the extended elevation gain array may be summed up to 180 degrees or summed up to -180 degrees. For example, a given index nel may specify an entry in the array eleGains that is associated with +45 degrees, and the same index nel will specify an entry in the array eleGainExtd that is associated with an angle of 135 degrees. Thus, the angles associated with a given index nel of the arrays eleGain and eleGainExtd may, in this example, be summed to +180 degrees. When such combinations are used, it can be ensured that appropriate scaling values are obtained with reasonable effort. Also, the fact that the elevation value is greater than 90 degrees does not overly increase the computational complexity of the concept.

[0254] The specific handling 1120 of the poles (i.e., the gain values associated with the poles) helps to avoid artifacts at the poles.

[0255] The function 1000 also includes normalization, indicated by reference numeral 1130, which may be considered optional. Thus, the spread gain is optionally normalized to keep the values within an appropriate range.

[0256] In conclusion, the function 1100 enables the derivation of the gain values associated with the spread support points based on the array of gain values associated with a single elevation and also based on the array of gain values associated with a single azimuth.

[0257] Note that the functions of functions 800, 900, 1000, and 1100 can be optionally introduced into any of the other embodiments, either individually or in combination. Also note that any of the functions described in the other embodiments can be optionally introduced into functions 800, 900, 1000, 1100, either individually or in combination.

[0258] 1.I. Method according to FIG. 14 FIG. 14 shows a flowchart of a method 1400 for determining a speaker gain that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information and object feature information or spread information.

[0259] The method includes obtaining a panned object speaker gain 1410 using point-source panning of the audio object.

[0260] The method also includes obtaining a spread object speaker gain 1420 considering the object position information and the object feature information or spread information.

[0261] This method also includes combining the panned object speaker gain and the spread object speaker gain 1430 such that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain.

[0262] Method 1400 is based on the same considerations as the above-described apparatus and can be optionally complemented, either individually or in combination, by any of the features, functions, and details described herein.

[0263] 1.J. Method according to FIG. 15 FIG. 15 shows a flowchart of a method 1500 for determining a speaker gain that describes a gain for including one or more audio object signals in a plurality of speaker signals based on object position information and object feature information or spread information.

[0264] The method includes obtaining a spread object speaker gain 1510 in consideration of object position information and object feature information.

[0265] The method also includes obtaining a spread gain 1520 using one or more polynomial functions having a degree of 3 or less that map an angular difference between an object position and a support point position to a spread gain value contribution.

[0266] Method 1530 also includes obtaining a spread object speaker gain using a spread gain, or using the spread gain as a spread object speaker gain, based on the spread gain contribution.

[0267] Method 1500 is based on the same considerations as the apparatus described above and may be optionally supplemented, individually or in combination, by any of the features, functions, and details described herein.

[0268] 2. Description of Further Embodiments Hereinafter, an object spread rendering algorithm according to an embodiment of the present invention will be described.

[0269] Note that this object spread rendering algorithm can be used independently, but can be supplemented, optionally, individually or in combination, by any of the features, functions, and details disclosed herein. Also, any of the features, functions, and details of the concepts described in this section can be optionally introduced, individually or in combination, into any of the apparatuses and methods described herein.

[0270] According to one aspect, the basic idea for realizing the spread effect is to activate additional speakers that play back the same object signal with a monotonically decreasing intensity starting from the object position. That is, for each speaker, a spread gain for generating the spread effect must be calculated for the spread of the object to be played back. The spread gain can be determined in the following way.

[0271] 2.1. Object spread calculation using additional spread support points In the case of an asymmetric and / or 2D speaker setup, in some cases, the speaker positions cannot be used as SSPs. The reason is the potential reduction in localization accuracy because the spread playback depends on SSPs that are evenly distributed (e.g., on a sphere). Therefore, a grid of equally spaced objects is created (e.g., in block 204 of the concept in FIG. 2), and they serve the role of SSPs. Each of these SSPs is panned (e.g., by VBAP) in, for example, block 205 (panning of spread support points). Block 201 (spread gain calculation) calculates the spread gain using the SSP positions (for details, see, for example, section 2.2), and thereby in block 206 (combination) they can be combined with the SSP panning gain (the simplest way to combine both types of gain is multiplication. Of course, other procedures are possible). To reproduce a small (e.g., an angle smaller than the SSP grid resolution) spread angle, the actual object is panned (e.g., by VBAP) in, for example, block 202, "object panning", and combined with the spread speaker gain (e.g., in block 203, "combination") (details see 2.5).

[0272] 2.2. Spread gain calculation (e.g., in block 201) In the calculation of the spread gain, for example, a monotonically decreasing function is used. For example, based on the spherical distance between the SSP and the object, the attenuation gain is determined from that function. For example, the function has a gain value of 1 at the position of the object and is 0 where no spreading effect is desired. For example, since amplification is not allowed, the attenuation gain can be restricted to a range between 0 and 1.

[0273] Without further processing steps, this procedure only allows the creation of a uniform spread pattern, such as those indicated by reference numerals 100, 101, 102 in FIG. 1 for example. To achieve a non-uniform spread pattern (such as those indicated by reference numerals 104 and 105 in FIG. 1), the spread angle (for example, [azimuth and elevation] or [width and height]) should be processed individually, for example. Therefore, in the following, the weight function is designed not only in one dimension (for example, one spread value) but also in two dimensions (for example, width and height).

[0274] For example, as depicted in FIG. 3, the two-dimensional gain function can be modeled as a combination of two one-dimensional gain functions. For example, to model a non-uniform horizontal spread pattern (for example, the upper plot in FIG. 3), a large spread angle is selected, which creates a wide one-dimensional gain function (illustrated as a projection onto the left wall). In parallel, for example, a narrow elevation spread angle is selected that creates a narrow one-dimensional gain function (illustrated as a projection onto the right wall). For illustrative purposes, the one-dimensional functions are normalized to have a maximum value of 1. For example, by combining both one-dimensional functions, a two-dimensional function is obtained. For example, the simplest way to combine two one-dimensional functions is multiplication. Of course, different other procedures are possible.

[0275] 2.2.1. Algorithm In the following, an example of spread gain calculation is described based on one SSP and one object. However, the algorithm can be well executed for all subsequent SSPs, and the resulting spread gains can be accumulated.

[0276] For example, in block 301, the absolute difference between the azimuth angle of the SSP and the object is calculated. In block 303, the spread value is restricted, for example, so as not to take a value smaller than the SSP grid resolution angle in the case of non-uniform spread. Otherwise, the "panning" of the non-uniform spread with respect to the moving object may cause a perceptible jump in some cases. For example, based on the difference angle from block 301 and the spread angle from block 303, one spread gain component is calculated, and the spread angle controls, for example, the shape / width of the one-dimensional gain curve, and the difference angle selects its value. An example is shown in FIG. 4.

[0277] The same procedure is repeated, for example, with the elevation angle values in blocks 304, 306, and 305. Both results are multiplied, for example, in block 313. This product already represents one value on the surface of the two-dimensional gain function, but since the elevation angle is naturally limited to [-90°, 90°], it is only within the elevation angle range. Note that the definition of the sphere in spherical coordinates typically (or conventionally) allows only the following two combinations. This is either an azimuth angle with a range of [-180°, 180°] and an elevation angle of [-90°, 90°], or the azimuth angle is limited to [-90°, 90°] and the elevation angle range can be extended to [-180°, 180°]. Otherwise, the rear part of the sphere is defined twice. However, this limitation can be overcome in some embodiments of the present invention.

[0278] For example, assume an object at an azimuth angle of 30° and an elevation angle of 80° in a front semi-sphere (hemisphere) with a vertical spread of 60°. Since the spread is, for example, symmetric with respect to the object, for example, a spread of 20° is placed in the rear semi-sphere (hemisphere) and the azimuth angle is, for example, 210° (or -150°). In that case, the horizontal one-dimensional gain function is, for example, selected to have a narrow shape and will mask the spread gain to be approximately zero so that the vertical spread can be achieved.

[0279] For these reasons, the elevation angle range is, in some embodiments, extended to [-180°, 180°] (more optional details are described, for example, in section 2.3) and mapped to the original range (covered, for example, by signal processing blocks 314 and 315). For example, similar to the procedures in blocks 301, 302, 304, 305, and 313, in blocks 308, 309, 310, 311, and 312, for example, the spread gain component is calculated for the extended elevation angle range. Finally, the results from blocks 312 and 313 are added in block 314 and normalized in block 315 (more optional details are described, for example, in section 2.4).

[0280] 2.3. Elevation Angle Range Extension (e.g., in block 307) For example, in some embodiments, all SSPs must be mirrored to an extended elevation angle range (i.e., from [0°, 90°] to [91°, 180°], and from [-90°, -1°] to [-180°, -91°]) in order to render a spread on the rear hemisphere (hemisphere) while the object is on the front hemisphere (hemisphere) (and vice versa). The procedure is described based on the following example.

[0281] As an example, assuming an SSP at (20°, 40°) (azimuth and elevation angles), its mirrored SSP is placed at (-160°, 140°). This simply means that, for example, the azimuth angle is shifted by 180° while maintaining the angular distance with respect to the horizontal plane, and the elevation angle is mirrored to the extended elevation range.

[0282] Another example for the lower hemisphere: An SSP at (120°, -70°) is mirrored to (-60°, -110°).

[0283] 2.4. Normalization (e.g., in block 315) In some cases, when the elevation range is extended, the gain in the extended range is calculated, and an additional gain component is added to the gain determined from the original range, it may be amplified. The maximum gain within the original range is always 1, for example, at the position of the object. The gain component from the extended range added to the maximum gain from the original range is at the mirrored position and depends on, for example, the spread value (e.g., in combination with azimuth and elevation angles).

[0284] For example, in the case of a monotonically decreasing function, this addition of both gain components results in the maximum possible gain for a given spread value. Therefore, in some cases, it is necessary (or advantageous) to normalize the spread gain of the object to this maximum value. The normalized gain is, for example

[0285]

Equation

[0286] is, SGCA is, for example, the spread gain component determined from block 309 with respect to the azimuth difference between an object and its mirrored position (the azimuth difference between the object and its mirrored position is always 180°), and SGCE is, for example, the spread gain component determined from block 311 with respect to the elevation difference between an object and its mirrored position.

[0287] 2.5. Combination of Spread Gain and Object Gain (e.g., in block 203) For example, one combination method is the sum of the object speaker gain and the spread speaker gain. The resulting vector including the speaker gain may need to be normalized to its Euclidean norm, for example. Depending on the SSP grid resolution, it may also be necessary to attenuate the spread speaker gain before combination. For example, when the SSP grid resolution is low (e.g., a low SSP grid resolution can reduce the computational complexity), rapidly changing spread values from 0° upwards may cause perceptible artifacts. For example, the attenuation can be determined from the formula

[0288]

Equation

[0289] (note that other formulas are also possible), res or g res is the SSP grid resolution, spread azi is the azimuth spread angle, and spread ele is the elevation spread angle.

[0290] 3. Efficient Implementations In the following, optional details of efficient implementations that can be optionally used, either individually or in combination, in combination with any of the embodiments disclosed herein are described.

[0291] The SSPs are selected, for example, to be equidistantly distributed on a sphere (spherical surface), so that, for example, each of the horizontal / vertical layers will have the same angular distance.

[0292] Figure 6 shows an example of an SSP grid with a resolution of 45°. As a result, 8 vertical layers and 5 horizontal layers are obtained, and the SSPs at elevation angles of + / -90° should be defined only once.

[0293] However, it is sufficient to calculate the difference azimuth angle (e.g., in blocks 301 and 308) between the object and the SSP only once on the horizontal layer, and the difference elevation angle (e.g., in blocks 304 and 310) only once on the vertical layer.

[0294] Furthermore, for example, it is possible to calculate the difference angle from the object to each SSP by calculating one difference angle (e.g., on the horizontal layer) in the clockwise direction, calculating one in the counterclockwise direction, and counting the SSP index.

[0295] Example: On the horizontal layer, the SPP (or SSP) azimuth angles are [0°, 45°, 90°, 135°, 180°, -135°, -90°, -45°]. Their already prepared (e.g., already prepared at initialization) indices can be, for example, [1, 2, 3, 4, 5, 6, 7, 8] stored in an index ring (using the index ring enables jumping, for example, from index 1 to 8 and from 8 to 1. For example, it is an infinite loop or at least approximates an infinite loop).

[0296] For example, let's assume an object azimuth angle of 30°. With respect to the SSP index, this is positioned between indices 1 and 2. The difference is 15° in the clockwise direction and 30° in the counterclockwise direction. For a given spread angle of 180° (i.e., ±90° symmetrically with respect to the object's position), the SSPs at azimuth angles [45°, 90°] (clockwise) and [0°, -45°] (counterclockwise) are activated. Thus, when calculating the difference angle, only two indices in the clockwise direction and two indices in the counterclockwise direction need to be considered.

[0297] This method has, for example, two advantages. First, by using an index ring, wrapping of angles outside the interval [-180°, 180°] is avoided. Additionally, this makes it possible to limit the calculations to the relevant SSPs. For small spread angles, as a result, the computational effort is significantly reduced.

[0298] Another possibility to make the algorithm computationally efficient is to appropriately select the design of the weighting function. On the one hand, a gain curve such as the previously introduced "Gaussian bell curve" uses an exponential function implemented as a power series expansion and is thus computationally inefficient. On the other hand, the gain function should ideally have the shape of a function obtained as a result determined from a panning algorithm (e.g., VBAP), as shown in FIG. 7. This is particularly desirable (or even necessary in some cases) when using a non-uniform spread pattern in combination with a small spread angle to ensure smooth object movement.

[0299] It has been found that an appropriate compromise is given when an appropriately extended and flipped parabola is used. This avoids exponential / trigonometric functions and well approximates the shape of the VBAP panning curve.

[0300] 4. Alternative Implementations Although some aspects are described in the context of an apparatus, these aspects also serve as descriptions of corresponding methods, and it is clear that a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also serve as descriptions of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps can be performed (or by using) by a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.

[0301] Depending on some implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation form can be executed using a digital storage medium, such as a floppy disk, a DVD, a Blu-Ray (registered trademark), a CD, a ROM, a PROM, an EPROM, an EEPROM, or a flash memory, in which an electronically readable control signal that cooperates (or can cooperate) with a programmable computer system in which each method is executed is stored. Therefore, the digital storage medium may be regarded as computer-readable.

[0302] Some embodiments according to the present invention include a data carrier containing an electronically readable control signal that can cooperate with a programmable computer system in which one of the methods described herein is executed.

[0303] Generally, embodiments of the present invention can be implemented as a computer program product with program code, and the program code is operable to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.

[0304] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable medium.

[0305] Thus, in other words, one embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is running on a computer.

[0306] Thus, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which is recorded a computer program for performing one of the methods described herein. A data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.

[0307] Thus, a further embodiment of the method of the invention is a sequence of data streams or signals representing a computer program for performing one of the methods described herein. The sequence of data streams or signals can be configured to be transferred, for example, via a data communication connection, such as the Internet.

[0308] A further embodiment includes processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.

[0309] A further embodiment includes a computer on which is installed a computer program for performing one of the methods described herein.

[0310] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0311] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0312] The apparatus described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0313] The apparatus described herein, or any component of the apparatus described herein, may be implemented at least partially in hardware and / or software.

[0314] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0315] The methods described herein, or any component of the apparatus described herein, may be performed at least partially by hardware and / or software.

[0316] The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and changes to the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, it is intended to be limited only by the scope of the following claims and not by the specific details presented in the description and explanation of the embodiments herein.

[0317] As an additional note, it should be noted that the phrase "considering" can have the meaning of, for example, but not necessarily, "based on" or "in response to (depending on)".

[0318] As an additional note, it should be noted that the phrase "describing" can have the meaning of, for example, but not necessarily, "representing" or "directly or indirectly representing" or "being a measure of" or "comprising". For example, a first quantity that "describes" another quantity may be equal to the other quantity, or proportional to the other quantity, or related to the other quantity using a predetermined (linear or non-linear) relationship.

[0319] As an additional note, it should be noted that the phrase "associated with an azimuth value" may, for example, have the meaning of "having an azimuth value".

[0320] As an additional note, it should be noted that the phrase "associated with an elevation value" may, for example, have the meaning of "having an elevation value".

Description of Reference Numerals

[0321] 200, 1300 Audio Object Renderer 201 Spread Gain Calculation 201a Spread Gain 202 Object Panning 202a, 1232, g Panned Object Speaker Gain 203 Combination 204 Grid Creation 204a information 205 panning 205a spread support point panning gain 206 combination 206a, 1242, gOS object feature information speaker gain 210, 1310, azi, ele object position information 212, 1312, spreadAngleAzi, spreadAngleEle spread information 214, 1214, 1214a~1214c speaker gain 301 azimuth difference angle calculation 302 azimuth gain function application 302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd object feature gain contribution 304 elevation difference angle calculation 305 elevation gain function application 305a elevation gain value 307 elevation range extension 308 azimuth difference angle calculation 309 azimuth gain function application 309a extended azimuth gain value 310, 320, 330 graphic representation 310 elevation difference angle calculation 311 elevation gain function application 311a extended elevation gain value 312a, 322a, 332a first axis 312b, 322b, 332b second axis 314 sum 314a, g_spd object feature information gain 314a, 324a, 334a azimuth spread function 314b, 324b, 334b elevation spread function 315 normalization 412 horizontal axis 416a, 416b, 416c, 416d curves 500 spread gain determination 502 Azimuth spread information 510a Azimuth object position information 510b Elevation object position information 512a Azimuth spread information 512b Elevation spread information 513a Azimuth spread support point position information 513b Elevation spread support point position information 514 Spread gain value 530 Azimuth gain determination 540 Elevation gain determination 610a, 612a, 612a, 612b, 612c, 612d, 612e, 612f, 612g, 614b, 614c, 614d, 614e, 614f, 616c, 616d, 616e Spread support points 712 Horizontal axis 714 Vertical axis 720 First curve 730 Second curve 800 Initialization 900 Main function 960 Function 1000 Function 1100 Function 1120 Specific handling 1200 Audio object renderer 1210 Object position information 1212 Object feature information 1214 Speaker gain 1214a, 1214b, 1214c Speaker gain 1230 Panned object speaker gain determination 1232 Panned object speaker gain 1240 Spread object speaker gain determination 1242 Spread object speaker gain 1250 Combiner or mixer 1260 Audio object signal 1262a~1262c Speaker signals 1300 Audio Object Renderer 1310 Object Location Information 1312 Object Feature Information 1314 Speaker Gain 1330 Spread Gain Determination 1332 Spread Gain 1334 Mapping 1336 Spread Gain Value Contribution 1338 Spread Gain Contribution Process 1340 Spread of Object Speaker Gain Determination 1400 Method 1500 Method

Claims

1. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a - 1214c) to include one or more audio object signals (1260) in a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain panned object speaker gains (202a, 1232, g) using point - source panning (202, 1230) of the audio object, the audio object being regarded as a point source in the point - source panning, and the object feature information being ignored in the point - source panning, and the object position information is used in the point - source panning, the audio object renderer is configured to expand the audio object over an extended region considering the object feature information (1212) to obtain object feature information speaker gains (206a, 1242, gOS), the audio object renderer is configured to combine the panned object speaker gains (202a, 1232, g) and the object feature information speaker gains (206a, 1242, gOS) such that the contribution of the panned object speaker gains is always present to obtain combined speaker gains (214, 1214, 1214a - 1214c), and the determination of the object feature information speaker gains takes into account the expansion of the audio object, The audio object renderer (200, 1200) is configured to evaluate one or more gain functions that map the difference between the positions (204a, aziSSP, eleSSP) of the support points (610a, 612a to 612g, 614b to 614f, 616c to 616e) and the object position (210, 1210, azi, ele) to one or more spread gain value contributions (302a, aziGain(naz), 305a, eleGain(nel)), and to determine a spread object speaker gain (206a, 1242, gOS) based on the one or more spread gain value contributions. **Claim 2** The audio object renderer (200, 1200) according to claim 1, wherein the object feature information is audio object spread information (212, 1212). **Claim 3** The audio object renderer is configured to determine a weight (attenGain, g) of a spread object speaker gain (206a, 1242, gOS) in a weighted combination with a panned object speaker gain (202a, 1232, g) according to a spread in a first direction (spreadAngleAzi, spread azi ) and a spread in a second direction (spreadAngleEle, spread ele ), the audio object renderer (200, 1200) according to claim 1 or 2. atten The audio object renderer (200, 1200) according to claim 1 or 2, wherein the audio object renderer is configured to determine a weight (attenGain, g) of a spread object speaker gain (206a, 1242, gOS) in a weighted combination with a panned object speaker gain (202a, 1232, g) according to a spread in a first direction (spreadAngleAzi, spread **Claim 4** The audio object renderer is configured to determine a spread angle (spreadAngleAzi, spread azi ) in a first direction and a spread angle (spreadAngleEle, spread eleThe weighted combination of the panned object speaker gain (202a, 1232, g) and the spread object speaker gain (206a, 1242, gOS) in accordance with the product with ( ) and the weight (attenGain, g) of the spread object speaker gain (206a, 1242, gOS) in the weighted combination with the panned object speaker gain (202a, 1232, g) according to the product with ( ) atten The audio object renderer (200, 1200) according to any one of claims 1 to 3, configured to determine ( ). **Claim 5** The audio object renderer, according to the spread angle (spreadAngleAzi, spread) in the first direction and the spread angle (spreadAngleEle, spread) in the second direction, adds the panned object speaker gain (202a, 1232, g) weighted by a fixed weight and the spread object speaker gain (206a, 1242, gOS) weighted by a variable weight (attenGain, g). azi The audio object renderer (200, 1200) according to any one of claims 1 to 4, configured to add ( ) and ( ) ele The audio object renderer, according to the spread angle (spreadAngleAzi, spread) in the first direction and the spread angle (spreadAngleEle, spread) in the second direction, adds the panned object speaker gain (202a, 1232, g) weighted by a fixed weight and the spread object speaker gain (206a, 1242, gOS) weighted by a variable weight (attenGain, g). atten The audio object renderer (200, 1200) according to any one of claims 1 to 4, configured to add ( ) and ( ) **Claim 6** The audio object renderer is configured to normalize the result of adding the panned object speaker gain (202a, 1232, g) weighted by the fixed weight and the spread object speaker gain (206a, 1242, gOS) weighted by the variable weight (attenGain, g). atten The audio object renderer (200, 1200) according to claim 5, configured to normalize the result of adding ( ) and ( ) **Claim 7** The audio object renderer attenGain = 0.89f * min(c 1 , max(spread azi , spread ele ) / g res1 ) + 0.11f * min(c 2 , min(spread azi , spread ele ) / g res2 ) configured to determine the weight (attenGain) of the spread object speaker gain (206a, 1242, gOS) in a weighted combination with the panned object speaker gain (202a, 1232, g), c 1 is a predetermined value, c 2 is a predetermined value, g res1 is a predetermined value, g res2 is a predetermined value, spread azi is the spread angle of the audio object in the azimuth direction, spread ele is the spread angle of the audio object in the elevation direction, min(.) is the minimum operator, max(.) is the maximum operator, the audio object renderer (200, 1200) according to any one of claims 1 to 6.

8. The audio object renderer is configured to increase the relative contribution of the spread object speaker gain (206a, 1242, gOS) compared to the panned object speaker gain (202a, 1232, g) as the spread angle (spreadAngleAz i, spread azi , spreadAngleEle, spread ele ) of the audio object increases, the audio object renderer (200, 1200) according to any one of claims 1 to 7.

9. The audio object renderer is configured to obtain a spread object speaker gain (206a, 1242, gOS) using a representation (204a, aziSSP, eleSSP) of a support point position in polar coordinates, taking into account the object position information (210, 1210, azi, ele) and spread information (212, 1212). The audio object renderer according to any one of claims 1 to 8, wherein the audio object renderer is configured to provide the speaker gain (214, 1214, 1214a to 1214c) based on the spread object speaker gain (206a, 1242, gOS). [

10. ] The audio object renderer evaluating one or more angular differences (diffCLKDir, diffAntiCLKDir) between the azimuth position (210, 1210, azi) of the audio object and one or more support points (204a, aziSSP), and / or evaluating one or more angular differences (diffCLKDir, diffAntiCLKDir) between the elevation position (210, 1210, azi) of the audio object and the elevation position (204a, eleSSP) of one or more support points whereby the audio object renderer (200, 1200) according to any one of claims 1 to 9 is configured to obtain a spread speaker gain (206a, 1242, gOS). [

11. ] The audio object renderer (200, 1200) according to any one of claims 1 to 10, wherein the support point positions (204a, aziSSP, eleSSP) are arranged and configured on a spherical surface within an allowable range of ±10% or ±20% of the radius of the spherical surface. [

12. ] The support point positions (204a, aziSSP, eleSSP) have a uniform azimuth interval along a circle having a constant elevation angle and a constant radius, and / or The support point positions (204a, aziSSP, eleSSP) are of the audio object renderer (200, 1200) according to any one of claims 1 to 11, having a uniform elevation angle along a circle having a certain azimuth angle and a certain radius. **Claim 13** The audio object renderer is configured to obtain a spread object speaker gain (206a, 1242, gOS) such that the audio object is spread over a region extending within a first hemispherical surface where the audio object is disposed and also extending within a second hemispherical surface having an azimuth position on the opposite side of the first hemispherical surface, of the audio object renderer (200, 1200) according to any one of claims 1 to 12. **Claim 14** The audio object renderer is configured to use an extended elevation angle range between -180 degrees and +180 degrees, of the audio object renderer (200, 1200) according to claim 13. **Claim 15** The audio object renderer is for a given object position (210, 1210, azi, ele) and a given spread (212, 1212, spreadAngleAzi, spread azi , spreadAngleEle, spread ele ), a first set of azimuth gain values (302a, aziGain) that describe the contribution to the spread gain (314a, g_spd) for a plurality of azimuth values related to a support point position or a support point azimuth index (naz) with respect to an elevation value within an original elevation value range that does not show the intersection of the poles of the spherical coordinate system, and a second set of azimuth gain values (309a, aziGainExtd) that describe the contribution to the spread gain (314a, g_spd) for a plurality of azimuth values related to a support point position or a support point azimuth index (naz) with respect to an elevation value within an extended elevation value range that shows the intersection of one of the poles of the spherical coordinate system, calculate, The audio object renderer (200, 1200) according to claim 13 or 14, configured to derive the spread gain (206a, 1242, gOS) using the azimuth gain value (302a, aziGain(naz)) of the first set and the azimuth gain value (309a, aziGainExtd) of the second set.

16. For a given object position (210, 1210, azi, ele) and a given spread (212, 1212, spreadAngleAzi, spreadAngleEle), the audio object renderer A first set of elevation gain values (305a, eleGain) that describe the contribution to the spread gain (314a, g_spd) for a plurality of elevation values related to a support point position or a speaker azimuth index or a support point elevation index (nel) within the original elevation value range that does not show the intersection of the poles of the spherical coordinate system, and A second set of elevation gain values (311a, eleGainExtd) that describe the contribution to the spread gain (314a, g_spd) for a plurality of elevation values related to a support point position or a speaker elevation index or a support point elevation index (e.g., nel) within the extended elevation value range that shows the intersection of one of the poles of the spherical coordinate system, calculates, The audio object renderer (200, 1200) according to claim 15, configured to derive the spread gain (206a, 1242, gOS) using the azimuth gain value (302a, aziGain(naz)) of the first set, the azimuth gain value (309a, aziGainExtd) of the second set, the elevation gain value (305a, eleGain(nel)) of the first set, and the elevation gain value (311a, eleGainExtd(nel)) of the second set.

17. The audio object renderer according to any one of claims 13 to 16, wherein the audio object renderer is configured to combine values of a first set of azimuth gain values and a first set of elevation gain values (302a, aziGain(naz), 305a, eleGain(nel)) and to combine values of a second set of azimuth gain values and a second set of elevation gain values (309a, aziGainExtd(naz), 311a, eleGainExtd(nel)).

18. The audio object renderer according to any one of claims 13 to 17, wherein the second set of azimuth gain values (309a, aziGainExtd) represents a gradual change in gain values over an azimuth angle that is shifted by 180 degrees when compared to a gradual change in gain values over an azimuth angle represented by the first set of azimuth gain values (302a, aziGaind).

19. The first set of azimuth gain values (302a, aziGain) represents a gradual change in gain values over a 360-degree range taking into account the azimuth object position (210, 1210, azi) and an azimuth spread angle (spreadAngleAzi, spread) having an angular accuracy determined by the number of speakers or the number of support points, and / or azi The second set of azimuth gain values (309a, aziGainExtd) represents a gradual change in gain values over a 360-degree range taking into account the azimuth object position (210, 1210, azi) rotated by 180 degrees and an azimuth spread angle (spreadAngleAzi, spread) having an angular accuracy determined by the number of speakers or the number of support points. The audio object renderer according to any one of claims 13 to 18. The audio object renderer according to any one of claims 13 to 18, wherein the second set of azimuth gain values (309a, aziGainExtd) represents a gradual change in gain values over a 360-degree range taking into account the azimuth object position (210, 1210, azi) rotated by 180 degrees and an azimuth spread angle (spreadAngleAzi, spread) having an angular accuracy determined by the number of speakers or the number of support points. azi The audio object renderer according to any one of claims 13 to 18, wherein the second set of azimuth gain values (309a, aziGainExtd) represents a gradual change in gain values over a 360-degree range taking into account the azimuth object position (210, 1210, azi) rotated by 180 degrees and an azimuth spread angle (spreadAngleAzi, spread) having an angular accuracy determined by the number of speakers or the number of support points.

20. The elevation gain values (305a, eleGain) of the first set represent a gradual change in gain values over an elevation range from -90 degrees to +90 degrees considering the elevation object position (210, 1210, ele) and the elevation spread angle (spreadAngleEle, spread ele ), and / or The elevation gain values (311a, eleGainExtd) of the second set represent a gradual change in gain values over elevation ranges from -180 degrees to -90 degrees and from +90 degrees to +180 degrees considering the elevation object position (210, 1210, ele) and the elevation spread angle (spreadAngleEle, spread ele ), the audio object renderer (200, 1200) according to any one of claims 13 to 19.

21. The audio object renderer is configured to determine a speaker gain (214, 1214, 1214a to 1214c) that describes a gain for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1310, azi, ele) and spread information (212, 1312, spreadAngleAzi, spreadAngleEle), The audio object renderer is configured to obtain a spread object speaker gain (206a, 1242, gOS) considering the object position information and the spread information, The audio object renderer is configured to obtain a spread gain (314a, g_spd) using one or more polynomial functions of degree three or less that map the angle difference between the object position (210, 1310, azi, ele) and the support point position (204a, aziSSP, eleSSP) to a spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd), The audio object renderer according to any one of claims 1 to 20, wherein the audio object renderer is configured to obtain the spread object speaker gain (206a, 1242, gOS) using a spread gain (314a, g_spd) based on the spread gain value contribution.

22. The audio object renderer (200, 1300) according to claim 21, wherein the width of the one or more polynomial functions is determined by spread information (212, 1312, spreadAngleAzi, spreadAngleEle).

23. The audio object renderer (200, 1300) according to claim 21 or 22, wherein the audio object renderer uses a first polynomial function that maps the azimuth difference between the object position (210, 1310, azi) and the support point position (204a, aziSSP) to a first spread gain value contribution (302a, aziGain(naz)), and a second polynomial function that maps the elevation difference between the object position (210, 1310, ele) and the support point position (204a, eleSSP) to a second spread gain value contribution (305a, eleGain(nel)) to obtain a spread gain value (314a, g_spd).

24. The audio object renderer (200, 1300) according to claim 23, wherein the audio object renderer is configured to combine the first spread gain value contribution (302a, aziGain(naz)) and the second spread gain value contribution (305a, eleGain(nel)) to obtain a spread gain value (314a, g_spd).

25. For a given object position (210, 1310, azi, ele) and a given spread (212, 1312, spreadAngleAzi, spreadAngleEle), the audio object renderer A set of azimuth gain values (302a, aziGain) that describe the contribution to the spread gain (314a, g_spd) for a plurality of azimuth values related to the support point position or the speaker azimuth angle index or the support point azimuth angle index (naz), and / or A set of elevation gain values (305a, eleGain) that describe the contribution to the spread gain (314a, g_spd) for a plurality of elevation values related to the support point position or the speaker elevation angle index or the support point elevation angle index (nel), are calculated, The audio object renderer (200, 1300) according to any one of claims 21 to 24, configured to derive the spread gain (314a, g_spd) using the azimuth gain values (302a, aziGain(naz), 305a, eleGain(nel)).

26. The audio object renderer is configured to combine an element (aziGain(naz)) of the set of azimuth gain values (302a, aziGain) related to the currently considered speaker or the currently considered support point with an element (eleGain(nel)) of the set of elevation gain values (305a, eleGain) related to the currently considered speaker or the currently considered support point to obtain spread gain values (spd(objNo)) related to a plurality of different speakers or a plurality of different support points. The audio object renderer (200, 1300) according to claim 25.

27. For a given object position (210, 1310, azi, ele) and a given spread (212, 1312, spreadAngleAzi, spreadAngleEle), the audio object renderer The first set of azimuth gain values (302a, aziGain) that describe the contribution to the spread gain (314a, g_spd) for a plurality of azimuth values related to a support point position or a speaker azimuth index or a support point azimuth index (naz) at an elevation value within the original elevation value range that does not show the intersection of the poles of the spherical coordinate system, and The second set of azimuth gain values (309a, aziGainExtd) that describe the contribution to the spread gain (314a, g_spd) for a plurality of azimuth values related to a support point position or a speaker azimuth index or a support point azimuth index (naz) at an elevation value within the extended elevation value range that shows the intersection of the poles of the spherical coordinate system, are calculated, The audio object renderer (200, 1300) according to any one of claims 21 to 26, configured to derive the spread gain (314a, g_spd) using the azimuth gain value (302a, aziGain(naz)) and / or the elevation gain value (305a, eleGain(nel)). **Claim 28** For a given object position (210, 1310, azi, ele) and a given spread (212, 1312, spreadAngleAzi, spreadAngleEle), the audio object renderer The first set of elevation gain values (305a, eleGain) that describe the contribution to the spread gain (314a, g_spd) for a plurality of elevation values related to a support point position or a speaker elevation index or a support point elevation index (nel) at an elevation value within the original elevation value range that does not show the intersection of the poles of the spherical coordinate system, and The second set of elevation gain values (311a, eleGainExtd) that describe the contribution to the spread gain (314a, g_spd) for a plurality of elevation values related to a support point position or a speaker elevation index or a support point elevation index (nel) at an elevation value within the extended elevation value range that shows the intersection of the poles of the spherical coordinate system, are calculated, The audio object renderer (200, 1300) according to claim 27, which is configured to derive the spread gain (314a, g_spd) using the azimuth gain value (302a, aziGain(naz), 309a, aziGainExtd(naz)) and the elevation gain value (305a, eleGain(nel), 311a, eleGainExtd(nel)). **Claim 29** The audio object renderer is configured to pre-calculate a support point panning gain (Spread.gainsSSP) for panning an audio signal associated with a plurality of support points to a plurality of speakers at initialization using panning, The audio object renderer is configured to obtain an object-to-support point spread gain (g_spd) that describes the contribution of an audio object signal to a plurality of support point signals using a polynomial function having a degree of 3 or less, The audio object renderer according to any one of claims 21 to 28, which is configured to obtain the spread object speaker gain by combining the object-to-support point spread gain and the support point panning gain. **Claim 30** The one or more polynomial functions having a degree of 3 or less are p = max(0, c1 * anglediff 2 + c2) is a parabola function that gives the return value p according to, c1 is a parameter that determines the width of the parabola function, c2 is a predetermined value, anglediff is the angle difference at which the parabola function is evaluated, max(.,.) is a maximum value operator that returns the maximum value of the operands, and the audio object renderer according to any one of claims 21 to 29. **Claim 31** The audio object renderer according to any one of claims 1 to 30, configured to provide a speaker gain combined based on both the point-source panning of the audio object and the spread of the audio object signal.

32. In determining the object feature information speaker gain, the audio object is spread over a larger number of speakers than in determining the panned object speaker gain, for the audio object renderer according to any one of claims 1 to 31.

33. A method for determining speaker gains (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), comprising: obtaining panned object speaker gains (202a, 1232, g) using the point-source panning (202, 1230) of the audio object, wherein the audio object is regarded as a point source in the point-source panning and the object feature information is ignored in the point-source panning, using the object position information in the point-source panning, spreading the audio object over an extended region taking into account the object feature information (1212) to obtain object feature information speaker gains (206a, 1242, gOS), and combining the panned object speaker gains (202a, 1232, g) and the object feature information speaker gains (206a, 1242, gOS) such that the contribution of the panned object speaker gain is always present to obtain combined speaker gains (214, 1214, 1214a to 1214c). In determining the object feature information speaker gain, the extension of the audio object is taken into consideration, and further includes a step of determining a weight (attenGain, g atten) of a spread object speaker gain (206a, 1242, gOS) in a weighted combination with a panned object speaker gain (202a, 1232, g) according to a spread in a first direction (spreadAngleAzi, spread azi) and a spread in a second direction (spreadAngleEle, spread ele).

34. The method (1400) according to claim 33, comprising: evaluating one or more gain functions that map a difference between a position of a support point and an object position to one or more spread gain value contributions (aziGain(naz), eleGain(nel)); and determining a spread object speaker gain (gOS) based on the one or more spread gain value contributions.

35. obtaining a spread object speaker gain (206a, 1242, gOS) in consideration of the object position information (210, 1210, azi, ele) and the object feature information (1212) (1510); obtaining a spread gain (314a, g_spd) using one or more polynomial functions having a degree of three or less that map an angular difference between an object position (210, 1310, azi, ele) and a support point position (204a, aziSSP, eleSSP) to spread gain value contributions (302a, aziGain, 305a, eleGain, 309a, aziGainExt d, 311a, eleGainExt d) (1520); obtaining the spread object speaker gain (206a, 1242, gOS) by using a spread gain (314a, g_spd) based on the spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExt, 311a, eleGainExt), or using the spread gain (314a, g_spd) as the spread object speaker gain (206a, 1242, gOS); step (1530) as recited in claim 33 or 34.

36. The method determines a speaker gain (214, 1214, 1214a - 1214c) that describes a gain for including one or more audio object signals (1260) in a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1310, azi, ele) and spread information (212, 1312, spreadAngleAzi, spreadAngleEle). obtaining the spread object speaker gain (206a, 1242, gOS) in consideration of the object position information and the spread information; step (1510) obtaining a spread gain (314a, g_spd) by using one or more polynomial functions having a degree of three or less that map an angular difference between an object position (210, 1310, azi, ele) and a support point position (204a, aziSSP, eleSSP) to a spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExt, 311a, eleGainExt, aziGain(naz), eleGain(nel)); step (1520) Step (1530) of obtaining the spread object speaker gain (206a, 1242, gOS) by using a spread gain (314a, g_spd) based on the spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExt d, 311a, eleGainExt d, aziGain(naz), eleGain(nel)), or by using the spread gain (314a, g_spd) as the spread object speaker gain, is included in the method according to any one of claims 33 to 35.

37. A computer program for executing the method according to any one of claims 33 to 36 when executed on a computer.

38. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) by using point source panning (202, 1230) of the audio object, the audio object is regarded as a point source in the point source panning, and the object feature information is ignored in the point source panning, The object position information is used in the point source panning, The audio object renderer is configured to spread the audio object over an extended area in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS), The audio object renderer is configured to obtain a combined speaker gain (214, 1214, 1214a to 1214c) by combining the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) such that the contribution of the panned object speaker gain is always present. In determining the object feature information speaker gain, the extension of the audio object is taken into account. The audio object renderer is configured to determine a weight (attenGain, g) of a spread object speaker gain (206a, 1242, gOS) in a weighted combination with the panned object speaker gain (202a, 1232, g) according to a spread in a first direction (spreadAngleAzi, spread) and a spread in a second direction (spreadAngleEle, spread). azi ele atten ), the audio object renderer (200, 1200). [

39. ] An audio object renderer (200, 1200) for determining a speaker gain (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) using point source panning (202, 1230) of the audio object, the audio object being regarded as a point source in the point source panning, and the object feature information being ignored in the point source panning. The object position information is used in the point source panning. ​​The audio object renderer is configured to spread the audio object over an extended area in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS). The audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that the contribution of the panned object speaker gain is always present, to obtain a combined speaker gain (214, 1214, 1214a to 1214c). In determining the object feature information speaker gain, the expansion of the audio object is taken into consideration. The audio object renderer has a spread angle (spreadAngleAzi, spread azi ) in a first direction and a spread angle (spreadAngleEle, spread ele ) in a second direction, and is configured to determine a weight (attenGain, g atten ) of the spread object speaker gain (206a, 1242, gOS) in a weighted combination with the panned object speaker gain (202a, 1232, g) according to the product of the two. The audio object renderer (200, 1200).

40. An audio object renderer (200, 1200) for determining a speaker gain (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) using point source panning (202, 1230) of the audio object, the audio object is regarded as a point source in the point source panning, and the object feature information is ignored in the point source panning. In the point source panning, the object position information is used. The audio object renderer is configured to spread the audio object over an extended area in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS). The audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that the contribution of the panned object speaker gain is always present, to obtain a combined speaker gain (214, 1214, 1214a to 1214c). In determining the object feature information speaker gain, the expansion of the audio object is considered. The audio object renderer has a spread angle (spreadAngleAzi, spread azi ) in a first direction and a spread angle (spreadAngleEle, spread ele ) in a second direction, and is configured to add a panned object speaker gain (202a, 1232, g) weighted with a fixed weight and a spread object speaker gain (206a, 1242, gOS) weighted with a variable weight (attenGain, g atten ). An audio object renderer (200, 1200).

41. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a - 1214c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain panned object speaker gains (202a, 1232, g) using point - source panning (202, 1230) of the audio object, where the audio object is regarded as a point source in the point - source panning and the object feature information is ignored in the point - source panning, The object position information is used in the point - source panning, The audio object renderer is configured to spread the audio object over an extended area considering the object feature information (1212) to obtain object feature information speaker gains (206a, 1242, gOS), The audio object renderer is configured to combine the panned object speaker gains (202a, 1232, g) and the object feature information speaker gains (206a, 1242, gOS) such that the contribution of the panned object speaker gains is always present to obtain combined speaker gains (214, 1214, 1214a - 1214c), The determination of the object feature information speaker gains takes into account the expansion of the audio object, The audio object renderer is, attenGain = 0.89f * min(c 1 , max(spread azi , spread ele ) / g res1 ) + 0.11f * min(c 2 , min(spread azi , spread ele) / g res2 ) and configured to determine a weight (attenGain) of a spread object speaker gain (206a, 1242, gOS) in a weighted combination with the panned object speaker gain (202a, 1232, g) according to c 1 is a predetermined value, c 2 is a predetermined value, g res1 is a predetermined value, g res2 is a predetermined value, spread azi is the spread angle of the audio object in the azimuth direction, spread ele is the spread angle of the audio object in the elevation direction, min(.) is the minimum operator, max(.) is the maximum operator, AudioObjectRenderer(200, 1200).

42. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a-1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a-1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), comprising: the audio object renderer is configured to obtain panned object speaker gains (202a, 1232, g) using point source panning (202, 1230) of an audio object, the audio object being considered as a point source in the point source panning, and the object feature information being ignored in the point source panning; The point sound source panning uses the object position information, The audio object renderer is configured to expand the audio object over an extended area in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS). The audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that there is always a contribution from the panned object speaker gain, to obtain a combined speaker gain (214, 1214, 1214a to 1214c). In determining the object feature information speaker gain, the expansion of the audio object is considered. The audio object renderer is configured to increase the relative contribution of the spread object speaker gain (206a, 1242, gOS) compared to the panned object speaker gain (202a, 1232, g) as the spread angle (spreadAngleAzi, spread azi , spreadAngleEle, spread ele ) of the audio object increases. The audio object renderer (200, 1200).

43. An audio object renderer (200, 1200) for determining a speaker gain (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) using point-source panning (202, 1230) of the audio object, the audio object is regarded as a point source in the point-source panning, and the object feature information is ignored in the point-source panning. In the point-source panning, the object position information is used. The audio object renderer is configured to spread the audio object over an extended area in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS). The audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain (214, 1214, 1214a to 1214c). In determining the object feature information speaker gain, the expansion of the audio object is considered. The audio object renderer is configured to obtain a spread object speaker gain (206a, 1242, gOS) using an expression (204a, aziSSP, eleSSP) of a support point position in polar coordinates in consideration of the object position information (210, 1210, azi, ele) and spread information (212, 1212). The audio object renderer (200, 1200) is configured to provide the speaker gain (214, 1214, 1214a to 1214c) based on the spread object speaker gain (206a, 1242, gOS).

44. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a - 1214c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), wherein the audio object renderer is configured to obtain panned object speaker gains (202a, 1232, g) using point - source panning (202, 1230) of the audio object, the audio object being regarded as a point source in the point - source panning and the object feature information being ignored in the point - source panning, wherein the object position information is used in the point - source panning, wherein the audio object renderer is configured to spread the audio object over an extended region considering the object feature information (1212) to obtain object feature information speaker gains (206a, 1242, gOS), wherein the audio object renderer is configured to combine the panned object speaker gains (202a, 1232, g) and the object feature information speaker gains (206a, 1242, gOS) such that the contribution of the panned object speaker gains is always present to obtain combined speaker gains (214, 1214, 1214a - 1214c), wherein the determination of the object feature information speaker gains takes into account the expansion of the audio object, the audio object renderer is, evaluating one or more angular differences (diffCLKDir, diffAntiCLKDir) between the azimuthal position (210, 1210, azi) of the audio object and one or more support points (204a, aziSSP), and / or Evaluating one or more angular differences (diffCLKDir, diffAntiCLKDir) between the elevation position (210, 1210, azi) of the audio object and the elevation position (204a, eleSSP) of one or more support points An audio object renderer (200, 1200) configured to obtain a spread speaker gain (206a, 1242, gOS) thereby.

45. An audio object renderer (200, 1200) for determining a speaker gain (214, 1214, 1214a - 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object characteristic information (1212), The audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) using point - source panning (202, 1230) of the audio object, the audio object is regarded as a point source in the point - source panning, and the object characteristic information is ignored in the point - source panning, The object position information is used in the point - source panning, The audio object renderer is configured to spread the audio object over an extended region considering the object characteristic information (1212) to obtain an object characteristic information speaker gain (206a, 1242, gOS), The audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object characteristic information speaker gain (206a, 1242, gOS) such that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain (214, 1214, 1214a - 1214c), In the determination of the object feature information speaker gain, the expansion of the audio object is taken into consideration. The support point positions (204a, aziSSP, eleSSP) are arranged and configured on the spherical surface within an allowable range of ±10% or ±20% of the radius of the spherical surface, for an audio object renderer (200, 1200).

46. An audio object renderer (200, 1200) for determining a speaker gain (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) using point source panning (202, 1230) of the audio object, the audio object being regarded as a point source in the point source panning, and the object feature information being ignored in the point source panning, In the point source panning, the object position information is used. The audio object renderer is configured to expand the audio object over an extended region considering the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS). The audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that the contribution of the panned object speaker gain is always present, to obtain a combined speaker gain (214, 1214, 1214a to 1214c). In the determination of the object feature information speaker gain, the expansion of the audio object is taken into consideration. The support point positions (204a, aziSSP, eleSSP) have a uniform azimuth interval along a circle having a constant elevation angle and a constant radius, and / or The audio object renderer (200, 1200) in which the support point positions (204a, aziSSP, eleSSP) have a uniform elevation angle along a circle having a constant azimuth angle and a constant radius.

47. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), The audio object renderer is configured to obtain panned object speaker gains (202a, 1232, g) using point source panning (202, 1230) of the audio object, the audio object is regarded as a point source in the point source panning, and the object feature information is ignored in the point source panning, The object position information is used in the point source panning, The audio object renderer is configured to expand the audio object over an extended area in consideration of the object feature information (1212) to obtain object feature information speaker gains (206a, 1242, gOS), The audio object renderer is configured to combine the panned object speaker gains (202a, 1232, g) and the object feature information speaker gains (206a, 1242, gOS) so that the contribution of the panned object speaker gains is always present to obtain combined speaker gains (214, 1214, 1214a to 1214c), The determination of the object feature information speaker gains takes into account the expansion of the audio object, The audio object renderer (200, 1200) is configured to obtain a spread object speaker gain (206a, 1242, gOS) such that the audio object is spread over a region that extends within a first hemispherical surface in which the audio object is disposed and also extends within a second hemispherical surface having an azimuthal position on the side opposite to the first hemispherical surface.

48. An audio object renderer (200, 1200) for determining a speaker gain (214, 1214, 1214a - 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object characteristic information (1212), the audio object renderer is configured to obtain a panned object speaker gain (202a, 1232, g) using point - source panning (202, 1230) of the audio object, the audio object is regarded as a point source in the point - source panning, and the object characteristic information is ignored in the point - source panning, the object position information is used in the point - source panning, the audio object renderer is configured to spread the audio object over an extended region in consideration of the object characteristic information (1212) to obtain an object characteristic information speaker gain (206a, 1242, gOS), the audio object renderer is configured to combine the panned object speaker gain (202a, 1232, g) and the object characteristic information speaker gain (206a, 1242, gOS) such that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain (214, 1214, 1214a - 1214c). In determining the object feature information speaker gain, the extension of the audio object is taken into consideration. The audio object renderer is configured to determine speaker gains (214, 1214, 1214a to 1214c) that describe gains for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1310, azi, ele) and spread information (212, 1312, spreadAngleAzi, spreadAngleEle). The audio object renderer is configured to obtain a spread object speaker gain (206a, 1242, gOS) in consideration of the object position information and the spread information. The audio object renderer is configured to obtain a spread gain (314a, g_spd) using one or more polynomial functions of degree three or less that map the angular difference between the object position (210, 1310, azi, ele) and the support point position (204a, aziSSP, eleSSP) to spread gain value contributions (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd). The audio object renderer is configured to obtain the spread object speaker gain (206a, 1242, gOS) using the spread gain (314a, g_spd) based on the spread gain value contribution.

49. A method for determining speaker gains (214, 1214, 1214a to 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a to 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212). A step of obtaining a panned object speaker gain (202a, 1232, g) using point source panning (202, 1230) of an audio object, wherein the audio object is regarded as a point source in the point source panning, and the object feature information is ignored in the point source panning, In the point source panning, the object position information is used, A step of expanding the audio object over an extended region in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS); Combining the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain (214, 1214, 1214a - 1214c), In determining the object feature information speaker gain, the expansion of the audio object is considered, Evaluating one or more gain functions that map the difference between the position of the support point and the object position to one or more spread gain value contributions (aziGain(naz), eleGain(nel)), and determining a spread object speaker gain (gOS) based on the one or more spread gain value contributions.

50. A method for determining a speaker gain (214, 1214, 1214a - 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), A step of obtaining a panned object speaker gain (202a, 1232, g) using point source panning (202, 1230) of an audio object, where the audio object is regarded as a point source in the point source panning, and the object feature information is ignored in the point source panning, In the point source panning, the object position information is used, A step of expanding the audio object over an extended region in consideration of the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS), A step of combining the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) so that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain (214, 1214, 1214a - 1214c), In the determination of the object feature information speaker gain, the expansion of the audio object is considered, A step (1510) of obtaining a spread object speaker gain (206a, 1242, gOS) in consideration of the object position information (210, 1210, azi, ele) and the object feature information (1212), A step (1520) of obtaining a spread gain (314a, g_spd) using one or more polynomial functions having a degree of three or less to map the angular difference between the object position (210, 1310, azi, ele) and the support point position (204a, aziSSP, eleSSP) to a spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExt, 311a, eleGainExt), A method including a step (1530) of obtaining the spread object speaker gain (206a, 1242, gOS) by using a spread gain (314a, g_spd) based on the spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd), or by using the spread gain (314a, g_spd) as the spread object speaker gain (206a, 1242, gOS).

51. A method for determining a speaker gain (214, 1214, 1214a - 1214c) for including one or more audio object signals (1260) in a plurality of speaker signals (1262a - 1262c) based on object position information (210, 1210, azi, ele) and object feature information (1212), including a step of obtaining a panned object speaker gain (202a, 1232, g) using the point - source panning (202, 1230) of the audio object, where the audio object is regarded as a point source in the point - source panning and the object feature information is ignored in the point - source panning, where the object position information is used in the point - source panning, a step of spreading the audio object over an extended region considering the object feature information (1212) to obtain an object feature information speaker gain (206a, 1242, gOS), a step of combining the panned object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) such that the contribution of the panned object speaker gain is always present to obtain a combined speaker gain (214, 1214, 1214a - 1214c), where the determination of the object feature information speaker gain takes into account the extension of the audio object, The method determines speaker gains (214, 1214, 1214a-1214c) that describe gains for including one or more audio object signals (1260) into a plurality of speaker signals (1262a-1262c) based on object position information (210, 1310, azi, ele) and spread information (212, 1312, spreadAngleAzi, spreadAngleEle). Step (1510) of obtaining a spread object speaker gain (206a, 1242, gOS) in consideration of the object position information and the spread information. Step (1520) of obtaining a spread gain (314a, g_spd) using one or more polynomial functions having a degree of three or less that map an angular difference between an object position (210, 1310, azi, ele) and a support point position (204a, aziSSP, eleSSP) to a spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExt d, 311a, eleGainExt d, aziGain(naz), eleGain(nel)). Step (1530) of obtaining the spread object speaker gain (206a, 1242, gOS) using the spread gain (314a, g_spd) based on the spread gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExt d, 311a, eleGainExt d, aziGain(naz), eleGain(nel)), or using the spread gain (314a, g_spd) as the spread object speaker gain. A method including this.

52. A computer program for executing the method according to any one of claims 49 to 51 when executed on a computer.

Citation Information

Patent Citations

  • Method for generating and consuming 3D sound scenes with spatially extended sound sources

    JP2006503491A

  • Rendering Audio Objects with Apparent Size to Arbitrary Loudspeaker Layouts

    JP2016511990A