Audio object renderer, method for determining loudspeaker gains

By combining an audio object renderer with translation and expansion of speaker gain, the problem of poor audio object localization and expansion in surround sound reproduction is solved, achieving accurate localization and appropriate expansion of audio objects with low computational complexity.

CN114902698BActive Publication Date: 2026-01-13FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080088819.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-20
Filing Date
2020-11-20
Publication Date
2026-01-13
Estimated Expiration
2040-11-20

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve good localization and expansion of audio objects in surround sound reproduction while maintaining low computational complexity, especially in multi-speaker environments where the user is not located in the "sweet spot" of the listening setup, resulting in poor perception of audio objects.

Method used

By combining the speaker gain of the translation object and the speaker gain of the extended object using an audio object renderer, and utilizing object position and feature information, the speaker gain is determined to achieve good localization and extension of the audio object. This audio object renderer obtains the speaker gain of the translation object by translating the audio object's point source, and obtains the speaker gain of the extended object by combining it with object feature information or extension information, thus combining them into a combined speaker gain.

Benefits of technology

It achieves good positioning and extension of audio objects with low computational complexity, ensuring accurate positioning and appropriate perceptual extension of audio objects in multi-speaker environments, especially suitable for users who are not in the "sweet spot" of the listening setup.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902698B_ABST
    Figure CN114902698B_ABST
Patent Text Reader

Abstract

An audio object renderer (200; 1200) for determining a loudspeaker gain (214, 1214, 1214a-c) based on object position information (210, 1210, azi, ele) and object characteristic information or extension information (1212), the loudspeaker gain describing a gain for including one or more audio object signals (1260) into a plurality of loudspeaker signals (1262a-c), the audio object renderer being configured to obtain a panned object loudspeaker gain (202a, 1232, g) using a point source panning (202, 1230) of the audio object. The audio object renderer is configured to obtain an extended object loudspeaker gain (206a, 1242, gOS) taking into account the object position information (210, 1210, azi, ele) and the object characteristic information or extension information (1212). The audio object renderer is configured to combine the panned object loudspeaker gain (202a, 1232, g) and the extended object loudspeaker gain (206a, 1242, gOS) in a way that the contribution of the panned object loudspeaker gain is always present in order to obtain a combined loudspeaker gain (214, 1214, 1214a-c). Methods and computer programs are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] An embodiment of the present invention relates to an audio object renderer.

[0002] Other embodiments of the invention relate to a method for determining loudspeaker gain.

[0003] Other embodiments of the present invention relate to computer programs.

[0004] Embodiments of the present invention generally involve the translation of an audio object having an extended source size. Background Technology

[0005] In the following sections, some background information on the invention will be described. However, it should be noted that the features, functions, and applications mentioned below may also be used optionally in conjunction with embodiments of the invention.

[0006] In the field of surround sound reproduction, loudspeakers are typically placed in specific locations within a room. A common surround sound system, "5.1," contains three loudspeakers in the front hemisphere and two in the rear hemisphere. If the intention is to reproduce a signal (e.g., a mono audio signal) in the space between two loudspeakers, the signal is proportionally distributed across these two adjacent loudspeakers. This process also applies to 3D loudspeaker setups, which additionally have loudspeakers above and / or below the horizontal plane. A well-known translation algorithm is the so-called "vector-based amplitude translation" (VBAP). After calculating the translation gain, the mono signal is reproduced from the relevant loudspeakers using corresponding weights.

[0007] It has been found that most translation techniques reproduce point-like sound signals (objects) in space. Furthermore, it has been found that it is often desirable to change the size of the object, to make its sound more diffuse, to alter the perceived distance, or to achieve other psychoacoustic effects. Therefore, the object should (or sometimes must) sound not just as a point, but rather as coming from a wider angle of reproduction.

[0008] Figure 1 Graphical representations of different object expansion configurations are shown. In the top row, at reference numerals 100, 101, and 102, objects with three different expansion values ​​are shown. In the bottom row, at reference numerals 104 and 105, objects expand non-uniformly on the reproducing sphere.

[0009] In other words, Figure 1 Different object expansion configurations, independent of the reproduction speaker setup, are depicted. At reference numeral 100, a point-like sound-emitting object is depicted. At reference numerals 101 and 102, the object expands uniformly at a wider / higher reproduction angle. At reference numeral 104, the object expands vertically, while at reference numeral 105, the object expands horizontally.

[0010] Given this situation, there is a need to create a concept that provides an improvement in auditory impression and computational complexity. Summary of the Invention

[0011] An audio object renderer is created according to embodiments of the present invention for determining speaker gain based on object location information and object feature information, wherein speaker gain describes the gain of including one or more audio object signals into multiple speaker signals. The audio object renderer is configured to obtain a translated object speaker gain using point source translation of the audio object. The audio object renderer is configured to take into account object feature information (1212) to obtain an object feature information speaker gain (e.g., speaker gain considering the extension and / or perceived extension and / or perceived angular extension and / or diffusion and / or blurring of one or more audio objects under consideration). For example, the object feature information may describe diffusion, such as the distribution of a source (or audio object or sound originating from an audio object) to multiple points, which may correspond, for example, to the broadening of the source. For example, the object or the perceived size of an object may be amplified based on the object feature information. Generally, for example, the object feature information may represent the extension and / or range and / or diffusion of an audio object, and the object feature information speaker gain may take into account such extension and / or range and / or diffusion of the audio object. Alternatively or additionally, the object feature information may, for example, describe the distance of the audio object, and this distance may be converted into an extension, for example, in a preparatory step, wherein this extension can then be taken into account when providing the speaker gain of the object feature information. However, as another option, the speaker gain of the object feature information may also be derived directly from the distance.

[0012] The audio object renderer is also configured to combine the translation object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) in a manner that always has the contribution of the translation object speaker gain, in order to obtain the combined speaker gain (214, 1214, 1214a~c).

[0013] This embodiment of the invention is based on the finding that a good trade-off between computational complexity and achievable auditory impression can be achieved by obtaining object feature information speaker gain (which can correspond to extended object speaker gain) based on both translation object speaker gain and extended object speaker gain, which describes the intensity of the object signal from different speaker signals associated with different speakers. Specifically, by using translation object speaker gain (which typically provides a “point source” auditory impression), the user’s localization of the audio object can be facilitated. For example, the derivation of translation object speaker gain can use point source translation of the audio object, which can, for example, select a single speaker for playback of the audio object, or can, for example, distribute the audio object to multiple speakers closest to the audio object (e.g., while not using those speakers that are not the closest to the audio object). Thus, “point source” translation of the audio object typically provides translation object speaker gain, where only the speaker gains of the few speakers closest to the object location are non-zero.

[0014] Furthermore, the audio object renderer also obtains object feature information object speaker gain, where the object extends over an extended region, such as an extension range in the azimuth angle and / or an extension range in the elevation angle. Therefore, the determination of object feature information object speaker gain takes into account the extension of the audio object, which can be derived, for example, from the object feature information. Compared to the determination of translation object speaker gain, the determination of object feature information object speaker gain typically results in the audio object extending across a larger number of speakers because the determination of object feature information object speaker gain takes into account the extension of the audio object and typically uses a relatively broad (e.g., even broader than the extension of the audio object under consideration) stable (e.g., steadily attenuating) distribution characteristic.

[0015] Therefore, by combining the translation object speaker gain based on the point source translation of the audio object with the object speaker gain that takes into account the object's extension, good localization of the audio object can be achieved (even for objects with relatively large extensions), while still allowing the extension of the audio object to be perceived. This is especially true when there are multiple users and / or users who are not located in the "sweet spot" of the listening arrangement. By always introducing the contribution of the translation object speaker gain, for example, independently of the object feature information, it is possible to always ensure the localization of the audio object, even for listeners who are not in the sweet spot.

[0016] Furthermore, it should be noted that object feature information typically allows for the identification or estimation of the extension of an audio object. For example, object feature information can indicate the type of object, which can imply parameters (e.g., extension parameters) used to determine the speaker gain of the extended object. For example, object feature information can allow for the differentiation between relatively small and relatively large objects. Alternatively or additionally, object feature information can allow for the differentiation between near and far objects, which can also imply one or more parameters used to determine the speaker gain of the extended object. Optionally, object feature information can describe the “blob” or “spread” extension of the audio object, or the distribution of the audio object to multiple local locations.

[0017] In summary, this audio object renderer can derive one or more parameters from object feature information to determine the object speaker gain. Therefore, the object feature information allows for appropriate adjustment of the derived extended object speaker gain, so that the object can be positioned by providing translational object speaker gain, and can also be perceived with an appropriate extension because the object feature information is taken into account when providing the speaker gain.

[0018] In a preferred embodiment, the audio object renderer is configured to also consider object position information to obtain object feature information, such as speaker gain. This allows both the extension and position of the audio object to be taken into account.

[0019] In a preferred embodiment, the object feature information is audio object extended information. This allows for particularly efficient computation because, in this case, it is not necessary to map the "abstract" object feature information to the object extended information.

[0020] An audio object renderer is created according to embodiments of the present invention for determining speaker gain based on object location information (e.g., azimuth and elevation angles) (which may be provided, for example, in spherical coordinates, such as using azimuth values ​​and elevation angle values ​​ele) and object feature information, wherein the speaker gain describes the gain of including one or more audio object signals into multiple speaker signals (e.g., combined speaker gain or resulting speaker gain). The object feature information may be, for example, information indicating whether the object is small or extended (e.g., object size values), or the object feature information may be, for example, object distance information that can be mapped to an extension value (e.g., mapped to extension angle information describing extension in the azimuth direction (e.g., spreadAngleAzi) and / or mapped to extension angle information describing extension in the elevation direction (e.g., spreadAngleEle)). However, other types of object feature information are also possible.

[0021] The audio object renderer is configured to use point-source translation of the audio object to obtain the translation object speaker gain (e.g., also referred to as "object speaker gain," or denoted by vector g). In point-source translation, the audio object can be treated, for example, as a point source, where extended information is ignored, and where the signal of the audio object is associated with two or more speakers in the environment of the audio object's object location by an appropriate selection of the translation object speaker gain.

[0022] The audio object renderer is configured to take into account object location information and object feature information to obtain extended object speaker gain (e.g., also referred to as extended speaker gain, or represented as a vector gOS).

[0023] The audio object renderer is configured to combine the translation object speaker gain (e.g., g) and the extended object speaker gain (e.g., gOS) in a manner that always has a contribution from the translation object speaker gain (e.g., independent of object feature information) in order to obtain a combined speaker gain.

[0024] This embodiment of the invention is based on the finding that a good trade-off between computational complexity and achievable auditory impression can be achieved by obtaining the extended object speaker gain based on both the translation object speaker gain and the extended object speaker gain, which describes the intensity of the object signal from different speaker signals associated with different speakers. Specifically, by using the translation object speaker gain (which typically provides a “point source” auditory impression), the user’s localization of the audio object can be facilitated. For example, the translation object speaker gain can be derived using a point source translation of the audio object, which can, for example, select a single speaker for playback of the audio object, or can, for example, distribute the audio object to multiple speakers closest to the audio object (e.g., without using those speakers that are not the closest to the audio object). Thus, the “point source” translation of the audio object typically provides a translation object speaker gain, where only the speaker gains of the few speakers closest to the object location are non-zero.

[0025] Furthermore, the audio object renderer also obtains extended object speaker gain, where the object extends over an extended region, such as an extension range in the azimuth angle and / or an extension range in the elevation angle. Therefore, the determination of extended object speaker gain takes into account the extension of the audio object, which can be derived, for example, from object feature information. Compared to the determination of translation object speaker gain, the determination of extended object speaker gain typically involves extending the audio object across a larger number of speakers because the determination of extended object speaker gain takes into account the extension of the audio object and typically uses a relatively broad (e.g., even broader than the extension of the audio object under consideration) stable (e.g., steadily attenuating) distribution characteristic.

[0026] Therefore, by combining the translation object speaker gain based on the point source translation of the audio object with the extended object speaker gain that takes into account the extension of the audio object, good localization of the audio object can be achieved (even for objects with relatively large extensions), while still allowing the extension of the audio object to be perceived. This is especially true when there are multiple users and / or users who are not located in the "sweet spot" of the listening arrangement. By always introducing the contribution of the translation object speaker gain, for example, independently of extension information or independent of object feature information, it is possible to always ensure the localization of the audio object, even for listeners who are not in the sweet spot.

[0027] Furthermore, it should be noted that object characteristic information typically allows for the identification or estimation of the extension of an audio object. For example, object characteristic information can indicate the type of object, which can imply parameters (e.g., extension parameters) used to determine the speaker gain of the extended object. For example, object characteristic information can allow for the differentiation between relatively small and relatively large objects. Alternatively or additionally, object characteristic information can allow for the differentiation between near and far objects, which can also imply one or more parameters used to determine the speaker gain of the extended object.

[0028] In summary, this audio object renderer can derive one or more parameters from object feature information to determine the extended object speaker gain. Therefore, the object feature information allows for appropriate adjustment of the derived extended object speaker gain so that the object can be positioned by providing translational object speaker gain, and also perceived with an appropriate extension because the object feature information is taken into account when providing the extended object speaker gain.

[0029] In summary, the aforementioned audio object renderer allows for the determination of speaker gain that provides a good auditory impression while maintaining reasonably low computational complexity.

[0030] To summarize further, the present invention generally creates an object renderer that uses VBAP to translate objects, then determines the object feature gain of the object, and always combines the object feature gain in relation to the VBAP translated object.

[0031] According to another embodiment of the invention, an audio object renderer is created for determining speaker gain (e.g., combined speaker gain or resulting speaker gain) based on object position information (e.g., azi and / or elevation angles) (which may be provided, for example, in spherical coordinates, such as using azi and elevation values ​​ele) and extension information (e.g., extension angle information describing the extension in the azi direction (e.g., spreadAngleAzi) and / or extension angle information describing the extension in the elevation direction (e.g., spreadAngleEle)). The speaker gain describes the gain of including one or more audio object signals into multiple speaker signals.

[0032] The audio object renderer is configured to use point source translation of the audio object (where, for example, the audio object is treated as a point source, where, for example, extended information is ignored, and where, for example, the signal of the audio object is associated with two or more speakers in the environment of the object location of the audio object by appropriate selection of translation object speaker gain) to obtain translation object speaker gain (also referred to as "object speaker gain", or denoted by vector g).

[0033] The audio object renderer is configured to take into account object position information and extension information to obtain the extended object speaker gain (e.g., also referred to as extended speaker gain, or, for example, represented as a vector gOS).

[0034] The audio object renderer is configured to combine the translation object speaker gain (e.g., g) and the extended object speaker gain (e.g., gOS) in a manner that always has a contribution from the translation object speaker gain (e.g., independent of the extended information) in order to obtain a combined speaker gain.

[0035] This audio object renderer is based on the same considerations as the audio object renderers described above. However, it evaluates extension information rather than object feature information, which directly describes how the object should extend. For example, extension information could be extension angle information describing extension in the azimuth direction and / or extension in the elevation direction. Alternatively, extension information could also be stereo angle information, or the size of the object could be specified in any other form (e.g., using absolute size information and / or distance information, etc.). Therefore, it is possible to obtain combined speaker gain such that the combined speaker gain allows the object signal to be represented with a good audible impression, enabling the listener to locate the object and perceive it with appropriate extension. This can be achieved, for example, by always using translation, even if the object's extension is sufficiently wide.

[0036] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to evaluate one or more gain functions (e.g., one or more polynomial functions or one or more parabolic functions; e.g., extended weighted curves) that map the difference between the location of a support point (e.g., an artificially generated support point; e.g., specified by SSP) and the object location to one or more extended gain value contributions (e.g., aziGain(naz) or eleGain(nel)), and determine the extended object speaker gain (e.g., gOS) based on the one or more extended gain value contributions.

[0037] By using the location of the support point and evaluating one or more gain functions that map the difference between the support point's location and the object's location to one or more extended gain value contributions, a uniform and computationally efficient scheme for determining the extended gain values ​​can be obtained. For example, the gain functions can weight the angular differences between the object's location and the support point's location (e.g., in both azimuth and elevation), allowing for a simple yet accurate determination of the extended gain value contributions. Furthermore, using the support point location instead of the speaker location ensures high uniformity, which generally helps reduce algorithmic complexity. For example, the mapping from the signal gain associated with the support point to the signal gain associated with the actual speaker can be pre-calculated, eliminating the need for recalculation for each audio object. However, algorithms with relatively low complexity can be used to determine the gain values ​​associated with support points that are typically geometrically regular. Therefore, the combined speaker gain can be determined with high efficiency.

[0038] In the preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to expand according to the spread in the first direction (e.g., spread). azi And according to the expansion in the second direction (e.g., spread) ele This determines the weight of the extended object speaker gain in combination with the translation object speaker gain.

[0039] Therefore, it can be determined how important (or weighted) the determination of the extended object speaker gain is compared to the determination of the translation object speaker gain. For example, for a large extension angle, the relative importance of the extended object speaker gain in the combination can increase compared to the case where the extension is relatively small. Furthermore, by adjusting the weight of the extended object speaker gain, it can be ensured that the (relative) contribution of the translation object speaker gain is reduced in cases where the extension is relatively wide, thus avoiding an unpleasant auditory impression. Conversely, when the extension is relatively small, the (relative) contribution of the translation object speaker gain can be increased, thus reflecting the localized nature of the audio object. Therefore, if, for example, the extension of the audio object can increase over time (perhaps, for example, in the case where the audio object under consideration is a moving audio object), a smooth transition is also possible. In summary, determining the weight of the extended object speaker gain in the combination with the translation object speaker gain allows for a good auditory impression, even in the case of a wide extension object.

[0040] In the preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to display an extension angle (or normalized extension angle or normalized extension range or weighted extension range) in a first direction (e.g., spread). azi ) and the spread angle in the second direction (or normalized spread angle or normalized spread range or weighted spread range) (e.g., spread) ele The product of ) determines the weight of the extended object loudspeaker gain in the combination with the translation object loudspeaker gain (at least at a predetermined threshold (e.g., g)). res (The following extended angle range).

[0041] It has been found that the product of the extension angle in the first direction and the angle in the second direction reflects the extension characteristics of the audio object well. Specifically, if the audio object includes a very small extension angle in one direction, the product will be relatively small, and the object will be considered relatively localized. However, it has been found that the product of the extension angle in the first direction and the extension angle in the second direction (which may be, for example, perpendicular to the first direction) plays a good role in adjusting the weight of the extended object speaker gain in combination with the translation object speaker gain, and can be determined with very high computational efficiency.

[0042] In the preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to apply a translation object speaker gain (e.g., a vector g of translation object speaker gains) weighted with a fixed weight (e.g., 1) and a variable weight (e.g., attenGain or g) to the translation object speaker gain. attenThe weighted extended object speaker gains (e.g., vector gOS) are added together, the variable weights depending on the spread angle in the first direction (e.g., spread). azi ) and the extension angle in the second direction (e.g., spreadad) ele ), and can optionally be restricted to no greater than a fixed weight, such as, for example, when used for g atten As shown in the equation.

[0043] Using this method, the translation object speaker gain can be given sufficient weight in the combination regardless of the extension angle, while still allowing for effective (relative) weighting between the translation object speaker gain and the extension object speaker gain. However, by ensuring that the contribution of the translation object speaker gain does not fall below a certain minimum weight, it ensures that the object can always be reasonably localized, regardless of the listener's position in the speaker environment. This allows for avoiding strong degradation of the audio impression while maintaining the possibility of adjusting the perceptual extension of the audio object.

[0044] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to normalize the result of adding a translation object speaker gain weighted with a fixed weight and an extended object speaker gain weighted with a variable weight (e.g., by dividing the sum by the norm of the sum).

[0045] Using this normalization, it is possible to achieve an expansion of the total energy or total perceived loudness that is largely independent of one or more audio objects. Therefore, loudness adjustment can be performed separately from the adjustment of the expansion.

[0046] In the preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to determine the weight attenGain of the extended object speaker gain in combination with the translated object speaker gain according to the following formula:

[0047] attenGain = 0.89f * min(c1, max(spread azi spread ele ) / g res1 ) + 0.11f *min(c2, min(spread azi spread ele ) / g res2 );

[0048] Where c1 is a predetermined value (e.g., 1), which can, for example, limit the contribution of attenGain to a first factor, and c2 is a predetermined value (e.g., 1), which can, for example, limit the contribution of attenGain to a second factor, and g res1It is a predetermined value (e.g., the azimuth interval of the support point location, such as 45 degrees, or, for example, the SSP grid resolution), where g res2 It is a predetermined value (e.g., the elevation angle interval of the support point location, such as 45 degrees, or, for example, the SSP grid resolution), where spread azi It is the spread angle of the audio object in the azimuth direction, where spread ele It is the expansion angle of the audio object in the elevation direction, where min(.) is the minimum operator and max(.) is the maximum operator.

[0049] By determining the weights of the extended speaker gain in combination with the translational object speaker gain, a particularly good auditory impression can be achieved. In this calculation, the weights depend on both the smaller and larger of the extension angles, with the larger angle being given a larger weight. It has been found that this evaluation of the two-dimensional extension (i.e., this determination of the weighting of the extended object speaker gain) results in a relative weighting of the extended object speaker gain and the translational object speaker gain, which leads to a good auditory impression.

[0050] However, it should be noted that the factors 0.89f and 0.11f can be changed, where applied to min(c1, max(spread)). azi spread ele ) / g res1 The first factor should be greater than the second factor (preferably, at least 50%).

[0051] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to increase the relative contribution of the extended object speaker gain as the extension angle of the audio object increases (e.g., as the product of the extension angle in the azimuth direction and the extension angle in the elevation direction increases) compared to the translation object speaker gain, for example until the extension angle in the azimuth direction and the extension angle in the elevation direction reach predetermined values.

[0052] It has been found that this concept leads to a particularly good auditory impression because both the extension angle in the first direction (e.g., azimuth direction) and the extension angle in the second direction (e.g., elevation direction) are considered and contribute to the weighting of the extension object loudspeaker gain in the aforementioned combination of translation object loudspeaker gain and extension object loudspeaker gain. However, by imposing specific constraints (which can be defined by a “predetermined value”), it is still possible to avoid overweighting the extension object loudspeaker gain, which would reduce the likelihood of sound source localization.

[0053] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to take into account object position information and extension information and use the representation of the support point position in polar coordinates (e.g., aziSSP and / or eleSSP) to obtain the extended object speaker gain (also referred to as extended speaker gain, or represented as a vector gOS), and the audio object renderer is configured to provide speaker gain based on the extended object speaker gain.

[0054] It has been found that representing the support point location in polar coordinates often greatly facilitates computation, as it is very helpful in representing extensions in both azimuth and elevation. Furthermore, when using polar coordinates, the radius component generally does not need to be considered, since it is usually appropriate to assume a predetermined radius for the support point location. Therefore, for example, the support point can be represented using only two polar coordinates (e.g., azimuth and elevation). Consequently, computational complexity is generally reduced.

[0055] In the preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured as follows:

[0056] - Evaluate one or more angular differences between the azimuth position of the audio object and the azimuth positions of one or more support points (e.g., diffCLKDir, diffAntiCLKDir, and optionally, diffCLKDir+(n-1)*Spread.openAngle and diffAntiCLKDir+(n-1)Spread.openAngle), and / or

[0057] - Evaluate one or more angular differences between the elevation position of the audio object and the elevation positions of one or more support points (e.g., diffCLKDir, diffAntiCLKDir).

[0058] In order to obtain extended speaker gain.

[0059] It has been found that the evaluation of such angular differences can be performed with minimal computational effort. Furthermore, it has been found that angular differences can be mapped to values ​​or gain values ​​associated with support points and describing, for example, how the audio signal of an audio object should be presented to the support points. Moreover, in a computationally efficient implementation, support points can be arranged in a highly regular manner such that, for example, the azimuth difference to be evaluated is the same for different elevation angles, and the elevation difference to be evaluated is the same for different azimuth angles. Using such a concept, particularly high computational efficiency can be achieved because it is only necessary to evaluate the azimuth difference for one elevation angle and the elevation difference for one azimuth angle. Therefore, the concept mentioned here can help reduce computational effort, facilitating its implementation in devices with limited computing resources.

[0060] In the preferred embodiment of the previously mentioned audio object renderer, the support points are positioned on the sphere within a tolerance of + / -10% or + / -20% of the sphere radius.

[0061] By using supports arranged on a sphere, the radius typically does not need to be explicitly considered. Furthermore, it has been found that such support placement is particularly helpful when expressing the extension of an audio object in angular values ​​(e.g., elevation and azimuth). In contrast, using Cartesian coordinates significantly increases the computational workload, as three spatial coordinates would need to be considered, and computationally expensive trigonometric functions would typically need to be applied in the calculations to convert between Cartesian coordinates and angular representations.

[0062] In a preferred embodiment of the previously mentioned audio object renderer, the support point location includes a circle with a constant elevation angle (e.g., -135 degrees, -90 degrees, -45 degrees, 0 degrees, 45 degrees, 90 degrees, or 135 degrees) and a constant radius (e.g., a normalized radius of 1) or even a uniform azimuth interval (e.g., 45 degrees) along multiple circles with different constant elevation angles and constant radii.

[0063] Alternatively or additionally, the location of the support point may include a circle with a constant azimuth (e.g., -135 degrees, -90 degrees, -45 degrees, 0 degrees, 45 degrees, 90 degrees, or 135 degrees) and a constant radius (e.g., a normalized radius of 1) or even a uniform elevation angle interval (e.g., 45 degrees) along multiple circles with different constant azimuths and constant radii.

[0064] The computational workload can be reduced by using uniform azimuth intervals along circles with constant elevation angles and uniform elevation angle intervals along circles with constant azimuth angles. Furthermore, if uniform azimuth intervals exist along multiple circles with different constant elevation angles and constant radii, the computational workload can be further reduced while still maintaining good coverage of the entire sphere surface, because if all multiple circles with different elevation angles and constant radii have uniform azimuth intervals, significant portions do not need to be recalculated for these circles. For example, the calculation of azimuth differences only needs to be performed for one of the circles, and the result can also be applied to other circles with the same uniform azimuth interval as the first circle under consideration. The same applies to multiple circles with different constant azimuth angles and constant radii, and preferably with the same uniform elevation angle interval. The elevation angle difference only needs to be calculated once, and its result can be applied to the support point positions on another circle with the same uniform elevation angle interval. Therefore, calculations can be performed on a large number of support points with a very small computational workload.

[0065] In a preferred embodiment of the aforementioned audio object renderer, the object renderer is configured to achieve extended object speaker gain such that the audio object extends over a region that extends in a first hemisphere where the audio object is located and also in a second hemisphere with an azimuth position opposite to that of the first hemisphere. For example, the first hemisphere could be the hemisphere in front of the listener's position, for example, having an azimuth angle between -90 degrees and +90 degrees, where 0 degrees is the listener's viewing direction, and the second hemisphere could be, for example, the hemisphere behind the listener's position, for example, having an azimuth angle between -180 degrees and -90 degrees or between 90 degrees and 180 degrees, or vice versa.

[0066] By extending the audio object over a region that extends in both the first and second hemispheres, the object can be effectively positioned above or below the user's head. Therefore, it is possible to create the auditory impression that the extended object is above the listener's head, so large that it appears both in front of and behind the listener's head. This provides a remarkably realistic auditory impression.

[0067] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to use an extended elevation range between -180 degrees and +180 degrees.

[0068] By extending the elevation angle range, discontinuities can be avoided and calculations can be simplified, reducing the number of case distinctions and angle corrections. For example, it is simpler to represent (and mathematically or algorithmically represent) that an object should be presented for a given azimuth and for elevation angles of 80 degrees and 100 degrees, rather than indicating that the object should be presented for an 80-degree elevation angle in two opposite azimuth angles (e.g., 0 degrees and 180 degrees). Therefore, by doubling the elevation angle range (which normally extends from -90 degrees to +90 degrees), calculations can be simplified, and the representation of the calculation results can be simplified as well.

[0069] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to calculate the following terms for a given object location (e.g., defined by the azimuth value azi and the elevation value ele) and for a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle):

[0070] - A first set of azimuth gain values ​​(e.g., aziGain) describes the contribution of multiple azimuth values ​​associated with the support point location or support point azimuth index (e.g., naz) to the spread gain (e.g., g_spd) (e.g., using a polynomial or parabolic function with a width appropriate to the object spread width in the azimuth direction). This first set of azimuth gain values ​​is associated with elevation values ​​in the original range of elevation values ​​(e.g., -90 degrees to +90 degrees) that do not cross the poles of the spherical coordinate system.

[0071] - A second set of azimuth gain values ​​(e.g., aziGainExtd) describes the contribution of multiple azimuth values ​​associated with the anchor point location or anchor point azimuth index (e.g., naz) to the extended gain (e.g., using a polynomial or parabolic function with a width suitable for the object's extension width in the azimuth direction). This second set of azimuth gain values ​​is associated with elevation values ​​in an extended range of elevation values ​​(e.g., -180 degrees to -90 degrees and +90 degrees to +180 degrees) indicating one of the poles of the spherical coordinate system (e.g., a pole at -90 degrees elevation or a pole at +90 degrees elevation), which may correspond, for example, to the extension of the audio object at one of the poles of the spherical coordinate system.

[0072] The audio object renderer is also configured to derive extended gain (which, for example, is used to determine the extended object speaker gain gOS) using a first set of azimuth gain values ​​(e.g., aziGain(naz)) and a second set of azimuth gain values ​​(e.g., aziGainExtd).

[0073] By calculating two sets of azimuth gain values ​​(one set for the original (or basic) range of elevation values ​​(or associated with it), and one set for the extended range of elevation values ​​(or associated with it)), the extension of an audio object above (or below) the listener's head can be calculated in a particularly efficient manner. Specifically, these sets of azimuth gain values ​​have been found to be useful for efficiently deriving the extended gain, for example, using predefined combination mappings. On the other hand, the computation and the representation of the results are facilitated because elevation values ​​in the extended range of elevation values ​​can be easily derived from those in the original range of elevation values ​​using addition or subtraction without distinguishing between multiple cases and without changing the azimuth values.

[0074] For example, when extending an audio object above a user's head, the extension above the user's head can be easily calculated using an elevation angle greater than 90 degrees. Therefore, an object with a given azimuth angle (e.g., within the range of -90 to +90 degrees) and a positive elevation angle between 0 and 90 degrees can be easily extended above the user's head using an elevation angle greater than 90 degrees, while maintaining the azimuth angle (within the range of -90 to +90 degrees) unchanged. Thus, the elevation gain value associated with the elevation angle values ​​within the range of extended elevation angle values ​​(in this example, between +90 and +180 degrees) can be obtained as an intermediate quantity and can later be mapped back to the support point or the coordinate system, for example, using an elevation angle only within the original elevation angle value range combined with the azimuth gain values ​​from the second set of azimuth gain values. In summary, the described concept significantly improves computational efficiency when extending an audio object above (or below) the user's head.

[0075] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to calculate the following terms for a given object location (e.g., defined by the azimuth value azi and the elevation value ele) and for a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle):

[0076] - A first set of elevation gain values ​​(e.g., eleGain) describes the contribution of multiple elevation values ​​associated with the support point location or speaker azimuth index or support point elevation index (e.g., nel) to the extended gain (e.g., using a parabolic function whose width is adapted to the object's extended width in the azimuth direction). This first set of elevation gain values ​​is associated with elevation values ​​in the original elevation value range (e.g., -90 degrees to +90 degrees) indicating that the poles of the spherical coordinate system are not crossed.

[0077] - A second set of elevation gain values ​​(e.g., eleGainExtd) describes the contribution of multiple elevation values ​​associated with the anchor point location or speaker elevation index or anchor point elevation index (e.g., nel) to the extended gain (e.g., using a parabolic function whose width is adapted to the object's extension width in the azimuth direction). This second set of elevation gain values ​​is associated with elevation values ​​in an extended range of elevation values ​​(e.g., -180 degrees to -90 degrees and +90 degrees to +180 degrees) indicating one of the poles of the spherical coordinate system (e.g., a pole at an elevation angle of -90 degrees or a pole at an elevation angle of +90 degrees), which may correspond, for example, to the extension of the audio object at one of the poles of the spherical coordinate system.

[0078] The audio object renderer is also configured to export extended gain using the first set of azimuth gain values ​​(aziGain(naz)), the second set of azimuth gain values ​​(aziGainExtd), the first set of elevation gain values ​​(eleGain(nel)), and the second set of elevation gain values ​​(eleGainExtd(nel)).

[0079] This embodiment is based on considerations similar to those of the embodiment that calculates the azimuth gain value.

[0080] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is further configured to combine (e.g., by multiplication) the values ​​of a first set of azimuth gain values ​​and a first set of elevation gain values ​​(e.g., corresponding values, such as aziGain(naz), eleGain(nel)), and combine the values ​​of a second set of azimuth gain values ​​and a second set of elevation gain values ​​(e.g., corresponding values, such as aziGainExtd(naz), eleGainExtd(nel)).

[0081] Using this combination, values ​​associated with the same location can be combined (e.g., added). For example, each support point location can be referenced by a combination of a first azimuth value and a first elevation value within the original elevation range, and a combination of a second azimuth value and a second elevation value within the extended elevation range. For example, a point (e.g., a support point location) can be specified by a +10 degree azimuth at +80 degrees elevation, and can also be specified by a -170 degree azimuth and a +100 degree elevation. In other words, it is possible to combine values ​​associated with the same point but referenced by different combinations of azimuth and elevation values, such as azimuth gain and elevation gain values.

[0082] For example, a second set of azimuth gain values ​​may include the same azimuth gain values ​​that are included in the first set of azimuth gain values ​​but associated with the opposite direction (or azimuth value). For instance, the second set of azimuth gain values ​​may indicate the same gain at an azimuth of -170 degrees (or typically x-180 degrees or x+180 degrees) as the first set of azimuth gain values ​​indicate at an azimuth of +10 degrees (or typically x degrees). However, the actual location specified by an azimuth of +10 degrees and an elevation of +80 degrees (or typically y degrees) is actually the same location specified by an elevation value within an extended range of elevation values ​​of -170 degrees and +100 degrees (or typically 180 degrees-y). Therefore, the elevation gain value associated with the elevation value within the extended range of elevation values ​​should be combined with the azimuth gain value associated with the "opposite" azimuth. Thus, the second set of azimuth gain values ​​effectively describes the contribution of the extended gain to the "opposite" azimuth value.

[0083] In summary, by combining the first set of azimuth gain values ​​with the first set of elevation gain values ​​and by combining the second set of azimuth gain values ​​with the second set of elevation gain values, two pairs of values ​​associated with the same location (described by different azimuth and elevation values) can be combined to obtain a meaningful gain value associated with the support point.

[0084] In the preferred embodiment of the previously mentioned audio object renderer, the second set of azimuth gain values ​​represents the evolution of the gain value at an azimuth angle shifted by 180 degrees compared to the evolution of the gain value at an azimuth angle represented by the first set of azimuth gain values.

[0085] By using this representation of the azimuth gain values, and employing two sets of azimuth gain values, it can be assumed that the elevation values ​​in the extended elevation range specify the same point as the elevation values ​​in the original elevation range when considering the modified azimuth. Therefore, representing the evolution of the gain value over two shifted azimuth ranges using two sets of azimuth gain values ​​allows for simple and computationally efficient processing of the elevation angles in the extended elevation range, as the "second set of azimuth gain values" can be easily combined with the elevation gain values ​​associated with the extended elevation range.

[0086] In the preferred embodiment of the previously mentioned audio object renderer, the first set of azimuth gain values ​​represents the evolution of the gain value over a 360-degree range, given the azimuth object position and the azimuth extension angle, wherein the angular accuracy is determined by the number of speakers or by the number of support points.

[0087] Alternatively or additionally, the second set of azimuth gain values ​​represents the evolution of the gain value over a 360-degree range, given the azimuth object position and azimuth extension angle after a 180-degree rotation, wherein the angular accuracy is determined by the number of speakers or by the number of support points.

[0088] By using a set of azimuth gain values ​​that represent the evolution of gain over a 360-degree range, the entire environment of the listener can be efficiently considered, taking into account both audio objects in front of and behind the listener. These sets of azimuth gain values, representing the evolution of gain over a 360-degree range, can also handle audio objects positioned above or below the user's head.

[0089] In the preferred embodiment of the previously mentioned audio object renderer, the first set of elevation gain values ​​represents the evolution of the gain values ​​over the elevation range of -90 degrees to +90 degrees, given the elevation object position (which is in the range of -90 degrees to +90 degrees) and the elevation extension angle (e.g., for cases where the audio object has not yet extended at the pole of the spherical coordinate system).

[0090] Alternatively or additionally, the second set of elevation gain values ​​represents the evolution of the gain values ​​over the elevation range between -180 degrees and -90 degrees and between +90 degrees, given the position of the elevation object (which is in the range between -90 degrees and +90 degrees) and the elevation extension angle (e.g., for cases where the audio object has extended to the pole of the spherical coordinate system).

[0091] By using these sets of elevation gain values, it is easy to handle the extension of an object above the user's head, because the angle values ​​can be simply added or subtracted without exceeding the range of elevation angles covered by the first set of elevation gain values ​​and the second set of elevation gain values.

[0092] According to embodiments of the present invention, a method is created for determining speaker gain (e.g., combined speaker gain or resulting speaker gain) based on object location information (e.g., azi, elevation angle) (which may be provided, for example, in spherical coordinates, such as using azi and elevation values ​​ele) and object feature information (e.g., extended angle information describing the extension in the azi direction (e.g., spreadAngleAzi) and / or extended angle information describing the extension in the elevation direction (e.g., spreadAngleEle)). The speaker gain describes the gain of including one or more audio object signals into multiple speaker signals.

[0093] The method includes obtaining a translation object speaker gain (also referred to as "object speaker gain", or denoted by vector g) by using a point source translation of an audio object (where the audio object is treated as a point source, extended information is ignored, and the signal of the audio object is associated with two or more speakers in the environment of the object location of the audio object by an appropriate selection of the translation object speaker gain).

[0094] The method also includes taking into account object location information and object feature information to obtain extended object speaker gain (also known as extended speaker gain, or represented as vector gOS).

[0095] The method includes combining the translation object speaker gain (e.g., g) and the extended object speaker gain (e.g., gOS) in a manner that always has a contribution from the translation object speaker gain (e.g., independent of the extended information) to obtain a combined speaker gain.

[0096] This method is based on the same considerations as the corresponding device described above.

[0097] Furthermore, the method may optionally be supplemented individually and in combination by any features, functions, and details described with respect to the corresponding apparatus described above.

[0098] According to embodiments of the present invention, a method is created for determining speaker gain (e.g., combined speaker gain or resulting speaker gain) based on object location information (e.g., azi, elevation angle) (which may be provided, for example, in spherical coordinates, such as using azi and elevation values ​​ele) and extension information (e.g., extension angle information describing the extension in the azi direction (e.g., spreadAngleAzi) and / or extension angle information describing the extension in the elevation direction (e.g., spreadAngleEle)). The speaker gain describes the gain of including one or more audio object signals into multiple speaker signals.

[0099] The method includes obtaining a translation object speaker gain (also referred to as "object speaker gain", or denoted by vector g) by using a point source translation of an audio object (where the audio object is treated as a point source, extended information is ignored, and the signal of the audio object is associated with two or more speakers in the environment of the object location of the audio object by an appropriate selection of the translation object speaker gain).

[0100] This method involves taking into account object position information and extension information to obtain the extended object speaker gain (also known as extended speaker gain, or represented as a vector gOS).

[0101] The method includes combining the translation object speaker gain (e.g., g) and the extended object speaker gain (e.g., gOS) in a manner that always has a contribution from the translation object speaker gain (e.g., independent of the extended information) to obtain a combined speaker gain.

[0102] This method is based on the same considerations as the corresponding device described above.

[0103] Furthermore, the method may optionally be supplemented individually and in combination by any features, functions, and details described with respect to the corresponding apparatus described above.

[0104] In a preferred embodiment of the previously mentioned method, the method includes evaluating one or more gain functions (e.g., one or more polynomial functions or one or more parabolic functions; e.g., extended weighted curves) that map the difference between the location of a support point (e.g., an artificially generated support point; SSP) and the object location to one or more extended gain value contributions (e.g., aziGain(naz) or eleGain(nel)), and determining an extended object speaker gain (e.g., gOS) based on the one or more extended gain value contributions.

[0105] According to an embodiment of the present invention, a computer program is created for executing the aforementioned method when the computer program is run on a computer.

[0106] This computer program can be supplemented individually and in combination by any of the features, functions and details described herein.

[0107] Other embodiments of the invention will be discussed below. These embodiments can be used alone or in combination with any other embodiments disclosed herein. In other words, alternatively, any features, functions, and details of the embodiments discussed below may be incorporated, either alone or in combination, into any other embodiments disclosed herein.

[0108] An audio object renderer (200, 1300) is created according to an embodiment of the present invention for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1310, azi, ele) and object feature information (1312). The speaker gains (214, 1214, 1214a~c) describe the gain of including one or more audio object signals (1260) into multiple speaker signals (1262a~1262c). The object renderer is configured to use one or more polynomial functions of degree less than or equal to 3 to obtain the object feature information gain (314a, g_spd).

[0109] In a preferred embodiment, the object renderer is configured to obtain object feature information speaker gain (206a, 1242, gOS) using object feature gain (314a, g_spd) based on object feature gain contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd).

[0110] In a preferred embodiment, the object feature information is extended information (212, 1312, spreadAngleAzi, spreadAngleEle).

[0111] An audio object renderer is created according to embodiments of the present invention for determining speaker gain (e.g., combined speaker gain) based on object location information (e.g., azi and ele) (which may be provided, for example, in spherical coordinates, such as using azi and ele values) and object feature information. The speaker gain describes the gain of including one or more audio object signals into multiple speaker signals. The object feature information may be, for example, information indicating whether the object is small or extended (e.g., object size values), or it may be, for example, object distance information that can be mapped to an extension value (e.g., mapped to extension angle information describing extension in the azimuth direction (e.g., spreadAngleAzi) and / or mapped to extension angle information describing extension in the elevation direction (e.g., spreadAngleEle)). However, other types of object feature information are also possible.

[0112] The object renderer is configured to take into account object position information and object feature information to obtain extended object speaker gain (also known as extended speaker gain, or denoted as vector gOS).

[0113] The object renderer is configured to use one or more polynomial functions of degree less than or equal to 3 (e.g., one or more parabolic functions or cubic polynomial functions (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1) to obtain the spread gain (e.g., object-to-support spread gain, such as g_spd), such as elements of the vector g_spd, which describes the contribution of the audio object signal to multiple speaker signals or to multiple support signals, and one or more polynomial functions, such as mapping the angular difference between the object position and the support position (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) to the spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)).

[0114] The object renderer is configured to use extended gain (g_spd) based on extended gain contribution to obtain extended object speaker gain.

[0115] This embodiment is based on the finding that polynomial functions of degree 2 or less are particularly well-suited for obtaining object extension gain based on the angular difference between the object's position and the support point's position. It has been recognized that polynomial functions of degree 3 or less can be evaluated with a modest computational effort and are well-suited for situations with limited computational resources. However, it has been found that such polynomial functions still provide a good approximation of the characteristics required to obtain object extension gain that delivers a good auditory impression of the extended audio object. Specifically, it has been found that polynomial functions of degree 3 or less can be evaluated more easily than other functions (e.g., exponential functions) that require very high computational effort or large lookup tables. Therefore, this audio object renderer can be implemented with a particularly small computational effort.

[0116] Furthermore, it should be noted that object characteristic information can be used to adjust the determination of extended object speaker gain. For example, object characteristic information can determine the extension width. For example, object characteristic information can be information indicating whether the object is small or extended (e.g., including object size values), or object characteristic information can include, for example, object distance information, which can be mapped to extension values ​​(e.g., mapped to extension angle information describing extension in the azimuth direction (e.g., spreadAngleAzi) and / or mapped to extension angle information describing extension in the elevation direction (e.g., spreadAngleEle)). In other words, it should be noted that object characteristic information generally allows the determination or estimation of the extension of an audio object. For example, object characteristic information can indicate the type of object, where the type of object can imply parameters used to determine the extended object speaker gain (e.g., extension parameters). For example, object characteristic information can allow the differentiation between relatively small and relatively large objects. Alternatively or additionally, object characteristic information can allow the differentiation between near and far objects, which can also imply one or more parameters used to determine the extended object speaker gain.

[0117] In summary, this audio object renderer can derive one or more parameters from object feature information to determine the extended object speaker gain. Therefore, the object feature information allows for appropriate adjustment of the derived extended object speaker gain, resulting in a pleasing auditory impression.

[0118] In summary, by using one or more polynomial functions of degree less than or equal to 3 and whose parameters may depend, for example, on object characteristic information, the extended object speaker gain can be determined efficiently, thus keeping the computational workload reasonably small.

[0119] However, it should be noted that the above method can optionally be applied in a more general form. Specifically, object expansion will not necessarily be performed. For example, if support points are pre-presented so that object characteristics (or object properties or object features) can be represented by applying a weighted curve to the support points, the weighted curve can be implemented, for example, for efficiency reasons, using a cubic polynomial.

[0120] An audio object renderer is created according to embodiments of the present invention for determining speaker gain (e.g., combined speaker gain) based on object position information (e.g., azi, ele) (which may be provided, for example, in spherical coordinates, such as using azi and ele values) and extension information (e.g., extension angle information describing the extension in the azi direction (e.g., spreadAngleAzi) and / or extension angle information describing the extension in the elevation direction (e.g., spreadAngleEle)). The speaker gain describes the gain of including one or more audio object signals into multiple speaker signals.

[0121] The object renderer is configured to take into account object position information and extension information to obtain the extended object speaker gain (also known as extended speaker gain, or denoted as vector gOS).

[0122] The object renderer is configured to use one or more polynomial functions of degree less than or equal to 3 (e.g., parabolic functions or cubic polynomial functions (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1) to obtain the spread gain (e.g., object-to-support spread gain, such as g_spd), such as spread gain values ​​(e.g., elements of the vector g_spd). The spread gain describes the contribution of the audio object signal to multiple speaker signals or multiple support signals. One or more polynomial functions map the angular difference between the object position and the support position (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) to the spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)).

[0123] The object renderer is configured to use extended gain (g_spd) based on extended gain contribution to obtain extended object speaker gain.

[0124] Generally, the function used (e.g., a polynomial function) should preferably (but not necessarily) have (at least roughly) a curve shape of point-source translation (in this example: VBAP). The function does not necessarily need to be parable. The idea is to translate extended components in the same way between support points (e.g., using VBAP to normally translate objects between two speakers).

[0125] This audio object renderer is based on the same considerations as the audio object renderer described above. However, it evaluates extension information rather than object feature information; extension information directly describes how the object should extend. For example, extension information could be extension angle information describing extension in the azimuth direction and / or extension in the elevation direction. Alternatively, extension information could also be stereo angle information, or the size of the object could be specified in any other form (e.g., using absolute size information and / or distance information, etc.). Therefore, the extended object speaker gain can be obtained computationally efficiently, where the extension information can, for example, be used to adjust the parameters of one or more polynomial functions (e.g., the width of the analog scale used to obtain the extended gain). Thus, calculations can be easily tuned to the actual extension indicated by the extension information and can be performed efficiently.

[0126] In a preferred embodiment of the audio object renderer mentioned above, the width of one or more polynomial functions (e.g., parabolic functions, which may be determined, for example, by a scaling value aziParable or a scaling value eleParable) is determined by extended information or by object feature information (wherein the audio object renderer may be configured, for example, to adapt the width of one or more polynomial functions or parabolic functions to an extended width associated with different audio objects).

[0127] It has been found that the width of a polynomial function (e.g., a parabolic function) can be easily adjusted because polynomial functions can be easily parameterized. However, evaluating a parabolic function (e.g., a polynomial function) that includes one or more parameters is often possible without excessive computational effort. Furthermore, by adjusting one or more parameters of the parabolic function based on extension information or object characteristic information, the extension width can be adjusted very smoothly, resulting in a very good perceptual impression.

[0128] In a preferred embodiment of the previously mentioned audio object renderer, the object renderer is configured to use a first polynomial function (e.g., a polynomial function of degree less than or equal to 3, such as a parabolic function) and a second polynomial function (e.g., a polynomial function of degree less than or equal to 3, such as a parabolic function) to obtain extended gain values ​​(e.g., the value of the vector g_spd). The first polynomial function maps the azimuth difference between the object position and the support point position to a first extended gain contribution (e.g., aziGain(naz)), and the second polynomial function maps the elevation difference between the object position and the support point position to a second extended gain contribution (e.g., eleGain(nel)).

[0129] This concept is based on the idea that a two-dimensional expansion function can be determined efficiently using a combination (e.g., multiplication) of two expansion functions. This two-dimensional expansion function can depend, for example, on both the difference in elevation angle between the audio object position and the expansion support point position, and the difference in azimuth angle between the audio object position and the expansion support point position, with one expansion function applied to the azimuth difference and the other to the elevation difference. In other words, it has been found that good expansion results can be achieved by applying two separate parabolic expansion functions to the azimuth and elevation differences (where the results are multiplied). Specifically, it has been found that this type of two-dimensional expansion results in a reasonably good auditory impression while keeping the computational workload relatively small. Specifically, evaluating the two expansion functions individually is generally much less computationally demanding than evaluating the joint two-dimensional function, while providing a good auditory impression.

[0130] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to combine (e.g., by multiplication) a first extended gain contribution (e.g., aziGain(naz)) and a second extended gain contribution (e.g., eleGain(nel)) to obtain an extended gain value (e.g., the value of the vector g_spd).

[0131] Two-dimensional expansion can be achieved by combining the extended gain contributions (which can be performed, for example, using multiplication), which depends on, for example... Figure 3 The two-dimensional extension of the extension function shown is well approximated. In other words, it has been found that multiplying the two extension gain contributions obtained based on two polynomial functions results in a two-dimensional extension characteristic that provides a reasonably good auditory impression.

[0132] In a preferred embodiment of the audio object renderer mentioned above, the object renderer is configured to calculate the following terms for a given object location (e.g., defined by the azimuth value azi and the elevation value ele) and for a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle):

[0133] - A set of azimuth gain values ​​(e.g., aziGain) describing the contribution of multiple azimuth values ​​associated with the support point location or speaker azimuth index or support point azimuth index (e.g., naz) to the extended gain (e.g., using a polynomial function with a degree less than or equal to 3 or a parabolic function with a width suitable for the object's extended width in the azimuth direction), and / or

[0134] - A set of elevation gain values ​​(e.g., eleGain) describing the contribution of multiple elevation values ​​associated with the anchor point location or speaker elevation index or anchor point elevation index (e.g., nel or naz) to the extended gain (e.g., using a polynomial function with a power less than or equal to 3 or a parabolic function with a width suitable for the object's extended width in the elevation direction).

[0135] And use the set of azi gain values ​​(aziGain(naz)) and / or use the set of elevation gain values ​​(eleGain(nel)) to derive the extended gain.

[0136] By determining a set of azimuth gain values ​​associated with multiple support point locations and / or multiple elevation gain values ​​associated with support point locations, the spread characteristics in one or two planes can be determined, and the spread gain can be derived from these spread characteristics in one or two planes. In the preferred case of determining both a set of azimuth gain values ​​and a set of elevation gain values ​​associated with support point locations, the spread gain can be readily obtained, for example, by multiplying a pair of elements (of the set) associated with the corresponding azimuth and elevation angles of the support points under consideration. Therefore, it is only necessary to evaluate a polynomial function for a set of support point locations with the same azimuth and different elevation angles and a set of support point locations with the same elevation and different azimuth angles, and then the spread gain can be derived (e.g., for all support points) by computationally simple multiplication of appropriate elements of the set of azimuth gain values ​​and the set of elevation gain values. Thus, high computational efficiency can be achieved.

[0137] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to combine (e.g., using multiplication) elements of the set of azi gain values ​​(e.g., aziGain(naz)) associated with the currently considered speaker or currently considered support point (e.g., specified by objNo and having associated values ​​naz and nel) and elements of the set of elevation gain values ​​(e.g., eleGain(nel)) associated with the currently considered speaker or currently considered support point to obtain extended gain values ​​(e.g., g_spd(objNo)) associated with multiple different speakers or multiple different support points (e.g., specified by different values ​​of objNo).

[0138] Therefore, by multiplying the corresponding elements of the set of azimuth gain values ​​with the corresponding elements of the set of elevation gain values ​​(e.g., elements associated with the azimuth and elevation of the respective support points), the extended gain values ​​associated with different support points can be obtained in a computationally efficient manner. Particularly high efficiency can be achieved if multiple support points include the same azimuth value, and if multiple support points include the same elevation value. In this case, the number of elements in the set of azimuth gain values ​​and the number of elements in the set of elevation gain values ​​can be kept reasonably small, and the elements of the set of azimuth gain values ​​and the set of elevation gain values ​​can be reused to determine the extended values ​​associated with multiple support points. In other words, this concept is particularly efficient when combined with uniformly spaced support points (in terms of azimuth and elevation).

[0139] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to calculate the following terms for a given object location (e.g., defined by the azimuth value azi and the elevation value ele) and for a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle):

[0140] - A first set of azimuth gain values ​​(e.g., aziGain) describes the contribution of multiple azimuth values ​​associated with the support point location or speaker azimuth index or support point azimuth index (e.g., naz) to the extended gain (e.g., using a polynomial function with a degree less than or equal to 3 or a parabolic function with a width suitable for the object's extended width in the azimuth direction). This first set of azimuth gain values ​​is associated with elevation values ​​within the original elevation value range (e.g., -90 degrees to +90 degrees, indicating that it does not cross the poles of the spherical coordinate system).

[0141] - A second set of azimuth gain values ​​(e.g., aziGainExtd) describes the contribution of multiple azimuth values ​​associated with the support point location or speaker azimuth index or support point azimuth index (e.g., naz) to the extended gain (e.g., using a polynomial function of degree less than or equal to 3 or a parabolic function with a width suitable for the object's extended width in the azimuth direction). This second set of azimuth gain values ​​is associated with elevation values ​​in an extended range of elevation values ​​indicating the poles of the spherical coordinate system (e.g., -180° to -90° and +90° to +180°).

[0142] Use this set of azi gain values ​​(aziGain(naz)) and / or use a set of elevation gain values ​​(eleGain(nel)) (or use a second set of azi gain values) to derive the extended gain.

[0143] By calculating two sets of azimuth gain values ​​(one for the original (or basic) elevation range and one for the extended elevation range), the extended gain of an audio object above (or below) the listener's head can be calculated in a particularly efficient manner. Specifically, it has been found that this set of azimuth gain values ​​can be used to derive the extended gain in an efficient manner (e.g., using predefined combination mappings). On the other hand, since elevation values ​​in the extended elevation range can be easily derived from those in the original elevation range using addition or subtraction (e.g., without distinguishing multiple cases and without changing the azimuth values), the elevation gain values ​​in a single set can also be calculated efficiently. For example, when extending an audio object above a user's head, the extension above the user's head can be easily calculated by using elevation values ​​greater than 90 degrees. Therefore, an object with a given azimuth angle (e.g., an azimuth angle within the range of -90 degrees to +90 degrees) and a positive elevation angle between 0 degrees and 90 degrees can be easily extended above the user's head using an elevation angle greater than 90 degrees while maintaining the azimuth angle (within the range of -90 degrees to +90 degrees). Thus, the azimuth gain value associated with the elevation angle values ​​within the extended elevation angle range (in this example, between +90 degrees and +180 degrees) can be obtained as an intermediate quantity and can later be mapped back to the support point or back to the coordinate system, for example, using elevation angles only within the original elevation angle range. The existence of a first set of azimuth gain values ​​and a second set of azimuth gain values ​​allows for efficient derivation of the extended values ​​because the second set of azimuth gain values ​​is suitable for combination with elevation gain values ​​within the extended elevation angle range.

[0144] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to calculate the following terms for a given object location (e.g., defined by the azimuth value azi and the elevation value ele) and for a given spread (e.g., defined by spreadAngleAzi or spreadAngleEle):

[0145] - A first set of elevation gain values ​​(e.g., eleGain) describes the contribution of multiple elevation values ​​associated with the support point location or speaker azimuth index or support point elevation index (e.g., nel) to the extended gain (e.g., using a polynomial function of degree less than or equal to 3 or a parabolic function with a width suitable for the object's extended width in the azimuth direction). This first set of elevation gain values ​​is associated with elevation values ​​within the original elevation value range (e.g., -90 degrees to +90 degrees) indicating that the poles of the spherical coordinate system are not crossed.

[0146] - A second set of elevation gain values ​​(e.g., eleGainExtd) describes the contribution of multiple elevation values ​​associated with the support point location or speaker elevation index or support point elevation index (e.g., nel) to the extended gain (e.g., using a polynomial function with a degree less than or equal to 3 or a parabolic function with a width suitable for the object's extended width in the azimuth direction). This second set of elevation gain values ​​is associated with elevation values ​​in an extended range of elevation values ​​indicating the poles of the spherical coordinate system (e.g., -180 degrees to -90 degrees and +90 degrees to +180 degrees).

[0147] Use the set of azi gain values ​​(aziGain(naz)) and the set of elevation gain values ​​(eleGain(nel)) (or use the first set of azi gain values, the second set of azi gain values, the first set of elevation gain values, and the second set of elevation gain values) to derive the extended gain.

[0148] This embodiment is based on similar considerations to the embodiment for calculating azimuth gain values. Specifically, the existence of a first set of elevation gain values ​​and a second set of elevation gain values ​​allows for simple and computationally efficient handling of cases where the object extends above the listener's head. Specifically, it has been found that using extended elevation values ​​(e.g., greater than +90 degrees) in this situation is much easier compared to immediate modifications of azimuth values ​​and "transformations" of elevation values.

[0149] As an example only, if the object's position includes an elevation angle of +80 degrees, it is easier to assume, for example, that the support point is at an elevation angle of 135 degrees for further calculations and for evaluating the parabolic function, since the elevation angle difference between the audio object and the support point can thus be stated as 55 degrees. In contrast, if a support point at 135 degrees is referenced as a support point at an elevation angle of 45 degrees, it becomes impossible to easily calculate the correct elevation angle difference between the audio object's position and the support point's position.

[0150] In summary, it has been found that using this extended elevation range is highly efficient because it avoids distinguishing a large number of cases in the calculation and facilitates the evaluation of polynomial functions. Furthermore, it has been found that, for example, the extended gain can be derived using a first set of azimuth gain values, a second set of azimuth gain values, a first set of elevation gain values, and a second set of elevation gain values. In this case, for example, entries for the first set of azimuth gain values ​​and entries for the first set of elevation gain values ​​can be combined, and entries for the second set of azimuth gain values ​​and entries for the second set of elevation gain values ​​can be combined efficiently to derive the extended value.

[0151] In a preferred embodiment of the audio object renderer mentioned above, the audio object renderer is configured to pre-calculate, during initialization (based on knowledge of the positions of the support points and the speakers), the support point translation gain (e.g., Spread.gainsSSP) used to translate the audio signals associated with the multiple support points to the multiple speakers.

[0152] The audio object renderer is configured to obtain the object-to-support spread gain (e.g., g_spd or spread gain values, such as elements of the vector g_spd) using a polynomial function of degree less than or equal to 3 (e.g., using a parabolic function). The object-to-support spread gain describes the contribution of the audio object signal to the signals of multiple support points.

[0153] The audio object renderer is configured to combine (e.g., multiply) objects to support point extended gain and support point translation gain in order to obtain extended object speaker gain.

[0154] It has been found that calculating the anchor point translation gain and the object-to-anchor point extension gain separately, and then combining the object-to-anchor point extension gain and the anchor point translation gain, yields particularly high computational efficiency, especially when there is more than one audio object. Since the anchor point is typically constant for processing multiple audio objects, the anchor point translation gain only needs to be calculated once, which can be done in the preparatory steps. In contrast, the object-to-anchor point extension gain typically depends on the object position and therefore needs to be calculated separately for each audio object.

[0155] Therefore, by reusing the anchor point translation gain for multiple objects, computational efficiency can be improved without degrading the achievable audio quality. Furthermore, using anchor points is particularly efficient because their spatial arrangement can be freely adjusted, focusing on computational efficiency without being constrained by the actual speaker location or setup. Thus, anchor points can be selected, for example, in a uniformly distributed manner (e.g., with uniform azimuth and elevation intervals), which significantly facilitates the determination of the object-to-anchor point extension gain. Therefore, using extended anchor points as intermediate extension targets can be said to actually contribute to improved efficiency because the first extension step can be performed independently of the actual speaker setup, and because the second extension step (from the extended anchor point to the speaker signal) only needs to be calculated once, even in the presence of multiple audio objects. Therefore, the process is extremely efficient.

[0156] In the preferred embodiment of the audio object renderer mentioned earlier, one or more polynomial functions of degree less than or equal to 3 are parabolic functions, which provide a return value p according to the following formula:

[0157] p = max(0, c1*anglediff 2 +c2),

[0158] Where c1 is the parameter that determines the width of the parabolic function, c2 is a predetermined value, angeldiff is the angular difference that the parabolic function is evaluated, and max(.,.) is the maximum value operator that returns the maximum value of its operands.

[0159] It has been found that such polynomial functions, restricted to non-negative values ​​(e.g., using maximum value operations), provide a good approximation of the desired extended properties and can be evaluated with minimal computational effort. Therefore, such polynomial functions have been found to be very good for determining extended values.

[0160] According to embodiments of the present invention, a method is created for determining speaker gain (e.g., combined speaker gain) based on object location information (e.g., azi, elevation angle) (which may be provided, for example, in spherical coordinates, such as using azi and elevation values ​​ele) and object feature information (e.g., extended angle information describing the extension in the azi direction (e.g., spreadAngleAzi) and / or extended angle information describing the extension in the elevation direction (e.g., spreadAngleEle)). The speaker gain describes the gain of including one or more audio object signals into multiple speaker signals.

[0161] This method involves taking into account object location information and object feature information to obtain extended object speaker gain (also known as extended speaker gain, or represented as vector gOS).

[0162] The method involves using one or more polynomial functions of degree less than or equal to 3 (e.g., parabolic functions or cubic polynomial functions (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1) to obtain the spread gain (e.g., object-to-support spread gain, such as g_spd), such as spread gain values ​​(e.g., elements of the vector g_spd), the spread gain describing the contribution of the audio object signal to multiple speaker signals or multiple support signals, and one or more polynomial functions mapping the angular difference between the object position and the support position (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) to the spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)).

[0163] The method includes obtaining the extended object speaker gain using an extended gain based on the extended gain contribution (e.g., g_spd), or using the extended gain (e.g., g_spd) as the extended object speaker gain.

[0164] This method is based on the same considerations as the corresponding device described above. Furthermore, this method may optionally be supplemented individually or in combination by any features, functions, and details described with respect to the corresponding device described above.

[0165] The embodiments create a method for determining speaker gain (e.g., combined speaker gain) based on object location information (e.g., azi, ele) (which may be provided, for example, in spherical coordinates, such as using azi and ele values) and extension information (e.g., extension angle information describing the extension in the azi direction (e.g., spreadAngleAzi) and / or extension angle information describing the extension in the elevation direction (e.g., spreadAngleEle)), wherein speaker gain describes the gain of including one or more audio object signals into multiple speaker signals.

[0166] This method involves taking into account object position information and extension information to obtain the extended object speaker gain (also known as extended speaker gain, or represented as a vector gOS).

[0167] The method involves using one or more polynomial functions of degree less than or equal to 3 (e.g., parabolic functions or cubic polynomial functions (e.g., parable*((diffCLKDir+(n-1)*Spread.openAngle).^2)+1 or parable*((diffAntiCLKDir+(n-1)*Spread.openAngle).^2)+1) to obtain the spread gain (e.g., object-to-support spread gain, such as g_spd), such as spread gain values ​​(e.g., elements of the vector g_spd), the spread gain describing the contribution of the audio object signal to multiple speaker signals or multiple support signals, and one or more polynomial functions mapping the angular difference between the object position and the support position (e.g., diffCLKDir+(n-1)*Spread.openAngle or diffAntiCLKDir+(n-1)*Spread.openAngle) to the spread gain value contribution (e.g., aziGain(naz) or eleGain(nel)).

[0168] The method includes obtaining the extended object speaker gain using an extended gain based on the extended gain contribution (e.g., g_spd), or using the extended gain (e.g., g_spd) as the extended object speaker gain.

[0169] This method is based on the same considerations as the corresponding device described above. Furthermore, this method may optionally be supplemented individually or in combination by any features, functions, and details described with respect to the corresponding device described above.

[0170] According to an embodiment of the present invention, a computer program is created for executing one of the aforementioned methods when the computer program is run on a computer.

[0171] This computer program can be supplemented individually and in combination by any of the features, functions and details described herein. Attached Figure Description

[0172] Embodiments of the invention will then be described with reference to the accompanying drawings, in which:

[0173] Figure 1 The diagram shows representations of different object extension configurations;

[0174] Figure 2 A signal flow diagram is shown for the object extension implementation of asymmetric and / or 2D loudspeaker setups;

[0175] Figure 3 The graphical representations of different extended gain functions are shown;

[0176] Figure 4 A graphical representation of the one-dimensional gain curves for different extension angles is shown;

[0177] Figure 5 A signal flow diagram for presenting extended gain is shown;

[0178] Figure 6 A graphical representation of an SSP grid with a resolution of 45 degrees is shown;

[0179] Figure 7 A comparison of the VBAP translation curve with a stretched and flipped parable shape is shown;

[0180] Figure 8 A MATLAB code example of the function spread_pannSSP is shown;

[0181] Figure 9 A MATLAB code example of the function spread_calculateGains is shown;

[0182] Figure 10 A MATLAB code example of the function calculateLayerGains is shown;

[0183] Figure 11 This shows a MATLAB code example for the function calculateSSPGains;

[0184] Figure 12 A schematic block diagram of an audio object renderer according to an embodiment of the present invention is shown;

[0185] Figure 13 A schematic block diagram of an audio object renderer according to an embodiment of the present invention is shown;

[0186] Figure 14 A flowchart of a method for determining loudspeaker gain according to an embodiment of the present invention is shown; and

[0187] Figure 15 A flowchart of a method for determining loudspeaker gain according to an embodiment of the present invention is shown. Detailed Implementation

[0188] 1. Discussion of some embodiments

[0189] 1. An audio object renderer according to Figure 12 2. The audio object renderer according to claim 1, wherein the audio object renderer is configured to perform the rendering of

[0190] Figure 12 A schematic block diagram of an audio object renderer 1200 according to an embodiment of the present invention is shown.

[0191] The audio object renderer 1200 is configured to receive object location information 1210 and object feature information 1212. Furthermore, the audio object renderer 1200 is configured to provide speaker gain 1214 (e.g., combined speaker gain or resulting speaker gain), which describes the gain used to include one or more audio object signals into multiple speaker signals.

[0192] Object position information 1210 may include, for example, object azimuth information (e.g., azi) and object elevation information (e.g., ele). For example, object position information may be provided in spherical coordinates, such as using azimuth values ​​and ele elevation values. Furthermore, object feature information or extension information 1212 may, for example, describe the characteristics of the audio object and indicate how the object should extend. Object feature information may, for example, be information indicating whether the object is small or extended (e.g., object size value). Object feature information may, for example, include object distance information, which may be mapped to extension values, such as mapping to extension angle information describing extension in the azimuth direction (e.g., spreadAngleAzi) and / or mapping to extension angle information describing extension in the elevation direction (e.g., spreadAngleEle). However, different types of object feature information are also possible. Alternatively, the audio object renderer may directly receive, for example, extension information describing the extension of the audio object in the azimuth direction and / or in the elevation direction.

[0193] The audio object renderer 1200 includes a translation object speaker gain determination 1230, which can be configured to use point-source translation of the audio object to obtain a translation object speaker gain 1232 (e.g., also referred to as "object speaker gain," or denoted by vector g). In point-source translation, the audio object can be treated as a point source, where extended information or object feature information 1212 is ignored, and the signal of the audio object is associated with two or more speakers in the environment of the object's location by appropriate selection of the translation object speaker gain 1232. In other words, the translation object speaker gain determination 1230 can, for example, perform a point-source translation of the audio object, treating the audio object as a point source, and can cause the audio object signal to be distributed (only) to those speakers closest to the audio object. However, no contribution of the audio object signal can be assigned to speakers farther from the audio object through point-source translation. However, the translation object speaker gain determination 1230 can use any point-source translation concept.

[0194] Furthermore, the audio object renderer 1200 may include an extended object speaker gain determination 1240, which takes into account object position information 1210 and object feature information or extension information 1212 to provide an extended object speaker gain 1242 (e.g., also referred to as extended speaker gain, and represented, for example, as a vector gOS). For example, the extended object speaker gain determination 1240 may determine a speaker gain that takes into account the extension of the audio object, for example, in the azimuth and elevation directions. Thus, the extended object speaker gain can be non-zero over a wide range because it is assumed that the audio object under consideration has a significant extension, and because it is generally assumed (and implemented) in the extension that the extended object speaker gain attenuates smoothly and steadily with increasing distance from the center position of the object.

[0195] In addition, the audio object renderer 1200 includes a combiner or combiner 1250 that combines the translation object speaker gain 1232 (e.g., g) and the extended object speaker gain (e.g., gOS) in a manner that always has a contribution from the translation object speaker gain (e.g., independent of object feature information or extended information 1212) in order to obtain a combined speaker gain.

[0196] Therefore, the audio object renderer 1200 can be said to provide a combined speaker gain 1214 based on both point source translation of the audio object signal and expansion of the audio object signal. For example, the contribution of the translation object speaker gain is always present in the combined speaker gain 1214, which ensures that the object can be reasonably localized even if the audio object is expanded over a relatively wide range.

[0197] Furthermore, it should be noted that the concepts described here can be implemented with particularly high computational efficiency, as will be described below.

[0198] In addition, it should be noted that Figure 12 The diagram also illustrates how the combined speaker gain 1214 can be used in further processing. For example, speaker gains 1214a, 1214b, and 1214c associated with different speakers in a speaker setup can be provided. A first speaker signal 1262a can be obtained by scaling the speaker gain 1214a associated with the first speaker as an audio signal 1260 associated with an audio object under consideration, and a second speaker signal 1262b can be obtained by scaling the audio object signal 1260 with the speaker gain 1214b associated with the second speaker, and a third speaker signal 1262c can be obtained by scaling the audio object signal 1260 with the third speaker gain 1214c, and so on. Speaker signals 1262a, 1262b, and 1262c can be naturally combined with speaker signals associated with other audio objects to obtain the actual speaker signals.

[0199] Therefore, the speaker signal can be obtained based on the combined speaker gain 1214, through which the audio object under consideration is represented in both point source translation and extended forms, which has been found to provide a particularly good auditory impression.

[0200] Furthermore, it should be noted that the audio object renderer 1200 may optionally be supplemented individually and in combination by any of the features, functions, and details described herein.

[0201] 1.B. According to Figure 13 Audio object renderer

[0202] Figure 13 A schematic block diagram of an audio object presenter 1300 according to an embodiment of the present invention is shown. The audio object presenter 1300 is configured to receive object location information 1310, which may correspond, for example, to object location information 1210. Furthermore, the audio object presenter 1300 is configured to receive object feature information or extended information 1312, which may correspond to object feature information or extended information 1212. Additionally, the audio object presenter 1300 provides (combined) speaker gain 1314, which may correspond to speaker gain 1214, and may be applied to the audio object signal in the same manner as speaker gains 1214a to 1214c. The audio object presenter is configured to take into account the object location information 1310 and the object feature information or extended information 1312 to obtain an extended object speaker gain (also referred to as extended speaker gain, or denoted as vector gOS).

[0203] The audio object renderer includes extended gain determination 1330, configured to obtain extended gain 1332, which may be, for example, an object-to-support point extended gain (e.g., g_spd). For example, extended gain determination may be configured to determine elements in the vector g_spd that describe the contribution of the audio object signal to multiple speaker signals or multiple support point signals. Specifically, extended gain determination 1330 may use one or more polynomial functions of degree less than or equal to 3 (e.g., one or more parabolic functions or cubic polynomial functions) to map one or more angular differences between the object position and one or more support point positions to one or more extended gain value contributions (e.g., mapped to "layer gain"), which may be represented, for example, by aziGain, aziGainExtd, eleGain, and eleGainExtd.

[0204] In other words, an extended gain contribution 1336 is obtained using mapping 1334, which employs a polynomial function of degree less than or equal to 3, wherein mapping 1334 maps one or more angular differences between the object position and one or more support point positions onto the extended gain contribution 1336. Furthermore, extended gain determination 1330 includes extended gain contribution processing 1338, which provides an extended gain 1332 (e.g., SPD) based on the extended gain contribution 1336. Additionally, the audio object renderer 1300 includes extended object speaker gain determination 1340, which obtains an extended object speaker gain 1340 based on the extended gain 1332, wherein the extended gain 1332 is based on the extended gain contribution 1336.

[0205] In other words, the extended gain contribution is derived using a polynomial function of degree less than or equal to 3, and then the extended gain contribution 1336 is mapped to an extended gain 1332, which may, for example, describe how the audio object signal should be distributed across multiple support points. The extended object speaker gain determination 1340 can map the extended gain (which is related to the extended support points) to the extended object speaker gain, which corresponds to the mapping from the support point location to the actual speaker (and the actual speaker location).

[0206] To summarize further, it should be noted that the derivation of the extended object loudspeaker gain uses a polynomial function (which can be evaluated with moderate computational complexity) rather than an exponential or trigonometric function that requires a relatively large amount of computation.

[0207] Therefore, the extended gain contribution of 1336 allows for the acquisition of extended gain 1332 in a very efficient manner. Furthermore, it has been found that no significant loss of auditory performance is achieved by using a polynomial function of degree less than or equal to 3.

[0208] However, it should be noted that the audio object renderer 1300 may optionally be supplemented individually and in combination by any of the features, functions and details described herein.

[0209] 1.C. According to Figure 2 Speaker gain calculation

[0210] Figure 2 A signal flow diagram is shown for an object extension implementation of asymmetric and / or 2D speaker setups.

[0211] Figure 2 The signal flow shown can, for example, be based on Figure 12 and Figure 13 Implemented in the audio object renderers 1200 and 1300.

[0212] In addition, refer to Figure 2 The concepts described may also be included, optionally individually and in combination. Figure 5 In the concept.

[0213] according to Figure 2 The speaker gain determination 200 receives object position information 210, which may correspond, for example, to object position information 1210, 1310. The object position information may describe the object position, for example, in terms of azimuth information (e.g., azi) and elevation information (e.g., ele). The speaker gain determination 200 also receives extension angle information 212, which may correspond, for example, to object feature information or extension information 1212, 1312. The extension angle information 212 may describe the extension angle, for example, in terms of azimuth and / or elevation information, or in terms of width and / or height information. Furthermore, the speaker gain determination 200 may, for example, provide a resulting speaker gain 214, which may correspond, for example, to speaker gains 1214, 1314.

[0214] The loudspeaker gain determination 200 includes grid creation 204 for extended support locations, which provides information 204a describing the location of the extended support points. Furthermore, the loudspeaker gain determination 200 may include translation 205 of the extended support points, which may, for example, provide an extended support point translation gain 205a.

[0215] Support point location information 204a may, for example, describe the locations of extended support points, which may, for example, be uniformly distributed on the sphere. For example, the locations of the extended support points described by information 204a may be selected independently of the actual speaker locations, and may, for example, form a grid defined by uniform azimuth and elevation intervals.

[0216] The extended support point translation gain described in information 205a can, for example, define how the audio object signal associated with the extended support point location is distributed to the actual speaker or speaker signal. Therefore, the extended support point translation gain described in information 205a can, for example, describe the distribution of the audio object signal associated with the extended support point to the actual speaker or speaker signal for all extended support points.

[0217] It should be noted that the grid creation 204 and translation 205 can be calculated only once, for example, and can be reused to extend multiple audio objects, since the extension support points usually remain unchanged for the extension of multiple audio objects.

[0218] The speaker gain determination 200 also includes an extended gain calculation 201, which takes into account the extended support point position 204a, object position information 210, and extended angle information 212. Based on this, the extended gain calculation 201 provides an extended gain 201a, which can, for example, describe the extension of the audio object signal to multiple extended support points.

[0219] The loudspeaker gain determination 200 also includes a combination 206 that combines the extended gain 201a with the extended support point translation gain 205a to obtain an extended loudspeaker gain 206a (e.g., gOS). For example, combination 206 can use a multiplication of the extended gain 201a and the extended support point translation gain 205a. For example, combination 206 can combine the mapping of the audio object signal to the extended support point (described by the extended gain 201a) with the mapping of the signal at (or associated with) the extended support point to the actual loudspeaker signal (described by the extended support point translation gain 205a), and thus provide an extended loudspeaker gain 206a such that the extended loudspeaker gain 206a describes the contribution of the audio object signal of the currently considered audio object to the actual loudspeaker signal (in the form of a weighted value to be applied to the audio object signal, thereby deriving the actual loudspeaker signal). However, it should be noted that the extended gain calculation 201 and combination 206 are typically performed separately for each audio object to be extended (while the extended support point position 204a and extended support point translation gain 205a can be reused without change).

[0220] The speaker gain determination 200 also includes object translation 202, which derives the object speaker gain or translates the object speaker gain 202a using the object position information 210. Object translation 202 can, for example, perform a point-source translation of the audio object, where it can be determined which actual speaker signal contributions of the audio object signal of the audio object under consideration are taken into account. Typically, when using point-source translation, only the object speaker gain (or translated object speaker gain) for speakers directly adjacent to the object position is non-zero.

[0221] The loudspeaker gain determination 200 also includes a combination 203, wherein the extended loudspeaker gain 206a and the translated target loudspeaker gain 202a are combined to obtain a resulting loudspeaker gain 214 (which may be specified, for example, by g). This combination may, for example, include a summation or weighted combination of the extended loudspeaker gain 206a and the translated target loudspeaker gain 202a.

[0222] This concept allows for the provision of a resulting speaker gain 214, which includes both the translational object speaker gain and the extended speaker gain 206a. It has been found that this concept can be computed with high computational efficiency and provides good quality audio perception.

[0223] It should be noted that, regarding Figure 2 The concept of speaker gain determination described herein may optionally be supplemented by any features, functions, and details described herein, individually and in combination.

[0224] 1.D. Extended Functions

[0225] The following sections will provide some details about possible extension functions that can be used in any audio object renderer described herein, as well as any audio object rendering concepts described herein.

[0226] For example, Figure 3 Graphical representations of different spread functions are shown. It should be noted that spread functions are typically defined relative to the object's position. Therefore, in graphical representations 310, 320, and 330, the first axes 312a, 322a, and 332a describe the azimuth difference between the currently considered spread support position and the object's azimuth position. The second axes 312b, 322b, and 332b describe the elevation difference between the currently considered spread support point position and the object's elevation position.

[0227] Graphic representation 310 shows horizontal extension (where the extension in the azimuth direction is greater than the extension in the elevation direction). Graphic representation 320 shows uniform extension (where the extension in the azimuth direction is equal to the extension in the elevation direction), and graphic representation 330 shows vertical extension (where the extension in the azimuth direction is less than the extension in the elevation direction).

[0228] These graphs illustrate the (relative) gain, which should be used to scale the audio object signal (e.g., to obtain the signal at the extended support point) based on the azimuth difference between the extended support point location and the azimuth object location, and based on the elevation difference between the extended support point location and the object elevation location. For example, the gain can be obtained by multiplying the corresponding azimuth extension functions 314a, 324a, 334a with the corresponding elevation extension functions 314b, 324b, 334b. It should be noted that this document describes in detail the concept of computationally efficient implementation of the extension function (or extended gain function), which is related to... Figure 3 Their extended gain functions are very similar. Furthermore, it should be noted that... Figure 3 The extended gain function can be used, for example, or substantially used in the audio object renderers 1200, 1300 described herein, or Figure 2 and Figure 5 The concept of audio object presentation.

[0229] In addition, it should be noted that other (optional) details regarding the extended gain function will also be discussed in this paper.

[0230] also, Figure 4 Another graphical representation of the one-dimensional gain curves for different spread angles is shown. It can be seen that the smaller the spread angle, the steeper the curve. It should be noted that... Figure 4In the diagram, the horizontal axis 412 describes the difference angle between the extended support point and the object (e.g., difference azimuth or difference elevation). The vertical axis 414 describes the gain applied to derive the extended support point signal from the audio object signal based on the difference angle (e.g., under the assumption of one-dimensional extension). Curves 416a, 416b, 416c, and 416d, expressing the gain as a function of the difference angle, are associated with different extension angles. For example, curve 416a is associated with a relatively narrow extension angle, and curve 416d is associated with a relatively wide extension angle.

[0231] As an example, for a 50-degree difference angle (azimuth or elevation) between the extended support point (SSP) and the object, the gain value is determined to be 0.37 from the gain curve 416c, which represents a fairly strong extension value.

[0232] In other words, Figure 4 The gain curves shown, or their approximations, can be used, for example, to derive extended gain 201a, extended target speaker gain 1242, or extended gain 1332. Further details regarding the gain curves and their possible approximations are also described herein. Furthermore, a more detailed discussion will be held regarding the... Figure 4 The use of gain curves.

[0233] 1.E According to Figure 5 Determining the extended gain

[0234] Figure 5 A signal flow diagram for presenting extended gain according to an embodiment of the present invention is shown. The azimuth processing path is covered by blue lines or lines having a first shadow type and a second shadow type, and the elevation angle (or more precisely, the elevation processing path) is covered by orange lines or lines having a third shadow type and a fourth shadow type.

[0235] The extended gain determination 500 receives object position information as input, including azimuth object position information 510a and elevation object position information 510b. The azimuth object position information 510a and elevation object position information 510b may, for example, correspond to the object position information 210 described above. Furthermore, the extended gain determination 500 also receives extended support point position information, including azimuth extended support point position information 513a and elevation extended support point position information 513b. The azimuth SSP position information 513a and elevation SSP position information 513b may, for example, correspond to the SSP position information 204a described above.

[0236] The extended gain determination 500 includes azimuth gain determination 530, which may include, for example, azimuth difference angle calculation 301 and azimuth gain function application 302. Furthermore, the extended gain determination 500 also includes elevation gain determination 540, which may include, for example, elevation difference angle calculation 304 and elevation gain function application 305. For example, in azimuth gain determination 530, one or more azimuth gain values ​​(e.g., aziGain) can be determined based on azimuth object position information 510a and azimuth SSP position information 513a. Furthermore, azimuth extended information 512a or its preprocessed version may also be considered in azimuth gain determination 530. For example, in azimuth difference angle calculation 301, the difference between the azimuth object position and one or more azimuth SSP positions can be calculated to obtain one or more angular differences, and a gain value can be determined for these one or more angular differences in azimuth gain function application 302. For example, the azimuth gain function application can evaluate a gain function for one or more angular differences (e.g., such as...). Figure 3 and Figure 4 The gain function shown herein, or a polynomial or parabolic gain function as disclosed herein. Therefore, one or more azimuth gain values ​​are obtained, which are associated with the extended support point position and determined by the value of the azimuth gain function for the corresponding angular difference between the corresponding extended support point position (azimuth position) and the corresponding object position (azimuth position), wherein the width of the azimuth gain function is adjusted taking into account azimuth extension information 502.

[0237] The elevation gain determination 540 performs a similar calculation. The elevation gain determination 540 receives elevation extension support point position information 513b and elevation object position information 510b, and provides one or more elevation gain values ​​305a based on these. For example, one or more differences between the object position elevation angle and the elevation angles of one or more extension support point positions can be calculated using difference angle calculation 304. An elevation gain function whose width can be determined by the elevation extension information 512b or its preprocessed version (e.g., it can be evaluated for one or more difference angles determined by difference angle calculation 304) can be applied to obtain one or more elevation gain values ​​305a. For example, refer to... Figure 3 Or refer to Figure 4 The described elevation gain function or its approximation (e.g., the polynomial gain function disclosed herein) can be used in elevation gain function application 305. The gain function can be evaluated, for example, for one or more difference angles determined in difference angle calculation 304 to obtain one or more elevation gain values ​​305a.

[0238] For example, an azimuth gain value 302a (which may be obtained, for example, for a given elevation angle and for azimuth angles of multiple extended support points) may be combined with multiple elevation gain values ​​305a (which may be obtained, for example, for a given azimuth angle value and for elevation angle values ​​of multiple extended support points) (e.g., by multiplication). This combination is specified by 313 and may be multiplicative, wherein different azimuth gain values ​​302a and elevation gain values ​​305a associated with different extended support points may be multiplied to obtain contributions to the gain values ​​associated with the different extended support points.

[0239] The extended gain determination 500 may also optionally include processing for extending the elevation range. For example, the extended gain determination may (optionally) include elevation range extension 307, which may, for example, adapt (or preprocess) the azimuth object position information 510a and / or azimuth SSP position information 513a and / or elevation SSP position information 513b and / or elevation object position information 510b for use in the extended elevation range calculation.

[0240] The calculation of the extended elevation range includes, for example, providing one or more extended azimuth gain values ​​309a. The extended azimuth gain value 309a may correspond, for example, to the azimuth gain value 302a, but may have a modified association with the azimuth. In other words, the extended azimuth gain value 309a (or a set of extended azimuth gain values) may, for example, be an angularly shifted version of the azimuth gain value 302a (or more precisely, a set of azimuth gain values). However, the extended azimuth gain value 309a may, for example, be derived from the azimuth gain value 302a, or may be obtained using the azimuth difference angle calculation 308 and the azimuth gain function application 309.

[0241] Similarly, the elevation range extension processing includes providing an extended elevation gain value 311a based on elevation object position information 510b and elevation SSP position information 513b. For example, elevation difference angle calculation 310 can calculate the angular difference between the elevation object position and the elevation extension support point position within the extended elevation range (e.g., between +90 degrees and +180 degrees or between -90 degrees and -180 degrees). Therefore, elevation gain function application 311 can apply a gain function to the angular difference determined in elevation difference angle calculation 310 to obtain an extended elevation gain value 311a (or more precisely, a set of extended elevation gain values), which is associated, for example, with the elevation SSP position within the extended elevation range. However, it should be noted that the extended elevation gain value 311a can also optionally be determined based on the elevation gain value 305a (e.g., using an appropriate mapping (or reordering)).

[0242] Furthermore, one or more extended azimuth gain values ​​309a can be combined (e.g., by multiplication) with one or more corresponding extended elevation gain values ​​311 to obtain contribution 312a to obtain a value 314a associated with the extended support point. For example, contributions 313a and 312a associated with the same extended support point can be summed in summation 314 to obtain a gain value 314a associated with the extended support point. Additionally, normalization 315 can optionally be applied to the gain value 314 to obtain an extended gain value 514. The extended gain value 514 can, for example, correspond to the extended gain 1332 described above. For example, normalization 315 can take into account the extended width and can help avoid changes in signal energy due to the extended gain.

[0243] Regarding the overall function of determining the extended gain 500, it should be noted that azimuth and elevation gain values ​​can be determined individually for multiple extended support point locations with different azimuth values ​​and multiple extended support point locations with different elevation values. Then, the gain values ​​for a larger number of extended support points are obtained by combining the azimuth and elevation gain values. Standard azimuth and extended azimuth gain values ​​can also be used to reflect the extension above or below the user's head. Combining the azimuth, elevation, extended azimuth, and extended elevation gain values ​​associated with the same extended support point efficiently yields the gain value 214 associated with the corresponding extended support point. This combination can be performed for different extended support points (or even for all extended support points, or for all extended support points except those located at poles).

[0244] Therefore, by using extended elevation range processing (e.g., including boxes 307, 308, 309, 310, and 311), overhead extension or under-listener extension can be easily implemented by allowing elevation angles greater than 90 degrees and / or less than -90 degrees within the extended elevation range. Computational efficiency can be improved by providing pairs of extended azimuth gain values ​​and extended elevation gain values ​​associated with the extension support point, since the extended azimuth gain values ​​and extended elevation gain value pairs can be combined (e.g., multiplied) to obtain an extension gain 314a (or a contribution 312a to the extension gain 314a).

[0245] In summary, the extended gain determination 500 is extremely efficient and allows for the determination of extended gain even when the audio object is above or below the listener's head.

[0246] Furthermore, it should be noted that the extended gain determination 500 may optionally be supplemented individually and in combination with any features, functions, and details disclosed herein.

[0247] 1.F According to Figure 6 Extended support point width

[0248] Figure 6 A representation of an example extended support point grid with a resolution of 45 degrees is shown. (Example:) Figure 6 As can be seen, extended support points 610a, 612a, 612a, 612b, 612c, 612d, 612e, 612f, 612g, 614b, 614c, 614d, 614e, 614f, 616c, 616d, and 616e are arranged on the sphere. Specifically, extended support points 612a to 612g, 614b to 614f, and 616c to 616e are arranged on a circle with a constant elevation angle. Extended support point 610a is located at the pole of the spherical coordinate system (e.g., at an elevation angle of +90 degrees). Furthermore, it should be noted that extended support points 612b and 614b are located on a semicircle with a constant azimuth angle. Similarly, extended support points 612c, 614c, and 616c are located on a semicircle with a constant azimuth angle.

[0249] Generally, extended support points are preferably arranged on a grid defined by a circle with a constant elevation angle and a semicircle or circle with a constant azimuth angle. Therefore, there are usually multiple extended support points with the same elevation angle, and there are also usually multiple extended support points with the same azimuth angle.

[0250] For further (optional) details, please also refer to the additional explanation regarding the location of the extended support points provided in this article.

[0251] However, it should be noted that the 45-degree resolution should only be considered as an example, and different resolutions can be chosen. For example, the resolution in the azimuth direction can naturally be different from the resolution in the elevation direction.

[0252] 1. G polynomial gain function

[0253] As described above, it is advantageous to use a polynomial gain function because such a polynomial gain function (e.g., including powers of 2 or 3) can be evaluated with low computational complexity.

[0254] Figure 7 This illustrates a shape comparison between the VBAP translation curve and a stretched and flipped analogue. For example, the horizontal axis 712 describes the angle difference (e.g., elevation difference angle or azimuth difference angle) between the currently considered extended support point and the object. The vertical axis 714 describes the gain or normalized gain. The first curve 720 describes the gain that will be obtained using VBAP. The second curve 730 describes the gain that can be obtained through the (stretched and flipped) analogue, whose values ​​are limited to non-negative. As can be seen, the analogue-based gain function is a very good approximation of the VBAP gain function.

[0255] Therefore, the analog gain function can be used in any apparatus and method described herein to map angular differences to gain values. In other words, the analog-based gain function can be used, for example, in extended gain calculation 201 or extended object loudspeaker gain determination 1240 or mapping 1334. For example, the analog-based gain function can replace (or approximate) Figure 3 The extended gain function shown and Figure 4 The gain curve shown. Furthermore, Figure 7 The analogy-based gain function shown can also be used in boxes 302, 309, 311, and 305 of the extended gain determination 500.

[0256] However, it should be noted that the analog can be naturally adapted to the expansion (e.g., based on azimuth or elevation expansion). Furthermore, the analog can be naturally scaled to suit the specific needs of the application, where the center value of the analog can be changed, and / or the width of the analog can be changed.

[0257] 1.H According to Figures 8 to 11 Implementation methods

[0258] Figures 8 to 11 This document presents a MATLAB code example illustrating the concept and method for determining speaker gain, which describes the gain of incorporating one or more audio object signals into multiple speaker signals.

[0259] It should be noted that, as a reference Figures 8 to 11 The concepts outlined herein, or parts thereof or details thereof, may optionally be used in any of the embodiments described herein.

[0260] This method includes initialization performed via the initialization function "spread_pannSSP". This initialization function uses a configuration structure containing VBAP parameters as input. The initialization function also receives information about the number of speakers. However, it should be noted that the initialization function does not necessarily need to use these input parameters.

[0261] However, initialization typically involves selecting (or defining) extended support points. For example, extended support points can be defined by a grid of azimuth and elevation angles in a spherical coordinate system (where, for example, all extended support points can have equal radii). For instance, the azimuth angle of an extended support point can be defined in an array `aziSSP`, and the elevation angle can be defined in an array `eleSSP`. The definition of an extended support point is shown at reference numeral 810. The definition of the extended support point shown at reference numeral 810 can, for example, correspond to grid creation 204.

[0262] Initialization also includes the translation of the extended support points, as shown at reference numeral 820 in the attached figure. The translation of the extended support points shown at reference numeral 820 can, for example, be related to... Figure 2The translation of the extended support point is shown at reference numeral 205 in the attached figure. For example, for each extended support point, the translation of the audio signal to be presented at the location of the corresponding extended support point onto the actual speaker signal is determined. In other words, for each extended support point, a scaling value is determined, which describes the translation of the signal to be presented at the location of the corresponding extended support point onto the actual speaker signal. These gain values ​​are stored in a data structure named "Spread.gainsSSP".

[0263] Special treatment can be applied to extended support points located at poles in a spherical coordinate system (e.g., at elevation angles of + / - 90 degrees). This is shown at reference numeral 830. However, it should be noted that specific details regarding these translation gain values ​​(used to translate the audio object at the extended support point location onto the speaker signal) are not particularly relevant to this invention. In the given example, the function vbap is used, but other functions (e.g., other translation functions) may also be used.

[0264] Initialization 800 also includes the initialization of some variables (or constants) used in further processing. This initialization is shown at reference numeral 840.

[0265] However, it should be noted that the details of initializing 800 should be considered optional.

[0266] The function calls, typically executed multiple times for different objects, will be described below. The main function is called "spread_calculateGains". This main function may, for example, receive a gain g provided by object translation (e.g., object translation 202) or by the speaker gain determination 1230 of the translated object, and provide an extended gain (also specified by g) based on the gain g, which may correspond to "derived speaker gain" 214 or speaker gains 1214, 1314. In addition, the main function receives object azi, object elevation information ele, object extended width spdAzi (or spreadAngleAzi), object extended height spdEle (or spreadAngleEle), and a data structure including the extended parameters (which may be provided, for example, by the initialization described above).

[0267] The main function 900 may include, for example, the determination of attenuation gain, as shown at reference numeral 910. For example, the attenuation gain attenGain can be determined based on the object extension width and object extension height, and also based on the extension grid resolution. For example, the calculation rules shown at reference numeral 910 can be used. However, in general, the attenuation gain can increase with the maximum object extension and also with the minimum object extension. Therefore, if the object extension is relatively large, the extended object speaker gain will be weighted relatively strongly (relative to the translation object speaker gain), while if the extension is relatively small, the extended object speaker gain will be weighted relatively weakly.

[0268] In a further preprocessing step 920, the object expansion width and / or object expansion height will be adjusted to ensure that a minimum object expansion width and / or minimum object expansion height are used in further processing. Specifically, if the smaller of the object expansion width and object expansion height is less than a minimum value, the smaller value will be adjusted to achieve the corresponding minimum value.

[0269] In a further preprocessing step shown at reference numeral 930, parameters for determining the gain value are calculated, for example, based on the corresponding extension angle.

[0270] Furthermore, in preparation step 940, the loop limit values ​​aziLoopLim and eleLoopLim are calculated, which determine the number of computational steps performed in the layer gain calculation. Computational complexity can be improved in some cases by (optionally) limiting the number of computational steps performed in the layer gain calculation.

[0271] Furthermore, the main function 900 also includes the determination of the azimuth layer gain, as shown at reference numeral 950. In the first sub-step 951, the azimuth layer gain array is determined by calling the function "calculateLayerGains". The azimuth layer gain calculated in step 951 is stored in the array aziGain. In the further sub-step 952, an angularly shifted version of the azimuth layer gain is obtained and stored in the array aziGainExtd. In other words, the entries of the array aziGain are copied into the array aziGainExtd in the order of modification.

[0272] It should be noted that step 951 can, for example, be related to... Figure 3 The functions 301 and 302 shown correspond to those shown. Furthermore, it should be noted that step 952 can be associated with… Figure 3 The functions indicated by reference numerals 308 and 309 in the accompanying drawings correspond to those functions. In other words, instead of performing... Figure 3Functions 301 and 302 shown can be used with the functions shown in reference numeral 951, or vice versa. Furthermore, instead of performing... Figure 3 The functions shown at reference numerals 308 and 309 in the attached figures can be used with the functions shown at reference numeral 952, or vice versa. For example, the arrays aziGain and aziGainExtd can represent layer gains associated with different azimuth angles (shifted by 180 degrees relative to each other). For example, the first element of the array aziGain can correspond to the azimuth angle ϕ1, and the first element of the array aziGainExtd can correspond to the azimuth angle ϕ1+180 degrees.

[0273] The main function 900 may also include the determination of the elevation layer gain, as shown at reference numeral 960. For this purpose, the function "calculateLayerGains" can be used (again), which returns an intermediate array of values, as shown at reference numeral 961. Based on the intermediate array of values ​​eleGainTMP, the elevation layer gain array eleGain can be determined by appropriately selecting the entries in the intermediate array of values, as shown at reference numeral 962. Similarly, the elevation layer gain extension array can be determined by appropriately selecting and sorting the entries in the intermediate array, as shown at reference numeral 963.

[0274] The functions shown at reference numerals 961 and 962 may correspond, for example, to the functions shown at boxes 304 and 305, and the functions shown at boxes 961 and 963 may correspond, for example, to the functions shown at boxes 310 and 311.

[0275] In other words, the functions shown at reference numerals 961 and 962 can replace function blocks 304 and 305, and the functions shown at reference numerals 961 and 963 can replace the functions shown at blocks 310 and 311. However, alternatively, the functions of blocks 304 and 305 and the functions of blocks 310 and 311 can be performed instead of function 960.

[0276] In a further step shown at reference numeral 970, the main function 900 calculates the extended support point extended gain based on the previously calculated azimuth layer gain and the previously calculated elevation layer gain. For this purpose, the function "calculateSSPGains", which will be described later, is invoked.

[0277] Therefore, the extended support point extension gain, specified by the array g_spd, is obtained, which describes which scaling should be used to render the audio object at the extended support point. However, since it is desirable to know which scaling should be used to render the audio object signal in the speaker signal, the extended support point extension gain is mapped to the speaker gain in step 980. This can be understood as a translation of the audio signal to be rendered at the location of the extended support point to the actual speaker signal (associated with the speaker at the actual speaker location, which is typically different from the extended support point).

[0278] For this purpose, the results of the previously performed extended support point translation (executed in initialization 800) are utilized. For example, the product of multiple sets of extended support point translation gains and extended support point extension gains is summed over all extended support points. In other words, the extended support point translation gain (or a set of extended support point translation gains) is associated with each extended support point (referenced by the moving variable obj), and the extended support point extension gain is also associated with each extended support point.

[0279] It should be noted that step 980 may correspond, for example, to the function shown at reference numeral 206. Therefore, the function of block 206 may be replaced by the function shown at reference numeral 980, and vice versa.

[0280] In step 990, the obtained (extended object) speaker gain gOS is combined with the input gain value g (which may be, for example, a translation object gain value). For example, scaling of the extended gain value gOS is determined by the attenuation gain attenGain described above. Furthermore, step 990 may optionally include normalizing the result of the combination of the translation gain value and the extended gain value.

[0281] For example, step 990 can correspond to the function of box 203.

[0282] Regarding the overall function of main function 900, it should be noted that the main functions are in steps 950, 960, 970, and 980. In step 950, the "layered gain value" array is calculated, which describes the extension of the audio object in the azimuth direction, more precisely, the gain value associated with the azimuth value of the extension support point. In this step, object extension in the azimuth direction is considered. Furthermore, the extended azimuth gain value, cyclically shifted relative to the azimuth gain value in the array aziGain, helps to form an "overhead" extension.

[0283] In step 960, an extension value is calculated, which is associated with a given azimuth value and different elevation values ​​associated with the extension support point. Here, the elevation position of the audio object and the object extension in the elevation direction are considered. Furthermore, an elevation layer gain extension array is obtained to support over-the-top extension of the audio object.

[0284] In step 970, the values ​​of the azimuth layer gain and the elevation layer gain are combined to calculate the gain value associated with all support point locations.

[0285] Therefore, in step 980, the gain value associated with the support point position is effectively mapped to the gain value associated with the speaker signal.

[0286] In the following text, reference will be made to Figure 10 and Figure 11 Describe some details of the functions "calculate layer gains" and "calculateSSPGains".

[0287] Figure 10 The MATLAB code for the function `calculateLayerGains` is shown. It should be noted that the function, specified by "gains", returns an array where the indices of the array elements are associated with either azimuth or elevation. Generally, the array entries comprise a roughly parabolic decay as the angular difference between the angular position (azimuth or elevation) of the audio object and the angle associated with the corresponding entry in the array (e.g., SSP azimuth or SSP elevation) increases.

[0288] Function 1000 includes the optional determination of the symbol value plumin, as shown at reference numeral 1010 in the figure.

[0289] Furthermore, function 1000 includes determining an array of indices (“multiple”) associated with the object’s position (e.g., the object’s elevation angle or the object’s azimuth angle). This determination is shown in reference numeral 1020.

[0290] Function 1000 also includes the calculation of the object position derived from the position (angle) of the adjacent extended support points, as shown in reference numeral 1030.

[0291] Furthermore, the function includes the calculation of gain values, as indicated by reference numeral 1040 in the attached figure. These gain values ​​are stored in an array “gains”, where array indices are associated with the angles (azimuth or elevation) of the extended support points. The gain values ​​themselves are determined using an evaluation of an analogy relating the position of the audio object under consideration to the corresponding angular difference between the position of the extended support points and the position of the corresponding extended support points. The gain values ​​are provided by the evaluation of the analogy, which is constrained so that the values ​​remain non-negative. Therefore, the array “gains” is populated with gain values ​​based on an evaluation of an analogy centered on the angle (azimuth or elevation) of the position of the audio object under consideration (where the constraint on non-negative values ​​is applied).

[0292] Therefore, the function “calculate layer gains” shown in reference numeral 1000 allows for the provision of an array of gain values, more precisely, extended gain values ​​associated with a constant azimuth angle (or alternatively, a constant elevation angle).

[0293] The details of the function "calculateSSPGains" will be described below.

[0294] The function shown at reference numeral 1100 includes the calculation of the extended gain, as shown at reference numeral 1110. Specifically, an extended gain value is calculated for each extended support point SSP, with specific treatment of extended support points located at the poles of the polar coordinate system shown at reference numeral 1120.

[0295] However, for each extended support point specified by the elevation index nel and the azimuth index naz, the azimuth gain value aziGain(naz) and the elevation gain value eleGain(nel) are multiplied together. This combination can correspond, for example, to the operation shown at box 313. Furthermore, the associated extended azimuth gain value aziGainExtd(naz) is also multiplied together with the associated extended elevation gain value eleGainExtd(nel), which can correspond to the operation shown at box 312.

[0296] Furthermore, the results of the multiplication combination are then added together, which corresponds to the operation shown in box 314. For example, the azimuth gain value and the extended azimuth gain value specified by the same index naz can correspond to angles that differ by 180 degrees. For example, when the azimuth gain value specified by a given index naz can be associated with an azimuth of +45 degrees, the extended azimuth gain value associated with the same index naz can be associated with an azimuth of -135 degrees. Additionally, the sum of the angles associated with the same index nel in the elevation gain array and the extended elevation gain array can be 180 degrees or -180 degrees. For example, a given index nel can specify an entry in the eleGains array associated with +45 degrees, and the same index nel will specify an entry in the eleGainExtd array associated with an angle of 135 degrees. Therefore, in this example, the sum of the angles associated with the given index nel in the eleGain and eleGainExtd arrays can be +180 degrees. When using this combination, it is possible to ensure that appropriate scaling values ​​are obtained with a moderate amount of work. Furthermore, the fact that the elevation angle is greater than 90 degrees does not excessively increase the computational complexity of this concept.

[0297] Specific processing of poles (i.e., the gain values ​​associated with poles) 1120 helps to avoid artifacts at poles.

[0298] Function 1000 also includes normalization, which is shown in reference numeral 1130 and can be considered optional. Therefore, the extended gain can be optionally normalized to bring it into the appropriate range of values.

[0299] In summary, function 1100 allows the derivation of gain values ​​associated with extended support points based on an array of gain values ​​associated with a single elevation angle and an array of gain values ​​associated with a single azimuth angle.

[0300] It should be noted that the functionality of functions 800, 900, 1000, and 1100 may optionally be incorporated into any other embodiment, either individually or in combination. It should also be noted that any functionality described in other embodiments may optionally be incorporated into functions 800, 900, 1000, and 1100, either individually or in combination.

[0301] 1I. According to Figure 14 Method

[0302] Figure 14 A flowchart is shown for a method 1400 for determining speaker gain based on object location information and object feature information or extended information, wherein speaker gain describes the gain used to include one or more audio object signals into multiple speaker signals.

[0303] This method involves using point source translation of an audio object to obtain a 1410 translation object speaker gain.

[0304] The method also includes taking into account object location information and object feature information or extended information to obtain 1420 extended object speaker gain.

[0305] The method also includes combining the 1430 translation object speaker gain and the extended object speaker gain in a manner that always contributes to the translation object speaker gain, in order to obtain a combined speaker gain.

[0306] Method 1400 is based on the same considerations as the apparatus described above, and may optionally be supplemented individually and in combination by any features, functions and details described herein.

[0307] 1J. According to Figure 15 Method

[0308] Figure 15 A flowchart is shown for a method 1500 for determining speaker gain based on object location information and object feature information or extended information, wherein speaker gain describes the gain used to include one or more audio object signals into multiple speaker signals.

[0309] This method involves taking into account object location information and object feature information to obtain 1510 extended object speaker gain.

[0310] The method also includes using one or more polynomial functions of degree less than or equal to 3 to obtain an extended gain of 1520, which maps the angular difference between the object position and the support point position to the extended gain value contribution.

[0311] The method also includes using the extended gain based on the extended gain contribution to obtain the 1530 extended object speaker gain, or using the extended gain as the extended object speaker gain.

[0312] Method 1500 is based on the same considerations as the apparatus described above, and may optionally be supplemented individually and in combination by any features, functions and details described herein.

[0313] 2. Discussion of other embodiments

[0314] The object expansion rendering algorithm according to an embodiment of the present invention will be described below.

[0315] It should be noted that this object extension rendering algorithm can be used independently, but can optionally be supplemented individually and in combination by any features, functions and details disclosed herein.

[0316] Furthermore, any features, functions, and details of the concepts described in this section may optionally be incorporated, individually and in combination, into any apparatus and method described herein.

[0317] According to one aspect, the basic idea for achieving the extended effect is to activate other speakers that reproduce the same object signal from the object's location with monotonically decreasing intensity. That is, for each speaker, an extension gain must be calculated to reproduce the object in order to create the extended effect. The extension gain can be determined as follows.

[0318] 2.1 Object Extension Computation Using Additional Extended Support Points

[0319] In asymmetrical and / or 2D speaker setups, speaker locations cannot be used as SSPs in some cases. This is because of the potential reduction in positioning accuracy, as extended reproduction relies on uniformly distributed (e.g., on a sphere) SSPs. Therefore, creating an equidistantly distributed object grid (e.g., on a sphere) is necessary. Figure 2In the concept box 204), these objects assume the role of SSPs. For example, each of these SSPs is translated (e.g., using VBAP) in box 205 (Translation of Extended Support Points). Box 201 (Extended Gain Calculation) uses the SSP position to calculate (for more details, see section 2.2, for example) the extended gain, so that the extended gain and the SSP translation gain can be combined in box 206 (Combination) (the simplest way to combine these two types of gains is by multiplication. Of course, other processes are possible). To reproduce small (e.g., angles smaller than the SSP grid resolution) extended angles, the actual objects are translated (e.g., "Translated Object" in box 202) and combined with the extended speaker gain (e.g., "Combined" in box 203) (for more details, see section 2.5).

[0320] 2.2 Extended gain calculation (e.g., in box 201)

[0321] To calculate the attenuation gain, a monotonically decreasing function can be used, for example. The attenuation gain can be determined from this function based on the spherical distance between the SSP and the object. For example, the function might have a gain value of 1 at the object's location and a gain value of 0 at locations where an attenuation effect is undesirable. For example, the attenuation gain can be limited to a range between 0 and 1, and amplification cannot be allowed.

[0322] Without further processing steps, this process only allows the creation of, for example, Figure 1 The uniformly spreading pattern is shown at reference numerals 100, 101, and 102 in the attached figures. To achieve a non-uniformly spreading pattern (e.g., as...), Figure 1 As shown in reference numerals 104 and 105 in the accompanying drawings, extended angles (e.g., [azimuth and elevation] or [width and height]) should be handled separately. Therefore, in the following text, the weighting function is designed not only to be one-dimensional (e.g., one extended value) but also two-dimensional (e.g., width and height).

[0323] For example, such as Figure 3 As described in the text, a two-dimensional gain function can be modeled as a combination of two one-dimensional gain functions. For example, to model a non-uniform horizontally extended pattern (e.g., Figure 3 The model is constructed using the upper part of the drawing, selecting a large expansion angle, which creates a wide one-dimensional gain function (shown as prominent on the left wall). Simultaneously, for example, a narrow elevation expansion angle is selected, which creates a narrow one-dimensional gain function (shown as prominent on the right wall). For illustrative purposes, the one-dimensional function is normalized to have a maximum value of 1. For example, combining two one-dimensional functions yields a two-dimensional function. For example, the simplest way to combine two one-dimensional functions is by multiplication. Of course, different other processes are possible.

[0324] 2.2.1 Algorithm

[0325] The following example illustrates the calculation of extended gain based on one SSP and one object. However, the algorithm can then be performed on all SSPs, and the resulting extended gains can be summed.

[0326] For example, in box 301, the absolute difference between the azimuth angle of the SSP and the object is calculated. In box 303, for example, the expansion value is constrained so that it does not take a value less than the SSP raster resolution angle in the case of non-uniform expansion. Otherwise, the "translation" of non-uniform expansion used to move the object may cause perceptible jumps in some cases. For example, an expansion gain component is calculated based on the difference angle from box 301 and the expansion angle from box 303, while the expansion angle controls, for example, the shape / width of a one-dimensional gain curve, and the difference angle selects a value. Figure 4 An example is shown in the figure.

[0327] For example, the same process is repeated for the elevation angle values ​​in boxes 304, 306, and 305. For example, the two results are multiplied in box 313. This product already represents a value on the surface of the two-dimensional gain function, but only within the elevation angle range, since the elevation angle is naturally limited to [-90°, 90°]. It should be noted that the definition of a sphere in spherical coordinates generally (or conventionally) only allows two combinations: either an azimuth angle with a range of [-180°, 180°] and an elevation angle of [-90°, 90°], or the azimuth angle can be limited to [-90°, 90°] and the elevation angle range extended to [-180°, 180°]. Otherwise, the latter part of the sphere would be defined twice. However, this limitation can be overcome in some embodiments of the invention.

[0328] For example, suppose an object located at an azimuth angle of 30° and an elevation angle of 80° in the foreground hemisphere has a vertical extension of 60°. Since the extension is, for example, symmetrical with respect to the object, an additional 20° of the extension would be located in the rear hemisphere, where the azimuth angle is, for example, 210° (or -150°). In this case, the horizontal one-dimensional gain function would, for example, mask the extension gain to near zero, because the horizontal gain function is chosen to have a narrow shape, allowing the vertical extension to be achieved.

[0329] For this reason, in some embodiments, the elevation range is extended (e.g., more optional details are discussed in section 2.3) to [-180°, 180°] and mapped back to the original range (e.g., covered by signal processing boxes 314 and 315). For example, similar to the processes in boxes 301, 302, 304, 305, and 313, the extended gain components are calculated for the extended elevation range, e.g., in boxes 308, 309, 310, 311, and 312. Finally, for example, the results from boxes 312 and 313 are added in box 314 and normalized in box 315 (e.g., more optional details are discussed in section 2.4).

[0330] 2.3 Elevation range extension (e.g., in box 307)

[0331] For example, in some embodiments, in order to present an extension in the rear hemisphere when the object is in the front hemisphere (and vice versa), all SSPs must be mirrored to the extended elevation range (i.e., from [0°, 90°] to [91°, 180°] and from [-90°, -1°] to [-180°, -91°]). This process is illustrated based on the following example.

[0332] As an example, assuming the SSP is at (20°, 40°) (azimuth and elevation), its mirror image SSP is at (-160°, 140°). This simply means, for example, that the azimuth is shifted by 180° and the elevation is mirrored to an extended elevation range while maintaining its angular distance from the horizontal plane.

[0333] Another example for the lower hemisphere: the SSP at (120°, -70°) is mirrored to (-60°, -110°).

[0334] 2.4 Normalization (e.g., in box 315)

[0335] In some cases, extending the elevation range, calculating the gain in the extended range, and adding additional gain components to the gain determined from the original range can lead to amplification. For example, the maximum gain in the original range is always 1 at the object's location. The gain component from the extended range added to the maximum gain from the original range is at a mirror position and depends, for example, on the extended value (e.g., azimuth and elevation in the combination).

[0336] For example, for a monotonically decreasing function, such addition of the two gain components will result in the maximum possible gain for a given extended value. Therefore, in some cases, it is necessary (or advantageous) to normalize the extended gain of an object to this maximum value. For example, the normalized gain is:

[0337] ,

[0338] Wherein, 𝑆𝐺𝐶𝐴 is, for example, the extended gain component determined from box 309 for the azimuth difference between the object and its mirror position (the azimuth difference between the object and its mirror position is always 180°), and 𝑆𝐺𝐶𝐸 is, for example, the extended gain component determined from box 311 for the elevation difference between the object and its mirror position.

[0339] 2.5 Combinations of expansion and object gain (e.g., in box 203)

[0340] For example, one approach to combining the target speaker gain and the extended speaker gain is to sum them. The resulting vector containing the speaker gain may need to be normalized to its Euclidean norm, for example. Depending on the SSP raster resolution, it may be necessary, for example, to attenuate the extended speaker gain before combining. For example, in the case of low SSP raster resolution (e.g., low SSP raster resolution allows for low computational complexity), a rapidly changing extended value from 0° upwards may cause noticeable artifacts. For example, the attenuation can be determined according to the following equation (where it should be noted that other equations are also possible):

[0341]

[0342] in 𝑟𝑒𝑠 or g 𝑟𝑒𝑠 It is SSP raster resolution, 𝑠𝑝𝑟𝑒𝑎d 𝑎𝑧𝑖 It is the azimuth extension angle, 𝑠𝑝𝑟𝑒𝑎d 𝑒l𝑒 It is the elevation angle extension angle.

[0343] 3. Efficient Implementation Methods

[0344] Optional details of efficient implementation methods will be described below, which may be used individually and in combination with any of the embodiments disclosed herein.

[0345] For example, since the SSPs are chosen to be distributed equidistantly on the sphere, they will have the same angular distance, for example, on each horizontal / vertical layer.

[0346] Figure 6 An example of an SSP grid with a resolution of 45° is depicted. It produces 8 vertical layers and 5 horizontal layers, where the SSP at an elevation angle of + / -90° should be defined only once.

[0347] However, it is sufficient to calculate the difference azimuth angle between the object and the SSP only once on the horizontal layer (e.g., in boxes 301 and 308) and the difference elevation angle only once on the vertical layer (e.g., in boxes 304 and 310).

[0348] Furthermore, for example, it is possible to calculate a difference angle in the clockwise direction (e.g., on a horizontal layer) and a difference angle in the counterclockwise direction, and to calculate the difference angle from the object to each SSP by taking into account the SSP index.

[0349] Example:

[0350] On the horizontal layer, the azimuth angles of the SPP (or SSP) are: [0°, 45°, 90°, 135°, 180°, -135°, -90°, -45°]. The indices they have prepared (e.g., prepared during initialization) can be, for example, [1, 2, 3, 4, 5, 6, 7, 8], and these indices are stored in an index ring (using an index ring allows, for example, jumping from index 1 to 8 and from 8 to 1. It is, for example, an infinite loop or at least approximately an infinite loop).

[0351] For example, assume the object's azimuth is 30°. Regarding the SSP index, it lies between indices 1 and 2. The difference is 15° clockwise and 30° counterclockwise. For a given extended angle of 180° (i.e., symmetrical about the object's position + / - 90°), activate the SSP at azimuth angles [45°, 90°] (clockwise) and [0°, -45°] (counterclockwise). Therefore, only the two indices in the clockwise direction and the two indices in the counterclockwise direction need to be considered during the calculation of the difference angle.

[0352] For example, this method has two advantages. First, by using an index ring, it avoids wrapping around angles outside the interval [-180°, 180°]. Furthermore, it allows computation to be limited to the relevant SSP. For small extended angles, this significantly reduces the computational workload.

[0353] Another possibility for making the algorithm more computationally efficient is to choose the design of the weighting function accordingly. On the one hand, the "Gaussian clock" shaped gain curve, as previously introduced, utilizes an exponential function, which is implemented as a power series expansion, and is therefore computationally inefficient. On the other hand, the gain function should ideally have the following properties: Figure 7 The shape shown is that of the function obtained according to a translation algorithm (e.g., VBAP). This is particularly ideal (or even required in some cases) when using a combination of non-uniform spreading patterns and small spreading angles to ensure smooth object movement.

[0354] It has been found that when using an analogy that is appropriately stretched and flipped, a suitable trade-off is made. This avoids exponential / trigonometric functions and approximates the shape of the VBAP translation curve well.

[0355] 4. Implement alternative solutions

[0356] Although some aspects have been described in the context of the apparatus, it will be clear that these aspects also represent a description of the corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of the features of the corresponding block or item or the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (such as microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0357] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) on which electronically readable control signals are stored, which cooperate with (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.

[0358] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0359] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium.

[0360] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.

[0361] In other words, embodiments of the method of the present invention are therefore computer programs having program code for performing one of the methods described herein when the computer program is run on a computer.

[0362] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded, the computer program being used to perform one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.

[0363] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence of a computer program used to perform one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).

[0364] Another embodiment includes a processing means, such as a computer or a programmable logic device, which is configured or adapted to perform one of the methods described herein.

[0365] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0366] Another embodiment of the invention includes an apparatus or system configured to transmit a computer program to a receiver (e.g., electronically or optically) for performing one of the methods described herein. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0367] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0368] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.

[0369] The apparatus described herein, or any component thereof, may be implemented, at least in part, in hardware and / or software.

[0370] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.

[0371] The methods or any components of the apparatus described herein may be performed, at least in part, by hardware and / or by software.

[0372] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the invention is intended to be limited only by the scope of the appended claims and not by the specific details given by way of the description and explanation of the embodiments herein.

[0373] Additionally, it should be noted that the term “consider” can, for example, but not necessarily must, have the meaning of “based on” or “depending on”.

[0374] Additionally, it should be noted that the term "describe" may, for example, but not necessarily, have the meaning of "representation," "direct or indirect representation," "as a measure of," or "constituting." For example, the first quantity that "describes" another quantity may be equal to the other quantity, or may be proportional to the other quantity, or may be related to the other quantity using a predetermined (linear or nonlinear) relationship.

[0375] Additionally, it should be noted that the phrase "associated with azimuth values" can, for example, mean "having azimuth values".

[0376] Additionally, it should be noted that the phrase "associated with elevation angle" can, for example, mean "having elevation angle".

Claims

1. An audio object renderer (200; 1200) for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and object feature information (1212), said speaker gains (214, 1214, 1214a~c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c), in, The audio object renderer is configured to obtain the speaker gain (202a, 1232, g) of the translated object using point source translation (202, 1230) of the audio object, wherein the audio object is regarded as a point source in the point source translation, and wherein the object feature information is ignored in the point source translation; The point source translation uses the object's position information; The audio object renderer is configured to take into account the object feature information (1212) to obtain the object feature information speaker gain (206a, 1242, gOS), wherein the audio object extends over the extended region; The audio object renderer is configured to combine the translation object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) in such a way that the contribution of the translation object speaker gain is always present, in order to obtain the combined speaker gain (214, 1214, 1214a~c). The determination of speaker gain based on the object feature information takes into account the extension of the audio object.

2. The audio object renderer (200; 1200) according to claim 1. in, The audio object renderer is configured to also consider the object location information (210, 1210, azi, ele) to obtain object feature information speaker gain (206a, 1242, gOS).

3. The audio object renderer (200; 1200) according to claim 1. in, The object feature information is the audio object extended information (212, 1212).

4. The audio object renderer (200; 1200) according to claim 2. in, The object feature information is the audio object extended information (212, 1212).

5. An audio object renderer (200; 1200) for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and object feature information (1212), said speaker gains (214, 1214, 1214a~c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c), in, The audio object renderer is configured to obtain the speaker gain (202a, 1232, g) of the translated object using point source translation (202, 1230) of the audio object, wherein the audio object is regarded as a point source in the point source translation, and wherein the object feature information is ignored in the point source translation; The point source translation uses the object's position information; The audio object renderer is configured to take into account the object location information (210, 1210, azi, ele) and the object feature information (1212) to obtain the extended object speaker gain (206a, 1242, gOS), wherein the audio object extends over the extended region; The audio object renderer is configured to combine the translation object speaker gain (202a, 1232, g) and the extended object speaker gain (206a, 1242, gOS) in such a way that the contribution of the translation object speaker gain is always present, in order to obtain the combined speaker gain (214, 1214, 1214a~c). The determination of the extended object speaker gain takes into account the extension of the audio object.

6. An audio object renderer (200; 1200) for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and extended information (212, 1212), said speaker gains (214, 1214, 1214a~c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c), in, The audio object renderer is configured to obtain the speaker gain (202a, 1232, g) of the translated object using point source translation (202, 1230) of the audio object, wherein the audio object is regarded as a point source in the point source translation, and wherein the extended information is ignored in the point source translation; The point source translation uses the object's position information; The audio object renderer is configured to take into account the object location information (210, 1210, azi, ele) and the extension information (212, 1212) to obtain the extended object speaker gain (206a, 1242, gOS), wherein the audio object extends over the extension region. The audio object renderer is configured to combine the translation object speaker gain (202a, 1232, g) and the extended object speaker gain (206a, 1242, gOS) in such a way that the contribution of the translation object speaker gain is always present, in order to obtain the combined speaker gain (214, 1214, 1214a~c). The determination of the extended object speaker gain takes into account the extension of the audio object.

7. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to evaluate one or more gain functions that map the difference between the positions (204a, aziSSP, eleSSP) of the support points (610a, 612a~612g, 614b~614f, 616c~616e) and the object positions (210, 1210, azi, ele) to one or more extended gain value contributions (302a, aziGain(naz), 305a, eleGain(nel)), and the audio object renderer is configured to determine the extended object speaker gain (206a, 1242, gOS) based on the one or more extended gain value contributions.

8. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to adjust according to the spread in the first direction (spreadAngleAzi, spread). azi And based on the extension in the second direction (spreadAngleEle, spread) ele To determine the weights (attenGain, g) of the extended object loudspeaker gain (206a, 1242, gOS) in combination with the translation object loudspeaker gain (202a, 1232, g). atten ), where the combination is a weighted combination.

9. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to adjust according to the spread angle (spreadAngleAzi, spread) in the first direction. azi ) and the spreading angle in the second direction (spreadAngleEle, spread) ele The product of the extended object loudspeaker gain (206a, 1242, gOS) and the translation object loudspeaker gain (202a, 1232, g) determines the weight (attenGain, g) in the combination. atten ), where the combination is a weighted combination.

10. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to apply a fixed-weighted translation object speaker gain (202a, 1232, g) and a variable-weighted gain (attenGain, g) to the target audio object. atten The weighted extended object speaker gains (206a, 1242, gOS) are summed, and the variable weights (attenGain, g) are added together. atten ) depends on the spread angle in the first direction (spreadAngleAzi, spread azi ) and the spreading angle in the second direction (spreadAngleEle, spread) ele ).

11. The audio object renderer (200, 1200) according to claim 10. in, The audio object renderer is configured to apply a fixed-weighted translation object speaker gain (202a, 1232, g) and a variable-weighted gain (attenGain, g). atten The result of adding the weighted extended object speaker gains (206a, 1242, gOS) is normalized.

12. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to determine the weight attenGain of the extended object speaker gain (206a, 1242, gOS) in a combination with the translation object speaker gain (202a, 1232, g), wherein the combination is a weighted combination: attenGain = 0.89f * min(c1, max(spread azi , spread ele ) / g res1 ) + 0.11f * min(c2, min(spread azi , spread ele ) / g res2 ); Where c1 is a predetermined value; Where c2 is a predetermined value; Where g res1 It is a pre-set value; Where g res2 It is a pre-set value; spread azi It is the extended angle of the audio object in the azimuth direction; spread ele It is the extended angle of the audio object in the elevation direction; Where min(.) is the minimum operator; and The max(.) operator is the maximum operator.

13. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to, compared to the speaker gain of the translation object (202a, 1232, g), increase with the spread angle (spreadAngleAzi, spread) of the audio object. azi ,spreadAngleEle,spread ele The relative contribution of increasing the loudspeaker gain (206a, 1242, gOS) of the extended object is increased.

14. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to consider the object position information (210, 1210, azi, ele) and the extended information (212, 1212) and use the representation of the support point position in polar coordinates (204a, aziSSP, eleSSP) to obtain the extended object speaker gain (206a, 1242, gOS); and in, The audio object renderer is configured to provide the speaker gain (214, 1214, 1214a~c) based on the extended object speaker gain (206a, 1242, gOS).

15. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured as follows: -Evaluate one or more angular differences (diffCLKDir, diffAntiCLKDir) between the azimuth position (210°, 1210°, azi) of the audio object and the azimuth positions (204°, aziSSP) of one or more support points, and / or - Evaluate one or more angular differences (diffCLKDir, diffAntiCLKDir) between the elevation position (210, 1210, azi) of the audio object and the elevation position (204a, eleSSP) of one or more support points. In order to obtain the extended object speaker gain (206a, 1242, gOS).

16. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The support points (204a, aziSSP, eleSSP) are arranged on the sphere within a tolerance of + / -10% or + / -20% of the sphere radius.

17. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The support point locations (204a, aziSSP, eleSSP) comprise uniform azimuth intervals along a circle with a constant elevation angle and a constant radius; and / or The support point locations (204a, aziSSP, eleSSP) include uniform elevation angles along a circle with a constant azimuth and a constant radius.

18. The audio object renderer (200, 1200) according to claim 1, 5 or 6. in, The audio object renderer is configured to obtain the extended object speaker gain (206a, 1242, gOS) such that the audio object extends over a region that extends in a first hemisphere where the audio object is located and also in a second hemisphere where the azimuth position is opposite to that of the first hemisphere.

19. The audio object renderer (200, 1200) according to claim 18. in, The audio object renderer is configured to use an extended elevation range between -180 degrees and +180 degrees.

20. The audio object renderer (200, 1200) according to claim 18. in, The audio object renderer is configured to target a given object position (210, 1210, azi, ele) and a given extension (212, 1212, spreadAngleAzi, spread). azi ,spreadAngleEle,spread ele Calculate the following: - The first set of azimuth gain values ​​(302a, aziGain) describes the contribution of multiple azimuth values ​​associated with the support point location or support point azimuth index (naz) to the extended gain (314a, g_spd). This first set of azimuth gain values ​​(302a, aziGain) is associated with elevation values ​​within the range of original elevation values ​​that do not cross the poles of the spherical coordinate system. - A second set of azimuth gain values ​​(309a, aziGainExtd) describes the contribution of multiple azimuth values ​​associated with the support point location or support point azimuth index (naz) to the extended gain (314a, g_spd). This second set of azimuth gain values ​​(309a, aziGainExtd) is associated with elevation values ​​within an extended range of elevation values ​​indicating passage through one of the poles of the spherical coordinate system. The extended gain (206a, 1242, gOS) is derived using the first set of azimuth gain values ​​(302a, aziGain(naz)) and the second set of azimuth gain values ​​(309a, aziGainExtd).

21. The audio object renderer (200, 1200) according to claim 20. in, The audio object renderer is configured to calculate the following for a given object position (210, 1210, azi, ele) and for a given extension (212, 1212, spreadAngleAzi, spreadAngleEle): - A first set of elevation gain values ​​(305a, eleGain) describes the contribution of multiple elevation values ​​associated with the support point location or speaker azimuth index or support point elevation index (nel) to the extended gain (314a, g_spd). This first set of elevation gain values ​​(305a, eleGain) is associated with elevation values ​​within the original range of elevation values ​​that do not cross the poles of the spherical coordinate system. - A second set of elevation gain values ​​(311a, eleGainExtd) describes the contribution of multiple elevation values ​​associated with the support point location or speaker elevation index or support point elevation index (nel) to the extended gain (314a, g_spd). This second set of elevation gain values ​​(311a, eleGainExtd) is associated with elevation values ​​within an extended range of elevation values ​​indicating passage through one of the poles of the spherical coordinate system. The extended gain (206a, 1242, gOS) is derived using the first set of azi angle gain values ​​(302a, aziGain(naz)), the second set of azi angle gain values ​​(309a, aziGainExtd), the first set of elevation gain values ​​(305a, eleGain(nel)), and the second set of elevation gain values ​​(311a, eleGainExtd(nel)).

22. The audio object renderer (200, 1200) according to claim 21. in, The audio object renderer is configured to combine the first set of azimuth gain values ​​(302a, aziGain(naz)) and the first set of elevation gain values ​​(305a, eleGain(nel)), and to combine the second set of azimuth gain values ​​(309a, aziGainExtd(naz)) and the second set of elevation gain values ​​(311a, eleGainExtd(nel)).

23. The audio object renderer (200, 1200) according to claim 20. in, The second set of azimuth gain values ​​(309a, aziGainExtd) represents the evolution of the gain value at an azimuth angle shifted by 180 degrees compared to the evolution of the gain value at the azimuth angle represented by the first set of azimuth gain values ​​(302a, aziGaind).

24. The audio object renderer (200, 1200) according to claim 20. in, The first set of azimuth gain values ​​(302a, aziGain) represents the azimuth object position (210, 1210, azi) and azimuth spread angle (spreadAngleAzi, spread) given the azimuth object position (210, 1210, azi) and the azimuth spread angle (spreadAngleAzi, spread). azi The gain value evolves over a 360-degree range, where angular accuracy is determined by the number of speakers or by the number of support points, and / or The second set of azimuth gain values ​​(309a, aziGainExtd) represents the azimuth object position (210, 1210, azi) and azimuth spread angle (spreadAngleAzi, spread) given a 180-degree rotation. azi The gain value evolves over a 360-degree range, where angular accuracy is determined by the number of speakers or the number of support points.

25. The audio object renderer (200, 1200) according to claim 21. in, The first set of elevation gain values ​​(305a, eleGain) represents the elevation gain based on the object's position (210, 1210, ele) and the elevation spread angle (spreadAngleEle, spread). ele The evolution of the gain value over the elevation angle range between -90 degrees and +90 degrees, and / or The second set of elevation gain values ​​(311a, eleGainExtd) represents the elevation gain values ​​based on the object's position (210, 1210, ele) and the elevation spread angle (spreadAngleEle, spread). ele The evolution of the gain value over the elevation angle range between -180 degrees and -90 degrees and between +90 degrees and +180 degrees.

26. The audio object renderer according to claim 1, 5, or 6, in, The audio object renderer is configured to determine speaker gains (214, 1214, 1214a~c) based on object location information (210, 1310, azi, ele) and spread information (212, 1312, spreadAngleAzi, spreadAngleEle). These speaker gains (214, 1214, 1214a~c) describe the gain used to include one or more audio object signals (1260) into multiple speaker signals (1262a~1262c). The audio object renderer is configured to take into account the object location information and the extension information to obtain the extended object speaker gain (206a, 1242, gOS). The audio object renderer is configured to use one or more polynomial functions of degree less than or equal to 3 to obtain extended gain (314a, g_spd). These polynomial functions map the angular difference between the object position (210, 1310, azi, ele) and the support point position (204a, aziSSP, eleSSP) to extended gain value contributions (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd). The audio object renderer is configured to obtain the extended object speaker gain (206a, 1242, gOS) using the extended gain (314a, g_spd) contributed based on the extended gain value.

27. The audio object renderer (200, 1300) according to claim 26. in, The width of the one or more polynomial functions is determined by the extended information (212, 1312, spreadAngleAzi, spreadAngleEle).

28. The audio object renderer (200, 1300) according to claim 26. in, The audio object renderer is configured to use a first polynomial function and a second polynomial function to obtain extended gain values ​​(314a, g_spd). The first polynomial function maps the azimuth difference between the object position (210, 1310, azi) and the support point position (204a, aziSSP) to a first extended gain contribution (302a, aziGain(naz)). The second polynomial function maps the elevation difference between the object position (210, 1310, ele) and the support point position (204a, eleSSP) to a second extended gain contribution (305a, eleGain(nel)).

29. The audio object renderer (200, 1300) according to claim 28. in, The audio object renderer is configured to combine the first extended gain contribution (302a, aziGain(naz)) and the second extended gain contribution (305a, eleGain(nel)) to obtain an extended gain value (314a, g_spd).

30. The audio object renderer (200, 1300) according to claim 26. in, The audio object renderer is configured to calculate the following for a given object position (210, 1310, azi, ele) and for a given extension (212, 1312, spreadAngleAzi, spreadAngleEle): - A set of azimuth gain values ​​(302a, aziGain) describing the contribution of multiple azimuth values ​​associated with the support point location or speaker azimuth index or support point azimuth index (naz) to the extended gain (314a, g_spd), and / or - A set of elevation gain values ​​(305a, eleGain) describing the contribution of multiple elevation values ​​associated with the support point location or speaker elevation index or support point elevation index (nel) to the extended gain (314a, g_spd). And the extended gain (314a, g_spd) is derived using the group azi angle gain values ​​(302a, aziGain(naz), 305a, eleGain(nel)).

31. The audio object renderer (200, 1300) according to claim 30. in, The audio object renderer is configured to combine the following two elements: an element of the group azi gain value (302a, aziGain) associated with the currently considered speaker or the currently considered support point (aziGain(naz)), and an element of the group elevation gain value (305a, eleGain) associated with the currently considered speaker or the currently considered support point (eleGain(nel)) to obtain an extended gain value (spd(objNo)) associated with multiple different speakers or multiple different support points.

32. The audio object renderer (200, 1300) according to claim 26. in, The audio object renderer is configured to calculate the following for a given object position (210, 1310, azi, ele) and for a given extension (212, 1312, spreadAngleAzi, spreadAngleEle): - A first set of aziance gain values ​​(302a, aziGain) describes the contribution of multiple aziance values ​​associated with the support point location or speaker aziance index or support point aziance index (naz) to the extended gain (314a, g_spd). This first set of aziance gain values ​​(302a, aziGain) is associated with elevation values ​​within the range of original elevation values ​​that do not cross the poles of the spherical coordinate system. - A second set of azimuth gain values ​​(309a, aziGainExtd) describes the contribution of multiple azimuth values ​​associated with the support point location or the speaker azimuth index or the support point azimuth index (naz) to the extended gain (314a, g_spd). This second set of azimuth gain values ​​(309a, aziGainExtd) is associated with elevation values ​​within an extended range of elevation values ​​indicating the poles of the spherical coordinate system. The extended gain (314a, g_spd) is derived using the set of azi angle gain values ​​(302a, aziGain(naz)) and / or using a set of elevation gain values ​​(305a, eleGain(nel)).

33. The audio object renderer (200, 1300) according to claim 32. in, The audio object renderer is configured to calculate the following for a given object position (210, 1310, azi, ele) and for a given extension (212, 1312, spreadAngleAzi, spreadAngleEle): - A first set of elevation gain values ​​(305a, eleGain) describes the contribution of multiple elevation values ​​associated with the support point location or speaker azimuth index or support point elevation index (nel) to the extended gain (314a, g_spd). This first set of elevation gain values ​​(305a, eleGain) is associated with elevation values ​​within the original range of elevation values ​​that do not cross the poles of the spherical coordinate system. - A second set of elevation gain values ​​(311a, eleGainExtd) describes the contribution of multiple elevation values ​​associated with the support point location or speaker elevation index or support point elevation index (nel) to the extended gain (314a, g_spd). This second set of elevation gain values ​​(311a, eleGainExtd) is associated with elevation values ​​in an extended range indicating the range of elevation values ​​across the poles of the spherical coordinate system. The extended gain (314a, g_spd) is derived using the set of azimuth gain values ​​(302a, aziGain(naz), 309a, aziGainExtd(naz)) and the set of elevation gain values ​​(305a, eleGain(nel), 311a, eleGainExtd(nel)).

34. The audio object renderer (200, 1300) according to claim 26. in, The audio object renderer is configured to use translation pre-calculation during initialization to translate the audio signal associated with multiple support points to multiple speakers using the support point translation gain (Spread.gainsSSP). The audio object renderer is configured to use a polynomial function of degree less than or equal to 3 to obtain the object-to-support point spread gain (g_spd), which describes the contribution of the audio object signal to multiple support point signals; and The audio object renderer is configured to combine the object-to-support point extension gain and the support point translation gain to obtain the extended object speaker gain.

35. The audio object renderer according to claim 26, in, The one or more polynomial functions of degree 3 or less are parabolic functions, which provide a return value p according to the following formula: p = max(0, c1*anglediff 2 +c2), Where c1 is a parameter that determines the width of the parabolic function; Where c2 is a predetermined value; Where angeldiff is the angular difference in which the parabolic function is evaluated; and The max(.,.) operator is the maximum value operator that returns the maximum value of its operands.

36. The audio object renderer according to claim 1, 5, or 6, in, The audio object renderer is configured to provide the combined speaker gain based on both point source translation and expansion of the audio object signal.

37. The audio object renderer according to claim 1, 5, or 6, in, Compared to determining the speaker gain of the translated object, determining the speaker gain of the object feature information extends the audio object to more speakers.

38. A method for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and object feature information (1212), said speaker gains (214, 1214, 1214a~c) being used to include one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c). in, The method includes obtaining a speaker gain (202a, 1232, g) of a translated object using point source translation (202, 1230) of an audio object, wherein the audio object is considered a point source in the point source translation, and wherein the object feature information is ignored in the point source translation; The point source translation uses the object's position information; The method includes taking into account the object feature information (1212) to obtain object feature information speaker gain (206a, 1242, gOS), wherein the audio object extends over the extended region; The method includes combining the translation object speaker gain (202a, 1232, g) and the object feature information speaker gain (206a, 1242, gOS) in a manner that always contributes to the translation object speaker gain, in order to obtain a combined speaker gain (214, 1214, 1214a~c). The determination of speaker gain based on the object feature information takes into account the extension of the audio object.

39. A method (1400) for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and object feature information (1212), said speaker gains (214, 1214, 1214a~c) being used to include one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c), in, The method includes obtaining (1410) the speaker gain of the translated object (202a, 1232, g) by using point source translation of the audio object. Wherein, the audio object is regarded as a point source in the point source translation, and wherein the object feature information is ignored in the point source translation; The point source translation uses the object's position information; The method includes obtaining (1420) extended object speaker gain (206a, 1242, gOS) by considering the object location information (210, 1210, azi, ele) and the object feature information (1212), wherein the audio object extends over the extended region; The method includes combining (1430) the translation object speaker gain (202a, 1232, g) and the extended object speaker gain (206a, 1242, gOS) in such a way that the contribution of the translation object speaker gain is always present, in order to obtain a combined speaker gain (214, 1214, 1214a~c). The determination of the extended object speaker gain takes into account the extension of the audio object.

40. A method (1400) for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and extended information (212, 1212, spreadAngleAzi, spreadAngleEle), said speaker gains (214, 1214, 1214a~c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c), in, The method includes obtaining (1410) a translated object speaker gain (202a, 1232, g) using a point source translation of an audio object, wherein the audio object is considered a point source in the point source translation, and wherein the extended information is ignored in the point source translation; The point source translation uses the object's position information; The method includes taking into account the object location information (210, 1210, azi, ele) and the extension information (212, 1212, spreadAngleAzi, spreadAngleEle) to obtain (1420) the extended object speaker gain (206a, 1242, gOS), wherein the audio object extends over the extension region; The method includes combining (1430) the translation object speaker gain (202a, 1232, g) and the extended object speaker gain (206a, 1242, gOS) in such a way that the contribution of the translation object speaker gain is always present, in order to obtain a combined speaker gain (214, 1214, 1214a~c). The determination of the extended object speaker gain takes into account the extension of the audio object.

41. The method according to claim 38, 39, or 40, in, The method includes evaluating one or more gain functions that map the difference between the location of the support point and the location of the object to one or more extended gain value contributions (aziGain(naz), eleGain(nel)), and the method includes determining the extended object speaker gain (gOS) based on the one or more extended gain value contributions.

42. The method according to claim 38, 39, or 40, in, The method includes taking into account the object location information (210, 1210, azi, ele) and the object feature information (1312) to obtain (1510) extended object speaker gain (206a, 1242, gOS). The method includes using one or more polynomial functions of degree less than or equal to 3 to obtain the (1520) extended gain (314a, g_spd), wherein the one or more polynomial functions map the angular difference between the object position (210, 1310, azi, ele) and the support point position (204a, aziSSP, eleSSP) to the extended gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd), and The method includes obtaining (1530) the extended target speaker gain (206a, 1242, gOS) using the extended gain (314a, g_spd) based on the extended gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd), or using the extended gain (314a, g_spd) as the extended target speaker gain (206a, 1242, gOS).

43. The method according to claim 38, 39, or 40, in, The method determines speaker gains (214, 1214, 1214a~c) based on object location information (210, 1310, azi, ele) and spread information (212, 1312, spreadAngleAzi, spreadAngleEle). These speaker gains (214, 1214, 1214a~c) describe the gain of including one or more audio object signals (1260) into multiple speaker signals (1262a~1262c). The method includes taking into account the object location information and the extended information to obtain (1510) the extended object speaker gain (206a, 1242, gOS). The method includes using one or more polynomial functions of degree less than or equal to 3 to obtain the (1520) extended gain (314a, g_spd), wherein the one or more polynomial functions map the angular difference between the object position (210, 1310, azi, ele) and the support point position (204a, aziSSP, eleSSP) to the extended gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd, aziGain(naz), eleGain(nel)), and The method includes obtaining (1530) the extended target speaker gain (206a, 1242, gOS) using the extended gain (314a, g_spd) based on the extended gain value contribution (302a, aziGain, 305a, eleGain, 309a, aziGainExtd, 311a, eleGainExtd, aziGain(naz), eleGain(nel)), or using the extended gain (314a, g_spd) as the extended target speaker gain.

44. A computer program for performing the method according to claim 38, 39, or 40 when the computer program is run on a computer.

45. An audio object renderer (200, 1200) for determining speaker gains (214, 1214, 1214a~c) based on object location information (210, 1210, azi, ele) and extended information (212, 1212), said speaker gains (214, 1214, 1214a~c) for including one or more audio object signals (1260) into a plurality of speaker signals (1262a~1262c), in, The audio object renderer is configured to obtain the speaker gain (202a, 1232, g) of the translated object using point source translation (202, 1230) of the audio object, wherein the audio object is regarded as a point source in the point source translation, wherein the extended information is ignored, and wherein a single speaker is selected for playback of the audio object or wherein the audio object is distributed to multiple speakers closest to the audio object. The speaker gain of the translated object is based on the object's position information; The audio object renderer is configured to obtain the extended object speaker gain (206a, 1242, gOS) based on the object location information (210, 1210, azi, ele) and the extension information (212, 1212), wherein the audio object extends over the extension region. The audio object renderer is configured to combine the translation object speaker gain (202a, 1232, g) and the extended object speaker gain (206a, 1242, gOS) in such a way as to obtain a combined speaker gain (214, 1214, 1214a~c): the translation object speaker gain always contributes, that is, the contribution of the translation object speaker gain in the combination is non-zero; The determination of the extended object speaker gain takes into account the extension of the audio object; The audio object renderer is configured to provide the combined speaker gain based on both point source translation and expansion of the audio object signal; and In this context, compared to determining the speaker gain of the translation object, determining the speaker gain of the extended object extends the audio object to more speakers.

Citation Information

Patent Citations

  • Rendering of audio objects with apparent size to arbitrary loudspeaker layouts

    US20160007133A1