Rendering audio element
Adaptive placement of a center virtual speaker and attenuation coefficients address the challenge of rendering spatially heterogeneous audio elements in XR, ensuring consistent spatial representation across varying listener positions with reduced complexity.
Patent Information
- Application Number
- JP2025134482
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-11
- Filing Date
- 2025-08-12
- Publication Date
- 2025-12-23
AI Technical Summary
Existing methods for rendering spatially heterogeneous audio elements in extended reality (XR) environments face challenges in achieving optimal spatial representation with a limited number of virtual speakers, leading to psychoacoustic gaps and increased complexity, especially when the listener's position is not fixed.
Adaptive placement of a center virtual speaker based on the listener's position, using a combination of anchor points and midpoint calculations, along with attenuation coefficients to maintain even audio distribution across the audio element's range.
The method ensures consistent spatial representation of audio elements for any listener position with fewer virtual speakers, reducing complexity and minimizing psychoacoustic gaps.
Smart Images

Figure 2025186226000001_ABST
Abstract
Description
[Technical Field]
[0001] SUMMARY Embodiments related to rendering audio elements are disclosed. [Background technology]
[0002] Spatial audio rendering is a process used to present audio in an extended reality (XR) scene (e.g., a virtual reality (VR), augmented reality (AR), or mixed reality (MR) scene) that gives the listener the impression that the sound is coming from a physical source in the scene at a specific location and has a specific size and shape (i.e., extent). Presentation can be through headphone speakers or other speakers. When presentation is through headphone speakers, the process used is called binaural rendering and uses spatial cues of human spatial hearing that allow one to identify which direction a sound is coming from. The cues involve interaural time delay (ITD), interaural level difference (ILD), and / or spectral differences.
[0003] The most common form of spatial audio rendering is based on the concept of point sources, where each sound source is defined as emanating from one specific point. Because each sound source is defined as emanating from one specific point, the sound sources do not have any size or shape. Different methods have been developed to render sound sources that have a range (size and shape).
[0004] One such known method is to create multiple copies of a mono audio element at positions around the audio element. This configuration creates the perception of a spatially homogeneous object with a specific size. This concept is used, for example, in the "object spread" and "object divergence" features of the MPEG-H 3D Audio standard (see References [1] and [2]) and the "object divergence" feature of the EBU Audio Definition Model (ADM) standard (see Reference [4]). This idea using mono audio sources has been further developed as described in Reference [7], where the area-volume geometry of the sound object is projected onto a sphere around the listener, and the sound is rendered to the listener using a pair of head-related (HR) filters evaluated as the integral of all HR filters covering the geometric projection of the object on the sphere. For spherical volumetric sources, this integral has an analytical solution. However, for arbitrary area-volume source geometries, the integral is evaluated by sampling the projected source surface on the sphere using so-called Monte Carlo ray sampling.
[0005] Another rendering method is to render a spatially diffuse component in addition to the mono audio signal, which creates the perception of somewhat diffuse objects that, in contrast to the original mono audio elements, do not have a distinct pinpoint location. This concept is used, for example, in the "object diffuseness" feature of the MPEG-H 3D Audio standard (see reference [3]) and the "object diffuseness" feature of the EBU ADM (see reference [5]).
[0006] Combinations of the two methods above are also known, for example the "object extent" feature of the EBU ADM, which combines the creation of multiple copies of a mono audio element with the addition of a diffuse component (see reference [6]).
[0007] In many cases, the actual shape of an audio element can be adequately described by a basic shape (e.g., a sphere or a box), but the actual shape may be more complex and need to be described in a more detailed form (e.g., a mesh structure or a parametric description form).
[0008] However, these methods do not allow rendering of audio elements with distinct spatially heterogeneous characteristics, i.e., audio elements with a certain amount of spatial source variation within their spatial extent. Often, these sound sources are composed of the summation of multiple sound sources (e.g., forest sounds or cheering crowd sounds). Most of these known solutions are only able to create objects with spatially homogeneous (i.e., no spatial variation within the elements) or spatially diffuse characteristics, which are too limited to render some of the above examples in a convincing way.
[0009] For heterogeneous audio elements, as described in Reference [8], the audio element contains at least two audio channels (i.e., audio signals) to describe the spatial variation across its range. Techniques exist for rendering these heterogeneous audio elements, where the audio element is represented by a multi-channel audio recording and the rendering uses several virtual speakers to represent the audio element and the spatial variation within it. By placing the virtual speakers at positions corresponding to the range of the audio element, the illusion of audio emanating from the audio element can be conveyed.
[0010] The number of virtual speakers required to achieve reasonable spatial rendering of spatially heterogeneous audio elements depends on the range of the audio elements. If the spatially heterogeneous audio elements are small or located at a certain distance from the listener, a two-speaker setup may be sufficient. However, as shown in FIG. 1, for audio elements that are large and / or close to the listener, a two-speaker setup may be too sparse, causing a psychoacoustical hole between the left speaker (SP-L) and the right speaker (SP-R) because the speakers are placed far apart. Adding a third center speaker helps ameliorate this effect. For this reason, a center speaker is used in most standardized multi-channel speaker setups. The simplest way to render spatially heterogeneous audio elements is by representing each of their audio channels as a virtual speaker, although the number of speakers can be fewer or more than the number of audio channels. If the number of virtual speakers is fewer than the number of audio channels, a downmixing step is required to derive the signals for each virtual speaker. If the number of virtual speakers is greater than the number of audio channels, an upmixing step is required to derive a signal for each virtual speaker. One implementation is to simply use two virtual speakers at fixed positions. Summary of the Invention [Problem to be solved by the invention]
[0011] Currently, several challenges exist. For example, rendering spatially heterogeneous audio elements usually requires the use of several virtual speakers. While using a large number of speakers can be beneficial for having an evenly distributed audio representation of the range, when a source signal has a limited number of channels (e.g., a stereo signal), upsampling to a large number of speakers can cause the problem that using more speakers does not increase spatial quality. Furthermore, using a large number of virtual speakers results in undesirably high complexity. On the other hand, using too few virtual speakers can significantly impair the spatial characteristics of the audio elements, and the rendering may no longer be able to well represent the corresponding audio elements. Therefore, selecting the number of virtual speakers for rendering spatially heterogeneous audio elements is a trade-off between complexity and quality.
[0012] The aforementioned problem of the psychoacoustic gap between two speakers is well known and is particularly problematic when the listener is not positioned exactly in the sweet spot of the speakers. For example, typical multi-speaker setups designed for home theater use are built under the assumption that the listener will be positioned somewhere around the sweet spot. While these systems often place a center speaker centered between the left and right front speakers, in the case of XR audio rendering, where the listener is free to move around with six degrees of freedom, a static speaker setup is not ideal. The psychoacoustic gap problem can be particularly pronounced when the listener approaches the range of the sound source.
[0013] Therefore, designing a static loudspeaker setup with a limited number of loudspeakers that provides a good spatial representation that works for any listening position is problematic. [Means for solving the problem]
[0014]
[0006] Accordingly, in one aspect, a method is provided for rendering an audio element (e.g., a spatially heterogeneous audio element), the audio element having a range and represented using a set of virtual speakers including a center virtual speaker, the method including at least one of selecting a position of the center virtual speaker based on a position of a listener and calculating an attenuation coefficient of the center virtual speaker.
[0015] In another aspect, a computer program is provided that includes instructions that, when executed by a processing circuit of an audio renderer, cause the audio renderer to perform the above-described method. In one embodiment, a carrier is provided that includes the computer program. The carrier is one of an electrical signal, an optical signal, a wireless signal, and a computer-readable storage medium. In another aspect, a rendering apparatus is provided that is configured to perform the above-described method. The rendering apparatus may include a memory and a processing circuit coupled to the memory.
[0016] An advantage of the embodiments disclosed herein is that they provide an adaptive method for the placement of virtual speakers for rendering spatially heterogeneous audio elements having a range. The embodiments allow the position of a virtual speaker representing the center of the range to adapt to the current listener position, so that the spatial distribution over the range is preserved using a small number of virtual speakers. Compared to simpler methods that distribute speakers evenly over the range of an audio element, the embodiments provide a more efficient solution that works for all listening positions without using a large number of virtual speakers. [Brief explanation of the drawings]
[0017] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate various embodiments.
[0018] [Figure 1] A diagram showing a psychoacoustic hall.
[0019] [Figure 2A]FIG. 1 illustrates an exemplary virtual speaker setup.
[0020] [Figure 2B] FIG. 1 illustrates an exemplary virtual speaker setup that can create a psychoacoustic hall.
[0021] [Figure 3] FIG. 1 illustrates an exemplary virtual speaker setup.
[0022] [Figure 4] FIG. 1 illustrates an exemplary virtual speaker setup in which the center speaker is placed near the end speakers.
[0023] [Figure 5] FIG. 1 illustrates one embodiment of how the preferred position of the center speaker may be determined.
[0024] [Figure 6] FIG. 10 illustrates another embodiment of how the preferred position of the center speaker may be determined.
[0025] [Figure 7] FIG. 10 illustrates one embodiment of how the preferred position of the center speaker can be determined when the audio element area is rectangular in shape.
[0026] [Figure 8] 1 is a flowchart illustrating a process according to some embodiments.
[0027] [Figure 9] 1 is a flowchart illustrating a process according to some embodiments.
[0028] [Figure 10] 1 is a flowchart illustrating a process according to some embodiments.
[0029] [Figure 11A] FIG. 1 illustrates a system according to some embodiments. [Figure 11B] FIG. 1 illustrates a system according to some embodiments.
[0030] [Figure 12] FIG. 1 illustrates a system according to some embodiments.
[0031] [Figure 13] 1 illustrates a signal modifier according to one embodiment.
[0032] [Figure 14] 1 is a block diagram of an apparatus according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0033] This disclosure proposes, among other things, various methods for adapting the positions of virtual loudspeakers (or "speakers" for short) used to represent audio elements with ranges (e.g., spatially heterogeneous audio elements with a certain size and shape). By using several (e.g., two or more) speakers to represent the outer edges of the audio element and at least one speaker (hereinafter "middle speaker") adaptively positioned between the edges of the audio element or the edges of a simplified shape representing the audio element, optimized rendering can be achieved with a small number of speakers, with audio energy perceptually evenly distributed across the range of the audio element. Furthermore, the potential problem of undesirable excess energy due to overlap with the extra center speaker and one of the side speakers is addressed.
[0034] overview
[0035] The objective is to render audio elements with a range such that, for any listening position, the sound is perceived by the listener as being evenly distributed across the range. Embodiments avoid or reduce the psychoacoustic hall problem by using as few speakers as possible.
[0036] In some embodiments, a set of speakers representing the edges of an audio element and a center speaker adaptively positioned to represent the center of the audio element are used to render the audio element, and the placement of the center speaker and / or its attenuation factor (e.g., gain factor) take into account the listening position (e.g., the position of the listener in virtual space relative to the audio element).
[0037] An example of such a rendering setup is shown in Figure 2A, which shows a region 200 representing an audio element. The region 200 representing an audio element may be the region of the audio element (i.e., the region 200 has the same size and shape as the actual region of the audio element), or it may be a simplified region derived from the region of the audio element (e.g., a line or a rectangle). International Patent Application WO2021180820 describes various methods for generating such simplified regions.
[0038] FIG. 2A further shows that left speaker 202 is located at left endpoint 212 of range 200, right speaker 204 is located at right endpoint 214 of the range, and center speaker 203 is located somewhere between the left and right endpoints.
[0039] Advantageously, in one embodiment, the positioning of the center speaker 203 is controlled so that the center speaker 203 is located at or near the midpoint (MP) 220 between the first and second endpoints of the range when the listener is at least a certain distance (D) from the audio element. The distance D typically depends on the size of the range 200.
[0040] As shown in Figure 2B, when a listener moves closer to an audio element, keeping the center speaker 203 at or near the midpoint 220 can lead to problems with the psychoacoustic hall 240. Therefore, in this situation, the center speaker 203 is moved to a new position, as shown in Figure 3, so that a uniform spatial distribution is maintained.
[0041] Adaptive placement of center speaker
[0042] Some embodiments herein adaptively position the center speaker 203 based on the position of the listener. In many situations, the goal is to place the center speaker on a selected "anchor point" for the audio element, which can move with the listener. For example, in one embodiment, the anchor point is the point in range 200 closest to the listener. However, in other embodiments, the anchor point can be defined differently. International Patent Application WO2021180820 describes different methods for selecting anchor points for audio elements. However, there are situations in which placing the center speaker on the anchor point is not advantageous.
[0043] For example, when a listener is close to one of the ends of an audio element, if the center speaker is placed on the anchor point, the center speaker and the corresponding side speaker will overlap, resulting in an undesirable increase in energy from the corresponding side, as shown in Figure 4. In this situation, it is advantageous to place the center speaker closer to the midpoint between the left and right end points (e.g., the positions of the left and right speakers) in order to have more evenly distributed audio energy.
[0044] As another example, when the distance between the listener and an audio element is greater than a threshold (which may depend on the range of the audio element), experiments show that placing the center speaker closer to the midpoint 220 provides a more perceptually significant output.
[0045] In summary, if the listener is close to the range but not close to one of the range's ends, it is usually preferable to place the center speaker at the anchor point. For other listening positions, the preferred center speaker location is at or near the midpoint 220.
[0046] The proposed embodiment therefore provides adaptive placement of the center speaker, optimizing its position depending on the current listening position.
[0047] More specifically, in one embodiment, placing the center speaker at the anchor point is avoided when the anchor position is close to one of the end speakers. Thus, in addition to considering the anchor point (A) 590 (see FIG. 5), the positioning of the center speaker 203 may also depend on the midpoint (MP) 220.
[0048] In one embodiment, the preferred position of the center speaker 203 on the line 599 from one end point of the range 200 (e.g., the left end point 212) to another end point of the range 200 (e.g., the right end point 214), indicated by "M" 591, is calculated using equation (1). M=α*A+(1-α)*MP (Equation 1) where A is the position of the anchor point on line 599, MP is the position of the midpoint on line 599, and α∈[0,1] is a coefficient that controls the weight of the anchor point and the midpoint in positioning the center speaker on line 599.
[0049] The value of α is such that when the listener is near the range but not near the end of the range, M is near or above the anchor point (α→1), and when the listener is near the end of the range 200 or far from the range, M is near or above the midpoint (α→0). Thus, α simultaneously takes into account the movement of the listener in both the x and z directions. Once point M 591 is determined, the center speaker 203 may be "placed" at point M. That is, the position of the center speaker 203 is set to point M.
[0050] In one embodiment, α is a function of two variables, xm_w and zm_w, i.e., α=f(xm_w, zm_w) (Equation 2)
[0051] The first variable, xm_w, is a weight that reflects the listener's movement in the x direction, and zm_w, a weight that reflects the listener's movement in the z direction. The parameters β and λ, shown in Figure 5, are used to set the values of xm_w and zm_w.
[0052] In one embodiment, xm_w is a function of λ and β, i.e.,
number
[0053] From Equation 3 and Figure 5, we see that as the listener moves towards either end of range 200, λ → 0 or β → 0, and therefore xm_w → 0. Similarly, as the listener moves towards the center, λ and β approach each other and xm_w approaches 1.
[0054] The weight zm_w in the Z direction is also a function of λ and β, i.e. zm_w=f2(λ,β)=sin((λ+β) / 2) (Equation 4)
[0055] From equation (4) and Figure 5, we can see that as the listener gets closer to the range, λ and β become larger, and at very close distances, λ + β approaches 180 degrees, and therefore zm_w approaches 1, and as the listener gets further away from the range, λ and β become smaller, and therefore zm_w approaches 0.
[0056] In one embodiment, α is defined as: α=f3(xm_w,zm_w)=min(1,d*xm_w*zm_w) (Equation 5)
[0057] where d∈(0,+∞) is a tuning factor that controls the factor α. Experiments have shown that d≈2.2 produces desirable results. The above derivation of α is just one way of using information representing the relative position of the listener with respect to the range. Other methods using λ and β (e.g., the cosine or tangent of λ and β) or other parameters that can reflect the relative position of the listener with respect to the range (e.g., the coordinates of the listener, the midpoint, the anchor point, the left and right ends of the range) may also be used to derive α.
[0058] Another approach to addressing excess energy when the listener is at the edge of an audio element may be to attenuate the energy of the signal played through the center speaker instead of changing its position. That is, the center speaker 203 is placed at the anchor point, but the audio signal for the center speaker 203 is attenuated as the listener approaches the edge of the audio element. To do so, the angles described in Figure 5 may be used as follows:
number
[0059] where X is the original time-domain audio signal of the center speaker 203 and X′ is the time-domain signal reproduced by the center speaker 203. Although this approach alleviates the excess energy problem, it may not improve the spatial perception of the audio elements.
[0060] Adaptive placement of center loudspeaker based on equal angles
[0061] In another embodiment, the placement of the center speaker is controlled only by the angles relative to the left and right endpoints. In Figure 6, these angles are φ = λ + θ and Φ = β - θ, respectively. In one embodiment, point M 591 is selected so that Φ = φ. If the distance (dA) between the listener and the anchor point is known, the distance (dM) from point M 591 to point 212 can be found by calculating dM = dA * (tan(λ) + tan(θ)), so the x-coordinate of point M is equal to the x-coordinate of the left endpoint 212 + dM.
[0062] If dA is unknown, but v (the distance between the listener and the left endpoint 212) and w (the distance between the listener and the right endpoint 214) are known, then the location of point M can be found by calculating M=(v*Re+w*Le) / (v+w), where Re is the x-coordinate of the right endpoint 214 and Le is the x-coordinate of the left endpoint 212.
[0063] Adaptive placement of center loudspeaker in two dimensions
[0064] The examples provided above illustrate one-dimensional audio element ranges. The techniques described above also apply to two-dimensional audio element ranges, such as range 700 shown in FIG. 7 . In this example, range 700 has a rectangular shape and four endpoints are defined: top endpoint 701, right endpoint 702, bottom endpoint 703, and left endpoint 704. In this example, for each endpoint, the endpoint is located at the midpoint of the side on which it lies. Thus, left endpoint 704 is exactly halfway between the upper left and lower left corners of range 700; right endpoint 702 is exactly halfway between the upper right and lower right corners of range 700; top endpoint 701 is exactly halfway between the upper left and upper right corners of range 700; and bottom endpoint 703 is exactly halfway between the lower left and lower right corners of range 700.
[0065] In one embodiment, for each defined corner point 701-704, a speaker can be placed at that point. Thus, in one embodiment, four speakers are used to represent the top, bottom, left, and right edges of the two-dimensional plane 700. In another embodiment, a speaker is placed at each corner point of the plane 700.
[0066] Additionally, a center speaker can be used and can be placed using the same principles as already explained, i.e., the coordinates (Mx, My) of the center speaker can be found by calculating: Mx=αx*A1x+(1-αx)*MPx, My=αy*A2y+(1-αy)*MPy where αx=f3(xm_w,zm1_w), αy=f3(ym_w,zm2_w), A1x is the x coordinate of anchor point A1, A2y is the y coordinate of anchor point A2, MPx is the x coordinate of the midpoint, MPy is the y coordinate of the midpoint, xm_w=f1(λx,βx), ym_w=f1(λy,βy), zm1_w=f2(λx,βx), zm2_w=f2(λy,βy).
[0067] 8 is a flow chart illustrating a process 800 according to one embodiment for rendering an audio element (e.g., a spatially heterogeneous audio element). The audio element has a range and is represented using a set of virtual speakers, including a center virtual speaker. Process 800 begins at step s802, which includes selecting a position for the center virtual speaker based on the position of the listener and / or calculating an attenuation coefficient for the center virtual speaker.
[0068] 9 is a flow chart illustrating a process 900 according to one embodiment for rendering audio elements (e.g., spatially heterogeneous audio elements). The audio elements have a range and are represented using a set of virtual speakers, including a center virtual speaker. Process 900 begins at step s902.
[0069] Step s902 includes selecting an anchor point for the audio element. Typically, the anchor point is on a line passing through a first endpoint of a range associated with the audio element (e.g., the actual range of the audio element or the simplified range of the audio element), and the second endpoint of the range and the anchor point depend on the listener's position. In some embodiments where the audio element has a complex range, selecting the anchor point is part of a process for creating the simplified range of the audio element based on the listener's position and the range of the audio element.
[0070] Step s904 includes positioning a first speaker (e.g., a right speaker) and a second speaker (e.g., a left speaker). That is, the positions of the first and second speakers are determined. In one embodiment, the speakers are positioned at opposite ends of a range. That is, the left speaker is located at the left endpoint of the range and the right speaker is located at the right endpoint of the range. In an embodiment where the range associated with the audio element is rectangular, a speaker is positioned at each corner of the rectangle.
[0071] Step s906 involves determining the midpoint between the two speakers.
[0072] Step s908 includes determining a first angle (λ) between a line from the listener to the anchor point and a line from the listener to the left speaker, and a second angle (β) between a line from the listener to the anchor point and a line from the listener to the right speaker (examples of λ and β are shown in FIG. 5).
[0073] Step s910 includes determining the x weight values (xm_w) and z weight values (zm_w) (eg, calculating xm_w=f1(λ,β) and zm_w=f2(λ,β)).
[0074] Step s912 includes determining a coefficient (α) based on xm_w and zm_w, i.e., α is a function of xm_w and zm_w (e.g., α=f3(xm_w,zm_w)).
[0075] Step s914 involves calculating M=α*A+(1−α)*MP, where M is the x-coordinate of the preferred position of the center speaker, A is the x-coordinate of the anchor point, and MP is the x-coordinate of the midpoint between the left and right speakers.
[0076] 10 is a flow chart illustrating a process 1000 according to one embodiment for rendering an audio element (e.g., a spatially heterogeneous audio element). The audio element has a range and is represented using a set of virtual speakers including a center virtual speaker. Process 1000 begins at step s1002.
[0077] Step s1002 involves selecting an anchor point for the audio element (see step s902 above).
[0078] Step s1004 includes placing a first speaker (e.g., a right speaker) and placing a second speaker (e.g., a left speaker), i.e., the positions of the first and second speakers are determined (see step s904 above).
[0079] Step s1006 includes determining a first angle (λ) between a line from the listener to the anchor point and a line from the listener to the left speaker, and a second angle (β) between a line from the listener to the anchor point and a line from the listener to the right speaker (examples of λ and β are shown in FIG. 5).
[0080] Step s1008 includes calculating a gain factor (g) using λ and β, for example, g=f4(λ,β).
[0081] Step s1010 involves processing the center speaker signal (X) with a gain factor (g) to generate a modified signal X', eg, X'=g*X.
[0082] Usage example
[0083] 11A shows an XR system 1100 to which embodiments disclosed herein can be applied. The XR system 1100 includes speakers 1104 and 1105 (which may be speakers of headphones worn by a listener) and an XR device 1110 that includes a display for displaying images to a user and is configured to be worn by the listener in some embodiments. In the illustrated XR system 1100, the XR device 1110 has a display and is designed to be worn on the user's head, and is commonly referred to as a head-mounted display (HMD).
[0084] As shown in FIG. 11B, the XR device 1110 may include an orientation detector 1101, a position detector 1102, and a processor 1103 coupled (directly or indirectly) to an audio renderer 1151 to generate output audio signals (e.g., a left audio signal 1181 for a left speaker and a right audio signal 1182 for a right speaker, as shown).
[0085] The orientation sensing unit 1101 is configured to detect changes in the listener's orientation and provide information about the detected changes to the processing unit 1103. In some embodiments, the processing unit 1103 determines an absolute orientation (with respect to some coordinate system) taking into account the detected changes in orientation detected by the orientation sensing unit 1101. Also, different systems for determining orientation and position may exist, such as systems that use lighthouse trackers (lidars). In one embodiment, the orientation sensing unit 1101 may determine an absolute orientation (with respect to a coordinate system) given the detected changes in orientation. In this case, the processing unit 1103 may simply multiplex the absolute orientation data from the orientation sensing unit 1101 with the position data from the position sensing unit 1102. In some embodiments, the orientation sensing unit 1101 may comprise one or more accelerometers and / or one or more gyroscopes.
[0086] Audio renderer 1151 generates an audio output signal based on an input audio signal 1161, metadata 1162 about the XR scene the listener is experiencing, and information 1163 about the listener's position and orientation. The metadata 1162 for the XR scene includes metadata for each object and audio element included in the XR scene, and the metadata for an object may include information about the object's dimensions. The metadata 1162 may also include control information such as reverberation time values, reverberation level values, absorption parameters, etc. Audio renderer 1151 may be a component of XR device 1110 or may be remote from XR device 1110 (e.g., audio renderer 1151 or components thereof may be implemented in the so-called "cloud").
[0087] 12 shows an example implementation of an audio renderer 1151 for generating sound for an XR scene. The audio renderer 1151 includes a controller 1201 and a signal modifier 1202 for modifying an audio signal 1161 (e.g., an audio signal of a multi-channel audio element) based on control information 1210 from the controller 1201. The controller 1201 may be configured to receive one or more parameters and trigger the signal modifier 1202 to perform a modification of the audio signal 1161 (e.g., increase or decrease a volume level) based on the received parameters. The received parameters include information 1163 regarding the listener's position and / or orientation (e.g., direction and distance relative to the audio element) and metadata 1162 regarding the audio element (e.g., range 200) in the XR scene (in some embodiments, the controller 1201 itself generates the metadata 1162). Using the metadata and position / orientation information, the controller 1201 can calculate one or more gain factors (g) (also known as attenuation factors) for audio elements in the XR scene, as described herein.
[0088] 13 shows an exemplary implementation of the signal modifier 1202 according to one embodiment. The signal modifier 1202 includes a directional mixer 1304, a gain adjuster 1306, and a speaker signal generator 1308.
[0089] The directional mixer receives an audio input 1161, which in this example includes a pair of audio signals 1301 and 1302 associated with an audio element (e.g., an audio element associated with range 200 or 700), and generates a set of k virtual speaker signals (VS1, VS2, ..., VSk) based on the audio input and control information 1391. In one embodiment, the signal for each virtual speaker may be derived, for example, by appropriate mixing of signals including the audio input 1161. For example, VS1 = α × L + β × R, where L is the input audio signal 1301, R is the input audio signal 1302, and α and β are coefficients that depend, for example, on the position of the listener relative to the audio element and the position of the virtual speaker to which VS1 corresponds.
[0090] The gain adjuster 1306 may adjust the gain of any one or more of the virtual speaker signals based on control information 1392, which may include the above-described gain factors calculated by the controller 1301. That is, for example, when the center speaker 203 is placed near another speaker (e.g., the left speaker 202 shown in FIG. 4), the controller 1301 may control the gain adjuster 1306 to adjust the gain of the virtual speaker signal relative to the center speaker 203 by providing the gain adjuster 1306 with the gain factors calculated as described above.
[0091] Using the virtual speaker signals VS1, VS2, ..., VSk, speaker signal generator 1308 generates output signals (e.g., output signal 1181 and output signal 1182) for driving speakers (e.g., headphone speakers or other speakers). In an embodiment in which the speakers are headphone speakers, speaker signal generator 1308 can perform conventional binaural rendering to generate the output signals. In an embodiment in which the speakers are not headphone speakers, speaker signal generator 1308 can perform conventional speaker panning to generate the output signals.
[0092] FIG. 14 is a block diagram of an audio rendering device 1400 for performing methods disclosed herein, according to some embodiments (e.g., audio renderer 1151 may be implemented using audio rendering device 1400). As shown in FIG. 14 , audio rendering device 1400 includes a processing circuit (PC) 1402. Processing circuit (PC) 1402 may include one or more processors (P) 1455 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.). Multiple processors may be co-located within a single enclosure or a single data center, or may be geographically distributed (i.e., audio rendering device 1400 may be a distributed computing device). Audio rendering device 1400 includes at least one network interface 1448. At least one network interface 1448 may comprise a transmitter (Tx) 1445 and a receiver (Rx) 1447 to enable transmitting and receiving data to other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 1448 is connected (directly or indirectly). In that case, the network interface 1448 may be wirelessly connected to the network 110. The network interface 1448 is connected to an antenna arrangement and a storage unit (also referred to as a “data storage system”) 1408, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 1402 includes a programmable processor, a computer-readable storage medium (CRM) 1442 may be provided. The CRM 1442 stores a computer program (CP) 1443, which includes computer-readable instructions (CRI) 1444. The CRM 1442 may be a non-transitory computer-readable storage medium, such as a magnetic medium (e.g., a hard disk), an optical medium, or a memory device (e.g., a random access memory, a flash memory).In some embodiments, the CRI 1444 of the computer program 1443, when executed by the PC 1402, is configured to cause the audio-rendering device 1400 to perform the steps described herein (e.g., steps described herein with reference to flowcharts). In other embodiments, the audio-rendering device 1400 may be configured to perform the steps described herein without the need for code. That is, for example, the PC 1402 may consist solely of one or more ASICs. Thus, features of the embodiments described herein may be implemented in hardware and / or software.
[0093] Overview of Various Embodiments
[0094] A1. A method for rendering an audio element (e.g., a spatially heterogeneous audio element), the audio element having a range and represented using a set of virtual speakers including a center virtual speaker, the method including at least one of the following steps: selecting a position of the center virtual speaker based on the position of the listener; and calculating an attenuation coefficient of the center virtual speaker.
[0095] A2. The method of embodiment A1, wherein the method comprises selecting the position of the central virtual speaker based on the position of the listener, wherein selecting the position of the central virtual speaker based on the position of the listener includes determining a first angle based on the position of the listener and i) a first endpoint of the audio element or a range determined based on the range of the audio element, or ii) a first virtual speaker; determining a second angle based on the position of the listener and i) a second endpoint of the audio element or the range, or ii) a second virtual speaker; and calculating a first coordinate Mx of the central virtual speaker using the first angle and the second angle, wherein the selected position of the central virtual speaker is identified at least in part by the calculated first coordinate.
[0096] A3. The method of embodiment A2, wherein selecting the position of the central virtual speaker based on the position of the listener further includes: determining a third angle based on the position of the listener and i) the position of a third endpoint of the audio element or of the range, or ii) a third virtual speaker; determining a fourth angle based on the position of the listener and i) the position of a fourth endpoint of the audio element or of the range, or ii) a fourth virtual speaker; and calculating a second coordinate My of the central virtual speaker using the third angle and the fourth angle, wherein the selected position is specified at least in part by the calculated first and second coordinates.
[0097] A4. A method as described in embodiment A2 or A3, wherein the step of calculating the first coordinate Mx of the central virtual speaker using the first angle and the second angle includes the steps of: determining the first coefficient α1 using the first angle and the second angle; and calculating Mx=(α1*Ax)+((1-α1)*MPx), where Ax is the coordinate used to specify the position of the determined anchor point on a line extending between the first endpoint and the second endpoint or between the first virtual speaker and the second virtual speaker, and MPx is the coordinate used to specify the midpoint of the line.
[0098] A5. The method described in embodiment A4, wherein the step of determining α1 using the first angle and the second angle includes the steps of calculating a first weight xm_w using the first angle and the second angle, calculating a second weight zm_w using the first angle and the second angle, and determining α1 based on xm_w and zm_w.
[0099] A6. The method described in embodiment A5, wherein the step of determining α1 based on xm_w and zm_w includes the steps of determining whether d*xm_w*zm_w is less than 1 when a predetermined coefficient is d, and setting α1 to 1 if d*xm_w*zm_w is not less than 1, and otherwise setting α1 to d*xm_w*zm_w.
[0100] A7. A method as described in embodiment A5 or A6, wherein when the first angle is λ and the second angle is β, the step of calculating xm_w using the first angle and the second angle includes the step of calculating xm_w=sin(λ) / sin(β) or xm_w=sin(β) / sin(λ).
[0101] A8. The method of embodiment A5, A6, or A7, wherein calculating zm_w using the first angle and the second angle includes calculating zm_w=sin((λ+β) / 2).
[0102] A9. The method described in embodiment A3, wherein the step of calculating the second coordinate My of the central virtual speaker using the third angle and the fourth angle includes the steps of: determining a second coefficient α2 using the third angle and the fourth angle; and calculating My=(α2*Ay)+((1-α2)*MPy), where Ay is the coordinate used to specify the position of the determined anchor point on a line extending between the third endpoint and the fourth endpoint (or between the third virtual speaker and the fourth virtual speaker), and MPy is the coordinate used to specify the midpoint of the line.
[0103] A10. The method includes selecting the position of the central virtual speaker based on the position of the listener, wherein the step of selecting the position of the central virtual speaker based on the position of the listener includes selecting a position between 1) a first point (e.g., a first endpoint) of the audio element or of a range determined based on the range of the audio element and a second point (e.g., a second endpoint) of the audio element or of the range, or 2) between a first virtual speaker and a second virtual speaker. selecting a position point on a line such that an angle between i) a second line extending from the position of the listener to the first point (or the first virtual speaker) and ii) a third line extending from the position of the listener to the selected position point on the first line is equal to an angle between i) a fourth line extending from the position of the listener to the second point (or the second virtual speaker) and ii) the third line.
[0104] A11. The method described in embodiment A10, wherein the step of selecting the position point includes a step of calculating the coordinate M of the position point by calculating M=(v*Re+w*Le) / (v+w), where v is the length of the second line, w is the length of the third line, Re is the coordinate of the first point or the first virtual speaker, and Le is the coordinate of the second point or the second virtual speaker.
[0105] A12. The method of embodiment A10 or A11, further comprising positioning the central virtual speaker at the selected location point.
[0106] A13. The method of any one of embodiments A1 to A12, wherein the method includes a step of calculating an attenuation coefficient of the center virtual speaker based on the position of the listener, wherein the step of calculating an attenuation coefficient of the center virtual speaker based on the position of the listener includes a step of determining a first angle based on the position of the listener and i) a position of a first endpoint of the audio element or of a range determined based on the range of the audio element, or ii) a position of a first virtual speaker; a step of determining a second angle based on the position of the listener and a position of a second endpoint of the audio element or of the range, or ii) a position of a second virtual speaker; and a step of calculating ε=sin(λ) / sin(β) or ε=sin(β) / sin(λ), where λ is the first angle, β is the second angle, and ε is the attenuation coefficient.
[0107] A14. The method of embodiment A13, further comprising the steps of modifying the center virtual speaker signal X to generate a modified center virtual speaker signal X' such that X' = ε * X, and rendering the audio element using the modified center virtual speaker signal (e.g., generating an output signal using the center virtual speaker signal).
[0108] A15. A method according to any one of embodiments A2 to A14, wherein the set of virtual speakers further includes a first virtual speaker positioned at the first endpoint and a second virtual speaker positioned at the second endpoint, and the method further comprises rendering the audio element using information identifying the positions of the virtual speakers.
[0109] A16. The method of embodiment A3 or A9, wherein the set of virtual speakers further includes a first virtual speaker positioned at the first endpoint (e.g., a first corner point, or a first point that is the midpoint between a first pair of corner points), a second virtual speaker positioned at the second endpoint (e.g., a second corner point, or a second point that is the midpoint between another pair of corner points), a third virtual speaker positioned at the third endpoint (e.g., a third corner point, or a third point that is the midpoint between another pair of corner points), and a fourth virtual speaker positioned at the fourth endpoint (e.g., a fourth corner point, or a fourth point that is the midpoint between another pair of corner points), and the method further comprises a step of rendering the audio element using information identifying the positions of the virtual speakers.
[0110] A17. The method of embodiment A1, wherein the method includes selecting the position of the central virtual speaker based on the position of the listener, and selecting the position of the central virtual speaker based on the position of the listener includes determining a distance from the position of the listener to a position of the audio element, determining whether the determined distance is greater than a threshold, and selecting a position of the midpoint of the audio element or the midpoint of a range determined based on the range of the audio element as a result of determining that the determined distance is greater than the threshold.
[0111] A18. The method of embodiment A1, wherein the method includes selecting the position of the central virtual speaker based on the position of the listener, and selecting the position of the central virtual speaker based on the position of the listener includes: i) obtaining listener information indicating coordinates (e.g., x-coordinate) of the listener; ii) obtaining midpoint information indicating coordinates (e.g., x-coordinate) of a midpoint between a first point associated with the audio element (e.g., a first endpoint of a range associated with the audio element) and a second point associated with the audio element (e.g., a second endpoint of the range); and iii) selecting the position of the central virtual speaker based on the midpoint information and the listener information.
[0112] A19. The method described in embodiment A18, wherein the step of selecting the position of the central virtual speaker based on the midpoint information and the listener information includes the steps of: i) determining coordinates of an anchor point; and ii) selecting the position of the central virtual speaker based on the midpoint information and anchor information indicating the coordinates of the determined anchor point.
[0113] A20. A method as described in embodiment A19, wherein i) the midpoint information includes a midpoint value MP indicating the coordinates of the midpoint, ii) the anchor information includes an anchor value A indicating the coordinates of the anchor point, and iii) the step of selecting the position of the central virtual speaker based on the midpoint information and the anchor information includes a step of calculating a coordinate value M of the central speaker using MP and A.
[0114] A21. The method of embodiment A20, wherein the anchor value A depends on the indicated coordinate of the listener.
[0115] A22. The method of embodiment A21, wherein A=L, where L is the indicated coordinate of the listener (as shown in Figures 5 and 6, a coordinate system may be defined for the range such that the range extends along the x-axis).
[0116] A23. The method of embodiment A20, A21, or A22, wherein calculating M using MP and A includes calculating M=α*A+(1-α)*MP, where α is a coefficient that depends on the indicated coordinate of the listener.
[0117] A24. A method according to any one of embodiments A18 to A23, wherein the listener information further indicates a second coordinate (e.g., a y coordinate) of the listener, the midpoint information further indicates a second coordinate (e.g., a y coordinate) of the midpoint, and the step of selecting the position of the central virtual speaker based on the midpoint information and the listener information further includes a step of determining the second coordinate of the central virtual speaker based on the second coordinate of the listener and the second coordinate of the midpoint.
[0118] A25. A method according to any one of embodiments A10 or A18 to A24, wherein a first virtual speaker is placed at the first point and a second virtual speaker is placed at the second point, and the method further comprises rendering the audio element using information identifying the positions of the virtual speakers.
[0119] A26. A method according to any one of embodiments A1 to A25, further comprising: generating a center virtual speaker signal for the center virtual speaker based on the position of the center virtual speaker; and rendering the audio element using the center virtual speaker signal (e.g., generating an output signal using the center virtual speaker signal).
[0120] B1. A computer program comprising instructions that, when executed by processing circuitry of an audio renderer, cause the audio renderer to perform a method according to any one of the preceding embodiments.
[0121] B2. A carrier containing said computer program, said carrier being one of an electrical signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0122] C1. An audio rendering device configured to perform the method of any one of the preceding embodiments.
[0123] C2. The audio rendering device of embodiment C1, wherein the audio rendering device comprises a memory and a processing circuit coupled to the memory.
[0124] While various embodiments have been described herein, it should be understood that they are presented by way of example only, and not limitation. Thus, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments. Moreover, any combination of the above-described objects in all possible variations thereof is encompassed by the present disclosure unless otherwise indicated herein or clearly contradicted by context.
[0125] Additionally, while the processes described above and illustrated in the figures are shown as a series of steps, this is done for illustrative purposes only, and as such, steps may be added, steps may be omitted, the order of steps may be rearranged, and steps may be performed in parallel.
[0126] References
[0127] [1] MPEG-H 3D Audio, Section 8.4.4.7: "Spreading"
[0128] [2] MPEG-H 3D Audio, Section 18.1: "Element Metadata Preprocessing"
[0129] [3] MPEG-H 3D Audio, Section 18.11: "Diffuseness Rendering"
[0130] [4] EBU ADM Renderer Tech 3388, Section 7.3.6: "Divergence"
[0131] [5]EBU ADM Renderer Tech 3388, Clause 7.4: "Decorrelation Filters"
[0132] [6]EBU ADM Renderer Tech 3388, Clause 7.3.7: "Extent Panner"
[0133] [7]Efficient HRTF-based Spatial Audio for Area and Volumetric Sources, IEEE Transactions on Visualization and Computer Graphics 22(4):1 - 1 January 2016
[0134] [8]Patent Publication WO2020144062 "Efficient spatially - heterogeneous audio elements for Virtual Reality"
[0135] [9]Patent Publication WO2021180820 "Rendering of Audio Objects with a Complex Shape"
Claims
1. A method (800) for rendering an audio element, the audio element having a range (200) and represented using a set of virtual speakers (202, 203, 204) including a central virtual speaker (203), the method comprising: The method includes (s802) selecting a position for the center virtual speaker based on a position of a listener and / or calculating an attenuation factor for the center virtual speaker.
2. The method may include selecting the position of the central virtual speaker based on the position of the listener, wherein selecting the position of the central virtual speaker based on the position of the listener comprises: determining a first angle based on the position of the listener and the position of i) a first endpoint of the audio element or of a range determined based on the range of the audio element, or ii) a first virtual speaker; determining a second angle based on the position of the listener and i) a second endpoint of the audio element or of the range, or ii) a position of a second virtual speaker; calculating a first coordinate Mx of the central virtual speaker using the first angle and the second angle, wherein the selected position of the central virtual speaker is specified at least in part by the calculated first coordinate; The method of claim 1 , comprising:
3. selecting the position of the central virtual speaker based on the position of the listener, determining a third angle based on the position of the listener and i) a position of the audio element or of a third endpoint of the range, or ii) a third virtual speaker; determining a fourth angle based on the position of the listener and the position of i) a fourth endpoint of the audio element or of the range, or ii) a fourth virtual speaker; calculating a second coordinate My of the central virtual speaker using the third angle and the fourth angle, wherein the selected position is specified at least in part by the calculated first and second coordinates; The method of claim 2 further comprising:
4. The step of calculating the first coordinate Mx of the central virtual speaker using the first angle and the second angle includes: determining the first coefficient α1 using the first angle and the second angle; a step of calculating Mx=(α1*Ax)+((1-α1)*MPx), where Ax is a coordinate used when specifying the position of the determined anchor point on a line extending between the first endpoint and the second endpoint, or between the first virtual speaker and the second virtual speaker, and MPx is a coordinate used when specifying the midpoint of the line; The method of claim 2 or 3, comprising:
5. The step of determining α1 using the first angle and the second angle includes:
5. The method of claim 4, comprising: calculating a first weight xm_w using the first angle and the second angle; calculating a second weight zm_w using the first angle and the second angle; and determining α1 based on xm_w and zm_w.
6. The step of determining α1 based on xm_w and zm_w is determining whether d*xm_w*zm_w is less than 1, where d is a predetermined coefficient; setting α1 to 1 if d*xm_w*zm_w is not less than 1, otherwise setting α1 to d*xm_w*zm_w; The method of claim 5 , comprising:
7. When the first angle is λ and the second angle is β, The step of calculating xm_w using the first angle and the second angle includes: xm_w=sin(λ) / sin(β) or xm_w=sin(β) / sin(λ) 7. The method of claim 5, further comprising the step of calculating:
8. The step of calculating zm_w using the first angle and the second angle includes the steps of: zm_w=sin((λ+β) / 2) 8. The method of claim 5, 6 or 7, comprising the step of calculating:
9. The step of calculating the second coordinate My of the central virtual speaker using the third angle and the fourth angle includes: determining a second coefficient α2 using the third angle and the fourth angle; When a coordinate used when specifying the position of the determined anchor point on a line extending between the third endpoint and the fourth endpoint or between the third virtual speaker and the fourth virtual speaker is denoted by Ay, and a coordinate used when specifying the midpoint of the line is denoted by MPy, My=(α2*Ay)+((1-α2)*MPy) and calculating The method of claim 3, comprising:
10. The method may include selecting the position of the central virtual speaker based on the position of the listener, wherein selecting the position of the central virtual speaker based on the position of the listener comprises: selecting a position point on a first line between 1) a first point of the audio element or of a range determined based on the range of the audio element and a second point of the audio element or of the range, or 2) between a first virtual speaker and a second virtual speaker, selecting the location point such that an angle between i) a second line extending from the position of the listener to the first point or the first virtual speaker, and ii) a third line extending from the position of the listener to the selected location point on the first line, is equal to an angle between i) a fourth line extending from the position of the listener to the second point or the second virtual speaker, and ii) the third line; The method of claim 1 , comprising:
11. The step of selecting a location point comprises: The length of the second straight line is v, The length of the third straight line is w, The coordinates of the first point or the first virtual speaker are Re, The coordinates of the second point or the second virtual speaker are Le, When M=(v*Re+w*Le) / (v+w), 11. The method of claim 10, comprising calculating a coordinate M of the location point by calculating:
12. The method of claim 10 or 11, further comprising the step of positioning the central virtual speaker at the selected location point.
13. The method may include calculating an attenuation factor for the center virtual speaker based on the position of the listener, the step of calculating an attenuation factor for the center virtual speaker based on the position of the listener comprising: determining a first angle based on the position of the listener and i) a position of a first endpoint of the audio element or of a range determined based on the range of the audio element, or ii) a position of a first virtual speaker; determining a second angle based on the position of the listener and the position of the audio element or of a second endpoint of the range, or ii) the position of a second virtual speaker; Calculating ε=sin(λ) / sin(β) or ε=sin(β) / sin(λ), where λ is the first angle, β is the second angle, and ε is the damping coefficient; 13. The method of any one of claims 1 to 12, comprising:
14. 14. The method of claim 13, further comprising modifying the center virtual speaker signal X to generate a modified center virtual speaker signal X′ such that X′=ε*X, and rendering the audio element using the modified center virtual speaker signal.
15. The set of virtual speakers may include: a first virtual speaker located at the first endpoint and a second virtual speaker located at the second endpoint, the method further comprising rendering the audio element using information identifying the positions of the virtual speakers.
15. The method of any one of claims 2 to 14.
16. The set of virtual speakers may include: a first virtual speaker located at the first endpoint, a second virtual speaker located at the second endpoint, a third virtual speaker located at the third endpoint, and a fourth virtual speaker located at the fourth endpoint, the method further comprising rendering the audio element using information identifying the positions of the virtual speakers.
10. The method of claim 3 or 9.
17. The method may include selecting the position of the central virtual speaker based on the position of the listener, wherein selecting the position of the central virtual speaker based on the position of the listener comprises: determining a distance from the position of the listener to a position of the audio element; determining whether the determined distance is greater than a threshold; selecting a location of the midpoint of the audio element or the midpoint of a range determined based on the range of the audio element as a result of determining that the determined distance is greater than the threshold; The method of claim 1 , comprising:
18. The method may include selecting the position of the central virtual speaker based on the position of the listener, wherein selecting the position of the central virtual speaker based on the position of the listener comprises: i) obtaining listener information indicating the coordinates of the listener; ii) obtaining midpoint information indicating coordinates of a midpoint between a first point associated with the audio element and a second point associated with the audio element; iii) selecting the location of the center virtual speaker based on the midpoint information and the listener information; The method of claim 1 , comprising:
19. selecting the position of the center virtual speaker based on the midpoint information and the listener information, i) determining the coordinates of the anchor points; ii) selecting the position of the central virtual speaker based on the midpoint information and anchor information indicating the coordinates of the determined anchor point; 20. The method of claim 18, comprising:
20. 20. The method of claim 19, wherein i) the midpoint information includes a midpoint value MP indicating the coordinates of the midpoint, ii) the anchor information includes an anchor value A indicating the coordinates of the anchor point, and iii) selecting the position of the central virtual speaker based on the midpoint information and the anchor information includes calculating a coordinate value M of the central speaker using MP and A.
21. The method of claim 20 , wherein the anchor value A depends on the indicated coordinate of the listener.
22. 22. The method of claim 21, wherein A=L, where L is the indicated coordinate of the listener.
23. 23. The method of claim 20, 21, or 22, wherein calculating M using MP and A comprises calculating M=α*A+(1−α)*MP, where α is a coefficient that depends on the indicated coordinate of the listener.
24. 24. The method of claim 18, wherein the listener information further indicates a second coordinate of the listener, the midpoint information further indicates a second coordinate of the midpoint, and wherein selecting the position of the central virtual speaker based on the midpoint information and the listener information further comprises determining a second coordinate of the central virtual speaker based on the second coordinate of the listener and the second coordinate of the midpoint.
25. a first virtual speaker is placed at the first point, a second virtual speaker is placed at the second point, 25. The method of claim 10 or any one of claims 18 to 24, wherein the method further comprises rendering the audio element using information identifying the position of the virtual speaker.
26. generating a center virtual speaker signal for the center virtual speaker based on the position of the center virtual speaker; rendering the audio element using the center virtual speaker signal; 26. The method of any one of claims 1 to 25, further comprising:
27. A computer program (1443) comprising instructions (1444) that, when executed by a processing circuit (1402) of an audio renderer (1400), causes the audio renderer to perform the method of any one of claims 1 to 26.
28. 28. A carrier containing the computer program of claim 27, said carrier being one of an electrical signal, an optical signal, a radio signal, and a computer readable storage medium (1442).
29. An audio rendering device (1400) for rendering an audio element, the audio element having a range (200) and represented using a set of virtual speakers (202, 203, 204) including a central virtual speaker (203), 8. An audio rendering apparatus configured to: select a position of the central virtual speaker based on a position of a listener; and / or calculate an attenuation coefficient for the central virtual speaker (s802).
30. the audio rendering device, determining a first angle based on the position of the listener and the position of i) a first endpoint of the audio element or of a range determined based on the range of the audio element, or ii) a first virtual speaker; determining a second angle based on the position of the listener and i) a second endpoint of the audio element or of the range, or ii) a position of a second virtual speaker; calculating a first coordinate Mx of the central virtual speaker using the first angle and the second angle, wherein the selected position of the central virtual speaker is specified at least in part by the calculated first coordinate; 30. The audio rendering apparatus of claim 29, configured to select the position of the central virtual speaker based on the position of the listener by:
31. selecting the position of the central virtual speaker based on the position of the listener, determining a third angle based on the position of the listener and i) a position of the audio element or of a third endpoint of the range, or ii) a third virtual speaker; determining a fourth angle based on the position of the listener and the position of i) a fourth endpoint of the audio element or of the range, or ii) a fourth virtual speaker; calculating a second coordinate My of the central virtual speaker using the third angle and the fourth angle, wherein the selected position is specified at least in part by the calculated first and second coordinates; 31. The audio rendering device of claim 30, further comprising:
32. The step of calculating the first coordinate Mx of the central virtual speaker using the first angle and the second angle includes: determining the first coefficient α1 using the first angle and the second angle; a step of calculating Mx=(α1*Ax)+((1-α1)*MPx), where Ax is a coordinate used when specifying the position of the determined anchor point on a line extending between the first endpoint and the second endpoint, or between the first virtual speaker and the second virtual speaker, and MPx is a coordinate used when specifying the midpoint of the line; 32. An audio rendering device according to claim 30 or 31, comprising:
33. The step of determining α1 using the first angle and the second angle includes:
33. The audio rendering apparatus of claim 32, further comprising: calculating a first weight xm_w using the first angle and the second angle; calculating a second weight zm_w using the first angle and the second angle; and determining α1 based on xm_w and zm_w.
34. The step of determining α1 based on xm_w and zm_w is determining whether d*xm_w*zm_w is less than 1, where d is a predetermined coefficient; setting α1 to 1 if d*xm_w*zm_w is not less than 1, otherwise setting α1 to d*xm_w*zm_w; 34. An audio rendering apparatus according to claim 33, comprising:
35. When the first angle is λ and the second angle is β, The step of calculating xm_w using the first angle and the second angle includes: xm_w=sin(λ) / sin(β) or xm_w=sin(β) / sin(λ) 35. An audio rendering apparatus according to claim 33 or 34, comprising the step of calculating:
36. The step of calculating zm_w using the first angle and the second angle includes the steps of: zm_w=sin((λ+β) / 2) 36. An audio rendering apparatus according to claim 33, 34 or 35, comprising the step of calculating:
37. The step of calculating the second coordinate My of the central virtual speaker using the third angle and the fourth angle includes: determining a second coefficient α2 using the third angle and the fourth angle; When a coordinate used when specifying the position of the determined anchor point on a line extending between the third endpoint and the fourth endpoint or between the third virtual speaker and the fourth virtual speaker is denoted by Ay, and a coordinate used when specifying the midpoint of the line is denoted by MPy, My=(α2*Ay)+((1-α2)*MPy) and calculating 32. An audio rendering apparatus according to claim 31, comprising:
38. selecting a position point on a first line between 1) a first point of the audio element or of a range determined based on the range of the audio element and a second point of the audio element or of the range, or 2) between a first virtual speaker and a second virtual speaker, selecting the location point such that an angle between i) a second line extending from the position of the listener to the first point or the first virtual speaker, and ii) a third line extending from the position of the listener to the selected location point on the first line, is equal to an angle between i) a fourth line extending from the position of the listener to the second point or the second virtual speaker, and ii) the third line; 30. The audio rendering apparatus of claim 29, configured to select the position of the central virtual speaker based on the position of the listener by performing:
39. The step of selecting a location point comprises: The length of the second straight line is v, The length of the third straight line is w, The coordinates of the first point or the first virtual speaker are Re, The coordinates of the second point or the second virtual speaker are Le, When M=(v*Re+w*Le) / (v+w), 39. An audio rendering apparatus according to claim 38, comprising calculating a coordinate M of the location point by calculating:
40. 40. An audio rendering apparatus according to claim 38 or 39, wherein the audio rendering apparatus is further configured to position the central virtual speaker at the selected location point.
41. determining a first angle based on the position of the listener and i) a position of a first endpoint of the audio element or of a range determined based on the range of the audio element, or ii) a position of a first virtual speaker; determining a second angle based on the position of the listener and the position of the audio element or of a second endpoint of the range, or ii) the position of a second virtual speaker; Calculating ε=sin(λ) / sin(β) or ε=sin(β) / sin(λ), where λ is the first angle, β is the second angle, and ε is the damping coefficient; 41. An audio rendering apparatus according to any one of claims 29 to 40, configured to calculate the attenuation factor of the central virtual speaker based on the position of the listener by:
42. 42. The audio rendering apparatus of claim 41, further configured to modify the center virtual speaker signal X to generate a modified center virtual speaker signal X′ such that X′=ε*X, and to render the audio element using the modified center virtual speaker signal.
43. The set of virtual speakers may include:
43. An audio rendering apparatus according to any one of claims 30 to 42, further comprising a first virtual speaker located at the first endpoint and a second virtual speaker located at the second endpoint, the method further comprising the step of rendering the audio element using information identifying the positions of the virtual speakers.
44. The set of virtual speakers may include: a first virtual speaker located at the first endpoint, a second virtual speaker located at the second endpoint, a third virtual speaker located at the third endpoint, and a fourth virtual speaker located at the fourth endpoint, the method further comprising rendering the audio element using information identifying the positions of the virtual speakers.
38. An audio rendering apparatus according to claim 31 or 37.
45. determining a distance from the position of the listener to a position of the audio element; determining whether the determined distance is greater than a threshold; selecting a location of the midpoint of the audio element or the midpoint of a range determined based on the range of the audio element as a result of determining that the determined distance is greater than the threshold; 30. The audio rendering apparatus of claim 29, configured to select the position of the central virtual speaker based on the position of the listener by performing:
46. i) obtaining listener information indicating the coordinates of the listener; ii) obtaining midpoint information indicating coordinates of a midpoint between a first point associated with the audio element and a second point associated with the audio element; iii) selecting the location of the center virtual speaker based on the midpoint information and the listener information; 30. The audio rendering apparatus of claim 29, configured to select the position of the central virtual speaker based on the position of the listener by performing:
47. selecting the position of the center virtual speaker based on the midpoint information and the listener information, i) determining the coordinates of the anchor points; ii) selecting the position of the central virtual speaker based on the midpoint information and anchor information indicating the coordinates of the determined anchor point; 47. An audio rendering apparatus according to claim 46, comprising:
48. 48. The audio rendering apparatus of claim 47, wherein i) the midpoint information includes a midpoint value MP indicating the coordinates of the midpoint, ii) the anchor information includes an anchor value A indicating the coordinates of the anchor point, and iii) selecting the position of the central virtual speaker based on the midpoint information and the anchor information includes calculating a coordinate value M of the central speaker using MP and A.
49. 49. An audio rendering apparatus according to claim 48, wherein the anchor value A depends on the indicated coordinates of the listener.
50. 50. An audio rendering apparatus as claimed in claim 49, wherein A=L, where L is the indicated coordinate of the listener.
51. 51. An audio rendering apparatus as claimed in claim 48, 49 or 50, wherein the step of calculating M using MP and A comprises the step of calculating M = α * A + (1 - α) * MP, where α is a coefficient that depends on the indicated coordinate of the listener.
52. 52. The audio rendering apparatus of claim 46, wherein the listener information further indicates a second coordinate of the listener, the midpoint information further indicates a second coordinate of the midpoint, and wherein selecting the position of the central virtual speaker based on the midpoint information and the listener information further comprises determining the second coordinate of the central virtual speaker based on the second coordinate of the listener and the second coordinate of the midpoint.
53. a first virtual speaker is placed at the first point, a second virtual speaker is placed at the second point, 52. An audio rendering apparatus according to any one of claims 38 or 46 to 51, wherein the method further comprises rendering the audio element using information identifying the position of the virtual speaker.
54. generating a center virtual speaker signal for the center virtual speaker based on the position of the center virtual speaker; Rendering the audio element using the center virtual speaker signal.
54. An audio rendering apparatus according to any one of claims 29 to 53, further configured to: