A three-dimensional spatial sound playback method for layered loudspeaker arrays

CN120786280BActive Publication Date: 2026-09-18SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511231233.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-30
Publication Date
2026-09-18
Estimated Expiration
2045-08-30

AI Technical Summary

Technical Problem

[0006]针对目前Ambisonics三维重放在多层扬声器阵列中的稳定性与效率问题,本发明提出一种面向分层扬声器阵列的三维空间声重放方法,本发明使用两层物理扬声器阵列,通过二维Ambisonics合成一对位于同一方位角,不同仰角的虚拟源;然后,通过矢量调幅的方式调整两个虚拟源的相对强度,合成为目标虚拟源,目的在于降低算法对扬声器空间分布均匀性的要求,同时,在保证高仰角声源定位效果的前提下,尽可能提高具有实际扬声器的区域,特别是水平面区域的重放精度

Benefits of technology

[0030] 1. By using layered decoding, the requirements for the number and uniform distribution of speakers on non-horizontal surfaces in 3D Ambisonics playback are reduced, which better meets the needs of actual speaker arrays;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120786280B_ABST
    Figure CN120786280B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional space sound playback method for layered loudspeaker arrays, comprising the following steps: presetting the spatial coordinates of each unit of the loudspeaker array and a target virtual sound source; layering the loudspeaker array according to the elevation angle, calculating a two-dimensional Ambisonics decoding matrix for each layer of the loudspeaker array, combining two adjacent loudspeaker layers, and calculating an amplitude modulation inverse matrix; calculating the amplitude modulation coefficients of the two combined loudspeaker layers according to the elevation angle of the target virtual sound source, and selecting effective loudspeaker layers for playback; synthesizing two virtual sources in the effective loudspeaker layers according to the azimuth angle of the target sound source, and combining the horizontal azimuth angle decoding and the amplitude modulation coefficients to synthesize a three-dimensional virtual sound source. The application can improve the playback accuracy of the area with actual loudspeakers, especially the horizontal plane area, and reduce the requirements of the algorithm on the number of non-horizontal plane loudspeakers and the spatial uniformity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional spatial sound reproduction technology, specifically to a three-dimensional spatial sound reproduction method for layered speaker arrays, which optimizes the spatial resolution and timbre of the reproduction in the horizontal plane and the high elevation angle direction. Background Technology

[0002] Spatial sound reproduction technology based on loudspeaker arrays aims to preserve the directional sense of virtual sound sources to a certain extent using a limited number of loudspeakers. The most common spatial sound reproduction technologies are two-channel stereo and 5.1-channel surround sound, both of which can only reproduce virtual sound sources in a horizontal plane. In recent years, spatial sound technology has gradually expanded from two-dimensional to three-dimensional reproduction. Representative three-dimensional reproduction technologies include three-dimensional vector base amplitude modulation (VBAP), Ambisonics, and wave field synthesis. Among these, Ambisonics technology uses spherical harmonic functions to encode and decode spatial sound sources, achieving a relatively good balance between reproduction accuracy and the required number of loudspeakers.

[0003] Ambisonics technology first decomposes the sound field, composed of spatial plane waves or spherical waves, into spherical harmonic coefficients, obtaining the Nth-order spherical harmonic coefficients. Then, using sound field matching, based on the spatial position of the speaker array, these coefficients are decoded back into speaker gain coefficients. To ensure decoding stability, the number of speakers must not be less than the number of spherical harmonics, and the speakers should be spatially uniformly distributed. The order N of the spherical harmonics is directly related to the spatial resolution and timbre distortion of the reproduced sound source.

[0004] In a two-dimensional sound field, the Nth-order spherical harmonic function (which is then an angular harmonic function) contains 2N+1 harmonic components, requiring a relatively small number of loudspeakers. However, for a three-dimensional sound field, the Nth-order spherical harmonic function contains (N+1) harmonic components. 2 The number of loudspeakers required for reproducing each harmonic component increases exponentially with the order N. Furthermore, three-dimensional sound field reproduction requires loudspeakers to be uniformly distributed throughout the entire space. These requirements contradict common spatial loudspeaker arrays in practical applications (such as the NHK 22.2 system) and the perceptual characteristics of the auditory system: ① Practical systems typically deploy a larger number of loudspeaker arrays on the horizontal plane and only a small number at other heights; ② Considering the difficulty of loudspeaker installation, practical loudspeaker arrays are usually arranged in layers, such as several layers of two-dimensional loudspeaker arrays at specific elevation angles like 30° and 60° on the horizontal plane; ③ The auditory system is most sensitive to the direction of sound sources on the horizontal plane and has a relatively poor ability to distinguish the direction of sound sources at other heights.

[0005] Several methods have been attempted to address the problem caused by uneven spatial distribution of loudspeakers, such as the mixed-order Ambisonics method (Ahrens J., Spors S..An Analytical Approach to SoundField Reproduction Using Circular and Spherical Loudspeaker Distributions[J].Acta Acustica united with Acustica,2008,94(6):988-999.). This method improves the accuracy of two-dimensional reproduction by using more loudspeakers in the horizontal plane based on three-dimensional reproduction. However, this method can only improve the reproduction accuracy in the horizontal plane, and still has extremely high requirements for the number of arrays and the uniformity of spatial distribution in the high elevation direction. The All-Round Ambisonics Decoding algorithm (Zotter F., Frank M. All-Round Ambisonic Panning and Decoding[J]. Journal of the Audio Engineering Society, 2012, 60(10): 807-820.) combines VBAP and the Ambisonics algorithm to reduce the requirements for the three-dimensional layout of loudspeakers, but it does not improve the playback accuracy of specific areas (such as the horizontal plane). In areas with a large number of loudspeakers, it may even reduce the playback accuracy. In summary, there is currently no method that can improve the playback accuracy of areas with actual physical loudspeakers (including but not limited to the horizontal plane) while ensuring a certain vertical playback accuracy. In particular, loudspeaker arrays with common hierarchical structure characteristics have not received much attention from researchers, and there are no three-dimensional spatial sound playback algorithms for such arrays. Due to the above problems, three-dimensional spatial sound playback technology is difficult to deploy in actual loudspeaker array systems, and the playback effect is often poor. This invention proposes a three-dimensional spatial sound reproduction algorithm suitable for loudspeaker arrays with layered structure characteristics. While ensuring a certain high elevation angle sound source localization effect, it maximizes the reproduction accuracy of areas with actual loudspeakers, especially horizontal areas, and reduces the algorithm's requirements for the number of non-horizontal loudspeakers and spatial uniformity. Summary of the Invention

[0006] To address the stability and efficiency issues of current Ambisonics 3D reproduction in multi-layer speaker arrays, this invention proposes a 3D spatial sound reproduction method for layered speaker arrays. This invention uses two layers of physical speaker arrays and synthesizes a pair of virtual sources located at the same azimuth angle but different elevation angles using 2D Ambisonics. Then, the relative intensity of the two virtual sources is adjusted by vector amplitude modulation to synthesize a target virtual source. The aim is to reduce the algorithm's requirement for uniformity of speaker spatial distribution, while maximizing reproduction accuracy in areas with actual speakers, especially in the horizontal plane, while ensuring effective localization of high-elevation sound sources.

[0007] The present invention is achieved by at least one of the following technical solutions.

[0008] A three-dimensional spatial sound reproduction method for layered loudspeaker arrays includes the following steps:

[0009] A1. Set the spatial coordinates of each unit in the speaker array and set the spatial coordinates of the target virtual sound source;

[0010] A2. Divide the loudspeaker array into layers according to elevation angle, and calculate the two-dimensional Ambisonics decoding matrix based on the number of loudspeaker arrays in each layer and azimuth angle; combine loudspeaker layers in pairs according to the adjacent principle, and calculate the amplitude modulation inverse matrix;

[0011] A3. Based on the elevation angle of the target virtual sound source, calculate the vector amplitude modulation coefficient of the speaker layers in pairs, select the pairs with non-negative coefficients as the effective speaker layers, and retain the corresponding normalized amplitude modulation coefficients.

[0012] A4. Based on the azimuth angle of the target sound source, use two-dimensional Ambisonics to synthesize virtual sources on the effective loudspeaker layer, and use amplitude modulation coefficients to adjust the relative intensity of the two virtual sources to synthesize a three-dimensional virtual source.

[0013] Further, in step A2, all speakers are divided into L layers according to their elevation angle, where the l-th layer contains M layers. l There are 1 loudspeaker, and the coordinates of each loudspeaker are (θ). lm ,φ l ),l∈[1,L],m∈[1,M l ], φ l θ represents the elevation angle of the speaker in the l-th layer. lm This represents the azimuth angle of the m-th speaker in the l-th layer; K is calculated for each layer separately. l Two-dimensional Ambisonics decoding matrix D l ,in The method for calculating the decoding matrix corresponding to the l-th layer is as follows: (This indicates rounding down.)

[0014]

[0015] in, For K l The angular harmonic functions of order i, i = -1, i = 1 and i = 0 correspond to the sine term, cosine term and 0th order angular harmonic function, respectively.

[0016] Further, in step A2, the L layers of speaker units are combined in pairs according to their adjacent elevation angles, forming L-1 speaker unit combinations {(φ l ,φ l+1 )|1≤l≤L-1},φ l This indicates the elevation angle of the speaker on the l-th floor.

[0017] Further, in step A3, after determining the orientation of the target virtual sound source S0, the modulation index (φ) is calculated for all L-1 speaker layer combinations. l ,φ l+1 The modulation index of the 1st floor loudspeaker is:

[0018]

[0019] Among them (G′ l ,G′ l+1 ) is the speaker layer (φ l ,φ l+1 The amplitude modulation coefficient, φ s0 The elevation angle φ of the target virtual sound source S0 l This indicates the elevation angle of the speaker on the l-th floor.

[0020] Furthermore, in step A3, if the amplitude modulation coefficients corresponding to layers l and l+1 are both greater than 0, then layers l and l+1 are effective speaker layers, and the amplitude modulation coefficients are normalized to obtain the normalized gain.

[0021] Further, in step A4, two-dimensional Ambisonics playback is performed using the speakers of the effective speaker layer to generate a virtual source; the gain of each speaker in the effective speaker layer is calculated based on the decoding matrix corresponding to the effective speaker layer.

[0022] Further, in step A4, all speakers in the effective speaker layer are activated, and the target virtual sound source is synthesized using the virtual source of the effective speaker layer. The gain of the m-th speaker in the synthesized effective speaker layer is the product of the gain of the m-th speaker before synthesis and the gain after the amplitude modulation coefficient of this effective speaker layer is normalized.

[0023] Furthermore, in a speaker layer containing a single ceiling speaker, if the ceiling speaker is activated, the ceiling speaker is directly treated as the virtual sound source itself for gain allocation.

[0024] Furthermore, if the units in the activated speaker layer are not arranged in a ring, amplitude and delay alignment is performed in advance during playback.

[0025] The system for implementing the aforementioned three-dimensional spatial sound reproduction method for a layered loudspeaker array includes:

[0026] A layered speaker array, the speaker array containing at least two speaker layers;

[0027] The signal processing module is configured to implement the entire processing flow of the method.

[0028] The dynamic setting module, configured to implement the aforementioned process, can set the spatial coordinates of the sound source in real time and dynamically adjust the gain weights of each layer of speakers.

[0029] Compared with existing technologies, the beneficial effects of the present invention are as follows:

[0030] 1. By using layered decoding, the requirements for the number and uniform distribution of speakers on non-horizontal surfaces in 3D Ambisonics playback are reduced, which better meets the needs of actual speaker arrays;

[0031] 2. In areas with actual loudspeakers, especially in horizontal areas, minimize the energy dispersion of virtual sound sources to improve positioning accuracy and reduce timbre distortion;

[0032] 3. The speaker array can be flexibly designed according to the actual signal source requirements, and the playback spatial accuracy of different areas can be adjusted as needed. Attached Figure Description

[0033] Figure 1 This is a flowchart illustrating a three-dimensional spatial sound reproduction method for a layered loudspeaker array, as an example. Figure 2 This is a schematic diagram of the synthesis of the coordinate system and the 3D virtual source;

[0034] Figure 3 This is a diagram illustrating the sound field reconstruction effect of the high-elevation-angle speaker layer in the embodiment.

[0035] Figure 4a This is a diagram showing the results of the subjective positioning experiment in the embodiment.

[0036] Figure 4b The results of the subjective evaluation experiment on timbre were presented. Detailed Implementation

[0037] The technical solutions and advantages of the present invention will be further described below with reference to the embodiments and accompanying drawings, but the implementation of the present invention is not limited thereto.

[0038] like Figure 1 As shown, a three-dimensional spatial sound reproduction method for a layered loudspeaker array according to this embodiment includes the following steps:

[0039] A1. Data Input. Both the loudspeakers and the virtual source are in a spherical coordinate system. Preset the spatial coordinates (θ, φ) of each unit in the loudspeaker array, and set the spatial coordinates (θ, φ) of the target virtual sound source S0. s0 ,φ s0 ), where θ is the azimuth angle of the speaker and φ is the elevation angle of the speaker. Set the spatial coordinates (θ...) of the target virtual sound source S0. s0 ,φ s0 ), θ s0 It is the azimuth angle of the sound source, φ s0 This refers to the elevation angle of the sound source. The loudspeaker array should have spatial layering, meaning that multiple loudspeakers have the same or similar elevation angles.

[0040] A2. Pre-calculation. Group the loudspeaker arrays according to their elevation angles, and calculate the highest-order two-dimensional Ambisonics decoding matrix based on the number of loudspeaker arrays in each layer and their azimuth angles. Based on the elevation angle of each loudspeaker layer, combine the loudspeaker layers in pairs according to the adjacent principle, and calculate the amplitude modulation inverse matrix.

[0041] In step A2, all speakers are divided into L layers according to their elevation angle, and the l-th layer contains M. l There are [number] speakers, and the coordinates of each speaker on this layer are (θ). lm ,φ l ),l∈[1,L],m∈[1,M l ], where the elevation angle φ of the loudspeaker in the l-th layer is l Arrange them in ascending order of their numerical values ​​(i.e., φ) l+1 >φ l ). Calculate K for each layer. l Two-dimensional Ambisonics decoding matrix, Calculate the decoding matrices for all layers, and the decoding matrix D for the l-th layer. l for:

[0042]

[0043] in, For K l The angular harmonic function of order n. More generally, the angular harmonic function of the i-th term of order n. The definition of is:

[0044]

[0045] Specifically, if the l-th layer has only one speaker unit (such as a ceiling speaker), then step A2 is skipped. Furthermore, the decoding order of the l-th layer described in this invention is not limited to the highest order K. l The appropriate number should be selected flexibly based on the actual array design, and should not exceed the highest order K. l And while maintaining stability, try to use higher-order decoding.

[0046] At the same time, they are combined in pairs according to the principle of adjacent elevation angles to form L-1 speaker layer combinations {(φ l ,φ l+1 For each combination, calculate the amplitude modulation inverse matrix: |1≤l≤L-1}.

[0047]

[0048] A3. Two-dimensional VBAP: Based on the elevation angle of the target virtual sound source, calculate the vector amplitude modulation coefficients of the paired speaker layers. Select the pairs of combinations with non-negative coefficients as the effective speaker layers lay1 and lay2, and retain the corresponding normalized amplitude modulation coefficient G. lay1 G lay2 .

[0049] In step A3, after determining the location of the target sound source S0, the modulation index (MCI) is calculated for all L-1 speaker layer combinations. The MCI of speaker layers l and l+1 are:

[0050]

[0051] Among them (G′ l ,G′ l+1 ) is a speaker layer assembly (φ l ,φ l+1 The amplitude modulation coefficient of (G′). l ,G′ l+1 ) satisfies {(G′ l ,G′ l+1 )|G′ l ≥0,G′ l+1 If ≥0}, then layers l and l+1 are effective speaker layers, denoted as lay2 and lay1 respectively. Normalize this set of amplitude modulation coefficients:

[0052]

[0053] Among them (G′ lay1 ,G′ lay2 ) represents the amplitude modulation factor of the speaker layer combination (lay1,lay2).

[0054] A4. Two-dimensional Ambisonics decoding: Based on the azimuth angle of the target sound source, virtual sources S1 and S2 are synthesized using two-dimensional Ambisonics in layers 1 and 2 respectively, and the normalized amplitude modulation coefficient G is used. lay1 G lay2 The relative intensities of the two virtual sources are adjusted separately to synthesize a three-dimensional virtual source.

[0055] In step A4, two-dimensional Ambisonics playback is performed using the speakers on lay1 and lay2 layers respectively, generating virtual sources S1 and S2. The calculation of the speaker gains A and B on lay1 and lay2 layers is as follows:

[0056]

[0057] Where M and K represent the number of loudspeakers and the maximum order on the corresponding loudspeaker layer. For the Mth layer of layer lay1 lay1 The gain of each speaker, For the Mth layer on layer 2 lay2 The gain of each speaker, and K is the sound source S0 located at an azimuth angle θ. lay1 Order and K lay2 Angular harmonic function.

[0058] Finally, activate all speakers in layers lay1 and lay2, and synthesize S0 using virtual sources S1 and S2. The gain of the m-th speaker in layer lay1 is A. m G lay1 The gain of the m-th speaker in layer 2 is B. m G lay2 Specifically, if lay1 or lay2 contains only one speaker, then that speaker is directly considered as S1 or S2 itself. Specifically, if the units in the activated speaker layer are not arranged in a ring, amplitude and delay alignment should be performed in advance during playback.

[0059] As a specific embodiment, this embodiment provides a three-dimensional spatial sound reproduction method for layered loudspeaker arrays, such as... Figure 1 As shown, it includes the following steps:

[0060] A1. Data Input. Input the spatial coordinates (θ, φ) of each speaker unit in the speaker array to the computer or digital signal processor, where θ is the speaker azimuth angle and φ is the speaker elevation angle. Simultaneously, set the spatial coordinates (θ, φ) of the target sound source S0 of the virtual sound source. s0 ,φ s0 ), θ s0 It is the azimuth angle of the sound source, φ s0 It is the elevation angle of the sound source.

[0061] A2. Pre-calculation. The speaker array in this embodiment includes four layers of speakers. The elevation angles of the speakers in layers 1 to 4 are -20°, 0°, 30°, and 90°, respectively. The layers with elevation angles of -20° and 30° each contain 12 speakers evenly distributed according to their azimuth angles, the layer with an elevation angle of 0° contains 36 speakers evenly distributed, and the layer with an elevation angle of 90° contains one ceiling speaker. The four layers of speakers are numbered L = 1, 2, 3, and 4 in ascending order of elevation angle. The highest-order two-dimensional Ambisonics decoding matrix is ​​calculated based on the number of speakers in each layer and the azimuth angle. The highest-order matrices for the layers with elevation angles of -20°, 0°, and 30° are 5, 17, and 5, respectively. The coordinates of each speaker in layer l are (θ...). lm ,φ l ),l∈[1,L],m∈[1,M l The corresponding decoding matrix D l for:

[0062]

[0063] in, for Azimuth K l Angular harmonic function. Where M... l This indicates the number of speakers contained in the l-th layer. This is the angular harmonic function. When l=4, it contains only one loudspeaker, so there is no need to calculate the decoding matrix D4. Simultaneously, following the principle of adjacent elevation angles, they are combined in pairs to form a three-speaker layer combination (φ). l ,φ l+1 )∈{(φ1,φ2),(φ2,φ3),(φ3,φ4)}, where the elevation angles of the speakers in layers 1 to 4 are φ1=-20°, φ2=0°, φ3=30°, and φ4=90°, respectively. Calculate the amplitude modulation inverse matrix for each combination:

[0064]

[0065] Where φ l This indicates the elevation angle of the speaker on the l-th floor.

[0066] A3. Two-dimensional vector base amplitude modulation (VBAP): After determining the orientation of the target sound source S0, the modulation coefficient is calculated for all three speaker layers.

[0067]

[0068] In the formula, (G′ l ,G′ l+1 ) is a speaker layer assembly (φ l ,φl+1 The amplitude modulation coefficient.

[0069] In this embodiment, the azimuth of the target sound source S0 is set to (15°, 15°). The modulation index (G2′, G3′) of the speaker layer combination (0°, 30°) is calculated to be (0.5176, 0.5176). Since G2′≥0 and G3′≥0 are satisfied, the second layer (0° elevation angle) and the third layer (30° elevation angle) are effective speaker layers, denoted as lay2 and lay1 respectively. The modulation index (G2′, G3′) of this group is normalized as follows:

[0070]

[0071] The gain coefficients for this group are obtained as (G) lay2 G lay1 )=(G2,G3)=(0.7071,0.7071).

[0072] It should be noted that the method proposed in this invention is not limited to synthesizing virtual sound sources in that direction.

[0073] A4. Two-dimensional Ambisonics decoding: Two-dimensional Ambisonics playback is performed using speakers from layer 2 (0° elevation angle, layer 2) and layer 3 (30° elevation angle) (layer 1), generating virtual sources S1 and S2. The gain of each speaker on layer 1 is [A1, A2, ..., A...]. 12 [B1, B2, ..., B] and the speakers on layer 2. 36 The calculation is as follows:

[0074]

[0075] In the formula, D2 and D3 are the two-dimensional Ambisonics decoding matrices of the second and third layers, respectively.

[0076] Finally, as Figure 2 As shown, when reproducing the virtual sound source S0 in 3D, all speakers in layers lay1 and lay2 are activated, and the target virtual sound source S0 is synthesized using virtual sound source S1 and virtual sound source S2. The gain of speaker m in layer lay1 is A. m G lay1 The gain of speaker m on layer 2 is B. m G lay2 In the above process, since both virtual sound sources S1 and S2 are synthesized using high-order two-dimensional Ambisonics, the accuracy of their sound field synthesis is high. The simulation results of the sound field synthesis are as follows: Figure 3 As shown in (a) and (b).

[0077] Without loss of generality, in another embodiment, the elevation angle φ of the target virtual sound source S0 is set. s0 =60°, then the effective speaker layers are the 3rd and 4th layers. In this case, the gain calculation for the speaker in layer 1 is still calculated according to the above embodiment as A. m G lay1 The gain of the 2-layer ceiling speaker is G. lay2 .

[0078] Using the loudspeaker array in this embodiment, the algorithm proposed in this invention is compared with traditional three-dimensional Ambisonics and mixed-order Ambisonics algorithms. Ten subjects were recruited to conduct a sound source localization experiment. The localization error results are as follows: Figure 4a As shown. Ten subjects were recruited to participate in a timbre evaluation experiment. The timbre distortion scores are as follows. Figure 4b As shown. Subjective experiments demonstrate that, under this layered speaker array, the algorithm proposed in this invention outperforms traditional algorithms in both positioning accuracy and timbre distortion.

[0079] In any other embodiment, the specific implementation process can be flexibly adjusted according to the needs of the actual 3D signal program, adjusting the position of the speaker layers and the number of speakers in each layer. Typically, the horizontal plane layer should have a larger number of speakers to ensure the accuracy of the reproduction of the virtual sound source space on the horizontal plane.

[0080] The above embodiments are merely a few practical examples of the present invention, but the implementation of the present invention is not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A three-dimensional spatial sound reproduction method for layered loudspeaker arrays, characterized in that, Includes the following steps: A1. Set the spatial coordinates of each unit in the speaker array and set the spatial coordinates of the target virtual sound source; A2. Divide the loudspeaker array into layers according to elevation angle, and calculate the two-dimensional Ambisonics decoding matrix based on the number of loudspeaker arrays in each layer and azimuth angle; combine loudspeaker layers in pairs according to the adjacent principle, and calculate the amplitude modulation inverse matrix; A3. Based on the elevation angle of the target virtual sound source, calculate the vector amplitude modulation coefficient of the speaker layers in pairs, select the pairs with non-negative coefficients as the effective speaker layers, and retain the corresponding normalized amplitude modulation coefficients. A4. Based on the azimuth angle of the target sound source, use two-dimensional Ambisonics to synthesize virtual sources on the effective loudspeaker layer, and use amplitude modulation coefficients to adjust the relative intensity of the two virtual sources to synthesize a three-dimensional virtual source.

2. The three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 1, characterized in that, In step A2, all speakers are divided according to their elevation angle. L Layer, of which the first l Layer contains There are [number] speakers, each with coordinates [coordinate]. , Indicates the first l Speaker elevation angle of the floor, Indicates the first l Layer One speaker azimuth angle; Calculate for each layer separately Two-dimensional Ambisonics decoding matrix ,in , Indicates rounding down, the 1st... l The method for calculating the decoding matrix corresponding to the layer is as follows: in, for Angular harmonic function of order, , and These correspond to the sine term, cosine term, and 0th order angular harmonic function, respectively.

3. The three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 2, characterized in that, In step A2, the pre-divided sections are grouped in pairs according to the principle of adjacent elevation angles. L The speaker layers are combined in pairs according to their elevation angle to form... L -1 speaker layer combination , Indicates the first l The speaker elevation angle of the layer.

4. The three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 3, characterized in that, In step A3, the target virtual sound source is determined. S0 After determining the location, for all L -1 speaker layer combination to calculate amplitude modulation factor The modulation factor of the loudspeaker is: in For speaker layer amplitude modulation coefficient, Represents the target virtual sound source S0 The elevation angle of the sound source, Indicates the first l The speaker elevation angle of the layer.

5. A three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 4, characterized in that, In step A3, if the calculation yields... l and l The amplitude modulation coefficients corresponding to layer +1 are all greater than 0, then l and l The +1 layer is the effective speaker layer, and the amplitude modulation coefficient is normalized to obtain the normalized gain.

6. A three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 1, characterized in that, In step A4, two-dimensional Ambisonics playback is performed using the speakers of the effective speaker layer to generate a virtual source; The gain of each speaker in the effective speaker layer is calculated based on the decoding matrix corresponding to the effective speaker layer.

7. A three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 6, characterized in that, In step A4, all speakers in the effective speaker layer are activated, and the target virtual sound source is synthesized using the virtual source of the effective speaker layer. After synthesis, the first virtual sound source of the effective speaker layer is... m The gain of the speaker is the first one before synthesis. m The product of the speaker gain and the gain after normalization of the modulation index of the effective speaker layer.

8. A three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 7, characterized in that, In a speaker layer containing a single ceiling speaker, if the ceiling speaker is activated, the ceiling speaker is directly treated as the virtual sound source itself for gain allocation.

9. A three-dimensional spatial sound reproduction method for a layered loudspeaker array as described in claim 7, characterized in that, If the units in the active speaker layer are not arranged in a ring, amplitude and delay alignment should be performed in advance during playback.

10. A three-dimensional spatial sound reproduction system for a layered loudspeaker array, characterized in that, include: A layered speaker array, the speaker array containing at least two speaker layers; The signal processing module is configured to implement the entire processing flow of the method according to any one of claims 1-9; The dynamic setting module is configured to set the spatial coordinates of the sound source in real time and dynamically adjust the gain weight of each layer of speakers.

Citation Information

Patent Citations

  • Environment-adaptive sound field playback space decoding method

    CN113314129A

  • Sound playback method of object-based loudspeaker three-dimensional space array

    CN118828338A