Three-dimensional space sound playback method for layered loudspeaker array
By adopting the three-dimensional spatial sound reproduction method of layered speaker arrays, using two-dimensional Ambisonics to synthesize virtual sound sources and adjust the relative intensity, the problem of difficult speaker array deployment in the existing technology is solved, and the reproduction accuracy and timbre effect of the horizontal area are improved.
Patent Information
- Application Number
- CN202511231233.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-30
- Publication Date
- 2025-10-14
AI Technical Summary
Existing three-dimensional spatial sound reproduction technology is difficult to deploy in actual speaker arrays, especially for speaker arrays with a layered structure. The reproduction effect is poor, and high requirements are placed on the number of speakers and spatial uniformity. It is difficult to improve the reproduction accuracy of the horizontal plane area while ensuring the high-elevation sound source localization effect.
A three-dimensional spatial sound reproduction method for layered speaker arrays is adopted. By dividing the speaker array into multiple layers, two-dimensional Ambisonics is used to synthesize virtual sound sources, and the relative intensity is adjusted through vector amplitude modulation. This reduces the requirements for the uniformity of the speaker spatial distribution and improves the reproduction accuracy in the horizontal area.
Under the premise of ensuring the high-elevation-angle sound source localization effect, the requirements for the number of speakers and spatial uniformity are reduced, the playback accuracy of the horizontal area is improved and the sound distortion is reduced, and the speaker array can be flexibly designed to meet actual needs.
Smart Images

Figure CN120786280A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional spatial sound playback, and particularly relates to a three-dimensional spatial sound playback method for layered loudspeaker arrays, which optimizes the playback spatial resolution and timbre in the horizontal plane and high elevation directions. BACKGROUND
[0002] Spatial sound playback technology based on loudspeaker arrays aims to preserve the directionality of virtual sound sources to some extent through a limited number of loudspeakers. The most common spatial sound playback technologies are two-channel stereo and 5.1-channel surround sound, both of which can only playback virtual sound sources in the horizontal plane. In recent years, spatial sound technology has gradually expanded from two-dimensional playback to three-dimensional playback. Representative three-dimensional playback technologies include three-dimensional vector base amplitude panning (VBAP), Ambisonics, and wave field synthesis. Ambisonics technology encodes and decodes spatial sound sources using spherical harmonics, achieving a relatively good balance between playback accuracy and loudspeaker quantity requirements.
[0003] Ambisonics technology first decomposes a sound field composed of spatial plane waves or spherical waves into N-order spherical harmonic domain coefficients. Then, through sound field matching, the spherical harmonic domain coefficients are re-decoded into loudspeaker gain coefficients based on the spatial positions of the loudspeaker array. In this process, to ensure the stability of the decoding, the number of loudspeakers cannot be lower than the number of spherical harmonics, and the loudspeakers should be uniformly distributed in space. The order N of the spherical harmonics is directly related to the spatial resolution and timbre distortion of the playback sound source.
[0004] In a two-dimensional sound field, an N-order spherical harmonic function (in this case, an angular harmonic function) contains 2N+1 harmonic components, and the number of loudspeakers required is relatively low. However, for a three-dimensional sound field, an N-order spherical harmonic function contains (N+1) 2 harmonic components, and the number of loudspeakers required increases exponentially with the order N. In addition, three-dimensional sound field playback requires loudspeakers to be uniformly distributed in space, which is in conflict with the spatial loudspeaker arrays commonly used in practical application scenarios (such as the NHK 22.2 system) and the perceptual characteristics of the auditory system: ① Actual systems usually configure more loudspeakers in the horizontal plane and only a small number of loudspeakers at other heights; ② Considering the difficulty of loudspeaker installation engineering, actual loudspeaker arrays are usually arranged in layers, such as arranging two-dimensional loudspeaker arrays at specific elevation angles of 30°, 60°, etc. in the horizontal plane; ③ The auditory system is most sensitive to the direction of sound sources in the horizontal plane, and the directional resolution of sound sources at other heights is relatively poor.
[0005] Some methods have been proposed to solve the problem of non-uniform spatial distribution of loudspeakers, such as the hybrid-order Ambisonics method (Ahrens J., Spors S.. An Analytical Approach to Sound Field Reproduction Using Circular and Spherical Loudspeaker Distributions [J]. Acta Acustica united with Acustica, 2008, 94(6): 988-999.), which uses more loudspeakers in the horizontal plane to improve the accuracy of two-dimensional playback based on three-dimensional playback, but this method can only improve the accuracy of playback in the horizontal plane, and still requires a large number of arrays and spatial uniformity in the high-elevation direction. The All-Round Ambisonics Decoding algorithm (Zotter F., Frank M. All-Round Ambisonic Panning and Decoding [J]. Journal of the Audio Engineering Society, 2012, 60(10): 807-820.) combines VBAP and Ambisonics algorithms to reduce the requirements for three-dimensional layout of loudspeakers, but does not improve the accuracy of playback in specific areas (such as the horizontal plane), and may even reduce the accuracy of playback in areas with a large number of loudspeakers. In summary, there is currently no method that can improve the accuracy of playback in areas with actual physical loudspeakers (including but not limited to the horizontal plane) while ensuring a certain vertical playback accuracy. In particular, for common loudspeaker arrays with layered structure characteristics, researchers have not paid much attention to them, and there is no three-dimensional spatial sound playback algorithm for such arrays. Due to the above problems, the deployment of three-dimensional spatial sound playback technology in actual loudspeaker array systems is difficult, and the playback effect is often poor. The present application proposes a three-dimensional spatial sound playback algorithm for loudspeaker arrays with layered structure characteristics, which can improve the accuracy of playback in areas with actual loudspeakers, especially the horizontal plane, while ensuring a certain high-elevation sound source positioning effect, and reduce the requirements for the number of non-horizontal loudspeakers and spatial uniformity. SUMMARY
[0006] In view of the stability and efficiency of the current Ambisonics three-dimensional playback in a multi-layer loudspeaker array, the application provides a three-dimensional spatial sound playback method for a layered loudspeaker array, the application uses two layers of physical loudspeaker arrays, synthesizes a pair of virtual sources located at the same azimuth and different elevations through two-dimensional Ambisonics; then, the relative strength of the two virtual sources is adjusted through vector amplitude modulation to synthesize a target virtual source, aiming to reduce the requirement of the algorithm on the spatial uniformity of the loudspeakers, and to improve the playback accuracy of the area with actual loudspeakers as much as possible, especially the horizontal plane area, while ensuring the positioning effect of high-elevation sound sources.
[0007] The application is implemented by at least one of the following technical solutions.
[0008] A three-dimensional spatial sound playback method for a layered loudspeaker array, comprising the following steps:
[0009] A1, setting the spatial coordinates of each unit of the loudspeaker array and setting the spatial coordinates of the target virtual sound source;
[0010] A2, layering the loudspeaker array according to the elevation, calculating the two-dimensional Ambisonics decoding matrix according to the number of loudspeakers in each layer and the azimuth, combining the loudspeaker layers in pairs according to the adjacent principle, and calculating the amplitude modulation inverse matrix;
[0011] A3, calculating the vector amplitude modulation coefficient of the loudspeaker layers combined in pairs according to the elevation of the target virtual sound source, selecting the combination pair with non-negative coefficients as the effective loudspeaker layer, and retaining the corresponding normalized amplitude modulation coefficient;
[0012] A4, using two-dimensional Ambisonics to synthesize virtual sources in the effective loudspeaker layer according to the azimuth of the target sound source, and adjusting the relative strength of the two virtual sources using the amplitude modulation coefficient to synthesize a three-dimensional virtual source.
[0013] Further, in step A2, all loudspeakers are divided into L layers according to the elevation of the loudspeakers, wherein the lth layer contains M l loudspeakers, each loudspeaker coordinate is (θ lm ,φ l ), l∈[1,L], m∈[1,M l ], φ l represents the loudspeaker elevation of the lth layer, θ lm represents the azimuth of the mth loudspeaker of the lth layer; K l order two-dimensional Ambisonics decoding matrix D l is calculated for each layer, wherein represents the floor function, and the calculation method of the decoding matrix corresponding to the lth layer is:
[0014]
[0015] wherein, is K l The angular harmonic functions of order i = -1, i = 1 and i = 0 correspond to the sine term, the cosine term and the 0th order angular harmonic function, respectively.
[0016] Further, in step A2, according to the principle of elevation angle adjacency, the L layers of loudspeakers are combined two by two in the order of elevation angle to form L-1 loudspeaker layer combinations {(φ l ,φ l+1 )|1≤l≤L-1}, where φ l represents the loudspeaker elevation angle of the lth layer.
[0017] Further, in step A3, after the azimuth of the target virtual sound source S0 is determined, the amplitude modulation coefficients of all L-1 loudspeaker layer combinations are calculated, and the amplitude modulation coefficient of the (φ l ,φ l+1 ) layer loudspeaker is:
[0018]
[0019] where (G′ l ,G′ l+1 ) is the amplitude modulation coefficient of the (φ l ,φ l+1 ) layer loudspeaker, φ s0 is the sound source elevation angle of the target virtual sound source S0, and φ l represents the loudspeaker elevation angle of the lth layer.
[0020] Further, in step A3, if the amplitude modulation coefficients corresponding to the lth and (l+1)th layers are both greater than 0, the lth and (l+1)th layers are effective loudspeaker layers, and the amplitude modulation coefficients are subjected to amplitude normalization to obtain the normalized gain.
[0021] Further, in step A4, the loudspeakers of the effective loudspeaker layers are used for two-dimensional Ambisonics playback to generate a virtual source, and the gains of the loudspeakers in the effective loudspeaker layers are calculated according to the decoding matrix corresponding to the effective loudspeaker layers.
[0022] Further, in step A4, all the loudspeakers of the effective loudspeaker layers are activated, the virtual source of the effective loudspeaker layers is used to synthesize the target virtual sound source, and the gain of the mth loudspeaker of the effective loudspeaker layers after synthesis is the product of the gain of the mth loudspeaker before synthesis and the normalized gain of the amplitude modulation coefficient of the effective loudspeaker layer.
[0023] Further, in the speaker layer containing a single ceiling speaker, if the ceiling speaker is activated, the ceiling speaker is directly regarded as a virtual sound source itself for gain distribution.
[0024] Further, if the units in the activated speaker layer are not distributed in a ring shape, amplitude and delay alignment is performed in advance during playback.
[0025] The system for implementing the three-dimensional spatial sound playback method for a layered speaker array comprises:
[0026] The speaker array has layers, and comprises at least two speaker layers;
[0027] The signal processing module is configured to implement the entire processing procedure of the method.
[0028] The dynamic setting module is configured to implement the procedure, and can set the spatial coordinates of the sound source in real time and dynamically adjust the gain weights of the speaker layers.
[0029] Compared with the prior art, the method has the following beneficial effects:
[0030] 1. By layered decoding, the requirement of three-dimensional Ambisonics playback for the number and uniform distribution of the speakers on a non-horizontal plane is reduced, and the demand of an actual speaker array is met.
[0031] 2. In the area with actual speakers, especially the horizontal plane area, the energy dispersion of the virtual sound source is reduced to the maximum, the positioning accuracy is improved, and the timbre distortion is reduced.
[0032] 3. The speaker array can be flexibly designed according to the demand of an actual signal source, and the spatial accuracy of different areas can be adjusted as needed. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 FIG. 1 is a flowchart of an embodiment of the three-dimensional spatial sound playback method for a layered speaker array. Figure 2 FIG. 2 is a diagram of a coordinate system and three-dimensional virtual source synthesis.
[0034] Figure 3 FIG. 3 is a sound field reconstruction effect diagram of a high-elevation speaker layer in the embodiment.
[0035] Figure 4a FIG. 4 is a positioning subjective experiment evaluation result diagram in the embodiment.
[0036] Figure 4b FIG. 5 is a timbre subjective evaluation experiment result diagram in the embodiment. DETAILED DESCRIPTION
[0037] The technical solutions and advantages of the present application will be further described below in combination with embodiments and drawings, but the embodiments of the present application are not limited thereto.
[0038] As shown in the figure, the three-dimensional spatial sound playback method for layered loudspeaker array in the embodiment comprises the following steps: Figure 1
[0039] A1, data input. The loudspeakers and the virtual source are both in the spatial spherical coordinate system, the spatial coordinates (θ, φ) of each unit of the preset loudspeaker array are set, and the spatial coordinates (θ s0 ,φ s0 ) of the target virtual sound source S0 are set, wherein θ is the azimuth angle of the loudspeaker, and φ is the elevation angle of the loudspeaker. The spatial coordinates (θ s0 ,φ s0 ) of the target virtual sound source S0 are set, θ s0 is the azimuth angle of the sound source, and φ s0 is the elevation angle of the sound source. The loudspeaker array should have a layered characteristic in space, that is, multiple loudspeakers have the same or similar elevation angle.
[0040] A2, pre-computation. The loudspeaker array is grouped according to the elevation angle, the highest-order two-dimensional Ambisonics decoding matrix is calculated according to the number of loudspeaker arrays and the azimuth angle of each layer; the loudspeaker layers are combined two by two according to the adjacent principle according to the elevation angle of each layer, and the amplitude modulation inverse matrix is calculated.
[0041] In step A2, all loudspeakers are divided into L layers according to the elevation angle of the loudspeakers, the lth layer contains M l loudspeakers, and the coordinates of each loudspeaker on the layer are (θ lm ,φ l ), l∈[1,L], m∈[1,M l ], wherein the elevation angle φ l of the lth layer loudspeaker is arranged from small to large according to the numerical value (that is, φ l+1 >φ l ). The K l order two-dimensional Ambisonics decoding matrix is calculated for each layer, The decoding matrix of all layers is calculated, and the decoding matrix D l of the lth layer is:
[0042]
[0043] wherein, is the K l order angular harmonic function. More generally, the definition of the angular harmonic function of the nth order and the ith term is:
[0044]
[0045] In particular, if the lth layer has only one speaker unit (e.g. a ceiling speaker), step A2 is skipped. In addition, the decoding order of the lth layer is not limited to the highest order K l , which should be selected flexibly according to the actual array design, and the highest order K l is used as much as possible under the premise of maintaining stability.
[0046] Meanwhile, according to the principle of adjacent elevation angles, two-by-two combinations are formed to form L-1 speaker layer combinations {(φ l ,φ l+1 )|1≤l≤L-1}, and the amplitude modulation inverse matrix is calculated for each combination:
[0047]
[0048] A3, two-dimensional VBAP, according to the elevation angle of the target virtual sound source, the vector amplitude modulation coefficients of the two-by-two combined speaker layers are calculated, and the combination pair with non-negative coefficients is selected as the effective speaker layers lay1 and lay2, and the corresponding normalized amplitude modulation coefficients G lay1 ,G lay2 are retained.
[0049] In step A3, after the orientation of the target sound source S0 is determined, the amplitude modulation coefficients are calculated for all L-1 speaker layer combinations, and the amplitude modulation coefficients of the lth and (l+1)th layers are:
[0050]
[0051] where (G′ l ,G′ l+1 ) are the amplitude modulation coefficients of the speaker layer combination (φ l ,φ l+1 ). If (G′ l ,G′ l+1 ) satisfies {(G′ l ,G′ l+1 )|G′ l ≥0,G′ l+1 ≥0}, then the lth and (l+1)th layers are effective speaker layers, denoted as lay2 and lay1 respectively, and the amplitude modulation coefficients of this combination are normalized:
[0052]
[0053] where (G′ lay1 ,G′ lay2 ) represents the amplitude modulation coefficients of the speaker layer combination (lay1, lay2).
[0054] A4, 2D Ambisonics decoding, using 2D Ambisonics to synthesize virtual sources S1 and S2 respectively at lay1 and lay2 layers according to the target sound source azimuth angle, with normalized amplitude panning coefficients G lay1 lay2 Adjusting the relative intensity of the two virtual sources respectively, synthesizing a three-dimensional virtual source.
[0055] In step A4, 2D Ambisonics playback is performed using the lay1 and lay2 layers of loudspeakers respectively to produce virtual sources S1 and S2. The calculation of the gains A and B of each loudspeaker on the lay1 and lay2 layers is as follows:
[0056]
[0057] where M and K are the number of loudspeakers and the maximum order on the corresponding loudspeaker layer, is the gain of the M lay1 th loudspeaker on the lay1 layer, is the gain of the M lay2 th loudspeaker on the lay2 layer, and are the K lay1 th and K lay2 th angular harmonic functions of the sound source S0 located at the azimuth angle θ.
[0058] Finally, all loudspeakers on the lay1 and lay2 layers are activated, and virtual sources S1 and S2 are used to synthesize S0. The gain of the m m th loudspeaker on the lay1 layer is A lay1 G m , and the gain of the m lay2 th loudspeaker on the lay2 layer is B m G lay2 . In particular, if the lay1 or lay2 contains only one loudspeaker, the loudspeaker is directly considered as S1 or S2 itself. In particular, if the units in the activated loudspeaker layer are not distributed in a ring shape, amplitude and delay alignment should be performed in advance during playback.
[0059] As a specific embodiment, a three-dimensional spatial sound playback method for a layered loudspeaker array according to the present embodiment, as shown in Figure 1 , comprises the following steps:
[0060] A1, data input. Input the spatial coordinates (θ, φ) of each loudspeaker unit of the loudspeaker array to the computer or digital signal processor, θ is the azimuth angle of the loudspeaker, φ is the elevation angle of the loudspeaker, and set the spatial coordinates (θ s0 , φ s0 ) of the target virtual sound source S0, θ s0 is the azimuth angle of the sound source, and φ s0 is the elevation angle of the sound source.
[0061] A2. Pre-calculation. The speaker array of this embodiment includes four layers of speakers. The elevation angles of the speakers in the 1st to 4th layers are -20°, 0°, 30° and 90° respectively. The layers with elevation angles of -20° and 30° each contain 12 speakers evenly distributed in azimuth, the layer with elevation angle of 0° contains 36 evenly distributed speakers, and the layer with elevation angle of 90° contains one ceiling speaker. The four layers of speakers are numbered L=1, 2, 3, 4 in order from low to high elevation angles. The highest-order two-dimensional Ambisonics decoding matrix is calculated according to the number of speaker arrays and azimuth angles in each layer. The highest orders for the layers with elevation angles of -20°, 0° and 30° are 5, 17 and 5 respectively. The coordinates of each speaker on the lth layer are (θ lm ,φ l ),l∈[1,L],m∈[1,M l ], the corresponding decoding matrix D l for:
[0062]
[0063] in, for Azimuth K l The angular harmonic function. l represents the number of speakers contained in the lth layer, is the angular harmonic function. When l = 4, only one loudspeaker is included and the decoding matrix D4 does not need to be calculated. At the same time, according to the principle of adjacent elevation angles, three loudspeaker layer combinations are formed (φ l ,φ l+1 )∈{(φ1,φ2),(φ2,φ3),(φ3,φ4)}, where the elevation angles of the speakers in the 1st to 4th layers are φ1 = -20°, φ2 = 0°, φ3 = 30°, and φ4 = 90°. Calculate the inverse amplitude modulation matrix for each combination:
[0064]
[0065] where φ l Indicates the elevation angle of the loudspeaker on the lth layer.
[0066] A3, 2D Vector Base Amplitude Panning (VBAP), after determining the direction of the target sound source S0, calculates the amplitude modulation coefficient for all three speaker layer combinations:
[0067]
[0068] In the formula, (G′ l ,G′ l+1 ) is the speaker layer combination (φ l ,φl+1 ) of the amplitude modulation coefficient.
[0069] In this embodiment, the orientation of the target sound source S0 is set to (15°, 15°), and the amplitude modulation coefficients (G2′, G3′) of the speaker layer combination (0°, 30°) are calculated to be (0.5176, 0.5176). If the conditions G2′ ≥ 0, G3′ ≥ 0 are met, then the second layer (0° elevation angle) and the third layer (30° elevation angle) are effective speaker layers, denoted as lay2 and lay1 respectively. The amplitude modulation coefficients (G2′, G3′) of this group are normalized as follows:
[0070]
[0071] The gain coefficient of this group is (G lay2 ,G lay1 )=(G2,G3)=(0.7071,0.7071).
[0072] It should be noted that the method proposed in the present invention is not limited to synthesizing a virtual sound source in this orientation.
[0073] A4, 2D Ambisonics decoding, using the speakers of the second layer (0° elevation, lay2 layer) and the third layer (30° elevation) (lay1 layer) to perform 2D Ambisonics playback, generating virtual sources S1 and S2, and the gains of each speaker on the lay1 layer [A1, A2, ..., A 12 ] and each speaker on the lay2 layer [B1, B2, ..., B 36 ]The calculation is as follows:
[0074]
[0075] Where D2 and D3 are the two-dimensional Ambisonics decoding matrices of the second and third layers.
[0076] Finally, if Figure 2 As shown, when reproducing the virtual sound source S0 in three dimensions, all the speakers in the lay1 and lay2 layers are activated, and the target virtual sound source S0 is synthesized using the virtual sound source S1 and the virtual sound source S2. The gain of the m-th speaker in the lay1 layer is A m G lay1 , the gain of the m-th speaker in layer 2 is B m G lay2 In the above process, since the virtual sound source S1 and the virtual sound source S2 are synthesized by higher-order two-dimensional Ambisonics, the accuracy of the sound field synthesis is high. The simulation results of the sound field synthesis are as follows: Figure 3 As shown in (a) and (b).
[0077] Without loss of generality, in another embodiment, the elevation angle φ of the target virtual sound source S0 is set to 60°, then the effective speaker layers are the 3rd layer and the 4th layer, and the gain of the lay1 layer speaker is still calculated as A s0 m lay1 lay2 .
[0078] With the speaker array in the embodiment, the algorithm proposed in the present application is compared with the conventional three-dimensional Ambisonics and mixed-order Ambisonics algorithm, 10 subjects are recruited to conduct sound source positioning experiments, and the positioning error results are as shown in Figure 4a . 10 subjects are recruited to conduct timbre evaluation experiments, and the timbre distortion score results are as shown in Figure 4b . The subjective experiments show that under the layered speaker array, the algorithm proposed in the present application is superior to the conventional algorithm in positioning accuracy and timbre distortion.
[0079] In any other embodiment, the specific implementation process can flexibly adjust the position of the speaker layer and the number of speakers in each layer according to the actual needs of the three-dimensional signal program. Generally, the horizontal layer should have a larger number of speakers to ensure the accuracy of the horizontal virtual sound source spatial playback.
[0080] The above embodiments are only a few practical cases of the present application, but the implementation of the present application is not limited by the above embodiments, and any changes, modifications, substitutions, combinations and simplifications made without departing from the spirit and principles of the present application shall be equivalent replacement methods and shall be within the scope of protection of the present application.
Claims
1. A three-dimensional spatial sound reproduction method for a layered loudspeaker array, characterized in that: The following steps are involved: A1. Set the spatial coordinates of each unit in the speaker array and the spatial coordinates of the target virtual sound source; A2. Layer the speaker array according to elevation angles and calculate the two-dimensional Ambisonics decoding matrix based on the number of speakers in each layer and the azimuth angle. Combine the speaker layers in pairs according to the adjacent principle and calculate the inverse amplitude modulation matrix. A3. Calculate the vector amplitude modulation coefficients of the two-pair loudspeaker layers according to the elevation angle of the target virtual sound source, select the pairs whose coefficients are both non-negative as the effective loudspeaker layers, and retain the corresponding normalized amplitude modulation coefficients; A4. Based on the azimuth of the target sound source, two-dimensional Ambisonics is used to synthesize virtual sources at the effective speaker layer. The relative intensities of the two virtual sources are adjusted using the amplitude modulation coefficient to synthesize a three-dimensional virtual source.
2. The three-dimensional sound reproduction method for a layered speaker array according to claim 1, wherein: In step A2, all the speakers are divided into L layers according to their elevation angles, where the lth layer contains M l speakers, and the coordinates of each speaker are (θ lm ,φ l ),l∈[1,L],m∈[1,M l ],φ l represents the elevation angle of the loudspeaker on the lth layer, θ lm represents the azimuth angle of the mth loudspeaker in the lth layer; Calculate K for each layer separately l The 2D Ambisonics decoding matrix D l ,in Indicates rounding down. The calculation method of the decoding matrix corresponding to the lth layer is: in, K l The angular harmonic function of the order, i = -1, i = 1 and i = 0 correspond to the sine term, cosine term and 0th order angular harmonic function respectively.
3. The three-dimensional sound reproduction method for a layered speaker array according to claim 1, wherein: In step A2, the divided L speaker layers are combined in pairs according to the principle of adjacent elevation angles, and L-1 speaker layer combinations are formed {(φ l ,φ l+1 )|1≤l≤L-1},φ l Indicates the elevation angle of the loudspeaker on the lth layer.
4. The three-dimensional sound reproduction method for a layered speaker array according to claim 1, wherein: In step A3, after determining the orientation of the target virtual sound source S0, the amplitude modulation coefficients (φ l ,φ l+1 ) layer loudspeaker amplitude modulation coefficient is: Where (G′ l ,G′ l+1 ) is the speaker layer (φ l ,φ l+1 ) of the amplitude modulation coefficient, φ s0 The sound source elevation angle of the target virtual sound source S0, φ l Indicates the elevation angle of the loudspeaker on the lth layer.
5. The three-dimensional sound reproduction method for a layered speaker array according to claim 4, wherein: In step A3, if the calculated amplitude modulation coefficients corresponding to layers 1 and 1+1 are both greater than 0, layers 1 and 1+1 are effective speaker layers, and the amplitude modulation coefficients are normalized to obtain normalized gains.
6. The three-dimensional sound reproduction method for a layered speaker array according to claim 1, wherein: In step A4, two-dimensional Ambisonics playback is performed using the speakers of the two effective speaker layers to generate a virtual source; The gain of each speaker in the effective speaker layer is calculated according to the decoding matrix corresponding to the effective speaker layer.
7. The three-dimensional sound reproduction method for a layered speaker array according to claim 6, wherein: In step A4, all speakers of the effective speaker layer are activated, and the target virtual sound source is synthesized using the virtual source of the effective speaker layer. The gain of the mth speaker of the effective speaker layer after synthesis is the product of the gain of the mth speaker before synthesis and the gain after normalization of the amplitude modulation coefficient of this effective speaker layer.
8. The three-dimensional sound reproduction method for a layered loudspeaker array according to claim 7, wherein: In a speaker layer containing a single ceiling speaker, if the ceiling speaker is activated, the gain is distributed directly as if it were a virtual sound source itself.
9. The three-dimensional sound reproduction method for a layered speaker array according to claim 7, wherein: If the units in the activated speaker layer are not distributed in a ring shape, align the amplitude and delay in advance during playback.
10. A system for implementing the three-dimensional spatial sound reproduction method for a layered speaker array as claimed in claim 1, characterized in that: include: A speaker array having a layered structure, the speaker array comprising at least two speaker layers; A signal processing module configured to implement the entire processing flow of the method according to any one of claims 1 to 8; The dynamic setting module is configured to implement the process described in claims 1-8, and can set the spatial coordinates of the sound source in real time and dynamically adjust the gain weights of the speakers in each layer.
Citation Information
Patent Citations
Environment-adaptive sound field playback space decoding method
CN113314129A
Sound playback method of object-based loudspeaker three-dimensional space array
CN118828338A
Method and device for rendering an audio soundfield representation for audio playback
US20150163615A1
Method and device for decoding an audio soundfield representation for audio playback
WO2011117399A1