Early reflection energy controlled injection method based on DirAC parameter

By performing direction encoding and energy allocation within the DirAC parameter domain, the problems of directional flicker and energy overflow in early reflection injection under the DirAC parameter domain are solved, achieving controlled injection of direction and energy and consistent rendering, adapting to different audio scenarios.

CN121284477APending Publication Date: 2026-01-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511397579.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

In immersive audio scenarios, existing technologies struggle to achieve directional selectivity and energy-controlled early reflection injection in the DirAC parameter domain, resulting in directional flickering, energy overflow, and color distortion. Furthermore, it is difficult to maintain equivalence and consistency of direction and energy across back-end systems.

Method used

Directional coding and energy allocation are performed within the DirAC parameter domain. The directional index n(f,t) is adaptively adjusted through hard clipping of the frame-level budget ρ(f,t) and soft projection of the angular difference, and scheduled between path-level, hybrid, or merge modes to ensure consistency of energy and direction.

Benefits of technology

It improves directional clarity and energy concentration, reduces directional entropy, maintains energy and directional consistency between different back-ends, and adapts to differences in array layout and room materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284477A_ABST
    Figure CN121284477A_ABST
Patent Text Reader

Abstract

The invention discloses an early reflection injection method of a surface DirAC. The method comprises the following steps: performing time-frequency analysis on an FOA signal to obtain a DirAC arrival direction and diffusivity psi (f, t), and receiving a path set comprising a direction rk, a time delay tauk and a gain gk; and constructing a direction coding impulse response corresponding to the rk in an FOA / DirAC domain, projecting according to an angular difference weight ((1 + cos delta theta) / 2) n (f, t), and executing energy normalization on each path. And constructing a frame-level injection budget rho (f, t) based on the psi (f, t) and the intensity, constraining the injection energy, and scheduling between path-level / mixing / merging modes. According to the method, the direction selectivity and the external allelopathy are improved while the energy and physical consistency is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of audio signal processing and spatial audio rendering, specifically to early reflection injection and rendering based on DirAC. Background Technology

[0002] In real-world immersive audio applications (VR / AR headsets, interactive live streaming, and 3GPP EVS-IVAS codec links), the 50–80ms "early window" directly behind the sound often determines the listener's first impression of externalization, directional stability, and distance. We conducted several small experiments in a 10×8×3m rectangular room: when only simple gain superposition of the channel domains was performed without upper limits on the injected energy, the headphones were more prone to "chilly" sounds at high frequencies; switching to an AllRAD array resulted in a wider forward energy diffusion. These phenomena suggest that engineering implementation requires not only "reflections," but also directionally focused and energy-controlled early reflections.

[0003] Traditional geometric acoustic methods (image source method, light / sound ray tracing) generate multi-order reflections based on room geometry and materials, providing a complete path-delay-direction-gain quadruple, and the physical interpretation is also very direct (Reference: Allen, Jont B., and David A. Berkley. "Image method for efficiently simulating small-roomacoustics." The Journal of the Acoustical Society of America 65.4(1979):943-950.). However, after the third order, we often see a significant "spreading out" of the angular domain distribution: the energy is spread out on the sphere, which is inconsistent with the auditory tendency to "focus around the dominant direction of arrival (DoA)"; if these are mixed indiscriminately into the decoding backend, externalization and front-to-back localization can easily become "floaty". Conversely, perceptual / statistical schemes (such as directional FDN) are quick to implement and have controllable computational power, but they are more about "synthesizing impressions", making it difficult to provide a path-level structure that can be reused across systems, and they also lack verifiable energy constraints with the backend decoding.

[0004] Within a parametric analysis framework that falls between the two, Directional Audio Coding (DirAC) is used in the time-frequency domain to estimate DoA and diffusion, distinguishing between directional and diffusion components. It is very useful as a real-time front-end, and we have long used it as the entry point of our analysis chain (Reference: Pulkki, Ville, et al. "First-Order Directional Audio Coding (DirAC)." Parametric Time-Frequency Domain Spatial Audio (2017): 89-140.). However, DirAC's "default behavior" is still analysis and allocation: it doesn't directly produce multipath sets, nor does it have a set of rules for "quantitatively writing back (injecting) early energy" in the FOA / DirAC analysis domain. We tried channel-domain weighting based solely on DirAC parameters, and the results were: in strong diffusion or multi-source scenarios, the direction flickers slightly; and if over-injected, overflow and colorization occur. When mapping the same result to the VBAP / binaural back-end, the energy and direction don't always match—mainly due to the lack of a unified budget and normalization constraints.

[0005] Another frequently mentioned approach is rendering based on measured RIR. For example, SIRR performs DoA / diffusion analysis on the room's impulse response and then redistributes energy according to direction, often achieving very "accurate" directional localization (Reference: Merimaa, Juha, and Ville Pulkki. Spatial impulse response rendering I: Analysis and synthesis. Journal of the Audio Engineering Society 53.12(2005):1115-1127.). The problem is that this approach is heavily reliant on measured data: the RIR needs to be updated whenever the user or sound source moves; maintaining this data in mobile scenarios is costly, and the decoupling between the front-end and back-end is not always smooth.

[0006] Based on the above experience, we would prefer a component that takes a set of paths as input, from any source (geometric, statistical, or labeled), and includes direction, delay, and gain; distributes path energy in the FOA / DirAC domain using soft projection based on angular difference, and normalizes the energy of each path to avoid artificially inflating the total energy due to "multi-channel writing"; and sets a budget ρ(f,t) for each time-frequency unit using the diffusivity ψ(f,t) and the active intensity ||I||, which is the upper limit of the injected energy.

[0007] In a 48kHz, 1024 / 512STFT configuration, we adjusted ρ(f,t) from 0.2 to 0.6 and found that the externalization boost was not linear; values ​​exceeding 0.4 were more likely to induce comb effects. We introduced an adaptive directional exponent n(f,t): tightening the focus when ψ(f,t) is low (clear direction) and widening it when ψ(f,t) is high (significant diffusion) to suppress directional flicker caused by time-frequency jitter. We provided three scheduling methods: path-level, hybrid, and merging. We performed AllRAD / VBAP / binary decoding after controlled injection in the parameter domain, making it easier to achieve equivalence and consistency of direction and energy between different back-ends (our actual tests showed that after unifying the budget, the energy difference across back-ends in the same frame could converge to 1–2dB).

[0008] From an implementation perspective, this type of "controlled injection for DirAC" establishes a closed loop of "analysis → injection → decoding": the front end remains a stable DoA / diffusion estimation, the middle section uses angle difference weighting and budget limiting to write back the path-level energy, and only at the end does it enter the specific rendering and playback. Compared with simple geometric composition or channel domain reweighting, it is easier to find a balance between physical reachability, perceptual focus, and engineering real-time; moreover, because the input is an "estimator-independent" set of paths, it can maintain loose coupling with existing COMPASS, DirAC front end, and AllRAD / VBAP / binaural back end, making it easy to implement as a plug-in in VR / AR and EVS-IVAS links. It should be noted that in high diffusion scenarios, the settings of ρ and n still include empirical priors, and we currently tend to set ρ... max Keep it between 0.3 and 0.5 to avoid over-injection. Summary of the Invention

[0009] The goal of this invention is to achieve direction-selective and energy-controlled early reflection injection and rendering within the DirAC parameter domain, for a given set of early reflection paths, while maintaining estimator decoupling and backend consistency. Typically, only time-frequency DoA and diffusion are analyzed and allocated, without outputting the set of injectable paths, and there is a lack of mechanisms for quantitative budgeting and energy conservation in the FOA / DirAC analysis domain. In strong diffusion or multi-source scenes, directional flicker, energy overflow, and color distortion easily occur, and maintaining equivalence of direction and energy across backends (AllRAD / VBAP / binaural) is difficult. The specific process is as follows:

[0010] S1: Perform a Short Time Fourier Transform (STFT) on the input first-order immersive audio (FOA) signal to obtain the time-frequency active intensity vector ||I(f,t)|| and smoothing energy, and calculate the direction of arrival. The diffusion ψ(f,t) is related to the frequency jitter. To suppress time-frequency jitter, IIR / exponential smoothing can be used to smooth I(f,t) and energy in the time domain / band, thereby improving the efficiency. The stability of the window is not limited to a specific window length or step size. In practice, a Hann window can be selected, and the number of FFT points can be 1 to 2 times the window length.

[0011] S2: Receive the early reflection path set PathSet = {(r k ,τ k ,g k )}, where r k τ represents the path direction (unit vector). k For arrival delay, g k The path gain includes distance attenuation, material absorption, and optional sensing index tuning. Early reflection paths are unrestricted and can originate from geometric tracing, statistical generation, or manual annotation, coupled with specific DoA / diffusion estimators (such as DirAC / COMPASS).

[0012] S3: Set of output directions in the FOA / DirAC domain Construct directional encoded impulse responses one by one; for any path r k With output direction v i The included angle θ ki Calculate the angle difference soft weight:

[0013]

[0014] Normalize the weighted energy of each channel to obtain Based on τ k In each channel, a short window kernel (preferably a Hann window) is formed to create a channelized τ. k And it is aligned with the time delay. The earlyIR can be a set of several sub-cores that are "path-bound" or a single sub-core that is "channel-merged," depending on the injection mode. To ensure that "focus tightens when direction is clear and relaxes when diffusion is high," the direction focus index is adjusted according to...

[0015] n(f,t)=n0+λ·(1-ψ(f,t))

[0016] Adaptive testing is performed, with n0∈[1,3] and λ∈[0,5]. Higher n is selected for the high-frequency sub-bands with a correspondingly lower injection budget, while the low-frequency sub-bands are relaxed to balance localization and tonal naturalness.

[0017] S4: Calculate the time-frequency injection budget -ρ(f,t) based on ψ(f,t) and ||I(f,t)||, and apply the energy limiting inequality:

[0018]

[0019] Where E in (f,t) represents the energy of the cell at the input FOA. The preferred budget function employs hard limiting and adjustable mapping.

[0020] ρ(f,t)=min{ρ max ,a(1-ψ(f,t)) b +c·‖I(f,t)‖}

[0021] Where ρ max ∈[0.2,0.6], and a,b,c are non-negative calibration coefficients. In implementation, this can be achieved by... Uniform scaling is used to ensure that all time-frequency units meet the amplitude limiting constraints. ρ(f,t) is appropriately reduced at high frequencies to control color enhancement and brightness increase, while low frequencies are relaxed to preserve spatial thickness.

[0022] S5: Convolve the channelized earlyIR and the multi-channel directional signals obtained from DirAC decoding (preferably implemented using FFTOLA / OLS) to inject early reflections, and adaptively schedule them among path-level, hybrid, and merging modes. Path-level mode retains the independent earlyIR and convolution for each path, offering the highest directional resolution but high computational cost; merging mode merges all paths into a single earlyIR within the channel before unified convolution, resulting in the lowest computational cost; hybrid mode retains the top-N energy paths per frame for path-level injection, merging the rest within the channel (N can be adaptively adjusted by ρ load). To reduce the comb effect in near-direction areas, when τ k When approaching the direct term, phase processing strategies such as minimum phase compensation, window center staggering, and phase peak shifting can be employed. When ψ(f,t)≥ψ hi (e.g. ψ) hi When the confidence level is below the threshold (∈[0.7,0.9]) or the DirAC confidence level is below the threshold, the system automatically switches to merging mode to ensure that directional flickering and distortion are not introduced in strong diffusion / low confidence scenarios.

[0023] S6: Output direction set {v i This can be mapped to backends such as AllRAD / VBAP / binaural rendering. To maintain consistent perception of energy and direction, a small-range energy equivalence table or real-time normalization matrix can be configured to control the energy deviation of the three types of backends at the main lobe pointing point within ±1 to 2 dB; in the binaural path, {v i The corresponding channels can be further stacked using HRTF convolution.

[0024] Because of the adoption of the above technical solution, the present invention has the following advantages:

[0025] 1. This method is a "controlled injection" approach for the DirAC domain, with optimization objectives consistent with physical behavior. Instead of linear superposition after channel decoding, this method performs directional coding and energy allocation within the FOA / DirAC domain; it utilizes hard clipping and soft projection of the frame-level budget ρ(f,t). The monotonic directional selectivity binds the optimization goal of "improving directional clarity" to the FOA energy constraint, which is different from the schemes that rely solely on geometry or back-end channel superposition (such as pure ISM or DirAC-only).

[0026] 2. This method automatically tightens or loosens according to diffusion, intensity, and signal-to-noise ratio. The directional index n(f,t) = n0 + λ·(1-ψ(f,t)) is adaptively adjusted with diffusion; the budget function ρ(f,t) = min{ρ max ,a(1-ψ(f,t)) b The strength is incorporated into the gating. This allows for adaptive updates of the injection quota and focus based on channel SNR and scene diffusion, providing good adaptability to different array layouts and room material variations.

[0027] 3. This method features a one-to-one mapping between parameters and effects, facilitating easy localization and optimization. Angle difference weights, budget upper limits, and frequency band strategies (increasing n and decreasing ρ at high frequencies) are all publicly disclosed using a family of functions and their ranges, and the energy from the path to the channel is normalized. Ensure that "how much energy a path contributes" can be directly audited; analyze the performance of the method by combining indicators such as directional energy heatmap / directional entropy / DoA hit rate.

[0028] 4. Significantly improved objective indicators with controllable safety. In representative simulations and spline tests, compared with DirAC-only and ISM baselines, this method can improve DoA hit rate, reduce directional entropy, and increase energy concentration; simultaneously, through... The hard limit ensures that the proportion of injected energy is controllable, avoiding colorization and over-enhancement. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be introduced below. The accompanying drawings described below are only some embodiments recorded in the embodiments of the present invention and are not to scale. The number and position of the components can be adjusted according to the application.

[0030] Figure 1 This is a schematic diagram of the overall process of the system of the present invention;

[0031] Figure 2 is a schematic diagram of the overall online operation process of this invention; it compares the channel energy distribution of different methods. (a) DirAC baseline; (b) ISM; (c) the method of this invention; (d) the normalized dB difference between ISM and this invention. All heatmaps are normalized to a uniform maximum value; a spherical grid of 300×300 with σ=1.0 is used for smoothing, and the rendering is done from the front view; to improve readability, channels 4 and 5 are omitted in the illustration.

[0032] Figure 3This diagram shows the RMS comparison of the original signal and the enhanced signal, as well as the injected energy increment diagram. The solid blue line represents the RMS value of the original signal, the dashed red line represents the RMS value of the enhanced signal, and the orange shaded area represents the injected energy increment (Δ) per frame, reflecting the energy injection of the enhanced signal over time. The overall injected energy percentage is 1.40%. The left Y-axis represents the RMS amplitude (au), used to measure the signal's energy intensity. The solid blue line and dashed red line in the diagram represent the energy intensity of the original signal and the enhanced signal, respectively. The right Y-axis represents the injected RMS increment (Δau), i.e., the energy increment injected into the signal during the enhancement process. The orange shaded area shows the energy contribution after signal enhancement. RMS represents the signal's energy, while the injected RMS increment (Δ) represents the energy increment injected after enhancement processing.

[0033] Figure 4 This diagram compares the directional entropy in this invention, showing the changes in directional entropy H(t) under three configurations. The blue solid line represents the signal enhanced using the traditional ISM method (ISM), the red solid line represents the signal enhanced using the proposed method, and the black dashed line represents the original signal without enhancement. The horizontal axis represents time (s), indicating the signal's temporal variation in seconds (s). The vertical axis represents the directional entropy H(t), reflecting the signal's spatial selectivity and directional consistency; a lower entropy value indicates stronger directional focusing. The results show that compared to the traditional ISM-enhanced signal, the method of this invention significantly reduces directional entropy, indicating that the enhanced signal has stronger directional focusing and spatial selectivity. The black dashed line shows the directional entropy of the original signal as a comparison benchmark. Detailed Implementation

[0034] The present invention will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.

[0035] Figure 1 This is a flowchart of the overall system of the present invention. The process includes: obtaining the time-frequency direction of arrival through DirAC analysis. With diffusion ψ(f,t); access to the set of paths {(r) including path direction, delay, and gain} k ,τ k ,g kBased on the path-output direction angle, calculate the angle difference soft weight and form a channelized directional encoded impulse response; S4 calculates the frame-level injection budget ρ(f,t) based on the diffusion and acoustic intensity and performs energy limiting; S5 convolves the impulse response with the DirAC decoded multi-channel signal to complete the injection, and schedules the output between path-level / merging / mixing modes; finally, it is mapped to AllRAD / VBAP or binaural back-end. The specific steps are as follows:

[0036] S1: (DirAC Analysis): Perform STFT (window length 1024, hop 512, Hann window) on the input FOA to obtain B(f,t); calculate the active intensity I and frame energy E. in Based on the DirAC front-end ψ(f,t), perform first-order smoothing (α) on ψ ψ ∈[0.7,0.98]).

[0037] S2: (Path set access): Receive set {(r k ,τ k ,g k Sources allowed: geometric tracing, statistical models, or offline annotation; coordinates are uniformly represented as unit vectors r. k ∈S 2 , τ k In seconds, g k ≥0 includes geometric dissipation and material absorption.

[0038] S3 (Direction Encoding and Angular Aberration Soft Projection): Preset Output Direction Set (Virtual speaker or binaural direction) Calculate angular difference Adaptive direction exponent:

[0039] n(f,t)=n0+λ·(1-ψ(f,t))

[0040] Where ψ(f,t)∈[0,1] is the DirAC diffusion degree, ψ→0 indicates strong orientation, and ψ→1 indicates strong diffusion. Parameter n_0 is the baseline focusing, preferably 1-3; λ controls the variation with ψ, preferably 0-5. To avoid jitter in the frequency band and time dimension, a first-order time-domain smoothing (α) can be applied to n(f,t). n ∈[0.7,0.98]) and / or subband constraints (high-frequency subbands raise the upper bound of n, e.g., n max =8~10); Assign weights to each output direction:

[0041]

[0042] Then, normalization and energy allocation are performed:

[0043]

[0044] With τ k Centered on the center, window length T = round(f s ·τ max (Default 50ms) Construct a Hann window and then obtain the impulse response:

[0045]

[0046] S4 (Frame-level Budget and Energy Limit): To ensure that the injected energy does not overflow and is consistent with perception, this embodiment sets a budget ρ(f,t) for each time-frequency unit, which is a non-subtractive combination of diffusion and intensity, and is subject to a hard upper limit ρ. max constraint:

[0047] ρ(f,t)=min{ρ max ,a(1-ψ(f,t)) b +c·‖I(f,t)‖}

[0048] Where ρ max Given ρ ∈ [0.2,0.6], a ∈ [0,1], b ∈ [0.5,3], c ∈ [0,1], and perform time-domain smoothing on ρ (α ρ ∈[0.5,0.98]), without changing the technical effect, ρ(f,t) can be replaced by any linear or sublinear combination of a monotone non-decreasing function of (1-ψ) and an intensity term (such as...). All of these are equivalent implementations of the present invention; the energy amplitude inequality is applied:

[0049]

[0050] If the limit is exceeded, all will be scaled proportionally. Therefore, update h k (t,i).

[0051] S5 (Convolutional Injection and Pattern Scheduling): Path-level injection pattern, preserving h k (t,i) Independent convolutions; merge injection patterns, aggregated into h by channel. early (t,i)=∑ k h k (t,i) post-convolution; hybrid injection mode, retaining the top-N (default N=5) energy per frame for path-level processing, merging the remaining channels; when ψ≥ψ hi If the DirAC confidence level is below the threshold, switch to merge injection mode.

[0052] Experimental setup and results

[0053] To further verify the technical effects of the present invention, the embodiments are described in conjunction with the accompanying drawings. The experimental environment and common parameters are as follows: The test was conducted in a 10×8×3 m rectangular room, with the receiving point located at the geometric center at a height of 1.2 m, and the incident DoA fixed at 0° azimuth and 0° elevation; both the ISM and the present invention use a third-order reflected signal as the injection; the signal processing is uniformly set to a sampling rate of 48kHz, an STFT window length of 1024, a hop of 512, and a Hann window; the output direction is a 9-channel virtual loudspeaker (consistent with the DirAC decoding direction); the path set of S2 access {(r k ,τ k ,g k In )}, r k Normalized to a unit vector, τ k In seconds, g k The algorithm already incorporates geometric scattering and material absorption. Candidate paths are derived from DoA-guided inverse tracing and selected through bi-angle filtering. The directional exponent is adaptively set to n0 = 2 and λ = 3 by default, and a first-order temporal smoothing (α) is applied to n(f,t). n =0.9) and sub-band upper limit constraint (high-frequency sub-band can take n) max =8~10), and finally execute clip([n min ,n max ]); Frame-level budget and limiting default parameters are ρ max =0.4, a=0.25, b=1.2, c=0.15, ρ(f,t) is smoothed in the time domain (α) ρ =0.9), which can reduce ρ in the high-diffusion subband. max (e.g., 0.2–0.4); Injection scheduling supports path-level (independent convolution), merging (convolution after channel aggregation), and hybridization (path-level injection of the top-N energy paths in each frame, merging of the remaining channels, default N=5), when ψ≥ψ hi Alternatively, if the DirAC confidence level is below the threshold, it can fall back to merge injection.

[0054] Figure 2 compares the spatial energy directivity of different methods (DirAC-only, ISM, and this invention), using "channel direction - channel energy" as the basic element for angular domain aggregation and rendering of channel energy heatmaps. This aperture is consistent with the injection and decoding structure of this invention (each output channel corresponds to a fixed unit direction vector v). i This avoids the "geometric-perceptual misalignment" bias caused by directly aggregating along the path direction. First, a set of output directions consistent with DirAC decoding is determined. Each channel corresponds to a unit direction vector v iIts azimuth / elevation angle is consistent with the system's default 9-channel layout; during the calculation phase, all channels can be retained to maintain energy conservation, and during the display phase, to improve readability, channels 4 and 5 can be omitted only at the display level (as shown in the figure caption). Subsequently, the total energy E of each channel is calculated for each of the three methods (DirAC-only, ISM, and this invention). i If assessing "early reflection distribution", then take The integral is windowed over a period of [0, 50 ms]; if evaluating the "overall orientation distribution", the final output is used. And integrate throughout. To construct a continuous directional energy density on the sphere, a 300×300 unit spherical grid is established, and the center direction of the grid is denoted as γ. Each "discrete channel energy" is then integrated along its center direction γ. i Using a kernel as the core, through diffusion and superposition of a monotonically decreasing angular kernel function with a single peak, we obtain S(γ)=Σ i E i ·K(∠(v i ,γ)), where ∠(v i ,γ)=cos -1 (v i ·γ), the recommended Gaussian kernel is K(θ)=exp(-θ 2 / 2σ 2 Or an equivalent vMF kernel, σ / κ is calibrated by the virtual speaker angle and the desired half-power angle (empirical practice: make the weights of adjacent channels in the middle direction approximately 0.5 each to ensure neither excessive smoothing nor excessive discretization). For horizontal comparability, S(γ) obtained by the three methods is combined with the global maximum value S. max Unification and normalization And fix the color stops and unify the forward viewing direction for rendering; to avoid color band jumps caused by minimum values, percentile clipping (such as 95th percentile normalization) can be performed before rendering and a minimum display threshold can be set. Difference numerator diagram 2(d) uses... The calculations show that positive values ​​correspond to the energy selectivity enhancement of this invention relative to ISM in that direction (expected to focus on the forward sector of the DoA neighborhood), while negative values ​​correspond to the suppression of off-axis diffused energy. The above process strictly aggregates in the corner domain with "direction and energy of each output channel" as the basic granularity, maintaining consistency with decoding / rendering while ensuring comparability, reproducibility, and layout readability of the three methods through unified mesh, kernel width calibration, global normalization, and fixed color scales. Figure captions should clearly indicate whether channels are omitted in the "calculation / display" stages to avoid ambiguity.

[0055] Figure 3To quantify the global energy proportion injected during early reflections, a frame-level root mean square (RMS) trajectory comparison method is employed. Specifically, the time-domain signal of each output channel is framed according to the STFT synchronization parameters (sampling rate 48kHz, window length N=1024 points, hop=512 points, Hann window). The frame RMS is calculated in the time domain, and the intra-channel RMS is then averaged across all channels to obtain the system RMS trajectory. The baseline signal containing only directional components is denoted as... The signal after injection is The RMS of frame t is defined as follows:

[0056]

[0057] in These are the baseline and enhanced frame sample sequences after channel averaging, respectively. To avoid negative increments caused by numerical noise, the frame-level injection increment is defined as:

[0058] ΔRMS(t)=max(0,RMS enh (t)-RMS org (t))

[0059] Considering that energy is proportional to the square of RMS, the ratio of the incremental square to the total enhancement square is used as the "injection percentage" (percentage):

[0060]

[0061] Figure 3 The shaded region is the envelope of ΔRMS(t), and the curves correspond to RMS. org (t)(blue solid line) and RMS enh (t)(red dashed line). Under the default parameters of this invention (see S4 budget-limiting configuration), the statistically obtained injection ratio is approximately 1.40%, which is consistent with the frame-level budget ρ(f,t) determined by the diffusivity ψ(f,t) and the active intensity ||I(f,t)|| in S4, indicating that the energy limiting mechanism is effectively effective: it satisfies the following for any time-frequency unit. This results in a limited and stable RMS injection ratio after time aggregation.

[0062] Figure 4 To measure the concentration of spatial energy in each output direction, this embodiment calculates the channel probability distribution for each analysis frame and obtains the Shannon entropy trajectory H(t) accordingly. Let x be the temporal sample of the i output directions (channels) in the t-th frame. i (t,n) (frame length N = 1024, hop = 512, Hann window, consistent with the aforementioned STFT parameters), first calculate the frame energy of each channel and normalize it to the directional probability:

[0063]

[0064] Where ε = 10 -12 To protect against division by zero and numerical stability, if some channels (such as channels 4 / 5) are omitted from the visualization of Figure 2 for ease of observation, the same channel set should be used in all methods for entropy calculation (all N channels are recommended). out (Channels) to ensure statistical comparability between different methods. Then, frame-level directional entropy is defined using a binary logarithm:

[0065]

[0066] ε = 10 -12 Same as above, used for p i Logarithmic protection is applied to the term with the extremely low probability of H(t)→0. To suppress inter-frame jitter, a moving average smoothing of length 8 frames is applied to H(t). The physical meaning of entropy is: at a fixed N... out The lower the value of H(t), the more concentrated the energy is in a few directions (stronger spatial selectivity and more precise directionality); conversely, the higher the value of H(t), the more uniform and diffuse the energy distribution. In the demonstration, the three curves correspond to DirAC-only (black dashed line), ISM (blue solid line), and the present invention (red solid line), respectively. The curve of the present invention is generally lower than that of the control method, indicating that under the same output channel layout and signal conditions, the present invention can significantly reduce directional entropy and enhance spatial focusing.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention, and all such modifications or substitutions should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for early reflection energy controlled injection based on DirAC parameters, characterized in that, Comprising: S1: Obtain time-frequency direction of arrival by DirAC framework first order ambisonic (FOA) signal analysis with the diffuseness ψ(f, t); S2: receive a path set P containing path directions r k , time delays τ k , and gains g k , which can be generated by any front-end; S3: construct r k corresponding directional encoding impulse response; for the kth path and output direction v i θ, the angle between the kth path and output direction v ki , compute the angular difference weight: and form the channelized early impulse response accordingly; S4: Compute frame-level injection budget p(f, t) based on ψ(f, t) and acoustic intensity or signal-to-noise ratio, such that each time-frequency cell satisfies the energy clipping inequality: where E in (f, t) is the energy of the input FOA or a smoothed estimate thereof; S5: convolve the early impulse response with the DirAC decoded multi-channel signal to inject early reflections, depending on ψ(f,t) and the computational load, schedule output among path-level, hybrid or combined mode.

2. The method of claim 1, wherein In the method of S3, the direction focusing index n(f,t) is adaptive to the diffusion degree ψ(f,t), satisfying: n(f,t) = n0 + λ·(1-ψ(f,t)) Where n0∈[1,3], λ∈[0,5].

3. The method of claim 1, wherein In the method of S4, the frame-level injection budget ρ(f,t) is determined by the diffusion degree and the intensity, satisfying: p(f, t) = min{p max , a(1 - ψ(f, t)) b + c · ||I(f, t)||} where p max ∈ [0.2, 0.6], and a, b, c are non-negative calibration coefficients, I(f, t) is the sound intensity intensity.

4. The method of claim 1, wherein In the methods of S3 and S4, higher n values are used in high-frequency subbands and ρ(f,t) is reduced, and in low-frequency subbands, it is relaxed.

5. The method of claim 1, wherein, The mixing injection mode in the method S5 refers to reserving the paths with the top N energy in the path level mode for separate injection, and merging the remaining paths by channel. When ψ ≥ ψ hi or the DirAC confidence is lower than a threshold, switching to the merged injection mode.

6. The method according to claim 1, wherein the impulse response of each path is windowed with a window function of duration T, preferably a Hanning window, and τ k max , preferably τ max = 50 ms.​​​