Conference acoustic simulation method and system based on virtual sound source method
By generating acoustic impulse responses using the virtual sound source method and establishing parameter correlation models, the parameter combination of multi-channel equalizers and feedback suppressors is optimized, solving the coupling optimization problem between speech intelligibility and system stability in complex conference halls, and realizing efficient acoustic simulation and parameter configuration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU ELITE DIGITAL TECH CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing acoustic simulations based on the virtual sound source method cannot effectively solve the coupling optimization problem between the gain of multi-channel equalizers and the notch frequency of feedback suppressors in conference halls with complex geometries, making it difficult to balance speech clarity and system stability.
The acoustic impulse response of the target acoustic measurement point in the conference room is generated using the virtual sound source method. A correlation model between the parameters of the adjustable parameter module and the acoustic performance index is established. The parameter combination of the multi-channel equalizer and feedback suppressor is optimized through a collaborative optimization algorithm to achieve the optimal synergy between speech intelligibility and system stability.
It achieves a collaborative optimal balance of maximizing voice clarity while ensuring system stability, improves the scientific nature and engineering efficiency of conference system design, and can adapt to real-time changes in the conference venue.
Smart Images

Figure CN122113693A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent conference system technology, specifically to a conference acoustic simulation method and system based on the virtual sound source method. Background Technology
[0002] In conference halls with complex geometries (especially international conference halls with curved walls and tiered seating), when using multi-speaker, multi-microphone conference systems, existing acoustic simulations based on the virtual source method can only output physical acoustic parameters such as sound pressure level and reverberation time. They cannot solve the coupling optimization problem between the gain of the multi-channel equalizer and the notch frequency of the feedback suppressor. As a result, the preset DSP parameters cannot simultaneously meet the dual requirements of improving speech clarity and ensuring feedback stability in actual operation.
[0003] Specifically, the sound focusing effect caused by the curved wall surface leads to excessively high sound pressure levels in certain frequency bands (such as the critical 2kHz-4kHz speech band) in local areas, requiring frequency attenuation through a multi-channel equalizer to correct the frequency response. However, attenuation of this frequency band directly reduces speech intelligibility, especially noticeable in shadowed areas obscured by tiered seating. Simultaneously, the distribution of multiple microphones makes this frequency band highly susceptible to feedback howling due to speaker-microphone coupling. Traditional design methods set equalization parameters and feedback suppression frequencies independently, neglecting their coupling relationship in the frequency domain: changes in the equalized frequency response alter the gain distribution of the feedback path, potentially making previously stable frequencies unstable or causing preset notch filters to fail. Existing virtual source simulations only provide static impulse responses and lack the ability to model the equalization-feedback coupling effect. Designers are forced to rely on trial and error based on experience, which is not only inefficient but also makes it difficult to obtain the optimal parameter configuration that balances sound quality and stability globally. Summary of the Invention
[0004] The purpose of this invention is to provide a conference acoustic simulation method and system based on the virtual sound source method to address the shortcomings in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a conference acoustic simulation method based on the virtual sound source method, applied to a conference system equipped with an audio processor, wherein the audio processor includes at least two adjustable parameter modules, including:
[0006] S1. The acoustic impulse response of the target acoustic measurement point in the conference room is generated using the virtual sound source method, and the audio processor is a programmable digital signal processor;
[0007] S2. Based on the acoustic impulse response, establish a correlation model between the parameters of at least two adjustable parameter modules in the audio processor and the acoustic performance indicators;
[0008] S3. Using the parameters of the at least two adjustable parameter modules as decision variables and the acoustic performance index as the objective function, a collaborative optimization algorithm is used to find the parameter combination that optimizes the overall performance of the conference system.
[0009] S4. Convert the parameter combination into configurable parameters of the audio processor and output them.
[0010] In a preferred embodiment, the at least two adjustable parameter modules include a multi-channel equalizer and a feedback suppressor, wherein the multi-channel equalizer is used to adjust the gain of different frequency bands, and the feedback suppressor is used to set notch parameters to suppress acoustic feedback.
[0011] In a preferred embodiment, the acoustic performance indicators include a first performance indicator and a second performance indicator, wherein the first performance indicator is a speech intelligibility indicator and the second performance indicator is a system stability margin indicator.
[0012] In a preferred embodiment, establishing the association model in S2 further includes:
[0013] Based on the acoustic impulse response, establish the feedback path transfer function matrix H(f) between multiple loudspeakers and multiple microphones;
[0014] Based on the feedback path transfer function matrix H(f), the gain parameter G of the multi-channel equalizer, and the notch parameter N of the feedback suppressor, the system open-loop transfer function L(f) = G(f)·N(f)·H(f) is constructed.
[0015] Where G(f) is the equalizer frequency response function generated by the gain parameter G, and N(f) is the feedback suppressor frequency response function generated by the notch parameter N;
[0016] The system stability margin index is calculated based on the open-loop transfer function L(f) of the system.
[0017] Based on the acoustic impulse response, the gain parameter G of the multi-channel equalizer, and the notch parameter N of the feedback suppressor, the equivalent impulse response after taking into account the effects of the equalizer and the feedback suppressor is calculated, and then the speech intelligibility index is calculated.
[0018] In a preferred embodiment, the collaborative optimization algorithm in S3 employs a game-theoretic optimization model, including:
[0019] S31. Define the gain parameter G of the multi-channel equalizer as the strategy set of the first player, and define the speech clarity index as the utility function of the first player. (G, N), where the utility function is used to quantify the payoff of the first player under the parameter combination (G, N); the notch parameter N of the feedback suppressor is defined as the strategy set of the second player, and the system stability margin index is defined as the utility function of the second player. (G, N), the utility function is used to quantify the payoff of the second player under the parameter combination (G, N);
[0020] S32. An iterative optimization is performed using a turn-based game mechanism. In the t-th round of iteration, the first player adopts the current feedback suppressor strategy. Choose to use under fixed conditions Maximization strategy The second player's current equalizer strategy Choose to use under fixed conditions Maximization strategy ;
[0021] S33. Repeat step S32 until the strategy converges, obtaining the optimal combination of equalizer gain parameters and feedback suppressor notch parameters. , );
[0022] The optimal collaborative combination satisfies the Nash equilibrium condition: for the first player, in When fixed, It is its utility function The maximum point; for the second player, in When fixed, It is its utility function The maximum value point.
[0023] In a preferred embodiment, the strategy selection in S32 employs a virtual game algorithm, including:
[0024] The first player chooses based on the average of the second player's historical strategies. Maximization strategy;
[0025] The second player chooses based on the average of the first player's historical strategies. Maximization strategy.
[0026] In a preferred embodiment, the strategy selection in S32 employs either the strategy hill-climbing method or the particle swarm optimization algorithm.
[0027] In a preferred embodiment, the generation of the acoustic impulse response using the virtual sound source method in S1 further includes:
[0028] For curved walls, a dynamic adaptive segmentation strategy is adopted, which dynamically determines the segmentation granularity based on the relative position of the sound source and the receiving point, and only generates virtual sound sources that contribute to the current receiving point.
[0029] For early reflections of order 1 to N, the virtual sound source method is used for accurate calculation, and for later reflections above order N, the ray tracking method is used for statistical calculation, where N is a preset reflection order threshold.
[0030] Based on the occlusion relationship of the tiered seating area, a three-dimensional occlusion model is established to pre-judge the visibility of virtual sound sources. Invalid virtual sound sources that are occluded are then eliminated through a reverse tracking algorithm.
[0031] In a preferred embodiment, a dynamic update step is also included:
[0032] S5. Acquire dynamic change information of the meeting venue in real time or near real time, including changes in the number of participants, adjustments to the table and chair layout, and changes in the speaker's position;
[0033] S6. Based on the dynamic change information, trigger incremental recalculation of S1 to S4, take the equilibrium solution of the previous game as the initial strategy of this iteration, use transfer learning to accelerate convergence, and generate updated configurable parameters for the audio processor.
[0034] This invention also provides a conference acoustic simulation system based on the virtual sound source method, comprising:
[0035] Geometric modeling module: used to acquire and store the geometric structural parameters and material acoustic parameters of the conference room;
[0036] Virtual sound source simulation module: used to generate acoustic impulse responses of various target measurement points in the conference room using the virtual sound source method;
[0037] Correlation modeling module: used to establish a correlation model between the parameters of at least two adjustable parameter modules and acoustic performance indicators based on the acoustic impulse response;
[0038] Collaborative optimization module: used to solve for the optimal combination of parameters that optimizes the overall performance of the conference system by using the parameters of the at least two adjustable parameter modules as decision variables and the acoustic performance index as the objective function;
[0039] Parameter mapping module: used to convert the parameter combination into configurable parameters for the audio processor and output them.
[0040] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0041] This invention is the first to model the parameter optimization problem of equalizer and feedback suppressor as a two-player game problem. By defining a strategy set and utility function and using a turn-based game mechanism to solve the Nash equilibrium, it fundamentally solves the coupling optimization problem of mutual interference between the two and difficulty in balancing speech intelligibility and system stability, as described in the background technology. It achieves a cooperative optimal balance that maximizes speech intelligibility while ensuring system stability.
[0042] This invention establishes a correlation model from the impulse response of the virtual sound source method to performance indicators, and optimizes the computational efficiency of the virtual sound source for curved walls and tiered seating. Finally, it directly converts the cooperative optimal parameters into configurable parameters for the audio processor, bridging the semantic gap between simulation results and engineering parameters. At the same time, it introduces dynamic update and transfer learning mechanisms, enabling the system to adapt to real-time changes in the meeting environment, significantly improving the scientific nature, engineering efficiency, and dynamic adaptability of the meeting system design. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0044] Figure 1 This is a flowchart of the method of the present invention.
[0045] Figure 2 This is a system block diagram of the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] To facilitate understanding of this invention, the following definitions are provided for specific technical terms appearing in the application documents:
[0048] Virtual source method: A computational method for modeling sound fields in enclosed spaces. Its principle is to treat the reflection of sound waves by reflective interfaces such as walls and ceilings as equivalent to the direct sound emitted by virtual sound sources symmetrical about those interfaces. By calculating the contributions of all visible virtual sound sources, the reflected sound patterns at the receiving point can be superimposed to construct the impulse response of the entire sound field.
[0049] Audio processor: refers to a programmable digital signal processor used to process audio signals in a conference system in real time, including at least adjustable parameter modules such as multi-channel equalizers and feedback suppressors.
[0050] Multi-channel equalizer: A filter bank used to adjust the gain of different frequency bands of an audio signal. Its gain parameter G is usually represented as a vector of gain values for each frequency band, such as G=[g1, g2, ..., gK], where gk is the gain value (in dB) of the k-th frequency band.
[0051] Feedback suppressor: An audio processing module used to suppress acoustic feedback howling. Its notch parameter N typically includes the notch center frequency, notch depth, and notch bandwidth, and can be expressed as N={(f1, d1, b1), (f2, d2, b2), ...}, where fl is the center frequency of the l-th notch, dl is the corresponding notch depth, and bl is the notch bandwidth.
[0052] Speech Proficiency Index (STI): Short for Speech Transmission Index, it's an objective metric used to quantify speech intelligibility. Its value ranges from 0 to 1, with higher values indicating higher speech proficiency. It's calculated based on the modulation transfer function method, by analyzing the modulation loss in each frequency band of the acoustic impulse response.
[0053] System stability margin: This metric measures the ability of a sound reinforcement system to resist feedback howling. It is usually taken as the minimum of the gain margin and phase margin, and is measured in dB. A higher value indicates a more stable system and a lower likelihood of howling.
[0054] Feedback path transfer function matrix H(f): A frequency domain transfer function matrix describing the acoustic feedback path between multiple loudspeakers and multiple microphones. Matrix elements... (f) represents the frequency domain transfer function of the acoustic feedback path from the i-th loudspeaker to the j-th microphone.
[0055] Game Theory Optimization Model: The parameter optimization problem of a multi-channel equalizer and a feedback suppressor is modeled as a two-player game. The equalizer acts as the first player, with speech clarity as its utility function; the feedback suppressor acts as the second player, with system stability margin as its utility function. The cooperative optimal parameter combination for both is obtained by solving the Nash equilibrium.
[0056] Strategy set: In a game optimization model, the set of all strategies that each player can choose. The strategy set of the first player is all possible values of the equalizer gain parameter G; the strategy set of the second player is all possible values of the feedback suppressor notch parameter N.
[0057] Utility function: In game optimization models, a function used to quantify the payoff of a player under a specific strategy combination. The utility function of the first player is: (G, N), where the value is the Speech Proficiency Index (STI(G, N)); the utility function of the second player is... (G, N) represents the system stability margin M(G, N).
[0058] Turn-based game mechanism: an iterative method for finding game equilibrium. In each round of iteration, the two players take turns choosing a strategy that maximizes their own utility function based on the other player's current strategy, and this process is repeated until the strategies converge.
[0059] Nash equilibrium: a core concept in game theory, referring to a strategy combination where, under this combination, unilaterally changing one player's strategy will not increase their utility. In this invention, the cooperative optimal combination (G, N) satisfies the Nash equilibrium condition.
[0060] Virtual game algorithm: an iterative algorithm for solving game equilibria. In each iteration, the player chooses the optimal response based on the statistical average of the opponent's historical strategies (rather than the current strategy), resulting in better convergence stability.
[0061] Dynamic adaptive segmentation strategy: A virtual sound source generation method used when processing curved walls. Based on the relative positions of the sound source and the receiver, the segmentation granularity of the curved wall is dynamically determined, generating only virtual sound sources that contribute to the current receiver, thus avoiding an explosion in the number of virtual sound sources.
[0062] Transfer learning: a machine learning technique that applies knowledge learned in a source domain to a target domain. In this invention, the equilibrium solution of the previous round of the game is used as the initial strategy for the current iteration, and transfer learning is used to accelerate the convergence process in dynamic scenarios.
[0063] Fast Fourier Transform (FFT): A mathematical transformation algorithm that converts a time-domain signal into a frequency-domain signal, used to obtain the frequency-domain transfer function H(f) from the impulse response h(t).
[0064] Argument arg: The phase angle of a complex number, in degrees or radians, used to calculate the phase of the open-loop transfer function.
[0065] Minimum value min: Takes the minimum value in a set of numbers.
[0066] Maximum value (max): Retrieves the maximum value from a set of numbers.
[0067] The multiplication symbol ∏ represents the product of multiple factors.
[0068] The summation symbol Σ: represents the sum of multiple numbers.
[0069] Exponential function exp: An exponential function with the natural constant e as its base, exp(x)=e^x.
[0070] Norm ||·||: A function that measures the size of a vector. In this invention, Euclidean distance is used, which is the square root of the sum of the squares of the vector's components.
[0071] Example 1, please refer to Figure 1 As shown in this embodiment, a conference acoustic simulation method based on the virtual sound source method is applied to a conference system equipped with an audio processor. The audio processor includes at least two adjustable parameter modules, including:
[0072] S1. The acoustic impulse response of the target acoustic measurement point in the conference room is generated using the virtual sound source method, and the audio processor is a programmable digital signal processor;
[0073] S2. Based on the acoustic impulse response, establish a correlation model between the parameters of at least two adjustable parameter modules in the audio processor and the acoustic performance indicators;
[0074] S3. Using the parameters of the at least two adjustable parameter modules as decision variables and the acoustic performance index as the objective function, a collaborative optimization algorithm is used to find the parameter combination that optimizes the overall performance of the conference system.
[0075] S4. Convert the parameter combination into configurable parameters of the audio processor and output them.
[0076] As described in S1-S4 above, by establishing a bridge between the acoustic simulation using the virtual sound source method and the parameter configuration of the conference audio processor, a direct mapping from the physical sound field to electroacoustic parameters is achieved. Based on the precise sound field impulse response, the system can establish a correlation model between the parameters of adjustable parameter modules such as equalizers and feedback suppressors and conference-specific performance indicators such as speech intelligibility and system stability margin. Furthermore, a collaborative optimization algorithm is used to simultaneously optimize the parameters of multiple adjustable parameter modules, solving the technical problem of mutual interference between equalizers and feedback suppressors and the difficulty in balancing intelligibility and stability in traditional methods. The final output parameter combination can be directly written into the audio processor's registers, realizing the automatic conversion from simulation results to engineering parameters, significantly improving the scientific rigor and efficiency of conference system design.
[0077] In one possible implementation, step S1 specifically includes:
[0078] S11. Obtain the geometric structural parameters and material acoustic parameters of the conference room.
[0079] Specifically, the first step is to obtain a precise geometric model of the conference room using architectural drawings or a 3D laser scanner. For large, stepped, curved international conference halls, the following information needs to be collected:
[0080] The radius of curvature, center position, and start and end angles of the curved wall; the height, depth, and row spacing of each row of seats in the tiered seating area; and the types of decorative materials used for each wall, ceiling, and floor (such as plasterboard, glass, wood sound-absorbing panels, carpet, etc.).
[0081] Material acoustic parameters include the absorption coefficient (usually provided in 1 / 3 octave or octave format) and scattering coefficient for each material. These parameters can be obtained by consulting a material acoustics handbook or by conducting impedance tube tests.
[0082] S12. Configure the location of the sound source and receiver.
[0083] Specifically, based on the actual deployment plan of the conference system, the position, directivity, and power parameters of multi-channel speakers, as well as the pickup positions and directivity patterns of multiple microphones (including gooseneck microphones, ceiling array microphones, etc.) are set in the 3D model. Simultaneously, target acoustic measurement points are evenly distributed in the audience area to evaluate the acoustic performance of each seat. For example, a measurement point can be set in the center of each seating area to ensure coverage of all seating areas.
[0084] S13. Generate acoustic impulse response using the virtual sound source method.
[0085] Specifically, an improved virtual sound source method is used for sound field calculation. This method is specifically optimized for complex geometric features such as curved walls.
[0086] For curved walls, a dynamic adaptive segmentation strategy is employed. Based on the relative positions of the sound source and receiver, the reflection point and reflection path length of the sound wave on the curved surface are calculated. Based on the path length and incident angle, it is dynamically determined whether the curved segment needs further subdivision. Only virtual sound sources that contribute to the current receiver are generated, avoiding the indiscriminate subdivision of the entire curved surface into numerous small flat panels, thus effectively suppressing the explosion of virtual sound source numbers.
[0087] For early reflections of orders 1-3, the virtual source method is used for accurate calculation. The system enumerates all possible combinations of reflection paths (including walls, ceilings, floors, etc.) and generates corresponding mirrored virtual sound sources. For each path, energy attenuation is calculated based on the sound absorption coefficient of the reflecting interface, and propagation delay is calculated based on the path length.
[0088] For late reflections of order 3 and above, a ray-tracking method is used for statistical calculation. A large number of sound rays (e.g., 100,000) are emitted from the sound source, and their propagation paths in space are traced. When a sound ray encounters a wall, the reflection direction (specular or diffuse reflection) is determined based on the material's scattering coefficient, and energy attenuation is recorded. When a sound ray enters a detection sphere with radius r centered on the receiver point, it is considered to have contributed to that receiver point. The radius r of the detection sphere is typically taken as 0.5 to 1.0 meters, and the specific value can be dynamically adjusted according to the size of the conference room: a larger value is used for large conference halls, and a smaller value is used for small conference rooms, to ensure a sufficient statistical sample size of sound rays, recording their arrival time, energy, and direction.
[0089] A three-dimensional occlusion model is established based on the occlusion relationship of the tiered seating area. Each row of seats is modeled as a sound barrier unit, and the visibility of virtual sound sources is pre-judged. A reverse tracing algorithm is used: starting from the candidate virtual sound source location, it traces backward to the real sound source and the receiving point. If the number of intersections between the tracing path and the tiered seating area exceeds a preset threshold, or the energy attenuation exceeds a preset threshold, the virtual sound source is determined to be an invalid virtual sound source and is removed.
[0090] Finally, the contributions from each path are summed to obtain the acoustic impulse response h(t) for each target measurement point. The impulse response is stored in the form of a digital sampling sequence, with the sampling rate typically set to 48kHz or 96kHz to meet audio processing requirements.
[0091] S14, Define the audio processor.
[0092] Specifically, the simulation target is a programmable digital signal processor (DSP). The processor model can be selected according to actual engineering requirements, such as ADI's ADSP-21489 or TI's TMS320C6748. The adjustable parameter modules integrated within the processor include a multi-channel equalizer, feedback suppressor, delay unit, and automatic mixer. This embodiment mainly focuses on the coordinated optimization of the multi-channel equalizer and feedback suppressor.
[0093] In one possible implementation, step S2 specifically includes:
[0094] S21. Establish the feedback path transfer function matrix H(f) between multiple speakers and multiple microphones.
[0095] Specifically, the acoustic impulse response generated based on S1 (t) (with L sampling points), the frequency domain transfer function is obtained through Fast Fourier Transform:
[0096] (fk)=FFT[ (tn)],k=0,1,…,N-1;
[0097] Among them, the number of FFT points N is usually taken as 2048 or 4096. If L < N, then zeros are padded to the end of (t) to N points; if L > N, then the first N points are intercepted. The frequency resolution Δf = fs / N, where fs is the sampling rate of the impulse response. Finally, the transfer function values (fk) at the discrete frequency points fk = k · Δf are obtained.
[0098] S22. Define the equalizer gain parameter G and the feedback suppressor notch parameter N.
[0099] Specifically, the equalizer gain parameter G is expressed in vector form:
[0100] G = [g1, g2,..., gK]
[0101] where K is the number of equalizer frequency bands, usually taken as 31 bands (1 / 3 octave) or 15 bands (2 / 3 octave). The value range of each gk is [-12dB, +12dB], and the step precision is preset as 0.5dB or 0.1dB according to the hardware capabilities of the processor.
[0102] The feedback suppressor notch parameter N is expressed in set form:
[0103] N = {(f1, d1, b1), (f2, d2, b2),...}
[0104] where fl is the center frequency of the l-th notch, with a value range of 20Hz to 20kHz and can be continuously adjusted; dl is the notch depth, with a value range of [-40dB, 0dB]; bl is the notch bandwidth, and optional preset modes include 1 / 3 octave, 1 / 6 octave, etc. The number of notches usually does not exceed 12 to avoid excessive damage to the sound quality.
[0105] S23. Construct the open-loop transfer function L(f) of the system.
[0106] Specifically, generate the frequency response function G(f) according to the equalizer parameter G. For each frequency band k, its frequency response is piecewise constant, and a continuous frequency response curve is obtained through interpolation and smoothing. Common interpolation methods include linear interpolation, spline interpolation, etc.
[0107] Generate the frequency response function N(f) according to the notch parameter N. The frequency response model of each notch can be expressed as:
[0108] Nl(f) = 1 - · exp ;
[0109] where dl is the linear value of the notch depth, and its relationship with the dB value is: = 1 - . For example, a -20dB notch corresponds to =0.9;
[0110] Multiplying the frequency responses of all notch filters yields the overall frequency response function of the feedback suppressor:
[0111] N(f) = ;
[0112] The system's open-loop transfer function is the product of these three factors:
[0113] L(f) = G(f)·N(f)·H(f);
[0114] The multiplication is a matrix multiplication, which yields a matrix of dimension I×J.
[0115] S24. Calculate the system stability margin index.
[0116] Specifically, the system stability margin is defined as the minimum of the gain margin GM and the phase margin PM:
[0117] M(G, N) = min(GM, PM);
[0118] Calculation of gain margin GM: Finding the phase crossover frequency Even if we get arg(L( The frequency point is -180°.
[0119] For multi-channel systems, the combined effect of all speaker-microphone paths must be considered. A worst-case analysis is typically used, taking the minimum gain margin among all paths.
[0120] GM= [-20 ];
[0121] Calculation of phase margin (PM): Finding the gain crossover frequency dB, even if |L( The frequency point where |dB)|=1. Similarly, take the minimum phase margin among all paths:
[0122] PM= [180°+arg( ( dB))];
[0123] Gain margin GM represents the maximum amount of gain (dB) that the system is allowed to increase before reaching critical stability. The system is stable when GM > 0, and the larger the GM, the better the stability.
[0124] Phase margin (PM) represents the amount of phase lag (in degrees) that a system is allowed to maintain before reaching critical stability. The system is stable when PM > 0, and PM > 30° is typically required to ensure sufficient stability margin.
[0125] For multi-channel systems, the combined effect of all speaker-microphone paths must be considered.
[0126] This embodiment employs worst-case analysis, taking the minimum gain margin and phase margin among all paths as the overall stability margin index of the system, to ensure that all possible feedback paths are within the stable range.
[0127] S25. Calculate the speech intelligibility index.
[0128] Specifically, the equivalent impulse response h'(t) is first constructed based on the original impulse response h(t) and the current equalizer and feedback suppressor parameters. The equivalent impulse response can be obtained by convolving h(t) with the time-domain impulse responses of the equalizer and feedback suppressor:
[0129] h'(t) = h(t) * g(t) * n(t);
[0130] Where g(t) is the equalizer impulse response (obtained from G through inverse Fourier transform), and n(t) is the feedback suppressor impulse response (obtained from N through inverse Fourier transform).
[0131] The STI is calculated using the modulation transfer function method. The specific steps include:
[0132] Divide h'(t) into 7 octave bands with center frequencies of 125Hz, 250Hz, 500Hz, 1kHz, 2kHz, 4kHz, and 8kHz. For each band, calculate its modulation transfer function MTF(fm):
[0133] MTF(fm) = ;
[0134] The formula is adopted, where T is the effective length of the impulse response, which is usually taken as the duration corresponding to the reverberation time RT60.
[0135] For the discretely sampled impulse response h'[n], the integral is transformed into a summation:
[0136] MTF(fm) = ;
[0137] Where fs is the sampling rate and N is the number of sampling points;
[0138] Where fm is the modulation frequency, ranging from 0.63Hz to 12.5Hz (a total of 14 modulation frequencies). The apparent signal-to-noise ratio for each frequency band is calculated based on the MTF:
[0139] =10· ;
[0140] right Limiting amplitude: If If the value is less than -15dB, then take... =-15dB; if >+15dB, then take =+15dB; otherwise = The modulation signal-to-noise ratio is obtained. .
[0141] The Speech Proficiency Index (STIk) for each frequency band is as follows:
[0142] STIk= ;
[0143] The final STI is the weighted average of STIk for each frequency band:
[0144] STI= ·STIk;
[0145] The weighting coefficients αk follow the IEC60268-16 standard and correspond to weights of 125Hz to 8kHz: [0.13, 0.14, 0.11, 0.12, 0.19, 0.17, 0.14].
[0146] The Speech Proficiency Index (STI(G, N)) is the STI value calculated based on the current G and N parameters.
[0147] Through S21-S25, a correlation model was established between equalizer parameters G, feedback suppressor parameters N, system stability margin M(G,N), and speech intelligibility STI(G,N), providing a mathematical basis for subsequent collaborative optimization.
[0148] In one possible implementation, step S3 specifically includes:
[0149] S31. Construct a game theory optimization model.
[0150] Specifically, the gain parameter G of the multi-channel equalizer is defined as the strategy set of the first player, and the speech clarity index STI(G, N) is defined as the utility function of the first player. (G, N). This utility function quantifies the payoff of the first player under the parameter combination (G, N), with a larger value indicating higher speech clarity.
[0151] The notch parameter N of the feedback suppressor is defined as the strategy set of the second player, and the system stability margin M(G, N) is defined as the utility function of the second player. (G, N). This utility function is used to quantify the payoff of the second player under the parameter combination (G, N). The larger the value, the more stable the system.
[0152] The goal of the game is to find a strategy combination (G, N) such that neither side can increase their payoff by unilaterally changing their own strategy, i.e., to achieve Nash equilibrium.
[0153] S32. Iterative optimization is achieved by using a round-robin game mechanism.
[0154] Specifically, the initialization strategy ( , The default value is usually used (e.g., the gain of all frequency bands of the equalizer is 0dB, and all notches of the feedback suppressor are turned off) or the configuration saved last time.
[0155] In the t-th iteration (t=1, 2, 3, ...), the following operations are performed sequentially:
[0156] The first player uses the current feedback suppressor strategy. Under fixed conditions, choose to make Maximization strategy :
[0157] = (G, );
[0158] argmax: The value of the independent variable that maximizes the objective function. For example, y = f(x) represents the value of x that maximizes f(x);
[0159] This optimization can be achieved through various algorithms, such as grid search, gradient ascent, and particle swarm optimization. Considering computational efficiency, gradient ascent is typically used: [Calculation...] For the gradient of G, update G along the gradient direction until it converges to a local optimum.
[0160] The second player's current equalizer strategy Under fixed conditions, choose to make Maximization strategy :
[0161] = ( (N);
[0162] Since N contains discrete variables (number of notch filters, notch frequency) and continuous variables (notch depth), optimization is quite complex. A hybrid optimization strategy is usually adopted: first, potential howling frequencies are determined through peak search, and then the notch depth is continuously optimized.
[0163] S33. Determine if convergence has occurred and output the optimal collaborative combination.
[0164] Specifically, repeat step S32 until the policy converges. The convergence condition is that the change in the policy between two consecutive rounds is less than a preset threshold.
[0165] Δ=|| - ||+|| - ||<ε;
[0166] Where ||·|| represents the Euclidean distance between the vectors. For the equalizer gain parameter vector G, || - ||= ;
[0167] For the notch parameter set N of the feedback suppressor, it is necessary to first convert it into a fixed-length eigenvector, and then calculate the Euclidean distance.
[0168] ε is a preset threshold value, which can be set to 0.1dB or 0.01, etc., according to the engineering accuracy requirements.
[0169] For example, the change threshold for the gain parameter can be set to 0.1 dB; the change threshold for the notch frequency can be set to 1 Hz.
[0170] When the convergence condition is met, the optimal combination of equalizer gain parameters and feedback suppressor notch parameters is obtained. , This combination satisfies the Nash equilibrium condition:
[0171] For the first player, When fixed, It is its utility function The maximum value point, that is: ( , )≥ (G, This holds true for all G;
[0172] For the second player, When fixed, It is its utility function The maximum value point, that is: ( , )≥ ( (N) holds for all N;
[0173] This Nash equilibrium state means that, while ensuring system stability, speech intelligibility has reached its optimal level. Further adjusting the equalizer to improve intelligibility would inevitably lead to a decrease in stability margin; similarly, further adjusting the feedback suppressor to improve stability would inevitably lead to a decrease in intelligibility. The two have reached an optimal balance.
[0174] In one possible implementation, step S4 specifically includes:
[0175] S41. Convert the optimal combination of collaborations into audio processor register parameters.
[0176] Specifically, based on the selected audio processor model, the equalizer gain parameter G1 is converted into the write value of each frequency band gain register.
[0177] For example, for the ADSP-21489 processor, equalizer parameters are typically stored in a specific parameter RAM area, with each gain value corresponding to a 16-bit fixed-point number. The conversion formula is:
[0178] =round( )+offset;
[0179] Where: round() is the rounding function; step is the gain step, which is determined by the processor hardware specifications, for example, 0.5dB; offset is the offset, which makes the 0dB gain correspond to a specific register value, for example, 0x8000.
[0180] The specific step and offset values can be found in the processor datasheet. This system has a built-in parameter mapping table for mainstream processors, and users only need to select the processor model for automatic matching.
[0181] Notch parameters for feedback suppressor This is then converted into a configuration sequence for the notch control registers. Each notch typically requires setting a center frequency register, depth register, bandwidth register, etc. Frequency values need to be converted into corresponding index values or coefficients, depending on the processor implementation.
[0182] S42, Output configurable parameters.
[0183] Specifically, the converted register parameters are output in a standard format. Output methods include:
[0184] Generate parameter configuration files (such as XML or JSON format) for field engineers to import into the audio processor;
[0185] The audio can be written directly to the audio processor via a serial interface (RS-232 / RS-485) or a network interface (TCP / IP);
[0186] The simulation software interface displays a list of parameters for users to view and configure manually.
[0187] Through S41-S42, a seamless connection from simulation optimization to actual deployment is achieved, avoiding errors that may be introduced by manual calculations in traditional methods and significantly improving engineering efficiency.
[0188] Example 2, please refer to Figure 2 As shown in this embodiment, a conference acoustic simulation system based on the virtual sound source method includes:
[0189] Geometric modeling module: used to acquire and store the geometric structural parameters and material acoustic parameters of the conference room;
[0190] Virtual sound source simulation module: used to generate acoustic impulse responses of various target measurement points in the conference room using the virtual sound source method;
[0191] Correlation modeling module: used to establish a correlation model between the parameters of at least two adjustable parameter modules and acoustic performance indicators based on the acoustic impulse response;
[0192] Collaborative optimization module: used to solve for the optimal combination of parameters that optimizes the overall performance of the conference system by using the parameters of the at least two adjustable parameter modules as decision variables and the acoustic performance index as the objective function;
[0193] Parameter mapping module: used to convert the parameter combination into configurable parameters for the audio processor and output them.
[0194] In one specific implementation, the geometry modeling module is a 3D modeling software module based on OpenGL or Direct3D, supporting the import of AutoCAD drawings, Revit models, or point cloud data. Users can interactively define the geometry and material properties of each surface in the conference room through a graphical interface. The module internally uses boundary representation to store the 3D model, with each face containing attributes such as vertex coordinates, normal vectors, and material indexes. The material library includes pre-set sound absorption and scattering coefficient data for common building materials, and users can also define their own new materials.
[0195] The virtual sound source simulation module is the computational core of the system and is typically deployed on a high-performance server or multi-core workstation. This module implements the aforementioned dynamic adaptive virtual sound source segmentation method, hybrid computation model, and occlusion pre-judgment algorithm. After calculation, the acoustic impulse response of each target measurement point is stored in WAV format or a dedicated binary format for easy access by subsequent modules.
[0196] The correlation modeling module is implemented as a series of mathematical calculation function libraries, including FFT, numerical optimization, and signal processing libraries. This module receives the impulse response output from the virtual sound source simulation module, as well as the user-specified speaker and microphone configurations, and builds the correlation model according to steps S21-S25. The calculation results (H(f) matrix, intermediate STI calculation results, etc.) are stored in shared memory for access by the collaborative optimization module.
[0197] The collaborative optimization module is the core innovative module of this system. In a preferred embodiment, this module further includes a game optimization submodule, used to construct and solve the game optimization model between the multi-channel equalizer and the feedback suppressor.
[0198] Specifically, the game optimization submodule defines the equalizer gain parameter as the strategy set of the first player and the speech clarity index as the utility function of the first player; it defines the feedback suppressor notch parameter as the strategy set of the second player and the system stability margin index as the utility function of the second player; it uses a turn-based game mechanism to iteratively optimize until Nash equilibrium is reached, and outputs the cooperative optimal parameter combination.
[0199] The parameter mapping module contains a processor parameter database that records the register mapping rules and communication protocols of mainstream audio processors (such as Biamp, QSC, Symetrix, etc.). This module receives the optimal parameter combination, queries the database to obtain the parameter conversion rules for the corresponding model, and generates executable control instructions. Instruction formats include Modbus protocol frames, CobraNet control packets, and Dante controller API calls.
[0200] In another preferred embodiment, the system further includes a communication interface module for establishing a real-time communication connection with the conference audio processor. The communication interface module can employ communication protocols such as TCP / IP, RS-485, Bluetooth, or Wi-Fi, directly writing the configurable parameters generated by the parameter mapping module into the audio processor's registers to achieve real-time synchronization between simulation parameters and the actual system.
[0201] In another preferred embodiment, the collaborative optimization module further includes a dynamic update submodule to respond to dynamic changes in the meeting environment. When a change in the number of participants, an adjustment in the table and chair layout, or a change in the speaker's position is detected, the dynamic update submodule triggers incremental recalculation, using the equilibrium solution from the previous round of the game as the initial strategy for the current iteration. Transfer learning techniques are employed to accelerate convergence, generating updated configurable parameters for the audio processor. This mechanism enables the system to adapt to the constantly changing acoustic environment during the meeting, maintaining optimal performance at all times.
[0202] In the strategy selection step S32, various optimization algorithms can be employed. In addition to the basic turn-based game mechanism, the preferred embodiment of the present invention also provides the following algorithm implementations:
[0203] Virtual game algorithm: In each iteration, the players do not directly base their optimal response on their opponent's current strategy, but rather on the statistical average of the opponent's historical strategies. Specifically, the first player selects the optimal response based on the average of the second player's historical strategies. The strategy that maximizes the value; the second player chooses based on the average of the first player's historical strategies. The strategy is to maximize the convergence. This algorithm has better convergence stability and avoids oscillations between multiple local optima.
[0204] Policy hill climbing: In each iteration, a new policy that maximizes the utility function is searched within the neighborhood of the current policy. The neighborhood size gradually decreases with the number of iterations, achieving a transition from coarse search to fine search.
[0205] Particle Swarm Optimization (PSO) algorithm: It maintains multiple candidate strategies (particles) at the same time. Each particle adjusts its strategy based on its own historical best and the global historical best, and finds the global optimum through group cooperation.
[0206] In this invention, the superscript (t) represents the strategy value after the t-th iteration. This represents the feedback suppressor strategy determined after the (t-1)th iteration. This represents the new equalizer strategy chosen by the first player in the t-th iteration. This indicates the feedback suppressor strategy newly chosen by the second player in round t. This superscript notation clearly illustrates the order of strategy updates and dependencies during the turn-based game.
[0207] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A conference acoustic simulation method based on the virtual sound source method, applied to a conference system equipped with an audio processor, wherein the audio processor includes at least two adjustable parameter modules, characterized in that, include: S1. The acoustic impulse response of the target acoustic measurement point in the conference room is generated using the virtual sound source method, and the audio processor is a programmable digital signal processor; S2. Based on the acoustic impulse response, establish a correlation model between the parameters of at least two adjustable parameter modules in the audio processor and the acoustic performance indicators; S3. Using the parameters of the at least two adjustable parameter modules as decision variables and the acoustic performance index as the objective function, a collaborative optimization algorithm is used to find the parameter combination that optimizes the overall performance of the conference system. S4. Convert the parameter combination into configurable parameters of the audio processor and output them.
2. The conference acoustic simulation method based on the virtual sound source method according to claim 1, characterized in that: The at least two adjustable parameter modules include a multi-channel equalizer and a feedback suppressor. The multi-channel equalizer is used to adjust the gain of different frequency bands, and the feedback suppressor is used to set notch parameters to suppress acoustic feedback.
3. The conference acoustic simulation method based on the virtual sound source method according to claim 2, characterized in that: The acoustic performance indicators include a first performance indicator and a second performance indicator. The first performance indicator is a speech intelligibility indicator, and the second performance indicator is a system stability margin indicator.
4. The conference acoustic simulation method based on the virtual sound source method according to claim 3, characterized in that: The establishment of the association model in S2 further includes: Based on the acoustic impulse response, establish the feedback path transfer function matrix H(f) between multiple loudspeakers and multiple microphones; Based on the feedback path transfer function matrix H(f), the gain parameter G of the multi-channel equalizer, and the notch parameter N of the feedback suppressor, the system open-loop transfer function L(f) = G(f)·N(f)·H(f) is constructed. Where G(f) is the equalizer frequency response function generated by the gain parameter G, and N(f) is the feedback suppressor frequency response function generated by the notch parameter N; The system stability margin index is calculated based on the open-loop transfer function L(f) of the system. Based on the acoustic impulse response, the gain parameter G of the multi-channel equalizer, and the notch parameter N of the feedback suppressor, the equivalent impulse response after taking into account the effects of the equalizer and the feedback suppressor is calculated, and then the speech intelligibility index is calculated.
5. The conference acoustic simulation method based on the virtual sound source method according to claim 4, characterized in that: The collaborative optimization algorithm in S3 adopts a game-theoretic optimization model, including: S31. Define the gain parameter G of the multi-channel equalizer as the strategy set of the first player, and define the speech clarity index as the utility function of the first player. (G, N), where the utility function is used to quantify the payoff of the first player under the parameter combination (G, N); the notch parameter N of the feedback suppressor is defined as the strategy set of the second player, and the system stability margin index is defined as the utility function of the second player. (G, N), the utility function is used to quantify the payoff of the second player under the parameter combination (G, N); S32. An iterative optimization is performed using a turn-based game mechanism. In the t-th round of iteration, the first player adopts the current feedback suppressor strategy. Choose to use under fixed conditions Maximization strategy The second player's current equalizer strategy Choose to use under fixed conditions Maximization strategy ; S33. Repeat step S32 until the strategy converges, obtaining the optimal combination of equalizer gain parameters and feedback suppressor notch parameters. , ); The optimal collaborative combination satisfies the Nash equilibrium condition: for the first player, in When fixed, It is its utility function The maximum point; for the second player, in When fixed, It is its utility function The maximum value point.
6. The conference acoustic simulation method based on the virtual sound source method according to claim 5, characterized in that: The strategy selection in S32 employs a virtual game algorithm, including: The first player chooses based on the average of the second player's historical strategies. Maximization strategy; The second player chooses based on the average of the first player's historical strategies. Maximization strategy.
7. The conference acoustic simulation method based on the virtual sound source method according to claim 5, characterized in that: The strategy selection in S32 adopts either the strategy hill climbing method or the particle swarm optimization algorithm.
8. The conference acoustic simulation method based on the virtual sound source method according to claim 1, characterized in that: The generation of acoustic impulse response using the virtual sound source method in S1 further includes: For curved walls, a dynamic adaptive segmentation strategy is adopted, which dynamically determines the segmentation granularity based on the relative position of the sound source and the receiving point, and only generates virtual sound sources that contribute to the current receiving point. For early reflections of order 1 to N, the virtual sound source method is used for accurate calculation, and for later reflections above order N, the ray tracking method is used for statistical calculation, where N is a preset reflection order threshold. Based on the occlusion relationship of the tiered seating area, a three-dimensional occlusion model is established to pre-judge the visibility of virtual sound sources. Invalid virtual sound sources that are occluded are then eliminated through a reverse tracking algorithm.
9. A conference acoustic simulation method based on the virtual sound source method according to claim 5, characterized in that: It also includes a dynamic update step: S5. Acquire dynamic change information of the meeting venue in real time or near real time, including changes in the number of participants, adjustments to the table and chair layout, and changes in the speaker's position; S6. Based on the dynamic change information, trigger incremental recalculation of S1 to S4, take the equilibrium solution of the previous game as the initial strategy of this iteration, use transfer learning to accelerate convergence, and generate updated configurable parameters for the audio processor.
10. A conference acoustic simulation system based on the virtual sound source method, used to implement the conference acoustic simulation method based on the virtual sound source method as described in any one of claims 1-9, characterized in that, include: Geometric modeling module: used to acquire and store the geometric structural parameters and material acoustic parameters of the conference room; Virtual sound source simulation module: used to generate acoustic impulse responses of various target measurement points in the conference room using the virtual sound source method; Correlation modeling module: used to establish a correlation model between the parameters of at least two adjustable parameter modules and acoustic performance indicators based on the acoustic impulse response; Collaborative optimization module: used to solve for the optimal combination of parameters that optimizes the overall performance of the conference system by using the parameters of the at least two adjustable parameter modules as decision variables and the acoustic performance index as the objective function; Parameter mapping module: used to convert the parameter combination into configurable parameters for the audio processor and output them.