A microphone array layout optimization method and a sound reinforcement system

By optimizing the microphone array layout using the particle swarm optimization algorithm, the contradiction between array design and signal processing in the sound reinforcement system was resolved. This enabled high-performance sound source separation and howling suppression in compact devices, improving the system's real-time performance and signal quality.

CN122491010APending Publication Date: 2026-07-31JUSHENG (YANGJIANG) TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JUSHENG (YANGJIANG) TECHNOLOGY CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing sound reinforcement systems suffer from several problems in microphone array design and voice processing, including difficulty in balancing low-frequency directivity maintenance and high-frequency spatial aliasing suppression, performance conflicts between noise suppression and feedback control in traditional beamforming methods, and contradictions in signal processing delay accumulation and hardware resource consumption. These issues limit the sound pickup quality and deployment flexibility.

Method used

The microphone array layout is optimized using a particle swarm optimization algorithm. By combining the minimum microphone spacing and the maximum array aperture, a non-uniform array design is achieved. Voice enhancement and feedback suppression functions are integrated to reduce signal processing latency. A sparse array and integrated processing architecture are adopted to optimize full-band performance.

Benefits of technology

This invention enhances the array's directional sensitivity and noise suppression capabilities in a compact device, reduces hardware resource consumption, enables effective separation of multiple target sound sources and suppression of howling signals, and improves system real-time performance and output signal fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491010A_ABST
    Figure CN122491010A_ABST
Patent Text Reader

Abstract

This invention provides a microphone array layout optimization method and a sound reinforcement system. The method includes: obtaining initialization parameters for a particle swarm optimization algorithm, including physical constraints, a total fitness function, and the position of a reference microphone; and calculating and determining the globally optimal microphone array layout based on the initialization parameters using the particle swarm optimization algorithm. The physical constraints include the minimum microphone spacing and the maximum array aperture. The total fitness function includes the fitness function contribution corresponding to the average directivity factor at representative frequency points for at least two target directions, the fitness function contribution corresponding to the average crosstalk power in the target directions, the fitness function contribution corresponding to the howling suppression potential, and the fitness function contribution corresponding to the average white noise gain at representative frequency points. This invention significantly improves the array's directivity sensitivity and noise suppression capability within a limited size, greatly reduces signal processing latency, and improves processing efficiency and output signal fidelity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound reinforcement system technology, specifically to a microphone array layout optimization method and a sound reinforcement system. Background Technology

[0002] Existing public address systems still have shortcomings in microphone array design and speech processing. In terms of array architecture, mainstream solutions mostly use uniform linear arrays, which, limited by the half-wavelength rule (d=λ / 2), struggle to balance low-frequency directivity maintenance and high-frequency spatial aliasing suppression within a limited physical size. Simultaneously, traditional beamforming methods (such as MVDR) inherently conflict between noise suppression and feedback control, often requiring systems to sacrifice directivity for algorithm stability. At the signal processing level, discrete modular architectures force speech enhancement, feedback suppression, and other components to be configured and processed independently and in series, leading to accumulated link delays, increased phase distortion, and a lack of scene-adaptive sound field rendering capabilities, failing to dynamically adjust acoustic zones based on speaker identity. In terms of hardware implementation, multi-channel independent processing schemes rely on large-scale DSP resources for parallel computation, contradicting the trend towards miniaturized and highly integrated devices. These issues collectively limit the pickup quality, real-time processing performance, and deployment flexibility of existing public address systems in complex acoustic environments.

[0003] Existing technology discloses a linear differential microphone array based on geometry optimization, specifically disclosing that: the array's microphones are divided into microphone subarrays, and then the optimal subarray geometry in different frequency bands is identified based on specified performance targets. Then, based on the evaluation of the specified performance targets on the frequency bands, the union of the optimal subarray geometries constitutes the complete array geometry, wherein the optimal subarray geometry for each of the 80 frequency bands can be optimized using a particle swarm optimization (PSO) algorithm across multiple (80) frequency bands. However, existing solutions lack specificity for the core requirements of conference sound reinforcement systems, making it difficult to achieve clear pickup, effective separation, and active howling suppression of dual-target sound sources within a compact size, and the optimization results have poor engineering practicality due to neglecting physical constraints. Summary of the Invention

[0004] The primary objective of this invention is to provide a microphone array layout optimization method that overcomes the constraints of physical size and performance, thereby achieving integrated acoustic processing.

[0005] The second objective of this invention is to provide a sound reinforcement system that can significantly reduce signal processing delay and improve system real-time performance and output signal fidelity.

[0006] To achieve the aforementioned first objective, this invention provides a microphone array layout optimization method, comprising the following steps: obtaining initialization parameters for a particle swarm optimization algorithm, the initialization parameters including physical constraints, a total fitness function, and the position of a reference microphone; based on the initialization parameters, calculating and determining the globally optimal microphone array layout using the particle swarm optimization algorithm; the physical constraints include the minimum microphone spacing and the maximum array aperture; the total fitness function includes the fitness function contribution corresponding to the average directivity factor of at least two target directions at representative frequency points, the fitness function contribution corresponding to the average crosstalk power of the target directions, the fitness function contribution corresponding to the howling suppression potential, and the fitness function contribution corresponding to the average white noise gain at representative frequency points.

[0007] As can be seen from the above scheme, this invention generates a non-uniform array mop based on the particle swarm optimization algorithm, and combines the minimum microphone spacing, maximum array aperture settings, and overall fitness function settings to enable compact devices to simultaneously achieve the directionality and noise suppression capabilities of commercially available large arrays over a wide bandwidth. This invention breaks through the physical limitations of traditional uniform arrays, achieving a compact design through intelligent optimization algorithms. It significantly improves the array's directionality sensitivity and noise suppression capabilities within a limited size, while also considering wideband acoustic performance. Compared to traditional solutions, it can achieve the core indicators of commercial-grade sound reinforcement systems with a smaller hardware scale. This invention can effectively separate multiple target sound sources, especially dual-target sound sources, and suppress feedback signals. This invention innovatively integrates voice enhancement and feedback suppression functions into a single-stage processing link, significantly reducing signal processing latency, avoiding phase distortion problems caused by multi-module cascading, and improving system real-time performance and output signal fidelity. This invention adopts a sparse array design and integrated processing architecture, maintaining high performance while reducing hardware resource consumption, providing a cost-effective sound reinforcement solution for miniaturized intelligent devices.

[0008] A further approach is to express the total fitness function as follows: .

[0009] in, .in, This represents the fitness function contribution corresponding to the average directivity factor of at least two target directions at representative frequency points. This represents the contribution of the fitness function to the average crosstalk power in the target direction. This represents the contribution of the fitness function to the whistle suppression potential. The fitness function contribution corresponding to the average white noise gain at representative frequency points is represented by a weighting coefficient set for each part of the fitness function contribution.

[0010] Therefore, this invention facilitates the adjustment of the contribution of each part of the fitness function by setting weight coefficients. This invention also divides the entire frequency band into multiple sub-bands, requiring the particle swarm optimization algorithm to simultaneously optimize the directionality factor and noise suppression capability of each band, avoiding performance optimization in a single band leading to degradation in other bands.

[0011] A further approach is to determine the globally optimal microphone array layout using a particle swarm optimization algorithm, including: maximizing the directivity factor based on the fitness function contribution corresponding to the average directivity factor at representative frequency points for at least two target directions, and maximizing the white noise gain based on the fitness function contribution corresponding to the average white noise gain at representative frequency points under the constraint of a set target value.

[0012] Therefore, it is possible to optimize white noise gain control as much as possible while ensuring the optimization of the directional factor.

[0013] A further approach is to express the fitness function contribution corresponding to the howling suppression potential as: , ,in, Indicates the direction of loudspeaker radiation. This indicates the beam power corresponding to the direction of loudspeaker radiation. For the corresponding weighting coefficients, For the first Particles in the PSO algorithm.

[0014] Therefore, it is evident that optimizations can be specifically tailored for scenarios using speakers to achieve noise suppression directly within the microphone array.

[0015] A further approach is to express the fitness function contribution corresponding to the average white noise gain at representative frequency points as follows: , , , , To set a target value, For frequency weights, The number of representative frequency points, For the corresponding weighting coefficients, For the first Particles in the PSO algorithm.

[0016] A further proposed approach involves two target directions, corresponding to the main speaking area and the audience interaction area, respectively; the fitness function contribution of the average directivity factor of the two target directions at representative frequency points is expressed as follows: , ,in, For the corresponding weighting coefficients, As a directional factor, The number of representative frequency points, For frequency weights, For the first One target direction, For the first A representative frequency point, For the first Particles in the PSO algorithm.

[0017] A further approach is to express the fitness function contribution corresponding to the average crosstalk power in the target direction as follows: .

[0018] .

[0019] in, ( The signal represents the response power of the beam from the first target direction at the second target direction. ( The symbol represents the response power of the beam in the second target direction at the first target direction. These are the corresponding weighting coefficients.

[0020] This demonstrates that it can effectively separate dual-target sound sources and suppress howling signals, especially in classroom and conference room scenarios, while ensuring speech clarity in the target area and significantly suppressing acoustic crosstalk in non-target areas.

[0021] A further proposed approach is to include the following initialization parameters: the number of particles can be selected from 50 to 100, the maximum number of iterations can be selected from 100 to 500, the inertia weight can be decreased from 0.9 to 0.4, and the learning factor can be selected from 1.5 to 2.0.

[0022] A further approach is to calculate the overall fitness function based on the beamforming filter corresponding to the particle in the particle swarm optimization algorithm.

[0023] Therefore, it can be seen that by utilizing the space-frequency joint control characteristics of differential filters, voice enhancement and howling suppression functions can be integrated into a single-stage processing link.

[0024] To achieve the second objective mentioned above, the present invention provides a sound reinforcement system, including a microphone array, a voice processing device, and a loudspeaker. The microphone array is connected to the voice processing device, and the voice processing device is connected to the loudspeaker. The array layout of the microphone array is determined according to a microphone array layout optimization method. The voice processing device includes a beamforming module, a voice enhancement module, and a feedback suppression module.

[0025] As can be seen from the above scheme, the microphone array used in this invention achieves clear pickup, effective separation, and active feedback suppression of the target sound source, providing a better front-end signal for the speech processing device. This improves the processing efficiency of differential beamforming, speech enhancement, and feedback suppression in the subsequent speech processing device, thereby enhancing the processing efficiency and sound reinforcement effect of the entire sound reinforcement system. This invention achieves zoned control of the acoustic space through dynamic sound field rendering technology. The system can automatically adjust the direction of sound energy projection according to the sound source identity, significantly suppressing acoustic crosstalk in non-target areas while ensuring speech clarity in the target area. Attached Figure Description

[0026] Figure 1 This is a flowchart of an embodiment of the microphone array layout optimization method of the present invention.

[0027] Figure 2 yes Figure 1 The detailed flowchart of step S13.

[0028] Figure 3 This is a system framework diagram of an embodiment of the sound reinforcement system of the present invention.

[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments. Detailed Implementation

[0030] The microphone array layout optimization method of this invention employs the Particle Swarm Optimization (PSO) algorithm to obtain the globally optimal microphone array layout. This enables the microphone array based on the globally optimal layout to possess speech enhancement and feedback suppression capabilities, achieving effective separation of multiple target sound sources and suppression of feedback signals. Furthermore, in conjunction with the speech enhancement and feedback suppression of the speech processing device in the sound reinforcement system of this invention, overall optimization is achieved. In addition, the sound reinforcement system utilizes directional sound field rendering technology to precisely control the acoustic space, thereby improving the robustness, spatial resolution, and overall listening quality of the sound reinforcement system.

[0031] Example of microphone array layout optimization method: This embodiment describes the optimization of a non-uniform linear microphone array containing M microphones. This non-uniform linear microphone array includes one reference microphone and M-1 non-reference microphones, with the reference microphone positioned at the origin. Then the rest of M The position of a non-reference microphone This refers to the target parameters optimized using the particle swarm optimization algorithm in this embodiment. The globally optimal microphone array layout, i.e., the optimal M, is obtained through iterative processing using the particle swarm optimization algorithm. The position of a non-reference microphone .

[0032] See Figure 1 This embodiment is implemented by a computer program and specifically includes the following steps: S11: Obtain the initialization parameters for the particle swarm optimization algorithm.

[0033] S12: Initialize the particle swarm.

[0034] S13: Based on the initialization parameters, iteratively perform the particle swarm optimization algorithm.

[0035] S14: Determine if the required number of iterations has been met. If the required number of iterations has been met, proceed to step S15; otherwise, return to step S13 to continue the iteration.

[0036] S15: Outputs the globally optimal microphone array layout.

[0037] In step S11 above, the initial parameters of the particle swarm optimization algorithm include physical constraints, the overall fitness function, and the position of the reference microphone.

[0038] Physical constraints are used to indicate the physical constraints on the search space OS of the particle swarm optimization algorithm, and are expressed as:

[0039] .in, Minimum microphone spacing, This represents the maximum aperture of the array.

[0040] The overall fitness function includes the fitness function contribution corresponding to the average directivity factor of at least two target directions at representative frequency points, the fitness function contribution corresponding to the average crosstalk power of the target directions, the fitness function contribution corresponding to the howling suppression potential, and the fitness function contribution corresponding to the average white noise gain at representative frequency points.

[0041] In this embodiment, the total fitness function is expressed as:

[0042] .

[0043] in, This represents the fitness function contribution corresponding to the average directivity factor of at least two target directions at representative frequency points. This represents the contribution of the fitness function to the average crosstalk power in the target direction. This represents the contribution of the fitness function to the whistle suppression potential. This represents the fitness function contribution corresponding to the average white noise gain at representative frequency points. Each fitness function contribution is assigned a corresponding weighting coefficient.

[0044] The fitness function contribution corresponding to the average directivity factor at representative frequency points for at least two target directions is used to measure the target source pickup and directivity of the microphone array, specifically reflected by the microphone array's directivity factor (DF). The directivity factor quantifies the microphone array's ability to concentrate sound energy in the target direction and is defined as the ratio of the beam response energy in the target direction to the average energy across the entire space. In the overall fitness function, the goal is to maximize the fitness function contribution corresponding to the average directivity factor at representative frequency points for at least two target directions.

[0045] The fitness function contribution corresponding to the average crosstalk power in the target direction is used to measure the source separation of the microphone array, aiming to minimize crosstalk between target directions.

[0046] The fitness function contribution corresponding to the feedback suppression potential is used to measure the feedback suppression potential of a microphone array: it is evaluated by creating an acoustic null in the main radiation direction of the loudspeaker.

[0047] The contribution of the fitness function corresponding to the average white noise gain at representative frequency points is used to measure the robustness of the microphone array to the sensor's own noise, preventing excessive noise amplification. This is specifically reflected by the white noise gain (WNG). White noise gain is a parameter characterizing the array's sensitivity to the microphone's inherent noise floor, reflecting the risk level of array noise amplification. In the overall fitness function, the contribution of the fitness function corresponding to the average white noise gain at representative frequency points is expected to be as large as possible. In practical applications, this can be set to be no less than a predetermined threshold (e.g., 0dB). Since maximizing white noise gain may conflict with maximizing the directionality factor, a satisfaction term or soft constraint term can be set. In this embodiment, a penalty is imposed when the white noise gain falls below the set target value.

[0048] This embodiment is illustrated using two target directions. The first target direction corresponds to the main speaking area, and the second target direction corresponds to the audience interaction area. For example, in a teaching scenario, the podium is the main speaking area, and the students are the audience interaction area; in a conference scenario, the speaker is the speaking area, and the attendees are the audience interaction area.

[0049] The contribution of the fitness function corresponding to the average directivity factor of the two target directions at representative frequency points is expressed as follows: , ,in, These are the corresponding weighting coefficients; It is a directional factor; The number of representative frequency points, For the corresponding representative frequency point The frequency weight can be set according to the importance of that frequency to the overall sound quality or a specific target (such as speech intelligibility); For the first One target direction; For the first Representative frequency points; For the first Particles in the PSO algorithm.

[0050] The fitness function contribution corresponding to the average crosstalk power in the target direction (negative sign because we want to minimize crosstalk) is expressed as: .

[0051] .

[0052] in, ( The signal represents the response power of the beam from the first target direction at the second target direction. ( The symbol represents the response power of the beam in the second target direction at the first target direction. For the corresponding representative frequency point Frequency weights, For the first Particles in the PSO algorithm.

[0053] The fitness function contribution to the feedback suppression potential (the objective is to minimize the pickup direction of the speaker, i.e., maximize its negative value) is: , ,in, Indicates the direction of loudspeaker radiation. This indicates the beam power corresponding to the direction of loudspeaker radiation. For the corresponding weighting coefficients, For the first Particles in the PSO algorithm.

[0054] The contribution of the fitness function to the average white noise gain at representative frequency points is expressed as follows: , , , , To set a target value, For the corresponding representative frequency point Frequency weights, The number of representative frequency points, For the corresponding weighting coefficients, For the first Particles in the PSO algorithm.

[0055] For a particle A linear differential microphone array (LDMA) with a defined specific geometry, at an angular frequency... and direction The steering vector can be represented as: ,in, The speed of sound.

[0056] For an Nth-order linear differential microphone array, its desired beam pattern BN(θ) typically has specific zeros, and the beamforming filter... The design goal is to make its array responsive Approximating the desired beam pattern. Common methods include maximum WNG differential beamformers. ,in, It is a matrix based on the desired direction and the zero-point direction steering vector. This embodiment utilizes the characteristics of a second-order differential array (N=2, requiring at least 3 microphones to form a subarray), by optimizing the microphone positions. In conjunction with the beamforming method described above, the overall fitness function is calculated. , , , The beamformer design method in this embodiment is not limited to the maximum WNG differential beamformer, but can also be other existing methods, such as minimum variance distortionless response (MVDR), linearly constrained minimum variance (LCMV), and generalized sidelobe canceller (GSC).

[0057] The initialization parameters also include the selection of the number of particles, the maximum number of iterations, the inertia weight, and the learning factor. Specifically, the number of particles can be selected from 50 to 100, the maximum number of iterations can be selected from 100 to 500, the inertia weight decreases from 0.9 to 0.4, and the learning factor can be selected from 1.5 to 2.0.

[0058] In step S12 above, the initialization process of a single particle in the particle swarm is explained. First, the initial position of the particle is randomly generated within the physical constraints to obtain the initial position of the current particle. Next, the velocity of the currently generated particle is randomly determined. Then, the individual optimal position of the current particle is set as its initial position. The current fitness of the current particle is calculated based on the global fitness function, and its individual optimal fitness is set as its current fitness. Finally, the current fitness of the current particle is compared with the global optimal fitness. If the current fitness is greater than the global optimal fitness, the initial fitness of the current particle is assigned to the global optimal fitness. The global optimal fitness can be empty. If the global optimal fitness is empty, the current fitness of the current particle is used as the global optimal fitness. The global optimal fitness can also be set to a large value according to the actual situation. In this way, the initialization of all particles can be completed.

[0059] In step S13 above, see Figure 2 Specifically, it includes the following steps: S131: Calculate the inertia weight for the current iteration.

[0060] S132: j=1.

[0061] S133: Update the velocity and position of the j-th particle according to the PSO formula and physical constraints, and determine the current fitness of the j-th particle.

[0062] S134: Determine whether the current fitness of the j-th particle is better than the corresponding individual's optimal fitness. If yes, continue to step 135; otherwise, jump to step S138.

[0063] S135: Update the individual optimal position and individual optimal fitness of the j-th particle.

[0064] S136: Determine whether the current fitness of the j-th particle is better than the global optimal fitness. If yes, continue to step S137; otherwise, jump to step S138.

[0065] S137: Update the global optimal fitness using the current fitness of the j-th particle.

[0066] S138: j = j + 1.

[0067] S139: Determine if j is greater than the number of particles. If yes, continue to step S14; otherwise, return to step S133.

[0068] Therefore, after iteration, a globally optimal microphone array layout is obtained. Based on this globally optimal layout, an optimal non-uniform linear microphone array can be designed, enabling the array to achieve optimal overall performance (including target speech enhancement, crosstalk suppression, feedback suppression potential, and broadband robustness) in a specific scenario (such as a conference room or classroom). This embodiment utilizes the space-frequency joint control characteristics of differential filters to integrate speech enhancement and feedback suppression functions into a single-stage processing link, reducing latency to less than 8ms and phase consistency error to <0.05π.

[0069] As can be seen, this embodiment performs end-to-end geometric layout optimization on the entire non-uniform linear array containing M microphones, rather than focusing on the combination or selection of subarrays. The Particle Swarm Optimization (PSO) algorithm directly adjusts the precise physical positions [ρ2, ρ3, ..., ρM] of the M-1 movable microphones (assuming the first microphone ρ1 is fixed at the origin). To ensure that the optimized array has excellent and consistent performance in practical applications (especially when processing broadband signals such as voice), this scheme adopts a multi-band collaborative optimization strategy. First, representative frequency points are selected. In this embodiment, a set of discrete and representative frequency points are selected within the target operating bandwidth. The selection of these frequency points can be based on logarithmic intervals to better match the characteristics of human hearing, or on more intensive sampling in key frequency bands (such as speech formant regions and areas prone to feedback). Then, for the full-band evaluation of the overall fitness function, this embodiment evaluates each particle in the PSO algorithm (representing a complete layout scheme of M microphones) ), its fitness function The calculation is not based on the performance at a single frequency point, but rather on the overall performance of the configuration across all selected representative frequencies. Specifically, the fitness function incorporates various frequency-related performance metrics, such as beam power in the target direction. directional factors White noise gain Wait, it will first be at each representative frequency point Each frequency point index is calculated independently. Subsequently, these single-frequency indexes are aggregated into a comprehensive index for the entire frequency band through methods such as weighted averaging, summation, or worst-case selection.

[0070] The pseudocode based on the particle swarm optimization algorithm is as follows: PROCEDURE: PSO_Array_Optimization INPUT: Np: Number of particles MaxIter: Maximum number of iterations w_max, w_min, c1, c2: PSO parameters OS: Search Space and Physical Constraints FitnessFunction: Fitness Function OUTPUT: g_best_position: Globally optimal microphone array layout INITIALIZATION: FOR i = 1 TO Np DO particle[i].position = RandomPosition(OS) particle[i].velocity = RandomInitialVelocity() particle[i].p_best_position = particle[i].position particle[i].p_best_fitness = FitnessFunction(particle[i].position) UpdateGlobalBest(particle[i]) / / Compare and update the global best g_best END FOR ITERATIVE OPTIMIZATION: FOR iter = 1 TO MaxIter DO w = CalculateInertiaWeight(iter, MaxIter, w_max, w_min) / / Dynamically adjust inertia weight FOR i = 1 TO Np DO / / Update speed and location particle[i].velocity = UpdateVelocity(particle[i], g_best_position,w, c1, c2) particle[i].position = particle[i].position + particle[i].velocity particle[i].position = ApplyConstraints(particle[i].position, OS) / / Handle boundary / constraints / / Assessment and Update current_fitness = FitnessFunction(particle[i].position) IF current_fitness is better than particle[i].p_best_fitness THEN particle[i].p_best_position = particle[i].position particle[i].p_best_fitness = current_fitness UpdateGlobalBest(particle[i]) / / Compare and update the global best g_best END IF END FOR END FOR RETURN g_best_position Sub-function explanation in pseudocode (conceptual): •RandomPosition(OS): Generates a random layout within the constraint OS.

[0071] •RandomInitialVelocity(): Generates a random initial velocity.

[0072] •FitnessFunction(position): Calculates the fitness value of the layout.

[0073] •UpdateGlobalBest(particle): Updates the global best based on the current individual best of each particle.

[0074] •CalculateInertiaWeight(...): Calculates the inertia weight for the current iteration.

[0075] •UpdateVelocity(...): Calculates the new velocity based on the PSO formula.

[0076] Example of a sound reinforcement system: See Figure 3The sound reinforcement system 1 in this embodiment includes a microphone array 10, a voice processing device 20, and a speaker 30. The microphone array 10 is connected to the voice processing device 20, and the voice processing device 20 is connected to the speaker 30. The array layout of the microphone array 10 in this embodiment is determined according to the microphone array layout optimization method of the above embodiment.

[0077] The voice processing device 20 includes a voice activity detection module, a sound source localization module, a beamforming module, a voice enhancement module, and a howling suppression module.

[0078] The speech activity detection module, which is a detection module that distinguishes speech segments from non-speech segments in real time based on signal feature analysis, is used to implement speech activity detection (VAD). The core indicators include short-time energy and spectral characteristics.

[0079] In this embodiment, the microphone array includes a main microphone, which serves as the signal acquisition source for voice activity detection. Its omnidirectional pickup characteristics and excellent acoustic performance ensure effective capture of voice features. Due to the single-microphone detection architecture, the high-sensitivity omnidirectional pickup characteristics of the main microphone guarantee reliable capture of voice components while avoiding the computational resource consumption associated with multi-channel signal analysis, achieving an optimized balance between detection accuracy and system power consumption. In the real-time processing flow, the voice activity detection module executes the following judgment logic for each frame of audio signal: when a valid voice feature is detected, subsequent processing links are triggered, including sequentially triggering the sound source localization module, beamforming module, voice enhancement module, and howling suppression module; if no voice component is detected, a mute output mechanism is immediately activated, and the speaker signal is reset to zero.

[0080] The core detection algorithm for speech activity detection is based on frequency domain feature analysis technology. After converting the time-domain signal into a frequency-domain representation through Fourier transform, it focuses on analyzing the spectral characteristics within the 50-4000Hz main speech frequency band. By identifying fundamental frequency characteristics and harmonic distribution patterns, combined with short-time energy change patterns of the speech signal, a multi-dimensional criterion is constructed to achieve accurate speech / non-speech classification.

[0081] The sound source localization module is used to achieve sound source localization (SSL), determining the direction of the sound source by processing the audio signals collected by the microphone array. Specifically, sound source localization can be achieved using either the MUSIC method or the GCC-PHAT method. The MUSIC method, due to its relatively low computational cost, is suitable for relatively open environments. While the GCC-PHAT method has a higher computational cost, it performs well in environments with significant reverberation and noise.

[0082] After obtaining the location and direction information of the sound source, a beamforming filter can be designed to perform beamforming (BF). A beamforming module is used to enhance signals from a specific direction by adjusting the delay (or phase) and gain (or amplitude) of the signals in each channel of the microphone array to form a directional pickup beam in space. The number of delay points can be obtained by combining the direction information calculated by the sound source localization module with the transfer function of the incident angle.

[0083] The beamforming module generates two spatially filtered audio signals, corresponding to the two target pickup areas in this embodiment, namely the first target direction and the second target direction described in the previous embodiment. Through spatial coherence processing of the array signals, the system can dynamically separate sound source signals from different directions. The output channel expression is: , This indicates the enhanced speech signal corresponding to the direction of the first target. This indicates the enhanced speech signal corresponding to the direction of the second target. This represents the spatial response vector of the microphone array to the direction of the first target. This represents the spatial response vector of the microphone array to the direction of the second target. The spatial response vector includes phase compensation and amplitude weighting parameters. This refers to the original input signals of each microphone in the microphone array.

[0084] In addition, the beamforming module supports real-time adjustment of beam direction, which can adapt to typical meeting scenarios such as speaker position movement and multiple people taking turns speaking. It effectively suppresses more than 60% of lateral reverberation and background noise through spatial filtering, thereby improving the speech signal-to-noise ratio in the target area.

[0085] The speech enhancement module improves the clarity and intelligibility of the target speech. The output of the beamforming module is fed into the speech enhancement module for processing. The speech enhancement module uses a Dynamic Range Compressor (DRC) algorithm to reduce the dynamic range of the speech and then increase the overall volume. This method effectively prevents the digital audio signal from exceeding its maximum amplitude, avoiding clipping. The output of the speech enhancement module is represented as follows: .

[0086] Feedback suppression modules are used to eliminate feedback phenomena caused by acoustic feedback in sound reinforcement systems. Since sound is amplified, played out by speakers, and then captured and amplified again by microphones, this process is repeated multiple times, causing the sound to continuously intensify and eventually produce a high-frequency feedback sound. Feedback suppression modules identify the frequency bands prone to feedback and design a notch filter to reduce the sound in that band, for example, using the LMS adaptive algorithm. The feedback frequency band depends on the impulse response of the listening environment and can be calculated and designed in advance. The output of the feedback suppression module is expressed as: .

[0087] Because the microphone array is a differential array (DMA), a microphone system that utilizes the signal differences between multiple microphone units to enhance sound from a specific direction and suppress noise and interference, designing nulls pointing towards the speaker in a differential microphone array essentially achieves acoustic feedback path attenuation through spatial filtering. When the array's beam pattern satisfies: ,in, Represents the array weight vector. Indicates the direction of the speaker The array manifold vector. A differential microphone array is used, and the null direction is controlled by adjusting the element spacing d and the delay parameter τ: ,in When the array elements are interspersed ,(in (As a critical spacing condition), a wideband null can be formed in the direction of the speaker.

[0088] Therefore, the microphone array is physically designed with nulls pointing towards the speaker to reduce the energy input in the speaker direction and reduce the spatial suppression of feedback signals entering the system. On this basis, the feedback suppression module performs frequency domain suppression by notch filtering on the residual feedback frequency. The two work together to enhance and suppress the audio signal so that it can be played better by the speaker, reducing feedback and enhancing the overall feedback suppression capability of the sound reinforcement system.

[0089] In summary, the microphone array layout optimization method of this invention directly optimizes the specific positions of microphones in a non-uniform linear array, fully releasing geometric degrees of freedom and enabling it to achieve excellent broadband performance even in a compact size. This overcomes the limitations of traditional uniform arrays. The optimization process considers physical constraints such as minimum microphone spacing and maximum array length, allowing the optimization results to be directly applied to actual product design. By optimizing the array structure, it provides a better front-end signal for beamforming, voice enhancement, and feedback suppression in subsequent sound reinforcement systems, thereby improving the processing efficiency and final effect of the entire system.

[0090] Finally, it should be emphasized that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for optimizing the layout of a microphone array, characterized in that, Includes the following steps: Obtain the initialization parameters of the particle swarm optimization algorithm, including physical constraints, total fitness function, and the position of the reference microphone; Based on the initialization parameters, the globally optimal microphone array layout is calculated and determined using a particle swarm optimization algorithm. The physical constraints include minimum microphone spacing and maximum array aperture; The total fitness function includes the fitness function contribution corresponding to the average directivity factor of at least two target directions at representative frequency points, the fitness function contribution corresponding to the average crosstalk power of the target directions, the fitness function contribution corresponding to the howling suppression potential, and the fitness function contribution corresponding to the average white noise gain at representative frequency points.

2. The microphone array layout optimization method as described in claim 1, characterized in that: The total fitness function is expressed as: ; , in, This represents the fitness function contribution corresponding to the average directivity factor of the at least two target directions at representative frequency points. This represents the fitness function contribution corresponding to the average crosstalk power in the target direction. This represents the fitness function contribution corresponding to the squeal suppression potential. The fitness function contribution corresponding to the average white noise gain at the representative frequency point is represented by a weighting coefficient for each fitness function contribution.

3. The microphone array layout optimization method as described in claim 1, characterized in that: Determining the globally optimal microphone array layout using a particle swarm optimization algorithm includes: maximizing the directional factor based on the fitness function contribution corresponding to the average directional factor at representative frequency points of the at least two target directions; and maximizing the white noise gain based on the fitness function contribution corresponding to the average white noise gain at the representative frequency points, under the constraint of a set target value.

4. The microphone array layout optimization method as described in claim 1, characterized in that: The fitness function contribution corresponding to the howling suppression potential is expressed as follows: , ,in, Indicates the direction of loudspeaker radiation. This indicates the beam power corresponding to the direction of loudspeaker radiation. For the corresponding weighting coefficients, For the first Particles in the PSO algorithm.

5. The microphone array layout optimization method as described in claim 1, characterized in that: The fitness function contribution corresponding to the average white noise gain at the representative frequency points is expressed as follows: , , , , To set a target value, For frequency weights, The number of representative frequency points, For the corresponding weighting coefficients, For the first Particles in the PSO algorithm.

6. The microphone array layout optimization method as described in claim 1, characterized in that: The target directions are two, corresponding to the main speaking area and the audience interaction area respectively; The fitness function contribution corresponding to the average directivity factor of the two target directions at representative frequency points is expressed as follows: , ,in, For the corresponding weighting coefficients, As a directional factor, The number of representative frequency points, For frequency weights, For the first One target direction, For the first A representative frequency point, For the first Particles in the PSO algorithm.

7. The microphone array layout optimization method as described in claim 6, characterized in that: The fitness function contribution corresponding to the average crosstalk power in the target direction is expressed as follows: , ; in, ( The signal represents the response power of the beam from the first target direction at the second target direction. ( The signal represents the response power of the beam from the second target direction at the first target direction. These are the corresponding weighting coefficients.

8. The microphone array layout optimization method as described in claim 1, characterized in that: The initialization parameters include: The number of particles can be selected from 50 to 100, the maximum number of iterations can be selected from 100 to 500, the inertia weight decreases from 0.9 to 0.4, and the learning factor can be selected from 1.5 to 2.

0.

9. The microphone array layout optimization method according to any one of claims 1 to 8, characterized in that: The overall fitness function is calculated based on the beamforming filter corresponding to the particle in the particle swarm optimization algorithm.

10. A sound reinforcement system, comprising a microphone array, a voice processing device, and a loudspeaker, wherein the microphone array is connected to the voice processing device, and the voice processing device is connected to the loudspeaker, characterized in that: The array layout of the microphone array is determined according to the microphone array layout optimization method according to any one of claims 1 to 9; The voice processing device includes a beamforming module, a voice enhancement module, and a howling suppression module.