On-site sound field adaptive generation system based on multi-modal stage data
The on-site sound field adaptive generation system based on multimodal stage data solves the problems of sound field inhomogeneity and howling in the sound reinforcement system under dynamic stage obstruction and changes in sound source posture, and achieves stable and uniform sound field coverage and clarity maintenance.
Patent Information
- Application Number
- CN202610187935.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing sound reinforcement systems suffer from uneven coverage of the listening area, large fluctuations in clarity, and acoustic feedback when faced with dynamic stage obstruction, changes in sound source posture, and insufficient utilization of non-line-of-sight propagation paths.
An adaptive sound field generation system based on multimodal stage data is adopted. It collects image depth data, raw audio signals and environmental spatial geometry data through a multimodal sensor network, uses a central processing unit to perform timestamp alignment and acoustic voxel map construction, calculates sound source attitude in real time, performs sound ray tracing and calculates adaptive filter parameters and beamforming weights to generate an on-site sound field that matches the stage environment.
It achieves continuous and uniform sound field coverage of the target listening area in complex and ever-changing stage scenes, maintains the clarity of speech and consistency of timbre during the performance, effectively suppresses acoustic feedback howling, and improves the robustness and stability of the system.
Smart Images

Figure CN122054046A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio signal processing and sound reinforcement technology, in particular to a live sound field adaptive generation system based on multi-modal stage data. BACKGROUND
[0002] Modern large-scale stage performances have increasingly high requirements for live sound reinforcement systems. Not only are high sound pressure levels required, but the clarity and uniformity of sound coverage in complex sound field environments are also required. However, in practical applications, traditional sound reinforcement systems face multiple technical challenges.
[0003] Firstly, complex stage mechanical devices, moving scenery, and dynamic scheduling of actors often result in the direct sound propagation path between the sound source and the receiving end being physically blocked. Existing sound reinforcement systems are usually configured based on static sound field models or preset coverage strategies, lacking real-time perception ability for changes in physical space geometry. When the direct path is dynamically blocked, the system cannot automatically identify and utilize the reflecting surfaces in the environment to construct an alternative transmission channel, resulting in severe fluctuations in sound energy in the target listening area, and even sound coverage blind areas, seriously affecting the coherence of the performance listening experience.
[0004] Secondly, the directional variation of the sound source is another important factor affecting sound clarity. Human voice and musical instruments are not ideal point sound sources in physics, especially with directional characteristics in the high frequency band. During dynamic performances, actors will frequently turn around, lower their heads, or quickly move, causing the main radiation axis of the sound source to change rapidly relative to the angle of the sound pickup device or the target area. Existing signal processing techniques cannot obtain real-time three-dimensional pose information of the sound source, and cannot dynamically compensate for off-axis frequency response distortion caused by changes in posture. This causes a large amount of high-frequency information to be lost when the actor is facing away from the target direction, resulting in a decrease in sound clarity and inconsistent brightness of the tone.
[0005] In addition, the stability of the sound reinforcement system is always threatened by acoustic feedback howling. The stage sound field is a time-varying system, and changes in environmental layout, fluctuations in air temperature and humidity will cause the drift of sound propagation path and loop gain characteristics. Traditional feedback suppression methods often rely on static system debugging or post-detection based on signal characteristics, making it difficult to make predictive parameter adjustments when environmental changes cause feedback path changes. This lag causes the system to approach a critical oscillation state when pursuing sufficient loudness, triggering sudden howling, or having to significantly reduce the available gain to maintain stability, limiting the dynamic range and sound reinforcement effect of the system. SUMMARY
[0006] To address the shortcomings of existing technologies, this invention provides a live sound field adaptive generation system based on multimodal stage data. This system solves the problems of uneven sound field coverage, large clarity fluctuations, and acoustic feedback howling caused by the difficulty of live sound reinforcement systems in coping with dynamic stage obstruction, changes in sound source posture, and the lack of effective utilization of non-line-of-sight propagation paths.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solution: an adaptive generation system for live sound field based on multimodal stage data, comprising: a multimodal sensor network, a central processing unit, and a multi-channel sound field generation array.
[0008] The multimodal sensor network is used to collect image depth data, raw audio signals, and environmental spatial geometry data of the stage area and the audience area; The central processing unit is communicatively connected to the multimodal sensor network and is used to perform timestamp alignment operations on the collected data, establish a unified spatiotemporal reference, and construct a dynamic acoustic voxel map containing acoustic material properties. The central processing unit uses visual algorithms to calculate the real-time attitude of the sound source, performs ray tracing calculations based on the dynamic acoustic voxel map, and calculates the physical occlusion loss of the direct path. The central processing unit performs routing decisions between the direct path mode and the reflection path mode based on the physical occlusion loss, and calculates the corresponding adaptive filter parameters and beamforming weights. The multi-channel sound field generation array is connected to the central processing unit and is used to perform acoustic rendering on the digital audio signal according to the adaptive filter parameters and beamforming weights, radiate sound waves outward, and generate a live sound field that matches the stage environment.
[0009] Preferably, the central processing unit is logically divided into multiple execution modules to achieve parallel processing and closed-loop control of the data stream. Specifically, the stage data acquisition module receives the raw data stream and performs timestamp-based frame matching to establish the spatiotemporal reference; the sound field analysis and modeling module constructs the dynamic acoustic voxel map based on the environmental spatial geometry data and calculates the effective radiation spectrum and the cumulative occlusion transfer function of the direct path in conjunction with the real-time attitude of the sound source; the multipath decision module compares the total transmission loss of the direct path with a preset threshold, and searches for an effective reflection plane when the direct path is deemed invalid; the sound field generation and rendering module generates compensation filters and beamforming weights according to the selected transmission path and synthesizes multi-channel drive signals; and the closed-loop feedback correction module monitors the error between the actual sound field response and the target response and dynamically adjusts the system parameters.
[0010] Preferably, the sound field analysis and modeling module employs voxel-based spatial modeling and visually assisted sound source characteristic analysis. The sound field analysis and modeling module discretizes the physical space into voxel units, determines the occupancy state of the voxels based on the point cloud density, and assigns acoustic attribute vectors containing transmission and reflection coefficients to the occupied voxels based on the material recognition results. Simultaneously, the sound field analysis and modeling module uses a visual skeleton key point extraction algorithm to solve the three-dimensional skeleton posture and orientation vector of the sound source, calculates the deviation angle between the main axis direction of the sound source and the direction of the target point, and uses a frequency-dependent directivity transfer function to calculate the effective radiation spectrum radiated by the sound source towards the target direction, compensating for the high-frequency energy loss caused by the rotation of the sound source.
[0011] Preferably, the multipath routing decision module executes path decision logic based on energy cost. The multipath routing decision module calculates the total transmission loss of the direct path, which is the frequency domain superposition of physical obstruction loss calculated by ray tracing, directivity loss caused by sound source attitude, and distance attenuation loss caused by propagation distance. The multipath routing decision module calculates the weighted energy cost index of the total transmission loss within the key acoustic frequency band and compares the weighted energy cost index with dual-threshold hysteresis logic including entry and exit thresholds to generate a mode switching signal. This controls the system to switch to the reflection path mode, avoiding frequent switching of the system in critical states.
[0012] Preferably, when switching to the reflection path mode, the system uses environmental reflective surfaces to construct a non-line-of-sight transmission channel. The multipath routing decision module filters candidate reflective planes in the dynamic acoustic voxel map, calculates the mirror source position of the sound source with respect to the candidate reflective plane and the corresponding physical reflection point using the mirror source method, performs multi-level validity verification on the non-line-of-sight path via the physical reflection point, constructs a total energy cost function including geometric diffusion loss and interface absorption loss, and selects the path with the smallest total energy cost function that does not exceed the maximum available gain of the system as the target transmission channel.
[0013] Preferably, the sound field generation and rendering module adopts differentiated signal processing strategies for different transmission modes. In the direct path mode, the sound field generation and rendering module constructs an inverse filter containing directional deviation compensation components and occlusion loss compensation components based on the physical transmission model of the direct path using the Tikhonov regularization method. In the reflection path mode, the sound field generation and rendering module calculates the spatial phase delay from each unit of the multi-channel sound field generation array to the selected physical reflection point, introduces material inverse filtering parameters containing material absorption loss compensation components of the reflective surface, generates beamforming weights pointing to the physical reflection point, and uses acoustic reflection to complete the sound energy coverage of the occluded area.
[0014] Preferably, the system includes a dynamic beamforming unit to ensure the continuity of sound field changes. The dynamic beamforming unit performs coordinate smoothing tracking, uses a first-order hysteresis filter to smooth the target position calculated by vision, and adaptively adjusts the smoothing factor according to the sound source moving speed. The dynamic beamforming unit also performs soft switching transition. Within the transition time window of path mode switching, the beam weights corresponding to the old and new modes are calculated in parallel, and the instantaneous beam weights are synthesized using constant power interpolation logic based on trigonometric functions to ensure a smooth transition of sound energy during mode switching.
[0015] Preferably, the sound field generation and rendering module adopts a frequency domain parallel processing architecture. The sound field generation and rendering module uses a block convolution architecture with overlapping addition to synthesize signals. In the frequency domain, for each physical channel of the multi-channel sound field generation array, parallel multiplication operations of the original signal spectrum, sound field compensation filter response and beamforming weights are performed, and the signal is restored to a time domain signal sequence through inverse fast Fourier transform to reduce the computational delay of large-scale array processing.
[0016] Preferably, the system achieves active howling suppression through a stability monitoring unit. The stability monitoring unit uses a reference microphone to collect feedback signals, estimates the transfer function of the acoustic feedback path, and calculates the loop gain spectrum across the entire frequency band by combining the system's positive gain. By calculating the difference between the peak value of the closed-loop gain amplitude-frequency response and the critical oscillation condition, the acoustic feedback gain margin index is obtained. When the acoustic feedback gain margin index is lower than a preset threshold, the closed-loop feedback correction module inserts a notch filter at the peak frequency point or performs global gain attenuation to ensure the stability of the sound reinforcement system.
[0017] Preferably, the closed-loop feedback correction module introduces an adaptive update mechanism based on confidence level. The closed-loop feedback correction module obtains the measured sound pressure level spectrum of the audience area and calculates the residual between the measured sound pressure level spectrum and the predicted sound pressure level spectrum. Based on the residual, the frequency domain correction is calculated using negative feedback update logic that includes a signal-to-noise ratio confidence function. The frequency domain correction is used to update the output filter gain or update the material parameters in the dynamic acoustic voxel map, and the model error is corrected according to the actual environmental response.
[0018] This invention provides an adaptive sound field generation system based on multimodal stage data. It has the following advantages: 1. This invention achieves real-time perception of stage dynamic occlusion and sound source orientation by constructing a dynamic acoustic voxel map containing material properties and combining it with visual posture calculation. When the direct sound path is blocked by stage props or people, the system can automatically switch to the reflection path mode based on routing decision logic, and use walls or reflectors in the environment to construct a non-line-of-sight transmission channel. Through the multipath complementary mechanism, it ensures that the target listening area can still obtain continuous and uniform sound field coverage in complex and ever-changing stage scenes, avoiding the sound shadow zone or listening discontinuity caused by physical obstruction in traditional sound reinforcement systems.
[0019] 2. This invention utilizes visual skeleton key point extraction technology to accurately obtain the three-dimensional posture of the sound source and introduces a frequency-dependent directivity transfer function for compensation. Based on the real-time deviation angle between the main axis direction of the sound source and the target direction, an inverse filter and beamforming weight are dynamically generated to correct the high-frequency attenuation of the sound source directivity caused by the actor turning his head, turning around, or moving quickly. This ensures that the sound reinforcement system can still maintain the flatness of the frequency response at the receiving end when the posture of the sound source changes drastically, and improves the speech intelligibility and timbre consistency during dynamic performances.
[0020] 3. This invention integrates a stability monitoring and closed-loop feedback correction mechanism based on loop gain spectrum estimation. By calculating the acoustic feedback gain margin in real time, it can automatically insert a notch filter or perform gain attenuation when the system approaches critical oscillation conditions, thereby effectively suppressing howling while ensuring the maximum usable gain of the system. In addition, the model parameters are updated online using the measured sound pressure level residuals, enabling the system to adapt to acoustic field drift caused by changes in stage setting or temperature and humidity fluctuations, thus improving the robustness and long-term operational stability of the system. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the system functional architecture of the present invention; Figure 2 This is a schematic diagram of the system operation process of the present invention.
[0022] The module includes: 10, multimodal sensor network; 20, central processing unit; 30, multi-channel sound field generation array; 100, stage data acquisition module; 200, sound field analysis and modeling module; 300, multipath routing decision module; 400, sound field generation and rendering module; and 500, closed-loop feedback correction module. Detailed Implementation
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Please see the appendix Figure 1 This invention provides an adaptive generation system for live sound field based on multimodal stage data. The system is configured to be deployed in theaters, concert halls or multifunctional performance spaces for real-time calculation and physical reconstruction of the live sound field.
[0025] The system physically comprises a multimodal sensor network 10, a central processing unit 20, and a multi-channel sound field generation array 30. The multimodal sensor network 10 is distributed across the stage and audience areas to acquire real-time status data of the physical environment. The multimodal sensor network 10 communicates with the central processing unit 20 via a high-bandwidth, low-latency data bus. The central processing unit 20, as the system's computational core, executes sound field analysis, decision-making, and signal rendering algorithms. The multi-channel sound field generation array 30, connected to the central processing unit 20, performs acoustic rendering on digital audio signals, radiates physical sound waves, and generates a live sound field that matches the stage environment.
[0026] The central processing unit 20 is logically divided into several cooperating modules. These modules include: stage data acquisition module 100, sound field analysis and modeling module 200, multipath routing decision module 300, sound field generation and rendering module 400, and closed-loop feedback correction module 500.
[0027] The stage data acquisition module 100 is configured to receive a raw data stream from the multimodal sensor network 10. This raw data stream includes image depth data from a vision sensor, raw audio signals from a microphone array, and spatial geometry data from an environmental sensor. The stage data acquisition module 100 performs timestamp alignment to establish a unified spatiotemporal reference, ensuring synchronization between visual and audio frames.
[0028] The sound field analysis and modeling module 200 is connected to the stage data acquisition module 100. The sound field analysis and modeling module 200 is configured to receive aligned data and use computer vision algorithms to calculate the three-dimensional skeletal pose and orientation vector of the sound source. The sound field analysis and modeling module 200 also constructs a dynamic acoustic voxel map containing absorption and reflection coefficients based on environmental scan data. The sound field analysis and modeling module 200 performs ray tracing calculations in the voxel map, quantifying the radiation directivity characteristics of the sound source and the physical obstruction loss along the direct path.
[0029] The multipath routing decision module 300 is connected to the sound field analysis and modeling module 200. The multipath routing decision module 300 is the core logic unit of the system, configured to compare physical occlusion loss with a preset gain threshold. When it is determined that a direct path is deeply occluded, the multipath routing decision module 300 searches for effective reflection planes in the acoustic voxel map, calculates the energy transmission cost of non-line-of-sight paths, and selects the optimal reflection path that can bypass the obstacle as the target transmission channel.
[0030] The sound field generation and rendering module 400 is connected to the multipath routing decision module 300. This module is configured to calculate the corresponding inverse filter parameters and beamforming weights based on the selected transmission path. In direct path mode, the sound field generation and rendering module 400 generates an inverse filter to compensate for directivity deviation; in reflection path mode, the sound field generation and rendering module 400 generates a relocation filter to compensate for reflection loss and distance attenuation, and sets the beam focus of the speaker array to the selected physical reflection point. The sound field generation and rendering module 400 ultimately synthesizes the multiple digital audio signals driving the multi-channel sound field generation array 30.
[0031] The closed-loop feedback correction module 500 is connected between the output of the multi-channel sound field generation array 30 and the input of the sound field generation and rendering module 400. The closed-loop feedback correction module 500 acquires the actual sound pressure level and frequency response data of the audience area, calculates the error between the actual response and the target response, and dynamically adjusts the filter parameters and the smooth transition coefficient during the path switching process based on the error to maintain the stability of the system.
[0032] See attached document Figure 2 First, the synchronous acquisition and preprocessing of multimodal data is performed. The stage data acquisition module 100 acquires the audio signal and kinematic data of the sound source, while scanning the environmental geometry to update the acoustic element map, and mapping all data to the world coordinate system.
[0033] Subsequently, sound source radiation characteristic modeling and occlusion detection are performed. The sound field analysis and modeling module 200 calculates the effective radiation spectrum based on the real-time attitude of the sound source, and calculates the cumulative occlusion transfer function of the direct path by emitting virtual sound rays in the voxel map.
[0034] Next, a multipath routing decision based on reflector pathfinding is executed. The multipath routing decision module 300 determines whether the occlusion loss of the direct path exceeds the physical compensation limit. If it exceeds the limit, it searches for high-impedance reflectors in the environment, calculates the transmission cost of the non-line-of-sight path via the reflector, and determines the optimal projection path.
[0035] Next, adaptive sound field parameter generation is performed. Based on the decision results, the sound field generation and rendering module 400 calculates the corresponding predistortion filter and array beam weights. If a direct path is selected, a beam pointing to the sound source location is generated; if a reflection path is selected, a beam pointing to the physical reflecting surface is generated to construct a virtual sound source.
[0036] Next, audio signal synthesis and real-time rendering are performed. The sound field generation and rendering module 400 performs frequency domain convolution on the original audio signal and the calculated filters and weights to synthesize a multi-channel driving signal and drive the multi-channel sound field generation array 30 to output and generate a live sound field that matches the stage environment.
[0037] Finally, feedback correction is performed. The closed-loop feedback correction module 500 monitors the sound field state of the target area and fine-tunes the system parameters to ensure that the actual listening experience matches the acoustic model's expectations.
[0038] The stage data acquisition module 100 acquires the physical state information of the sound source through the multimodal sensor network 10. This process mainly includes frequency domain conversion of the audio signal stream and vector calculation of visual kinematic features, and data alignment based on a unified time reference.
[0039] For audio data acquisition and processing, the system utilizes a microphone array distributed across the stage area to acquire raw sound wave signals. The microphone array consists of multiple omnidirectional or cardioid electret or MEMS microphone units used to pick up the time-domain sound pressure level signal emitted by the sound source. The analog signal is preamplified for gain adjustment and then converted into a discrete digital signal via an analog-to-digital converter. To meet the computational requirements of subsequent frequency-domain filtering and beamforming, the system processes the time-domain signal... Perform a Short-Time Fourier Transform (STFT). This process uses a sliding window function (such as a Hamming or Hanning window) to frame the signal, decomposing the non-stationary time-varying audio signal into a series of short-time stationary signals. For the first... Frame signal, its frequency domain expression for: ; in, Represents the discrete-time input signal. For sliding window functions, Angular frequency, The unit is the imaginary unit. After transformation, the system extracts the spectral amplitude and phase information of the sound source at that moment, which are used as the baseband input for subsequent sound source radiation characteristic modeling. For preprocessing operations such as microphone array calibration, noise suppression, and echo cancellation, those skilled in the art can use existing digital signal processing algorithms, which are well-known technologies in the field and will not be elaborated upon here.
[0040] Simultaneously, the system acquires and processes sound source kinematic data through a visual sensing unit. This unit uses a depth camera (RGB-D camera) or an infrared optical motion capture system to acquire point cloud data or infrared marker data containing depth information of the stage area. After acquiring the raw visual data, the system runs a skeletal keypoint extraction algorithm to identify and locate the human topology of the sound source target. This algorithm uses a deep convolutional neural network to regress and predict human joints in the image, outputting a set of skeletal keypoints containing three-dimensional spatial coordinates. ,in Indicates the first The position vector of a joint (such as the center of the head, neck, shoulder, spine, etc.) in the camera coordinate system.
[0041] Based on the extracted skeletal key points, the system further calculates the instantaneous pose vector of the sound source to quantify its physical orientation. Since the sound radiation characteristics of a sound source (such as a singer) are primarily influenced by the head and torso orientation, the system constructs a local reference coordinate system for description. The neck joint position is selected. With the central joint of the head Define the principal axis vector of the head orientation. To improve orientation robustness, the system incorporates orthogonalization correction using the vector connecting the two shoulders. Let the coordinates of the left shoulder be... The coordinates of the right shoulder are Calculate the shoulder lateral vector The direction the sound source faces is vector-like. Calculate and normalize using the following formula: ; Wherein, the vector The system accurately characterizes the normal direction of the sound source's face, which is used to subsequently calculate the deflection angle of the sound source relative to the listener. If the sound source is a musical instrument, the system uses a vector construction method to calculate the direction of the instrument's principal axis of sound radiation by identifying specific geometric feature points on the instrument's surface (such as the headstock and chin rest of a violin).
[0042] To address the issues of inconsistent sampling rates and transmission delays between the visual and audio acquisition systems, the system performs synchronization alignment based on a global clock. A high-precision global timestamp generator is used to tag the acquisition time of both audio and visual data frames. Since the sampling rate of visual data (typically 30Hz to 60Hz) is much lower than that of audio data (typically 44.1kHz or 48kHz), a circular data buffer is established. When calculating sound field parameters, the system uses the timestamp of the current audio frame. Based on this, retrieve timestamps from the visual data buffer. closest The skeletal pose frames, or calculated using linear interpolation methods. The interpolated attitude vectors at each moment ensure a strict correspondence between the input spectral information and attitude information in physical time. All acquired position coordinates and orientation vectors are ultimately transformed to the stage world coordinate system through a pre-calibrated extrinsic parameter matrix. This is provided for subsequent sound field analysis modules to use.
[0043] The sound field analysis and modeling module 200 uses environmental geometric data to construct a three-dimensional spatial mesh model for acoustic calculations. This process includes discretization of the stage space, determination of obstacle occupancy status, and parameterization of acoustic material properties.
[0044] For the digital reconstruction of the stage physical space, the system acquires dense point cloud data of the stage area through environmental sensing units (such as LiDAR or wide-angle depth cameras). To meet the computational efficiency requirements of real-time sound ray tracking, the system does not directly use massive point cloud data, but instead employs spatial voxelization to divide the continuous physical space into discrete sets of cubic units. The system defines the boundary range of the stage space and sets the voxel resolution parameters. (For example, 10cm or 20cm). The system divides the three-dimensional space into... Individual voxel units are used to construct a global voxel map set. Any voxel in the set By its index in the grid coordinate system Unique identifier, its geometric center coordinates Obtained by mapping from the world coordinate system.
[0045] After completing the geometric partitioning, the system performs occupancy state detection. The system iterates through the real-time acquired point cloud data and counts the number of voxels falling into each voxel. The number of point clouds within a voxel. If the point cloud density within a voxel exceeds a preset occupancy threshold, the voxel is marked as occupied, indicating the presence of a physical entity (such as a wall, floor, or set prop); conversely, it is marked as free, indicating that the location is air, allowing sound waves to propagate freely. For voxels in the occupied state, the system further calculates their surface geometry and uses principal component analysis (PCA) or least squares to fit the local plane of the point cloud within the voxel, thereby extracting the surface normal vector of the voxel. The normal vector Used for subsequent calculations of the sound wave reflection angle on the obstacle surface.
[0046] To imbue the voxel map with acoustic physical meaning, the system performs the mapping and assignment of acoustic material properties. This process addresses the problem of traditional geometric maps lacking acoustic boundary conditions. The system utilizes RGB image data acquired by a visual sensor and, through texture analysis and material recognition algorithms, classifies occupied voxels into different material types (such as velvet curtains, wooden floors, concrete walls, metal supports, etc.). The material recognition algorithm can employ a support vector machine (SVM) classifier based on texture feature extraction or a convolutional neural network model, which are techniques readily available to those skilled in the art and will not be elaborated upon here.
[0047] The system has a pre-installed acoustic material database that stores the acoustic parameters of commonly used stage materials in the octave band. Based on the material classification results, the system assigns a value to each occupied voxel. Assign a set of frequency-dependent acoustic property vectors This attribute vector primarily contains the transmission coefficient. With reflection coefficient : ; in, Defined as a complex transmission coefficient, it characterizes the amplitude and phase changes of a sound wave as it passes through the voxel. For vacant voxels, the transmission coefficient modulus is set to 1 (or based on air absorptivity), and the reflection coefficient is set to 0. For occupied voxels; The modulus depends on the sound insulation and thickness of the material; Defined as the complex reflection coefficient, it characterizes the energy and phase response of a sound wave after reflection from the voxel surface. Its modulus is related to the material's sound absorption coefficient. There is a law of conservation of energy, that is Considering the dynamic changes in the stage scene (such as prop movement and curtain raising / lowering), the system implements a dynamic update mechanism for the voxel map. The system uses inter-frame differencing or background subtraction to monitor areas of change in the point cloud data. When a change in the point cloud distribution is detected in a local area, the system only resets the state and recalculates the attributes of the corresponding voxel subset, without reconstructing the entire map. This local update strategy ensures that the acoustic voxel map can reflect changes in the physical layout of the stage in real time, providing accurate physical boundary conditions for subsequent ray tracing and multipath routing calculations.
[0048] To achieve precise fusion of visual position data and acoustic physical parameters, the system establishes a unified global spatial reference system, namely the stage world coordinate system. It also performs spatial rigid body transformation of heterogeneous data and synchronizes it with the time base.
[0049] The system defines the stage world coordinate system. The origin Located at the center of the stage proscenium arch, the ground projection point or a pre-set reference measurement point, its The axis is perpendicular to the stage floor and points upwards. The axis points towards the audience seating area. The axis points to the right side of the stage. All distributed sensor nodes and actuators must map their local measurement data to this coordinate system. For the visual sensing unit, its output skeletal keypoints or point cloud data are initially based on the camera's local coordinate system. The system obtains the rotation matrix of the camera's local coordinate system relative to the world coordinate system through a pre-executed external parameter calibration process. With translation vector For any spatial point captured by a vision sensor Its corresponding coordinates in the world coordinate system Calculated using the following rigid body transformation formula: ; This transformation ensures that the calculated sound source location remains within a consistent physical spatial description regardless of the camera's mounting position and pitch angle. Similarly, for each speaker unit in the sound field generation array... Its position in the local coordinate system of the array Also through the corresponding array extrinsic matrix and Mapped to world coordinates These coordinates are used for subsequent beamforming weight phase delay calculations. For obtaining the extrinsic parameter matrix, those skilled in the art can use the PnP (Perspective-n-Point) algorithm based on a checkerboard calibration board or obtain it through on-site point measurement using a total station.
[0050] Regarding time alignment, the system addresses the clock drift and transmission delay issues between visual data (low sampling rates, such as 30fps or 60fps) and audio data (high sampling rates, such as 48kHz). The system deploys a master clock server based on the IEEE 1588 Precision Time Protocol (PTP), with each sensor node acting as a slave clock device to synchronize its local oscillator via the network, ensuring that the timestamp error of all acquisition devices is controlled within the microsecond range.
[0051] During the data fusion processing phase, the system performs timestamp-based frame matching. Since the arrival times of visual and audio frames often do not perfectly coincide, the system establishes a circular buffer with a time index. Let the start timestamp of the current audio processing frame be... The system searches for two adjacent visual frames in the visual data buffer, whose timestamps are respectively and ,satisfy .
[0052] In order to obtain The system maintains precise physical states of the sound source at all times, and analyzes the skeletal pose vectors in two visual frames. Performs spherical linear interpolation (Slerp) or linear interpolation. The interpolated attitude vector... The calculation is as follows: ; This step eliminates time jitter caused by non-integer multiples of the sampling rate, ensuring that the geometric position referenced in the acoustic model calculation corresponds strictly to the current acoustic wave phase state in physical time, thereby avoiding beam pointing deviation or phase cancellation caused by spatiotemporal misalignment.
[0053] The Sound Field Analysis and Modeling Module 200 abandons the idealized assumption of treating the sound source as an omnidirectional point source in traditional sound field control. Instead, it accurately calculates the effective sound energy radiated by the sound source in a specific spatial direction based on real-time orientation vectors obtained from visual calculations. This process involves establishing spatial geometric relationships, mapping frequency-dependent directivity functions, and synthesizing the effective radiation spectrum.
[0054] When constructing spatial geometric relationships, the system first determines the relative position vector between the sound source and the target point. Let the real-time position of the sound source in the world coordinate system be... This location is determined by the aforementioned visual skeleton keypoints. Let the location of any target point to be evaluated in space be... The target point could be a location in the audience seating area. It could also be the center point of a potential reflective surface in a voxel map. The system uses vector algebra to calculate the normalized unit direction vector from the sound source to the target point. To quantify the deviation between the sound source radiation direction and the target direction, the system utilizes the unit vector of the sound source orientation output by the vision module. (That is, the direction of the acoustic principal axis of the sound source). Through calculation... and The angle between them gives the instantaneous deviation angle. The deviation angle characterizes the degree to which the target point deviates from the front of the sound source: when When it is close to 0 degrees, it indicates that the target point is located in the direction of the main axis of the sound source; when When the angle is close to 180 degrees, it indicates that the target point is located behind the sound source.
[0055] Based on the calculated deviation angle, the system calls a pre-defined directivity model library to determine the attenuation characteristics. This library stores normalized directivity transfer functions for different types of sound sources (such as tenor, soprano, and musical instruments). This function describes the amplitude of the sound wave as a function of the deviation angle. and angular frequency The variation pattern reflects the physical characteristics of a sound source: strong directivity in the high-frequency range and omnidirectionality in the low-frequency range.
[0056] This invention employs a parameterized frequency-dependent model to calculate directional gain in real time. For a human voice source, the directional transfer function... Expressed as: ; In the formula, The deviation angle is the one calculated above; Defined as a frequency-dependent lobe sharpness factor. This sharpness factor It is angular frequency The monotonically increasing function is used to simulate the physical phenomenon that the shorter the wavelength of a sound wave, the weaker its diffraction ability, resulting in a narrower sound beam. For low-frequency components, The value is small, making The value changes smoothly from all angles; for high-frequency components, The value is large, making It decays rapidly as the angle deviates.
[0057] After obtaining the instantaneous directivity transfer function, the system combines it with the acquired original sound source spectrum. Calculate the effective spectrum of the sound source actually radiated in the direction of the target. This calculation is achieved by multiplying the original spectrum by the directivity transfer function point-by-point in the frequency domain: ; This effective spectrum accurately quantifies the natural changes in sound and tone caused by actors turning, facing sideways to the audience, or turning their backs to the audience, such as the roll-off of high-frequency energy. This calculation provides a physical basis for subsequent system decisions, enabling the system to distinguish between attenuation caused by distance and directional attenuation caused by the orientation of the sound source, thereby achieving more accurate sound field compensation.
[0058] After constructing a voxel map containing acoustic properties and establishing the spatial relationship between the sound source and the target point, the sound field analysis and modeling module 200 further quantifies the degree of physical obstruction on the direct sound path. This process simulates the absorption and blocking effects experienced by sound waves on a straight propagation path by performing a ray casting algorithm in a discretized voxel grid, thereby generating a cumulative obstruction transfer function that reflects the acoustic transparency of the physical environment.
[0059] The system first establishes the connection points of the sound sources in the stage world coordinate system. Location relative to the target audience The line segment connecting the line of sight, i.e., the line-of-sight path, is geometrically described by a spatial straight-line equation. To trace this path in a discrete voxel map, the system employs either a 3D spatial digital differential analysis (3D DDA) algorithm or the Bresenham straight-line algorithm to rasterize the line-of-sight path. Starting from the initial voxel containing the sound source, the algorithm iteratively traverses all voxel unit sequences traversed by the line segment in physical space along the direction vector of the line-of-sight path. This process identifies the set of voxels actually traversed in physical space by the direct sound path.
[0060] During the traversal of the voxel set, the system checks the occupancy status of each voxel on the path. For voxels marked as vacant, the system treats them as air medium, ignoring their influence or only including standard air absorption attenuation when calculating occupancy loss. For voxels marked as occupied, the system identifies them as obstacles on the path and extracts the frequency-dependent transmission coefficient corresponding to that voxel based on the data stored in the voxel map. Since obstacles in the stage environment (such as curtains, scenery, and actors themselves) have different blocking capabilities for sound waves of different frequencies, the system addresses each occupied voxel... Extract its complex transmission coefficient .
[0061] Based on the acoustic properties of all obstacle object pixels along the path, the system calculates the cumulative occlusion transfer function of the direct path. This function models the continuous physical occlusion effect along the path as a cascaded linear filter system, characterizing the spectral changes of the sound wave after it traverses the entire path solely due to physical occlusion. It assumes that the total number of occupied voxels identified along the path is... The cumulative occlusion transfer function is calculated using the following formula: ; In the formula, This is the complex transmission coefficient corresponding to that voxel. If there are no occupied voxels on the path (i.e., ...), ... ),but A modulus of 1 indicates that the direct path is unobstructed.
[0062] To facilitate threshold determination in subsequent routing decision logic, the system further calculates the amplitude of the aforementioned cumulative occlusion transfer function and converts it into a sound pressure level loss value in the logarithmic domain, i.e., occlusion loss. (Unit: decibel). This calculation follows the standard signal processing definition, which is to take the common logarithm of the cumulative occlusion transfer function magnitude and multiply it by -20. This loss value... It directly reflects the amount of energy attenuation of a specific frequency component during transmission.
[0063] Furthermore, this computational model implicitly incorporates the diffraction effect of sound waves through frequency-dependent transmission coefficients. During the attribute assignment phase of the voxel map, for voxels of the same material, the transmission coefficient modulus in the low-frequency band is set higher than that in the high-frequency band. This results in the calculation showing that low-frequency sound waves bypass small obstacles, while high-frequency sound waves exhibit blocking characteristics. Through the above calculations, the system obtains a loss curve that varies with frequency. This curve accurately describes the acoustic transparency of the direct path, providing a quantitative physical basis for the system to determine whether to continue direct path transmission or switch to a reflection path.
[0064] The multipath routing decision module 300, as the central logic unit of the system, is configured to receive quantized data from the sound field analysis and modeling module 200 and make a decision between the direct compensation mode and the reflection relay mode based on the physical energy criterion. This process is achieved by calculating the total transmission loss of the direct path and comparing it with the physical compensation capability limit of the system.
[0065] The multipath routing decision module 300 first performs a comprehensive calculation of the total transmission loss. This calculation integrates the blocking loss caused by physical obstruction, the directional loss caused by the deflection of the sound source attitude, and the distance attenuation loss caused by the propagation distance. For the distance attenuation loss, the system calculates it based on the well-known principle of spherical sound wave diffusion (i.e., the inverse square law), and this loss value is proportional to the logarithm of the distance from the sound source to the target point. For the directional loss, the system converts the normalized directional transfer function obtained in the previous steps into a decibel value in the logarithmic domain. Based on this, the system obtains the total loss spectrum of the direct path across the entire frequency band through linear superposition. : ; In the formula, This refers to the occlusion loss calculated based on ray tracing. For directional loss, This represents the distance attenuation loss. This summation process is performed one by one at each discrete frequency point in the system's frequency domain processing unit, generating a complete frequency response curve that reflects the cost of acoustic energy transmission along the direct path.
[0066] To transform multi-dimensional spectral data into a single, decision-governing scalar metric, the system performs a band-weighted energy cost assessment. Considering the characteristics of human hearing and the effective operating frequency bands of the sound reinforcement system, the system selects key acoustic frequency bands for weighted integration to generate an energy cost metric for the direct path. The specific calculation formula for this indicator is as follows: ; In the formula, and These are the lower and upper angular frequencies of the evaluation band, respectively. This is a preset frequency weighting function. The distribution pattern is set according to the actual application scenario to highlight the frequency band of speech intelligibility or the fundamental frequency band of a specific instrument. This integral operation is implemented in the digital system through discrete summation, and the calculation result... It characterizes the weighted average signal attenuation of the current direct path within the key auditory frequency band.
[0067] Obtaining energy cost indicators The system then compares it with a preset maximum system compensation gain threshold. To avoid frequent path mode switching due to data fluctuations in critical states, the system includes an entry threshold. With exit threshold The dual-threshold hysteresis comparison logic, and satisfies .
[0068] When the system is currently in direct compensation mode, the multipath routing decision module 300 determines the real-time calculated... Is it greater than the entry threshold? If the judgment result is yes, it indicates that the direct path is severely blocked, and simply increasing the gain is no longer sufficient to compensate or may lead to system instability. The system determines that the direct path has failed and switches the control state to reflection relay mode, triggering the subsequent reflection surface search process. When the system is currently in reflection relay mode, the multipath routing decision module 300 continuously monitors the loss of the direct path and makes a judgment. Is it less than the exit threshold? If the judgment result is yes, it indicates that the obstruction has been removed or the sound source attitude has been restored, the direct path has been restored to unobstructed, the system determines that the direct path is available, and switches the control state back to direct compensation mode. Through this hysteresis logic, the system ensures the stability of the sound field reconstruction strategy during the dynamic changes of the physical environment.
[0069] When the system determines that a direct path is unavailable and switches to reflection relay mode, the multipath routing decision module 300 initiates a reflective surface search program based on the physical environment. This program aims to identify surfaces capable of effectively transmitting acoustic energy from a dynamic acoustic voxel map and calculate the precise geometric reflection point locations.
[0070] The multipath routing decision module 300 performs a candidate reflective surface filtering operation based on the material attribute data in the acoustic voxel map. The system traverses the set of occupied voxels in the map and sets a filtering threshold based on the modulus of the acoustic reflection coefficient. The system retains only voxels with an acoustic reflection coefficient greater than the preset threshold and filters out areas covered by sound-absorbing materials.
[0071] The system uses a region growing algorithm or a normal vector clustering algorithm to aggregate qualified voxels that are spatially continuous and have the same normal vector direction into an independent set of candidate reflection planes. For each candidate reflection plane The system extracts its geometric center point. Unit normal vector And the physical boundary contour.
[0072] For each candidate reflecting plane in the set, the system performs geometric path calculations based on the image source method. The system first calculates the sound source location. Regarding planes virtual image source location This calculation is based on the principles of geometric optics, specifically the location of the mirror source. Location of the sound source The formula for calculating the vector of a symmetric projection along the direction of the plane's normal vector is as follows: ; In the formula, Let be the position vector of the sound source in the world coordinate system. Let be the position vector of any point on the plane (here, the geometric center of the plane is used). Let be the unit normal vector of the plane.
[0073] After determining the location of the mirror source, the system builds a connection to the mirror source. With the target audience A virtual straight path. This virtual path is related to the candidate reflection plane. The geometric intersection is the theoretical physical reflection point. The system solves for the coordinates of the intersection point by simultaneously solving the parametric equations of the line and the plane in space.
[0074] Obtain the theoretical reflection point Then, the system executes multi-level validity verification logic to eliminate physically infeasible paths.
[0075] The first level of verification is physical boundary constraint verification. The system uses a point-within-a-polygon determination algorithm from computational geometry to determine the calculated reflection point. Is it located in the candidate plane? Within the actual physical contour range. If the intersection point falls outside the plane boundary, it indicates that the plane cannot provide specular reflection under the current geometric relationship, and the system discards the candidate plane.
[0076] The second level of verification is the occlusion accessibility verification. This verification process is broken down into incident path verification and reflection path verification. The system again calls the aforementioned voxel-based occlusion detection algorithm to perform ray projection checks on the path segment from the sound source to the reflection point and the path segment from the reflection point to the target point, respectively. Only when the cumulative occlusion loss of these two path segments is lower than a preset accessibility threshold is the system determined that the reflection path is physically connected.
[0077] The third level of verification is the validity verification of the reflection angle. The system calculates the incident angle of the sound wave at the reflection point. , that is, the angle between the incident vector and the plane normal vector. The system has a preset critical grazing angle threshold; if the incident angle... If the threshold is exceeded, it indicates that the sound wave is passing over the surface at a near-parallel angle, making it difficult to form effective specular reflection energy, and the system will ignore this path.
[0078] After verifying all candidate planes, if multiple feasible reflection paths exist, the system calculates the combined transmission cost of each path. To select the optimal reflective surface. Overall transmission cost. Taking into account both the spherical diffusion attenuation caused by the total path length and the absorption attenuation caused by the reflective surface material, the calculation formula is as follows: ; In the formula, The distance is the Euclidean distance from the sound source to the reflection point; The distance is the Euclidean distance from the reflection point to the target point; Let be the magnitude of the complex reflection coefficient of the candidate plane at the current incident angle. In the formula... It is a standard mathematical operator in the fields of acoustics and signal processing that converts the ratio of field quantities (such as sound pressure, voltage, distance, etc.) into decibel (dB) scale.
[0079] Specifically, since the sound pressure amplitude of a point source is inversely proportional to the distance ( Its sound pressure level attenuation follows 20 Therefore, the first term of the formula uses a coefficient of 20 to calculate the sound pressure level transmission loss caused by geometric diffusion; similarly, the reflection coefficient... Essentially, it's the ratio of reflected sound pressure to incident sound pressure; therefore, the second term of the formula also uses a coefficient of 20 to calculate the sound pressure level loss (i.e., reflection loss) during the reflection process. Multipath routing decision module 300 selects... The shortest path is used as the final target reflection channel, and the corresponding reflection point coordinates are... Output to the sound field generation and rendering module 400.
[0080] After the multipath routing decision module 300 identifies geometrically connected candidate reflection paths, the system needs to further quantify and evaluate these paths from the perspective of energy transmission efficiency. This evaluation process aims to calculate the residual acoustic energy level of each non-line-of-sight path when it reaches the target point after overcoming geometric diffusion and interface absorption, and convert it into a normalized energy cost index so as to match it with the system gain capability.
[0081] The system establishes a frequency domain transmission model for non-line-of-sight paths. This model decomposes the propagation process of sound waves from the sound source through the reflection point to the receiver into a free-field propagation stage and an interface reflection stage. For any given path... For each candidate non-line-of-sight path, the system first calculates its total acoustic path length. This length is the sum of the Euclidean distance from the sound source to the reflection point and the Euclidean distance from the reflection point to the receiving point.
[0082] For the free field propagation stage, the system calculates the geometric diffusion loss. This calculation is based on the well-known principle of spherical wave diffusion from a point source (i.e., the inverse square law), reflecting the physical fact that sound energy density attenuates with increasing propagation distance. The system converts this physical attenuation into a decibel value in the logarithmic domain, the value of which is equal to the total sound path length. 20 times the common logarithm of the ratio to the reference distance.
[0083] Regarding the interface reflection stage, the system evaluates the impact of the reflecting interface on the spectral characteristics of the acoustic signal. Since the reflecting surface is not an ideal rigid surface, its surface acoustic impedance causes partial energy absorption and abrupt phase changes during sound wave reflection. Based on the material parameters stored in the aforementioned acoustic voxel map, the system extracts the frequency-dependent complex reflection coefficients at the reflection points. .in, Angular frequency, For the first The incident angle of the sound wave along the path on the reflecting surface. Based on this, the system calculates the interface absorption loss spectrum caused by the reflection. The calculation formula is as follows: ; In the formula, Indicates the modulus of a complex number; The loss in decibels used to convert the magnitude of the reflection coefficient (typically ranging from 0 to 1) into a positive value. This loss spectrum... This typically manifests as a low-pass filter with higher loss in the high-frequency band than in the low-frequency band.
[0084] To obtain a comprehensive index reflecting the overall transmission quality of the path, the system constructs a total energy cost function. This function combines the broadband attenuation caused by distance and the frequency-selective attenuation caused by material properties, and introduces a weighting factor based on human auditory sensitivity to quantify the system compensation cost required if this path is adopted. The calculation formula is as follows: ; In the formula, The geometric diffusion loss value obtained from the aforementioned calculation; and The effective operating bandwidth of the system is defined (e.g., 20 Hz to 20 kHz). This is a frequency weighting function, which is set according to the spectral importance of the human ear's equal loudness curve or specific application scenarios (such as voice amplification or music playback) to give higher evaluation weight to the core frequency band. This is a preset stability margin constant used to compensate for the decrease in signal-to-noise ratio caused by environmental noise or measurement errors. The integral term represents the weighted average material loss of the reflection path within the effective bandwidth.
[0085] The system calculates The numerical value directly represents the result of choosing this number. The additional gain compensation required by the sound reinforcement system when using a non-line-of-sight path as the main transmission channel. The multipath routing decision module 300 will consider all candidate paths... Sort the values and remove those. The value exceeds the system's maximum available gain. The path is selected. This step, from the perspective of physical energy conservation, ensures that the selected reflection path is not only geometrically accessible but also feasible within the power playback capability of the electroacoustic system, avoiding speaker overload or acoustic feedback howling problems caused by forcibly compensating for high-loss paths. For the remaining paths after screening, the system selects... The shortest path is the optimal non-line-of-sight transmission link.
[0086] Based on the path mode decision results output by the multipath routing decision module 300, the sound field generation and rendering module 400 activates the direct sound inverse filter design sub-logic or the reflection virtual sound source relocation filter design sub-logic respectively to generate the digital filter coefficients finally loaded onto the speaker array processor.
[0087] When the system operates in direct compensation mode, its purpose is to eliminate spectral distortion caused by sound source attitude deflection or slight physical obstruction, so that the timbre perceived by the listener is close to the ideal state when the sound source is directly facing them. The system is based on the deviation angle obtained from the visual sensor. and the occlusion transfer function calculated by ray tracing Construct a physical transmission model to be compensated. The model is represented as the product of the directional transfer function and the occlusion transfer function, i.e. .
[0088] To recover the original signal, the system needs to design a compensation filter whose ideal frequency response is the reciprocal of the physical transmission model. Considering that the physical transmission model has minimum values at certain frequency points (such as high-frequency notch filtering or severe blockage), directly taking the reciprocal will lead to excessive compensation gain, which in turn will cause system overload or amplify background noise.
[0089] Therefore, this invention employs the Tikhonov regularization method to design the inverse filter. This method limits the filter's output energy while minimizing the reconstruction error. The frequency response of the compensated filter is then considered. The calculation formula is as follows: ; In the formula, Angular frequency, Represents the complex conjugate of the physical transport model, used to correct phase distortion; Indicate its power spectrum; This is a frequency-dependent regularization parameter. The value of is related to the current signal-to-noise ratio of the system: in frequency bands with high signal-to-noise ratios, Choose a smaller value to make the filter approximate the characteristics of an ideal inverse filter; in frequency bands with low signal-to-noise ratios or where deep notches appear, A larger value is chosen to limit the maximum compensation gain and prevent noise amplification.
[0090] Obtaining the frequency domain response The system then uses the well-known Inverse Fast Fourier Transform (IFFT) algorithm to convert it into finite-length unit impulse response (FIR) filter coefficients in the time domain. To eliminate the Gibbs effect (i.e., ringing) caused by frequency domain truncation, the system applies a Hanning window or Hamming window to smooth the impulse response in the time domain. The resulting FIR filter coefficients are then dynamically updated in the digital signal processing unit of the speaker array for real-time convolution processing of the input audio signal.
[0091] When the system is operating in reflection relay mode, the sound field generation and rendering module 400 no longer attempts to directly cover the audience area, but instead redirects the focus of the speaker array to the optimal reflection point found in the aforementioned steps. Virtual sound sources are generated by using walls or ceilings as acoustic mirrors.
[0092] The system first calculates the spatial phase delay required for beamforming. This is done for each speaker unit in the array. The system calculates its distance to the reflection point. Euclidean distance To ensure that the sound waves radiated by all units are superimposed in phase at the reflection point, forming an energy caustic point, the system calculates the compensation time delay for each unit based on the well-known time-delay summation beamforming principle. This delay is determined by distance. Divide by the speed of sound This ensures that the wavefronts emitted from each unit arrive at the reflection point simultaneously.
[0093] While achieving spatial focusing, the system must compensate for the spectral loss caused by the reflective surface material. The system reads the complex reflection coefficient corresponding to the optimal reflection point. Because the reflection process is usually accompanied by energy absorption (i.e., The spectrum of the reflected sound signal will change. In order to make the reflected sound that finally reaches the listener's ear have natural frequency response characteristics, the system introduces a material inverse filtering component into the beamforming weight.
[0094] Taking into account both spatial focusing and material compensation, the system is the first in the array Each speaker unit generates a complex frequency domain weighting coefficient. The calculation formula is as follows: ; In the formula, the first part For material compensation items, among which To prevent small positive numbers with a denominator of zero, this term serves to increase the frequency components absorbed by the reflecting surface; Part Two For spatial phase alignment terms, where The imaginary unit controls the direction of the main lobe of the beam towards the reflection point; Part Three These are the coefficients of the spatial apodized window function, used to weight the array aperture (such as Chebyshev weighting) to suppress sidelobe leakage caused by beamforming and reduce acoustic interference to non-target areas.
[0095] Through the above processing, the sound waves emitted by the speaker array are reflected at the point... The sound waves converge and are reflected, and the pre-compensated spectrum they carry cancels out the absorption characteristics of the reflecting surface. To the listener, the sound waves appear to originate from a mirrored virtual sound source located behind the wall, while retaining the original timbre characteristics of the sound source, thus achieving high-fidelity sound reinforcement under non-line-of-sight conditions.
[0096] After the sound field generation and rendering module 400 determines the target filtering coefficients and spatial directionality, the dynamic beamforming unit is responsible for converting these mathematical parameters into physical control signals to drive the speaker array. This process involves not only static beamforming, but more importantly, handling the dynamic transitions when the sound source moves or the path mode switches, in order to prevent perceived spatial jumps or signal breaks.
[0097] The system implements a target tracking strategy based on coordinate smoothing. Considering the high-frequency jitter or measurement noise in the coordinate data output by the visual sensor, directly using it for beam steering calculation would cause frequent fluctuations in the phase center of the speaker array, producing auditory artifacts resembling mechanical noise. Therefore, a first-order hysteresis filter is introduced before beam weight calculation. Let... The original target position calculated by vision at any given time is The system calculates the smoothed control target position based on the following formula. ; In the formula, The control position at the previous moment; This is a smoothing factor with a value range of (0,1). This factor is adaptively adjusted based on the sound source's movement speed: when the sound source is detected to be stationary or moving slowly, the system decreases the input value to enhance positional stability; when the sound source is detected to be moving rapidly, the system increases the input value. This value is used to reduce tracking delay and ensure that the beam can cover the sound source in a timely manner.
[0098] For the path mode switching process (i.e., switching from direct compensation mode to reflection relay mode, or vice versa), this invention designs a soft handover mechanism that preserves energy. When the multipath routing decision module 300 issues a mode switching command, the system does not instantly change the beam pointing, but instead initiates a transition time window of a preset duration. Within this window, the system calculates the beam weight vector corresponding to the old path mode in parallel. and the beam weight vector corresponding to the new path mode .
[0099] The system uses dynamic weighted interpolation to synthesize instantaneous beam weights. To ensure a constant total radiated energy during beam pointing deflection and avoid sudden changes in volume, the system employs constant power interpolation logic based on trigonometric functions, calculated using the following formula: ; In the formula, For normalized time progress variables, as of time The beam weights increase linearly from 0 to 1 within the transition window. This formula utilizes the property that the sum of the squares of the sine and cosine functions is 1, ensuring that the energy superposition of the two beam weights remains normal during the transition. This achieves a smooth deformation of the sound field focus rather than an abrupt jump, making the audience feel that the sound naturally slides from the center of the stage to the wall reflection point, rather than a hard cut-off of the sound source.
[0100] After obtaining the real-time beam weights and compensation filter coefficients, the signal synthesis unit performs the final digital audio processing to generate a voltage signal to drive the multi-channel amplifier.
[0101] The system employs a block convolutional architecture with overlapping addition or overlapping preservation to process continuous audio streams. For the input raw audio signal... The system divides the data into fixed-length frames and performs windowing to reduce spectral leakage. Then, the system uses the Fast Fourier Transform (FFT) algorithm, which is common in the field, to convert the time-domain signal into a frequency-domain representation. In the frequency domain, the system targets each physical channel in the array. Parallel multiplication is performed. This operation fuses the original signal spectrum, the acoustic field compensation filter response, and the dynamic beamforming weights. Frequency domain output signal of each channel The calculation is as follows: ; In the formula, The spectrum of the source signal; The frequency response compensation function for direct or reflected sound generated in the preceding steps; This is a beamforming coefficient that incorporates spatial phase delay and aperture apodization weighting. This frequency domain multiplication is equivalent to a linear convolution operation in the time domain, enabling synchronous control of the signal's time-domain waveform and spatial pointing.
[0102] After completing frequency domain synthesis, the system performs frequency domain synthesis on each channel. The system performs an inverse fast Fourier transform to restore the signal to a time-domain sequence. During this stage, the system superimposes the end of the current frame with the end of the previous frame in the time domain, according to the rules of the overlap-add method, thereby reconstructing a continuous and seamless output waveform.
[0103] Before the signal is output to the digital-to-analog converter, the system incorporates peak limiting and dynamic range control. Considering that deep compensation at certain frequencies can cause digital signal amplitude overflow, a look-ahead limiter is deployed. This limiter monitors the peak level of the synthesized signal in real time. Once it detects that the signal amplitude is about to exceed the system's allowable full-scale threshold, it introduces attenuation gain with an extremely fast start-up time, ensuring that the output signal remains physically within its linear operating range and protecting the speaker units from overload damage. The multi-channel digital signal, after the above processing, is finally converted into an analog voltage, driving the phased array to reconstruct a sound field with specific directivity, frequency characteristics, and spatial location in physical space.
[0104] To ensure stability under high-gain sound reinforcement conditions, the system includes an independently operating stability monitoring unit. This unit uses a reference microphone deployed near the speaker array or directly utilizes microphone units in the array that are not emitting signals to acquire feedback signals from the ambient sound field in real time, and estimates the transfer function of the acoustic feedback path based on adaptive filtering principles.
[0105] The stability monitoring unit compares the drive signal of the loudspeaker array with the feedback signal collected by the reference microphone, and uses well-known algorithms in the art, such as the normalized minimum mean square error algorithm or the recursive least squares algorithm, to identify the impulse response of the acoustic feedback channel from the loudspeaker to the microphone in real time. The system calculates the closed-loop gain in the frequency domain to quantify the stability risk of the current system. The system defines a certain moment... Estimation of the transfer function of the acoustic feedback loop Combined with the current system positive gain Calculate the loop gain spectrum across the entire frequency band.
[0106] To quantify the safe distance between the system and the occurrence of howling (self-oscillation), the system calculates the acoustic feedback gain margin (GBF) index. This index is defined as the difference between the peak value of the closed-loop gain amplitude-frequency response and the critical oscillation condition (i.e., gain of 1 or 0 dB), and its calculation formula is as follows: ; In the formula, The complex transfer function of the estimated acoustic feedback path; The complex transfer function of the total system gain, which includes beamforming weights and compensation filters; This represents the modulo operation; This indicates that the minimum value is sought within the system's effective operating bandwidth. The value indicates how many decibels of gain the system needs to increase at the most dangerous frequency point to trigger a howling sound. If If the value is below the preset safety threshold, it indicates that the system is in a critically unstable state.
[0107] When insufficient acoustic feedback gain margin or deviation from the expected sound field coverage is detected, the closed-loop feedback correction module 500 is triggered to execute parameter fine-tuning or emergency suppression strategies.
[0108] To address the insufficient acoustic feedback gain margin, the system employs a hierarchical control strategy combining notch suppression and global gain attenuation. When the value is below the warning threshold but above the danger threshold, the system identifies the peak frequency point in the loop gain spectrum. A high-Q digital notch filter is dynamically inserted at this frequency point. The center frequency of this notch filter is set to... The attenuation depth is adaptively set based on the current margin gap to disrupt the positive feedback condition without affecting the overall listening experience. When the value falls below the danger threshold, the system immediately performs a step-by-step reduction in the total broadband gain until the margin is restored to a safe range.
[0109] To address sound field model biases, the system performs model calibration based on measurement errors. Because the sound absorption coefficients of materials in the voxel map differ from the actual physical environment, the predicted sound pressure level may not match the sound pressure level actually received at the listener's location. The system uses calibration microphones placed at key locations in the sound field (or utilizes mobile terminals as crowdsourced sensors) to acquire the actual sound pressure level spectrum. The system calculates the measured value and the predicted value from the sound field analysis module. The residuals between them are used to correct the material parameters in the voxel map or directly correct the output filter gain.
[0110] The update logic of the gain correction filter follows the principle of negative feedback control, and its frequency domain correction is... The calculation is as follows: ; In the formula, To update the step size factor, which is used to control the correction speed to avoid system oscillation; and The measured and predicted sound pressure level logarithmic spectra are shown below; The confidence function is set based on the signal-to-noise ratio of the measurement location. It takes a value close to 1 in the frequency band with high signal-to-noise ratio and a value close to 0 in the frequency band with large environmental noise interference, thereby preventing noise interference from introducing incorrect gain correction.
[0111] Through the aforementioned closed-loop correction mechanism, the system can not only prevent the risk of acoustic feedback howling, but also continuously accumulate data and correct the acoustic parameters of the physical environment during operation. For example, when the system detects that the signal reflected by a certain wall is consistently weaker than expected, it will automatically reduce the reflection coefficient value of the corresponding area in the voxel map, and reduce the priority of that reflecting surface or increase the compensation gain in subsequent path planning, thereby realizing the adaptive evolution of the sound field control system to changes in the physical environment.
Claims
1. A live sound field adaptive generation system based on multimodal stage data, characterized in that, include: Multimodal sensor networks are used to acquire image depth data, raw audio signals, and environmental spatial geometry data of the stage area and audience area; The central processing unit is communicatively connected to the multimodal sensor network and is used to perform timestamp alignment operations on the collected data, establish a unified spatiotemporal reference, construct a dynamic acoustic voxel map containing acoustic material properties, calculate the real-time attitude of the sound source, and perform ray tracing operations based on the dynamic acoustic voxel map to calculate the physical occlusion loss of the direct path. The central processing unit performs routing decisions between the direct path mode and the reflection path mode based on the physical occlusion loss, and calculates the corresponding adaptive filter parameters and beamforming weights. A multi-channel sound field generation array, connected to the central processing unit, is used to perform acoustic rendering on digital audio signals based on the adaptive filter parameters and beamforming weights, radiate physical sound waves, and generate a live sound field that matches the stage environment.
2. The adaptive generation system for live sound field based on multimodal stage data according to claim 1, characterized in that, The central processing unit is logically divided into: The stage data acquisition module is used to receive the raw data stream and perform a timestamp-based frame matching operation to establish the spatiotemporal reference. The sound field analysis and modeling module is used to construct the dynamic acoustic voxel map based on the environmental spatial geometry data, and to calculate the effective radiation spectrum and the cumulative occlusion transfer function of the direct path in combination with the real-time attitude of the sound source. The multipath routing decision module is used to compare the total transmission loss of the direct path with a preset threshold, and to search for a valid reflection plane when the direct path is determined to be invalid. The sound field generation and rendering module is used to generate compensation filters and beamforming weights based on the selected transmission path, and to synthesize multi-channel drive signals. The closed-loop feedback correction module is used to monitor the error between the actual sound field response and the target response, and to dynamically adjust the system parameters.
3. The adaptive generation system for live sound field based on multimodal stage data according to claim 2, characterized in that, The sound field analysis and modeling module is used for: The physical space is discretized into voxel units, and the occupancy state of the voxels is determined based on the point cloud density. For voxels in the occupancy state, an acoustic attribute vector containing the transmission coefficient and the reflection coefficient is assigned based on the material recognition result. The sound field analysis and modeling module uses a visual skeleton key point extraction algorithm to solve the three-dimensional skeleton pose and orientation vector of the sound source, and calculates the deviation angle between the main axis direction of the sound source and the direction of the target point. It also uses a frequency-dependent directivity transfer function to calculate the effective radiation spectrum radiated by the sound source towards the target direction.
4. The adaptive generation system for live sound field based on multimodal stage data according to claim 2, characterized in that, The multipath routing decision module is used to calculate the total transmission loss of the direct path. The total transmission loss is the frequency domain superposition of physical occlusion loss calculated by ray tracing, directivity loss caused by sound source attitude, and distance attenuation loss caused by propagation distance. The multipath routing decision module calculates the weighted energy cost index of the total transmission loss in the key acoustic frequency band, compares the weighted energy cost index with the dual-threshold hysteresis logic including entry threshold and exit threshold, generates a mode switching signal, and controls the system to switch to the reflection path mode.
5. The adaptive generation system for live sound field based on multimodal stage data according to claim 4, characterized in that, When switching to the reflection path mode, the multipath routing decision module filters candidate reflection planes in the dynamic acoustic voxel map and uses the mirror source method to calculate the mirror source position of the sound source with respect to the candidate reflection plane and the corresponding physical reflection point. The multipath routing decision module performs multi-level validity verification on the non-line-of-sight path via the physical reflection point, and constructs a total energy cost function that includes geometric diffusion loss and interface absorption loss. The path with the smallest total energy cost function that does not exceed the maximum available gain of the system is selected as the target transmission channel.
6. The adaptive generation system for live sound field based on multimodal stage data according to claim 2, characterized in that, The sound field generation and rendering module is used for: In the direct path mode, based on the physical transmission model of the direct path, an inverse filter containing a directional deviation compensation component and an occlusion loss compensation component is constructed using the Tikhonov regularization method. In reflection path mode, the spatial phase delay from each unit of the multi-channel sound field generation array to the selected physical reflection point is calculated, and material inverse filtering parameters including the material absorption loss compensation component of the reflective surface are introduced to generate beamforming weights pointing to the physical reflection point.
7. The adaptive generation system for live sound field based on multimodal stage data according to claim 6, characterized in that, The sound field generation and rendering module further includes a dynamic beamforming unit, which is used for: Perform coordinate smoothing tracking, use a first-order hysteresis filter to smooth the target position calculated by vision, and adaptively adjust the smoothing factor according to the sound source movement speed. A soft handover transition is performed. Within the transition time window of the path mode switching, the beam weights corresponding to the old and new modes are calculated in parallel, and the instantaneous beam weights are synthesized using constant power interpolation logic based on trigonometric functions.
8. The adaptive generation system for live sound field based on multimodal stage data according to claim 6, characterized in that, The sound field generation and rendering module uses a block convolution architecture with overlapping addition to synthesize signals; In the frequency domain, for each physical channel of the multi-channel sound field generation array, parallel multiplication operations are performed on the original signal spectrum, the sound field compensation filter response, and the beamforming weights, and then restored to the time domain signal sequence through inverse fast Fourier transform.
9. The adaptive generation system for live sound field based on multimodal stage data according to claim 2, characterized in that, The system also includes a stability monitoring unit, which is used to collect feedback signals using a reference microphone, estimate the transfer function of the acoustic feedback path, and calculate the loop gain spectrum of the entire frequency band in combination with the system's positive gain. The stability monitoring unit obtains the acoustic feedback gain margin index by calculating the difference between the peak value of the closed-loop gain amplitude-frequency response and the critical oscillation condition. When the acoustic feedback gain margin index is lower than a preset threshold, the closed-loop feedback correction module inserts a notch filter at the peak frequency point or performs global gain attenuation.
10. The adaptive generation system for live sound field based on multimodal stage data according to claim 2, characterized in that, The closed-loop feedback correction module is used to obtain the measured sound pressure level spectrum of the audience area and calculate the residual between the measured sound pressure level spectrum and the predicted sound pressure level spectrum. The closed-loop feedback correction module calculates the frequency domain correction based on the residual using negative feedback update logic that includes a signal-to-noise ratio confidence function, and uses the frequency domain correction to update the output filter gain or update the material parameters in the dynamic acoustic voxel map.