Dynamic frequency domain balanced AI noise reduction conference power amplification method
By constructing an audio model and setting an AI noise reduction model, and dynamically adjusting the power amplification module, the problems of noise adaptability and amplification strategy in traditional conference systems are solved, achieving synergistic optimization of noise reduction and amplification, and improving the system's adaptability and stability.
Patent Information
- Application Number
- CN202510917446.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional conferencing systems cannot adapt to dynamically changing noise characteristics, resulting in incomplete or excessive noise reduction. Furthermore, the power amplification strategy cannot be dynamically adjusted according to real-time noise energy, leading to voice signal distortion and increased energy consumption.
An audio model for the meeting scenario is constructed, an AI noise reduction model is set up, and the activation duration and gain of the power amplification module are dynamically adjusted by calculating the total noise energy and frequency domain equalization parameters to achieve synergistic optimization of noise reduction and amplification. A power amplification effect evaluation system is also established.
It improves the comprehensiveness and accuracy of noise suppression, enhances the dynamic adaptability and stability of the system, avoids resource waste and performance degradation, and ensures a balance between audio quality and energy efficiency.
Smart Images

Figure CN120998221A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, in particular to an AI noise reduction conference power amplification method based on dynamic frequency domain equalization. BACKGROUND
[0002] From the perspective of noise suppression, the traditional conference system mostly adopts the frequency domain equalization method with fixed parameters, which is difficult to adapt to the dynamic noise characteristics in the conference scene. For example, there may be stable noise of air conditioning running, transient noise of door and window opening and closing, and transitional noise of personnel walking in the conference room. The traditional method cannot adjust the noise reduction strategy in real time for different types of noise, resulting in incomplete or excessive noise reduction, causing distortion of the voice signal. In addition, the existing noise reduction technology lacks accurate modeling of the spatial characteristics of the conference scene, and does not fully consider the energy storage capacity and noise interference factors of each node in the acoustic space, making it difficult to build accurate acoustic balance equations, resulting in low matching degree of the noise reduction algorithm and the actual scene.
[0003] In terms of power amplification control, the traditional scheme usually adopts a static power amplification strategy, which cannot dynamically adjust the amplification gain according to the real-time noise energy. When the noise energy of the conference scene fluctuates, the fixed amplification gain may cause two problems: when the noise energy is high, insufficient power amplification will cause the voice signal to be overwhelmed by noise; when the noise energy is low, excessive amplification will increase the system energy consumption and introduce distortion. At the same time, the traditional method does not combine the noise reduction performance and power amplification demand to calculate the frequency domain equalization parameters, making it difficult to realize the collaborative optimization of noise reduction and amplification, resulting in unstable audio quality.
[0004] The existing technology lacks a dynamic regulation mechanism for the power amplification module, which cannot adjust the activation duration and gain according to the real-time noise reduction demand, making it difficult to adaptively process noise changes at different times. Moreover, the traditional scheme lacks a quantitative evaluation system for the power amplification effect, which cannot accurately measure the difference between the actual amplification demand and the theoretical demand, resulting in a lack of data support for system optimization, and performance degradation after long-term use. SUMMARY
[0005] The present application aims to provide an AI noise reduction conference power amplification method based on dynamic frequency domain equalization to solve the problems raised in the background.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical scheme: an AI noise reduction conference power amplification method based on dynamic frequency domain equalization, the method comprising:
[0007] Collecting audio influence nodes of the conference scene, and constructing an audio model of the conference scene based on the audio influence nodes;
[0008] Setting an AI noise reduction model, and testing the AI noise reduction model to obtain the noise reduction performance of the AI noise reduction model;
[0009] adding the AI noise reduction model into the audio model of the conference scene;
[0010] calculating total noise energy of the conference scene in a preset time period, calculating theoretical amplification requirement of the AI noise reduction model in the preset time period, and calculating frequency domain equalization parameters of the AI noise reduction model in the preset time period in combination with noise reduction performance of the AI noise reduction model;
[0011] generating dynamic control information of a power amplification module in the AI noise reduction model based on the frequency domain equalization parameters of the AI noise reduction model in the preset time period;
[0012] dynamically controlling the power amplification module based on the dynamic control information of the power amplification module in the AI noise reduction model;
[0013] evaluating power amplification effect of the AI noise reduction model on the conference scene.
[0014] Preferably, the audio influence nodes of the conference scene are collected, and the audio model of the conference scene is constructed based on the audio influence nodes, which includes:
[0015] acquiring spatial information of the conference scene, and dividing the conference scene into a plurality of acoustic spaces;
[0016] extracting spatial nodes of each acoustic space to form the audio influence nodes of the conference scene;
[0017] calculating acoustic energy storage capacity of each audio influence node, and extracting noise interference factors acting on the audio influence nodes;
[0018] constructing an acoustic balance equation of the conference scene based on the acoustic energy storage capacity of each audio influence node and the noise interference factors.
[0019] Preferably, the acoustic balance equation of the conference scene is constructed based on the acoustic energy storage capacity of each audio influence node and the noise interference factors, which includes:
[0020] extracting interaction effects between each audio influence node and action effects of the noise interference factors on each audio influence node respectively, and constructing the acoustic balance equation of the conference scene based on the acoustic energy storage capacity of each audio influence node, the interaction effects between each audio influence node, and the action effects of the noise interference factors on each audio influence node.
[0021] Preferably, the AI noise reduction model is set, and the noise reduction performance of the AI noise reduction model is acquired by testing the AI noise reduction model, which includes:
[0022] respectively test the AI noise reduction model under the conditions of steady-state signals, transient signals and transition signals to obtain the noise reduction performance of the AI noise reduction model;
[0023] The respective testing of the AI noise reduction model under the conditions of steady-state signals, transient signals and transition signals comprises:
[0024] The continuous noise attenuation of the AI noise reduction model is tested under the condition of steady-state signals, the burst noise response delay of the AI noise reduction model is tested under the condition of transient signals, and the frequency domain switching stability of the AI noise reduction model is tested under the condition of transition signals, and the noise reduction performance of the AI noise reduction model is comprehensively obtained.
[0025] Preferably, the AI noise reduction model is added to the audio model of the conference scene, which comprises:
[0026] The audio influence nodes in the conference scene are analyzed, the audio influence node with the greatest influence degree is obtained, and the AI noise reduction model is set at the position of the audio influence node with the greatest influence degree in the audio model of the conference scene.
[0027] Preferably, the total noise energy of the conference scene in a preset period is calculated, and the theoretical amplification requirement of the AI noise reduction model in a preset period is calculated, and the frequency domain equalization parameter of the AI noise reduction model in a preset period is calculated in combination with the noise reduction performance of the AI noise reduction model, which comprises:
[0028] The energy values of the noise sources in each audio influence node in a preset period are calculated and summarized to obtain the total noise energy of the conference scene in a preset period;
[0029] The total noise energy of the conference scene in a preset period is calculated, and the theoretical amplification requirement of the AI noise reduction model in a preset period is calculated in combination with the total noise energy of the conference scene in a preset period;
[0030] The frequency domain equalization parameter of the AI noise reduction model in a preset period is obtained by combining the theoretical amplification requirement of the AI noise reduction model in a preset period with the noise reduction performance of the AI noise reduction model.
[0031] Preferably, the dynamic control information of the power amplification module in the AI noise reduction model is generated based on the frequency domain equalization parameter of the AI noise reduction model in a preset period, which comprises:
[0032] The unit amplification gain of the power amplification module in the AI noise reduction model in the activated state is tested, and the activation duration of the power amplification module is obtained in combination with the frequency domain equalization parameter of the AI noise reduction model in a preset period.
[0033] Preferably, the dynamic control information based on the power amplification module in the AI noise reduction model dynamically controls the power amplification module, including:
[0034] The preset period is divided into several discrete time periods, and the activation duration of the power amplification module is evenly distributed in the several discrete time periods.
[0035] After the end of the previous discrete time period, it is judged whether the amplification gain of the power amplification module in the discrete time period reaches the expectation.
[0036] The activation duration of the power amplification module in the next discrete time period is adaptively adjusted until the dynamic control of the power amplification module in the entire preset period is completed.
[0037] Preferably, the AI noise reduction model evaluates the power amplification effect of the conference scene, including:
[0038] Obtain the audio quality change value of the conference scene before and after power amplification.
[0039] Based on the audio quality change value, the actual amplification demand of the conference scene is obtained.
[0040] The actual amplification demand and the theoretical amplification demand are analyzed to obtain the regulation accuracy of the power amplification module in the AI noise reduction model.
[0041] Based on the regulation accuracy, the AI noise reduction model evaluates the power amplification effect of the conference scene.
[0042] Preferably, the space nodes of each acoustic space are extracted to form the audio influence nodes of the conference scene, including:
[0043] The sound field distribution of each acoustic space is scanned, and the nodes with sound field energy density higher than a preset threshold are extracted as space nodes to form the audio influence nodes of the conference scene.
[0044] Compared with the prior art, the beneficial effects of the present application are:
[0045] The present application provides an accurate scene modeling basis for dynamic noise reduction and power amplification by constructing a conference scene audio model and accurately analyzing the energy storage capacity and noise interference factors of each acoustic space node. This way of dividing acoustic spaces and extracting key nodes based on spatial information can capture the acoustic characteristics in the conference environment in detail, making the subsequent AI noise reduction model and frequency domain equalization processing more suitable for actual scene requirements, and fundamentally improving the pertinence of noise reduction and amplification.
[0046] The AI noise reduction model is set and tested for its noise reduction performance under steady-state, transient and transition signals to ensure that the model has the ability to cope with complex noise environment. By quantifying the continuous noise attenuation, burst noise response delay and frequency domain switching stability, the AI model can automatically adjust the strategy according to different noise types, effectively solve the problem of insufficient dynamic noise processing of traditional fixed parameter noise reduction, and significantly improve the comprehensiveness and accuracy of noise suppression.
[0047] The AI noise reduction model is precisely added to the node position with the greatest impact on the audio model, realizing the deep integration of the noise reduction module and the scene model. This targeted deployment allows the AI model to directly act on the areas that have the most significant impact on audio quality, maximizing the noise reduction effect while avoiding resource waste and improving system processing efficiency.
[0048] Through the linkage mechanism of total noise energy, theoretical amplification requirement and frequency domain equalization parameters, the collaborative optimization of noise reduction and power amplification is realized. According to the real-time noise energy, the activation duration and gain of the power amplification module are dynamically adjusted, which can provide sufficient amplification capacity to ensure voice clarity when the noise is strong, and can reduce energy consumption when the noise is weak, avoiding distortion caused by excessive amplification, effectively balancing audio quality and energy efficiency.
[0049] The power amplification module is dynamically controlled and adaptively adjusted in discrete periods, so that the system can respond to noise changes in real time. By evaluating whether the amplification gain meets the requirements after the end of the previous discrete period and adjusting the activation duration of the subsequent period, a closed-loop feedback control is formed to ensure that the power amplification is always in the optimal state, significantly improving the dynamic adaptability and stability of the system.
[0050] An evaluation system for power amplification effect is established, which quantifies the control accuracy by comparing the audio quality change value with the theoretical amplification requirement, providing data support for system optimization. This evaluation mechanism can continuously optimize the AI model and power amplification strategy, so that the system can maintain good performance in long-term use, avoiding the performance degradation problem caused by the lack of feedback mechanism in traditional solutions. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 The working principle diagram of the AI noise reduction conference power amplification method of dynamic frequency domain equalization according to the present application;
[0052] Figure 2 The design diagram for frequency domain equalization parameter calculation;
[0053] Figure 3 The design diagram for dynamic control of the power amplification module;
[0054] Figure 4 The design diagram for power amplification effect evaluation. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present application.
[0056] Please refer to Figures 1-4 The application relates to an AI noise reduction conference power amplification method based on dynamic frequency domain equalization, and the implementation steps are as follows:
[0057] Audio influence nodes of a conference scene are collected, and an audio model of the conference scene is constructed based on the audio influence nodes. First, spatial information of the conference scene is acquired, and the conference scene is divided into a plurality of acoustic spaces. Then, spatial nodes of each acoustic space are extracted to form audio influence nodes of the conference scene. When the spatial nodes are extracted, the sound field distribution of each acoustic space is scanned, and nodes with sound field energy density higher than a preset threshold are extracted as spatial nodes. Then, the acoustic energy storage capacity of each audio influence node is calculated, and noise interference factors acting on the audio influence nodes are extracted. Finally, the interaction effect between each audio influence node and the effect of the noise interference factors on each audio influence node are extracted, respectively, and an acoustic balance equation of the conference scene is constructed based on the acoustic energy storage capacity of each audio influence node, the interaction effect between each audio influence node, and the effect of the noise interference factors on each audio influence node.
[0058] An AI noise reduction model is set, and the AI noise reduction model is tested to obtain the noise reduction performance of the AI noise reduction model. The noise reduction of the AI noise reduction model under steady-state signals, transient signals and transition signals is tested. The continuous noise attenuation of the AI noise reduction model under steady-state signals is tested, the burst noise response delay of the AI noise reduction model under transient signals is tested, and the frequency domain switching stability of the AI noise reduction model under transition signals is tested, and the noise reduction performance of the AI noise reduction model is obtained by comprehensively testing these test results.
[0059] The AI noise reduction model is added to the audio model of the conference scene. The audio influence nodes in the conference scene are analyzed, the audio influence node with the largest influence degree is obtained, and the AI noise reduction model is set at the position of the audio influence node with the largest influence degree in the audio model of the conference scene.
[0060] The total noise energy of the conference scene in the preset period is calculated, and the theoretical amplification requirement of the AI noise reduction model in the preset period is calculated, and the frequency domain equalization parameters of the AI noise reduction model in the preset period are calculated based on the noise reduction performance of the AI noise reduction model. The energy value of the noise source in each audio influence node in the preset period is calculated and summarized to obtain the total noise energy of the conference scene in the preset period. The processing capacity of the AI noise reduction model in the preset period is obtained, and the total noise energy of the conference scene in the preset period is obtained, and the theoretical amplification requirement of the AI noise reduction model in the preset period is obtained. The frequency domain equalization parameters of the AI noise reduction model in the preset period are obtained by combining the theoretical amplification requirement of the AI noise reduction model in the preset period with the noise reduction performance of the AI noise reduction model.
[0061] The dynamic control information of the power amplification module in the AI noise reduction model is generated based on the frequency domain equalization parameters of the AI noise reduction model in the preset period. The unit amplification gain of the power amplification module in the AI noise reduction model in the active state is tested, and the active duration of the power amplification module is obtained based on the frequency domain equalization parameters of the AI noise reduction model in the preset period.
[0062] The power amplification module is dynamically controlled based on the dynamic control information of the power amplification module in the AI noise reduction model. The preset period is divided into a plurality of discrete time periods, and the active duration of the power amplification module is evenly distributed in the plurality of discrete time periods. After the end of the previous discrete time period, it is judged whether the amplification gain of the power amplification module in the discrete time period reaches the expectation. The active duration of the power amplification module in the next discrete time period is adaptively adjusted until the dynamic control of the power amplification module in the entire preset period is completed.
[0063] The power amplification effect of the AI noise reduction model on the conference scene is evaluated. The audio quality change value of the conference scene before and after power amplification is obtained, and the actual amplification requirement of the conference scene is obtained based on the audio quality change value. The actual amplification requirement and the theoretical amplification requirement are analyzed to obtain the regulation precision of the power amplification module in the AI noise reduction model, and the power amplification effect of the AI noise reduction model on the conference scene is evaluated based on the regulation precision.
[0064] Embodiment 1:
[0065] When constructing the audio model of the conference scene, the spatial information of the conference scene needs to be obtained first. Specifically, the geometric dimensions of the conference room are accurately measured by a measuring device, such as recording the length, width and height of the conference room, and the layout inside the conference room is carefully surveyed, including the positions and material properties of fixed or movable obstacles such as walls, doors and windows, tables and chairs, projectors, etc. These spatial information is the basis for subsequent division of acoustic space, because different spatial structures and materials will have different effects on the propagation and reflection of sound.
[0066] Based on the acquired spatial information, the conference scene is divided into several acoustic spaces. The division method can be based on functional areas or differences in acoustic characteristics. For example, a conference room can be divided into speaking areas, audience areas, equipment areas, and other functional areas. Each area has different acoustic environments due to differences in personnel activities, equipment distribution, and other factors. Alternatively, the space can be divided based on acoustic characteristics such as sound pressure level, reverberation time, and other parameters. Regions with similar acoustic characteristics are divided into an acoustic space. During the division process, the geometry of the space, the distribution of materials, and the characteristics of sound propagation must be considered to ensure that each acoustic space is reasonably divided and accurately reflects the acoustic characteristics of the region.
[0067] After completing the division of the acoustic space, the sound field distribution of each acoustic space is scanned. Professional acoustic measurement instruments such as sound level meters, microphone arrays, and other devices are used to measure multiple points within each acoustic space at a certain grid spacing to obtain the sound field energy density data of each measurement point. During measurement, the stability and accuracy of the measurement instrument must be maintained to avoid the influence of human factors or environmental interference on the measurement results. Through analysis of the measurement data, nodes in each acoustic space with sound field energy density higher than the preset threshold are determined. The setting of the preset threshold must be determined based on the actual needs and acoustic characteristics of the conference scene. Usually, relevant acoustic standards or preliminary experimental data can be referred to for setting. Extracting these nodes above the threshold as space nodes, these space nodes collectively constitute the audio impact nodes of the conference scene, which are key locations that significantly affect audio quality in the conference scene.
[0068] Next, the acoustic energy storage capacity of each audio impact node is calculated. The acoustic energy storage capacity is closely related to the spatial characteristics of the node's location, including the volume of the space where the node is located, the sound absorption coefficient and reflection coefficient of the material, and other factors. For example, for a node located near a wall, if the wall material is a material with good sound absorption performance, such as soundproof cotton, the acoustic energy storage capacity of the node is relatively weak. If the wall is made of hard concrete material with strong reflection performance, the energy storage capacity is relatively strong. By establishing a mathematical model, these factors are considered to calculate the amount of acoustic energy that each audio impact node can store.
[0069] At the same time, noise interference factors acting on the audio impact nodes are extracted. Noise interference factors come from a wide range of sources, including external traffic noise, indoor air conditioning, computer and other device-generated noise, and noise generated by personnel activities. For each audio impact node, the noise sources it may be exposed to are analyzed, and the impact of each noise source on the node is determined. For example, a node near a window may be mainly disturbed by external traffic noise, while a node near an air conditioner outlet is mainly disturbed by air conditioner noise. Through identification and analysis of noise sources, a basis is provided for subsequent noise processing.
[0070] After obtaining the acoustic energy storage capacity, noise interference factors of each audio impact node, and the interaction effects between nodes, an acoustic balance equation for the conference scene is constructed. The interaction effects between nodes mainly manifest in the process of sound propagation and reflection between different nodes. The sound energy of one node can propagate through the air or reflect off objects to reach other nodes, thereby affecting the acoustic energy distribution of other nodes. It is necessary to analyze the sound propagation paths and energy transfer efficiency between each audio impact node and other nodes to determine their interaction relationship. At the same time, the effects of noise interference factors on each audio impact node, i.e., the energy input of noise sources to each node, are considered. Based on these factors, an acoustic balance equation is established that can describe the sound energy distribution and changes in the conference scene. This equation combines the acoustic energy storage capacity of each audio impact node, the interaction effects between nodes, and the effects of noise interference factors, providing a theoretical basis for subsequent addition of AI noise reduction models and power amplification control.
[0071] Embodiment 2:
[0072] After setting up the AI noise reduction model, its noise reduction performance needs to be comprehensively tested, covering three different types of input scenarios: steady-state signals, transient signals, and transition signals, to ensure the model's adaptability in various audio environments.
[0073] For the test of steady-state signals, noise signals with sustained stable characteristics need to be generated, such as white noise or pink noise. The energy of such signals is uniformly distributed over a wide frequency range, which can simulate stable background noise scenarios in conference rooms, such as continuous operation of air conditioners, low-frequency noise of ventilation systems, etc. When inputting steady-state signals into the AI noise reduction model, the continuous input state of the signal needs to be maintained, and the output signal processed by the model is obtained in real time through the audio acquisition device. During the test, the intensity and frequency range of the input signal need to be accurately controlled to ensure the consistency of the test conditions. At the same time, professional audio analysis software is used to compare the input signal and the output signal in real time, focusing on the changes in noise attenuation. The noise attenuation here refers to the difference between the noise energy in the input signal and the residual noise energy in the output signal. By continuously recording the attenuation data at multiple time points, the continuous noise reduction capability of the model in a steady-state noise environment can be evaluated. For example, within a 30-minute test period, attenuation data is collected every 1 minute, and the stability and fluctuation range of the data are observed to determine whether the model can maintain effective noise reduction effects in a long-time stable noise environment.
[0074] For the test of transient signals, it is necessary to simulate the sudden, short-duration noise scenarios such as the impact sound of a conference room door closing suddenly, the click sound of a device starting suddenly, etc. The characteristics of transient signals are that the energy is concentrated in a short time and the response speed of the noise reduction model is required to be high. When testing, generate transient noise signals with specific time domain characteristics, such as setting the rise time, peak amplitude, and decay time of the signal, etc. After inputting the transient signal into the AI noise reduction model, use a high-speed audio recording device to capture the output signal of the model at a sampling rate of microseconds. By analyzing the time axis data of the input signal and the output signal, determine the time interval from the input of the transient noise signal to the start of effective noise suppression by the model, i.e. the response delay of the burst noise. In order to ensure the reliability of the test results, it is necessary to input different types of transient signals repeatedly, such as transient noise with different peak amplitudes and different frequency components, measure the response delay respectively, and calculate the average value and dispersion degree. For example, perform 20 tests of different transient signals, record the response delay time of each test, analyze the distribution of the data, and evaluate the response speed and stability of the model to various burst noises.
[0075] For the test of transient signals, it is necessary to simulate the sudden, short-duration noise scenarios such as the impact sound of a conference room door closing suddenly, the click sound of a device starting suddenly, etc. The characteristics of transient signals are that the energy is concentrated in a short time and the response speed of the noise reduction model is required to be high. When testing, generate transient noise signals with specific time domain characteristics, such as setting the rise time, peak amplitude, and decay time of the signal, etc. After inputting the transient signal into the AI noise reduction model, use a high-speed audio recording device to capture the output signal of the model at a sampling rate of microseconds. By analyzing the time axis data of the input signal and the output signal, determine the time interval from the input of the transient noise signal to the start of effective noise suppression by the model, i.e. the response delay of the burst noise. In order to ensure the reliability of the test results, it is necessary to input different types of transient signals repeatedly, such as transient noise with different peak amplitudes and different frequency components, measure the response delay respectively, and calculate the average value and dispersion degree. For example, perform 20 tests of different transient signals, record the response delay time of each test, analyze the distribution of the data, and evaluate the response speed and stability of the model to various burst noises.
[0076] Throughout the test process, it is necessary to strictly control the acoustic conditions of the test environment, ensure that the test site has good sound insulation effect, and avoid external noise interference on the test results. At the same time, the calibration and accuracy of the test equipment are also crucial, such as the sensitivity of the microphone, the frequency response range of the audio analyzer, etc. all need to be calibrated regularly to ensure the accuracy of the collected data. In addition, the running parameters of the AI noise reduction model need to be kept stable. In the test of different types of signals, except for the type of input signal, other configuration parameters of the model should be kept consistent to ensure that the test results can truly reflect the differences in noise reduction performance of the model under different signal types.
[0077] By testing the continuous noise attenuation under steady-state signals, we can understand the model's continuous processing capability in a stable noise environment; by testing the response delay of burst noise under transient signals, we can evaluate the model's rapid response capability to burst noise; by testing the frequency domain switching stability under transition signals, we can master the model's adaptability in frequency change scenarios. By integrating the test results of these three aspects, we can comprehensively and systematically obtain the noise reduction performance of the AI noise reduction model, providing key basis for subsequent addition of the model to the conference scene audio model and determination of the frequency domain equalization parameters. In the process of recording and analyzing test data, objective and scientific methods should be used to record various test data accurately, without adding any hypothetical experimental effect description, and only through data statistics and comparison to reflect the actual performance of the model.
[0078] Example 3:
[0079] When adding the AI noise reduction model to the audio model of the conference scene, it is necessary to systematically analyze the audio impact nodes in the conference scene to determine the best deployment location of the model. First, it is necessary to comprehensively sort out the attributes of each audio impact node in the conference scene. These nodes are determined through sound field distribution scanning when building the audio model in the early stage, and their distribution is closely related to the acoustic characteristics of the conference space. For example, in a rectangular conference room, audio impact nodes may be distributed near the speaking platform, different areas of the audience seat, around doors and windows, and equipment placement, etc. The sound field energy density, noise interference degree, and influence on the overall audio quality of each node are different.
[0080] To obtain the audio impact node with the greatest impact, a scientific evaluation system needs to be established. This evaluation system needs to consider multiple dimensions of factors, including the sound field energy concentration of the node's location, the intensity of noise interference, and the impact of the node on the sound propagation path. Specifically, for each audio impact node, first analyze its sound field energy density. The higher the energy density, the greater the impact of the strength and weakness of the sound signal on the overall audio quality. Second, evaluate the noise interference factors that the node is subjected to. For example, nodes near windows may be more severely disturbed by external traffic noise. If such nodes do not undergo effective noise reduction, they may have a greater impact on conference audio. In addition, the position of the node in the sound propagation network also needs to be considered. Some nodes may be on the key path of sound propagation, and their signal processing effects will directly affect the audio quality of other nodes.
[0081] Taking a specific conference room scenario as an example, assume that the conference room contains 10 audio impact nodes, labeled N1 to N10. Among them, N1 is located directly in front of the podium, N2 to N4 are distributed in the front row of the audience, N5 to N7 are in the middle and back rows of the audience, N8 and N9 are near the two side walls, and N10 is near the air conditioner outlet. During the analysis process, the sound field energy density values of each node are obtained through preliminary measurement. For example, the energy density of N1 is 85 dB, N2 is 78 dB, N3 is 75 dB, and N10 is near a noise source, with a noise energy ratio of 60%. At the same time, through acoustic simulation analysis of the sound propagation path between nodes, it is found that N1 is the main starting node for the sound propagation of the speaker to the audience, and its signal quality directly affects the audio input of subsequent nodes such as N2 to N7. Although N10 has a low energy density, it is continuously disturbed by air conditioner noise, and its noise signal will affect nodes such as N7 and N8 through air propagation.
[0082] In the comprehensive evaluation, a weight coefficient is set for each influencing factor. For example, the weight of sound field energy density is set to 40%, the weight of noise interference intensity is set to 35%, and the weight of propagation path importance is set to 25%. For N1, the sound field energy density score is 85, the noise interference intensity score is 60 (since it is in the speaking area, it mainly receives the sound of the speaker, and the noise interference is relatively small), the propagation path importance score is 90, and the comprehensive score is 85 x 40% + 60 x 35% + 90 x 25% = 79.5. The sound field energy density score of N10 is 65, the noise interference intensity score is 90, the propagation path importance score is 50, and the comprehensive score is 65 x 40% + 90 x 35% + 50 x 25% = 72. Through similar calculations, it is determined that N1 has the highest comprehensive score, i.e., it is the audio impact node with the greatest impact.
[0083] After determining the node with the greatest impact, the node needs to be mapped to the audio model of the conference scene. The audio model of the conference scene is described by the acoustic balance equation constructed in advance, where each audio impact node has a corresponding parameter representation in the model, such as acoustic energy storage capacity, interaction coefficient with other nodes, etc. The node with the greatest impact is usually represented by a larger acoustic energy storage capacity parameter or a larger absolute value of the corresponding row or column element in the interaction coefficient matrix in the model. For example, in the matrix representation of the acoustic balance equation, the row vector corresponding to N1 may contain multiple large interaction coefficients, indicating that it has a significant impact on the energy transmission of other nodes.
[0084] Next, the AI noise reduction model is set at the position where the node with the greatest impact is mapped in the audio model. In actual operation, this means that in the algorithm implementation of the audio model, the processing logic of the AI noise reduction model is embedded into the parameter calculation process corresponding to the node. For example, when performing iterative calculation based on the acoustic balance equation, for the energy calculation step of the node, the AI noise reduction model is first introduced to perform noise reduction processing on the signal input to the node, and then subsequent energy storage and transmission calculation is performed. In hardware implementation, if the conference system adopts a distributed audio processing architecture, it may be necessary to deploy AI noise reduction module hardware devices near the physical location corresponding to the node, such as installing an audio processor with AI noise reduction function near the podium, so that it can directly perform real-time processing on the audio signal of the node.
[0085] During the setting of the AI noise reduction model, compatibility between the model and the audio model needs to be ensured. For example, the input and output signal format of the AI noise reduction model needs to be consistent with the signal representation method in the audio model, the processing delay of the model needs to be controlled within a reasonable range to avoid significant lag in overall audio transmission. At the same time, preliminary debugging of the parameters of the model is also needed to make it work best in the specific acoustic environment of the node. For example, according to the noise characteristics of the node, adjust the filter coefficients, neural network weights and other parameters of the AI noise reduction model to more specifically suppress the main noise source of the node.
[0086] Embodiment 4:
[0087] When calculating the total noise energy of the conference scene in the preset time period, and combining the processing capacity of the AI noise reduction model to obtain the theoretical amplification requirement and frequency domain equalization parameters, a specific conference scene needs to be taken as an example to carry out detailed implementation. Assuming that a medium-sized conference room is 10 meters long, 8 meters wide, and 3.5 meters high, with a podium, 10 rows of audience seats, and air conditioning, projector and other equipment inside, the preset time period is set to 10 minutes, and the specific implementation process is described taking this scene as an example.
[0088] The total noise energy of the conference scenario in the preset period is calculated. The audio influence nodes of the conference room include nodes N1 near the speaking platform, nodes N2-N4 in the front row of the audience seats, nodes N5-N7 in the middle and back rows, node N8 near the air conditioner outlet, node N9 near the projector equipment, and nodes N10-N12 near the doors and windows. In the preset 10-minute period, the energy values of the noise sources in each audio influence node are collected one by one.
[0089] Taking node N8 (air conditioner outlet) as an example, its noise source is low-frequency noise generated by the operation of the air conditioner fan. The noise energy value of this node is recorded by the noise sensor every 10 seconds, and a total of 60 groups of data are obtained in 10 minutes, with each group of data corresponding to an energy value (unit: decibel) at a time point. These data are summarized by accumulating the energy values at each time point and taking the average value to obtain the average noise energy value of this node in the preset period. Similarly, for node N9 (projector), its noise source is fan rotating noise, and the same collection method is used to record the noise energy data of this node in 10 minutes and calculate the average value. For nodes N10-N12 near the doors and windows, their noise sources may include external traffic noise, personnel walking noise, etc., and the noise energy data of each node in the preset period needs to be collected.
[0090] When summarizing the noise energy values of all audio influence nodes, it needs to be noted that the noise energies of different nodes may have superposition effects. For example, the air conditioner noise of node N8 and the projector noise of node N9 will propagate to the audience seat nodes N5-N7 at the same time, so when calculating the total noise energy, the energy values of each node cannot be simply added together, but the principle of sound superposition needs to be considered. Specifically, for each time point, the noise energy values of all nodes are converted into sound pressure levels and then superimposed, and the superposition results of all time points in 10 minutes are summarized to finally obtain the total noise energy of the conference scenario in the preset period. Assuming that the total noise energy of the conference room in 10 minutes is calculated by the above method, it is equivalent to the total noise energy of a continuous 80 decibel noise.
[0091] The AI noise reduction model's own processing capability in the preset period is obtained. The AI noise reduction model's own processing capability is related to its algorithm architecture, hardware performance, and other factors. For example, a certain AI noise reduction model uses a deep learning architecture and contains multiple layers of neural networks. Its processing capability can be measured by the range of noise energy it can process per unit time. In this conference room scenario, the model can process noise energy equivalent to 70 decibels per minute under ideal conditions (such as a single noise source and stable input signals). Therefore, in a 10-minute preset period, the model can theoretically process a total of 70 decibels x 10 minutes of noise energy. However, in actual applications, the noise in the conference scenario is often multi-source and complex, and the model's processing capability will be affected to some extent, so its processing capability needs to be corrected based on the test data obtained in the early stage. Assuming that through early testing in a similar conference scenario, it is determined that the model's actual processing capability in a multi-source noise environment is 80% of the ideal state, then the model's own processing capability in the preset period is 70 decibels x 10 minutes x 80%.
[0092] By combining the total noise energy of the conference scenario and the AI noise reduction model's own processing capability, the theoretical amplification requirement of the model in the preset period can be calculated. The essence of the theoretical amplification requirement is the energy gain required by the model to effectively suppress the noise in the conference scenario so that its output audio signal meets certain quality standards. For example, if the total noise energy of the conference scenario is 80 decibels x 10 minutes, and the model's own processing capability in the preset period is 56 decibels x 10 minutes (i.e., 70 x 10 x 80%), then in order to process the remaining noise energy (80-56 = 24 decibels x 10 minutes), the model needs a certain amount of energy amplification to enhance the noise reduction effect. The amount of energy amplification is the theoretical amplification requirement. When calculating, the remaining noise energy needs to be converted into the corresponding amplification gain value based on the conversion relationship between noise energy and amplification gain.
[0093] After obtaining the theoretical amplification requirement, the frequency domain equalization parameters are calculated in combination with the noise reduction performance of the AI noise reduction model. The noise reduction performance of the AI noise reduction model is obtained through early testing under steady-state, transient, and transitional signals. For example, the model has strong noise reduction capability in the low frequency band (20-200 Hz), with an attenuation of 25 decibels, while in the medium and high frequency band (200-5000 Hz), the attenuation is 15-20 decibels, and in the high frequency band (above 5000 Hz), the attenuation is 10-15 decibels. The theoretical amplification requirement needs to be allocated according to the noise reduction performance of different frequency bands to achieve dynamic frequency domain equalization.
[0094] For example, assuming that the theoretical amplification requirement is a total gain of 10 decibels, the 10 decibels of gain needs to be allocated according to the noise reduction requirements of different frequency bands. Since low-frequency noise (such as air conditioner noise) has a greater impact on conference audio, and the model has better noise reduction performance in this frequency band, 4 decibels of gain can be allocated to enhance the noise reduction capability of the low-frequency band; 3 decibels of gain is allocated to the medium-high frequency band (such as background noise of people talking); 3 decibels of gain is allocated to the high frequency band (such as high-frequency noise of equipment). In this way, the frequency domain equalization parameters of each frequency band are obtained, which are used to guide the gain adjustment of the power amplification module in different frequency bands.
[0095] In the implementation process, the noise frequency characteristics of different audio influence nodes in the conference scene also need to be considered. For example, the noise of the air outlet node N8 of the air conditioner is mainly concentrated in the low frequency band, the noise of the projector node N9 is mainly concentrated in the medium-high frequency band, and the external traffic noise may be distributed in the medium-high frequency band and the high frequency band. Therefore, when calculating the frequency domain equalization parameters, the allocation of the theoretical amplification requirement in different frequency bands needs to be further optimized according to the noise frequency distribution of each node. For example, for the low-frequency noise of node N8, the gain allocation of the low frequency band can be appropriately increased to more effectively suppress the noise of this node; for the medium-high frequency noise of node N9, the gain allocation of the medium-high frequency band is correspondingly increased.
[0096] In addition, the distribution of useful components (such as the speaker's voice) in the audio signal in different frequency bands also needs to be considered to avoid excessive attenuation of useful signals caused by the setting of frequency domain equalization parameters. For example, the speaker's voice is mainly distributed in the medium frequency band (200-3000Hz), so when allocating the gain of the medium frequency band, a balance needs to be struck between noise suppression and useful signal preservation to ensure that the adjustment of the gain does not significantly affect the quality of the speaker's voice.
[0097] Example 5:
[0098] When dynamically controlling the power amplification module in the AI noise reduction model, a medium-sized conference room in a certain enterprise is taken as an example for specific implementation. The conference room is 12 meters long and 9 meters wide, has a speaking platform, 8 rows of audience seats, is equipped with air conditioners, projectors and other equipment, has a preset time period of 15 minutes, and the power amplification module adopts a digital power amplifier architecture and is integrated in the AI noise reduction processing host.
[0099] The preset 15-minute period is divided into several discrete time segments. Considering the real-time processing requirements of the audio signal and the response speed of the power amplification module, the 15 minutes is divided into 30 discrete time segments, each with a duration of 30 seconds. This division can ensure fine control of the power amplification module without increasing computational complexity due to too short time segments. After division, the power amplification module activation time based on the frequency domain equalization parameters needs to be evenly distributed to each discrete time segment. Assuming that through preliminary calculation, the total activation time of the power amplification module in the 15-minute preset period is 9 minutes, then the average activation time of each discrete time segment is 9 minutes x 60 seconds / 30 = 18 seconds.
[0100] In the first discrete time segment (0-30 seconds), the power amplification module is controlled according to the average activation time of 18 seconds. Specifically, at the beginning of this time segment, the power amplification module is started and placed in the active state, and after 18 seconds, it is turned off, leaving the remaining 12 seconds in the inactive state. During the module activation period, the output audio signal is collected in real time, and the amplification gain data in this time segment is recorded by the audio monitoring device. For example, using a high-precision audio sampling device, the output signal is sampled at a sampling rate of 44.1 kHz to obtain the amplification gain curve in this time segment, focusing on the average value and fluctuation range of the gain.
[0101] When the first discrete time segment ends, the amplification gain of the power amplification module in that time segment is judged. The basis for judgment is the preset expected amplification gain value, which is determined by preliminary theoretical calculation and model testing. For example, the expected amplification gain is 12 dB, and by analyzing the gain data collected in this time segment, the actual average gain is calculated to be 11.5 dB, with a deviation of 0.5 dB from the expected gain. At this time, the activation time of the power amplification module in the next discrete time segment (30-60 seconds) needs to be adjusted adaptively. Since the actual gain does not reach the expected value, it is considered appropriate to increase the activation time to increase the amplification gain. The adjustment range is determined according to the size of the deviation and the gain characteristics of the module. Assuming that according to the gain-activation time curve of the module, the average gain can be increased by 0.2 dB for every 1 second increase in activation time, then to make up for the 0.5 dB deviation, the activation time of the next time segment is adjusted to 18 seconds + 3 seconds = 21 seconds.
[0102] In the second discrete time period (30-60 seconds), the power amplification module is controlled according to the adjusted activation duration of 21 seconds, i.e., activated for 21 seconds and turned off for 9 seconds. Similarly, after the end of this time period, the output signal is collected and the actual amplification gain is calculated. Assuming that the actual average gain this time is 12.2 dB, which exceeds the expected gain by 0.2 dB, the activation duration of the third discrete time period (60-90 seconds) needs to be adjusted inversely, i.e., reduced to lower the gain. According to the gain characteristics, the activation duration is adjusted to 21 seconds - 2 seconds = 19 seconds.
[0103] In this way, the amplification gain after the end of each discrete time period is judged in real time, and the activation duration of the next time period is adjusted adaptively according to the judgment result, gradually optimizing the working state of the power amplification module. During the adjustment process, the following points need to be noted: first, the adjustment amplitude should not be too large to avoid sudden changes in the audio signal caused by drastic changes in the activation duration, affecting the quality of the conference audio; second, the heat loss of the power amplification module needs to be considered, and the adjustment of the activation duration should not exceed the safe working range of the module to avoid overheating caused by long-term high-load operation; third, the actual audio signal changes in the conference scene need to be considered, for example, when the speaker's voice suddenly increases, the activation duration may need to be temporarily increased to ensure sufficient amplification gain, and when the voice decreases, the activation duration should be reduced accordingly.
[0104] Taking this conference room as an example, at the 10th discrete time period (4 minutes 30 seconds-5 minutes), the speaker is near the door and window, and the external traffic noise suddenly increases, causing the noise energy in this time period to increase. Through real-time monitoring, it is found that the actual amplification gain in this time period reaches the expected 12 dB, but the noise in the output audio is still noticeable. After analysis, it is found that the sudden increase in noise energy makes the original activation duration insufficient to effectively suppress the added noise. Therefore, in the 11th discrete time period (5 minutes-5 minutes 30 seconds), according to the real-time noise energy detection result, the activation duration is increased by 5 seconds based on the original adjustment value to enhance the gain of the power amplification module and effectively suppress the sudden noise.
[0105] During the entire 15-minute preset period, by continuously judging and adjusting the activation duration of each discrete time period, the amplification gain of the power amplification module is always maintained at a level close to the expected level, achieving dynamic control of the power amplification module. This dynamic control method can adapt to the real-time changes of noise energy in the conference scene, ensuring that the power amplification module can provide appropriate amplification gain in different noise environments, thereby effectively suppressing noise while ensuring clear transmission of the speaker's voice.
[0106] In the implementation process, a perfect feedback mechanism also needs to be established. The feedback mechanism includes real-time monitoring of the amplification gain, real-time detection of the noise energy, and subjective evaluation of the audio quality, etc. For example, in addition to monitoring the amplification gain through instruments, professional audio engineers can also be arranged to listen to the conference audio in real time, evaluate the control effect of the power amplification module according to the subjective listening feeling, and feed back the evaluation results to the control algorithm to further optimize the adjustment strategy of the activation duration.
[0107] In addition, the cooperative work between the AI noise reduction model and the power amplification module also needs to be considered. For example, when the AI noise reduction model detects that the noise energy of a certain frequency band suddenly increases, the frequency domain equalization parameters will be adjusted in time, and at this time, the dynamic control information of the power amplification module also needs to be adjusted accordingly to cooperate with the work of the noise reduction model, so as to achieve more efficient noise suppression and power amplification.
[0108] Through the detailed implementation of the above specific examples, from the division of discrete time periods, the initial allocation of the activation duration, to the gain judgment and duration adjustment of each time period, to the flexible handling of sudden situations and the establishment of the feedback mechanism, a complete power amplification module dynamic control process is formed. Based on the needs of the actual conference scene, through real-time monitoring and adaptive adjustment, the power amplification module can accurately meet the amplification needs of the conference audio and improve the quality of the conference audio.
[0109] It should be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment.
[0110] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A dynamic frequency domain equalization method for AI noise reduction conference power amplification, characterized in that, The method includes: Collect audio impact nodes in the meeting scene, and construct an audio model of the meeting scene based on the audio impact nodes; Set up an AI noise reduction model and test the AI noise reduction model to obtain its noise reduction performance; Add the AI noise reduction model to the audio model of the meeting scenario; Calculate the total noise energy of the meeting scene in a preset time period, and calculate the theoretical amplification requirement of the AI noise reduction model in the preset time period. Combine the noise reduction performance of the AI noise reduction model to calculate the frequency domain equalization parameters of the AI noise reduction model in the preset time period. Based on the frequency domain equalization parameters of the AI noise reduction model in a preset time period, dynamic control information of the power amplification module in the AI noise reduction model is generated. The power amplification module is dynamically controlled based on the dynamic control information of the power amplification module in the AI noise reduction model. The power amplification effect of the AI noise reduction model on the meeting scenario is evaluated.
2. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 1, characterized in that, The process of collecting audio impact nodes in the meeting scene and constructing an audio model of the meeting scene based on the audio impact nodes includes: Acquire spatial information of the meeting scene and divide the meeting scene into several acoustic spaces; Extract the spatial nodes of each acoustic space to form the audio impact nodes of the meeting scene; Calculate the acoustic energy storage capacity of each audio-affecting node and extract the noise interference factors acting on the audio-affecting node; The acoustic balance equation for the meeting scenario is constructed based on the acoustic energy storage capacity of each audio-affecting node and the noise interference factors.
3. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 2, characterized in that, The acoustic balance equation for the conference scenario, constructed based on the acoustic energy storage capacity of each audio-affecting node and the noise interference factors, includes: The interaction effects between each audio-affecting node and the effect of the noise interference factor on each audio-affecting node are extracted respectively. Based on the acoustic energy storage capacity of each audio-affecting node, the interaction effects between each audio-affecting node, and the effect of the noise interference factor on each audio-affecting node, the acoustic balance equation of the conference scene is constructed.
4. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 1, characterized in that, The process of setting up an AI noise reduction model and testing the AI noise reduction model to obtain its noise reduction performance includes: The denoising performance of the AI denoising model was tested under steady-state, transient, and transitional signals to obtain the denoising performance of the AI denoising model. The tests performed on the AI noise reduction model under steady-state, transient, and transitional signals respectively include: The continuous noise attenuation of the AI noise reduction model is tested under steady-state signals, the sudden noise response delay of the AI noise reduction model is tested under transient signals, and the frequency domain switching stability of the AI noise reduction model is tested under transitional signals. The noise reduction performance of the AI noise reduction model is obtained by combining the results.
5. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 1, characterized in that, Adding the AI noise reduction model to the audio model of the meeting scenario includes: Analyze the audio influencing nodes in the meeting scenario, identify the audio influencing node with the greatest influence, and set the AI noise reduction model at the position of the audio influencing node with the greatest influence in the audio model of the meeting scenario.
6. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 1, characterized in that, The calculation of the total noise energy of the meeting scenario within a preset time period, and the calculation of the theoretical amplification requirement of the AI noise reduction model within the preset time period, combined with the noise reduction performance of the AI noise reduction model, calculates the frequency domain equalization parameters of the AI noise reduction model within the preset time period, including: Calculate and summarize the energy value of noise sources in each audio-affected node during a preset time period to obtain the total noise energy of the meeting scene during the preset time period; The processing power of the AI noise reduction model within a preset time period is obtained, and combined with the total noise energy of the meeting scenario within the preset time period, the theoretical amplification requirement of the AI noise reduction model within the preset time period is obtained. By combining the theoretical amplification requirements of the AI noise reduction model within a preset time period with the noise reduction performance of the AI noise reduction model, the frequency domain equalization parameters of the AI noise reduction model within the preset time period are obtained.
7. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 1, characterized in that, The process of generating dynamic control information for the power amplification module in the AI noise reduction model based on the frequency domain equalization parameters of the AI noise reduction model during a preset time period includes: The unity amplification gain of the power amplification module in the AI noise reduction model in the active state is tested, and the activation duration of the power amplification module is obtained by combining the frequency domain equalization parameters of the AI noise reduction model in a preset time period.
8. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 7, characterized in that, The dynamic control of the power amplification module based on the dynamic control information of the power amplification module in the AI noise reduction model includes: The preset time period is divided into several discrete time periods, and the activation duration of the power amplifier module is evenly distributed among the several discrete time periods. After the previous discrete time period ends, determine whether the amplification gain of the power amplifier module in that discrete time period has reached the expected level. The activation duration of the power amplifier module in the next discrete time period is adaptively adjusted until dynamic control of the power amplifier module for the entire preset time period is completed.
9. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 1, characterized in that, The evaluation of the power amplification effect of the AI noise reduction model on the meeting scene includes: Obtain the audio quality change values of the conference scene before and after power amplification; The actual amplification requirements for the meeting scenario are obtained based on the audio quality change values. By analyzing the actual amplification requirements and the theoretical amplification requirements, the control accuracy of the power amplification module in the AI noise reduction model is obtained. The power amplification effect of the AI noise reduction model on the meeting scene is evaluated based on the aforementioned control accuracy.
10. The AI noise reduction conference power amplification method with dynamic frequency domain equalization as described in claim 2, characterized in that, The step of extracting spatial nodes from each acoustic space to form audio impact nodes for the conference scene includes: A sound field distribution scan is performed on each acoustic space, and nodes with sound field energy density higher than a preset threshold are extracted as spatial nodes to form the audio impact nodes of the meeting scene.