Building indoor acoustic noise reduction and sound absorption optimization system based on deep learning

By constructing a closed-loop optimization system and utilizing a high-precision sensor array and meta-learning algorithm, the problem of insufficient generalization ability of deep learning acoustic optimization system in different building environments was solved. This enabled fast and accurate acoustic environment adaptive optimization, reduced cost and time consumption, and improved the robustness and optimization effect of the system.

CN121747544APending Publication Date: 2026-03-27CHONGQING IND POLYTECHNIC COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing deep learning-based acoustic optimization systems suffer from insufficient model generalization ability due to differences in building environments, high retraining costs, and low deployment efficiency, making it difficult to meet the needs of modern buildings for rapid deployment and adaptive optimization of acoustic systems.

Method used

A closed-loop optimization system is constructed, comprising a sound field perception and data acquisition module, a cross-environment acoustic feature extraction module, a meta-learning-driven model rapid adaptation module, a multi-objective acoustic optimization decision module, and a distributed sound-absorbing actuator control module. Rapid adaptation and accurate optimization are achieved through a high-precision sensor array, a deep convolutional neural network, and a meta-learning algorithm.

Benefits of technology

Without requiring large-scale data re-collection and complete retraining, it achieves rapid and accurate adaptive optimization for different building interior acoustic environments, reduces data acquisition and computing resource consumption, improves the robustness of the system and the accuracy and consistency of the optimization strategy, and meets the rapid deployment needs of modern buildings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747544A_ABST
    Figure CN121747544A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, and particularly discloses a building indoor acoustic noise reduction and sound absorption optimization system based on deep learning. The system comprises a sound field perception and data acquisition module, a cross-environment acoustic feature extraction module, a meta-learning driving model rapid adaptation module, a multi-target acoustic optimization decision module and a distributed sound absorption actuator control module, rapid model adaptation is realized through meta-learning, parameters are finely adjusted by using a small amount of field data, and the sound absorption performance is improved. And in combination with a multi-objective optimization algorithm, acoustic indexes such as voice definition, noise suppression and space sense maintenance are balanced, and finally, the sound absorption unit is dynamically regulated and controlled through the distributed actuator, so that efficient and accurate indoor acoustic environment optimization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech recognition technology, specifically relating to a deep learning-based building interior acoustic noise reduction and sound absorption optimization system. Background Technology

[0002] Optimizing the acoustic environment of building interiors is an important research direction in the fields of building physics and audio engineering. Its core objective is to improve indoor sound quality, enhance speech intelligibility, and reduce noise interference through active or passive means. Among these, acoustic noise reduction and sound absorption treatment are key technical pathways to achieve the above objectives.

[0003] Acoustic optimization methods based on deep learning have become a research hotspot in recent years. These methods learn nonlinear characteristics in complex acoustic environments through a data-driven approach, thereby achieving intelligent noise suppression and precise control of sound absorption performance.

[0004] Current technologies generally employ deep learning models trained on fixed datasets for acoustic optimization. However, these models suffer from a significant decline in generalization ability when faced with varying impulse response characteristics due to differences in room geometry, interface materials, and spatial layout in different architectural environments. To adapt to new acoustic scenarios, existing solutions typically require the re-collection of hundreds of hours of on-site audio data and the execution of a complete model retraining process. This process is not only time-consuming and costly but also heavily reliant on specialized measurement equipment and manual intervention, failing to meet the urgent needs of modern buildings for rapid deployment and adaptive optimization of acoustic systems. Therefore, developing an intelligent acoustic optimization system that can overcome cross-environment generalization bottlenecks and reduce data dependence has become a pressing technical challenge in this field. Summary of the Invention

[0005] The technical problem this invention aims to solve is to overcome the shortcomings of existing deep learning-based acoustic optimization systems, such as insufficient model generalization ability due to differences in building environments, high retraining costs, and low deployment efficiency. This invention provides a deep learning-based building interior acoustic noise reduction and sound absorption optimization system. This system can achieve rapid and accurate adaptive optimization for different building interior acoustic environments without the need for large-scale data re-collection and complete retraining.

[0006] To achieve the above objectives, the technical solution adopted by this invention is to construct a complete system-level solution. This system includes a sound field perception and data acquisition module, a cross-environment acoustic feature extraction module, a meta-learning-driven model rapid adaptation module, a multi-objective acoustic optimization decision-making module, and a distributed sound-absorbing actuator control module. These modules work collaboratively through preset data interfaces and communication protocols, forming a closed-loop optimization system from environmental perception to decision execution.

[0007] The sound field perception and data acquisition module is responsible for acquiring real-time acoustic environment data within the building. This module deploys a high-precision sound pressure sensor array with a density of at least one sensor node per 10 square meters. Each sensor node integrates a condenser microphone with a frequency response range of 20 Hz to 20,000 Hz and is equipped with a 24-bit analog-to-digital converter to acquire raw sound pressure signals. Furthermore, the module includes an inertial measurement unit for synchronously recording changes in the spatial position of the sound source and the sensors. The sound field perception and data acquisition module transmits the acquired multi-channel time-domain sound pressure signals and their corresponding spatiotemporal coordinate information to the cross-environment acoustic feature extraction module via wired or wireless communication networks.

[0008] The cross-environment acoustic feature extraction module receives raw acoustic data from the sound field perception and data acquisition module and performs the core feature abstraction process. This module first preprocesses the raw sound pressure signal, including windowing using a Hanning window and resampling at a sampling rate of 48,000 Hz. Subsequently, the module calculates the short-time Fourier transform of each channel signal, converting the time-domain signal into a time-frequency domain representation. The core of feature extraction lies in constructing a generalized room impulse response feature encoder. This encoder is based on a deep convolutional neural network architecture, with its input being a multi-channel time-frequency map and its output being a 512-dimensional acoustic scene embedding vector. This embedding vector condenses key acoustic properties such as the current room's geometry, the sound absorption coefficient of interface materials, and the main reflection paths. Its calculation process is independent of specific noise types or speech content, thus achieving a unified representation of different acoustic environments.

[0009] The meta-learning-driven model rapid adaptation module is the core of the system's cross-environment generalization. This module pre-loads a deep neural network model optimized by a meta-learning training strategy as the basic acoustic optimizer. The meta-learning training process is completed offline, aiming to teach the model how to quickly adjust its internal parameters based on a small number of new environmental samples. Specifically, the network weights of the basic acoustic optimizer are initialized to meta-states with rapid adaptability. When the system is deployed to a new building environment, this module receives acoustic scene embedding vectors from the cross-environment acoustic feature extraction module, along with a small number of noise-target acoustic spectrum pairs collected on-site. The module performs an internal loop optimization process, which fine-tunes specific layer parameters of the basic acoustic optimizer using only 2 to 5 gradient descent steps, requiring less than 30 minutes of dual-channel audio recording. The fine-tuned model then possesses a high-precision modeling capability for the acoustic characteristics of the new environment, thereby generating noise reduction filter coefficients and impedance adjustment suggestions for sound-absorbing materials tailored to that environment.

[0010] The multi-objective acoustic optimization decision-making module is responsible for comprehensively balancing multiple, sometimes conflicting, acoustic performance metrics. This module receives initial optimization suggestions from the meta-learning-driven model's rapid adaptation module and makes refined decisions based on a set of preset optimization objective functions. The optimization objectives include maximizing the speech transmission index to at least 0.75, minimizing the background noise level to below 35 dB A-weighted, and maintaining the spatial perception parameter within the ideal range of 0.5 to 0.8. The decision-making process is implemented through a weighted summation multi-objective optimization algorithm, which dynamically adjusts the weight coefficients of each objective function based on real-time calculated acoustic comfort metrics and the user's preset preference profile. Finally, the module outputs a set of globally optimal acoustic processing parameters, including the required noise attenuation for each frequency band, the ideal reverberation time distribution, and the preferred reflection suppression path.

[0011] The distributed sound-absorbing actuator control module transforms the abstract parameters output by the multi-objective acoustic optimization decision module into specific physical control commands. This module connects and controls multiple active sound-absorbing units and tunable acoustic panels deployed indoors. Each active sound-absorbing unit integrates a speaker array and a feedback controller, which generates an anti-noise field with the opposite phase to the incident noise based on the received control commands. The tunable acoustic panels are driven by miniature servo motors, dynamically adjusting their surface perforation rate within the range of 5% to 30%, thereby altering their Helmholtz resonant frequency to achieve targeted absorption of noise in specific frequency bands. The control module forms a closed-loop control circuit by real-time monitoring of the actuator status and residual noise level, ensuring the continuous and stable acoustic optimization effect.

[0012] In a preferred embodiment of the present invention, the sensor array in the sound field sensing and data acquisition module adopts a self-organizing network topology. When some sensor nodes fail due to faults or occlusion, the remaining nodes can dynamically reconstruct the sensing network through negotiation and recover complete sound field information from sparse sampled data using a compressed sensing algorithm, ensuring robust operation of the system under non-ideal conditions.

[0013] Furthermore, the deep convolutional neural network architecture employed in the cross-environment acoustic feature extraction module comprises eight convolutional layers and three fully connected layers. The convolutional layers use 3x3 convolutional kernels and max-pooling operations with a stride of 2 to extract abstract representations of acoustic features layer by layer. The fully connected layers prevent overfitting through a dropout technique with a dropout rate of 0.5, ultimately outputting an acoustic scene embedding vector of dimension 512. The cosine similarity of this vector is used to quantify the differences between different acoustic environments.

[0014] In another preferred embodiment of the present invention, the meta-learning-driven model rapid adaptation module employs a model-independent meta-learning algorithm as its core learning paradigm. During the meta-training phase, the model is trained on a diverse task set covering 100 different room acoustic configurations. Each task simulates a specific acoustic environment adaptation problem. The model optimizes its initial parameters in the outer loop, enabling it to rapidly adapt to new tasks with only a small number of samples and iterations in the inner loop. This process significantly improves the generalization ability of the base acoustic optimizer to unseen architectural environments.

[0015] Furthermore, the multi-objective acoustic optimization decision-making module integrates a multi-objective solver based on a non-dominated sorting genetic algorithm. This solver maintains a population of 100 candidate solutions and performs iterative evolution through selection, crossover, and mutation operations, ultimately approximating the Pareto optimal solution set. Decision-makers can select the most suitable acoustic processing strategy from the solution set according to actual needs, achieving a flexible trade-off between optimization objectives.

[0016] The distributed acoustic actuator control module also features energy management. This function monitors the power consumption of each actuator in real time and, provided that the acoustic performance meets preset thresholds, prioritizes scheduling actuator combinations with high energy efficiency ratios, keeping the total system power consumption below 80% of the rated power, thus achieving a balance between acoustic optimization and energy consumption.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs a fundamental acoustic optimizer with inherent rapid adaptability by introducing a meta-learning-driven model rapid adaptation module. This design enables the system to achieve excellent optimization performance when facing new architectural acoustic environments, without having to undergo a time-consuming and costly complete model retraining process. It only requires fine-tuning the gradient step size using less than 30 minutes of on-site audio data. This reduces system deployment and adaptation time from weeks in traditional methods to hours, significantly reducing data acquisition and computational resource consumption, and solving the core generalization bottleneck of deep learning models in cross-environment applications.

[0018] 2. The cross-environment acoustic feature extraction module designed in this invention can abstract embedding vectors from the original acoustic signal that represent the essential acoustic properties of the room, independent of the specific sound source content. This feature representation method unifies complex acoustic environmental differences into a quantifiable low-dimensional space, providing a stable and universal input benchmark for subsequent rapid model adaptation. This module effectively decouples environmental characteristics from noise characteristics, enhances the system's robustness to different noise types and speaker variations, and improves the accuracy and consistency of the optimization strategy.

[0019] 3. This invention achieves comprehensive balance and synergistic optimization of multiple key acoustic indicators, such as speech intelligibility, noise suppression, and spatial awareness, through a multi-objective acoustic optimization decision module. This module employs an advanced evolutionary multi-objective optimization algorithm that automatically searches for Pareto optimal solutions that satisfy various performance constraints, avoiding potential acoustic quality degradation issues that might result from single-objective optimization. This enables the system to provide customized, globally optimal acoustic solutions based on the specific functional requirements of the architectural space, significantly improving the overall comfort and practicality of the indoor acoustic environment. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall technical solution architecture of the deep learning-based building interior acoustic noise reduction and sound absorption optimization system proposed in this invention. Figure 2 This is a schematic diagram of the core principle framework of the meta-learning-driven model rapid adaptation module in this invention; Figure 3 This is a logical flowchart of the cross-environment acoustic feature extraction module in this invention; Figure 4 This is a flowchart of the optimization decision-making process of the multi-objective acoustic optimization decision-making module in this invention; Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow between the sound field perception and data acquisition module and the distributed sound-absorbing actuator control module in this invention. Detailed Implementation

[0021] Example 1: Please refer to the appendix Figure 1 This embodiment details the technical implementation of a deep learning-based building interior acoustic noise reduction and sound absorption optimization system. The system aims to address the problem of insufficient model generalization ability caused by differences in building environments, achieving rapid and accurate adaptive optimization for different building interior acoustic environments without the need for large-scale data re-collection and complete retraining. The system constitutes a closed-loop optimization system from environmental perception to decision execution. Its core modules include a sound field perception and data acquisition module, a cross-environment acoustic feature extraction module, a meta-learning-driven model rapid adaptation module, a multi-objective acoustic optimization decision module, and a distributed sound-absorbing actuator control module. These modules work collaboratively through a pre-set high-speed data bus and a network following a real-time transmission protocol, ensuring low latency and high reliability of data flow.

[0022] The sound field perception and data acquisition module is the front end of the system for perceiving the physical acoustic environment. This module deploys a high-precision sound pressure sensor array in the interior space of the target building. The array's deployment density strictly adheres to the principle of having at least one sensor node per 10 square meters.

[0023] Each sensing node is a highly integrated intelligent sensing unit, its core component being a condenser microphone with a frequency response range covering 20 Hz to 20,000 Hz. This microphone connects to a 24-bit high-precision analog-to-digital converter, continuously acquiring raw sound pressure signals at a sampling rate of 48,000 Hz to ensure the capture of full-band acoustic information. The sensing node also integrates a six-axis inertial measurement unit, containing a three-axis accelerometer and a three-axis gyroscope, used to synchronously and accurately record the spatial position and attitude changes of the sound source and the sensor itself. All sensing nodes transmit the acquired multi-channel time-domain sound pressure signals and their corresponding spatiotemporal coordinate information into data frames in real time to the cross-environment acoustic feature extraction module via wired gigabit Ethernet or a wireless communication network compliant with the Wireless Fidelity 6 standard.

[0024] The data frame format includes a frame header, sensor node identifier, timestamp, multi-channel sound pressure data block, and inertial measurement unit data block. The frame header includes the data length and checksum to ensure transmission integrity. As a preferred implementation, the sensor array employs a self-organizing network topology. When some sensor nodes in the system fail due to physical obstruction, power failure, or communication interruption, the remaining healthy nodes can dynamically reconstruct the sensing network through a distributed negotiation protocol. The reconstruction process is based on neighbor discovery and link quality assessment algorithms, and utilizes sparse signal recovery techniques from compressed sensing theory to recover the complete spatial distribution information of the sound field from the reduced-dimensional, non-uniform sampled data with high probability. This ensures the robustness of the system's sensing function even under non-ideal conditions where some nodes fail.

[0025] Please refer to the attached document. Figure 3 With appendix Figure 5 The cross-environment acoustic feature extraction module receives raw multi-channel acoustic data streams from the sound field perception and data acquisition module. The core task of this module is to perform high-level acoustic feature abstraction, transforming the raw sound pressure signals into low-dimensional embedding vectors that characterize the essential acoustic properties of the room. The module first initiates a data preprocessing pipeline.

[0026] The first step in preprocessing is to window the input raw sound pressure signal using a Hanning window function to reduce spectral leakage. The window length is set to 1024 sampling points, and the overlapping region is set to 512 sampling points. The second step in preprocessing is to resample the windowed signal, standardizing the sampling rate to 48000 Hz to eliminate inconsistencies that may arise from different sampling sources. Subsequently, the module performs a short-time Fourier transform on the preprocessed signal for each channel, converting the time-domain signal into a time-frequency domain representation, generating a time-spectrum. The frequency resolution of the time-spectrum is 46.875 Hz, and the time frame shift is 10.67 milliseconds. The core of feature extraction is a generalized room impulse response feature encoder, which is implemented based on a deep convolutional neural network architecture. This network architecture specifically includes 8 convolutional layers and 3 fully connected layers. The input to this network is a three-dimensional data tensor composed of stacked multi-channel time-spectrums. The first convolutional layer uses 64 3x3 convolutional kernels with a stride of 1 to perform convolution operations and extract low-order acoustic features. It is then connected to a max pooling layer with a stride of 2 for downsampling.

[0027] Subsequent convolutional layers progressively increase the number of channels to 512, continuously using 3x3 convolutional kernels and max pooling with a stride of 2 to gradually extract more abstract and representative acoustic features. Following the convolutional layers are three fully connected layers with 2048, 1024, and 512 neurons, respectively. In these fully connected layers, a random deactivation technique with a dropout rate of 0.5 is used to randomly disable some neurons during forward propagation, effectively preventing overfitting of the model to the training data. Finally, the network outputs a floating-point vector with a fixed dimension of 512, namely the acoustic scene embedding vector. This vector encapsulates key acoustic properties such as the geometry of the current room, the distribution of the sound absorption coefficient of the interface materials, the main early reflected sound paths, and the distribution of normal modes. The calculation of this embedding vector is completely independent of specific noise types or speech content, achieving a unified and quantitative representation of different acoustic physical environments.

[0028] To quantify the differences between different acoustic environments, the system calculates the cosine similarity between the embedding vectors of two acoustic scenes. The cosine similarity value ranges from -1 to +1; the closer to +1, the more similar the two acoustic environments are, and the closer to -1, the greater the difference. This similarity metric provides an important reference for subsequent model adaptation strategies.

[0029] Please refer to the attached document. Figure 2The meta-learning-driven model rapid adaptation module is the intelligent core of the system's ability to generalize across environments. This module pre-loads a deep neural network model, called the basic acoustic optimizer, which has been deeply optimized through a meta-learning training strategy. The meta-training process of this basic acoustic optimizer is completed before the system is deployed offline. The meta-training phase utilizes a diverse task dataset covering 100 different room acoustic configurations. Each task simulates a unique acoustic environment adaptation problem, including the acoustic scene embedding vector of that environment and the corresponding noise target acoustic spectrum pair.

[0030] The training objective is to optimize the initial network weight parameters of the basic acoustic optimizer, enabling the model to possess an inherent rapid adaptability. This means that when faced with a completely new acoustic environment, it can quickly adjust its internal parameters using only a small number of samples from that environment and a minimal number of gradient updates, thus efficiently modeling the acoustic characteristics of the new environment. This module preferentially employs a model-independent meta-learning algorithm as its core learning paradigm. This algorithm comprises two loop optimization processes: an outer loop optimization process responsible for adjusting the model's initial parameters during the meta-training phase, ensuring the model has a good starting point for all task distributions. When the system is officially deployed to a new building interior environment, the meta-learning-driven model rapid adaptation module begins online operation. It receives the acoustic scene embedding vector of the current new environment calculated by the cross-environment acoustic feature extraction module, and simultaneously receives a small number of noise target acoustic spectrum pairs collected on-site, corresponding to the current environment. The total amount of these sample data is strictly controlled within less than 30 minutes of dual-channel audio recording data. The module then immediately initiates the inner loop optimization process.

[0031] This process fine-tunes only a few specific fully connected layer parameters in the basic acoustic optimizer using stochastic gradient descent. The number of fine-tuning steps is strictly controlled to 2 to 5, with a learning rate set to 0.001. After this highly efficient fine-tuning process, the parameters of the basic acoustic optimizer are rapidly updated, instantly enabling it to model the acoustic characteristics of the current environment with high precision. The fine-tuned model can generate a set of highly customized noise reduction filter coefficients for digital signal processing based on the real-time input acoustic scene embedding vector. Simultaneously, it outputs suggested values ​​for adjusting the surface impedance of sound-absorbing materials optimized for this environment, providing precise guidance for subsequent physical control.

[0032] The core algorithm of the meta-learning-driven model fast adaptation module involves the rapid updating of model parameters on a limited amount of data. This process can be described as follows: given the initial parameter vector of the basic acoustic optimizer and the support set data of the new environmental task, the model updates its parameters by performing several gradient descent iterations. The parameter update rule can be formally expressed as: in, Indicates the initial parameters of the model. This represents the model parameters after inner loop optimization. This represents the learning rate of the inner loop. This represents the gradient of the loss function with respect to the parameters on the new environment task support set. This formula describes how the model utilizes a small number of samples from the new environment to achieve rapid adaptation by calculating the loss gradient and updating the parameters in the reverse direction of the gradient. After this efficient fine-tuning, the model can generate noise reduction filter coefficients and sound-absorbing material impedance tuning parameters optimized for the new environment.

[0033] Please refer to the attached document. Figure 4 The multi-objective acoustic optimization decision module bears the crucial responsibility of comprehensively balancing multiple acoustic performance indicators. This module receives preliminary optimization suggestions from the meta-learning-driven model's rapid adaptation module, including noise reduction filter coefficients and sound absorption control suggestions. Internally, the module maintains a set of preset optimization objective functions, which quantify different dimensions of acoustic quality. The primary optimization objective is to maximize the speech transmission index, with a target threshold set at no less than 0.75 to ensure excellent speech intelligibility. The second key objective is to minimize the background noise level, requiring the A-weighted sound pressure level of indoor background noise to be stably controlled below 35 dB. The third important objective is to maintain or optimize the spatial perception parameter, precisely controlling it within the ideal range of 0.5 to 0.8 to preserve the naturalness and spatial immersion of the indoor sound field.

[0034] These objectives may conflict with each other; for example, excessive noise suppression might weaken spatial perception. Therefore, the decision-making process is implemented through a weighted summation multi-objective optimization algorithm. This algorithm dynamically adjusts the weight coefficients of each objective function. The adjustment strategy for the weight coefficients is based on two key inputs: one is a real-time calculated comprehensive acoustic comfort index, which is a normalized weighted combination of the speech transmission index, noise level, and spatial perception parameters; the other is a personalized preference profile preset by the user through the system interface, which explicitly specifies the user's priority for different acoustic attributes. As a preferred implementation, this module integrates a multi-objective solver based on a non-dominated sorting genetic algorithm to efficiently handle this complex optimization problem.

[0035] The solver initializes a candidate solution population of 100, each representing a complete set of acoustic processing parameters. The solver iteratively performs genetic operations such as selection, crossover, and mutation. Selection is based on non-dominated sorting and crowding calculation, prioritizing the retention of superior individuals that are not dominated by other solutions and are distributed in sparse regions of the target space. Crossover uses a simulated binary crossover operator, exchanging some parameters of parent individuals with a certain probability. Mutation introduces small perturbations through polynomial mutation to maintain population diversity. After several generations of evolution, the algorithm eventually approximates and outputs a Pareto optimal solution set.

[0036] This solution set contains all non-dominated solutions that cannot improve any objective without harming others. The system decision-maker or automated decision logic can select the most suitable acoustic processing strategy from this Pareto-optimal solution set based on real-time acoustic requirements and energy constraints. Finally, the multi-objective acoustic optimization decision module outputs a globally optimal set of acoustic processing parameter instructions. This instruction set specifies in detail the required noise attenuation at the center frequency of each one-third octave band, the target reverberation time frequency response curve, and the identifiers of specific reflected sound paths that need to be prioritized for suppression.

[0037] Please refer to the attached document. Figure 5 The distributed sound-absorbing actuator control module is the execution terminal of the system, translating digital optimization decisions into physical acoustic control effects. This module connects and precisely controls multiple active sound-absorbing units and tunable acoustic panel arrays deployed in the indoor space via industrial Ethernet or a dedicated control bus. Each active sound-absorbing unit is an independent intelligent entity integrating a speaker array and a high-performance digital signal processor. The feedback controller inside the unit receives noise reduction filter coefficients and target noise attenuation spectra from the multi-target acoustic optimization decision module in real time. Based on these parameters, the controller generates an anti-noise signal with opposite phase and matched amplitude to the incident noise sound wave through an adaptive filtering algorithm, driving the speaker array to emit this anti-noise field, thereby achieving destructive interference of sound energy in a specific area and achieving active noise reduction. The tunable acoustic panel is a device that dynamically adjusts its sound absorption characteristics by changing its physical structure.

[0038] Each panel incorporates a precision mechanical adjustment mechanism driven by a miniature servo motor. The control module sends angle or displacement control commands to these servo motors, causing changes in the micro-perforation array structure on the panel surface, allowing for continuous and precise adjustment of its effective perforation rate within the range of 5% to 30%. By altering the perforation rate, the panel's Helmholtz resonant frequency changes, thereby achieving targeted and efficient absorption of noise in specific frequency bands, especially mid-to-low frequency noise. A closed-loop control circuit is constructed within the distributed sound-absorbing actuator control module.

[0039] This circuit continuously monitors two types of key status information: first, the operating status of each actuator, including servo motor position feedback, speaker impedance, and unit temperature; second, the residual noise level and real-time acoustic parameters fed back by the sound field sensing and data acquisition module. The control module compares the residual noise with the optimization target, calculates the error signal, and dynamically adjusts the control commands sent to the actuators using a proportional-integral-derivative control algorithm to ensure that the acoustic optimization effect can be continuously and stably maintained within the preset target range. Furthermore, this module also has advanced energy management functions. Internally, it maintains an actuator energy consumption database, recording the real-time power consumption of each active sound-absorbing unit and tunable panel under different operating modes.

[0040] The system periodically evaluates current acoustic performance indicators. Once it is confirmed that all key acoustic indicators meet or exceed preset thresholds, the energy management logic is activated. This logic, based on an energy efficiency ratio ranking algorithm, prioritizes the operation of combinations of high-efficiency actuators that can generate greater noise attenuation per unit of power consumption. At the same time, it appropriately reduces or suspends the output of some low-efficiency actuators, thereby strictly controlling the total power consumption of the entire distributed actuator system to below 80% of the system's rated maximum power, achieving an optimal balance between high acoustic optimization performance and low energy consumption.

[0041] Data interaction between modules within the system follows strict timing and protocol specifications. The sound field perception and data acquisition module sends data packets to the cross-environment acoustic feature extraction module every 100 milliseconds. The feature calculation latency of the cross-environment acoustic feature extraction module is less than 50 milliseconds. The fine-tuning process of the meta-learning-driven model rapid adaptation module is completed within 500 milliseconds after receiving new data. The optimization solution cycle of the multi-objective acoustic optimization decision module is approximately 1 second.

[0042] The control command update frequency of the distributed acoustic absorber actuator control module is as high as 100 Hz. This precise timing design ensures that the total system delay from the occurrence of an acoustic event to the actuator's corresponding control action is less than 2 seconds, meeting the application requirements of real-time interactive acoustic control. All modules are equipped with a watchdog timer and heartbeat detection mechanism. Any abnormal timeout or failure of any module will trigger a system-level exception handling process, which includes faulty module isolation, system function degradation, and administrator alarms, maximizing the availability and robustness of the system.

[0043] Example 2: This example focuses on another specific implementation of a deep learning-based building interior acoustic noise reduction and sound absorption optimization system in a specific application scenario, emphasizing the deployment and optimization strategy adjustment of the system in an open office environment.

[0044] In open-plan office environments, the sensor array deployment strategy for the sound field perception and data acquisition module needs targeted optimization. The array density is increased to one sensor node per 5 square meters to address the more complex sound field distribution and variable sound source locations in open spaces. The sensor nodes also integrate directional microphone arrays, using beamforming technology to enhance the capture of voice signals from specific conversation areas while suppressing background noise interference from other directions. The inertial measurement unit's data update frequency is increased to 200 Hz for more accurate tracking of acoustic environmental transients caused by personnel movement.

[0045] In this scenario, the cross-environment acoustic feature extraction module's deep convolutional neural network undergoes enhanced training specifically tailored to the acoustic features of the office environment. The training dataset includes a significant amount of impulse response data and acoustic parameter labels for open-plan offices. An attention mechanism layer is introduced into the network structure, enabling the model to adaptively focus on feature regions that have the greatest impact on acoustic comfort in office environments, such as speech frequency bands from 500 Hz to 4000 Hz. The output acoustic scene embedding vector maintains the same dimensionality, but its feature space distribution is adjusted to better distinguish common acoustic problems in open-plan offices, such as excessively high levels of speech privacy or low levels of team collaboration.

[0046] In the application of meta-learning-driven model rapid adaptation modules in open-plan office scenarios, the rapid adaptation process introduces a prior knowledge transfer mechanism based on task similarity. When the system is deployed to a new open-plan office, the module first calculates the cosine similarity between the new environment's acoustic scene embedding vector and the embedding vectors of all open-plan office environment tasks in the meta-training task library. The three historical tasks with the highest similarity are selected as reference tasks, and their model fine-tuning trajectories are used as prior knowledge to initialize the fine-tuning starting point and direction for the new task. This further reduces the number of steps in the inner loop gradient descent, typically requiring only two fine-tuning steps to achieve the target performance, reducing the on-site data requirement from 30 minutes to 15 minutes, and significantly improving deployment efficiency.

[0047] The multi-objective acoustic optimization decision module has made significant adjustments to the weighting of optimization objectives in open-plan office environments. Since speech clarity and focus are paramount in an office environment, the speech transmission index (STI) has been prioritized, with its threshold increased from 0.75 to 0.8. Background noise level control remains important, but a weighting of up to 40 dB A is allowed during non-core work periods. The weighting of spatial perception parameters has been relatively reduced, with the target range adjusted to 0.4 to 0.7, focusing more on reducing unnecessary spatial reverberation to avoid distractions. The population initialization strategy of the multi-objective solver has also been optimized, prioritizing the injection of candidate solutions historically proven effective in similar open-plan office environments as part of the initial seed, accelerating the evolutionary convergence process.

[0048] The distributed acoustic actuator control module, deployed in an open-plan office, exhibits zoned collaborative control characteristics. The entire office area is divided into several independent acoustic control zones, each managed by a dedicated set of active sound-absorbing units and tunable acoustic panels. The control module implements a zoned collaborative control strategy. When sensors detect high-intensity conversations or noise events within a zone, the actuators in that zone immediately activate a powerful noise reduction mode. Simultaneously, actuators in adjacent zones receive coordination commands, adjusting their operating parameters to form an acoustic buffer zone, preventing noise from spreading and interfering with other areas. The tunable acoustic panels' adjustment strategy focuses on absorbing mid-to-high frequency noise to quickly attenuate typical office noises such as keyboard clicks and telephone rings. Energy management in this scenario incorporates a time-of-use electricity pricing strategy. During peak electricity price periods, the system appropriately relaxes noise control targets in non-core areas, prioritizing acoustic quality in the core work area while reducing overall system energy consumption for economical operation.

[0049] The system's communication protocols in open office environments have also been enhanced, employing time-sensitive networking technology to ensure extremely low jitter and deterministic latency transmission of control commands in complex office network environments. This ensures the synchronization of distributed actuator actions and avoids acoustic performance fluctuations caused by control asynchrony. Furthermore, the system provides an application programming interface (API) that allows integration with enterprise office management systems. Based on contextual data such as meeting room reservation status and personnel presence information, it can dynamically preload corresponding acoustic optimization configuration files, achieving more intelligent, scenario-adaptive acoustic environment management.

[0050] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish an entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A deep learning-based building interior acoustic noise reduction and sound absorption optimization system, characterized in that, include: The sound field perception and data acquisition module is used to acquire acoustic environment data of the building interior in real time; The cross-environment acoustic feature extraction module is used to receive raw acoustic data from the sound field perception and data acquisition module and perform feature abstraction process; The meta-learning-driven model rapid adaptation module is used to achieve cross-environment generalization. The meta-learning-driven model rapid adaptation module is pre-configured with a deep neural network model optimized by the meta-learning training strategy as the basic acoustic optimizer. When the system is deployed to a new building environment, the module receives the acoustic scene embedding vector from the cross-environment acoustic feature extraction module and a small number of noise target acoustic spectrum pairs collected on site. The module performs an internal loop optimization process. A multi-objective acoustic optimization decision module is used to comprehensively balance multiple acoustic performance indicators. The multi-objective acoustic optimization decision module receives preliminary optimization suggestions from the meta-learning driven model fast adaptation module and makes refined decisions based on a set of preset optimization objective functions. The distributed sound-absorbing actuator control module is used to convert the abstract parameters output by the multi-objective acoustic optimization decision module into specific physical control commands. The distributed sound-absorbing actuator control module connects to and controls multiple active sound-absorbing units and tunable acoustic panels deployed indoors.

2. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The sound field sensing and data acquisition module also includes an inertial measurement unit for synchronously recording the spatial position changes of the sound source and the sensor. The sound field sensing and data acquisition module transmits the acquired multi-channel time-domain sound pressure signals and their corresponding spatiotemporal coordinate information to the cross-environment acoustic feature extraction module through a wired or wireless communication network.

3. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The cross-environment acoustic feature extraction module first preprocesses the original sound pressure signal, including windowing using a Hanning window and resampling at a sampling rate of 48,000 Hz. Then, it calculates the short-time Fourier transform of each channel signal to convert the time-domain signal into a time-frequency domain representation. The core of feature extraction lies in constructing a generalized room impulse response feature encoder. The encoder is based on a deep convolutional neural network architecture. Its input is a multi-channel time-frequency map, and its output is a 512-dimensional acoustic scene embedding vector. This embedding vector condenses the current room's geometry, the sound absorption coefficient of the interface material, and the key acoustic properties of the reflection path.

4. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The sensor array in the sound field perception and data acquisition module adopts a self-organizing network topology. When some sensor nodes fail due to faults or occlusion, the remaining nodes can dynamically reconstruct the perception network through negotiation and use compressed sensing algorithms to recover complete sound field information from sparse sampled data.

5. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The deep convolutional neural network architecture used in the cross-environment acoustic feature extraction module includes 8 convolutional layers and 3 fully connected layers. The convolutional layers use 3x3 convolutional kernels and max pooling operations with a stride of 2 to extract abstract representations of acoustic features layer by layer. The fully connected layers prevent overfitting by using a random deactivation technique with a dropout rate of 0.

5.

6. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The meta-learning-driven model rapid adaptation module adopts a model-independent meta-learning algorithm as its core learning paradigm. In the meta-training phase, the model is trained on a diverse task set covering 100 different room acoustic configurations, with each task simulating a specific acoustic environment adaptation problem.

7. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 4, characterized in that, The parameter update process of the model-independent meta-learning algorithm is as follows: Given the initial parameter vector of the basic acoustic optimizer and the support set data of the new environment task, the model updates the parameters by performing gradient descent several times. The parameter update rule is expressed as the initial parameters minus the learning rate multiplied by the gradient of the loss function with respect to the parameters on the support set of the new environment task.

8. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The multi-objective acoustic optimization decision module integrates a multi-objective solver based on a non-dominated sorting genetic algorithm. The solver maintains a population of 100 candidate solutions and performs selection, crossover, and mutation operations to iteratively evolve, eventually approximating the Pareto optimal solution set.

9. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The distributed sound-absorbing actuator control module also has an energy management function. This function monitors the power consumption of each actuator in real time, and, provided that the acoustic performance meets the preset threshold, prioritizes scheduling the combination of actuators with high energy efficiency ratio, so as to control the total power consumption of the system to below 80% of the rated power.

10. The deep learning-based building interior acoustic noise reduction and sound absorption optimization system according to claim 1, characterized in that, The decision-making process of the multi-objective acoustic optimization decision module is realized through a weighted summation multi-objective optimization algorithm. The algorithm dynamically adjusts the weight coefficients of each objective function, and the adjustment of the weight coefficients is based on the acoustic comfort index calculated in real time and the user's preset preference configuration file.