Multi-mode driving sound field control method and system for waiting area

By equivalently dividing the waiting area and collecting multimodal data, a dynamic correlation model of people-noise-sound effects was constructed, and the sound field control parameters were optimized. This solved the problems of poor adaptability and insufficient stability of the sound field control in the waiting area, and realized intelligent regulation and stability improvement of the zoned sound field.

CN121832399APending Publication Date: 2026-04-10PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The existing sound field control in the waiting area lacks the ability to accurately adapt to different zones, and is easily affected by dynamic changes in people and noise, resulting in poor adaptability and insufficient stability.

Method used

By dividing the waiting area into equivalent sections, deploying multimodal sensing devices to collect multimodal data streams, constructing a dynamic correlation model of people, noise, and sound effects, performing sound field control analysis, and dynamically optimizing the linkage through a preset smooth adjustment strategy, the closed-loop control of the sound field in the waiting area is achieved.

Benefits of technology

It realizes intelligent control of the zoned sound field in the waiting area, improves the adaptability and stability of the sound field, and ensures that the sound field always adapts to the current environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121832399A_ABST
    Figure CN121832399A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal driving sound field control method and system for a waiting area, and relates to the technical field of waiting area sound field control, and the method comprises the steps: carrying out the equivalent division of a target waiting area, obtaining N waiting sub-areas, and deploying a multi-modal sensing device; multi-modal data streams of N waiting areas are collected, and a noise feature set of the number of people in the N areas is obtained; driving to construct a people number-noise-sound effect dynamic correlation model, and outputting N initial area sound field control parameters; and dynamic linkage optimization is carried out, sound field control parameters of the N target areas are determined, and closed-loop control of the sound field of the waiting area is carried out. The technical problems that in the prior art, sound field control of the waiting area lacks the partition precise adaptation capacity and is prone to being affected by dynamic changes of crowds and noise, so that the adaptability is poor, and the stability is insufficient are solved, and the technical effects that partition sound field intelligent regulation and control of the waiting area are achieved, and the adaptability and stability of the sound field are improved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound field control technology in waiting areas, specifically to a multimodal driven sound field control method and system for waiting areas. Background Technology

[0002] In public places such as waiting areas, the acoustic environment directly affects people's experience and comfort. However, existing acoustic control in waiting areas mostly adopts a unified fixed parameter control mode, which lacks the ability to accurately perceive and adapt to the dynamic changes in population density and environmental noise fluctuations in the area, and does not consider the mutual interference of acoustic propagation in different areas. This results in problems such as poor adaptability, insufficient stability, and lack of regional coordination in acoustic control, making it difficult to meet the personalized acoustic needs of different waiting areas and different population sizes.

[0003] Existing technologies for sound field control in waiting areas lack the ability to precisely adapt to different zones, and are susceptible to dynamic changes in crowds and noise, resulting in poor adaptability and insufficient stability. Summary of the Invention

[0004] This application provides a multimodal driven sound field control method and system for waiting areas, which addresses the technical problems of existing waiting area sound field control lacking precise zone adaptation capability and being easily affected by dynamic changes in crowds and noise, resulting in poor adaptability and insufficient stability.

[0005] In view of the above problems, this application provides a multimodal driven sound field control method and system for waiting areas.

[0006] A first aspect of this application provides a multimodal driven sound field control method for a waiting area, the method comprising: The target waiting area is equivalently divided into N sub-regions. Multimodal sensing devices are deployed in these N sub-regions, each integrating an audio sensor, an infrared / thermal imaging module, a visible light camera, millimeter-wave radar, and a ToF depth sensor. Multimodal data streams from the N waiting areas are collected using these devices. These data streams are then spatiotemporally aligned, and features are extracted and fused to obtain N sets of noise features related to the number of people in each region. A dynamic correlation model of people, noise, and sound effects is constructed. This model is used to analyze the sound field of the N regions, outputting N initial sound field control parameters. The N initial sound field control parameters are dynamically optimized according to a preset smoothing adjustment strategy to determine the sound field control parameters for the N target regions. Finally, closed-loop sound field control of the waiting area is implemented based on these target region sound field control parameters.

[0007] A second aspect of this application provides a multimodal driven sound field control system for a waiting area, the system comprising: A multimodal sensing device deployment module is used to equally divide the target waiting area into N waiting sub-areas, and deploy multimodal sensing devices on the N waiting sub-areas. The multimodal sensing devices integrate audio sensors, infrared / thermal imaging modules, visible light cameras, millimeter-wave radar, and ToF depth sensors. A noise feature set acquisition module is used to collect multimodal data streams from the N waiting areas through the multimodal sensing devices, perform spatiotemporal alignment and feature extraction and fusion on the N waiting area multimodal data streams, and obtain N area noise feature sets of people. A sound field control parameter output module is used to drive the construction of a dynamic correlation model of people-noise-sound effects, use the dynamic correlation model of people-noise-sound effects to perform sound field control analysis on the noise feature sets of people in the N areas, and output N initial area sound field control parameters. A closed-loop control module is used to dynamically link and optimize the sound field control parameters of the N initial areas according to a preset smooth adjustment strategy, determine the sound field control parameters of the N target areas, and perform closed-loop control of the waiting area sound field based on the sound field control parameters of the N target areas.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: The target waiting area is equivalently divided into N sub-regions. Multimodal data streams from these N waiting areas are collected using a multimodal sensing device. These data streams are then spatiotemporally aligned, and features are extracted and fused to obtain N sets of noise features per person per region. A dynamic correlation model of people, noise, and sound effects is constructed, outputting N initial sound field control parameters for each region. These initial sound field control parameters are dynamically optimized according to a preset smoothing adjustment strategy to determine the sound field control parameters for the N target regions. Based on these parameters, closed-loop sound field control is implemented in the waiting area. This achieves intelligent regional sound field control of the waiting area, improving sound field adaptability and stability. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A schematic flowchart of a multimodal driven sound field control method for a waiting area is provided in an embodiment of this application; Figure 2This is a schematic diagram of a multimodal driven sound field control system for a waiting area, provided as an embodiment of this application.

[0011] Figure reference numerals: Multimodal sensing device deployment module 10, noise feature set acquisition module 20, sound field control parameter output module 30, closed-loop control module 40. Detailed Implementation

[0012] This application provides a multimodal driven sound field control method and system for waiting areas, which addresses the technical problems in existing waiting area sound field control that lack precise zone adaptation capabilities and are easily affected by dynamic changes in crowds and noise, resulting in poor adaptability and insufficient stability.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] Example 1, as Figure 1 As shown, this application provides a multimodal driven sound field control method for a waiting area, the method comprising: Step S100: Divide the target waiting area into N equivalent waiting sub-areas, and deploy a multimodal sensing device on the N waiting sub-areas. The multimodal sensing device integrates an audio sensor, an infrared / thermal imaging module, a visible light camera, a millimeter-wave radar, and a ToF depth sensor.

[0015] Specifically, the process begins by acquiring waiting area division rules that include spatial continuity, functional consistency, population density balance, and sound field independence. These rules are then prioritized to determine a region division priority sequence. Following this sequence, correlation data mining is performed on the target waiting area to obtain association rule attribute data. Based on the region division priority sequence, these attribute data are then evaluated and divided into N waiting sub-areas. Subsequently, multimodal sensing devices are deployed in each of these N sub-areas. These devices integrate audio sensors, infrared / thermal imaging modules, visible light cameras, millimeter-wave radar, and ToF depth sensors, enabling comprehensive data collection of environmental noise, population status, and other multi-dimensional data. This provides all-round data support for the accurate determination of subsequent sound field control parameters.

[0016] Step S200: Collect N multimodal data streams from waiting areas using the multimodal sensing device, perform spatiotemporal alignment and feature extraction fusion on the N multimodal data streams from waiting areas, and obtain N regional noise feature sets of people.

[0017] Specifically, leveraging multimodal sensing devices deployed in N waiting areas, multimodal data streams from each sub-area are simultaneously collected, covering multiple dimensions such as environmental noise, crowd density, and population flow. First, the collected N waiting area multimodal data streams undergo spatiotemporal alignment processing to eliminate time differences and spatial location deviations between different sensors, resulting in N usable waiting area multimodal data streams. Then, based on the structural characteristics of the data collected by each sensor in the multimodal sensing device, a suitable set of sensor data feature extraction algorithms is selected. Using this algorithm set, crowd characteristics, such as density and flow trends, and noise characteristics, such as intensity and frequency distribution, are extracted from the N usable waiting area multimodal data streams, resulting in N multimodal crowd noise feature sets. Finally, decision-level fusion is performed on the different modal features in each multimodal crowd noise feature set, integrating the effective information from various dimensions to ultimately obtain N regional population noise feature sets that comprehensively and accurately reflect the status of each sub-area.

[0018] Step S300: Drive the construction of a dynamic correlation model of number of people-noise-sound effect, and use the dynamic correlation model of number of people-noise-sound effect to perform sound field control analysis on the sound field of the N regional noise feature sets, and output N initial regional sound field control parameters.

[0019] Specifically, the process begins by collecting a historical sound field dataset of the waiting area, including crowd density trend data, regional noise data, sound field effect data, and corresponding sound field control effect data. This dataset is then filtered based on the sound field control effect data to obtain a usable waiting area sound field dataset. The crowd density trend data and regional noise data from this dataset are used as input data, and the sound field effect data is used as output data, with correlation and labeling applied to construct a waiting area sound field sample set. A deep neural network is then employed, specifically a hybrid model combining convolutional neural networks and long short-term memory networks. This model utilizes convolutional neural networks to extract spatial features of the data, while the long short-term memory network captures the temporal relationship between crowd density and noise. The model uses the sequence variation law to train and fit the sound field control of the sample set. The network weights and biases are iteratively adjusted through backpropagation algorithm, and cross-validation is used for iterative verification and optimization until the model prediction error is lower than the preset threshold or the convergence condition of the maximum training rounds is reached, thus completing the construction of the dynamic correlation model of people-noise-sound effect. Then, the obtained N regional people and noise feature sets are input into the model. The model accurately maps and adapts the broadcast volume, audio output mode and other control schemes by hierarchically analyzing the features such as the size of people, noise intensity and frequency distribution in each region. Finally, it outputs N initial regional sound field control parameters, which provide a basis for subsequent parameter optimization.

[0020] Step S400: Dynamically optimize the sound field control parameters of the N initial regions according to the preset smooth adjustment strategy, determine the sound field control parameters of the N target regions, and perform closed-loop control of the sound field in the waiting area based on the sound field control parameters of the N target regions.

[0021] Specifically, firstly, the rule system corresponding to the preset smooth adjustment strategy is clearly defined, covering adjustment amplitude limits, adjustment frequency limits, and linkage coordination rules, providing a constraint basis for parameter optimization; then, based on the spatial location characteristics of N waiting sub-regions, a geometric acoustic model is used to conduct sound field propagation simulation, constructing a sound field propagation model for the waiting area. By simulating and evaluating the propagation impact of each sub-region, a set of sound field propagation impact coefficients for N regions and the remaining regions is obtained. Each waiting sub-region is used as a regional sound field node, and a visual connection is made based on the coefficient set to form a sound field propagation impact network for the waiting area; subsequently, according to the smooth adjustment strategy rules, combined with the sound field... The propagation influence network analyzes the interaction relationships of the sound fields in each sub-region, and dynamically optimizes the sound field control parameters of N initial regions to avoid interference to surrounding regions caused by adjustments in a single region. Finally, it determines the sound field control parameters of N target regions that take into account both regional independence and overall coordination. Based on these target parameters, the sound field control of the waiting area is implemented, while the sound field status of each region is monitored in real time to obtain the sound field control feedback parameters of N regions. The feedback parameters are used to continuously optimize the sound field control parameters of the target regions. Through a closed-loop control mechanism, the sound field of the waiting area is always at the optimal level that adapts to the current environmental conditions, achieving predictive and continuous control of the sound field.

[0022] In one possible implementation, step S100 further includes: Step S110: Obtain the waiting area division rules, which include spatial continuity, functional consistency, crowd density balance and sound field independence.

[0023] Step S120: Perform priority analysis and sorting on each division rule in the waiting area division rules to determine the area division priority sequence rule.

[0024] Step S130: Perform association data mining on the target waiting area according to the waiting area division rules to obtain the waiting area association rule attribute data.

[0025] Step S140: Based on the region division priority sequence rule, perform equivalent evaluation and division on the association rule attribute data of the waiting area to obtain N waiting sub-regions.

[0026] Specifically, based on the core need for precise sound field control in waiting areas, a key rule system for the division of waiting areas was identified and obtained. This system covers four core dimensions: spatial continuity, ensuring that the divided waiting sub-areas are physically connected and seamless, avoiding blind spots in cross-regional data collection and control; functional consistency, ensuring that the same sub-area has the same or related medical service functions, such as waiting areas for the same department, aligning with the functional orientation of personnel activities; balanced population density, ensuring that the expected or historical population distribution of each sub-area is relatively balanced through division, avoiding data collection deviations or sound field control failures due to excessive concentration of people in certain areas; and sound field independence, reducing mutual interference between sound fields in different sub-areas, providing a basis for personalized sound field control in each area. The above four rules together constitute the core basis for the scientific division of waiting areas, laying the foundation for subsequent priority ranking and equivalent division.

[0027] Based on the core objective of precise sound field control in the waiting area and the needs of practical application scenarios, a priority analysis was conducted on the four division rules already acquired: spatial continuity, functional consistency, crowd density balance, and sound field independence. Spatial continuity, as the physical foundation, determines the integrity of data collection and control in sub-regions and must be prioritized. Functional consistency aligns with the functional orientation of personnel activities in the waiting area and directly affects the targeting of sound field requirements, thus having a lower priority. Sound field independence is a key prerequisite for achieving personalized sound field control in different zones, avoiding interference between zones, and has a close second priority. Crowd density balance ensures a balance between the control load and effect in each sub-region, facilitating subsequent parameter optimization, and has a relatively lower priority. Through the above multi-dimensional evaluation and ranking, the priority sequence rules for regional division of spatial continuity, functional consistency, sound field independence, and crowd density balance were finally determined, providing a clear priority guide for subsequent equivalent evaluation and division of the waiting area.

[0028] Guided by the waiting area zoning rules of spatial continuity, functional consistency, population density balance, and sound field independence, this study conducts multi-dimensional correlation data mining on target waiting areas. It focuses on collecting and analyzing physical spatial layout data of the waiting areas, such as wall locations, passageway distribution, and area boundaries. This data is matched with spatial continuity rules and functional zoning planning data, such as waiting areas corresponding to departments and the distribution of service facilities. It also aligns with functional consistency rules and historical and real-time population flow and density statistics, such as peak-hour population distribution and average number of people staying in each area. Furthermore, it supports population density balance rules and sound field propagation characteristics data, such as acoustic parameters of building materials, obstacle distribution, and the location and coverage of existing audio equipment. Finally, it adapts to sound field independence rules. Through the integration, filtering, and correlation analysis of this data, key attribute information that accurately reflects whether the target waiting area conforms to various zoning rules is extracted, ultimately forming the waiting area correlation rule attribute data.

[0029] Guided by the established priority sequence of regional division rules—namely, spatial continuity > functional consistency > sound field independence > population density balance—and combined with the associated rule attribute data of the waiting area, a systematic equivalence evaluation and division was conducted, covering dimensions such as spatial layout, functional zoning, population distribution, and sound field characteristics. First, based on the highest priority rule of spatial continuity, candidate areas with no physical spatial fragmentation and complete connectivity were selected, excluding division schemes with spatial breaks. Next, the functional consistency rule was used for further selection, ensuring that the same candidate area possessed the same or related medical service functions, eliminating functionally mixed areas. Then, referring to the sound field independence rule, the degree of sound field interference between candidate areas was assessed, and the boundaries were adjusted to reduce the mutual influence of sound fields between areas to below a preset threshold. Finally, supplemented by the population density balance rule, the expected population distribution of each candidate area was calibrated to avoid situations with significant disparities in the number of people. After multiple rounds of rule matching and equivalence verification, the target waiting area was ultimately accurately divided into N sub-areas that met all division rules and were suitable for subsequent multimodal perception and sound field control requirements.

[0030] In one possible implementation, step S200 further includes: Step S210: Perform spatiotemporal alignment processing on the N waiting area multimodal data streams to obtain N available waiting area multimodal data streams.

[0031] Step S220: Select a set of sensor data feature extraction algorithms based on the data structure characteristics of each sensor in the multimodal sensing device.

[0032] Step S230: Use the sensor data feature extraction algorithm set to extract crowd and noise features from the N available waiting area multimodal data streams to obtain N multimodal crowd noise feature sets.

[0033] Step S240: Perform decision-level fusion on the modal features of the N multimodal crowd noise feature sets to obtain N regional crowd noise feature sets.

[0034] Specifically, for the multimodal data streams collected by multimodal sensing devices in N waiting sub-areas, covering various data types such as audio, temperature distribution, images, distance-velocity, and 3D space, systematic spatiotemporal alignment processing is carried out. In the time dimension, a unified time base is used as a reference, and the acquisition timestamps of each sensor data are calibrated to eliminate time deviations caused by differences in response speeds between different sensors, ensuring the time synchronization of data from different modalities at the same moment. In the spatial dimension, the spatial position of the data collected by different sensors is matched and calibrated by combining the spatial coordinate mapping relationship of each waiting sub-area. For example, the coordinates of infrared / thermal imaging data, visible light image data and 3D spatial data from ToF depth sensors are aligned to eliminate data misalignment caused by spatial position offset. Through the dual processing of time synchronization and spatial calibration, the consistency and correlation of the multimodal data streams in each waiting sub-area in the spatiotemporal dimensions are ensured, ultimately resulting in N usable multimodal data streams for waiting areas.

[0035] Based on the structural characteristics of the data acquired by each sensor in the multimodal sensing device, a suitable set of sensor data feature extraction algorithms is selected. For the noise time-domain / frequency-domain sequence data output by the audio sensor, the Mel-frequency cepstral coefficient (MFCC) algorithm and the power spectral density analysis algorithm are selected to extract key information such as noise intensity, frequency distribution, and spectral characteristics. For the temperature distribution matrix data generated by the infrared / thermal imaging module, the threshold segmentation algorithm and the region growing algorithm are used to assist in identifying areas where people gather and density-related features. For the image pixel array data acquired by the visible light camera, the YOLO target detection algorithm combined with the density clustering algorithm is used to realize the counting of people and the extraction of movement trajectory features. For the distance-velocity two-dimensional data output by the millimeter-wave radar, the Kalman filtering algorithm and the peak detection algorithm are used to extract dynamic features such as the movement speed and position changes of people. For the three-dimensional spatial point cloud data acquired by the ToF depth sensor, the voxelization algorithm and the spatial clustering algorithm are used to mine the regional space occupancy rate and personnel distribution density features. By integrating the above algorithms adapted to the data types of each sensor, a complete set of sensor data feature extraction algorithms is formed.

[0036] A set of defined sensor data feature extraction algorithms was applied to N available waiting area multimodal data streams. Targeted crowd and noise feature extraction was performed for different modalities. For audio data streams, the Mel-frequency cepstral coefficient (MFCC) algorithm and power spectral density analysis algorithm were used to extract core noise features such as noise intensity, frequency distribution, and spectral peaks. For infrared / thermal imaging data streams, threshold segmentation and region growing algorithms were used to identify crowd gathering hotspots and extract crowd density-related features. For visible light image data streams, the YOLO target detection algorithm was used to count the number of people, combined with density clustering algorithms to capture crowd movement trajectories and clustering trends. For millimeter-wave radar data streams, Kalman filtering and peak detection algorithms were used to extract dynamic features such as movement speed and position change amplitude. For ToF depth data streams, voxelization and spatial clustering algorithms were used to obtain three-dimensional spatial features such as area space occupancy and crowd distribution density. By performing the above feature extractions on the multimodal data streams of each waiting area sub-region, N multimodal crowd noise feature sets containing multidimensional noise features and multidimensional crowd status features were finally formed.

[0037] For the multimodal features including audio, infrared / thermal imaging, visible light, millimeter-wave radar, and ToF depth contained in N multimodal crowd noise feature sets, a decision-level fusion method is adopted for integration and optimization. First, based on the acquisition accuracy, environmental adaptability, and correlation with crowd / noise status of each modality feature, the weight coefficients of different modal features are determined by the analytic hierarchy process (AHP) or entropy weighting method, such as the weight of audio modality on noise features and the weight of visible light and ToF depth modes on crowd density features. Then, combined with Bayesian decision theory or weighted voting method, the consistency of similar features extracted from each modality, such as crowd density data and noise intensity data obtained from different sensors, is checked and redundancy is eliminated to strengthen high-confidence features and correct low-reliability data. Finally, the fused crowd features and noise features are structurally integrated to form a unique regional population noise feature set for each waiting sub-region that can comprehensively and accurately reflect the core status of the current crowd and noise. Finally, N regional population noise feature sets corresponding one-to-one with each sub-region are output.

[0038] In one possible implementation, step S300 further includes: Step S310: Collect historical sound field dataset of the waiting area through data-driven acquisition. The historical sound field dataset of the waiting area includes crowd density trend data, regional noise data, sound field effect data, and corresponding sound field control effect data.

[0039] Step S320: Optimize the historical sound field dataset of the waiting area according to the sound field control effect data to obtain an usable sound field dataset for the waiting area.

[0040] Step S330: Use a deep neural network to control and train the available waiting area sound field dataset to construct a dynamic correlation model of number of people, noise and sound effects.

[0041] Specifically, a comprehensive collection of historical sound field data for waiting areas was conducted using a data-driven approach to construct a historical sound field dataset covering all factors related to sound field regulation. This dataset focuses on four core data dimensions: crowd density trend data includes real-time statistics of the number of people in each waiting area at different times, the distribution of peak periods, the duration of crowd flow, and the rate of population movement; regional noise data includes time-series variation curves of noise intensity in each area, noise frequency distribution characteristics, main noise source types, and noise fluctuation ranges under different scenarios; sound field effect data includes parameters such as broadcast volume levels and audio output modes set for different environmental conditions, such as voice broadcasts, prompt tone types, sound effect frequency ranges, and differentiated settings for zoned sound effects; and corresponding sound field control effect data includes related data such as auditory comfort ratings of the sound field environment, accuracy of medical information transmission, assessment results of sound field interference levels between different areas, and user feedback. By integrating the above multi-dimensional data, a comprehensive foundational data support is provided for subsequent model training and parameter optimization.

[0042] Using the sound field control effect data from the historical sound field dataset of the waiting area as the core selection criterion, a dataset optimization process was carried out. First, the quantitative evaluation criteria for sound field control effect were defined, including key indicators such as the threshold for auditory comfort scores, the minimum requirement for the accuracy of medical information transmission, and the upper limit of sound field interference between areas. Then, each data sample in the historical sound field dataset was verified one by one, and invalid control samples with comfort scores below the threshold, information transmission accuracy below the standard, and excessive regional interference were removed. At the same time, inferior data with incomplete data fields and abnormal values ​​deviating from the normal range were filtered out. Then, through correlation analysis, high-quality samples that clearly reflect the effective correspondence between crowd density, regional noise, and sound field effects and have stable control effects were retained to ensure that the selected data can provide reliable support for model training. Finally, all qualified samples that passed the verification were integrated to form a usable waiting area sound field dataset with complete structure, reliable data, and significant patterns.

[0043] The structure of the deep neural network is clearly defined, employing a hybrid network architecture of Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), and fully connected layers. The CNN layers are responsible for extracting spatial features from the input data. This is achieved through three convolutional layers with kernel sizes of 3×3, 5×5, and 3×3, combined with max pooling layers (2×2 kernels), to hierarchically extract spatial distribution correlation features from crowd density trend data and regional noise data. The LSTM layers contain two bidirectional LSTM units to capture the temporal variation patterns and long-term / short-term dependencies of crowd density and noise intensity. Each LSTM layer has 128 units and uses a Dropout layer with a dropout rate of 0.3 to prevent overfitting. The fully connected layers consist of three hidden layers with 256, 128, and 64 neurons respectively, and one output layer. The output layer uses a linear activation function to map and obtain sound field and sound effect related parameters. Subsequently, a suitable waiting area sound field dataset is used. Using crowd density trend data and regional noise data as network input data and sound field effect data as output data, the input and output data are correlated and labeled to construct a standardized sound field sample set for the waiting area. The sample set is divided into training, validation, and test sets in a 7:2:1 ratio. The Adam optimizer is used with an initial learning rate of 0.001, dynamically adjusted according to an exponential decay strategy, and mean squared error (MSE) as the loss function to train and fit the deep neural network for sound field control. During training, the model performance is monitored in real time through the validation set. An early stopping method is used, stopping training if the loss on the validation set does not decrease after 10 consecutive rounds to avoid overfitting. Iterative validation and optimization are performed using cross-validation until the model reaches the preset convergence condition, with the loss function value below 0.005 and the prediction accuracy on the test set above 95%. Finally, a dynamic correlation model of people, noise, and sound effects that accurately maps the relationship between crowds, noise, and sound effects is constructed.

[0044] In one possible implementation, step S330 further includes: Step S331: Use the crowd density trend data and regional noise data in the available waiting area sound field dataset as input data, and the sound field effect data as output data.

[0045] Step S332: Associate and identify the input data and output data to obtain the sound field sample set of the waiting area.

[0046] Step S333: Use a deep neural network to train and fit the sound field control of the sound field sample set in the waiting area, iteratively verify and optimize it until the preset convergence condition is met, and construct a dynamic correlation model of number of people-noise-sound effect.

[0047] Specifically, in the process of constructing the dynamic correlation model of people-noise-sound effects, the dimensions of the input and output data of the model are clearly defined. The population density trend data and regional noise data that reflect the core information of the waiting area environment status in the available waiting area sound field dataset are defined as the model input data. At the same time, the sound field control scheme parameters, i.e., the sound field sound effect data, corresponding to the environmental status in the dataset are determined as the model output data. Through the precise definition of input and output data, a clear data mapping relationship is established for the subsequent sample set construction and model training fitting.

[0048] For the identified input and output data, association and labeling work was carried out. Using environmental state information composed of population density trend data and regional noise data as the core association basis, each set of input data was precisely bound one-to-one with the corresponding sound field effect data. This ensured that each set of input data reflecting the environmental state of the waiting area could clearly correspond to the appropriate sound field control output data. At the same time, unified labeling information was added to each set of bound data to facilitate data retrieval and traceability during subsequent model training. By eliminating abnormal data with mismatched associations, missing data, or logical contradictions, a standardized waiting area sound field sample set with clear mapping relationships and reliable data quality was finally integrated, providing a standardized data foundation for the training and fitting of deep neural networks.

[0049] A hybrid deep neural network architecture consisting of Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), and fully connected layers is employed. The CNN layers extract spatial correlation features from the input data in the waiting area sound field sample set through multiple convolutional operations and pooling with varying kernel sizes. The LSTM layers capture the temporal variation patterns and long-term / short-term dependencies between crowd density trends and regional noise through bidirectional units. The fully connected layers map the extracted features to corresponding output parameters. The waiting area sound field sample set is divided into training, validation, and test sets in a reasonable proportion. Mean squared error is used as the loss function, and the Adam optimizer is used to dynamically adjust the learning rate for sound field control training and fitting of the deep neural network. During training, the model's prediction performance is monitored in real-time using the validation set. Cross-validation is used to iteratively optimize network weights, biases, and other parameters. Dropout layers and early stopping mechanisms are implemented to prevent overfitting. Training continues until the model meets preset convergence conditions, such as the loss function value falling below a set threshold and the test set prediction accuracy reaching the target standard. Ultimately, a dynamic correlation model of people, noise, and sound effects is constructed, accurately depicting the relationship between crowds, noise, and sound effects, and possessing stable prediction capabilities.

[0050] In one possible implementation, step S400 further includes: Step S410: Determine smooth adjustment strategy rules according to the preset smooth adjustment strategy. The smooth adjustment strategy rules include adjustment range limits, adjustment frequency limits, and linkage coordination rules.

[0051] Step S420: Based on the spatial location characteristics of the N waiting areas, perform sound field propagation influence analysis to obtain the sound field propagation influence network of the waiting area.

[0052] Step S430: Based on the sound field propagation influence network of the waiting area, dynamically optimize the sound field control parameters of the N initial regions according to the smooth adjustment strategy rules, and determine the sound field control parameters of the N target regions.

[0053] Specifically, guided by the core objective of a pre-set smooth adjustment strategy, a complete set of smooth adjustment strategy rules is defined. This rule system covers three key aspects: adjustment range limits, which define the maximum range of change for core parameters such as broadcast volume and sound effect frequency during a single sound field adjustment to avoid auditory discomfort to waiting patients due to sudden parameter changes; adjustment frequency limits, which set the maximum number of adjustments to sound field control parameters per unit time to ensure the stability and continuity of the sound field environment in the waiting area and prevent interference caused by frequent adjustments; and linkage and coordination rules, which define the collaborative logic of sound field parameter adjustments between different waiting sub-areas, and clarify the adaptation adjustment principles for surrounding related areas when parameters in one area change, ensuring that the sound field adjustments in each area cooperate with each other without significant interference.

[0054] Based on the spatial distribution, wall layout, and passageway characteristics of N waiting area sub-regions, a geometric acoustic model is first used to simulate the sound field propagation process of each sub-region. By simulating the direct, reflected, and refracted paths of sound waves and the energy attenuation law, a sound field propagation model of the waiting area that accurately reflects the sound field interaction between regions is constructed. Then, based on this model, each sub-region is treated as a sound field emission source, and its sound field coverage, intensity attenuation, and interference impact on all other sub-regions are simulated. The corresponding sound field propagation influence coefficients between each region are obtained through quantitative calculation, forming a set of sound field propagation influence coefficients for N regions and the remaining regions. Finally, each waiting area sub-region is set as an independent regional sound field node. The correlation strength between nodes is determined based on the sound field propagation influence coefficients between regions. Nodes with correlation strength reaching a preset threshold are visually connected. The higher the correlation strength, the greater the weight of the connection line. Finally, a sound field propagation influence network of the waiting area that clearly presents the correlation and influence degree of sound field propagation in each sub-region is constructed.

[0055] Using the smooth adjustment strategy rule as the core constraint, and combining the sound field correlation and influence degree of each sub-region presented by the sound field propagation influence network of the waiting area, dynamic linkage optimization is carried out on the sound field control parameters of N initial regions. First, based on the adjustment range limit, the maximum range of single change of parameters such as broadcast volume and sound effect frequency in each region is defined to avoid auditory discomfort caused by parameter abrupt changes. Then, according to the adjustment frequency limit, the number of parameter adjustments per unit time is controlled to ensure the stability of the sound field environment. At the same time, following the linkage coordination rule, when adjusting the parameters of any region, the propagation influence coefficient of that region and the surrounding regions in the sound field propagation influence network is referenced to simultaneously calculate the adaptation adjustment range of the surrounding regions to ensure that the sound fields between regions are interference-free and mutually compatible. Through multiple rounds of iterative calibration, the parameters of each region are continuously optimized to balance the adaptability of individual regions and the overall regional coordination. Finally, N target region sound field control parameters are determined that can accurately match the population and noise status of each sub-region and ensure the harmonious and stable overall sound field of the waiting area.

[0056] In one possible implementation, step S420 further includes: Step S421: Use a geometric acoustic model to simulate the sound field propagation of the spatial location characteristics of the N waiting areas, and construct a sound field propagation model for the waiting area.

[0057] Step S422: Based on the sound field propagation model of the waiting area, simulate the propagation impact and assess the degree of propagation impact on each of the N waiting area sub-regions to obtain the sound field propagation impact coefficient set of N regions and the remaining regions.

[0058] Step S423: Based on the set of sound field propagation influence coefficients of the N regions and the remaining regions, obtain the sound field propagation influence network of the waiting area.

[0059] Specifically, taking the spatial location distribution, physical structure layout, wall material, passageway direction, and obstruction of N waiting sub-areas as core inputs, a geometric acoustic model is used to conduct a refined sound field propagation simulation. By simulating the propagation paths of sound waves within and between each sub-area, including direct, reflected, and refracted propagation paths, key parameters such as energy attenuation, propagation delay, and phase change during sound wave propagation are accurately calculated. At the same time, combined with the spatial boundary conditions and acoustic characteristics of each sub-area, the sound field distribution patterns under different scenarios are restored, and finally, a waiting area sound field propagation model that can truly reflect the sound field propagation characteristics of the waiting area and is highly consistent with the actual spatial environment is constructed.

[0060] Using the established sound field propagation model of the waiting area as the core analysis tool, the sound field propagation impact simulation and degree assessment are carried out for each of the N waiting area sub-regions. Each sub-region is treated as an independent sound field emission source. The model simulates the complete process of its sound waves propagating to all other sub-regions, accurately capturing key features such as energy attenuation and interference superposition along the propagation path. At the same time, combined with the spatial characteristics and acoustic environment of each receiving sub-region, the sound field impact intensity of the emission source sub-region on each receiving sub-region is quantitatively evaluated. Finally, a dataset containing the quantitative values ​​of the impact degree on all other sub-regions is generated for each emission source sub-region, thus forming N sets of sound field propagation impact coefficients that correspond one-to-one with each sub-region.

[0061] First, each of the N waiting sub-areas is assigned as an independent regional sound field node, clearly defining each node as the core of sound field propagation in its corresponding waiting sub-area. Then, based on the set of sound field propagation influence coefficients between the N regions and the rest of the regions, the sound field propagation influence coefficient between each regional sound field node and all other regional sound field nodes is extracted, and this coefficient is used as the basis for determining the strength of the association between nodes. Finally, the regional sound field nodes whose association strength reaches a preset threshold are visualized and connected. The higher the coefficient, the more significant the weight of the connection link. By integrating all nodes and links with effective associations, a waiting area sound field propagation influence network is finally formed, which can intuitively present the relationship and degree of mutual propagation influence between the sound fields of each waiting sub-area.

[0062] In one possible implementation, step S423 further includes: Each of the N waiting area sub-regions is taken as an N regional sound field node.

[0063] Based on the set of sound field propagation influence coefficients of the N regions and the remaining regions, the sound field nodes of the N regions are visualized and connected to obtain the sound field propagation influence network of the waiting area.

[0064] Specifically, in the process of constructing the sound field propagation influence network in the waiting area, the N pre-divided waiting sub-areas are first defined as nodes. Each waiting sub-area is assigned as an independent regional sound field node. Each node is precisely associated with the spatial location, acoustic environment, and core sound field propagation attributes of the corresponding sub-area, clarifying the one-to-one correspondence between the nodes and the actual waiting sub-areas. This lays the foundation for subsequent node association and construction of a visualized influence network based on the inter-regional sound field propagation influence coefficient.

[0065] Based on the quantitative data of the sound field propagation influence coefficients of N regions and the remaining regions, a visual connection operation is performed on the N pre-defined sound field nodes. For each region's sound field node, the sound field propagation influence coefficient between it and all other region's sound field nodes is extracted. It is then determined whether this coefficient reaches a preset association threshold, retaining only valid associations that meet the threshold. Simultaneously, the difference in association strength between nodes is defined based on the magnitude of the influence coefficient, and this difference in strength is reflected through differentiated visualization methods, such as line thickness and color depth. The higher the coefficient, the thicker the connecting line or the darker the color, and vice versa. All region sound field nodes that meet the conditions and their corresponding association links are integrated and presented, ultimately forming a waiting area sound field propagation influence network that intuitively and clearly reflects the sound field propagation influence relationships and degrees between each waiting area sub-region.

[0066] In one possible implementation, step S400 further includes: Step S440: Based on the sound field control parameters of the N target areas, perform sound field control monitoring in the waiting area to obtain sound field control feedback parameters for the N areas.

[0067] Step S450: Optimize the sound field control parameters of the N target areas using the N area sound field control feedback parameters, and perform closed-loop control of the sound field in the waiting area using the optimized N target area sound field control parameters.

[0068] Specifically, after initiating sound field control in the waiting area based on sound field control parameters for N target areas, multimodal sensing devices deployed in each waiting sub-area, integrating audio sensors, infrared / thermal imaging modules, visible light cameras, millimeter-wave radar, and ToF depth sensors, are used to conduct real-time sound field control monitoring. This continuously collects actual sound field status data for each area, including actual broadcast volume, sound effect propagation intensity, and comprehensive acoustic characteristics after environmental noise superposition. Subsequently, the collected monitoring data undergoes spatiotemporal alignment calibration and validity verification, eliminating abnormal interference data. Then, a feature extraction algorithm is used to extract core parameters that accurately reflect the actual effect of sound field control in each area, ultimately forming N area sound field control feedback parameters that correspond one-to-one with the N waiting sub-areas.

[0069] The sound field control feedback parameters of N regions are compared dimension-by-dimensionally with the preset sound field control target parameters. The difference between the actual sound field state and the target state of each waiting sub-region is quantified by the deviation analysis algorithm to locate the root cause of the parameter deviation. Then, combined with the spatial characteristics of each sub-region, the correlation in the sound field propagation influence network, and the smooth adjustment strategy rules, the adaptive PID optimization algorithm is used to accurately correct the sound field control parameters of the original N target regions, ensuring that the adjusted parameters meet the amplitude limit, frequency limit, and regional linkage coordination requirements. Subsequently, the optimized target region sound field control parameters are sent to the sound field regulation execution units of each sub-region in real time, such as smart speakers and zone amplifiers, to drive the equipment to complete the dynamic adjustment of the sound field. At the same time, new regional sound field control feedback parameters are continuously collected through multimodal sensing devices, repeating the closed-loop process of deviation analysis-parameter optimization-regulation execution-feedback collection to achieve dynamic iterative optimization and stable control of the sound field in the waiting area.

[0070] Example 2, based on the same inventive concept as the multimodal driven sound field control method for waiting areas in the foregoing examples, such as... Figure 2 As shown, this application provides a multimodal driven sound field control system for a waiting area. The system and method embodiments in this application are based on the same inventive concept. The system includes: The multimodal sensing device deployment module 10 is used to equally divide the target waiting area into N waiting sub-areas, and deploy multimodal sensing devices on the N waiting sub-areas. The multimodal sensing devices integrate an audio sensor, an infrared / thermal imaging module, a visible light camera, a millimeter-wave radar, and a ToF depth sensor.

[0071] The noise feature set acquisition module 20 is used to collect N multimodal data streams from waiting areas through the multimodal sensing device, perform spatiotemporal alignment and feature extraction and fusion on the N multimodal data streams from waiting areas, and obtain N noise feature sets of the number of people in the area.

[0072] The sound field control parameter output module 30 is used to drive the construction of a dynamic correlation model of number of people, noise and sound effects. The dynamic correlation model of number of people, noise and sound effects is used to perform sound field control analysis on the sound field of the N regional noise feature sets and output N initial regional sound field control parameters.

[0073] The closed-loop control module 40 is used to dynamically optimize the sound field control parameters of the N initial regions according to a preset smooth adjustment strategy, determine the sound field control parameters of the N target regions, and perform closed-loop control of the sound field in the waiting area based on the sound field control parameters of the N target regions.

[0074] Furthermore, the system is also used to implement the following functions: The process involves obtaining waiting area division rules, including spatial continuity, functional consistency, population density balance, and sound field independence; prioritizing and ranking each division rule to determine a region division priority sequence rule; performing correlation data mining on the target waiting area according to the waiting area division rules to obtain waiting area correlation rule attribute data; and performing equivalent evaluation and division on the waiting area correlation rule attribute data based on the region division priority sequence rule to obtain N waiting sub-areas.

[0075] Furthermore, the system is also used to implement the following functions: Spatiotemporal alignment processing is performed on the N multimodal data streams of the waiting areas to obtain N available multimodal data streams of the waiting areas; based on the data structure characteristics of each sensor in the multimodal sensing device, a set of sensor data feature extraction algorithms is selected; the set of sensor data feature extraction algorithms is used to extract crowd and noise features from the N available multimodal data streams of the waiting areas to obtain N multimodal crowd noise feature sets; decision-level fusion is performed on the modal features in the N multimodal crowd noise feature sets to obtain N regional crowd noise feature sets.

[0076] Furthermore, the system is also used to implement the following functions: A historical sound field dataset of the waiting area is collected through data-driven methods. The historical sound field dataset of the waiting area includes crowd density trend data, regional noise data, sound field effect data, and corresponding sound field control effect data. The historical sound field dataset of the waiting area is optimized according to the sound field control effect data to obtain a usable sound field dataset of the waiting area. A deep neural network is used to control, train, and fit the usable sound field dataset of the waiting area to construct a dynamic correlation model of people-noise-sound effect.

[0077] Furthermore, the system is also used to implement the following functions: The population density trend data and regional noise data in the available waiting area sound field dataset are used as input data, and the sound field effect data is used as output data. The input data and output data are associated and identified to obtain the waiting area sound field sample set. A deep neural network is used to train and fit the sound field control of the waiting area sound field sample set, and iteratively verify and optimize it until the preset convergence condition is met, so as to construct a dynamic correlation model of people-noise-sound effect.

[0078] Furthermore, the system is also used to implement the following functions: Based on the preset smoothing adjustment strategy, smoothing adjustment strategy rules are determined, including adjustment amplitude limits, adjustment frequency limits, and linkage coordination rules; sound field propagation impact analysis is performed based on the spatial location characteristics of the N waiting sub-regions to obtain the waiting area sound field propagation impact network; according to the smoothing adjustment strategy rules, the sound field control parameters of the N initial regions are dynamically linked and optimized based on the waiting area sound field propagation impact network to determine the sound field control parameters of the N target regions.

[0079] Furthermore, the system is also used to implement the following functions: A geometric acoustic model is used to simulate the sound field propagation of the spatial location characteristics of the N waiting areas, thus constructing a sound field propagation model for the waiting area. Based on the sound field propagation model, the propagation impact is simulated and the degree of impact is evaluated for each of the N waiting areas, resulting in a set of sound field propagation impact coefficients between the N areas and the remaining areas. Based on the set of sound field propagation impact coefficients between the N areas and the remaining areas, a sound field propagation impact network for the waiting area is obtained.

[0080] Furthermore, the system is also used to implement the following functions: Each of the N waiting area sub-regions is taken as an N regional sound field node; the N regional sound field nodes are visualized and connected based on the set of sound field propagation influence coefficients between the N regions and the remaining regions to obtain the sound field propagation influence network of the waiting area.

[0081] Furthermore, the system is also used to implement the following functions: Based on the sound field control parameters of the N target areas, the sound field control monitoring of the waiting area is carried out to obtain the sound field control feedback parameters of the N areas; the sound field control feedback parameters of the N areas are used to optimize the sound field control parameters of the N target areas, and the sound field closed-loop control of the waiting area is carried out through the optimized sound field control parameters of the N target areas.

[0082] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0083] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0084] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.

Claims

1. A multimodal driven sound field control method for a waiting area, characterized in that, The method includes: The target waiting area is divided into N equivalent waiting sub-areas. Multimodal sensing devices are deployed on the N waiting sub-areas. The multimodal sensing devices integrate an audio sensor, an infrared / thermal imaging module, a visible light camera, a millimeter-wave radar, and a ToF depth sensor. The multimodal sensing device collects multimodal data streams from N waiting areas, performs spatiotemporal alignment and feature extraction and fusion on the N multimodal data streams from the waiting areas, and obtains N regional noise feature sets of people. Drive the construction of a dynamic correlation model of number of people, noise and sound effects, and use the dynamic correlation model of number of people, noise and sound effects to perform sound field control analysis on the sound field control of the N regional number of people and noise feature sets, and output N initial regional sound field control parameters; According to the preset smooth adjustment strategy, the sound field control parameters of the N initial regions are dynamically linked and optimized to determine the sound field control parameters of the N target regions, and the sound field closed-loop control of the waiting area is performed based on the sound field control parameters of the N target regions.

2. The multimodal driven sound field control method for a waiting area as described in claim 1, characterized in that, We obtain N waiting sub-regions, including: Obtain the waiting area division rules, which include spatial continuity, functional consistency, population density balance, and sound field independence; The priority of each division rule in the waiting area division rule is analyzed and sorted to determine the priority sequence rule for area division; According to the waiting area division rules, the target waiting area is subjected to association data mining to obtain the waiting area association rule attribute data. Based on the region division priority sequence rule, the association rule attribute data of the waiting area is evaluated and divided into N waiting sub-regions.

3. The multimodal driven sound field control method for a waiting area as described in claim 1, characterized in that, Obtain noise feature sets of population in N regions, including: Spatiotemporal alignment processing is performed on the N multimodal data streams of the waiting areas to obtain N available multimodal data streams of the waiting areas; Based on the structural characteristics of the data collected by each sensor in the multimodal sensing device, a set of sensor data feature extraction algorithms is selected; The sensor data feature extraction algorithm set is used to extract crowd and noise features from the N available waiting area multimodal data streams to obtain N multimodal crowd and noise feature sets; Decision-level fusion is performed on the modal features of the N multimodal population noise feature sets to obtain N regional population noise feature sets.

4. The multimodal driven sound field control method for a waiting area as described in claim 1, characterized in that, Driven by the construction of a dynamic correlation model of people, noise, and sound effects, including: The historical sound field dataset of the waiting area is collected through data-driven methods. The historical sound field dataset of the waiting area includes crowd density trend data, regional noise data, sound field effect data, and corresponding sound field control effect data. The historical sound field dataset of the waiting area is optimized based on the sound field control effect data to obtain a usable sound field dataset for the waiting area. A deep neural network was used to control and train the available waiting area sound field dataset to construct a dynamic correlation model of people-noise-sound effects.

5. The multimodal driven sound field control method for a waiting area as described in claim 4, characterized in that, A deep neural network is used to control and train the available waiting area sound field dataset to construct a dynamic correlation model of people, noise, and sound effects, including: The population density trend data and regional noise data in the available waiting area sound field dataset are used as input data, and the sound field effect data is used as output data. The input and output data are associated and identified to obtain a sound field sample set for the waiting area; A deep neural network is used to train, fit, iteratively verify, and optimize the sound field sample set in the waiting area until the preset convergence condition is met, thereby constructing a dynamic correlation model of people-noise-sound effects.

6. The multimodal driven sound field control method for a waiting area as described in claim 1, characterized in that, Determine the sound field control parameters for N target regions, including: Based on the preset smoothing adjustment strategy, smoothing adjustment strategy rules are determined, including adjustment range limits, adjustment frequency limits, and linkage coordination rules. Based on the spatial location characteristics of the N waiting areas, a sound field propagation influence analysis is performed to obtain the sound field propagation influence network of the waiting area. According to the smooth adjustment strategy rules, the sound field control parameters of the N initial regions are dynamically linked and optimized based on the sound field propagation influence network of the waiting area, so as to determine the sound field control parameters of the N target regions.

7. The multimodal driven sound field control method for a waiting area as described in claim 6, characterized in that, The sound field propagation influence network in the waiting area was obtained, including: A geometric acoustic model was used to simulate the sound field propagation of the spatial location characteristics of the N waiting areas, and a sound field propagation model of the waiting area was constructed. Based on the sound field propagation model of the waiting area, the propagation impact is simulated and the degree of assessment is evaluated for each of the N waiting area sub-regions, resulting in a set of sound field propagation impact coefficients for the N regions and the remaining regions. Based on the set of sound field propagation influence coefficients for the N regions and the remaining regions, the sound field propagation influence network of the waiting area is obtained.

8. The multimodal driven sound field control method for a waiting area as described in claim 7, characterized in that, Based on the set of sound field propagation influence coefficients for the N regions and the remaining regions, the sound field propagation influence network of the waiting area is obtained, including: Each of the N waiting area sub-regions is taken as an N regional sound field node; Based on the set of sound field propagation influence coefficients of the N regions and the remaining regions, the sound field nodes of the N regions are visualized and connected to obtain the sound field propagation influence network of the waiting area.

9. A multimodal driven sound field control method for a waiting area as described in claim 1, characterized in that, Closed-loop control of the waiting area sound field is performed based on the sound field control parameters of the N target areas, including: Based on the sound field control parameters of the N target areas, sound field control monitoring of the waiting area is performed to obtain sound field control feedback parameters of the N areas; The sound field control parameters of the N target areas are optimized using the sound field control feedback parameters of the N areas, and closed-loop control of the sound field in the waiting area is performed using the optimized sound field control parameters of the N target areas.

10. A multimodal driven sound field control system for a waiting area, characterized in that, The system is used to implement the multimodal driven sound field control method for a waiting area according to any one of claims 1-9, the system comprising: A multimodal sensing device deployment module is used to equally divide the target waiting area into N waiting sub-areas, and deploy multimodal sensing devices on the N waiting sub-areas. The multimodal sensing devices integrate an audio sensor, an infrared / thermal imaging module, a visible light camera, a millimeter-wave radar, and a ToF depth sensor. The noise feature set acquisition module is used to collect N multimodal data streams from waiting areas through the multimodal sensing device, perform spatiotemporal alignment and feature extraction and fusion on the N multimodal data streams from waiting areas, and obtain N noise feature sets of people in the area. The sound field control parameter output module is used to drive the construction of a dynamic correlation model of number of people, noise and sound effects. The dynamic correlation model of number of people, noise and sound effects is used to perform sound field control analysis on the number of people and noise feature sets of the N regions, and outputs N initial region sound field control parameters. The closed-loop control module is used to dynamically optimize the sound field control parameters of the N initial regions according to a preset smooth adjustment strategy, determine the sound field control parameters of the N target regions, and perform closed-loop control of the sound field in the waiting area based on the sound field control parameters of the N target regions.