A noise automatic monitoring system and method based on multi-sensor data fusion
Through multi-sensor data fusion technology, accurate separation and prediction of multi-source noise in complex urban noise environments are achieved, solving the data distortion and insufficient prediction problems of existing noise monitoring systems and improving the accuracy and management efficiency of noise monitoring.
Patent Information
- Application Number
- CN202510983221.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing noise monitoring systems lack the ability to automatically determine the validity of data, cannot identify data distortion caused by environmental factors and equipment failures, cannot adaptively handle complex urban noise environments, and lack the ability to intelligently predict and proactively analyze noise pollution sources.
By adopting multi-sensor data fusion technology, through timestamp synchronization calibration, noise spectrum singular value decomposition, independent component analysis and spatiotemporal convolutional neural network, the accurate separation and prediction of multi-source noise data can be achieved, and the spatial propagation characteristics of noise sources are constructed in combination with urban geographic information.
It has achieved accurate positioning of various noise pollution sources and estimation of responsibility weights in complex urban environments, improved the accuracy of noise monitoring and management efficiency, and provided scientific and forward-looking decision-making support.
Smart Images

Figure CN120493030B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a noise automatic monitoring system and method based on multi-sensor data fusion. Background Art
[0002] Traditional noise monitoring systems typically rely on a single sound level meter for data collection, primarily assessing ambient noise levels by measuring sound pressure levels. These systems can provide basic noise intensity information and determine whether levels exceed specified limits based on preset thresholds. Existing multi-sensor noise monitoring methods are also beginning to incorporate a variety of devices, including voiceprint sensors, vibration sensors, and meteorological sensors, using data fusion technology to conduct comprehensive noise monitoring. These methods have, to a certain extent, improved monitoring accuracy and reliability.
[0003] However, existing technologies have significant shortcomings: first, traditional methods lack the ability to automatically determine the validity of data and are unable to identify data distortion caused by environmental factors (rainfall, wind speed, snowfall, etc.) and equipment failures; second, existing multi-sensor data fusion methods mainly use simple rule-based judgment and fixed weight distribution, which cannot adaptively handle complex urban noise environments; third, they lack the ability to intelligently predict and forward-lookingly analyze noise pollution sources, and are unable to accurately predict the spatiotemporal distribution trends and development patterns of different types of pollution sources such as traffic noise, construction noise, and industrial noise.
[0004] Based on the analysis of the above technical defects, the core technical problem faced by existing technologies is: how to build an intelligent multi-sensor data fusion system to achieve automatic prediction, accurate positioning and responsibility weight estimation of various noise pollution sources in complex urban environments, and at the same time have the ability to automatically judge data validity and predict noise propagation patterns, so as to solve the problem of transitioning from traditional passive monitoring to active prediction and prevention. Summary of the Invention
[0005] The present application provides a noise automatic monitoring system and method based on multi-sensor data fusion, which is used to solve the technical problem that existing noise monitoring technology cannot achieve intelligent identification of multi-source noise and accurate tracing of pollution sources.
[0006] In a first aspect, the present application provides an automatic noise monitoring system based on multi-sensor data fusion, the automatic noise monitoring system based on multi-sensor data fusion comprising:
[0007] The calibration module is used to perform time stamp synchronization calibration on the original noise monitoring data collected by the voiceprint sensor, vibration sensor, and meteorological sensor to obtain a multi-source noise dataset;
[0008] An extraction module is used to perform frequency domain feature extraction processing on the multi-source noise data set using a noise spectrum singular value decomposition algorithm to obtain a noise event feature matrix including a frequency domain feature vector, a time domain feature vector and an environmental parameter vector;
[0009] A separation module is used to perform signal source separation processing on the noise event feature matrix using an independent component analysis algorithm based on attention weights to obtain independent source signal data of traffic noise components, construction noise components, and industrial noise components;
[0010] A construction module is used to perform time-frequency-spatial feature construction processing on the independent source signal data through noise propagation Hilbert transform, and generate noise source spatial propagation feature data in combination with urban geographic information data;
[0011] The classification module is used to perform pattern prediction through a spatiotemporal convolutional neural network based on the spatial propagation characteristic data of the noise source, and output the noise pollution source type prediction result, spatial location coordinates and pollution source responsibility weight distribution data.
[0012] In a second aspect, the present application provides a noise automatic monitoring method based on multi-sensor data fusion, the noise automatic monitoring method based on multi-sensor data fusion comprising:
[0013] The original noise monitoring data collected by the voiceprint sensor, vibration sensor, and meteorological sensor are time-stamped and synchronized to obtain a multi-source noise dataset.
[0014] A noise spectrum singular value decomposition algorithm is used to perform frequency domain feature extraction processing on the multi-source noise data set to obtain a noise event feature matrix including a frequency domain feature vector, a time domain feature vector and an environmental parameter vector;
[0015] Performing signal source separation processing on the noise event feature matrix using an independent component analysis algorithm based on attention weights to obtain independent source signal data of traffic noise components, construction noise components, and industrial noise components;
[0016] The independent source signal data is subjected to a noise propagation Hilbert transform to perform time-frequency-space feature construction processing, and noise source spatial propagation feature data is generated in combination with urban geographic information data;
[0017] According to the spatial propagation characteristic data of the noise source, pattern prediction is performed through a spatiotemporal convolutional neural network, and the noise pollution source type prediction result, spatial location coordinates and pollution source responsibility weight distribution data are output.
[0018] In a third aspect, an automatic noise monitoring device based on multi-sensor data fusion is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the automatic noise monitoring device based on multi-sensor data fusion executes the above-mentioned automatic noise monitoring method based on multi-sensor data fusion.
[0019] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned automatic noise monitoring method based on multi-sensor data fusion.
[0020] In the technical solution provided by this application, the technical problem of inconsistent data timing in traditional noise monitoring is solved through multi-sensor timestamp synchronization calibration processing, ensuring that the data collected by voiceprint sensors, vibration sensors, and meteorological sensors are synchronized within millisecond-level accuracy, laying a reliable data foundation for subsequent accurate prediction and analysis. At the same time, the noise spectrum singular value decomposition algorithm is used to extract features from multi-source noise data sets. Compared with the traditional fast Fourier transform, this algorithm has stronger noise adaptability and spectrum decomposition capabilities, can effectively separate the aliased frequency components and extract comprehensive feature information including frequency domain feature vectors, time domain feature vectors and environmental parameter vectors, providing rich features for noise source trend prediction. Description, signal source separation processing is performed through the independent component analysis algorithm based on attention weights. This algorithm breaks through the limitations of traditional fixed weight distribution, and can adaptively focus on the contribution of different frequency components to various noise sources. It successfully realizes the precise separation of traffic noise components, construction noise components and industrial noise components, and significantly improves the prediction accuracy of multi-source noise in complex urban environments. The noise propagation Hilbert transform combined with the processing method of urban geographic information data has realized the quantitative description of the noise propagation characteristics in three-dimensional urban space for the first time, and can accurately calculate the propagation path and attenuation law of sound waves in complex environments such as buildings and ground, providing a scientific basis for the spatial distribution prediction of noise pollution sources.
[0021] The introduction of a spatiotemporal convolutional neural network (STN) provides powerful pattern prediction and spatiotemporal evolution analysis capabilities. This network architecture is specifically optimized for the spatiotemporal characteristics of noise events. Through the collaborative work of temporal, spatial, and frequency-domain convolutional layers, the system can simultaneously extract deep features of noise data across the three dimensions of time, space, and frequency. Compared to traditional single-dimensional analysis methods, the multidimensional feature fusion mechanism of the present invention significantly improves the accuracy and stability of noise source type prediction. Specifically, in the field of urban noise monitoring, the algorithmic features of the present invention enable the system to accurately predict the development trends of noise sources with similar spectral characteristics but different spatiotemporal distributions, addressing the low prediction accuracy of traditional methods in complex acoustic environments. By outputting noise pollution source type prediction results, spatial location coordinates, and pollution source responsibility weight distribution data, the present invention achieves a technological leap from simple exceedance alarms to intelligent pollution source prediction and control, providing environmental management departments with a scientific, quantitative, and forward-looking decision-making support tool. In specific application scenarios such as smart city noise control, environmental impact assessment, and pollution source supervision, the combination of the present invention's technical features produces significant synergistic effects, improving prediction accuracy and management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 This is a schematic diagram of an embodiment of an automatic noise monitoring system based on multi-sensor data fusion in an embodiment of the present application;
[0024] Figure 2 This is a schematic diagram of an embodiment of a noise automatic monitoring method based on multi-sensor data fusion in an embodiment of the present application;
[0025] Figure 3 It is a schematic block diagram of the structure of an automatic noise monitoring device based on multi-sensor data fusion in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The embodiments of the present application provide a noise automatic monitoring system and method based on multi-sensor data fusion. The terms "first," "second," "third," "fourth," etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products, or devices.
[0027] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, an automatic noise monitoring system based on multi-sensor data fusion includes:
[0028] The calibration module 101 is used to perform time stamp synchronization calibration processing on the original noise monitoring data collected by the voiceprint sensor, vibration sensor, and meteorological sensor to obtain a multi-source noise data set;
[0029] Extraction module 102, configured to perform frequency domain feature extraction processing on a multi-source noise data set using a noise spectrum singular value decomposition algorithm to obtain a noise event feature matrix comprising a frequency domain feature vector, a time domain feature vector, and an environmental parameter vector;
[0030] Separation module 103, configured to perform signal source separation processing on the noise event feature matrix using an attention weighted independent component analysis algorithm to obtain independent source signal data of traffic noise components, construction noise components, and industrial noise components;
[0031] A construction module 104 is used to construct time-frequency-space features of the independent source signal data through a noise propagation Hilbert transform, and generate noise source spatial propagation feature data in combination with urban geographic information data;
[0032] The classification module 105 is used to perform pattern prediction based on the spatial propagation characteristic data of the noise source through a spatiotemporal convolutional neural network, and output the noise pollution source type prediction result, spatial location coordinates and pollution source responsibility weight distribution data.
[0033] It is understandable that the execution subject of this application can be a noise automatic monitoring system based on multi-sensor data fusion, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0034] Specifically, calibration module 101 implements time synchronization processing for multi-sensor data. After the voiceprint sensor collects acoustic signals from the environment, it uses a bandpass filter to remove power frequency interference. The filter's passband is set to cover the effective frequency band for noise monitoring, while also blocking the 50Hz power frequency noise generated by the power system, resulting in pure raw acoustic data. When the vibration sensor detects ground vibration signals, it uses a low-pass filter to eliminate high-frequency noise interference. The cutoff frequency is set within the effective frequency band of the vibration signal, removing sensor circuit noise and environmental electromagnetic interference to obtain true raw vibration data. Meteorological sensors simultaneously monitor environmental parameters such as wind speed, humidity, and temperature. Data validation algorithms are used to check that parameter values are within reasonable ranges, eliminating anomalous data beyond physical probability and ensuring the validity of the raw environmental parameter data. The hardware clock synchronization mechanism uses a unified time base to mark all sensor data, eliminating delay differences between sensors and controlling time errors to the millisecond level. The data is then precisely aligned according to the time series to generate a multi-source noise dataset with a unified time base.
[0035] Extraction module 102 uses a specially designed noise spectrum singular value decomposition (SVD) algorithm for feature extraction. First, a short-time Fourier transform (SFT) is performed on the acoustic data in the multi-source noise dataset, converting the time-domain signal into a time-frequency domain representation. This generates an acoustic spectrum matrix containing both frequency information and temporal evolution information. The noise-specific SVD algorithm is optimized for the unique properties of noise signals, reorganizing the acoustic spectrum matrix with frequency components as rows and time windows as columns, forming a standardized matrix structure. The decomposition process uses eigenvalue decomposition of the covariance matrix and eigenvector orthogonalization to decompose the original spectrum matrix into three components: a left singular vector matrix describing the frequency domain spatial basis, a singular value sequence representing the energy distribution, and a right singular vector matrix describing the time domain spatial basis. Based on the frequency component weights corresponding to each column vector in the left singular vector matrix, a frequency domain eigenvector matrix containing the frequency characteristics of traffic noise, construction noise, and industrial noise is extracted. Based on the weight coefficients for each of the traffic, construction, and industrial noise components, the singular value diagonal matrix is weighted and reconstructed to highlight the spectral features associated with the specific noise type, forming an enhanced spectral feature matrix. The vibration data is subjected to feature dimensionality reduction using a principal component analysis algorithm to extract the main vibration mode information. Finally, the enhanced spectrum feature matrix, time domain feature vector matrix, vibration frequency domain feature vector, and environmental parameter vector are concatenated to form a noise event feature matrix containing multidimensional information. The separation module 103 uses an independent component analysis algorithm based on attention weights to achieve signal source separation. First, the noise event feature matrix is modeled as a linear mixture of the observed signal and the source signal, assuming that the observed mixed signal is formed by a linear combination of multiple independent noise source signals. The introduction of the attention mechanism enables the algorithm to adaptively focus on the contribution of different frequency components to various noise sources. By calculating the dot product of the query vector and the key vector and performing softmax normalization, an attention weight distribution matrix is obtained. This matrix quantifies the importance weight of each frequency component to traffic noise sources, construction noise sources, and industrial noise sources. Guided by the attention weight distribution matrix, the algorithm sets specialized nonlinear separation functions for different noise source types. For traffic noise, an exponential function is used to address its continuity; for construction noise, a hyperbolic tangent function is used to address its burstiness; and for industrial noise, a cubic function is used to address its periodicity. The parameters of these functions are adaptively adjusted based on the attention weights. The separation process uses an iterative optimization algorithm to minimize the mutual information between source signals and find the optimal demixing matrix. When the cross-correlation coefficient between source signals drops below a preset threshold, the signal separation is considered ideal and the iterations cease. The separated signals undergo statistical independence verification to confirm that the noise components are statistically independent. Finally, the independent source signal data for traffic, construction, and industrial noise components are output.
[0036] Construction module 104 constructs time-frequency-spatial features using the Hilbert transform of noise propagation. Empirical mode decomposition is performed on the separated traffic, construction, and industrial noise components, decomposing each noise component into multiple intrinsic mode components and residual components. The eigenmode components represent the signal's oscillation modes at different time scales. Taking into account the impact of factors such as buildings and ground materials on sound wave propagation in urban environments, the algorithm introduces a corresponding spatial attenuation factor to each eigenmode component based on building distribution and ground material information from urban geographic information data, generating modified modal components that reflect the impact of the actual urban environment. The Hilbert transform is applied to the modified modal components, calculating the instantaneous amplitude and instantaneous frequency of each component. The instantaneous amplitude reflects the temporal variation of sound energy, while the instantaneous frequency describes the time-varying characteristics of the frequency components. Combined with information on the propagation distance of the noise from the source to the receiver, a mapping process is performed from the time-frequency domain to the spatial domain, generating a noise energy time-frequency distribution matrix that describes the distribution of sound energy in three-dimensional space. The matrix is associated with the spatial coordinates of urban geographic information data to construct the tracking path of sound rays in the urban environment. Taking into account propagation mechanisms such as building reflection and ground absorption, after three-dimensional spatial interpolation processing, it finally generates complete noise source spatial propagation characteristic data including propagation path and impact range.
[0037] Classification module 105 performs pattern prediction and trend analysis based on a spatiotemporal convolutional neural network. It constructs a four-dimensional feature tensor based on spatial coordinate, time series, and frequency component axes to transform the spatial propagation characteristics of noise sources into a four-dimensional feature tensor. This tensor comprehensively describes the complex correlations and evolutionary patterns of noise events across the three dimensions of time, space, and frequency. The network's temporal convolutional layer specifically processes the temporal evolutionary characteristics and future development trends of noise events. The spatial convolutional layer extracts the geographic distribution patterns and diffusion predictions of noise. The frequency convolutional layer analyzes the spectral characteristics and frequency variation trends of noise. These three convolutional layers work in parallel to extract feature representations and evolutionary patterns in their respective dimensions. A multi-head attention mechanism is used to perform cross-modal semantic fusion of features from different dimensions, calculating the correlation weights and temporal dependencies between different features to generate a unified, comprehensive semantic feature vector that incorporates the multi-dimensional information and predictive characteristics of the noise event. The fully connected prediction layer receives the comprehensive semantic feature vector and uses a softmax activation function to predict and classify the noise source type. It also performs position regression to predict the spatial coordinates of the noise source and the diffusion trend of its impact range. Finally, combined with the predicted noise pollution source type, spatial location, impact range evolution and duration estimate, the responsibility weight estimate data of each pollution source is obtained through weighted calculation to quantify the potential contribution of different noise sources to future environmental pollution.
[0038] In a specific embodiment, the calibration module 101 is configured to:
[0039] Perform bandpass filtering preprocessing on the ambient acoustic signal collected by the voiceprint sensor to obtain the original acoustic data with power frequency interference filtered out;
[0040] The ground vibration signal collected by the vibration sensor is pre-processed by low-pass filtering to obtain the original vibration data with high-frequency noise removed;
[0041] Perform data verification and preprocessing on the wind speed, humidity, and temperature parameters collected by meteorological sensors to obtain the original data of environmental parameters within the valid range;
[0042] Based on the hardware clock synchronization mechanism, the acoustic raw data, vibration raw data and environmental parameter raw data are marked with a unified time reference to obtain synchronized time stamped data;
[0043] The synchronized timestamp marked data are aligned according to the time series to obtain a multi-source noise dataset with a unified time base.
[0044] Specifically, the soundprint sensor signal preprocessing in the calibration module 101 removes power frequency interference through a bandpass filter. The bandpass filter sets a lower cutoff frequency and an upper cutoff frequency to form a frequency window that allows signals to pass through. When the ambient acoustic signal passes through the filter, noise components with frequencies within the passband are allowed to pass through with their original amplitude, while power frequency interference signals outside the passband are significantly attenuated. Power frequency interference mainly comes from fixed-frequency noise generated by the power system. Through the frequency domain selective filtering mechanism, the true environmental noise information is retained in the acoustic raw data while the regular interference generated by electrical equipment is eliminated. The vibration sensor signal preprocessing uses a low-pass filter to eliminate high-frequency noise components. The low-pass filter allows signal components below the cutoff frequency to pass while suppressing noise components above the cutoff frequency. The effective frequency band of the ground vibration signal is mainly concentrated in the lower frequency range, while sensor circuit noise and environmental electromagnetic interference usually appear as high-frequency random noise. Through low-pass filtering, the vibration raw data retains the true ground vibration information caused by the noise source and removes high-frequency interference components that are not related to noise monitoring. Meteorological sensor data validation preprocessing detects and removes outliers by setting physically reasonable ranges for each parameter. Wind speed parameter validation checks whether the value is within the meteorologically acceptable range, humidity parameter validation confirms that the percentage value is between zero and one hundred, and temperature parameter validation verifies that the value conforms to the reasonable range for local climate conditions. When abnormal data outside the preset range is detected, the validation algorithm marks it as invalid and removes it from the data stream. The raw environmental parameter data undergoes validation to ensure that each value has physical meaning and conforms to actual environmental conditions. The hardware clock synchronization mechanism timestamps the acoustic, vibration, and environmental parameter raw data using a unified time base. The clock synchronizer sends a synchronization pulse signal to each sensor. Upon receiving the synchronization pulse, each sensor records the timestamp of the currently collected data. Timestamp accuracy reaches milliseconds, eliminating latency differences between different sensors. Each data point in the synchronized timestamp data carries a unified time identifier.
[0045] Data alignment processing arranges and matches multi-source sensor data with timestamps in a time series. The alignment algorithm merges data from different sensors at the same time into the same time slice based on the timestamp information. When data from multiple sensors exists at a certain point in time, the algorithm combines these data into a composite data record containing acoustic, vibration, and environmental parameters. When specific sensor data is missing at a certain point in time, the algorithm uses time interpolation to fill in the missing values. Each time slice of the multi-source noise dataset contains acoustic, vibration, and environmental information, forming a time-synchronized multidimensional sensor data array.
[0046] In a specific embodiment, the extraction module 102 includes:
[0047] A transform unit is used to perform short-time Fourier transform processing on the acoustic data in the multi-source noise data set to obtain an acoustic spectrum matrix;
[0048] A decomposition unit is used to perform a noise-specific singular value decomposition operation based on the acoustic spectrum matrix to obtain a frequency domain eigenvector matrix, a singular value diagonal matrix, and a time domain eigenvector matrix;
[0049] The reconstruction unit is used to perform weighted reconstruction on the singular value diagonal matrix according to the weight coefficients of traffic noise, construction noise, and industrial noise to obtain an enhanced spectrum feature matrix related to the noise type. The weighted reconstruction is achieved by the following formula:
[0050]
[0051] in represents the enhanced spectrum feature matrix, represents the number of singular values retained, 、 、 Represent the basic weight coefficients of traffic noise, construction noise and industrial noise respectively, 、 、 denotes the specific weights of the j-th singular value component to the three noise types, represents the jth singular value, represents the jth left singular vector, represents the transpose of the jth right singular vector;
[0052] A dimension reduction unit is used to perform feature dimension reduction processing on the vibration data in the multi-source noise data set through principal component analysis to obtain a vibration frequency domain feature vector;
[0053] The splicing unit is used to perform matrix splicing and combination processing based on the enhanced spectrum feature matrix, the time domain feature vector matrix, the vibration frequency domain feature vector and the environmental parameters to obtain a noise event feature matrix including the frequency domain feature vector, the time domain feature vector and the environmental parameter vector.
[0054] Specifically, the transformation unit in the extraction module 102 converts the acoustic data in the multi-source noise dataset into a frequency domain representation through a short-time Fourier transform. The short-time Fourier transform algorithm divides the continuous time-domain acoustic signal into multiple overlapping time windows. The signal in each time window is regarded as a quasi-static signal and is Fourier transformed. During the transformation process, the length of the time window determines the balance between time and frequency resolution. A shorter time window provides better time resolution but lower frequency resolution, while a longer time window provides better frequency resolution but lower time resolution. The rows of the acoustic spectrum matrix correspond to frequency components, and the columns correspond to time window sequences. Each element in the matrix represents the complex amplitude and phase information of a specific frequency in a specific time window. The decomposition unit performs noise-specific singular value decomposition operations based on the acoustic spectrum matrix. The algorithm reconstructs the spectrum matrix into a standard matrix format with frequency components as rows and time windows as columns, and then calculates the covariance of the matrix and performs eigenvalue decomposition. The decomposition process produces a left singular vector matrix describing the frequency domain spatial basis, a singular value sequence representing the energy distribution, and a right singular vector matrix describing the time domain spatial basis. Each column of the left singular vector matrix represents a frequency pattern, and each column of the right singular vector matrix represents a time pattern. The size of the singular value sequence reflects the contribution of the corresponding pattern to the original spectrum matrix.
[0055] The reconstruction unit performs weighted reconstruction on the singular value diagonal matrix according to the weight coefficients of traffic noise, construction noise, and industrial noise. The reconstruction process is implemented by the following formula:
[0056]
[0057] in represents the enhanced spectrum feature matrix, represents the number of singular values retained, 、 、 Represent the basic weight coefficients of traffic noise, construction noise and industrial noise respectively, 、 、 denotes the specific weights of the j-th singular value component to the three noise types, represents the jth singular value, represents the jth left singular vector, Represents the transpose of the jth right singular vector. Through this weighted reconstruction mechanism, the algorithm can highlight spectral features related to specific noise types and suppress irrelevant components. The dimensionality reduction unit inputs the vibration data from the multi-source noise dataset into the principal component analysis algorithm for feature dimensionality reduction. The principal component analysis first calculates the covariance matrix of the vibration data and then solves the eigenvalues and eigenvectors of the covariance matrix. The size of the eigenvalue represents the variance contribution of the corresponding principal component, and the eigenvector represents the direction of the principal component. The algorithm selects the first few principal components in descending order of eigenvalue. These principal components contain the main change patterns in the vibration data. The vibration frequency domain eigenvector is composed of a linear combination of the selected principal components, which retains the key characteristic information of the vibration signal while reducing the data dimension.
[0058] The splicing unit performs matrix splicing and combination processing on the enhanced spectrum feature matrix, time domain feature vector matrix, vibration frequency domain feature vector and environmental parameters in a predefined dimension order. The splicing process first ensures that each feature matrix and vector has the same time dimension, that is, each time point has a corresponding eigenvalue. Then the frequency domain feature vector is used as the first part, the time domain feature vector is used as the second part, the vibration frequency domain feature vector is used as the third part, and the environmental parameter vector is used as the fourth part. The splicing operation is performed in the column direction. Each row of the noise event feature matrix corresponds to a time slice, and each column corresponds to a feature dimension. The matrix contains acoustic frequency domain information, time domain evolution information, vibration mode information and environmental background information.
[0059] Taking industrial park noise monitoring as an example to illustrate the relevance of data processing, when the monitoring station detects complex noise including the sound of mechanical equipment operation, vehicle transportation, and construction work, the transformation unit performs short-time Fourier transform on the acoustic data to convert the time domain mixed signal into a time-frequency spectrum representation. The low-frequency band in the spectrum matrix mainly reflects the vehicle engine noise, the mid-frequency band contains the impact sound of construction equipment, and the high-frequency band shows the operation noise of mechanical equipment. The decomposition unit performs singular value decomposition on the spectrum matrix. The left singular vector matrix identifies the patterns in different frequency ranges, and the right singular vector matrix extracts the activity patterns in different time periods. The size of the singular value reflects the energy distribution of various noise components. The reconstruction unit is based on the industrial noise weight coefficient. The higher characteristics strengthen the frequency components related to mechanical equipment and weaken the occasional traffic and construction noise components. The enhanced spectrum feature matrix highlights the main noise characteristics of the industrial park. The dimensionality reduction unit processes the synchronously collected ground vibration data. The principal component analysis extracts the main vibration modes related to the operation of heavy equipment. The vibration frequency domain feature vector captures the periodic vibration characteristics of the equipment operation. The splicing unit combines the enhanced acoustic spectrum characteristics, time domain evolution characteristics, vibration characteristics and the wind speed, humidity and temperature parameters at that time into a unified feature matrix. This matrix fully describes the noise event characteristics of the industrial park at a specific moment, contains multi-dimensional related information and establishes data associations between acoustic vibration environment parameters.
[0060] In a specific embodiment, the decomposition unit is used to:
[0061] The acoustic spectrum matrix is reconstructed in a way that the frequency components are rows and the time windows are columns, and a standardized acoustic spectrum data matrix that meets the input requirements of the singular value decomposition is obtained;
[0062] The standardized acoustic spectrum data matrix is subjected to singular value decomposition through covariance matrix eigenvalue decomposition and eigenvector orthogonalization operations to obtain the left singular vector matrix describing the frequency domain space basis, the singular value sequence describing the energy distribution and the right singular vector matrix describing the time domain space basis;
[0063] The noise spectrum feature extraction process is performed based on the frequency component weights corresponding to each column vector in the left singular vector matrix to obtain a frequency domain feature vector matrix containing the frequency characteristics of traffic noise, construction noise, and industrial noise;
[0064] The time series pattern feature extraction process is performed according to the time component weights corresponding to each column vector in the right singular vector matrix to obtain the time domain feature vector matrix reflecting the time evolution law of the noise event;
[0065] The singular value sequence is arranged in descending order of energy contribution from large to small to construct a diagonal matrix, and a singular value diagonal matrix representing the importance of each frequency-time pattern is obtained.
[0066] Specifically, the matrix reconstruction processing in the decomposition unit reorganizes the acoustic spectrum matrix according to the standard format of frequency components as rows and time windows as columns. The reconstruction process checks the data arrangement of the original spectrum matrix. When it is found that the data arrangement does not meet the input requirements of the singular value decomposition algorithm, the matrix transposition or reindexing operation is performed to ensure that the row index of the matrix corresponds to the frequency bins and the column index corresponds to the time window sequence. Each element in the reconstructed standardized acoustic spectrum data matrix represents the amplitude information of a specific frequency in a specific time window. The matrix size is determined by the frequency resolution and time resolution. The standardization processing includes amplitude normalization of the matrix elements to eliminate the scale differences between different frequency components, so that the subsequent singular value decomposition can process each frequency component equally. The core calculation process of singular value decomposition is composed of eigenvalue decomposition of the covariance matrix and orthogonalization of the eigenvectors. The algorithm first calculates the product of the normalized acoustic spectrum data matrix and its transpose to obtain the covariance matrix that reflects the correlation between frequencies. The covariance matrix is then eigenvalue decomposition is performed to solve the eigenvalues and corresponding eigenvectors. The size of the eigenvalue indicates the importance of the corresponding eigenvector. The eigenvectors are orthogonalized to ensure that the vectors are perpendicular to each other. The orthogonalization process uses the Gram-Schmidt algorithm to gradually construct an orthogonal vector group. The left singular vector matrix is composed of the eigenvectors of the frequency domain covariance matrix, and the right singular vector matrix is composed of the eigenvectors of the time domain covariance matrix. The singular value sequence is composed of the square roots of the eigenvalues. These three components fully describe the decomposition results of the original spectrum matrix. The noise spectrum feature extraction process is based on the frequency component weights corresponding to each column vector in the left singular vector matrix to identify the frequency characteristics of different types of noise. Each left singular vector represents a frequency pattern, and the numerical value of each element in the vector reflects the contribution of different frequencies to the pattern. The frequency characteristics of traffic noise are mainly concentrated in the medium and low frequency bands, and the corresponding left singular vectors have larger weight values at these frequency positions. The frequency characteristics of construction noise are manifested as wide-band impact characteristics, and the corresponding left singular vectors have significant weights in multiple frequency bands. The frequency characteristics of industrial noise show periodic and harmonic structures, and the corresponding left singular vectors have larger weights at the fundamental frequency and harmonic positions. By analyzing the weight distribution pattern of the left singular vectors, the algorithm can identify and extract frequency domain feature vectors related to different noise types. Each column of the frequency domain feature vector matrix corresponds to the frequency characteristic pattern of a noise type, and the numerical value in the matrix quantifies the contribution of each frequency component to the identification of a specific noise type.The temporal pattern feature extraction process captures the temporal evolution of noise events based on the time component weights corresponding to each column vector in the right singular vector matrix. The right singular vector describes the change pattern of the spectrum matrix in the time dimension. Each element in the vector corresponds to the weight of a different time window. The temporal characteristics of traffic noise are continuous and gradual, and the corresponding right singular vectors have a smooth weight change between adjacent time windows. The temporal characteristics of construction noise show suddenness and intermittency, and the corresponding right singular vectors have prominent peaks in specific time windows. The temporal characteristics of industrial noise reflect periodicity and stability, and the corresponding right singular vectors show regular periodic fluctuations. The time domain feature vector matrix records the evolution patterns of various noise events on the time axis, and the values of the matrix elements represent the importance weights of different time periods for noise event identification.
[0067] The singular value diagonal matrix construction process arranges the singular value sequence in descending order according to energy contribution, and then places the sorted singular values in the diagonal positions of the diagonal matrix, and sets the non-diagonal elements to zero. The size of the diagonal matrix matches the number of columns of the left singular vector matrix and the right singular vector matrix. The numerical value of each diagonal element in the matrix directly reflects the contribution of the corresponding frequency-time pattern to the reconstruction of the original spectrum matrix. Larger singular values correspond to the main noise pattern, and smaller singular values correspond to secondary noise components or noise interference. Through the sorting and selection of singular values, the algorithm can distinguish between the main noise source and background noise. The singular value diagonal matrix provides a quantitative importance index for subsequent weighted reconstruction. The distribution characteristics of the singular values in the matrix reflect the complexity of the noise environment and the energy distribution of the dominant noise source.
[0068] In a specific embodiment, the separation module 103 is configured to:
[0069] The noise event feature matrix is constructed as a linear mixture relationship between the observation signal and the source signal to obtain a multi-source signal mixing matrix;
[0070] Based on the attention mechanism, the contribution weight of each frequency component to different noise source types is calculated, and the attention weight distribution matrix is obtained by normalizing the query vector and the key vector through the dot product operation.
[0071] According to the attention weight distribution matrix, exponential, hyperbolic tangent, and cubic nonlinear separation functions are set for traffic noise, construction noise, and industrial noise respectively to obtain noise source-specific separation parameters.
[0072] The demixing matrix is solved by iterative optimization operation that minimizes the mutual information between source signals. When the mutual correlation coefficient of source signals is less than a threshold, the iteration is terminated to obtain independent source signal data.
[0073] The independent source signal data were statistically verified to obtain the independent source signal data of traffic noise component, construction noise component and industrial noise component.
[0074] Specifically, the linear mixing relationship construction in the separation module 103 models the noise event feature matrix as a mathematical relationship between the observation signal and the source signal. The observation signal represents the mixed noise data actually collected by the sensor, and the source signal represents various independent noise sources such as traffic noise, construction noise, and industrial noise. The linear mixing hypothesis assumes that the observed noise signal is the linear superposition result of multiple independent noise sources after passing through different propagation paths and attenuation coefficients. During the mixing process, each noise source signal will be affected by factors such as propagation distance, building obstruction, and atmospheric absorption. These influencing factors are mathematically expressed as mixing coefficients. Each element in the multi-source signal mixing matrix represents the contribution weight of a specific noise source to a specific observation point. The number of rows in the matrix is equal to the number of observation signals, and the number of columns is equal to the number of hypothetical noise sources. The values in the matrix reflect the intensity and phase relationship of each noise source during the mixing process. The introduction of the attention mechanism enables the algorithm to adaptively identify the importance of each frequency component to different noise source types. The query vector represents the frequency feature template of the noise source type that needs to be analyzed, and the key vector represents the actual characteristics of each frequency component in the noise event feature matrix. The dot product operation calculates the similarity between the query vector and each key vector. Frequency components with high similarity indicate that they have a strong indicative effect on the current noise source type. The normalization process uses the softmax function to convert the dot product result into a probability distribution to ensure that the sum of the weights of all frequency components is one. The rows of the attention weight distribution matrix correspond to different noise source types, and the columns correspond to different frequency components. The values in the matrix quantify the contribution of each frequency component to the identification of a specific noise source. The setting of noise source-specific separation parameters is based on the guidance of the attention weight distribution matrix. The algorithm selects a corresponding nonlinear separation function for each noise type. The exponential nonlinear separation function for traffic noise is suitable for processing its energy concentration and relatively gentle variation characteristics. The function parameters are adjusted according to the distribution pattern of traffic noise in the attention weight matrix. The hyperbolic tangent nonlinear separation function for construction noise can effectively handle its suddenness and high dynamic range characteristics. The saturation characteristics of the function help suppress abnormally high amplitude components in construction noise. The cubic nonlinear separation function for industrial noise is specifically designed to process its rich harmonics and strong periodicity. The nonlinear characteristics of the cubic function can enhance the ability to identify the harmonic components of industrial noise. The parameters of each separation function are optimized and adjusted according to the weight distribution of the corresponding noise source in the attention weight matrix. The parameter optimization process takes into account the spectral characteristics, time-varying characteristics and energy distribution laws of the noise source.The iterative optimization operation solves the demixing matrix by minimizing the mutual information between the source signals. Mutual information measures the degree of statistical dependence between two signals. Independent noise sources should have minimal mutual information. The optimization algorithm uses a gradient descent method to gradually adjust the elements of the demixing matrix. Each iteration calculates the mutual information of the current separation result and updates the matrix parameters. When the change in the mutual correlation coefficient of the source signals between adjacent iterations is less than a preset threshold, the algorithm is considered to have converged and the iteration is stopped. The final result of the demixing matrix represents the linear transformation relationship of the independent source signals recovered from the observed signal. Each row in the matrix corresponds to an independent noise source, and each column corresponds to an observed signal channel.
[0075] The statistical independence verification process performs a quality check on the independent source signal data obtained by separation. The verification process includes calculating the correlation coefficient, mutual information and high-order statistics between each source signal. The correlation coefficient tests linear correlation, the mutual information tests nonlinear dependence, and high-order statistics such as skewness and kurtosis test the probability distribution characteristics of the signal. Truly independent noise source signals should show independence in all these statistical indicators. The verification algorithm also checks whether the spectral characteristics of each separated signal meet the expected noise source characteristics. The traffic noise component should have a strong energy concentration in the low-frequency band, the construction noise component should show a wide-band impact characteristic, and the industrial noise component should show an obvious harmonic structure. The signal that passes the verification is marked as reliable independent source signal data. The signal that fails the verification needs to readjust the separation parameters and execute the separation process again.
[0076] In one embodiment, the construction module 104 is configured to:
[0077] Perform empirical mode decomposition on traffic noise, construction noise, and industrial noise components to obtain the intrinsic mode components and residual components of each noise source;
[0078] Based on the building distribution and ground material in urban geographic information data, the spatial attenuation factor is introduced into the intrinsic mode component to correct the process, and the corrected mode component considering the influence of urban environment is obtained;
[0079] The instantaneous amplitude and instantaneous frequency of the modified modal component are calculated by Hilbert transform, and the time-frequency domain space mapping is performed in combination with the noise propagation distance to obtain the noise energy time-frequency distribution matrix;
[0080] Based on the spatial coordinate correlation processing of the noise energy time-frequency distribution matrix and urban geographic information data, the sound ray tracing path calculation is constructed to obtain the noise propagation trajectory data;
[0081] The noise propagation trajectory data is interpolated in three dimensions to generate noise source spatial propagation characteristic data including the propagation path and impact range.
[0082] Specifically, the empirical mode decomposition process in the construction module 104 decomposes the traffic noise component, the construction noise component and the industrial noise component into time domain signals respectively. The empirical mode decomposition algorithm determines the intrinsic oscillation mode of the signal by identifying the local extreme points in the signal. The algorithm first identifies all local maxima and minima in each noise component, and then performs cubic spline interpolation on these extreme points respectively. The upper envelope connects all local maxima, and the lower envelope connects all local minima. The mean of the envelope is subtracted from the original signal to obtain the first candidate intrinsic mode component. If the candidate component meets the condition of the intrinsic mode function, that is, the difference between the number of local extreme points and the number of zero crossings is If the value of the first intrinsic mode component does not exceed one and the mean of the upper and lower envelopes tends to zero, it is determined to be the first intrinsic mode component. Otherwise, the screening process is repeated for the candidate components until the conditions are met. The first intrinsic mode component is subtracted from the original signal to obtain the residual signal. The same decomposition process is continued on the residual signal to obtain subsequent intrinsic mode components. The decomposition process continues until the residual component becomes a monotonic function or contains less than two extreme points. The intrinsic mode components decomposed from traffic noise mainly reflect the changes in engine speed and the acceleration and deceleration process of the vehicle. The intrinsic mode components of construction noise capture the impact and intermittent characteristics. The intrinsic mode components of industrial noise show the periodicity and harmonic components of equipment operation. The spatial attenuation factor correction process compensates for the environmental impact of the intrinsic mode components based on the building distribution and ground material in the urban geographic information data. The building distribution data contains the height, width, material and location coordinates of each building. The ground material data describes the acoustic absorption characteristics of road types in different areas, such as asphalt, concrete, and grass. The correction algorithm calculates the obstacles and reflection surfaces encountered by the sound wave based on the propagation path from the noise source to the monitoring point. For the propagation path blocked by buildings, the algorithm introduces the diffraction attenuation factor. The degree of attenuation depends on the height of the building and the frequency of the sound wave. For the propagation path through different ground materials, the algorithm applies the corresponding ground absorption coefficient. The absorption coefficient of hard road surface is smaller, and the absorption coefficient of soft ground is larger. The correction process applies these attenuation factors to each intrinsic mode component. The high-frequency component is more strongly attenuated, and the attenuation of the low-frequency component is relatively small. The corrected modal component reflects the amplitude and phase changes of the sound wave after propagation in the real urban environment. This correction makes the subsequent propagation analysis more in line with the actual urban acoustic environment.
[0083] The Hilbert transform calculation converts the modified modal components into analytical signals to extract instantaneous amplitude and frequency information. The Hilbert transform converts a real signal into a complex analytical signal through convolution. The real part of the analytical signal is the original modified modal component, and the imaginary part is its Hilbert transform result. The instantaneous amplitude is obtained by calculating the modulus of the analytical signal and reflects the temporal variation of the signal energy. The instantaneous frequency is obtained by calculating the time derivative of the analytical signal phase and describes the time-varying characteristics of the signal frequency component. The time-frequency domain spatial mapping process associates the instantaneous amplitude and instantaneous frequency with the noise propagation distance. The propagation distance is calculated based on the geometric relationship between the noise source location and the monitoring point location. The mapping process establishes a correspondence between the time axis, frequency axis, and spatial axis. The rows of the noise energy time-frequency distribution matrix correspond to the time sampling points, and the columns correspond to the frequency components. The values of the matrix elements represent the energy density at a specific time and frequency. At the same time, each matrix element is associated with a corresponding spatial position coordinate. This three-dimensional mapping relationship fully describes the distribution of noise energy in the time-frequency space.
[0084] The spatial coordinate association processing matches and fuses the noise energy time-frequency distribution matrix with the urban geographic information data. The association algorithm first corresponds each energy value in the time-frequency distribution matrix to a specific geographic coordinate point, and then calculates the sound line propagation path in combination with the city's three-dimensional terrain model. The sound line tracking algorithm simulates the process of sound waves starting from the noise source and passing through different propagation paths to reach various spatial positions. The tracking process takes into account the direct, reflection, diffraction and scattering phenomena of sound waves. The direct path represents the shortest propagation distance of the sound wave. The reflection path takes into account the mirror reflection of the building surface. The diffraction path processes the propagation of sound waves around the edge of obstacles. The scattering phenomenon takes into account the random scattering effect of rough surfaces on sound waves. The propagation time of each sound line is calculated based on the path length and sound speed. The propagation loss is determined based on factors such as distance attenuation, atmospheric absorption, and obstacle attenuation. The noise propagation trajectory data records the propagation path, propagation time and energy attenuation information of the sound wave energy from the source point to each receiving point.
[0085] Three-dimensional spatial interpolation processing performs spatial continuity processing on noise propagation trajectory data to generate a propagation characteristic description. The interpolation algorithm adopts a multi-point interpolation method in three-dimensional space to calculate the noise propagation characteristics of unknown locations between known propagation trajectory points. The interpolation process takes into account the spatial distance weight and the propagation path similarity weight. Known points with closer distances have a greater impact on the interpolation results, and points with similar propagation paths have higher correlation weights. The interpolation results generate noise propagation parameters for any position in three-dimensional space, including propagation time, energy attenuation, arrival angle, etc. The spatial propagation characteristic data of the noise source contains three-dimensional propagation field information, which describes the distribution pattern of noise energy in urban space, the propagation path network and the boundary of the affected range. These data provide a spatial physical basis for the positioning, identification and impact assessment of noise pollution sources.
[0086] In one embodiment, the classification module 105 is configured to:
[0087] The noise source spatial propagation characteristic data is constructed into a four-dimensional feature tensor according to the spatial coordinates, time series, and frequency components to obtain the space-time frequency domain fusion feature data;
[0088] Based on the time-space frequency domain fusion feature data, multi-dimensional feature extraction processing is performed through the time convolution layer, space convolution layer, and frequency domain convolution layer to obtain time series features, spatial distribution features, and spectrum features;
[0089] The temporal features, spatial distribution features and spectral features are cross-modal semantically fused through a multi-head attention mechanism to obtain a comprehensive semantic feature vector.
[0090] Based on the comprehensive semantic feature vector, the noise source type is discriminated and the position regression calculation is performed through the fully connected classification layer to obtain the noise pollution source type prediction result and spatial location coordinates;
[0091] Based on the noise pollution source type prediction results and spatial location coordinates combined with the impact range and duration, weight calculation is performed to obtain the pollution source responsibility weight distribution data.
[0092] Specifically, the four-dimensional feature tensor construction in the classification module 105 reorganizes the spatial propagation feature data of the noise source according to three dimensions: spatial coordinates, time series, and frequency components. The spatial coordinate dimension contains the latitude, longitude, and altitude information of all geographical locations involved in the noise propagation process and the predicted diffusion path nodes. The time series dimension records the time evolution process of the noise event from the beginning to the end and contains the predicted sampling points of the future time window. The frequency component dimension describes the energy distribution and frequency evolution trend of the noise signal at different frequencies. The first dimension of the four-dimensional feature tensor corresponds to the spatial position index, the second dimension corresponds to the time sampling points and the predicted time window, the third dimension corresponds to the frequency bins and the frequency change pattern, and the fourth dimension corresponds to the feature type such as amplitude, phase, propagation distance, evolution rate, etc. Each element in the tensor represents the noise propagation feature value and prediction parameter at a specific spatial position, specific time, and specific frequency. The spatiotemporal frequency domain fusion feature data integrates the propagation information originally scattered in different data structures into a unified tensor representation. This representation method facilitates the subsequent parallel processing of convolutional neural networks and multi-dimensional feature prediction and extraction.
[0093] Multidimensional feature extraction processing uses three specially designed convolutional layers to process features in the three dimensions of time, space, and frequency. The temporal convolution layer uses a one-dimensional convolution kernel that slides along the time axis. After training, the convolution kernel's weight parameters can identify the temporal patterns of noise events, such as suddenness, persistence, and periodicity, and predict their future evolution trends and temporal development patterns. The convolution operation weightedly combines the feature values of adjacent time points to extract local and global variation patterns in the temporal dimension and temporal prediction features. The spatial convolution layer uses a two-dimensional convolution kernel to extract features on the geographic coordinate plane. The convolution kernel can identify the spatial distribution pattern of noise, such as geometric features such as point sources, line sources, and area sources, while capturing the directionality and attenuation of noise propagation and predicting spatial diffusion trends and propagation path evolution. The frequency domain convolution layer uses a one-dimensional convolution kernel to process spectral features, identify frequency characteristics of different noise sources, such as fundamental frequency, harmonics, and bandwidth, and analyze the frequency variation trend and spectral evolution pattern over time. The three convolutional layers work in parallel to output temporal evolution feature vectors, spatial diffusion feature vectors, and spectral change feature vectors, respectively.
[0094] Cross-modal semantic fusion processing associates and integrates features from different dimensions through a multi-head attention mechanism. The multi-head attention mechanism contains multiple parallel attention heads, each of which focuses on different types of correlations and temporal dependency patterns between features. The first attention head specifically processes the association between temporal evolution features and spatial diffusion features, identifies the spatiotemporal coupling patterns of noise events such as the trajectory characteristics of mobile noise sources, and predicts future movement trends and spatial evolution paths. The second attention head processes the association between temporal evolution features and spectral change features, captures the modulation characteristics and frequency evolution laws of noise frequency over time, and the third attention head processes the association between temporal evolution features and spectral change features. The head processes the association between spatial diffusion features and spectral change features, analyzes the spectral differences at different spatial positions and the evolution trend of propagation dispersion effects. The attention calculation process uses the temporal evolution feature vector as the query vector, the spatial diffusion feature vector and the spectral change feature vector as the key-value vector. The attention weight is obtained by calculating the similarity between the query vector and the key-value vector. The weight value reflects the degree of correlation and temporal dependency between different features. The weighted sum operation fuses multiple feature vectors into a comprehensive prediction semantic feature vector, which contains the comprehensive feature information and evolution prediction information of the noise event in the three dimensions of time, space and frequency.
[0095] The noise source type prediction and location evolution regression calculation process maps the comprehensive prediction semantic feature vector into prediction classification and evolution regression results through a fully connected prediction layer. The fully connected layer contains multiple neuron nodes, each of which is connected to all elements of the input feature vector. The node activation value is calculated through weighted summation and a nonlinear activation function. The number of output nodes in the prediction classification part is equal to the number of noise source types. Each output node corresponds to the future development trend of a noise source type such as traffic noise, construction noise, and industrial noise. The softmax activation function converts the output value into a probability distribution, and the category with the highest probability is determined as the prediction result. The location evolution regression part contains three output nodes, corresponding to the time evolution trend of the longitude, latitude, and altitude coordinates of the noise source, respectively. The regression calculation uses a linear activation function to directly output the coordinate evolution value and change rate. The loss function considers both the prediction error and the evolution regression error. The network parameters are trained through the backpropagation algorithm. The noise pollution source type prediction result gives the future development confidence of various noise source types in the form of a probability distribution. The spatial location evolution coordinates give the predicted geographic location trajectory of the noise source in the form of three-dimensional coordinate change trend and evolution rate.
[0096] The calculation of pollution source responsibility weight allocation data is based on a comprehensive evaluation of the noise pollution source type prediction results and spatial position evolution coordinates combined with the impact range diffusion prediction and duration estimation. The impact range diffusion prediction predicts the geographical coverage area diffusion trend and evolution boundary of noise pollution based on the spatial position and propagation characteristic data of the noise source. The calculation process takes into account the temporal evolution and environmental change influence of factors such as the noise propagation distance, attenuation law, and shielding effect. The impact range is represented by an irregular area centered on the noise source. The regional boundary is determined by the position where the noise intensity drops to the environmental standard limit, and the dynamic diffusion and contraction trend of the boundary is predicted. The duration estimation predicts the development time, peak time and total duration of the noise event based on the time series evolution characteristic data. Evolution model, the weight estimation calculation formula comprehensively considers the basic weight of the noise source type, the spatial weight of the impact range diffusion, the time weight of the duration estimation and the amplitude weight of the noise intensity evolution. The basic weight is set according to the severity of the environmental impact of different noise source types and the future development trend. The spatial weight is related to the evolution of the area of the impact range diffusion prediction and the degree of overlap of the sensitive area. The time weight is related to the duration estimation and the sensitivity prediction of the occurrence period. The amplitude weight is related to the evolution trend and peak prediction of the degree of noise intensity exceeding the standard. The product of the four weights constitutes the total responsibility weight estimation of each noise source. The pollution source responsibility weight estimation allocation data quantifies the potential contribution of each noise source to future environmental pollution and the impact evolution trend in the form of percentage.
[0097] The above describes the noise automatic monitoring system based on multi-sensor data fusion in the embodiment of the present application. The following describes the noise automatic monitoring method based on multi-sensor data fusion in the embodiment of the present application. Figure 2 In one embodiment of the present application, an automatic noise monitoring method based on multi-sensor data fusion includes:
[0098] S201, performing time stamp synchronization calibration processing on the original noise monitoring data collected by the voiceprint sensor, vibration sensor, and meteorological sensor to obtain a multi-source noise data set;
[0099] S202, using a noise spectrum singular value decomposition algorithm to perform frequency domain feature extraction processing on the multi-source noise data set to obtain a noise event feature matrix including a frequency domain feature vector, a time domain feature vector, and an environmental parameter vector;
[0100] S203, performing signal source separation processing on the noise event feature matrix using an independent component analysis algorithm based on attention weights to obtain independent source signal data of traffic noise components, construction noise components, and industrial noise components;
[0101] S204, performing a time-frequency-spatial feature construction process on the independent source signal data through a noise propagation Hilbert transform, and generating noise source spatial propagation feature data in combination with the urban geographic information data;
[0102] S205. Perform pattern prediction using a spatiotemporal convolutional neural network based on the spatial propagation characteristic data of the noise source, and output noise pollution source type prediction results, spatial location coordinates, and pollution source responsibility weight distribution data.
[0103] above Figure 2 The automatic noise monitoring system based on multi-sensor data fusion in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The automatic noise monitoring device based on multi-sensor data fusion in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0104] Reference Figure 3 In an embodiment of the present invention, a noise automatic monitoring device based on multi-sensor data fusion is also provided. The noise automatic monitoring device based on multi-sensor data fusion can be a server, and its internal structure can be as follows: Figure 3As shown. The noise automatic monitoring device based on multi-sensor data fusion includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the computer-designed processor is used to provide computing and control capabilities. The memory of the noise automatic monitoring device based on multi-sensor data fusion includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the noise automatic monitoring device based on multi-sensor data fusion is used to store the corresponding data in this embodiment. The network interface of the noise automatic monitoring device based on multi-sensor data fusion is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0105] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the automatic noise monitoring device based on multi-sensor data fusion to which the solution of the present invention is applied.
[0106] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the automatic noise monitoring system based on multi-sensor data fusion.
[0107] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0108] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling an automatic noise monitoring device based on multi-sensor data fusion (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0109] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A noise automatic monitoring system based on multi-sensor data fusion, characterized in that: The system includes: The calibration module is used to perform time stamp synchronization calibration on the original noise monitoring data collected by the voiceprint sensor, vibration sensor, and meteorological sensor to obtain a multi-source noise dataset; An extraction module is used to perform frequency domain feature extraction processing on the multi-source noise data set using a noise spectrum singular value decomposition algorithm to obtain a noise event feature matrix including a frequency domain feature vector, a time domain feature vector, and an environmental parameter vector. The extraction module includes: A transformation unit is configured to perform short-time Fourier transform processing on the acoustic data in the multi-source noise data set to obtain an acoustic spectrum matrix; a decomposition unit is configured to perform noise-specific singular value decomposition operation based on the acoustic spectrum matrix to obtain a frequency domain eigenvector matrix, a singular value diagonal matrix, and a time domain eigenvector matrix; and a reconstruction unit is configured to perform weighted reconstruction processing on the singular value diagonal matrix according to weight coefficients of traffic noise, construction noise, and industrial noise to obtain an enhanced spectrum feature matrix related to the noise type. The weighted reconstruction is achieved by the following formula: Where Hnew represents the enhanced spectrum feature matrix, k represents the number of retained singular values, β1, β2, and β3 represent the basic weight coefficients of traffic noise, construction noise, and industrial noise respectively, and w 1j 、w 2j 、w 3j denotes the specific weight of the j-th singular value component to the three noise types, λ j represents the jth singular value, e j represents the jth left singular vector, represents the transpose of the jth right singular vector; a dimensionality reduction unit, used to perform feature dimensionality reduction processing on the vibration data in the multi-source noise data set through principal component analysis to obtain a vibration frequency domain eigenvector; a splicing unit, used to perform matrix splicing and combination processing based on the enhanced spectrum feature matrix, the time domain eigenvector matrix, the vibration frequency domain eigenvector and the environmental parameters to obtain a noise event feature matrix including the frequency domain eigenvector, the time domain eigenvector and the environmental parameter vector; A separation module is used to perform signal source separation processing on the noise event feature matrix using an independent component analysis algorithm based on attention weights to obtain independent source signal data of traffic noise components, construction noise components, and industrial noise components; A construction module is used to perform time-frequency-spatial feature construction processing on the independent source signal data through noise propagation Hilbert transform, and generate noise source spatial propagation feature data in combination with urban geographic information data; The classification module is used to perform pattern prediction through a spatiotemporal convolutional neural network based on the spatial propagation characteristic data of the noise source, and output the noise pollution source type prediction result, spatial location coordinates and pollution source responsibility weight distribution data.
2. The noise automatic monitoring system based on multi-sensor data fusion according to claim 1 is characterized in that: The calibration module is used to: Perform bandpass filtering preprocessing on the ambient acoustic signal collected by the voiceprint sensor to obtain the original acoustic data with power frequency interference filtered out; The ground vibration signal collected by the vibration sensor is pre-processed by low-pass filtering to obtain the original vibration data with high-frequency noise removed; Perform data verification and preprocessing on the wind speed, humidity, and temperature parameters collected by meteorological sensors to obtain the original data of environmental parameters within the valid range; Performing unified time reference marking processing on the acoustic raw data, the vibration raw data, and the environmental parameter raw data based on a hardware clock synchronization mechanism to obtain synchronized time stamped data; The synchronized time stamp marked data are aligned according to the time series to obtain a multi-source noise data set with a unified time reference.
3. The noise automatic monitoring system based on multi-sensor data fusion according to claim 1 is characterized in that: The decomposition unit is used to: Reconstructing the acoustic spectrum matrix in a manner in which frequency components are arranged in rows and time windows are arranged in columns to obtain a standardized acoustic spectrum data matrix that meets the input requirements of singular value decomposition; Performing singular value decomposition on the standardized acoustic spectrum data matrix by covariance matrix eigenvalue decomposition and eigenvector orthogonalization operations to obtain a left singular vector matrix describing a frequency domain spatial basis, a singular value sequence describing energy distribution, and a right singular vector matrix describing a time domain spatial basis; Perform noise spectrum feature extraction based on the frequency component weights corresponding to each column vector in the left singular vector matrix to obtain a frequency domain feature vector matrix containing the frequency characteristics of traffic noise, construction noise, and industrial noise; Performing time series pattern feature extraction processing according to the time component weights corresponding to each column vector in the right singular vector matrix to obtain a time domain feature vector matrix reflecting the temporal evolution law of the noise event; The singular value sequences are arranged in descending order of energy contribution from large to small to construct a diagonal matrix, thereby obtaining a singular value diagonal matrix representing the importance of each frequency-time mode.
4. The noise automatic monitoring system based on multi-sensor data fusion according to claim 1 is characterized in that: The separation module is used to: The noise event characteristic matrix is constructed as a linear mixed relationship between the observation signal and the source signal to obtain a multi-source signal mixing matrix; Based on the attention mechanism, the contribution weight of each frequency component to different noise source types is calculated, and the attention weight distribution matrix is obtained by normalizing the query vector and the key vector through the dot product operation. According to the attention weight distribution matrix, exponential, hyperbolic tangent, and cubic nonlinear separation functions are set for traffic noise, construction noise, and industrial noise, respectively, to obtain noise source-specific separation parameters; The demixing matrix is solved by iterative optimization operation that minimizes the mutual information between source signals. When the mutual correlation coefficient of source signals is less than a threshold, the iteration is terminated to obtain independent source signal data. Statistical independence verification processing is performed on the independent source signal data to obtain independent source signal data of traffic noise components, construction noise components and industrial noise components.
5. The noise automatic monitoring system based on multi-sensor data fusion according to claim 1 is characterized in that: The building blocks are used to: Performing empirical mode decomposition on the traffic noise component, construction noise component, and industrial noise component to obtain an intrinsic mode component and a residual component of each noise source; Based on the building distribution and ground material in the urban geographic information data, a spatial attenuation factor is introduced into the intrinsic mode component to correct the process, so as to obtain a corrected mode component that takes into account the influence of the urban environment; The instantaneous amplitude and instantaneous frequency of the modified modal component are calculated by Hilbert transform, and the time-frequency domain space mapping is performed in combination with the noise propagation distance to obtain the noise energy time-frequency distribution matrix; Perform spatial coordinate correlation processing based on the noise energy time-frequency distribution matrix and urban geographic information data, construct a sound ray tracing path calculation, and obtain noise propagation trajectory data; The noise propagation trajectory data is subjected to three-dimensional spatial interpolation processing to generate noise source spatial propagation characteristic data including a propagation path and an impact range.
6. The noise automatic monitoring system based on multi-sensor data fusion according to claim 1 is characterized in that: The classification module is used to: Constructing a four-dimensional feature tensor from the noise source spatial propagation feature data according to spatial coordinates, time series, and frequency components to obtain time-space frequency domain fusion feature data; Based on the time-space-frequency domain fusion feature data, multi-dimensional feature extraction processing is performed through a time convolution layer, a space convolution layer, and a frequency domain convolution layer to obtain time series features, spatial distribution features, and spectrum features; The temporal features, spatial distribution features and spectral features are subjected to cross-modal semantic fusion processing through a multi-head attention mechanism to obtain a comprehensive semantic feature vector; According to the comprehensive semantic feature vector, noise source type discrimination and position regression calculation processing are performed through a fully connected classification layer to obtain a noise pollution source type prediction result and spatial position coordinates; Based on the noise pollution source type prediction results and spatial position coordinates combined with the impact range and duration, weight calculation processing is performed to obtain pollution source responsibility weight distribution data.
7. A noise automatic monitoring method based on multi-sensor data fusion, characterized in that: The automatic noise monitoring system based on multi-sensor data fusion according to any one of claims 1 to 6 is implemented, and the automatic noise monitoring method based on multi-sensor data fusion comprises: The original noise monitoring data collected by the voiceprint sensor, vibration sensor, and meteorological sensor are time-stamped and synchronized to obtain a multi-source noise dataset. A noise spectrum singular value decomposition algorithm is used to perform frequency domain feature extraction processing on the multi-source noise data set to obtain a noise event feature matrix including a frequency domain feature vector, a time domain feature vector, and an environmental parameter vector. The method includes: performing short-time Fourier transform processing on the acoustic data in the multi-source noise data set to obtain an acoustic spectrum matrix; performing a noise-specific singular value decomposition operation based on the acoustic spectrum matrix to obtain a frequency domain feature vector matrix, a singular value diagonal matrix, and a time domain feature vector matrix; and performing weighted reconstruction processing on the singular value diagonal matrix according to weight coefficients of traffic noise, construction noise, and industrial noise to obtain an enhanced spectrum feature matrix related to the noise type. The weighted reconstruction is achieved by the following formula: Where Hnew represents the enhanced spectrum feature matrix, k represents the number of retained singular values, β1, β2, and β3 represent the basic weight coefficients of traffic noise, construction noise, and industrial noise respectively, and w 1j 、w 2j 、w 3j denotes the specific weight of the j-th singular value component to the three noise types, λ j represents the jth singular value, e j represents the jth left singular vector, represents the transpose of the jth right singular vector; the vibration data in the multi-source noise data set are subjected to feature dimensionality reduction processing through principal component analysis to obtain vibration frequency domain eigenvectors; matrix splicing and combination processing is performed based on the enhanced spectrum feature matrix, time domain eigenvector matrix, vibration frequency domain eigenvector and environmental parameters to obtain a noise event feature matrix including frequency domain eigenvectors, time domain eigenvectors and environmental parameter vectors; Performing signal source separation processing on the noise event feature matrix using an independent component analysis algorithm based on attention weights to obtain independent source signal data of traffic noise components, construction noise components, and industrial noise components; The independent source signal data is subjected to a noise propagation Hilbert transform to perform time-frequency-space feature construction processing, and noise source spatial propagation feature data is generated in combination with urban geographic information data; According to the spatial propagation characteristic data of the noise source, pattern prediction is performed through a spatiotemporal convolutional neural network, and the noise pollution source type prediction result, spatial location coordinates and pollution source responsibility weight distribution data are output.
Citation Information
Patent Citations
Environmental noise detection device and method
CN120101924A
Method and apparatus for multi-channel active control of noise or vibration or of multi-channel separation of a signal from a noisy environment
US5917919A