AI modeling and pollution tracing system based on regional noise big data

By using distributed gridded sensing terminals and multi-source data fusion, a physically constrained spatiotemporal evolution model of noise is constructed. Combined with edge-cloud collaborative processing, the problems of insufficient coverage and low accuracy of traditional noise monitoring are solved, realizing intelligent and precise full-process regional noise governance.

CN121963422APending Publication Date: 2026-05-01吉林省生态环境监测中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
吉林省生态环境监测中心
Filing Date
2025-12-31
Publication Date
2026-05-01

Smart Images

  • Figure CN121963422A_ABST
    Figure CN121963422A_ABST
Patent Text Reader

Abstract

The invention discloses an AI modeling and pollution tracing system based on regional noise big data, and aims to solve the problems that traditional noise monitoring coverage is fragmented, the modeling generalization ability is weak, and pollution tracing is low in efficiency. The system comprises a data acquisition module, a preprocessing module, an AI modeling module, a pollution traceability module, a decision output module and an edge-cloud cooperative processing module, noise data and associated environment data are acquired through a distributed gridding sensing terminal, after cross-modal feature fusion, an AI model integrated with an acoustic propagation physical law is input, and the noise data and the associated environment data are subjected to cross-modal feature fusion. Through combination of space-time correlation map construction and a propagation law reverse deduction technology, accurate positioning of a noise source and propagation path tracking are realized. Based on edge-cloud collaborative optimization resource scheduling and data transmission, automatic output of target protection area early warning and graded treatment suggestions is realized, a global perception-precise modeling-efficient traceability-intelligent decision full-process solution is formed, and reliable technical support is provided for scientific management and control of regional noise pollution.
Need to check novelty before this filing date? Find Prior Art

Description

AI Modeling and Pollution Source Tracing System Based on Regional Noise Big Data Technical Field

[0001] This invention relates to the fields of digital data processing, environmental monitoring and governance, and in particular to an AI modeling and pollution source tracing system based on regional noise big data. Background Technology

[0002] The ongoing urbanization process has led to increasingly prominent regional noise pollution. Multiple noise sources, including traffic, industrial production, and human activities, are intertwined and superimposed, and their spatial and temporal distribution is significantly influenced by environmental factors such as geography and meteorological conditions. This has become a key issue restricting the improvement of residents' quality of life and disrupting the ecological balance. Traditional noise monitoring relies heavily on single-point equipment for data collection, and source tracing methods are primarily based on empirical inference. This makes it difficult to achieve comprehensive regional coverage and multi-factor collaborative analysis, failing to meet the core demands of current environmental governance for precise noise monitoring and intelligent source tracing. Therefore, upgrading noise control technologies towards the integration of big data and AI is an inevitable trend.

[0003] Existing noise-related technologies have significant shortcomings: at the modeling level, they mostly adopt a purely data-driven approach, failing to incorporate the physical laws of acoustic propagation. This results in weak generalization ability of the models in complex environments and large deviations between predicted results and actual conditions. At the data processing and source tracing level, the fusion of multi-source heterogeneous data is insufficient, and source tracing relies solely on simple inferences from single data points, lacking a comprehensive consideration of spatiotemporal correlation characteristics and noise propagation laws. This leads to inaccurate source location and unclear propagation path tracing. Given this technological status quo, there is an urgent need for an integrated solution that combines multi-source data acquisition, physical constraint modeling, and spatiotemporal correlation source tracing to address the core pain points of existing technologies. Summary of the Invention

[0004] This invention provides an AI modeling and pollution source tracing system based on regional noise big data, which solves the problems of limited coverage of traditional noise monitoring, insufficient fusion of multi-source data, lack of physical constraints in modeling and weak generalization, and low accuracy of pollution source tracing.

[0005] To address the aforementioned technical problems, this invention provides an AI modeling and pollution source tracing system based on regional noise big data, comprising:

[0006] The data acquisition module is used to acquire noise data and related environmental data within the area. The related environmental data are relevant data that affect the propagation or distribution of noise, including at least one of geospatial data, meteorological data, traffic operation data, and human activity data.

[0007] The data preprocessing module, connected to the data acquisition module, is used to perform abnormal data removal and data format unification standardization processing on the noise data and related environmental data, and to extract multi-dimensional fusion features covering temporal, spatial and heterogeneous data association attributes through feature engineering technology.

[0008] The AI ​​modeling module is connected to the data preprocessing module and is used to construct a physically constrained spatiotemporal evolution model of noise based on the multi-dimensional fusion features. The model uses the physical laws related to acoustic propagation as constraints.

[0009] The pollution source tracing module is connected to the AI ​​modeling module and is used to reverse-engineer the noise distribution data output by the noise spatiotemporal evolution model through spatiotemporal correlation feature analysis and noise propagation law.

[0010] The decision output module, connected to the pollution source tracing module, is used to generate and output noise pollution control plans and risk warning information based on the pollution source, propagation path, and scope of impact.

[0011] Preferably, the data acquisition module includes a distributed, grid-deployed array of sensing terminals. The array of sensing terminals contains at least two or more sensing devices, including acoustic sensing devices, geographic information acquisition devices, meteorological monitoring devices, traffic flow monitoring devices, and human activity status acquisition devices. Each sensing terminal transmits data synchronously through an Internet of Things (IoT) communication protocol.

[0012] Preferably, the specific implementation steps of the data preprocessing module include:

[0013] The collected noise data and related environmental data are cleaned, and outlier detection algorithms are used to identify and remove abrupt outliers in the noise data. Missing data is filled in by interpolation algorithms, and redundant and duplicate data is deleted.

[0014] Then, the different types of data are standardized, the geographic coordinate data are converted into a unified projected coordinate system, and the numerical data is normalized to the same order of magnitude.

[0015] Finally, feature extraction and fusion are performed. Fourier transform is used to extract frequency domain features for noisy data, sliding window method is used to extract trend features for time series data, graph structure analysis is used to extract topological relationship features for spatial data, and cross-modal features are fused through attention mechanism to generate multi-dimensional fused feature vectors.

[0016] Preferably, the noise spatiotemporal evolution model constructed by the AI ​​modeling module is a multi-feature fusion model based on deep learning. The acoustic propagation physical laws include at least one of the following: the propagation attenuation law of noise in different media, the influence law of topography on noise propagation, and the law of obstacle reflection and diffraction. The acoustic propagation physical laws are incorporated into the model training process in the form of constraints.

[0017] Preferably, the processing steps of the pollution source tracing module include: constructing a regional noise spatiotemporal correlation map based on noise distribution data, and identifying suspected pollution source areas through map analysis; and performing reverse deduction on the suspected areas in combination with noise propagation patterns to determine the precise pollution source, propagation range, and main path.

[0018] Preferably, the noise pollution control solution generated by the decision output module includes suggestions for pollution source control, suggestions for blocking the propagation path, and suggestions for optimizing regional noise protection. The risk warning information includes warnings for excessive noise, warnings for pollution spread, and warnings for protection of key areas. The control suggestions and warning information are displayed through a visualization platform.

[0019] As a preferred option, an edge-cloud collaborative processing module is also included;

[0020] The edge-cloud collaborative processing module is used to deploy some feature extraction tasks of the data preprocessing module and real-time inference tasks of the AI ​​modeling module on edge nodes, and to deploy model training tasks and big data storage tasks on cloud nodes.

[0021] The deployed edge node tasks include partial feature extraction tasks of the data preprocessing module and real-time inference tasks of the AI ​​modeling module. The deployed cloud node tasks include model training tasks and big data storage tasks. The edge and cloud linkage is achieved through task division, resource scheduling, data transmission optimization and collaborative synchronization mechanisms.

[0022] Preferably, the IoT communication protocol includes 5G or LoRa protocol, and the data transmission process adopts data integrity verification and time synchronization mechanism.

[0023] Compared with related technologies, the AI ​​modeling and pollution source tracing system based on regional noise big data provided by this invention has the following beneficial effects:

[0024] 1. This solution achieves full coverage of noise data in the target area through the deployment of distributed grid-based sensing terminals; at the same time, it integrates heterogeneous data from multiple sources such as acoustics, traffic, topography, and meteorology, and completes cross-modal feature fusion through a preprocessing module, providing comprehensive data support for noise analysis and solving the problems of the one-sidedness and isolation of traditional monitoring.

[0025] 2. This solution incorporates the physical laws of acoustic propagation into the AI ​​modeling module and balances data fitting and physical consistency through a hybrid loss function; the pollution source tracing module combines the construction of spatiotemporal correlation maps with reverse inference of propagation laws to first identify suspected source areas and then accurately locate them, effectively solving the pain points of weak generalization ability and low source tracing accuracy of traditional technologies.

[0026] 3. This solution uses an edge-cloud collaboration module to distribute tasks through a greedy load balancing algorithm, deploying fast-response tasks such as real-time inference on edge nodes and heavy-response tasks such as model training on the cloud. Combined with fault tolerance and task migration mechanisms, it reduces data transmission and processing latency, avoids the impact of single-point failures on the system, and improves resource utilization efficiency and system stability.

[0027] In summary, through technological innovations such as full-domain perception, physical augmentation modeling, and cloud-edge collaboration, the core problems of incomplete monitoring, inaccurate prediction, and unstable operation in traditional noise control have been systematically solved. This has built a highly efficient and intelligent noise control system covering the entire process of perception, analysis, source tracing, and decision-making, providing reliable technical support for the precise control and scientific management of regional noise pollution. Attached Figure Description

[0028] Figure 1 is a block diagram of the core module principle of the AI ​​modeling and pollution source tracing system based on regional noise big data provided by the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “group,” “class,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0031] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0032] Please refer to Figure 1. The AI ​​modeling and pollution source tracing system based on regional noise big data includes:

[0033] An AI modeling and pollution source tracing system based on regional noise big data is characterized by including:

[0034] The data acquisition module is used to acquire noise data and related environmental data within the area. The related environmental data are relevant data that affect the propagation or distribution of noise, including at least one of geospatial data, meteorological data, traffic operation data, and human activity data.

[0035] Specifically, the data acquisition module deploys a grid-like array of sensing terminals within a preset target area. These terminals include acoustic sensing devices, geographic information acquisition devices, traffic flow monitoring devices, and human activity status acquisition devices.

[0036] Obtain the preset target area, and divide the target area into grids according to the preset grid side length 'a', with the number of grids... Where S is the area of ​​the target region. This indicates rounding up, and deploying a sensor terminal array at each grid node;

[0037] Acoustic sensing equipment is used to collect noise data at different times within a target area. The noise data includes noise intensity and frequency characteristic data. The noise intensity is equivalent continuous A-weighted sound level data, and the frequency characteristic data covers the 20Hz-20kHz audio frequency range. The acquisition method is to continuously sample at preset time intervals, or to extract data in segments according to preset duration after continuous acquisition. The sampling frequency is not less than twice the highest frequency of the noise signal and meets the Nyquist sampling criterion to ensure that the signal is collected without distortion.

[0038] The geographic information acquisition equipment is used to acquire geospatial data of each monitoring point. The geospatial data includes latitude and longitude based on the WGS-84 coordinate system, elevation of the geoid reference, and topographic and landform type data divided according to preset standards (the preset standards include slope grades: ≤5° is gentle, 5°-15° is gentle slope, and >15° is steep slope; the land cover type includes built-up areas, green spaces, water bodies, and roads).

[0039] Meteorological monitoring equipment is used to collect meteorological data for a target area. The meteorological data includes the average wind speed v over a set period of time. w The data collected must include the wind direction azimuth (0°-360°) measured clockwise from 0° due north, ambient temperature, and relative humidity, and must comply with the requirements of the "Specifications for Ground Meteorological Observation" (GB / T35226-2017).

[0040] Traffic flow monitoring equipment is used to acquire road traffic-related data for a target area. The road traffic-related data includes the number of vehicles passing through a unit of time (which can be preset to every minute or hour), vehicle type (classified by size as small, medium, and large, or by power type as fuel or electric), and average speed of vehicles in the interval.

[0041] Human activity status acquisition equipment is used to statistically analyze the time period distribution characteristics of human activities in a target area through video surveillance or mobile terminal data. Specifically, it collects image data through video surveillance equipment with a preset frame rate (such as 15fps) and 720P resolution, or obtains anonymized mobile terminal activity data in the area; it divides a single day of 24 hours into 24 equal time intervals of 1 hour, counts the number of valid human activity records in each interval, and generates time period activity distribution characteristic data.

[0042] Geospatial data, meteorological data, road traffic-related data, and the distribution characteristics of human activities during different time periods are recorded as associated environmental data;

[0043] Each sensing terminal integrates a unique device identifier (ID) and transmits the collected multi-source data, device identifier (ID), and timestamp information to the data processing center in real time via 5G or LoRa IoT communication protocols. During transmission, a CRC (Cyclic Redundancy Check) algorithm is used to verify data integrity; the verification formula is H... chk =CRC(D), where D is the transmitted data segment, H is the CRC(D) value, and D is the CRC(D) value. chk This is the checksum; if the receiver's checksum result matches H... chk If there is a discrepancy, a retransmission is triggered; at the same time, the network time protocol is used to calibrate the timestamps of each terminal to ensure that the time synchronization error of multi-source data does not exceed the preset error threshold; all collected data, together with device identification, timestamp, and monitoring point coordinate information, are transmitted to the data processing center in real time and stored in a directory structure of "device ID-collection time-data type";

[0044] The data preprocessing module, connected to the data acquisition module, is used to perform abnormal data removal and data format standardization on noisy data and related environmental data. It also extracts multi-dimensional fusion features covering temporal, spatial, and heterogeneous data correlation attributes through feature engineering techniques. Specific implementation steps include:

[0045] Receive noise data and associated environmental data within the target area sent by the data acquisition module;

[0046] Numerical data, such as noise intensity, number of vehicles, driving speed, wind speed, and number of valid records of human activities, are extracted from noise data and related environmental data.

[0047] The 3σ outlier detection algorithm is used to identify abrupt outliers. The mean of a single-class numerical dataset X is defined as μ and the standard deviation as σ. If a data sample x satisfies |x-μ|>3σ, it is determined to be an outlier and removed.

[0048] For missing values ​​in the above numerical data, a linear interpolation algorithm is used to fill in the missing values. The position of the missing data is denoted as x, and its adjacent valid data points are (x1, y1) and (x2, y2). The interpolation result is calculated using the following formula: Where y fill x represents the missing value after imputation. miss For missing location data; interpolate the result y fill Fill in the missing positions;

[0049] For each data sample in the noise data and associated environmental data, its device identifier, timestamp, and core data items are concatenated into a string in a fixed order, and the hash value is calculated using the SHA-256 hash algorithm; if the hash values ​​of two samples are completely identical, they are determined to be redundant duplicate data, the first data is retained and the remaining duplicates are deleted;

[0050] Next, the noise data and related environmental data are standardized. For the latitude and longitude (x, y) in the geospatial data, a preset projection transformation function is used to convert them into planar coordinates (X, Y) in a unified projection coordinate system, where X = f(lon, lat) and Y = g(lon, lat), where f and g represent transformation functions conforming to national projection standards to ensure the consistency of spatial data. For numerical data, min-max normalization is used, with the normalization formula being x' = (xx... min ) / (x max -x min ), where x represents the original data in the noise data and associated environmental data, x' is the normalized data, and x max x min These are the minimum and maximum values ​​of the corresponding data set, respectively, so that all numerical data are normalized to the [0,1] level;

[0051] Then, feature extraction and fusion are performed on the noise data and related environmental data:

[0052] Frequency domain feature extraction: For frequency domain data in noisy data, the Discrete Fourier Transform (DFT) is used to extract frequency domain features. The transform formula is as follows: Where X(j) is the j-th frequency domain feature vector, f(n) is the sampled value of the time-domain noise signal, and Nsamp Where j is the length of the sampled data, j is the frequency domain feature index, n is the index of the time domain sampling point, representing the sampling point of the nth time domain noise signal, and e is the natural constant (Euler number), a mathematical constant.

[0053] Temporal feature extraction identifies time-series data within noisy and associated environmental data, such as vehicle numbers, driving speeds, and human activity frequency. A sliding window method is used to extract trend features. With a window size of 10 and a step size of 5, multiple windows are generated by sliding the window. Each window's features are the mean and standard deviation of the data within that window, forming a temporal feature vector.

[0054] Spatial feature extraction involves constructing a graph structure G = (V, E) for each monitoring point corresponding to the planar coordinates, where V is the set of monitoring points and E is the line (edge) connecting monitoring points within a set spatial distance (e.g., spatial distance ≤ 1 km); and calculating the topological features of each node. Where C(v) is the topological feature of node v, d(v) is the number of edges (degree) of node v, and V node The set of monitoring points is composed of spatial feature vectors formed by the topological features of each node;

[0055] Cross-modal feature fusion calculates the weights of each modality feature through an attention mechanism, with the weight calculation formula being α. j =softmax(e j1 ), where e j1 The importance score for the j1th modality feature, and ∑α j1 =1, j1 is the modal feature index; finally, a multi-dimensional fused feature vector Z = α1·X(j) + α2·F is generated. t +α3·F s F t For time series feature vectors, F s For the spatial feature vector, α1, α2, and α3 represent the weight coefficients corresponding to the frequency domain feature, temporal feature vector, and spatial feature vector, respectively. By organically combining frequency domain, temporal, and spatial features, the effectiveness of modeling input is improved. After preprocessing, the multi-dimensional fused feature vector Z will be transmitted to the AI ​​modeling module in real time as input data for model training and inference.

[0056] The AI ​​modeling module, connected to the data preprocessing module, is used to construct a physically constrained spatiotemporal evolution model of noise based on multi-dimensional fusion features Z. The model embeds acoustic propagation-related physical laws as constraints to accurately characterize the distribution characteristics and trends of regional noise in both time and space. Specific implementation steps are as follows:

[0057] A CNN-LSTM fusion architecture is adopted as the basis for the data-driven model, in which the CNN layer is used to extract spatial dimension features and the LSTM layer is used to capture temporal dimension dependencies, ensuring that the model can learn the spatiotemporal correlation features of noise at the same time.

[0058] The physical laws of acoustic propagation are selected as constraints, including the attenuation laws of noise propagation in different media and the influence of topography on noise propagation. These physical laws are transformed into mathematical constraints and incorporated into the model loss function in the form of regularization terms to construct a hybrid loss function, the formula of which is L = L data +λ·L phys ,in, The mean squared error loss is given by Ntr, where Ntr is the number of samples in the training set. Let y be the model prediction value for the j2th sample. j2 Let j2 be the true value of the j2-th sample, and j2 be the sample index. Calculate the physical consistency loss using the following formula: P is the predicted propagation noise intensity, P is the theoretical value calculated according to acoustic laws, and λ is the constraint weight coefficient used to balance data fitting and physical consistency.

[0059] Then perform model training and optimization:

[0060] The training data is divided into training, validation, and test sets according to a set ratio (e.g., 7:2:1). The training set is used for model parameter learning, the validation set is used for hyperparameter tuning and overfitting monitoring during training, and the test set is used for final generalization performance evaluation after model training is completed. The selection of this ratio is based on the balance between the data sample size and the model's generalization ability, and is suitable for the multi-source data scale of this embodiment.

[0061] The training parameters are set using the Adam optimization algorithm, which is suitable for efficient optimization of non-convex objective functions and applicable to the scenario of optimizing hybrid loss functions with physical constraints in this embodiment. The learning rate is set to 0.001, the number of iterations is set to a number of rounds (e.g., 100 rounds, which is determined by the amount of data samples and the computing resource capacity), and the batch size is set to 32. This value is adapted to the memory capacity of the training hardware and ensures the stability of parameter updates and training efficiency.

[0062] The early stopping strategy calculates the mixed loss function value on the validation set in real time after each iteration during model training. If the mixed loss function value on the validation set does not show a decreasing trend within a preset number of iterations (e.g., 10 iterations), the model training is terminated. This strategy is used to avoid the model from overfitting on the training set and to ensure the model's generalization performance on unknown data.

[0063] Model evaluation: After training, two metrics are used to evaluate model performance:

[0064] The formula for calculating the mean absolute error (MAE) is as follows: Where Q is the number of samples in the test set. Let y be the predicted value for the q-th test sample. q Let q be the true value of the q-th test sample, where q is the index of the test sample; the MAE is required to be ≤3dB, and this threshold meets the actual accuracy requirements of regional noise monitoring.

[0065] Calculate the post-propagation noise intensity predicted by the model The absolute difference between the theoretical value P obtained according to the laws of acoustic propagation, i.e. The error must be ≤2dB to ensure that the model output conforms to the actual acoustic propagation law;

[0066] A well-trained noise spatiotemporal evolution model can output predicted noise intensity values ​​for the target region at any time and any spatial coordinates (x, y). Forming a spatiotemporal distribution map of noise;

[0067] The pollution source tracing module, connected to the AI ​​modeling module, is used to accurately locate the source of noise pollution and trace its propagation path based on noise distribution data output from the noise spatiotemporal evolution model, through spatiotemporal correlation feature analysis and reverse deduction of noise propagation patterns. Specific implementation steps include:

[0068] Spatiotemporal correlation graph construction: Each monitoring point is used as a node set V = {v1, v2, ..., v...} n1 Let the i-th monitoring point be denoted as v. i Where i is the index of the monitoring point, n1 is the number of nodes within the monitoring point, and its coordinates are (x i ,y i Using the spatiotemporal correlation strength of noise as the edge weight, a global noise spatiotemporal correlation graph G=(V,E) is constructed; the correlation strength w i0j0 The formula for calculating spatiotemporal similarity through fusion is: w i0j0 =α1·Sim t (i0,j0)+β1·Sim s (i0,j0), where α1 is the time weighting coefficient and β1 is the spatial weighting coefficient, which are preset according to the spatiotemporal characteristics of regional noise; For time similarity, t i0 t j0 , i and j0 are the timestamps of the noise data at nodes i and j0, respectively, and τ is the time decay coefficient, used to quantify the impact of time difference on the correlation strength; For spatial similarity, d i0j0ρ is the spatial straight-line distance between nodes i0 and j0, ρ is the spatial attenuation coefficient, adapted to the regional geographical scale; E is the edge set, retaining only the association strength w. i0j0 Edges ≥ w0 (where w0 is a preset correlation strength threshold) ensure that the graph focuses on highly correlated noisy data;

[0069] Suspected pollution source area identification: The Louvain community detection algorithm is used to perform cluster analysis on the spatiotemporal correlation graph G. Using modularity Q1 as the optimization objective, highly correlated node clusters are selected as suspected pollution source areas. The formula for calculating modularity Q1 is: Where c is the community index formed by clustering, ∑ in ∑ is the sum of the weights of all edges within the community. tot The sum of the weights of all edges (including those inside and outside the community) connected to all nodes within the community is given by , and m is the sum of the weights of all edges in the graph G. The community with the largest modularity Q1 is selected as the suspected pollution source region R, and the region boundary is determined by the circumscribed polygon of all nodes within the community.

[0070] Noise Propagation Simulation: Reusing the physical laws of acoustic propagation and modifying the propagation model based on regional environmental characteristics, the simulation demonstrates the propagation process of noise from a suspected area to its surroundings. The propagation attenuation formula is used to quantify the change in noise intensity: P = P0·e -k·r·s1 Where P is the noise intensity at a certain point after propagation, P0 is the initial noise intensity of the sound source, k is the medium attenuation coefficient (preset according to the type of propagation medium such as air and vegetation), r is the propagation distance from the sound source to the target point, and s1 is the terrain blocking coefficient, which satisfies 0≤s1≤1. For flat terrain, s1=1. The higher the terrain complexity, the smaller the value of s1, which quantifies the blocking effect of terrain on noise propagation.

[0071] Simultaneously, the propagation direction is corrected by incorporating meteorological data. The correction formula is: θ'=θ+γ wd ·v w ·sin(φ-θ), where θ is the noise propagation direction without environmental influence, θ' is the corrected actual propagation direction, and γ wd V is the wind direction influence coefficient. w φ represents wind speed, and φ represents wind direction (measured clockwise with true north as 0°).

[0072] Precise localization through reverse inference: The gradient descent reverse inference algorithm is used to determine the noise prediction values ​​of each monitoring point within the suspected area R. Compared with the actual monitored value P re,i To minimize the error, we inversely determine the initial position of the sound source. The objective function is: in, The predicted noise value for this monitoring point is output by the AI ​​modeling module, where R represents the actual noise intensity collected at the monitoring point, i represents the suspected pollution source area, and i represents the index of the monitoring point within the area. P re,i Let x0 and y0 represent the model prediction value and measured noise intensity of the i-th monitoring point, respectively. The candidate positions of the sound source are adjusted through iterative calculation until the error meets the preset accuracy requirements, and finally the precise coordinates (x0, y0) of the pollution source are determined.

[0073] Propagation range and main path determination: Centered on the precisely located sound source (x0, y0), the propagation attenuation formula is used to calculate the range and main path that satisfy P ≥ P. th (P th The set of all spatial points (for a preset noise exceedance threshold) forms the pollution propagation range S. Using Dijkstra's shortest path algorithm, with the noise intensity attenuation as the path weight (the smaller the attenuation, the lower the path weight), the optimal path from the sound source to each exceedance monitoring point is selected. This path is the main path of pollution propagation. Key nodes on the path are output simultaneously, such as terrain obstacles and noise intensity abrupt change points along the propagation route.

[0074] The decision output module, connected to the pollution source tracing module, is used to generate and output noise pollution control plans and risk warning information based on the pollution source, propagation path, and impact range. The noise pollution control plan includes suggestions for pollution source control, propagation path blocking, and regional noise protection optimization. Risk warning information includes warnings for excessive noise, pollution spread, and protection of key areas. The control suggestions and warning information are displayed through a visualization platform. Specific implementation steps are as follows:

[0075] Construct a pollution source feature vector F = (P, T, D, Freq), where P is the peak noise intensity, T is the duration of noise, D is the spatial distribution density, and Freq is the dominant frequency feature of noise.

[0076] Preset three types of standard feature template sets: industrial noise, traffic noise, and human activity noise, M = {M1, M2, M3} (M k1 =(P k1 ,T k1 D k1 ,Freq k1 () represents the typical feature thresholds for various types of noise, and k1 is the template index, taking values ​​of 1, 2, and 3, corresponding to industrial noise, traffic noise, and human activity noise, respectively; the feature matching degree is calculated using the weighted cosine similarity algorithm, with the formula as follows: Where, ω j3 For the feature weight of the j3rd term, based on the preset regional noise control priorities, F j2 M is the j3rd feature value of the source to be classified. k1j3 is the j3rd feature value of the k1th template, k1∈[1,3], corresponding to industrial / traffic / human activity noise, and j3 is the feature index (j3∈[1,4], corresponding to peak intensity / duration / spatial density / dominant frequency feature);

[0077] Select the maximum similarity value max(Sim(F,M) k1 The type corresponding to ))≥Sim0 (Sim0 is the preset matching threshold) is taken as the final source type. If the preset matching threshold is not met, it is marked as a composite noise source.

[0078] The system calls upon a pre-defined governance suggestion library to generate a candidate suggestion set G for different source types (industrial noise candidate suggestions include "optimize production processes", "install sound insulation equipment", "adjust production hours", etc.; traffic noise includes "adjust traffic hours", "optimize road design", "install sound barriers", etc.; human activity noise includes "delineate quiet areas", "restrict activity hours", "strengthen publicity and guidance", etc.).

[0079] The score for each suggestion is calculated using a priority evaluation algorithm, as shown in the formula: Where f1 and f2 represent the effect weighting coefficient and cost weighting coefficient, respectively, E b Let C be the governance effectiveness coefficient of suggestion b. b The implementation cost coefficient for suggestion b (the higher the cost, the larger the coefficient); Screening score (Score(G)) b The top 3 suggestions with scores ≥ Score0 (Score0 is the preset score threshold) are used as the final governance suggestions and output in descending order of score.

[0080] Two preset noise exceedance thresholds are established, including P1 and P2. P1 is the secondary warning threshold (e.g., 60 dB(A), corresponding to the daytime standard for Class 2 areas in the "Environmental Quality Standard for Noise" GB3096-2008), and P2 is the primary warning threshold (e.g., 70 dB(A), corresponding to the nighttime standard for Class 2 areas or the upper limit of exceedance for Class 1 areas). Combined with actual noise measurement data, P... real Compared with the AI ​​modeling module's predicted data P pred Warning triggered:

[0081] When P real >P2 or P pred When P2 is reached, a Level 1 exceedance warning (i.e., an emergency warning) is triggered and simultaneously pushed to the responsible person in the area;

[0082] When P1 <P real ≤P2 or P1 <P pred When the value is ≤P2, a level 2 exceeding warning (i.e., a reminder warning) is triggered;

[0083] When P real ≤P1 and P predWhen P1 is ≤, no over-limit warning is triggered, and only normal monitoring data is recorded;

[0084] Based on the propagation path length L and propagation time t output by the pollution source tracing module, the average diffusion velocity is calculated. The pollution coverage area is predicted within the next τ1 time period (τ1 can be preset to 1h, 2h, 3h, etc.) by the formula: S(τ1)=S0∪{(x,y)|d((x,y),O)≤v·τ1}, where S0 is the current noise exceeding area (the set of spatial points that satisfy P≥P1), O=(x0,y0) is the precise coordinate of the pollution source, and d((x,y),O) is the spatial straight-line distance between the target point (x,y) and the source O;

[0085] Simultaneously calculate the target protection area (x) g ,y g The warning time window: Where t w This indicates that the protected area needs to be within t w Take protective measures within a short period of time (such as closing windows, setting up temporary sound barriers, etc.);

[0086] Target protection area (x) g ,y g The generation logic of ) is as follows:

[0087] A list of noise-sensitive areas within the target area (including residential areas, schools, hospitals, and concentrated office areas, etc., and recording the coordinate range of each sensitive area) is pre-entered into the system to form a sensitive area set Sd; based on the current noise-exceeding area S0 output by the pollution source tracing module and the pollution diffusion range Sd(τ1) predicted by the decision output module within the future τ1 time period, sensitive areas that spatially overlap with Sd(τ1) or are ≤ a preset safety distance (e.g., 50 meters) from the boundary of Sd(τ1) are selected; the selected sensitive areas are merged with the areas within the current noise-exceeding area that have not yet implemented noise protection measures to form a target protection area set {(x g ,y g )|g=1,2,...,Gd}(Gd is the total number of target protection areas, (x g ,y g (where ) represents the center coordinates of the g-th target protection area;

[0088] A layer overlay algorithm is used to construct independent vector data layers for pollution source coordinates (labeled by type and intensity), propagation path (labeled by attenuation nodes), areas exceeding standards (yellow for level II warning, red for level I warning), diffusion warning range (orange gradient label), and remediation recommendation labels (sorted by priority). The layer display priority is optimized by adjusting the transparency formula: Trans(I) = 0.5 + 0.4·I, where I is the layer importance coefficient (0 < I ≤ 1, e.g., source coordinates and warning area I = 1, propagation path I = 0.8, remediation recommendation I = 0.7), and Trans(I) is the layer transparency (the larger the coefficient, the clearer the display). An integrated interactive map of "source-path-warning-recommendation" is generated on the visualization platform, and a text-based decision report (including data tables, formula calculation process, and implementation priority) is output simultaneously and pushed to the terminal devices of relevant management departments through the API interface.

[0089] The present invention also includes an edge-cloud collaborative processing module;

[0090] The edge-cloud collaborative processing module is used to deploy some feature extraction tasks of the data preprocessing module and real-time inference tasks of the AI ​​modeling module on edge nodes, and to deploy model training tasks and big data storage tasks on cloud nodes, so as to realize the collaborative linkage between edge computing and cloud computing.

[0091] The deployed edge node tasks include some feature extraction tasks of the data preprocessing module and real-time inference tasks of the AI ​​modeling module. The deployed cloud node tasks include model training tasks and big data storage tasks. The edge and cloud linkage is achieved through task division, resource scheduling, data transmission optimization and collaborative synchronization mechanisms.

[0092] The edge-cloud collaborative processing module is specifically used for:

[0093] The collaborative linkage between edge computing and cloud computing is achieved through task division, resource scheduling, data transmission optimization and collaborative synchronization mechanism. Specifically, the task division involves deploying the preliminary extraction of time series features, the preliminary detection of outliers in the data preprocessing module and the short-term noise prediction task (prediction duration not exceeding 1 hour) for a single monitoring point in the AI ​​modeling module to edge nodes, and deploying the model training task, the full-area noise spatiotemporal evolution simulation task and the historical big data storage task with a storage duration of not less than 3 months to cloud nodes.

[0094] Resource scheduling adopts a greedy load balancing algorithm, which matches and allocates resources based on the remaining computing resources of each edge node and the resource requirements of the edge tasks to be assigned, so as to minimize the load difference of all edge nodes and avoid processing delays caused by overload of a single node.

[0095] During data transmission optimization, the multi-dimensional fusion feature vector generated by the data preprocessing module transmitted from the edge node to the cloud adopts the LZ77 data compression algorithm to reduce the transmission bandwidth usage. The cloud sends model parameters to the edge node using an incremental update strategy, transmitting only the part that differs from the current model parameters of the edge node to reduce transmission time.

[0096] The collaborative synchronization adopts a heartbeat packet combined with incremental synchronization mode. Every 5 minutes, the edge node sends a heartbeat packet to the cloud containing data on the node's online status, current task execution progress, and resource utilization. Every hour, the edge task processing results are synchronized to the cloud. Every 24 hours, the cloud sends model parameter update packets to the edge node. During the synchronization process, the MD5 checksum algorithm is used to verify data integrity. If the verification result at the receiving end is inconsistent with the verification value at the sending end, resynchronization is triggered.

[0097] At the same time, a fault tolerance mechanism is set up. If the cloud fails to receive a heartbeat packet from an edge node three times in a row (with a cumulative duration of not less than 15 minutes), the node is judged to be faulty. Then, the task migration mechanism is started to redistribute the unfinished tasks of the faulty node to other normal edge nodes according to the greedy load balancing algorithm mentioned above, ensuring that the task interruption time during the migration process does not exceed 30 seconds, and ensuring the continuous and stable operation of the system.

[0098] It should be noted that the formulas for noise spatiotemporal evolution models, pollution source tracing, and early warning time window calculations in this system all use dimensionless values ​​for computation. Dimensionlessness can be achieved through conventional methods such as standardization, which will not be elaborated here. The relevant formulas are obtained through software simulation and calibration using massive amounts of regional noise data and supporting multi-source correlated data (geographical, meteorological, etc.), ensuring maximum approximation to real-world scenarios. The preset parameters in the formulas can be flexibly configured by those skilled in the art based on the actual situation, such as the monitoring area, noise type, and treatment needs.

[0099] This system embodiment can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. The hardware can employ distributed noise sensing terminals, edge computing nodes, and cloud servers; the core software includes data preprocessing algorithms, AI modeling programs, and pollution source tracing logic modules.

[0100] When implemented as software, this system can be represented as a computer program product, containing several computer instructions. When these instructions are loaded or executed, they will generate the system's core processes (data acquisition - feature fusion - model inference - source tracing decision-making) and corresponding functions. The instructions can be stored on computer-readable storage media or transmitted between websites, servers, and data centers via wired or wireless means.

[0101] It should be clarified that the sequence number of each processing step in the system does not represent the order of execution. The order is determined by the functional logic of "noise data input - model calculation - result output" and does not constitute an implementation limitation. Those skilled in the art can choose to implement the functions of modules such as data acquisition units and algorithm processing units using electronic hardware or a combination of hardware and software, depending on the application scenario and design constraints. Such implementations are all within the scope of this application.

[0102] The unit division of the system device is only a logical functional division, and can be flexibly adjusted in actual implementation. For example, data preprocessing can be integrated with the edge computing unit, or unnecessary auxiliary function modules can be omitted. Each unit can be physically separated or centrally deployed, and some or all units can be selected to achieve the solution objectives according to the requirements; functional units can also be integrated into a single processing unit or maintain independent physical existence.

[0103] If the core functions of the system are sold or used independently as software units, they may be stored on computer-readable storage media. This software product contains several instructions used to drive computer devices (personal computers, servers, network devices, etc.) to execute all or part of the system's steps. Common storage media include USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, and optical disks.

[0104] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0105] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. An AI modeling and pollution source tracing system based on regional noise big data, characterized in that, include: The data acquisition module is used to acquire noise data and related environmental data within the area. The related environmental data are relevant data that affect the propagation or distribution of noise, including at least one of geospatial data, meteorological data, traffic operation data, and human activity data. The data preprocessing module is connected to the data acquisition module and is used to perform abnormal data removal, data format unification and standardization processing on the noise data and related environmental data, and to extract multi-dimensional fusion features covering temporal, spatial and heterogeneous data correlation attributes through feature engineering technology. The AI ​​modeling module, connected to the data preprocessing module, is used to construct a physically constrained spatiotemporal evolution model of noise based on the multi-dimensional fusion features. The model uses embedded physical laws related to acoustic propagation as constraints. The pollution source tracing module, connected to the AI ​​modeling module, is used to reverse-engineer the noise distribution data output by the noise spatiotemporal evolution model through spatiotemporal correlation feature analysis and noise propagation laws. The decision output module, connected to the pollution source tracing module, is used to generate and output noise pollution control plans and risk warning information based on the pollution source, propagation path, and scope of impact.

2. The AI ​​modeling and pollution source tracing system based on regional noise big data as described in claim 1, characterized in that, The data acquisition module includes a distributed, grid-deployed array of sensing terminals. The array contains at least two or more sensing devices, including acoustic sensing devices, geographic information acquisition devices, meteorological monitoring devices, traffic flow monitoring devices, and human activity status acquisition devices. Each sensing terminal transmits data synchronously via an Internet of Things (IoT) communication protocol.

3. The AI ​​modeling and pollution source tracing system based on regional noise big data as described in claim 1, characterized in that, The specific implementation steps of the data preprocessing module include: cleaning the collected noise data and related environmental data; using an outlier detection algorithm to identify and remove abrupt outliers in the noise data; filling in missing data using an interpolation algorithm; and deleting redundant and duplicate data; then standardizing different types of data, converting geographic coordinate data into a unified projected coordinate system, and normalizing numerical data to the same magnitude; finally, performing feature extraction and fusion, using Fourier transform to extract frequency domain features from noise data, using the sliding window method to extract trend features from time-series data, using graph structure analysis to extract topological relationship features from spatial data, and fusing cross-modal features through an attention mechanism to generate a multi-dimensional fused feature vector.

4. The AI ​​modeling and pollution source tracing system based on regional noise big data according to claim 1, characterized in that, The noise spatiotemporal evolution model constructed by the AI ​​modeling module is specifically a multi-feature fusion model based on deep learning. The acoustic propagation physical laws include at least one of the following: the propagation attenuation law of noise in different media, the influence law of topography on noise propagation, and the law of obstacle reflection and diffraction. The acoustic propagation physical laws are incorporated into the model training process in the form of constraints.

5. The AI ​​modeling and pollution source tracing system based on regional noise big data according to claim 1, characterized in that, The processing steps of the pollution source tracing module include: constructing a regional noise spatiotemporal correlation map based on noise distribution data, and identifying suspected pollution source areas through map analysis; and performing reverse deduction on the suspected areas in combination with noise propagation patterns to determine the precise pollution source, propagation range, and main path.

6. The AI ​​modeling and pollution source tracing system based on regional noise big data according to claim 1, characterized in that, The decision output module generates noise pollution control solutions including suggestions for pollution source control, suggestions for blocking transmission paths, and suggestions for optimizing regional noise protection. The risk warning information includes warnings for excessive noise, pollution spread, and protection of key areas. The control suggestions and warning information are displayed through a visualization platform.

7. The AI ​​modeling and pollution source tracing system based on regional noise big data according to claim 1, characterized in that, It also includes an edge-cloud collaborative processing module; the edge-cloud collaborative processing module is used to deploy some feature extraction tasks of the data preprocessing module and real-time inference tasks of the AI ​​modeling module on edge nodes, and to deploy model training tasks and big data storage tasks on cloud nodes; wherein, the deployed edge node tasks include some feature extraction tasks of the data preprocessing module and real-time inference tasks of the AI ​​modeling module, and the deployed cloud node tasks include model training tasks and big data storage tasks, and the edge and cloud linkage is realized through task division, resource scheduling, data transmission optimization and collaborative synchronization mechanism.

8. The AI ​​modeling and pollution source tracing system based on regional noise big data according to claim 2, characterized in that, The IoT communication protocol includes 5G or LoRa protocol, and the data transmission process adopts data integrity verification and time synchronization mechanism.