Urban noise map generation method based on machine learning

By utilizing open-source data and machine learning algorithms, combined with environmental noise acoustic modeling and multi-source data clustering regression models, urban noise maps are generated, solving the problems of high data costs and insufficient interpretability in existing technologies, and achieving low-cost, high-precision noise distribution prediction and enhanced interpretability.

CN120910178APending Publication Date: 2025-11-07DALIAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511047020.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing noise map generation methods suffer from high data collection costs, difficulty in covering complex urban spaces, especially densely built-up areas, and insufficient interpretability of existing machine learning models, resulting in a lack of targeted planning strategies.

Method used

By combining open-source data with machine learning, urban noise maps are generated through environmental noise acoustic modeling, multi-source data clustering, and regression models. Clustering and regression algorithms are used for urban area classification and noise prediction, and feature weights are revealed through the SHAP interpreter.

Benefits of technology

It achieves low-cost, high-precision prediction of urban noise spatial distribution, enhances the interpretability and applicability of the model, accurately characterizes urban spatial heterogeneity, and provides scientific noise control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910178A_ABST
    Figure CN120910178A_ABST
Patent Text Reader

Abstract

The invention discloses a city noise map generation method based on machine learning, and belongs to the technical field of city sound environment monitoring and noise treatment. According to the method, multi-source open-source data such as remote sensing images, spatial syntax and planning elements are integrated through 75m grids, an urban area is divided into five types such as a hub area and a green area through Gaussian mixture model (GMM) clustering, a noise value is predicted in combination with a random forest (RF) regression model, a feature contribution mechanism is analyzed through SHAP, and finally a noise map is generated. The actual effect of the method on noise prediction is reflected by verifying the urban area trained by the model and selecting a new urban area for testing, and the use requirements are met. According to the method, the existing city element features in the built-up area grid units are recognized, a noise prediction system suitable for a city updating scene is constructed, and accurate improvement and efficient treatment of sound environment quality in the city updating process are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of urban sound environment monitoring and noise control, in particular to a method and system for generating a noise map of an urban built-up area based on multi-source open data and machine learning algorithms, which is suitable for spatial distribution prediction, visual expression and planning decision support of traffic noise. BACKGROUND

[0002] Traffic noise has become the second largest pollution source threatening the health of urban residents. As a core tool for noise management, noise map reflects noise distribution through visualization, providing decision support for regional control and regulation formulation. Existing noise map drawing methods rely on noise prediction models, such as the FHWA model in the United States and the CNOSSOS-EU model in the European Union, which have limitations in data acquisition. Some methods are based on monitoring data, such as CN118885975B and CN116844572B, which require the deployment of noise sensors in the city, resulting in high data costs. Other methods are based on multi-source data of built-up cities, such as CN202410572844.7, which uses a generative adversarial network to draw a city noise map, but the model lacks interpretability.

[0003] In summary, the actual measurement method requires a large number of monitoring points, resulting in high data collection costs and difficulty in covering complex urban spaces, especially for noise distribution representation in densely built areas. Simulation models rely on professional data such as traffic flow and acoustic parameters, which have high input thresholds and are not adaptable to urban spatial heterogeneity (such as road network density and green space distribution), making it difficult to accurately depict local noise characteristics. Existing machine learning models have "black box" problems, making it difficult to determine the influence weight of features on noise distribution, resulting in a lack of targeted planning strategies. SUMMARY

[0004] The present application provides a noise map generation method based on multi-source open data and machine learning, which realizes low-cost, high-precision and interpretable urban noise spatial distribution prediction, and supports precise management of urban sound environment.

[0005] Technical solution of the present application:

[0006] A method for generating a city noise map based on machine learning, comprising the following steps:

[0007] S1: Obtain target city area data using an open source map platform;

[0008] S2: Perform environmental noise acoustic modeling using an environmental noise simulation software to obtain environmental noise data;

[0009] S3: Generate grid cells using the geographic information system (GIS) platform, and assign the multi-source target urban area data and environmental noise data to the grid cells;

[0010] S4: Construct a clustering model using the multi-source target urban area data of the grid cells and machine learning clustering algorithms, and train the clustering model to achieve target urban area classification, summarize the spatial form and traffic conditions of each urban area, and assign labels accordingly;

[0011] S5: Construct a noise prediction regression model using the multi-source target urban area data and regression algorithms, and predict the noise of each urban area classified in step S4, and train the respective noise prediction regression model;

[0012] S6: Integrate the noise prediction regression model fitting results of step S5 into ArcGIS to generate a noise map of the target urban area, in order to compare and verify the effect of the noise prediction regression model;

[0013] Further: Further comprising the following steps:

[0014] S7: Integrate the trained clustering model and noise prediction regression model to facilitate subsequent model verification, adjustment, and rapid generation of traffic noise maps for other urban areas;

[0015] S8: Call the SHAP interpreter (SHapley Additive exPlanations) to calculate the marginal effect (Shapley value) of each input feature on the predicted variable, i.e., the noise value, to represent the feature weight in terms of contribution, and reveal the key features of the dominant noise problem.

[0016] In the above technical solution, in step S1, the target urban area data includes remote sensing image data (normalized vegetation index NDVI, urban green space ratio UGSR, etc.), map open platform data (night light data ND, bus route density TND, urban points of interest POI, etc.), spatial syntax data (choice degree Choice, integration degree Integration), and urban planning index data (such as road length index RLF, road area density RAF, etc.).

[0017] In the above technical solution, in step S2, commercial noise prediction software Soundplan is used to construct an environmental noise acoustic model for the target urban area, and the model is verified by actual measurement data to ensure the accuracy and effectiveness of the software simulation results.

[0018] In the above technical solution, in step S3, the grid cell size is preferably 75m x 75m.

[0019] In the above technical solution, in step S4, the following steps are included.

[0020] S401: Data preprocessing: For multi-source target city area data and environmental noise data, the minimum maximum standardization is used to eliminate the dimensional difference, the recursive feature elimination method is used to screen out redundant features, and the grid search method is used to optimize the clustering model hyperparameters.

[0021] S402: Select Gaussian mixture model (GMM) as the clustering model, train the clustering model, and evaluate based on Silhouette Coefficient coefficient and Calinski-Harabasz index. The trained clustering model divides the target city area into five categories: hub area, public building area, green area, dense road network small street area and sparse road network large street area. This division of city area integrates spatial form, functional attribute and public resource distribution elements, and can effectively identify spatial heterogeneity.

[0022] In the above technical solution, in step S5, the following steps are included.

[0023] S501: Through Pearson correlation analysis of multi-source target city area data, the features with weak correlation with noise value and redundancy (such as FSI and LSI-S) are removed, and the remaining data is divided into training set and validation set.

[0024] S502: Adopting random forest (RF) regression algorithm, a noise prediction regression model is constructed for each type of city area divided in step S4, the performance of each noise prediction regression model is trained and evaluated using the training set and the validation set, the structure or hyperparameters of each noise prediction regression model is adjusted according to the demand, and the loss change in the training process is monitored to ensure that each noise prediction regression model gradually converges during training.

[0025] In the above technical solution, in step S6, based on the operation results of the clustering model and the noise prediction regression model, visualization is realized, and ArcGIS technology is used to generate a noise map of the target city area.

[0026] The machine learning city noise map generation method of the present application has the following advantages compared with the prior art:

[0027] The cost problem of noise prediction is solved by using open source data combined with a small amount of measured values, and the difficulty and cost of data acquisition are reduced; multi-source data features are effectively fused to perform spatial clustering on urban built-up areas, deeply analyze urban spatial heterogeneity, reduce the differences between each city area ignored by the regression model due to the "global average rule", and improve the fitting effect; the regression model can explain 76.3% of the noise change, improving the model's ability to predict noise levels in complex urban areas; RF regression combined with SHAP analysis enhances the model's interpretability and comprehensive evaluation ability, and can reveal the key features of different urban areas that affect noise distribution, providing a scientific basis for precise noise reduction strategies; wide applicability: the model has cross-regional applicability and can be applied to generate noise maps for new land blocks, providing support for noise control in different cities. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 is a model framework diagram;

[0029] Figure 2 is a GMM clustering result of urban areas;

[0030] Figure 3 is a noise RF regression fitting result of each type of urban area, wherein (a) is a hub area, (b) is a public building area, (c) is a green area, (d) is a dense road network small street area, and (e) is a sparse road network large street area;

[0031] Figure 4 is a visual comparison chart of software simulation and model prediction of the model training area, wherein (a) is a noise map based on model prediction, and (b) is a noise map based on software simulation;

[0032] Figure 5 is a visual comparison chart of software simulation and model prediction of the selected new urban area block, wherein (a) is a new urban area block simulated by software, and (b) is a new urban area block of the model;

[0033] Figure 6 is a SHAP weight interpretation chart of each type of urban area, wherein (a) is a hub area, (b) is a public building area, (c) is a green area, (d) is a dense road network small street area, and (e) is a sparse road network large street area. DETAILED DESCRIPTION

[0034] The specific implementation cases and the accompanying drawings will be described below. Figures 1-6 The present application is further described, but the present application is not limited to these embodiments.

[0035] The application develops a city built-up area noise "clustering-regression" model constructed by machine learning. In the embodiment of the application, the trained model is simulated cooperatively with ArcGIS by writing a program in Python, a noise prediction system suitable for urban renewal scenarios is constructed, and the system has cross-regional applicability and can be applied to generate noise maps of new plots. The specific generation process is shown in FIG. 1.

[0036] The constructed model mainly includes two functions: city division and noise prediction. In the new plot application, the feature data is preprocessed according to the unified standard, and is summarized to the same 75m grid; the model trained based on the plot data is used to construct GMM clustering of the built-up area, and the city division is completed; then, the RF regression model is constructed for each type of region to predict the noise value, and the noise map is drawn. The process test has been completed in the verification plot, and the results show that the model results are effective.

[0037] The five types of city heterogeneity expressed by GMM clustering division are clear, the SC, CHI and DBI obtained by GMM clustering effect evaluation are 0.22, 287.77 and 2.42 respectively, the performance is relatively excellent, and the feature expression of each type of city is obvious. FIG. 2 shows that the clustering results are labeled according to the city characteristics. The "traffic core area" is closely related to the city main road. The "dense road network small street area" and the "sparse road network large street area" are two types, the former has high road density and single function, and the latter has low road density, and the dominant factors of noise distribution are also different. The "public building area" has large public building size, and the noise characteristics are significantly different from other areas, which is suitable for separate division. The "green area" is highly related to the city natural resource density, and directly affects the life quality of urban residents, so a separate model is constructed to predict the noise.

[0038] The RF model is constructed for each type of division to regress and predict the noise value of each city area. FIG. 3 shows the fitting results of the regression prediction of the five types of divisions after the RF model is constructed for each type of division. The goodness of fit of the five types of city models is as follows, R 2 枢纽区 =0.731; R 2 密路网小街区 =0.690; R 2 疏路网大街区 =0.743; R 2 公建区 =0.735; R 2 绿地区 =0.917, and the weighted average goodness of fit of the model is R 2= 0.763, indicating that about 76.3% of the noise variation can be explained in the noise prediction, and the prediction effect is improved by about 49.0% compared with the data fitting without partitioning. MAP = 68.0%, MAE = 2.56dB, which proves that the model also has better performance in robustness. Considering the influence of small data size and high-precision spatial resolution on the accuracy of the model, the prediction effect of the model in this study is still significantly competitive, which verifies the effectiveness and universality of the method. The "fitting scatter plot" reflects the fitting trend and problems through the deviation at different noise levels, and it is found that the data distribution of the "public building area" model has more outliers than other urban areas. The "absolute error histogram" shows the accumulation of samples in different error ranges, among which the "sparsely road network street area" and "public building area" have more absolute errors in the high error area, and the model still has room for improvement; the rest of the urban areas are basically half-normal distribution, and the prediction ability is relatively stable.

[0039] The visualization of the noise map based on Soundplan software simulation and the prediction of the model is shown in Figure 4 , which provides an intuitive reference for model performance evaluation. From the analysis of the noise prediction results, the difference between the two is quantitatively compared based on 3dB, and the number of grid cells with a difference less than 3dB accounts for 86.5%. Further combined with the verification results of the field measurement points, the overall accuracy is 76.2%. This shows that the model exhibits high prediction accuracy and can provide scientific and effective reference for practical applications such as urban noise control and planning layout. Compared with traditional commercial software that generates noise maps based on average traffic volume on road segments, it is difficult to reflect local traffic changes in real time. The model in this study can sensitively capture the dynamic characteristics of road segments through comprehensive analysis of multiple data sources, and has more advantages in applicability and interpretability.

[0040] The trained model is integrated into two main modules of the tool: the clustering module for urban division and the regression module for noise prediction. The module mainly generates the best parameter configuration and feature importance based on the target plot model training results, and applies it to the generation of new urban plot noise maps. The visualization of the noise map based on Soundplan software simulation and the prediction of the model is shown in Figure 5 , and the difference between the two is quantitatively compared based on 3dB, and the number of grid cells with a difference less than 3dB accounts for 68.2%. Based on the model feature contribution of each urban area extracted by SHAP, it is proved that the key features affecting noise distribution in different urban areas have significant differences.

[0041] The research combines the bee chart and column chart to visualize the feature contribution of the "SHAP weight explanation chart" as Figure 6Among them, the horizontal coordinate of the bee chart shows the influence intensity of the data points on the prediction results. The farther the data points are from the zero value line, the greater the influence on the model prediction. The color of the data points represents the size of the feature value, thus revealing the influence mechanism of feature changes on the model output. The column chart part represents the weight of each feature in the global model through the absolute average SHAP value. The SHAP reveals that the key feature difference of the influence noise distribution of different urban areas is significant. In the "hub area", the weights of "8-lane distance" and "6-lane distance" are dominant, and the dominant role of the sound source distance on the noise of the developed traffic area. The "dense road network small block" has a rich road network, and the weight of "6-lane distance" is the highest. The key features of "green area" and "sparse road network large block" are "TND" and "traffic POI" respectively, reflecting the influence of public transport flow on the noise of sparse road areas. "Public building area" is affected by "6-lane distance" and "TND", which is consistent with the functional positioning of traffic-oriented.

Claims

1. A method for generating a city noise map based on machine learning, characterized in that, Comprise the following steps: S1: Obtain target city area data using an open source map platform; S2: Obtain environmental noise data by environmental noise simulation software for environmental noise acoustic modeling; S3: Generate grid cells using a geographic information system (GIS) platform, and assign city multi-source target city area data and environmental noise data to the grid; S4: Use multi-source target city area data of the grid cells and machine learning clustering algorithm to build a clustering model, and train the clustering model to achieve target city area classification, and assign labels to each city space form and traffic condition; S5: Use city multi-source target city area data and regression algorithm to build a noise prediction regression model, and predict the noise of each city area divided in step S4, and train the respective noise prediction regression model; S6: Integrate the noise prediction regression model fitting results of step S5 into ArcGIS to generate a noise map of the target city area, so as to compare and verify the effect of the noise prediction regression model. 2.The method of claim 1, wherein, In step S1, the target city area data includes remote sensing image data, map open platform data, spatial syntax data and city planning index data. 3.The method of claim 1, wherein, In step S2, commercial noise prediction software Soundplan is used to build an environmental noise acoustic model for the target city area, and the model is verified by measured data to ensure the accuracy and effectiveness of the software simulation results. 4.The method of claim 1, wherein, In step S3, the grid cell size is preferably 75m x 75m. 5.The method of claim 1, wherein, In step S4, the following steps are included: S401: Data preprocessing: For multi-source target city area data and environmental noise data, use the minimum-maximum standardization to eliminate dimensional differences, use the recursive feature elimination method to exclude redundant features, and use the grid search method to optimize the clustering model hyperparameters; S402: Select Gaussian Mixture Model (GMM) as the clustering model, train the clustering model, and evaluate it based on Silhouette Coefficient and Calinski-Harabasz index. The trained clustering model divides the target city area into five categories: hub area, public building area, green area, dense road network small street area, and sparse road network large street area. This division of city area integrates spatial form, functional attribute and public resource distribution elements, and can effectively identify spatial heterogeneity. 6.The method of claim 1, wherein, In step S5, the following steps are included: S501: Perform Pearson correlation analysis on multi-source target city area data, remove features with weak correlation with noise values and redundancy, and divide the remaining data into training set and validation set; S502: Use random forest (RF) regression algorithm to build noise prediction regression models for each type of city area divided in step S4, train and evaluate the performance of each noise prediction regression model using the training set and validation set, adjust the structure or hyperparameters of each noise prediction regression model according to the requirements, and monitor the loss change during training to ensure that each noise prediction regression model gradually converges during training.

7. The method of claim 1, wherein, In step S6, based on the clustering model and the noise prediction regression model operation results, visualization is implemented, and an ArcGIS technique is used to generate a noise map of the target urban area.

Citation Information

Patent Citations

  • A method for constructing urban noise maps based on clustering and machine learning

    CN116844572B

  • Method for drawing urban noise map by using generative adversarial network

    CN118799437A

  • A method and system for integrating noise map and automatic monitoring data

    CN118885975B