Spatial distribution prediction method, device and equipment for heavy metal emission intensity of nonferrous metal enterprise and medium
Through the geo-weighted random forest model combined with multiple characteristic data, the inefficiency problem of spatial distribution of heavy metal emission intensity research in non-ferrous metal enterprises is solved, high-precision prediction and environmental risk assessment are achieved, and environmental governance and enterprise emission reduction are supported.
Patent Information
- Application Number
- CN202510366858.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises relies on field sampling, resulting in low efficiency and high cost, making it difficult to achieve efficient and accurate spatial analysis.
The geo-weighted random forest model is used to combine geospatial characteristics, non-ferrous metal enterprise production characteristics and environmental management characteristic data, and the heavy metal emission intensity is predicted through grid division and data preprocessing, and the random forest model is adjusted to adaptively capture regional feature differences.
It improves the prediction accuracy and efficiency of heavy metal emission intensity, provides more comprehensive data support, assists in environmental supervision and enterprise emission reduction planning.
Smart Images

Figure CN120298002A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of environmental science and technology, and particularly to a method, device, electronic device and computer-readable storage medium for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises. Background Art
[0002] The heavy metal emissions of non-ferrous metal enterprises are the main sources of heavy metal pollution in environmental media. Accurately mastering the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises in a region provides data support for the comprehensive management of the regional environment and has important practical significance for preventing and controlling the environmental governance of non-ferrous metal enterprises.
[0003] Currently, the research on the heavy metal emission intensity of non-ferrous metal enterprises mainly relies on on-site sampling and analysis, which requires a large amount of manpower, material resources and financial resources, resulting in low efficiency of spatial analysis of heavy metal emission intensity of non-ferrous metal enterprises. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, electronic device and computer-readable storage medium for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises, which are used to improve the efficiency and accuracy of predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises.
[0005] In the first aspect, the embodiments of the present application provide a method for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises, including:
[0006] Dividing a target area into grids, and obtaining the geospatial feature data, production feature data of non-ferrous metal enterprises and environmental management feature data within each grid to obtain the basic data of each grid;
[0007] Inputting the basic data of each grid into a pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
[0008] In the second aspect, the embodiments of the present application provide a device for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises, including:
[0009] An acquisition module, configured to divide a target area into grids, and obtain the geospatial feature data, production feature data of non-ferrous metal enterprises and environmental management feature data within each grid to obtain the basic data of each grid;
[0010] A processing module, configured to input the basic data of each grid into a pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
[0011] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for predicting the spatial distribution of the heavy metal emission intensity of non-ferrous metal enterprises provided in the first aspect of the embodiment of the present application are implemented.
[0012] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for predicting the spatial distribution of the heavy metal emission intensity of non-ferrous metal enterprises provided in the first aspect of the embodiment of the present application are implemented.
[0013] The technical solution provided by the embodiment of the present application obtains the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data in each grid of the target area, obtains the basic data of each grid, and inputs the basic data of each grid into the pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid. By coupling the geographically weighted model into the random forest model, the geographically weighted random forest model can fully consider the diffusion effects of various factors such as the geographical features, spatial relationships, enterprise production features, and enterprise management features of different grids on the heavy metal emission intensity, thereby improving the prediction accuracy and prediction efficiency. Description of the Drawings
[0014] Figure 1 It is a schematic flowchart of a method for predicting the spatial distribution of the heavy metal emission intensity of non-ferrous metal enterprises provided by the embodiment of the present application;
[0015] Figure 2 It is a schematic flowchart of a heavy metal environmental risk prediction process provided by the embodiment of the present application;
[0016] Figure 3 It is a schematic flowchart of a geographically weighted random forest model training process provided by the embodiment of the present application;
[0017] Figure 4 It is a schematic structural diagram of a device for predicting the spatial distribution of the heavy metal emission intensity of non-ferrous metal enterprises provided by the embodiment of the present application;
[0018] Figure 5 It is a schematic structural diagram of an electronic device provided by the embodiment of the present application. Detailed Embodiments
[0019] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the technical solutions in the embodiments of the present application will be further described in detail through the following embodiments in combination with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Those skilled in the art can adjust it as needed to suit the specific application scenarios. Additionally, it should be noted that for the convenience of description, only the parts related to the present application rather than all the structures are shown in the drawings.
[0020] Figure 1 This is a schematic flowchart of a method for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises provided in an embodiment of the present application. As Figure 1 shown, the method may include:
[0021] S101. Divide the target area into grids, and obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data within each grid to obtain the basic data of each grid.
[0022] Among them, the target area is the area where the spatial distribution of heavy metal emission intensity needs to be predicted, which can be an aggregation area of non-ferrous metal enterprises. Specifically, the geographical boundary of the target area can be determined through a geographic information system or administrative division data, and according to the specified grid size and the geographical boundary of the target area, the target area is divided into grids by using the corresponding grid division tool. The above-mentioned specified grid size can be set according to the actual data accuracy requirements. For example, in the case of high data accuracy requirements, a smaller grid size can be set, and when conducting macroscopic analysis, a larger grid size can be set. By dividing the target area into multiple grids, based on grid analysis, it provides a basis for the subsequent prediction of the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises.
[0023] For each grid, obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data within the grid. Among them, the geospatial feature data is used to describe the natural geography and environmental conditions within the grid, and these natural environmental conditions have an important impact on the emission, diffusion, and accumulation of heavy metals. Optionally, the geospatial feature data may include at least one of the following: land use type (such as rural settlements, cultivated land, and urban areas), terrain (such as altitude, slope, and elevation), vegetation index, river information (such as river density and distance from the river), climate information (such as rainfall, wind speed, and wind direction), socioeconomic information (such as population density, gross domestic product, road network density), and enterprise geographical location information (such as longitude and latitude). The above geospatial feature data can be obtained through meteorological stations, remote sensing technology, and resource and environmental science data platforms, etc. For example, obtain the detection data of meteorological stations, focus on screening climate data such as rainfall, wind speed, and wind direction from the detection data, and perform interpolation processing on the screened climate data through spline functions, etc. Another example is to obtain the current land use type, geographical element data, etc. of each grid through remote sensing technology, and use some distance analysis tools to obtain the distances from each grid to roads, rivers, and settlements, etc. Another example is to obtain data such as gross domestic product and population density through the resource and environmental science data platform.
[0024] The above production feature data is used to describe the production activities and processes of non-ferrous metal enterprises, and these features directly affect the emission intensity of heavy metals. Optionally, the above production feature data includes at least one of the following: enterprise raw material features and energy consumption features. Among them, the enterprise raw material features can include the types and sources of raw materials, and the energy consumption features can include the types of energy consumption (such as coal, natural gas, electricity) and energy utilization efficiency.
[0025] The above environmental management feature data is used to describe the environmental management and treatment measures of non-ferrous metal enterprises, and these measures can reduce the emission of heavy metals. Optionally, the above environmental management feature data includes at least one of the following: enterprise illegal records and industrial pollutant treatment efficiency. Among them, the enterprise illegal records refer to the illegal records of enterprises in environmental management, such as illegal emissions and non-compliant emissions, etc. Industrial pollutants mainly include industrial three wastes, such as waste gas, waste water, and waste residue. The above production feature data and environmental management feature data can be obtained based on the environmental protection data platform.
[0026] S102. Input the basic data of each grid into the pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
[0027] Among them, the geographically weighted random forest model is an extension of the random forest in the field of geography. The geographically weighted random forest model embeds and couples geographically weighted regression into the random forest, and adjusts the sample sampling and node splitting rules of the random forest through the geographically weighted mechanism, enabling the geographically weighted random forest model to adaptively capture the differences in characteristic relationships in different regions and realize the spatial distribution prediction of the heavy metal emission intensity of non-ferrous metal enterprises.
[0028] Specifically, after obtaining the basic data of each grid, the basic data of each grid can be input into the pre-trained geographically weighted random forest model. The geographically weighted random forest model generates a spatial weight matrix based on the spatial information (or geographical location information) of non-ferrous metal enterprises in each grid, and uses this spatial weight matrix and the basic data of each grid to make predictions on each decision tree, obtaining the local heavy metal emission intensity prediction values of each decision tree. Then, the local heavy metal emission intensity prediction values of all decision trees are weighted averaged or voted to obtain the final predicted heavy metal emission intensity of each grid. Among them, the simplified formula of the geographically weighted random forest model can be:
[0029] PD i =A (ui,vi) X i +ε;
[0030] Among them, PD i represents the heavy metal emission intensity of the i-th grid, A (ui,vi) X i is the random forest prediction value (implied spatial weight) calibrated on the i-th grid, (ui, vi) is the centroid coordinate of the i-th grid, X i is the basic data of the i-th grid (such as geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data), ε is the error term, i = 1, 2,... n, and n is the number of grids.
[0031] After obtaining the predicted heavy metal emission intensity of each grid, a heatmap or other visualization methods can be used to display the predicted heavy metal emission intensity of each grid, providing a basis for environmental supervision and enterprise emission reduction planning.
[0032] Optionally, before inputting the basic data of each grid into the pre-trained geographically weighted random forest model, the basic data of each grid can also be preprocessed. Among them, the preprocessing includes at least one of the following: data cleaning, data standardization, and encoding and classifying data.
[0033] Data cleaning: Correct or eliminate data that obviously does not conform to the actual situation. At the same time, check the integrity of the geospatial feature data, fill in missing data through spatial interpolation; remove redundant information; for those with inconsistent coordinate systems, unify the spatio-temporal benchmarks.
[0034] Data standardization: Unify the scales of data with different dimensions, perform a linear transformation on the basic data using deviation standardization, and map the results to the interval [0, 1]. For example, the following formula can be used for data standardization processing:
[0035]
[0036] In the formula, x i is a value in the basic data, x min represents the minimum value of the basic data, x max represents the maximum value of the basic data, and y is the value of the standardized data.
[0037] Encoding categorical data: For categorical variables, encoding processing operations can be performed on the categorical variables to convert them into binary vectors. For example, through encoding processing, categorical variables such as land use types and production processes are converted into binary vectors.
[0038] It should be noted that after dividing the target area into grids, some grids contain non-ferrous metal enterprises, and some grids do not contain non-ferrous metal enterprises. For grids that do not contain non-ferrous metal enterprises, when using the geographically weighted random forest model for prediction, the production feature data and environmental management feature data in the basic data of the grid can be set to default values. For example, the default value can be 0.
[0039] The spatial distribution prediction method for the heavy metal emission intensity of non-ferrous metal enterprises provided by the embodiments of the present application obtains the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data in each grid of the target area, obtains the basic data of each grid, inputs the basic data of each grid into the pre-trained geographically weighted random forest model, and obtains the predicted heavy metal emission intensity of each grid. By coupling the geographically weighted model into the random forest model, the geographically weighted random forest model can fully consider the diffusion effects of various factors such as the geographical features, spatial relationships, enterprise production features, and enterprise management features of different grids on the heavy metal emission intensity. Moreover, the data used is more accessible, the data source reliability is higher, the data is more comprehensive and accurate, thereby improving the prediction accuracy and prediction efficiency.
[0040] After obtaining the predicted heavy metal emission intensity of each grid, optionally, as Figure 2 shown, the method further includes:
[0041] S201. For each grid, obtain the population density and farmland production potential value within the grid, and determine the predicted result of the heavy metal environmental risk of the grid according to the population density, farmland production potential value, and predicted heavy metal emission intensity.
[0042] Among them, the population density can be obtained by decomposing census data according to grid cells. The farmland production potential value can be evaluated by comprehensively considering the influence of various factors such as sunlight, temperature, water, carbon dioxide concentration, pests and diseases, agricultural climate limitations, soil, and terrain within the grid. The heavy metal environmental risk prediction result is used to characterize the harm degree caused by heavy metal emissions to the risk receptors within the grid. Here, the population and farmland are selected as risk receptors, and by comprehensively considering the influence of different factors such as the predicted heavy metal emission intensity, population density, and farmland production potential value on the environment, different weight coefficients can be assigned to different factors to balance the contributions of different factors, so as to obtain the heavy metal environmental risk prediction result of the grid.
[0043] Optionally, the process of determining the heavy metal environmental risk prediction result of the grid according to the population density, farmland production potential value, and predicted heavy metal emission intensity can be as follows: Determine the first environmental risk factor according to the predicted heavy metal emission intensity and population density; Determine the second environmental risk factor according to the predicted heavy metal emission intensity and farmland production potential value; Determine the heavy metal environmental risk prediction result of the grid according to the first environmental risk factor and the second environmental factor.
[0044] Among them, the first environmental risk factor is used to characterize the influence of heavy metal emissions on the population, and the second environmental risk factor is used to characterize the influence of heavy metal emissions on the farmland. By synthesizing the environmental risk factors of heavy metal emissions on the population and farmland, the final heavy metal environmental risk prediction result is determined.
[0045] Optionally, in combination with the population density, farmland production potential value, and predicted heavy metal emission intensity, the heavy metal environmental risk prediction result of the grid can also be determined by the following formula:
[0046] R = N(HD)×N(P)+N(HD)×N(AL);
[0047] Among them, R is the heavy metal environmental risk prediction result of the grid, HD is the predicted heavy metal emission intensity of the grid, P is the population density of the grid, AL is the farmland production potential value of the grid, and N represents normalization.
[0048] In this embodiment, by performing environmental risk assessment through the population density, farmland production potential value, and predicted heavy metal emission intensity of the grid, environmental risks can be more comprehensively identified and managed, which helps to protect human health and the ecosystem.
[0049] S202. Visually display the heavy metal environmental risk prediction results of each grid.
[0050] Specifically, visual methods such as heat maps can be used to display the heavy metal environmental risk prediction results of each grid, providing a basis for environmental supervision and enterprise emission reduction planning.
[0051] Furthermore, it is also possible to screen out target grids with environmental risks greater than a preset threshold based on the heavy metal environmental risk prediction results of each grid, determine environmental rectification suggestions according to the geospatial feature data, current production feature data, and environmental management feature data of non-ferrous metal enterprises in the target grids, and send the environmental rectification suggestions to grid management personnel to assist relevant personnel in environmental monitoring management, thereby improving the efficiency and accuracy of environmental management. For example, for grids with high heavy metal environmental risks, it is recommended to optimize production processes, control pollution sources, strengthen pollution control, ecological restoration, and health monitoring, etc.
[0052] In some examples, it is also possible to determine the risk level of each grid in combination with the heavy metal environmental risk prediction results of each grid, and output environmental rectification suggestions corresponding to the risk level in combination with the risk level of each grid. Optionally, the sending methods of environmental rectification suggestions include but are not limited to: email, text message, online platform, etc.
[0053] In one embodiment, optionally, as Figure 3 shown, the training process of the geographically weighted random forest model may include:
[0054] S301. Obtain a training data set.
[0055] Among them, the training data set includes the sample feature data of multiple sample non-ferrous metal enterprises and the corresponding historical heavy metal emission intensities. The sample feature data includes the geospatial feature data, production feature data, and environmental management feature data of the sample non-ferrous metal enterprises.
[0056] The historical heavy metal emission intensities of the above sample non-ferrous metal enterprises can be obtained through the following process:
[0057] (1). Calculate the historical product output of the sample non-ferrous metal enterprises
[0058] The product output can be obtained according to the production information data of the sample non-ferrous metal enterprises. If it cannot be directly obtained, the historical product output of the sample non-ferrous metal enterprises can be calculated through the following formula:
[0059]
[0060] In the formula, P y,i,j represents the output of sample non-ferrous metal enterprise y for product j in the i-th year, P x,i,j represents the output of product j in the x administrative region where sample non-ferrous metal enterprise y is located in the i-th year, IV y represents the industrial output value of sample non-ferrous metal enterprise y, and IV total,x,i,j represents the industrial output value of product j in the x administrative region where sample non-ferrous metal enterprise y is located in the i-th year. P x,i,jIt can be obtained from the statistical yearbooks or statistical bulletins published by the x administrative region.
[0061] (2) Determine the heavy metal production and pollutant discharge coefficients
[0062] The heavy metal production and pollutant discharge coefficients of the sample non-ferrous metal enterprises can be determined with reference to information such as products, raw materials, processes, sections, scales, and end-of-treatment technologies in the "Handbook of Industrial Pollution Source Production and Pollutant Discharge Coefficients".
[0063] (3) Calculate the historical heavy metal emissions of the sample non-ferrous metal enterprises
[0064] Construct a heavy metal emissions accounting model for the sample non-ferrous metal enterprises (as shown in the following formula), and determine the historical heavy metal emissions generated by the sample non-ferrous metal enterprises through the atmosphere and wastewater based on the historical product output and heavy metal production and pollutant discharge coefficients obtained above:
[0065]
[0066] Among them, H 排 represents the historical heavy metal emissions of the sample non-ferrous metal enterprises, H 产 represents the historical heavy metal generation of the sample non-ferrous metal enterprises, H 除 represents the historical heavy metal removal of the sample non-ferrous metal enterprises, P y,i,j represents the output of product j of sample non-ferrous metal enterprise y in the i-th year, M y,i,j represents the heavy metal production and pollutant discharge coefficient of sample non-ferrous metal enterprise y for product j in the i-th year, η represents the heavy metal removal rate, κ represents the equipment operation rate, and n represents the number of product types.
[0067] (4) Account for the historical heavy metal emission intensity of the sample non-ferrous metal enterprises
[0068] Based on the historical heavy metal emissions calculated above and the enterprise area, calculate the historical heavy metal emission intensity of the sample non-ferrous metal enterprises through the following formula:
[0069]
[0070] Among them, H 排 represents the historical heavy metal emissions of the sample non-ferrous metal enterprises, S 企业 is the enterprise area of the sample non-ferrous metal enterprise, and HD is the historical heavy metal emission intensity of the sample non-ferrous metal enterprise.
[0071] S302. Conduct a correlation check on the obtained multiple sample characteristic data, and determine the target sample characteristic data from the multiple sample characteristic data according to the correlation check results.
[0072] Specifically, according to the data type of the sample feature data, the obtained multiple sample feature data can be classified to obtain at least one data group. For each data group, according to the corresponding correlation verification algorithm of the data group, the multiple sample feature data in the data group are subjected to correlation verification, and the sample feature data whose correlation verification results meet the preset conditions are selected from the multiple sample feature data as the target sample feature data. For example, for numerical sample feature data (such as altitude, slope, elevation, vegetation index, river network density, distance from river, road network density, distance from road, rainfall, wind speed, population density, gross domestic product, energy utilization rate, waste water / waste gas / waste residue treatment efficiency, etc.), the Pearson correlation coefficient can be directly used for correlation test. Another example is that for ordinal categorical sample feature data (such as the frequency of enterprise illegal records), the biserial correlation coefficient test can be used. Another example is that the distance from road is negatively correlated with the road network density. Since it is easier to obtain dynamic data for the road network density, the distance from road can be deleted and the road network density can be retained.
[0073] By performing correlation verification on the obtained multiple sample feature data, redundant sample feature data are removed, reducing the computational cost and time in the model training process, thereby improving the training efficiency and generalization ability of the model.
[0074] S303. Based on the geographical location information of all sample non-ferrous metal enterprises, construct a spatial weight matrix through a kernel function.
[0075] Based on the geographical location information of all sample non-ferrous metal enterprises, determine the Euclidean distance between each sample non-ferrous metal enterprise to form a distance matrix, and construct a spatial weight matrix using a kernel function according to the distance matrix and the spatial bandwidth parameter. Among them, the elements in the spatial weight matrix are used to represent the spatial weights between sample non-ferrous metal enterprises. The above spatial bandwidth parameter can be adaptively adjusted according to the spatial distribution density of the sample non-ferrous metal enterprises. For example, a smaller spatial bandwidth parameter is used in dense areas, and a larger spatial bandwidth parameter is used in sparse areas.
[0076] S304. Train the geographically weighted random forest model obtained according to the spatial weight matrix based on the target sample feature data and the historical heavy metal emission intensity.
[0077] Optionally, the optimization parameters of the geographically weighted random forest model include at least one of the following: spatial bandwidth parameter, number of decision trees, maximum depth of decision trees, minimum number of samples in leaf nodes and node splitting rules, number of features for splitting nodes, and minimum number of samples in internal nodes.
[0078] In the sample sampling stage of constructing decision trees in the geographically weighted random forest model, a spatial weight matrix is introduced, and samples are drawn according to the magnitude of the weights. The probability of a sample being drawn is proportional to its weight, that is, neighboring points are more likely to be drawn. For each sample set obtained by weighted sampling, a decision tree is independently constructed. At each node split, the spatial weight matrix is added to adjust the calculation results of purity metrics (such as root mean square error, information gain, etc.), and the feature with the optimal purity metric is selected for splitting, thereby improving the accuracy of the geographically weighted random forest model in learning spatial features in local regions.
[0079] Furthermore, the performance of the trained geographically weighted random forest model can also be evaluated by the coefficient of determination (R2), relative standard deviation (RMSE), and mean absolute error (MAE).
[0080] Specifically,
[0081]
[0082] where n is the number of sample feature data, y m is the statistical value (i.e., historical heavy metal emission intensity), y p is the predicted value (i.e., predicted heavy metal emission intensity), y pa is the average of the predicted values, and y ma is the average of the statistical values.
[0083] Quantitatively evaluate the model according to the selected metrics, compare the evaluation results under different model versions or parameter settings. If R 2 is low and MAE is high, indicating that the model needs to be optimized, then the model parameters can be adjusted, and the geographically weighted random forest model can be continuously trained based on the adjusted model parameters until the preset model convergence condition is reached.
[0084] In this embodiment, the spatial weight matrix is introduced into the training process of the geographically weighted random forest model, enabling the geographically weighted random forest model to fully consider the diffusion effects of various factors such as geographical features, spatial relationships, enterprise production features, and enterprise management features of different grids on heavy metal emission intensity, thereby improving the prediction accuracy of the model.
[0085] Taking the prediction of the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises in the Yangtze River Basin as an example for introduction:
[0086] Collect the basic information of non-ferrous metal enterprises in the Yangtze River Basin through the industrial and commercial registration information and the environmental protection record-filing enterprise list in the Yangtze River Basin, and combine technologies such as web crawling. Divide the Yangtze River Basin into grids according to the preset grid size, and combine the above basic information to obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data in each grid, so as to obtain the basic data of each grid. Through the pre-trained geographically weighted random forest model and the basic data of each grid, predict the heavy metal emission intensity of each grid in the Yangtze River Basin. Further, normalize the predicted heavy metal emission intensity of each grid, delimit the high and low areas of pollution accumulation according to the quantile method, and visually display the normalized results. Further, it is also possible to obtain the population density and farmland production potential value of each grid in the Yangtze River Basin, use the predicted heavy metal emission intensity, population density and farmland production potential value obtained above to predict the heavy metal environmental risk score of each grid, and determine the risk level corresponding to the heavy metal environmental risk score by the natural break method. For example, the risk level can be a low-risk area, a medium-risk area, and a high-risk area; when the risk level exceeds the preset level, trigger an alarm and output corresponding environmental rectification suggestions. Figure 4 This is a schematic structural diagram of a device for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises provided by an embodiment of the present application. As Figure 4 shown, the device may include: an acquisition module 401 and a processing module 402.
[0087] Specifically, the acquisition module 401 is used to divide the target area into grids, and obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data in each grid, so as to obtain the basic data of each grid;
[0088] The processing module 402 is used to input the basic data of each grid into the pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
[0089] On the basis of the above embodiment, optionally, the acquisition module 401 is further used to obtain the population density and farmland production potential value in each grid for each grid;
[0090] The processing module 402 is further used to determine the predicted result of the heavy metal environmental risk of each grid according to the population density, the farmland production potential value, and the predicted heavy metal emission intensity, and visually display the predicted results of the heavy metal environmental risk of each grid.
[0091] Based on the above embodiments, optionally, the processing module 402 is specifically configured to determine a first environmental risk factor according to the predicted heavy metal emission intensity and the population density; determine a second environmental risk factor according to the predicted heavy metal emission intensity and the farmland production potential value; and determine the heavy metal environmental risk prediction result of the grid according to the first environmental risk factor and the second environmental factor.
[0092] Based on the above embodiments, optionally, the acquisition module 401 is further configured to acquire a training data set; wherein, the training data set includes sample feature data of multiple sample non-ferrous metal enterprises and corresponding historical heavy metal emission intensities, and the sample feature data includes geospatial feature data, production feature data, and environmental management feature data of the sample non-ferrous metal enterprises.
[0093] The processing module 402 is further configured to perform correlation verification on the obtained multiple sample feature data, and determine target sample feature data from the multiple sample feature data according to the correlation verification result; construct a spatial weight matrix through a kernel function based on the geographical location information of all sample non-ferrous metal enterprises; and train a geographically weighted random forest model obtained according to the spatial weight matrix based on the target sample feature data and the historical heavy metal emission intensity; wherein, the elements in the spatial weight matrix are used to represent the spatial weights between the sample non-ferrous metal enterprises.
[0094] Based on the above embodiments, optionally, the spatial weight matrix is determined by the following formula:
[0095]
[0096] where d ij is the Euclidean distance between sample non-ferrous metal enterprise i and sample non-ferrous metal enterprise j, h i is the spatial bandwidth parameter, used to control the attenuation speed of the spatial weight, and w ij is the element in the spatial weight matrix, representing the spatial weight between sample non-ferrous metal enterprise i and sample non-ferrous metal enterprise j.
[0097] Based on the above embodiments, optionally, the optimization parameters of the geographically weighted random forest model include at least one of the following:
[0098] Spatial bandwidth parameter, number of decision trees, maximum depth of decision trees, minimum number of samples in leaf nodes and node splitting rules, number of features for splitting nodes, and minimum number of samples in internal nodes.
[0099] Based on the above embodiments, optionally, the geospatial feature data includes at least one of the following: land use type, terrain, vegetation index, river information, climate information, socioeconomic information, and enterprise geographical location information;
[0100] The production feature data includes at least one of the following: enterprise raw material characteristics and energy consumption characteristics;
[0101] The environmental management feature data includes at least one of the following: enterprise violation records and industrial pollutant treatment efficiency.
[0102] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present application is shown in Figure 5 As shown, the electronic device includes a processor 50, a memory 51, an input device 52, and an output device 53; the number of processors 50 in the electronic device can be one or more, Figure 5 Taking one processor 50 as an example; the processor 50, the memory 51, the input device 52, and the output device 53 in the electronic device can be connected through a bus or other means, Figure 5 Taking the connection through a bus as an example.
[0103] The memory 51, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the method for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises in the embodiments of the present application. The processor 50 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 51, that is, implements the method for predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises provided in any of the above embodiments.
[0104] The memory 51 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created during the process of predicting the spatial distribution of heavy metal emission intensity of non-ferrous metal enterprises, etc. In addition, the memory 51 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 51 may further include a memory remotely set relative to the processor 50, and these remote memories can be connected to the device / terminal / server through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0105] The input device 52 can be used to receive input digital or character information and generate key signal inputs related to user settings and function controls of the electronic device. The output device 53 can include display devices such as a display screen.
[0106] In one embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0107] Perform grid division on the target area, and obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data within each grid to obtain the basic data of each grid;
[0108] Input the basic data of each grid into a pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
[0109] In one embodiment, a computer program product is further provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0110] Perform grid division on the target area, and obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data within each grid to obtain the basic data of each grid;
[0111] Input the basic data of each grid into a pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
[0112] The spatial distribution prediction device, electronic device, computer-readable storage medium, and computer program product for the heavy metal emission intensity of non-ferrous metal enterprises provided in the above embodiments can execute the spatial distribution prediction method for the heavy metal emission intensity of non-ferrous metal enterprises provided in any embodiment of the present application, and have corresponding functional modules and beneficial effects for executing this method. Technical details not described in detail in the above embodiments can be referred to the spatial distribution prediction method for the heavy metal emission intensity of non-ferrous metal enterprises provided in any embodiment of the present application.
[0113] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0114] It should be noted that the various units and modules included in the above embodiments are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present application.
[0115] Note that the above is only the preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for predicting the spatial distribution of heavy metal emission intensity in non-ferrous metal enterprises, characterized in that, Including: Performing grid division on the target area, and obtaining the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data within each grid to obtain the basic data of each grid; Inputting the basic data of each grid into the pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
2. The method according to claim 1, wherein It also includes: For each grid, obtaining the population density and farmland production potential value within the grid, and determining the predicted heavy metal environmental risk result of the grid according to the population density, the farmland production potential value, and the predicted heavy metal emission intensity; Visually displaying the predicted heavy metal environmental risk results of each grid.
3. The method according to claim 2, wherein Determining the predicted heavy metal environmental risk result of the grid according to the population density, the farmland production potential value, and the predicted heavy metal emission intensity includes: Determining a first environmental risk factor according to the predicted heavy metal emission intensity and the population density; Determining a second environmental risk factor according to the predicted heavy metal emission intensity and the farmland production potential value; Determining the predicted heavy metal environmental risk result of the grid according to the first environmental risk factor and the second environmental factor.
4. The method according to claim 1, wherein The training process of the geographically weighted random forest model includes: Obtaining a training data set; wherein, the training data set includes the sample feature data of multiple sample non-ferrous metal enterprises and the corresponding historical heavy metal emission intensity, and the sample feature data includes the geospatial feature data, production feature data, and environmental management feature data of the sample non-ferrous metal enterprises; Performing correlation verification on the obtained multiple sample feature data, and determining target sample feature data from the multiple sample feature data according to the correlation verification result; Based on the geographical location information of all sample non-ferrous metal enterprises, constructing a spatial weight matrix through a kernel function; wherein, the elements in the spatial weight matrix are used to represent the spatial weights between the sample non-ferrous metal enterprises; Training the geographically weighted random forest model obtained according to the spatial weight matrix based on the target sample feature data and the historical heavy metal emission intensity.
5. The method according to claim 4, wherein The spatial weight matrix is determined by the following formula: Among them, d ij is the Euclidean distance between sample non-ferrous metal enterprise i and sample non-ferrous metal enterprise j, and h i is the spatial bandwidth parameter, which is used to control the attenuation rate of the spatial weight, and w ij is an element in the spatial weight matrix, representing the spatial weight between sample non-ferrous metal enterprise i and sample non-ferrous metal enterprise j.
6. The method according to claim 4, characterized in that, The optimization parameters of the geographically weighted random forest model include at least one of the following: Spatial bandwidth parameter, number of decision trees, maximum depth of decision trees, minimum number of samples in leaf nodes and node splitting rules, number of features for splitting nodes, and minimum number of samples in internal nodes.
7. The method according to any one of claims 1 to 6, characterized in that The geospatial feature data includes at least one of the following: land use type, terrain, vegetation index, river information, climate information, social and economic information, and enterprise geographical location information; The production feature data includes at least one of the following: enterprise raw material characteristics and energy consumption characteristics; The environmental management feature data includes at least one of the following: enterprise violation records and industrial pollutant treatment efficiency.
8. A device for predicting the spatial distribution of the heavy metal emission intensity of non-ferrous metal enterprises, characterized in that, Including: An acquisition module, configured to perform grid division on the target area, and obtain the geospatial feature data, production feature data of non-ferrous metal enterprises, and environmental management feature data within each grid to obtain the basic data of each grid; A processing module, configured to input the basic data of each grid into a pre-trained geographically weighted random forest model to obtain the predicted heavy metal emission intensity of each grid.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Geochemical variable space prediction method based on geostatistical weighted random forest
CN114139819A
Large-area soil heavy metal detection and spatial and temporal distribution characteristic analysis method and system
CN114443982A
Regional soil heavy metal net input flux investigation method
CN118211189A
Urban plant-network-river-based integrated optimal scheduling method and system
CN119047742A