A method and system for monitoring soil arsenic distribution
By constructing the optimal spatial inference model of soil arsenic content and historical natural factors and human factor variables, the problem of inefficient monitoring of soil arsenic distribution is solved, and rapid and accurate prediction of soil arsenic content is achieved, and monitoring efficiency and reliability are improved.
Patent Information
- Application Number
- CN202411512238.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The prior art is inefficient in monitoring the spatial distribution of soil arsenic elements, relies on a large number of field sampling and laboratory analysis, which is time-consuming and laborious and costly.
By constructing the optimal spatial inference model of soil arsenic content, historical natural factor variables and human factor variables, using symbolic regression method and maximum mutual information coefficient to screen variables, an expression binary tree is constructed to optimize the model, and a rapid and accurate prediction of soil arsenic content is achieved.
It improves the efficiency of soil arsenic distribution monitoring, reduces dependence on field sampling, reduces workload and cost, and improves the reliability of predicted results.
Smart Images

Figure CN119537343B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of arsenic monitoring, and in particular to a method and system for monitoring the distribution of arsenic in soil. Background Art
[0002] Arsenic (As) is a highly carcinogenic non-metallic element. Its spatial distribution characteristics in soil are of great significance for assessing ecosystem impacts, monitoring ecological health of polluted areas, and formulating environmental governance policies. However, there are significant challenges in estimating the spatial distribution of soil arsenic.
[0003] Traditional methods rely on a large number of field sampling and laboratory analysis to determine the arsenic content in soil, which is not only time-consuming and labor-intensive, but also costly. In addition, the data processing and analysis stages also require a lot of time and human resources, resulting in inefficient monitoring of soil arsenic distribution. Summary of the invention
[0004] In response to the problems in the prior art, the present application provides a soil arsenic distribution monitoring method that can quickly and accurately predict the arsenic content in the soil and effectively improve the efficiency of soil arsenic content distribution monitoring by constructing an optimal soil arsenic spatial inference model of soil arsenic content and historical natural factor variables and human factor variables.
[0005] A method for monitoring the distribution of arsenic in soil comprises the following steps:
[0006] Acquire several groups of historical natural factor variables, historical human factor variables and historical arsenic content data of the target monitoring area from a preset variable database, wherein the historical natural factor variables include at least: temperature, precipitation, soil gravel content and soil pH value; the historical human factor variables include at least: population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points;
[0007] Respectively analyzing and screening the correlation between the historical natural factor variables and the historical human factor variables and the historical arsenic content data to obtain screened historical natural factor variables and historical human factor variables;
[0008] Respectively analyzing the linear correlation between the screened historical natural factor variables and historical human factor variables and the historical arsenic content data to obtain the partial correlation coefficient of the natural factor and the partial correlation coefficient of the human factor;
[0009] Respectively dividing the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors by the sum of the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors to obtain the weight of natural factors and the weight of human factors;
[0010] According to the historical natural factor variables, historical human factor variables and historical arsenic content data, a soil arsenic distribution monitoring model is constructed based on the symbolic regression method;
[0011] The actual historical natural factor variables and the actual human factor variables of the target monitoring area are obtained from the variable database, the actual historical natural factor variables and the actual human factor variables are substituted into the soil arsenic distribution monitoring model, and weighted summation is performed according to the natural factor weight and the human factor weight to obtain the actual soil arsenic content.
[0012] The present application also provides a soil arsenic distribution monitoring system, comprising:
[0013] Variable data acquisition module: used to obtain several groups of historical natural factor variables, historical human factor variables and historical arsenic content data of the target monitoring area from a preset variable database, wherein the historical natural factor variables include at least: temperature, precipitation, soil gravel content and soil pH value; the historical human factor variables include at least: population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points;
[0014] Variable screening module: used to analyze and screen the correlation between the historical natural factor variable and the historical human factor variable and the historical arsenic content data, respectively, to obtain the screened historical natural factor variable and historical human factor variable;
[0015] Partial correlation coefficient calculation module: used to analyze the linear correlation between the historical natural factor variables and the historical human factor variables after screening and the historical arsenic content data, and obtain the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors;
[0016] A weight calculation module: used for respectively dividing the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors by the sum of the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors to obtain the weight of natural factors and the weight of human factors;
[0017] Model building module: used to build a soil arsenic distribution monitoring model based on the symbolic regression method according to the historical natural factor variables, historical human factor variables and historical arsenic content data;
[0018] Actual arsenic content calculation module: used to obtain the actual historical natural factor variables and actual human factor variables of the target monitoring area from the variable database, substitute the actual historical natural factor variables and actual human factor variables into the soil arsenic distribution monitoring model, and perform weighted summation according to the natural factor weight and the human factor weight to obtain the actual soil arsenic content.
[0019] Compared with the existing technology, this scheme uses the maximum mutual information coefficient and significance test to screen explanatory variables, determines the weights through the partial correlation coefficient, and then constructs and optimizes the expression binary tree based on the screened explanatory variables, and finally obtains the optimal function fitting expression, that is, the soil arsenic element distribution monitoring model to predict the soil arsenic content. It can more accurately predict the soil arsenic content and improve the reliability of the prediction results. On the basis of the known historical natural factor variables and human factor variables, the constructed soil arsenic element distribution monitoring model can reduce the dependence on field sampling to a certain extent, thereby reducing the workload of manual field sampling, etc., and effectively improving the monitoring efficiency of soil arsenic content.
[0020] In order to provide a clearer understanding of the present application, the specific implementation of the present application will be described below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flow chart of a method for monitoring the distribution of arsenic in soil for this application;
[0022] Figure 2 A flow chart of a method for screening historical natural factor variables and historical human factor variables in a soil arsenic distribution monitoring method for this application;
[0023] Figure 3 A flow chart of a method for obtaining the maximum mutual information coefficient of natural factors in a soil arsenic distribution monitoring method of this application;
[0024] Figure 4 A flow chart of a method for obtaining the maximum mutual information coefficient of human factors in a soil arsenic distribution monitoring method for this application;
[0025] Figure 5 A flow chart of a method for obtaining significant probability values of natural factors and significant probability values of human factors in a soil arsenic distribution monitoring method of the present application;
[0026] Figure 6 A flow chart of a method for obtaining partial correlation coefficients of natural factors and partial correlation coefficients of human factors in a soil arsenic distribution monitoring method for this application;
[0027] Figure 7 A flow chart of a method for constructing a soil arsenic distribution monitoring model in a soil arsenic distribution monitoring method of the present application;
[0028] Figure 8 A flow chart of a method for iteratively optimizing a model to be optimized in a soil arsenic distribution monitoring method of the present application;
[0029] Fig. 9 This is a schematic diagram of a soil arsenic distribution monitoring system for the present application. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0031] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of the application. It should be understood that the operations of the flowcharts may be implemented out of order, and steps without logical contextual relationships may be reversed in order or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowcharts, or may remove one or more operations from the flowcharts, under the guidance of the content of this application.
[0032] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0033] Example 1
[0034] See also Figure 1 , Figure 1 This is a flow chart of a soil arsenic distribution monitoring method of the present application. The present application provides a soil arsenic distribution monitoring method, which specifically includes the following steps:
[0035] S1: Acquire several groups of historical natural factor variables, historical human factor variables and historical arsenic content data of the target monitoring area from a preset variable database, wherein the historical natural factor variables at least include: temperature, precipitation, soil gravel content and soil pH value; the historical human factor variables at least include: population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points;
[0036] S2: respectively analyzing and screening the correlation between the historical natural factor variable and the historical human factor variable and the historical arsenic content data to obtain screened historical natural factor variables and historical human factor variables;
[0037] S3: respectively analyzing the linear correlation between the screened historical natural factor variables and historical human factor variables and the historical arsenic content data to obtain the partial correlation coefficient of the natural factor and the partial correlation coefficient of the human factor;
[0038] S4: respectively dividing the partial correlation coefficient of the natural factor and the partial correlation coefficient of the human factor by the sum of the partial correlation coefficient of the natural factor and the partial correlation coefficient of the human factor to obtain the weight of the natural factor and the weight of the human factor;
[0039] S5: constructing a soil arsenic distribution monitoring model based on the symbolic regression method according to the historical natural factor variables, historical human factor variables and historical arsenic content data;
[0040] S6: Obtain the actual historical natural factor variables and actual human factor variables of the target monitoring area from the variable database, substitute the actual historical natural factor variables and actual human factor variables into the soil arsenic distribution monitoring model, and perform weighted summation according to the natural factor weight and the human factor weight to obtain the actual soil arsenic content.
[0041] The soil arsenic distribution monitoring method of the present invention can be executed by the following computer system, which includes a variable database server, a data acquisition server and a soil arsenic distribution analysis server. The variable database server is used to construct the variable database, and store several groups of historical natural factor variables, historical human factor variables, historical arsenic content data, actual historical natural factor variables, actual human factor variables and variable data sets in the target monitoring area.
[0042] The data acquisition server is used to obtain several groups of historical natural factor variables, historical human factor variables, historical arsenic content data, actual historical natural factor variables and actual human factor variables in the target monitoring area from the variable database server, and send them to the soil arsenic element distribution analysis server for processing.
[0043] The soil arsenic distribution analysis server executes the soil arsenic distribution monitoring method of the present invention to calculate the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors to obtain the natural factor weight and the human factor weight, and constructs a soil arsenic distribution monitoring model based on the symbolic regression method according to the historical natural factor variables, the historical human factor variables and the historical arsenic content data, substitutes the actual historical natural factor variables and the actual human factor variables into the soil arsenic distribution monitoring model, and performs weighted summation in combination with the natural factor weight and the human factor weight to obtain the actual soil arsenic content, thereby completing the monitoring of soil arsenic.
[0044] Compared with the existing technology, this scheme uses the maximum mutual information coefficient and significance test to screen explanatory variables, and determines the weights through the partial correlation coefficient to ensure that in the process of monitoring soil arsenic content, more attention is paid to historical natural factor variables or human factor variables that have a significant impact on the monitoring results, and the impact of historical natural factor variables or human factor variables with less correlation with the monitoring results on the monitoring results is reduced, thereby improving the accuracy of the prediction.
[0045] Then, an expression binary tree is constructed and optimized based on the screened historical natural factor variables or human factor variables, and finally the optimal function fitting expression, that is, the soil arsenic element distribution monitoring model, is obtained to predict the soil arsenic content. The soil arsenic content can be predicted more accurately and the reliability of the prediction results can be improved. On the basis of the known historical natural factor variables and human factor variables, the constructed soil arsenic element distribution monitoring model can reduce the dependence on field sampling to a certain extent, thereby reducing the workload of manual field sampling, etc., and effectively improving the monitoring efficiency of soil arsenic content.
[0046] For step S1, in this embodiment, the historical natural factor variables and historical human factor variables can be obtained from the online database and platform of the target monitoring area. The historical natural factor variables include at least: temperature, precipitation, soil gravel content and soil pH value. For the temperature and precipitation, the monthly scale multi-year average value of temperature and the monthly scale multi-year average value of precipitation are preferentially obtained.
[0047] The monthly multi-year average of the temperature affects the migration and transformation rate of arsenic in the soil, and high temperature can accelerate the release and diffusion of arsenic; the monthly multi-year average of precipitation affects the leaching and migration of soil arsenic, and arsenic in the soil of high precipitation areas is more likely to be lost with water; the soil gravel content affects the soil's ability to adsorb and fix arsenic, and soil with a high gravel content has a stronger adsorption capacity for arsenic; the soil pH value affects the chemical form and solubility of arsenic in the soil, and the solubility of arsenic increases under acidic conditions. Therefore, in this embodiment, the temperature, precipitation, soil gravel content and soil pH value are selected as the historical natural factor variables.
[0048] The historical man-made and natural factor variables at least include: population density, land use type, inverse distance interpolation of transportation land and inverse distance interpolation of arsenic mine points.
[0049] Since densely populated areas are accompanied by more industrial, agricultural and domestic pollution, the sources of soil arsenic are increased; different land use types such as agricultural land, industrial land and residential land have an important impact on the distribution of soil arsenic, for example, agricultural land increases arsenic content due to the use of fertilizers and pesticides; the inverse distance interpolation of traffic land reflects the impact of traffic activities on soil arsenic distribution, and traffic-dense areas may increase soil arsenic content due to vehicle exhaust, tire wear, etc.; the inverse distance interpolation of arsenic mines reflects the impact of arsenic mineral mining and processing activities on soil arsenic distribution, and the soil arsenic content near arsenic mines increases significantly. Therefore, in this embodiment, the population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mines are selected as the historical human factor variables.
[0050] In other embodiments, since wind speed affects the migration and diffusion of soil particles (arsenic-containing particles), the risk of wind erosion and diffusion of soil arsenic increases when the wind speed is high; water vapor pressure is related to soil moisture and the activity of arsenic, and high water vapor pressure increases soil moisture, thereby affecting the solubility and mobility of arsenic; ultraviolet radiation and other radiation affect the photochemical transformation and degradation of soil arsenic; elevation changes lead to significant differences in climate and soil properties, thereby affecting the distribution of soil arsenic; slope and slope direction affect the leaching and migration of soil arsenic by affecting the direction and speed of water flow; curvature and terrain moisture index are related to soil moisture retention and drainage, affecting the leaching and accumulation of arsenic in the soil; the monthly scale multi-year average of soil moisture is a key factor affecting the solubility and mobility of soil arsenic, and high moisture content increases the dissolution and migration of arsenic; soil texture such as silt, clay and sand content affects the soil's ability to adsorb and fix arsenic, and soil with a high clay content has a stronger adsorption capacity for arsenic; parent rock parent material affects the original content and distribution of arsenic in the soil;
[0051] Therefore, the historical natural factor variables also include: wind speed, water vapor pressure, radiation, elevation, slope, aspect, curvature, terrain moisture index, soil moisture, soil bulk density, parent rock parent material, soil silt content, soil clay content and soil sand content. Among them, for wind speed, water vapor pressure, radiation and soil moisture, priority is given to obtaining the monthly multi-year average of wind speed, the monthly multi-year average of water vapor pressure, the monthly multi-year average of radiation and the monthly multi-year average of soil moisture.
[0052] Since night lights are an indirect indicator of human activities, areas with high night light intensity may be accompanied by more pollution emissions, affecting the soil arsenic content; mining and processing activities of gold, silver, copper, lead, zinc, pyrite and other minerals will also affect the distribution of soil arsenic, and the soil arsenic content near the mining sites will increase significantly.
[0053] Therefore, the historical human factor variables also include night lights, inverse distance interpolation of gold mine points, inverse distance interpolation of silver mine points, inverse distance interpolation of copper mine points, inverse distance interpolation of lead mine points, inverse distance interpolation of zinc mine points and inverse distance interpolation of pyrite mine points, etc.
[0054] Of course, the historical natural factor variables and historical human factor variables can also be obtained from soil research institutions and data centers in the target monitoring area.
[0055] Please also see Figure 2 , Figure 2 This is a flow chart of a method for screening historical natural factor variables and historical human factor variables in a soil arsenic distribution monitoring method of the present application. For step S2, the correlation between the historical natural factor variables and the historical human factor variables and the historical arsenic content data is analyzed and screened to obtain the screened historical natural factor variables and historical human factor variables, and the following steps are also included:
[0056] S21: performing mutual information maximization analysis on the historical natural factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of the natural factor;
[0057] S22: performing mutual information maximization analysis on the historical human factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of the human factor;
[0058] S23: performing significance tests on the multiple groups of historical natural factor variables, historical human factor variables and the historical arsenic content data respectively to obtain significant probability values of natural factors and significant probability values of human factors;
[0059] S24: Screen the historical natural factor variables and historical human factor variables corresponding to the natural factor maximum mutual information coefficient and the human factor maximum mutual information coefficient that are greater than the preset maximum mutual information coefficient threshold, and the natural factor significant probability value and the human factor significant probability value that are greater than the preset significant probability threshold.
[0060] For step S21, in one embodiment, the maximum mutual information coefficient of natural factors represents the correlation between the historical natural factor variables and the historical arsenic content data. The maximum mutual information coefficient of natural factors includes the maximum mutual information coefficient of air temperature, the maximum mutual information coefficient of precipitation, the maximum mutual information coefficient of soil gravel content, and the maximum mutual information coefficient of soil pH value, wherein the maximum mutual information coefficient of air temperature, the maximum mutual information coefficient of precipitation, the maximum mutual information coefficient of soil gravel content, and the maximum mutual information coefficient of soil pH value represent the correlation between air temperature, precipitation, soil gravel content, and soil pH value and the historical arsenic content data, respectively.
[0061] Of course, in other embodiments, the maximum mutual information coefficient of natural factors may also include the maximum mutual information coefficient of wind speed, the maximum mutual information coefficient of water vapor pressure, the maximum mutual information coefficient of radiation, the maximum mutual information coefficient of elevation, the maximum mutual information coefficient of slope, the maximum mutual information coefficient of aspect, the maximum mutual information coefficient of curvature, the maximum mutual information coefficient of terrain moisture index, the maximum mutual information coefficient of soil moisture, the maximum mutual information coefficient of soil bulk density, the maximum mutual information coefficient of parent rock and parent material, the maximum mutual information coefficient of soil silt content, the maximum mutual information coefficient of soil clay content and the maximum mutual information coefficient of soil sand content.
[0062] See also Figure 3 , Figure 3 This is a flow chart of a method for obtaining the maximum mutual information coefficient of natural factors in a soil arsenic distribution monitoring method of the present application. The method of maximizing the mutual information of the historical natural factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of natural factors also includes the following steps:
[0063] S211: Use the historical arsenic content data and any of the historical natural factor variables as the Y axis and X axis respectively to construct a natural factor scatter plot, and divide the X axis and Y axis of the natural factor scatter plot into a plurality of natural factor intervals according to the preset first horizontal axis dimension and first vertical axis dimension, where m1×n1 <s1 0.7 , m1 is the first horizontal axis dimension, n1 is the first vertical axis dimension, and s1 is the number of groups of the historical natural factor variables;
[0064] S212: creating a two-dimensional frequency matrix of natural factors according to the first horizontal axis dimension, the first vertical axis dimension and the natural factor scatter plot, wherein the elements of the two-dimensional frequency matrix of natural factors represent the number of groups in which any of the historical natural factor variables and the historical arsenic content data simultaneously fall within any of the natural factor intervals;
[0065] S213: Divide each element of the two-dimensional frequency matrix of natural factors by the number of groups of the historical natural factor variables to construct a natural factor joint probability distribution matrix, wherein the elements of the natural factor joint probability distribution matrix represent the joint probability of any of the historical natural factor variables and the historical arsenic content data;
[0066] S214: according to the natural factor joint probability distribution matrix, respectively accumulating the row data and the column data of the natural factor joint probability distribution matrix to obtain the probability distribution coefficient of any of the historical natural factor variables and the first arsenic content probability distribution coefficient;
[0067] S215: Calculate the maximum mutual information between the historical natural factor variable and the historical arsenic content data according to the following formula based on the joint probability of the historical natural factor variable and the historical arsenic content data, the probability distribution coefficient of the historical natural factor variable and the first arsenic content probability distribution coefficient:
[0068]
[0069] Wherein, MI(x;y) is the maximum mutual information between any historical natural factor variable and the historical arsenic content data, x is any historical natural factor variable, x∈X i , X i ={temperature, precipitation, soil gravel content, soil pH}, y is the historical arsenic content data, jp(x,y) is the joint probability of any of the historical natural factor variables and the historical arsenic content data, p(x) is the probability distribution coefficient of any of the historical natural factor variables, and p1(y) is the first arsenic content probability distribution coefficient;
[0070] S216: Calculate the maximum mutual information coefficient between any one of the historical natural factor variables and the historical arsenic content data according to the following formula based on the first horizontal axis dimension, the first vertical axis dimension, and the maximum mutual information between any one of the historical natural factor variables and the historical arsenic content data:
[0071] MIC(x;y)=maxMI(x;y) / log2min(m1,n1)
[0072] In the formula, MIC(x; y) is the maximum mutual information coefficient between any of the historical natural factor variables and the historical arsenic content data, MI(x; y) is the maximum mutual information between any of the historical natural factor variables and the historical arsenic content data, m1 and n1 are the first horizontal axis dimension and the first vertical axis dimension respectively, and min(m1, n1) represents the smaller value of m1 and n1.
[0073] For step S211, in this embodiment, the natural factor scatter plot includes: a temperature factor scatter plot, a precipitation factor scatter plot, a soil gravel content factor scatter plot and a soil pH value factor scatter plot, wherein the temperature factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the temperature as the X-axis; the precipitation factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the precipitation as the X-axis; the soil gravel content factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the soil gravel content as the X-axis; the soil pH value factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the soil pH as the X-axis.
[0074] Of course, in other embodiments, the natural factor scatter plot may also include a wind speed scatter plot, a water vapor pressure scatter plot, a radiation scatter plot, an elevation scatter plot, a slope scatter plot, an aspect scatter plot, a curvature scatter plot, a terrain moisture index scatter plot, a soil moisture scatter plot, a soil bulk density scatter plot, a parent rock and parent material scatter plot, a soil silt content scatter plot, a soil clay content scatter plot, and a soil sand content scatter plot.
[0075] For step S212, the two-dimensional frequency matrix of natural factors includes: a two-dimensional frequency matrix of temperature factors, a two-dimensional frequency matrix of precipitation factors, a two-dimensional frequency matrix of soil gravel content factors, and a two-dimensional frequency matrix of soil pH value factors.
[0076] Among them, the elements of the two-dimensional frequency matrix of the temperature factor represent the number of groups in which the temperature and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the temperature factor scatter diagram; the elements of the two-dimensional frequency matrix of the precipitation factor represent the number of groups in which the precipitation and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the precipitation factor scatter diagram; the elements of the two-dimensional frequency matrix of the soil gravel content factor represent the number of groups in which the soil gravel content and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the soil gravel content factor scatter diagram; the elements of the two-dimensional frequency matrix of the soil pH value factor represent the number of groups in which the soil pH value and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the soil pH value factor scatter diagram.
[0077] Of course, in other embodiments, the two-dimensional frequency matrix of natural factors may also include a two-dimensional frequency matrix of wind speed, a two-dimensional frequency matrix of water vapor pressure, a two-dimensional frequency matrix of radiation, a two-dimensional frequency matrix of elevation, a two-dimensional frequency matrix of slope, a two-dimensional frequency matrix of aspect, a two-dimensional frequency matrix of curvature, a two-dimensional frequency matrix of terrain moisture index, a two-dimensional frequency matrix of soil moisture, a two-dimensional frequency matrix of soil bulk density, a two-dimensional frequency matrix of parent rock and parent material, a two-dimensional frequency matrix of soil silt content, a two-dimensional frequency matrix of soil clay content, and a two-dimensional frequency matrix of soil sand content.
[0078] For step S213, in this embodiment, the natural factor joint probability distribution matrix includes: a temperature factor joint probability distribution matrix, a precipitation factor joint probability distribution matrix, a soil gravel content factor joint probability distribution matrix and a soil pH value factor joint probability distribution matrix.
[0079] Among them, the elements of the joint probability distribution matrix of the temperature factors represent the joint probability of the temperature and the historical arsenic content data; the elements of the joint probability distribution matrix of the precipitation factors represent the joint probability of the precipitation and the historical arsenic content data; the elements of the joint probability distribution matrix of the soil gravel content factors represent the joint probability of the soil gravel content and the historical arsenic content data; the elements of the joint probability distribution matrix of the soil pH value factors represent the joint probability of the soil pH value factors and the historical arsenic content data.
[0080] Of course, in other embodiments, the joint probability distribution matrix of natural factors may also include: a joint probability distribution matrix of wind speed, a joint probability distribution matrix of water vapor pressure, a joint probability distribution matrix of radiation, a joint probability distribution matrix of elevation, a joint probability distribution matrix of slope, a joint probability distribution matrix of aspect, a joint probability distribution matrix of curvature, a joint probability distribution matrix of terrain moisture index, a joint probability distribution matrix of soil moisture, a joint probability distribution matrix of soil bulk density, a joint probability distribution matrix of parent rock and parent material, a joint probability distribution matrix of soil silt content, a joint probability distribution matrix of soil clay content and a joint probability distribution matrix of soil sand content.
[0081] For step S214, in this embodiment, the probability distribution coefficients of the historical natural factor variables include: the probability distribution coefficient of temperature, the probability distribution coefficient of precipitation, the probability distribution coefficient of soil gravel content and the probability distribution coefficient of soil pH value.
[0082] Among them, the probability distribution coefficient of the temperature is obtained by accumulating the row data in the joint probability distribution matrix of the temperature factors; the probability distribution coefficient of the precipitation is obtained by accumulating the row data in the joint probability distribution matrix of the precipitation factors; the probability distribution coefficient of the soil gravel content is obtained by accumulating the row data in the joint probability distribution matrix of the soil gravel content factors; the probability distribution coefficient of the soil pH value is obtained by accumulating the row data in the joint probability distribution matrix of the soil pH value factors.
[0083] Of course, in other embodiments, the probability distribution coefficient of the historical natural factor variables may also include: the probability distribution coefficient of wind speed, the probability distribution coefficient of water vapor pressure, the probability distribution coefficient of radiation, the probability distribution coefficient of elevation, the probability distribution coefficient of slope, the probability distribution coefficient of slope aspect, the probability distribution coefficient of curvature, the probability distribution coefficient of terrain moisture index, the probability distribution coefficient of soil moisture, the probability distribution coefficient of soil bulk density, the probability distribution coefficient of parent rock and parent material, the probability distribution coefficient of soil silt content, the probability distribution coefficient of soil clay content and the probability distribution coefficient of soil sand content.
[0084] For step S215, in this embodiment, according to the joint probability of the temperature factor and the historical arsenic content data, the probability distribution coefficient of the temperature and the first arsenic content probability distribution coefficient, the maximum mutual information between the temperature and the historical arsenic content data is calculated according to the following formula:
[0085]
[0086] Wherein, MI(x;y) is the maximum mutual information between any historical natural factor variable and the historical arsenic content data, x is any historical natural factor variable, x∈X i , X i ={temperature, precipitation, soil gravel content, soil pH}, at this time MI(x; y) corresponds to the maximum mutual information between the temperature and the historical arsenic content data, jp(x, y) is the joint probability of any of the historical natural factor variables and the historical arsenic content data, at this time jp(x, y) corresponds to the joint probability of the temperature factor and the historical arsenic content data, and the maximum value of the joint probability of the temperature and the historical arsenic content data in the joint probability distribution matrix of the temperature factor is preferentially selected as jp(x, y), p(x) is the probability distribution coefficient of any of the historical natural factor variables, at this time p(x) corresponds to the probability distribution coefficient of the temperature, and p1(y) is the first arsenic content probability distribution coefficient.
[0087] The principles for calculating the maximum mutual information between the precipitation and the historical arsenic content data, the maximum mutual information between the soil gravel content and the historical arsenic content data, and the maximum mutual information between the soil pH value and the historical arsenic content data are the same as the principles for calculating the maximum mutual information between the air temperature and the historical arsenic content data, and are not repeated here.
[0088] For step S216, in this embodiment, the maximum mutual information coefficient between the air temperature and the historical arsenic content data can be calculated according to the first horizontal axis dimension, the first vertical axis dimension, and the maximum mutual information between the air temperature and the historical arsenic content data according to the following formula:
[0089] MIC(x;y)=maxMI(x;y) / log2min(m1,n1)
[0090] In the formula, MIC(x; y) is the maximum mutual information coefficient between any of the historical natural factor variables and the historical arsenic content data. At this time, MIC(x; y) corresponds to the maximum mutual information coefficient between the temperature and the historical arsenic content data. MI(x; y) is the maximum mutual information between any of the historical natural factor variables and the historical arsenic content data. At this time, MI(x; y) corresponds to the maximum mutual information between the temperature and the historical arsenic content data. m1 and n1 are the first horizontal axis dimension and the first vertical axis dimension respectively. Min(m1, n1) represents the smaller value of m1 and n1.
[0091] The principles for calculating the maximum mutual information coefficient between the precipitation and the historical arsenic content data, the maximum mutual information coefficient between the soil gravel content and the historical arsenic content data, and the maximum mutual information coefficient between the soil pH value and the historical arsenic content data are the same as the principles for calculating the maximum mutual information coefficient between the air temperature and the historical arsenic content data, and will not be repeated here.
[0092] For step S22, in one embodiment, the maximum mutual information coefficient of human factors represents the correlation between the historical human factor variables and the historical arsenic content data. The maximum mutual information coefficient of human factors includes the maximum mutual information coefficient of population density, the maximum mutual information coefficient of land use type, the maximum mutual information coefficient of inverse distance interpolation of traffic land, and the maximum mutual information coefficient of inverse distance interpolation of arsenic mine points, wherein the maximum mutual information coefficient of population density, the maximum mutual information coefficient of land use type, the maximum mutual information coefficient of inverse distance interpolation of traffic land, and the maximum mutual information coefficient of inverse distance interpolation of arsenic mine points represent the correlation between population density, land use type, inverse distance interpolation of traffic land, and inverse distance interpolation of arsenic mine points and the historical arsenic content data, respectively.
[0093] Of course, in other embodiments, the maximum mutual information coefficient of human factors may also include the maximum mutual information coefficient of nighttime lighting and the maximum mutual information coefficient of inverse distance interpolation of other mining points.
[0094] Please also see Figure 4 , Figure 4 This is a flow chart of a method for obtaining the maximum mutual information coefficient of human factors in a soil arsenic distribution monitoring method of the present application. The method of performing mutual information maximization analysis on the historical human factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of human factors also includes the following steps:
[0095] S221: Use the historical arsenic content data and any of the human factor variables as the Y axis and X axis respectively to construct a human factor scatter plot, and divide the X axis and Y axis of the human factor scatter plot into a number of human factor intervals using a preset second horizontal axis dimension and a second vertical axis dimension, where m2×n2 <s2 0.7, m2 is the second horizontal axis dimension, n2 is the second vertical axis dimension, and s2 is the number of groups of the historical human factor variables;
[0096] S222: creating a two-dimensional frequency matrix of human factors according to the second horizontal axis dimension and the second vertical axis dimension, wherein the elements of the two-dimensional frequency matrix of human factors represent the number of groups in which any of the historical human factor variables and the historical arsenic content data simultaneously fall within any of the human factor intervals;
[0097] S223: Divide each element of the two-dimensional frequency matrix of human factors by the number of groups of the historical human factor variables to obtain a joint probability distribution matrix of human factors, wherein the elements of the joint probability distribution matrix of human factors represent the joint probability of any of the historical human factor variables and the historical arsenic content data;
[0098] S224: according to the human factor joint probability distribution matrix, respectively accumulating the row data and the column data of the human factor joint probability distribution matrix to obtain the probability distribution coefficient of any of the historical human factor variables and the second arsenic content probability distribution coefficient;
[0099] S225: Calculate the maximum mutual information between the historical human factor variable and the historical arsenic content data according to the following formula based on the joint probability of the historical human factor variable and the historical arsenic content data, the probability distribution coefficient of the historical human factor variable and the second arsenic content probability distribution coefficient:
[0100]
[0101] Where MI(k; y) is the maximum mutual information between any historical human factor variable and the historical arsenic content data, x is any historical human factor variable, k∈K i , K i ={population density, land use type, inverse distance interpolation of traffic land, inverse distance interpolation of arsenic mine points}, y is the historical arsenic content data, jp(k,y) is the joint probability of any of the historical human factor variables and the historical arsenic content data, p(k) is the probability distribution coefficient of any of the historical human factor variables, and p2(y) is the second arsenic content probability distribution coefficient;
[0102] S226: Calculate the maximum mutual information coefficient between any one of the historical human factor variables and the historical arsenic content data according to the following formula based on the second horizontal axis dimension, the second vertical axis dimension, and the maximum mutual information between any one of the historical human factor variables and the historical arsenic content data:
[0103] MIC(k;y)=maxMI(k;y) / log2min(m2,n2)
[0104] In the formula, MIC(k; y) is the maximum mutual information coefficient between any of the historical human factor variables and the historical arsenic content data, MI(k; y) is the maximum mutual information between any of the historical human factor variables and the historical arsenic content data, m2 and n2 are the second horizontal axis dimension and the second vertical axis dimension respectively, and min(m2, n2) represents the smaller value of m2 and n2.
[0105] For step S221, in this embodiment, the human factor scatter plot includes: a population density factor scatter plot, a land use type factor scatter plot, a transportation land inverse distance interpolation factor scatter plot and an arsenic mine point inverse distance interpolation factor scatter plot, wherein the population density factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the population density as the X-axis; the land use type factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the land use type factor as the X-axis; the transportation land inverse distance interpolation factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the transportation land inverse distance interpolation as the X-axis; the arsenic mine point inverse distance interpolation factor scatter plot is constructed by taking the historical arsenic content data as the Y-axis and the arsenic mine point inverse distance interpolation as the X-axis.
[0106] Of course, in other embodiments, the human factor scatter plot may also include a night light scatter plot and other mining point inverse distance interpolation scatter plots.
[0107] For step S222, the two-dimensional frequency matrix of human factors includes: a two-dimensional frequency matrix of population density factors, a two-dimensional frequency matrix of land use type factors, a two-dimensional frequency matrix of inverse distance interpolation factors of transportation land, and a two-dimensional frequency matrix of inverse distance interpolation factors of arsenic mine points.
[0108] Among them, the elements of the two-dimensional frequency matrix of the population density factor represent the number of groups in which the population density and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the population density factor scatter diagram; the elements of the two-dimensional frequency matrix of the land use type factor represent the number of groups in which the land use type and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the land use type factor scatter diagram; the elements of the two-dimensional frequency matrix of the inverse distance interpolation factor of transportation land represent the number of groups in which the inverse distance interpolation of transportation land and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the inverse distance interpolation factor scatter diagram of transportation land; the elements of the two-dimensional frequency matrix of the inverse distance interpolation factor of arsenic mine point represent the number of groups in which the inverse distance interpolation of arsenic mine point and the corresponding historical arsenic content data simultaneously fall within any of the natural factor intervals in the inverse distance interpolation factor scatter diagram of arsenic mine point.
[0109] Of course, in other embodiments, the two-dimensional frequency matrix of human factors may also include a two-dimensional frequency matrix of nighttime light factors and a two-dimensional frequency matrix of other mining point inverse distance interpolation factors.
[0110] For step S223, in this embodiment, the joint probability distribution matrix of human factors includes: a joint probability distribution matrix of population density factors, a joint probability distribution matrix of land use type factors, a joint probability distribution matrix of inverse distance interpolation factors of transportation land, and a joint probability distribution matrix of inverse distance interpolation factors of arsenic mine points.
[0111] Among them, the elements of the joint probability distribution matrix of the population density factor represent the joint probability of the population density and the historical arsenic content data; the elements of the joint probability distribution matrix of the land use type factor represent the joint probability of the land use type and the historical arsenic content data; the elements of the joint probability distribution matrix of the inverse distance interpolation factors of transportation land represent the joint probability of the inverse distance interpolation of transportation land and the historical arsenic content data; the elements of the joint probability distribution matrix of the inverse distance interpolation factors of arsenic mine points represent the joint probability of the inverse distance interpolation factors of arsenic mine points and the historical arsenic content data.
[0112] Of course, in other embodiments, the human factor joint probability distribution matrix may also include: a nighttime light factor joint probability distribution matrix and other mining point inverse distance interpolation factor joint probability distribution matrices.
[0113] For step S224, in this embodiment, the probability distribution coefficients of the historical human factor variables include: the probability distribution coefficient of population density, the probability distribution coefficient of land use type, the probability distribution coefficient of inverse distance interpolation of transportation land, and the probability distribution coefficient of inverse distance interpolation of arsenic mine points.
[0114] Among them, the probability distribution coefficient of the population density is obtained by accumulating the row data in the joint probability distribution matrix of the population density factors; the probability distribution coefficient of the land use type is obtained by accumulating the row data in the joint probability distribution matrix of the land use type factors; the probability distribution coefficient of the inverse distance interpolation of transportation land is obtained by accumulating the row data in the joint probability distribution matrix of the inverse distance interpolation factors of transportation land; the probability distribution coefficient of the inverse distance interpolation of the arsenic mine point is obtained by accumulating the row data in the joint probability distribution matrix of the inverse distance interpolation factors of the arsenic mine point.
[0115] Of course, in other embodiments, the probability distribution coefficient of the man-made natural factor variable may also include: the probability distribution coefficient of night light and the probability distribution coefficient of inverse distance interpolation of other mining points.
[0116] For step S225, in this embodiment, the maximum mutual information between the population density and the historical arsenic content data can be calculated according to the joint probability of the population density and the historical arsenic content data, the probability distribution coefficient of the population density and the second arsenic content probability distribution coefficient according to the following formula:
[0117]
[0118] In the formula, MI(k; y) is the maximum mutual information between any historical human factor variable and the historical arsenic content data, x is any historical human factor variable, k∈K i , K i ={population density, land use type, inverse distance interpolation of transportation land, inverse distance interpolation of arsenic mine points}, at this time MI(k; y) corresponds to the maximum mutual information between the temperature and the historical arsenic content data, jp(k, y) is the joint probability of any of the historical human factor variables and the historical arsenic content data, at this time jp(k, y) corresponds to the joint probability of the population density and the historical arsenic content data, and the maximum value of the joint probability of the population density and the historical arsenic content data in the joint probability distribution matrix of the population density factor is preferentially selected as jp(k, y), p(k) is the probability distribution coefficient of any of the historical human factor variables, at this time p(k) corresponds to the probability distribution coefficient of the population density, and p2(y) is the second arsenic content probability distribution coefficient.
[0119] The principles for calculating the maximum mutual information between the land use type and the historical arsenic content data, the maximum mutual information between the inverse distance interpolation of transportation land and the historical arsenic content data, and the maximum mutual information between the inverse distance interpolation of arsenic mine points and the historical arsenic content data are the same as the principles for calculating the maximum mutual information between the population density and the historical arsenic content data, and will not be repeated here.
[0120] For step S226, in this embodiment, the maximum mutual information coefficient between the population density and the historical arsenic content data can be calculated according to the second horizontal axis dimension, the second vertical axis dimension, and the maximum mutual information between the population density and the historical arsenic content data according to the following formula:
[0121] MIC(k;y)=maxMI(k;y) / log2min(m2,n2)
[0122] In the formula, MIC(k; y) is the maximum mutual information coefficient between any of the historical human factor variables and the historical arsenic content data. At this time, MIC(k; y) corresponds to the maximum mutual information coefficient between the population density and the historical arsenic content data. MI(k; y) is the maximum mutual information between any of the historical human factor variables and the historical arsenic content data. At this time, MI(k; y) corresponds to the maximum mutual information between the population density and the historical arsenic content data. m2 and n2 are the second horizontal axis dimension and the second vertical axis dimension respectively. Min(m2, n2) represents the smaller value of m2 and n2.
[0123] The principles for calculating the maximum mutual information coefficient between the land use type and the historical arsenic content data, the maximum mutual information coefficient between the inverse distance interpolation of transportation land and the historical arsenic content data, and the maximum mutual information coefficient between the inverse distance interpolation of arsenic mine points and the historical arsenic content data are the same as the principles for calculating the maximum mutual information coefficient between the population density and the historical arsenic content data, and will not be repeated here.
[0124] See also Figure 5 , Figure 5 The present invention is a flow chart of a method for obtaining significant probability values of natural factors and significant probability values of human factors in a soil arsenic distribution monitoring method of the present invention. For step S23, the significant probability values of natural factors and significant probability values of human factors respectively represent the significance level of the historical natural factor variable and the historical arsenic content data and the significance level of the historical human factor variable and the historical arsenic content data.
[0125] The method of performing significance tests on the plurality of groups of historical natural factor variables, historical human factor variables and the historical arsenic content data to obtain significant probability values of natural factors and significant probability values of human factors also includes the following steps:
[0126] S231: Obtaining estimated values of multiple natural factor regression coefficients and estimated values of multiple human factor regression coefficients based on a fitting regression model according to the multiple groups of historical natural factor variables and the historical arsenic content data, and the multiple groups of historical human factor variables and the historical arsenic content data;
[0127] S232: Calculate the standard error of the natural factor regression coefficient, the standard error of the human factor regression coefficient, the average value of the natural factor regression coefficient, and the average value of the human factor regression coefficient according to the estimated values of the multiple natural factor regression coefficients and the estimated values of the multiple human factor regression coefficients;
[0128] S233: According to the average value of the natural factor regression coefficient and the average value of the human factor regression coefficient, as well as the standard error of the natural factor regression coefficient and the standard error of the human factor regression coefficient, the t statistic of the natural factor regression coefficient and the t statistic of the human factor regression coefficient are calculated according to the following formula, and they are used as the significant probability value of the natural factor and the significant probability value of the human factor, respectively:
[0129]
[0130] Where, t n is the t statistic of the natural factor regression coefficient, b1 is the average value of the natural factor regression coefficient, β1 is the overall value of the natural factor regression coefficient, under the null hypothesis, β1=0; σ1 is the standard error of the natural factor regression coefficient, t p is the t statistic of the human factor regression coefficient, b2 is the average value of the human factor regression coefficient, β2 is the overall value of the human factor regression coefficient, and under the null hypothesis, β2=0; σ2 is the standard error of the human factor regression coefficient.
[0131] For step S231, in this embodiment, the fitting regression model is preferably a linear function, and the best estimated value of the natural factor regression coefficient and the estimated value of the human factor regression coefficient are obtained by the least square method. The estimated value of the natural factor regression coefficient and the estimated value of the human factor regression coefficient respectively represent the intensity and direction of the influence of the historical natural factor variable on the historical arsenic content data, and the intensity and direction of the influence of the historical human factor variable on the historical arsenic content data.
[0132] The natural factor regression coefficients include: temperature factor regression coefficient, precipitation factor regression coefficient, soil gravel content factor regression coefficient and soil pH factor regression coefficient. The human factor regression coefficients include: population density factor regression coefficient, land use type factor regression coefficient, traffic land inverse distance interpolation factor regression coefficient and arsenic mine point inverse distance interpolation factor regression coefficient.
[0133] Of course, in other embodiments, the natural factor regression coefficient also includes: wind speed regression coefficient, water vapor pressure regression coefficient, radiation regression coefficient, elevation regression coefficient, slope regression coefficient, aspect regression coefficient, curvature regression coefficient, terrain moisture index regression coefficient, soil moisture regression coefficient, soil bulk density regression coefficient, parent rock and parent material regression coefficient, soil silt content regression coefficient, soil clay content regression coefficient and soil sand content regression coefficient.
[0134] The human factor regression coefficient also includes: night light regression coefficient and other mining point inverse distance interpolation regression coefficient.
[0135] The fitting regression model may also be selected from nonlinear functions such as quadratic nonlinear functions and cubic nonlinear functions, as well as functions such as polynomials, exponential functions, power functions, trigonometric functions and logarithmic functions.
[0136] For step S233, in this embodiment, the significant probability values of natural factors include: significant probability values of temperature factors, significant probability values of precipitation factors, significant probability values of soil gravel content factors, and significant probability values of soil pH factors. The significant probability values of human factors include: significant probability values of population density factors, significant probability values of land use type factors, significant probability values of inverse distance interpolation factors of traffic land, and significant probability values of inverse distance interpolation factors of arsenic mine points.
[0137] Of course, in other embodiments, the significant probability values of natural factors also include: significant probability value of wind speed, significant probability value of water vapor pressure, significant probability value of radiation, significant probability value of elevation, significant probability value of slope, significant probability value of slope aspect, significant probability value of curvature, significant probability value of terrain moisture index, significant probability value of soil moisture, significant probability value of soil bulk density, significant probability value of parent rock and parent material, significant probability value of soil silt content, significant probability value of soil clay content and significant probability value of soil sand content.
[0138] The significant probability values of human factors also include: significant probability values of night lights and significant probability values of inverse distance interpolation of other mining points.
[0139] For step S24, in one embodiment, the preset maximum mutual information coefficient threshold is preferably set to 0.5, and the preset significant probability threshold is preferably set to 95%.
[0140] The screening of the historical natural factor variables and the historical human factor variables corresponding to the maximum mutual information coefficient of the natural factor and the maximum mutual information coefficient of the human factor being greater than the preset maximum mutual information coefficient threshold, and the significant probability value of the natural factor and the significant probability value of the human factor being greater than the preset significant probability threshold, is specifically as follows:
[0141] The maximum mutual information coefficient of temperature, the maximum mutual information coefficient of precipitation, the maximum mutual information coefficient of soil gravel content, the maximum mutual information coefficient of soil pH value, the maximum mutual information coefficient of population density, the maximum mutual information coefficient of land use type, the maximum mutual information coefficient of inverse distance interpolation of transportation land and the maximum mutual information coefficient of inverse distance interpolation of arsenic mine points are respectively screened out, and the temperature, precipitation, soil gravel content, soil pH value, population density, land use type, inverse distance interpolation of transportation land and inverse distance interpolation of arsenic mine points corresponding to the maximum mutual information coefficient threshold are greater than the maximum mutual information coefficient threshold.
[0142] At the same time, the significant probability values of temperature factors, precipitation factors, soil gravel content factors, soil pH value factors, population density factors, land use type factors, transportation land inverse distance interpolation factors and arsenic mine point inverse distance interpolation factors are screened out respectively, and the significant probability values of the temperature factor, precipitation factor, soil gravel content, soil pH value, population density, land use type, transportation land inverse distance interpolation and arsenic mine point inverse distance interpolation are greater than the significant probability threshold corresponding to the temperature, precipitation, soil gravel content, soil pH value, population density, land use type, transportation land inverse distance interpolation and arsenic mine point inverse distance interpolation.
[0143] Of course, in other embodiments, the maximum mutual information coefficient threshold and the significant probability threshold may also be adaptively adjusted.
[0144] For step S3, in this embodiment, the partial correlation coefficient of natural factors represents the linear relationship between the historical natural factor variable and the historical arsenic content data; the partial correlation coefficient of human factors represents the linear relationship between the historical human factor variable and the historical arsenic content data.
[0145] The partial correlation coefficients of natural factors include: partial correlation coefficients of temperature factors, partial correlation coefficients of precipitation factors, partial correlation coefficients of soil gravel content factors, and partial correlation coefficients of soil pH value factors. The partial correlation coefficients of human factors include: partial correlation coefficients of population density factors, partial correlation coefficients of land use type factors, partial correlation coefficients of inverse distance interpolation factors of transportation land, and partial correlation coefficients of inverse distance interpolation factors of arsenic mine points.
[0146] See also Figure 6 , Figure 6 This is a flow chart of a method for obtaining partial correlation coefficients of natural factors and partial correlation coefficients of human factors in a soil arsenic distribution monitoring method of the present application. The linear correlation between the historical natural factor variables and historical human factor variables after screening and the historical arsenic content data is analyzed separately to obtain the partial correlation coefficients of natural factors and partial correlation coefficients of human factors, and the following steps are also included:
[0147] S31: constructing a factor set X according to the historical arsenic content data, temperature, precipitation, soil gravel content, soil pH value, population density, land use type, inverse distance interpolation of transportation land and inverse distance interpolation of arsenic mine points in sequence;
[0148] S32: According to the factor set X, the partial correlation coefficient of the temperature factor, the partial correlation coefficient of the precipitation factor, the partial correlation coefficient of the soil gravel content factor, the partial correlation coefficient of the soil pH value factor, the partial correlation coefficient of the population density factor, the partial correlation coefficient of the land use type factor, the partial correlation coefficient of the inverse distance interpolation factor of the traffic land, and the partial correlation coefficient of the inverse distance interpolation factor of the arsenic mine point are calculated respectively according to the following formula:
[0149]
[0150] Wherein, X1 represents the historical arsenic content data, To exclude X j After factoring, X i The partial correlation coefficient with the historical arsenic content data, that is, the partial correlation coefficient of the temperature factor, 1<i≤k, 1<j≤k, i≠j, k is the size of the factor set X; The historical arsenic content data and X i The correlation coefficient of The historical arsenic content data and X j The correlation coefficient of For X i With X j The correlation coefficient of .
[0151] For steps S31 and S32, in this embodiment, X = {historical arsenic content data, temperature, precipitation, soil gravel content, soil pH value, population density, land use type, inverse distance interpolation of traffic land, inverse distance interpolation of arsenic mine points}. Taking the calculation of the partial correlation coefficient of the temperature factor as an example, i is 2 at this time, and the influence of factors such as precipitation, soil gravel content, soil pH value, population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points on the partial correlation coefficient of the temperature factor needs to be excluded, that is, j = 3, 4..., k, and k is 9.
[0152] The historical arsenic content data and X i The correlation coefficient between the historical arsenic content data and X j The correlation coefficient and the X i With X j The correlation coefficients are preferably Pearson correlation coefficients, which can be calculated using the statistical software Python.
[0153] Of course, in other embodiments, the partial correlation coefficients of natural factors also include: partial correlation coefficient of wind speed, partial correlation coefficient of water vapor pressure, partial correlation coefficient of radiation, partial correlation coefficient of elevation, partial correlation coefficient of slope, partial correlation coefficient of aspect, partial correlation coefficient of curvature, partial correlation coefficient of terrain moisture index, partial correlation coefficient of soil moisture, partial correlation coefficient of soil bulk density, partial correlation coefficient of parent rock and parent material, partial correlation coefficient of soil silt content, partial correlation coefficient of soil clay content and partial correlation coefficient of soil sand content.
[0154] The human factor partial correlation coefficient also includes: night light partial correlation coefficient and other mining point inverse distance interpolation partial correlation coefficient.
[0155] The historical arsenic content data and X i The correlation coefficient between the historical arsenic content data and X j The correlation coefficient and the Xi With X j The correlation coefficients can also be the Spearman rank correlation coefficients.
[0156] See also Figure 7 , Figure 7 This is a flow chart of a method for constructing a soil arsenic distribution monitoring model in a soil arsenic distribution monitoring method of the present application. For step S5, the soil arsenic distribution monitoring model is constructed based on the symbolic regression method according to the historical natural factor variables, historical human factor variables and historical arsenic content data, and the following steps are also included:
[0157] S51: Setting a symbol search space, wherein the symbol search space includes mathematical operation symbols, input variable symbols, and constant symbols;
[0158] S52: taking the mathematical operation symbol as a parent node, and setting a corresponding number of leaf node vacancies on the parent node according to the operation type of the mathematical operation symbol;
[0159] S53: taking the variable symbols and constant symbols as leaf nodes and inserting them into the leaf node vacancies, thereby constructing an expression binary tree;
[0160] S54: Perform a pre-order traversal on the expression binary tree to generate a model to be optimized;
[0161] S55: Setting a maximum number of iterations, and iteratively optimizing the model to be optimized according to the historical natural factor variables and the historical human factor variables to obtain the soil arsenic element distribution monitoring model.
[0162] For step S51, in this embodiment, the mathematical operation symbols include four arithmetic operators such as addition, subtraction, multiplication and division, trigonometric function operators, natural exponential functions, logarithmic functions, square root operations and power operations. The input variable symbols are used to represent the historical natural factor variables and the historical human factor variables. The constant symbol represents a fixed value in the soil arsenic element distribution monitoring model, for example, the constant symbol can be an offset in the soil arsenic element distribution monitoring model. Of course, the mathematical operators can be adaptively adjusted according to actual needs.
[0163] For step S52, in one embodiment, the operation type of the mathematical operation symbol includes a unary operator and a binary operator. For example, when the mathematical operator is a four arithmetic operator, since the four arithmetic operators are binary operators, two corresponding leaf node spaces need to be set.
[0164] Please also see Figure 8 , Figure 8This is a flow chart of a method for iteratively optimizing a model to be optimized in a soil arsenic distribution monitoring method of the present application. For step S55, the maximum number of iterations is set, and the model to be optimized is iteratively optimized according to the historical natural factor variables and the historical human factor variables to obtain the soil arsenic distribution monitoring model, and the following steps are also included:
[0165] S551: Selecting a plurality of variable data sets from the variable database, and dividing the variable data sets into a training set and a validation set in a ratio of 7:3, wherein both the training set and the validation set include the historical natural factor variables, the historical human factor variables, and the historical arsenic content data;
[0166] S552: using the historical natural factor variables and the historical human factor variables of the training set as input data, and obtaining the arsenic content prediction data corresponding to the training set based on the model to be optimized;
[0167] S553: using a loss function, calculating the error between the historical arsenic content data of the training set and the predicted arsenic content data corresponding to the training set as a training set loss value;
[0168] S554: adjusting the parent node and the leaf node of the expression binary tree according to the training set loss value and the symbol search space, performing a pre-order traversal on the adjusted expression binary tree, and obtaining a soil arsenic element distribution monitoring adjustment model;
[0169] S555: using the historical natural factor variables and the historical human factor variables of the verification set as input data, and obtaining the arsenic content prediction data of the verification set based on the soil arsenic distribution monitoring and adjustment model;
[0170] S556: using a loss function, calculating the error between the historical arsenic content data of the validation set and the predicted arsenic content data of the validation set as a validation set loss value;
[0171] S557: When the validation set loss value is less than a predetermined validation loss threshold, the soil arsenic distribution monitoring model is obtained.
[0172] For steps S553-S557, in this embodiment, the loss function is preferably the root mean square error RMSE, and the training set loss value and the validation set loss value are both the mean square error of the training set and the mean square error of the validation set. When the validation set loss value is less than the predetermined validation loss threshold, steps S552-S557 are repeated, and when the number of executions is greater than the maximum number of iterations, the iterative optimization is completed. Among them, the validation loss threshold is preferably set to 1e-6. The maximum number of iterations is preferably set to 500 times.
[0173] The method of adjusting the parent node and leaf nodes of the expression binary tree according to the training set loss value and the symbol search space includes: obtaining mathematical operation symbols from the symbol search space and replacing the mathematical operation symbols in the original parent node of the expression binary tree; obtaining different constant symbols from the symbol search space and replacing the constant symbols in the original child nodes of the expression binary tree; taking the mathematical operation symbols as the parent node, and setting a corresponding number of leaf node vacancies on the parent node according to the operation type of the mathematical operation symbols, taking the variable symbols and constant symbols as leaf nodes, and inserting them into the leaf node vacancies, thereby constructing an expression sub-binary tree, and adding the expression sub-binary tree to the original expression binary tree; deleting the parent node or leaf node of the original expression binary tree.
[0174] Of course, in other embodiments, the verification loss threshold and the maximum number of iterations may also be adaptively modified.
[0175] Example 2
[0176] See also Fig. 9 , Fig. 9 This is a schematic diagram of a soil arsenic distribution monitoring system of the present application. The present application also provides a soil arsenic distribution monitoring system, including:
[0177] Variable data acquisition module 1: used to obtain several groups of historical natural factor variables, historical human factor variables and historical arsenic content data of the target monitoring area from a preset variable database, wherein the historical natural factor variables at least include: temperature, precipitation, soil gravel content and soil pH value; the historical human factor variables at least include: population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points;
[0178] Variable screening module 2: used to analyze and screen the correlation between the historical natural factor variable and the historical human factor variable and the historical arsenic content data, respectively, to obtain the screened historical natural factor variable and historical human factor variable;
[0179] Partial correlation coefficient calculation module 3: used to analyze the linear correlation between the historical natural factor variables and the historical human factor variables after screening and the historical arsenic content data, and obtain the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors;
[0180] Weight calculation module 4: used to divide the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors by the sum of the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors, respectively, to obtain the weight of natural factors and the weight of human factors;
[0181] Model building module 5: used to build a soil arsenic distribution monitoring model based on the symbolic regression method according to the historical natural factor variables, historical human factor variables and historical arsenic content data;
[0182] Actual arsenic content calculation module 6: used to obtain the actual historical natural factor variables and actual human factor variables of the target monitoring area from the variable database, substitute the actual historical natural factor variables and actual human factor variables into the soil arsenic distribution monitoring model, and perform weighted summation according to the natural factor weight and the human factor weight to obtain the actual soil arsenic content.
[0183] It should be noted that the data obtained by the soil arsenic distribution monitoring system provided in the present application when implementing a soil arsenic distribution monitoring method are stored one-to-one in the storage of the system. When relevant calculations are required, the data required for the calculation can be directly obtained from the storage.
[0184] It should also be noted that the soil arsenic distribution monitoring system provided in the above embodiment only uses the division of the above functional modules as an example when implementing a soil arsenic distribution monitoring method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the soil arsenic distribution monitoring system provided in the above embodiment and the soil arsenic distribution monitoring method of embodiment 1 belong to the same concept. The implementation process is detailed in the method embodiment, which will not be repeated here.
[0185] Based on the same inventive concept, the present application also provides an electronic device, which may be a terminal device such as a server, a desktop computing device or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). The device includes one or more processors and a memory, wherein the processor is used to execute a program to implement the soil arsenic distribution monitoring method; and the memory is used to store a computer program executable by the processor.
[0186] The present application may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be a computer-executable instruction, and the computer-executable instruction can execute the soil arsenic distribution monitoring method. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0187] The present application is not limited to the above-mentioned implementation modes. If various changes or modifications to the present application do not depart from the spirit and scope of the present application, and if these changes and modifications fall within the claims of the present application and the scope of equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for monitoring the distribution of arsenic in soil, characterized in that: The following steps are involved: Acquire several groups of historical natural factor variables, historical human factor variables and historical arsenic content data of the target monitoring area from a preset variable database, wherein the historical natural factor variables include at least: temperature, precipitation, soil gravel content and soil pH value; the historical human factor variables include at least: population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points; Respectively analyzing and screening the correlation between the historical natural factor variables and the historical human factor variables and the historical arsenic content data to obtain screened historical natural factor variables and historical human factor variables; Respectively analyzing the linear correlation between the screened historical natural factor variables and historical human factor variables and the historical arsenic content data to obtain the partial correlation coefficient of the natural factor and the partial correlation coefficient of the human factor; Respectively dividing the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors by the sum of the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors to obtain the weight of natural factors and the weight of human factors; According to the historical natural factor variables, historical human factor variables and historical arsenic content data, a soil arsenic element distribution monitoring model is constructed based on a symbolic regression method, including: setting a symbol search space, wherein the symbol search space includes mathematical operation symbols, input variable symbols and constant symbols; Taking the mathematical operation symbol as a parent node, and setting a corresponding number of leaf node vacancies on the parent node according to the operation type of the mathematical operation symbol; The variable symbols and constant symbols are used as leaf nodes and inserted into the leaf node vacancies, thereby constructing an expression binary tree; Performing a pre-order traversal on the expression binary tree to generate a model to be optimized; Setting a maximum number of iterations, iteratively optimizing the model to be optimized according to the historical natural factor variables and the historical human factor variables, and obtaining the soil arsenic element distribution monitoring model; The actual historical natural factor variables and the actual human factor variables of the target monitoring area are obtained from the variable database, the actual historical natural factor variables and the actual human factor variables are substituted into the soil arsenic distribution monitoring model, and weighted summation is performed according to the natural factor weight and the human factor weight to obtain the actual soil arsenic content.
2. The method for monitoring soil arsenic distribution according to claim 1, characterized in that: The setting of the maximum number of iterations, iteratively optimizing the model to be optimized according to the historical natural factor variables and the historical human factor variables to obtain the soil arsenic element distribution monitoring model, also includes the following steps: Selecting a plurality of variable data sets from the variable database, dividing the variable data sets into a training set and a validation set in a ratio of 7:3, wherein both the training set and the validation set include the historical natural factor variables, the historical human factor variables and the historical arsenic content data; Using the historical natural factor variables and the historical human factor variables of the training set as input data, and obtaining the arsenic content prediction data corresponding to the training set based on the model to be optimized; Using a loss function, calculating the error between the historical arsenic content data of the training set and the predicted arsenic content data corresponding to the training set as a training set loss value; According to the training set loss value and the symbol search space, the parent node and the leaf node of the expression binary tree are adjusted, and the adjusted expression binary tree is traversed in pre-order to obtain a soil arsenic element distribution monitoring adjustment model; Using the historical natural factor variables and historical human factor variables of the verification set as input data, and based on the soil arsenic element distribution monitoring and adjustment model, obtaining the arsenic content prediction data of the verification set; Using a loss function, calculating the error between the historical arsenic content data of the validation set and the predicted arsenic content data of the validation set as a validation set loss value; When the validation set loss value is less than a predetermined validation loss threshold, the soil arsenic element distribution monitoring model is obtained.
3. The method for monitoring soil arsenic distribution according to claim 1, characterized in that: The step of analyzing and screening the correlation between the historical natural factor variables and the historical human factor variables and the historical arsenic content data to obtain the screened historical natural factor variables and historical human factor variables also includes the following steps: Performing mutual information maximization analysis on the historical natural factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of the natural factor; Performing mutual information maximization analysis on the historical human factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of the human factor; Conducting significance tests on multiple groups of historical natural factor variables, historical human factor variables and the historical arsenic content data respectively, and obtaining significant probability values of natural factors and significant probability values of human factors; The historical natural factor variables and historical human factor variables corresponding to the maximum mutual information coefficient of natural factors and the maximum mutual information coefficient of human factors are greater than the preset maximum mutual information coefficient threshold, and the significant probability value of natural factors and the significant probability value of human factors are greater than the preset significant probability threshold are screened.
4. The method for monitoring soil arsenic distribution according to claim 3, characterized in that: The performing of mutual information maximization analysis on the historical natural factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of the natural factors also includes the following steps: The historical arsenic content data and any of the historical natural factor variables are used as the Y axis and the X axis respectively to construct a natural factor scatter plot, and the X axis and the Y axis of the natural factor scatter plot are divided into a number of natural factor intervals according to the preset first horizontal axis dimension and the first vertical axis dimension, where m1×n1 <s1 0.7 , m1 is the first horizontal axis dimension, n1 is the first vertical axis dimension, and s1 is the number of groups of the historical natural factor variables; Creating a two-dimensional frequency matrix of natural factors according to the first horizontal axis dimension, the first vertical axis dimension and the natural factor scatter plot, wherein the elements of the two-dimensional frequency matrix of natural factors represent the number of groups in which any of the historical natural factor variables and the historical arsenic content data simultaneously fall within any of the natural factor intervals; Dividing each element of the two-dimensional frequency matrix of natural factors by the number of groups of the historical natural factor variables, constructing a natural factor joint probability distribution matrix, wherein the elements of the natural factor joint probability distribution matrix represent the joint probability of any of the historical natural factor variables and the historical arsenic content data; According to the natural factor joint probability distribution matrix, respectively accumulating the row data and column data of the natural factor joint probability distribution matrix to obtain the probability distribution coefficient of any of the historical natural factor variables and the first arsenic content probability distribution coefficient; According to the joint probability of the historical natural factor variable and the historical arsenic content data, the probability distribution coefficient of the historical natural factor variable and the first arsenic content probability distribution coefficient, the maximum mutual information between the historical natural factor variable and the historical arsenic content data is calculated according to the following formula: Wherein, MI(x;y) is the maximum mutual information between any historical natural factor variable and the historical arsenic content data, x is any historical natural factor variable, x∈X i , X i ={temperature, precipitation, soil gravel content, soil pH}, y is the historical arsenic content data, jp(x,y) is the joint probability of any of the historical natural factor variables and the historical arsenic content data, p(x) is the probability distribution coefficient of any of the historical natural factor variables, and p1(y) is the first arsenic content probability distribution coefficient; According to the first horizontal axis dimension, the first vertical axis dimension and the maximum mutual information between any of the historical natural factor variables and the historical arsenic content data, the maximum mutual information coefficient between any of the historical natural factor variables and the historical arsenic content data is calculated according to the following formula: In the formula, MIC(x; y) is the maximum mutual information coefficient between any of the historical natural factor variables and the historical arsenic content data, MI(x; y) is the maximum mutual information between any of the historical natural factor variables and the historical arsenic content data, m1 and n1 are the first horizontal axis dimension and the first vertical axis dimension respectively, and min(m1, n1) represents the smaller value of m1 and n1.
5. The method for monitoring soil arsenic distribution according to claim 3, characterized in that: The performing of mutual information maximization analysis on the historical human factor variables and the historical arsenic content data to obtain the maximum mutual information coefficient of the human factor also includes the following steps: The historical arsenic content data and any of the human factor variables are used as the Y axis and the X axis respectively to construct a human factor scatter plot, and the X axis and the Y axis of the human factor scatter plot are divided into a number of human factor intervals by a preset second horizontal axis dimension and a second vertical axis dimension, where m2×n2 <s2 0.7 , m2 is the second horizontal axis dimension, n2 is the second vertical axis dimension, and s2 is the number of groups of the historical human factor variables; According to the second horizontal axis dimension and the second vertical axis dimension, a two-dimensional frequency matrix of human factors is created, wherein the elements of the two-dimensional frequency matrix of human factors represent the number of groups in which any of the historical human factor variables and the historical arsenic content data simultaneously fall within any of the human factor intervals; Dividing each element of the two-dimensional frequency matrix of human factors by the number of groups of the historical human factor variables, to obtain a joint probability distribution matrix of human factors, wherein the elements of the joint probability distribution matrix of human factors represent the joint probability of any of the historical human factor variables and the historical arsenic content data; According to the human factor joint probability distribution matrix, the row data and the column data of the human factor joint probability distribution matrix are accumulated respectively to obtain the probability distribution coefficient of any of the historical human factor variables and the second arsenic content probability distribution coefficient; According to the joint probability of the historical human factor variable and the historical arsenic content data, the probability distribution coefficient of the historical human factor variable and the second arsenic content probability distribution coefficient, the maximum mutual information between the historical human factor variable and the historical arsenic content data is calculated according to the following formula: Where MI(k; y) is the maximum mutual information between any historical human factor variable and the historical arsenic content data, x is any historical human factor variable, k∈K i , K i ={population density, land use type, inverse distance interpolation of traffic land, inverse distance interpolation of arsenic mine points}, y is the historical arsenic content data, jp(k,y) is the joint probability of any of the historical human factor variables and the historical arsenic content data, p(k) is the probability distribution coefficient of any of the historical human factor variables, and p2(y) is the second arsenic content probability distribution coefficient; According to the second horizontal axis dimension, the second vertical axis dimension and the maximum mutual information between any of the historical human factor variables and the historical arsenic content data, the maximum mutual information coefficient between any of the historical human factor variables and the historical arsenic content data is calculated according to the following formula: In the formula, MIC(k; y) is the maximum mutual information coefficient between any of the historical human factor variables and the historical arsenic content data, MI(k; y) is the maximum mutual information between any of the historical human factor variables and the historical arsenic content data, m2 and n2 are the second horizontal axis dimension and the second vertical axis dimension respectively, and min(m2, n2) represents the smaller value of m2 and n2.
6. The method for monitoring soil arsenic distribution according to claim 3, characterized in that: The method of performing significance tests on the plurality of groups of historical natural factor variables, historical human factor variables and the historical arsenic content data to obtain significant probability values of natural factors and significant probability values of human factors also includes the following steps: According to the multiple groups of historical natural factor variables and the historical arsenic content data, and the multiple groups of historical human factor variables and the historical arsenic content data, based on a fitting regression model, correspondingly obtain estimated values of multiple natural factor regression coefficients and estimated values of multiple human factor regression coefficients; Calculate the standard error of the natural factor regression coefficient, the standard error of the human factor regression coefficient, the average value of the natural factor regression coefficient, and the average value of the human factor regression coefficient according to the estimated values of the multiple natural factor regression coefficients and the estimated values of the multiple human factor regression coefficients; According to the average value of the natural factor regression coefficient and the average value of the human factor regression coefficient, as well as the standard error of the natural factor regression coefficient and the standard error of the human factor regression coefficient, the t statistic of the natural factor regression coefficient and the t statistic of the human factor regression coefficient are calculated according to the following formula, and are used as the significant probability value of the natural factor and the significant probability value of the human factor, respectively: Where, t n is the t statistic of the natural factor regression coefficient, b1 is the average value of the natural factor regression coefficient, β1 is the overall value of the natural factor regression coefficient, under the null hypothesis, β1=0; σ1 is the standard error of the natural factor regression coefficient, t p is the t statistic of the human factor regression coefficient, b2 is the average value of the human factor regression coefficient, β2 is the overall value of the human factor regression coefficient, and under the null hypothesis, β2=0; σ2 is the standard error of the human factor regression coefficient.
7. A soil arsenic distribution monitoring system, characterized in that: include: Variable data acquisition module: used to obtain several groups of historical natural factor variables, historical human factor variables and historical arsenic content data of the target monitoring area from a preset variable database, wherein the historical natural factor variables include at least: temperature, precipitation, soil gravel content and soil pH value; the historical human factor variables include at least: population density, land use type, inverse distance interpolation of traffic land and inverse distance interpolation of arsenic mine points; Variable screening module: used to analyze and screen the correlation between the historical natural factor variable and the historical human factor variable and the historical arsenic content data, respectively, to obtain the screened historical natural factor variable and historical human factor variable; Partial correlation coefficient calculation module: used to analyze the linear correlation between the historical natural factor variables and the historical human factor variables after screening and the historical arsenic content data, and obtain the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors; A weight calculation module: used for respectively dividing the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors by the sum of the partial correlation coefficient of natural factors and the partial correlation coefficient of human factors to obtain the weight of natural factors and the weight of human factors; Model building module: used to build a soil arsenic distribution monitoring model based on the symbolic regression method according to the historical natural factor variables, historical human factor variables and historical arsenic content data; Actual arsenic content calculation module: used to obtain the actual historical natural factor variables and actual human factor variables of the target monitoring area from the variable database, substitute the actual historical natural factor variables and actual human factor variables into the soil arsenic distribution monitoring model, and perform weighted summation according to the natural factor weight and the human factor weight to obtain the actual soil arsenic content.
8. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a method for monitoring the distribution of arsenic in soil as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing computer-executable instructions, characterized in that: The computer executable instructions are used in a soil arsenic distribution monitoring method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for constructing soil heavy metal environment rick prediction model
CN108647826A
Soil heavy metal spatial interpolation method and device and computer readable storage medium
CN113012771A