Large-scale range rice arsenic content high-resolution spatial distribution method and system
The arsenic content data of rice is standardized through a hierarchical adaptive standardization model and a multivariate spatiotemporal alignment function, and a high-resolution spatial distribution map of rice arsenic content is generated, which improves the comparability of the data and the reliability of the analysis results.
Patent Information
- Application Number
- CN202510927045.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-07
AI Technical Summary
The existing multi-source data fusion method faces fusion difficulties caused by differences in spatiotemporal resolution, data format and quality when processing heterogeneous data, and lacks uncertainty assessment of predicted results, which affects the reliability of management decisions.
The data in the rice arsenic content database were standardized using a hierarchical adaptive standardization model. Through multi-dimensional spatiotemporal alignment function and dynamic standardization technology, data from different sources were aligned to the same temporal and spatial resolution, and a spatial distribution map of the arsenic content in rice was generated in combination with GIS software.
It realizes effective integration and standardized processing of data from different sources, improves the integration and comparability of data, improves the accuracy and reliability of analysis results, and provides a high-reliability data foundation for subsequent arsenic pollution risk assessment.
Smart Images

Figure CN120408544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of spatial distribution generation, and particularly to a method and system for high-resolution spatial distribution of arsenic content in rice over a large-scale range. Background Art
[0002] With the intensification of global climate change and human activities, the problem of arsenic content in rice has become increasingly prominent. With the development of remote sensing technology, geographic information system, and data mining technology, multi-source data fusion methods have been widely applied in soil nutrient analysis. These methods have achieved large-scale and high-frequency monitoring and assessment of soil nutrient status by combining multi-source information such as satellite remote sensing images, terrain data, and meteorological data. However, existing multi-source data fusion methods still face many challenges in processing heterogeneous data. First, data from different sources have significant differences in spatio-temporal resolution, data format, and quality, resulting in difficulties in the data fusion process. Second, traditional data processing methods are difficult to effectively capture the spatio-temporal variation characteristics and non-linear relationships of soil nutrients. Third, existing analysis methods often lack an assessment of the uncertainty of prediction results, affecting the reliability of management decisions. Summary of the Invention
[0003] One of the purposes of the present invention is to provide a method for high-resolution spatial distribution of arsenic content in rice over a large-scale range, so as to solve the problem that data from different sources have significant differences in spatio-temporal resolution, data format, and quality in the prior art, resulting in difficulties in the data fusion process.
[0004] The present invention is realized by the following technical solutions. A method for high-resolution spatial distribution of arsenic content in rice over a large-scale range includes the following steps: S100, collecting data related to arsenic content in rice and integrating the collected data into a database of arsenic content in rice; S200, performing standardization processing on the data in the database of arsenic content in rice through a hierarchical adaptive standardization model, which fuses the data by considering the characteristics of different data and aligns them to the same temporal and spatial resolution through an interpolation method to achieve temporal synchronization and spatial matching of the data; S300, importing the data set processed by the hierarchical adaptive standardization model into GIS software, and combining with a spatial interpolation algorithm to generate a spatial distribution map of arsenic content in rice in the GIS software.
[0005] Furthermore, the data related to arsenic content in rice includes: climate change data, remote sensing data, environmental factor data, and historical arsenic content data in rice.
[0006] Furthermore, the climate change data includes: historical climate data, which consists of temperature, precipitation, humidity, CO2 concentration, frequency and intensity of extreme weather events; future climate scenario data, which is based on the global / regional climate model prediction data of institutions such as IPCC; surface hydrological data, which consists of surface runoff, soil moisture, and groundwater level data.
[0007] Furthermore, the remote sensing data includes: high-resolution optical remote sensing data, which is obtained by extracting rice physiological parameters such as vegetation index, chlorophyll content, and canopy temperature from relevant remote sensing data purchased from commercial satellite websites; synthetic aperture radar data, which consists of synthetic aperture radar data purchased from commercial satellite websites and is used to monitor rice planting area, growth stage, and flooding situation; hyperspectral remote sensing data, where specific bands are more sensitive to the physiological changes of plants that may be caused by arsenic stress.
[0008] Furthermore, the environmental factor data includes: soil data, which consists of soil type, pH value, organic matter content, redox potential, and heavy metal background value; water body data, which consists of irrigation water quality (arsenic content, pH, etc.) and arsenic content in groundwater; terrain data, including elevation, slope, and water system distribution, etc.
[0009] Furthermore, the historical rice arsenic content data includes: field sampling data, which is the arsenic content data of rice grains with clear geographical coordinates and sampling time; literature data / regional research reports, which are used to integrate existing regional survey data.
[0010] Furthermore, the hierarchical adaptive standardization model includes the following sub-steps: S210. By weighting data from different sources, align the data from different spatial positions and time points to a unified coordinate system, so that the data at each observation point is weighted according to the spatio-temporal distance. The observation points closer to the target point will have a greater impact on the final result. Combine the Gaussian kernel function to reduce the impact of the observation points far from the target point on the result by attenuation. By setting credibility weights for the observation points, different weights are given to data of different qualities during the alignment process to construct a multi-source spatio-temporal alignment function; S220. After completing the spatio-temporal alignment of the data through the multi-source spatio-temporal alignment function, perform dynamic standardization on the data. The dynamic standardization converts the rice arsenic content data of different types and dimensions of variables into dimensionless values by comprehensively considering the significant differences in the rice arsenic content data affected by climate, soil, vegetation cover, and human activities in different regions and at different times, and eliminates the differences in data dimension and scale.
[0011] Furthermore, the multi-source spatio-temporal alignment function is expressed by the following formula:
[0012] , where, is the Multi - temporal and spatial alignment function for class data sources is the target spatio - temporal point are the abscissa and ordinate in the spatial position is the time point is the data of a specific observation point is the set of data of all observation points is the th data credibility weight of the observation point is the Gaussian kernel function symbol is the weighted Euclidean distance is the anisotropic Gaussian kernel function is the original observation value matrix
[0013] Furthermore, the weighted Euclidean distance can be expressed by the following formula:
[0014] , where is the time - dimension scaling factor is the target spatio - temporal point is the original spatio - temporal point
[0015] Furthermore, the dynamic standardization is expressed by the following formula:
[0016] , where is the dynamic standardization value represents the standardized value of the th class variable at the spatial position ; is the original observation value represents the original observation value of the th class variable at the spatial position ; is the climate zone is the dynamic mean of the climate zone is the standard deviation of the climate zone is the vegetation coverage correction factor is the human activity compensation factor
[0017] Furthermore, the dynamic mean of the climate zone can be expressed by the following formula:
[0018] , where is the number of data points included in the climate zone is the transformation function symbol is the transformation function is the climate weight, used to measure the representativeness or importance of the spatial position in its affiliated climate zone ;
[0019] Furthermore, the standard deviation of the climate zoning can be expressed by the following formula:
[0020] , where is the soil weight, which is used to measure the influence of the soil properties at the spatial location on its variability.
[0021] Furthermore, the vegetation coverage correction factor can be expressed by the following formula:
[0022] , where is the normalized difference vegetation index, which reflects the growth and coverage rate of vegetation; is the maximum value of the normalized difference vegetation index, is the symbol of the logarithmic function, is the vegetation growth status parameter of the th type of variable; is the benchmark value of vegetation growth, representing the average growth level of vegetation.
[0023] Furthermore, the human activity compensation factor can be expressed by the following formula:
[0024] , where is the fertilizer application activity sensitivity coefficient, which is used to measure the sensitivity or influence intensity of fertilizer application activities on the standardized arsenic content, is the symbol of the hyperbolic tangent function, is the amount of fertilizer applied, is the critical fertilizer application level, is the hyperbolic tangent function, is the irrigation level sensitivity coefficient, is the activation function symbol, is the irrigation level; is the activation function, which is used to simulate the non-linear influence of irrigation level on arsenic content.
[0025] Furthermore, the hierarchical adaptive standardization model also includes: S230, the standardized process integration and optimization step, which is used to optimize the data after the multi-source spatio-temporal alignment function and dynamic standardization. Through data type discrimination, pre-conversion is performed according to the different characteristics of variables, and at the same time, considering the inherent measurement errors of the data, uncertainty compensation and correction are carried out on the data to improve the accuracy of the data, and local weighted smoothing is used to achieve spatial dependence correction by using the spatial correlation of the data.
[0026] Furthermore, the data type discrimination is realized by constructing a standardized path selection function, and the standardized path selection function is expressed by the following formula:
[0027] , where is a standardized path selection function, is the Box-Cox transformation, which is for continuous data and transforms the original non-normal distributed data into approximately normal distributed data by finding an optimal parameter; is the Fisher-z transformation, which is used to transform remote sensing data to make its distribution closer to normal distribution and stabilize the variance of remote sensing data; is the current observation value, is the current observation value of the th class variable, is the minimum value of all current observation values, is the interquartile range, is the interquartile range standardization function.
[0028] Furthermore, the uncertainty compensation correction is achieved through a measurement error function, and the measurement error function is expressed by the following formula: , where is the standardized value considering measurement error, is the standardized value from the dynamic standardization equation, is the symbol of the natural base function, is the attenuation factor coefficient, is the data error parameter.
[0029] Furthermore, the spatial dependence correction is expressed by the following formula: , where is the standardized value after spatial dependence correction, is the index of adjacent observation points, is the total number of adjacent observation points, is the observation point 's standardized value considering measurement error; is the geographically weighted kernel function.
[0030] Furthermore, the geographically weighted kernel function can be expressed by the following formula:
[0031] , where is the spatial location and the th observation point is the Euclidean distance between them. The spatial influence radius is used to characterize the speed of distance attenuation. The larger the spatial influence radius, the wider the influence range of an observation point, meaning that even data points at a relatively far distance can have a significant impact on the target point. The smaller the spatial influence radius, the more limited the influence range, and only very close data points have an obvious impact. is the weight decay exponent, which is used to control the non - linear degree of distance decay; when k = 1, the decay is exponential - linear decay; when k>1 (for example, k = 2 becomes Gaussian decay), the decay speed is faster, meaning that the weights assigned to data points at a greater distance decrease sharply. The typical value range of this weight decay exponent can be 1.5 - 2.5.
[0032] Furthermore, the spatial interpolation algorithm is Kriging method.
[0033] On the other hand, the present invention provides a high - resolution spatial distribution system for predicting the arsenic content in rice on a large - scale range, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the high - resolution spatial distribution method of the arsenic content in rice on a large - scale range as described above.
[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0035] 1. By collecting and integrating multi - source heterogeneous data such as climate change data, remote sensing data, environmental factor data, and historical arsenic content data in rice, the present invention realizes the effective fusion and standardized processing of multi - modal data from different sources, different times, and different spatial positions, overcomes problems such as data islands and data non - uniformity in the prior art, and improves the integration and comparability of data.
[0036] 2. By adopting a multi - variable spatio - temporal alignment function and dynamic standardization technology, the present invention enables variables of different types and different dimensions to be unified in dimension, eliminates dimension and scale differences, and through dynamic standardization, effectively improves the accuracy and reliability of the analysis results, overcoming the limitations of traditional standardization methods that assume data follows a certain global and single distribution.
[0037] 3. The integration of the standardization process of the present invention, including steps such as data type discrimination, pre - conversion, uncertainty compensation and correction, and spatial dependence correction, significantly improves the quality and reliability of data, and solves the problems of errors and uncertainties existing in the data processing process of the prior art.
[0038] 4. The present invention imports the processed data set into GIS software, and combines with a spatial interpolation algorithm to generate a spatial distribution map of the arsenic content in rice, providing a highly reliable data basis for advanced applications such as subsequent arsenic pollution risk assessment, spatial distribution mapping, and driving factor analysis, overcoming problems such as poor data timeliness and limited sampling points in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0040] Figure 1 This is the flowchart of the method provided in Embodiment 1 of the present invention. Detailed implementation manners
[0041] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0042] Embodiment 1
[0043] Currently, the common method for predicting the arsenic content in rice on a large scale is to collect data on the arsenic content of rice samples and soil arsenic content at the national scale and data on rice yields at the provincial scale across the country. Then, based on each square kilometer, the annual rice yields of each province are evenly distributed to the agricultural land in each province. Based on the rice yield and soil arsenic concentration per square kilometer of agricultural land, the rice yield on the land where the soil arsenic concentration exceeds 20 mg kg-1 is calculated. Based on the collected data on the inorganic arsenic concentration in rice and the proportion of arsenic-rich rice, the average value of the collected rice samples within different soil arsenic concentration ranges is obtained. However, the existing technology has low prediction resolution, only showing the overall national situation of rice arsenic content, lacking prediction results for different regions and high resolution.
[0044] This embodiment discloses a method for high-resolution spatial distribution of arsenic content in rice on a large scale range, Figure 1 showing the overall method flowchart in this embodiment. It can be seen from the figure that this embodiment includes the following steps:
[0045] Step 1: First, collect data related to arsenic content in rice, including climate change data, remote sensing data, environmental factor data, and historical arsenic content data in rice, and integrate the collected data into a complete database related to arsenic content in rice for convenient subsequent processing.
[0046] Specifically, the climate change data may include:
[0047] Historical climate data: temperature, precipitation, humidity, CO2 concentration, frequency and intensity of extreme weather events (such as floods, droughts), etc.
[0048] Future climate scenario data: prediction data based on global / regional climate models of institutions such as IPCC.
[0049] Surface hydrological data: surface runoff, soil moisture, groundwater level, etc.
[0050] The remote sensing data may include:
[0051] High-resolution optical remote sensing data: Relevant remote sensing data purchased from commercial satellite websites, used to extract rice physiological parameters such as vegetation indices (NDVI, EVI), chlorophyll content, canopy temperature, etc.
[0052] Synthetic aperture radar (SAR) data: Synthetic aperture radar data purchased from commercial satellite websites, used to monitor rice planting area, growth stage, flooding situation, etc.
[0053] Hyperspectral remote sensing data: Specific bands are more sensitive to plant physiological changes that may be caused by arsenic stress and can be used for early warning and precise diagnosis.
[0054] Environmental factor data may include:
[0055] Soil data: Soil type, pH value, organic matter content, redox potential, heavy metal background values (especially arsenic), etc.
[0056] Water body data: Irrigation water quality (arsenic content, pH, etc.), arsenic content in groundwater.
[0057] Topographic data: Elevation, slope, water system distribution, etc.
[0058] Historical rice arsenic content data may include:
[0059] Field sampling data: Rice grain arsenic content data with clear geographical coordinates and sampling time.
[0060] Literature data / regional research reports: Used to integrate existing regional survey data.
[0061] Step 2: Considering that the data related to rice arsenic content collected are multi-modal data from different sources, at different times, and in different spatial locations, it is very important to standardize these data to make subsequent calculations more accurate. In this embodiment, the collected data are standardized through a hierarchical adaptive standardization model.
[0062] The hierarchical adaptive standardization model in this embodiment realizes data time synchronization and spatial matching by considering the specific characteristics of different data and how to fuse them together, and by aligning them to the same time and spatial resolution through interpolation methods. And through hierarchical processing, such as first standardizing each different data type separately and then merging them into a unified spatial grid for fusion.
[0063] Specifically, the core of the hierarchical adaptive normalization model is the multivariate spatio-temporal alignment function. When dealing with complex data such as rice arsenic content-related data, a core problem often encountered is that the data sources are diverse, the observation times are different, and the sampling locations are not the same. For example, some of the arsenic content data may be obtained through satellite remote sensing in 2023, some are obtained through on-site sampling in 2024, and some are the average values of a certain area recorded in historical documents. These data are like information fragments scattered at different times, different locations, and described by different "languages" (measurement methods). The core role of the multivariate spatio-temporal alignment function in the hierarchical adaptive normalization model is to unify these fragments into a set of standard "spatio-temporal coordinate systems", so that data from different sources can be meaningfully compared and integrated. That is to say, the goal of the multivariate spatio-temporal alignment function is to map all information onto an accurate and unified global map.
[0064] In this embodiment, the multivariate spatio-temporal alignment function can be expressed by the following formula:
[0065] ,
[0066] where, is the multivariate spatio-temporal alignment function of the th type of data source, is the target spatio-temporal point, are the abscissa and ordinate in the spatial position, is the time point; it means that on the unified spatio-temporal coordinate , the estimated value of the th type of data source (for example, arsenic content measured by a specific type of sensor, soil arsenic concentration obtained through a certain analysis method, etc.), which is the result that the multivariate spatio-temporal alignment function attempts to output, that is, a standardized or interpolated data point at a specific spatio-temporal point. is the data of a specific observation point, is the set of data of all observation points; is the data credibility weight of the nd observation point; is the Gaussian kernel function symbol, is the weighted Euclidean distance, is the anisotropic Gaussian kernel function, is the original observation value matrix. The data in this matrix are the input data of the model. Here, using a matrix to represent each observation point means that it is not just a single arsenic content value, but may also contain other relevant observation features, such as soil type, moisture content, temperature, vegetation condition, etc. These are the original data obtained directly without any processing.
[0067] The weighted Euclidean distance can be expressed by the following formula:
[0068] ,
[0069] wherein, is the time dimension scaling factor; is the target spatio-temporal point; is the original spatio-temporal point.
[0070] It should be noted that the above formula aligns the data from different spatial positions and time points to a unified coordinate system through a weighted method. The data of each observation point is weighted according to its spatio-temporal distance, and the observation points closer to the target point will have a greater impact on the final result. At the same time, combined with the Gaussian kernel function, the impact of the observation points far from the target point on the result is reduced through attenuation. In this kernel function, determines the smoothness of space and time. A large standard deviation ( large) will make the influence of the distance in space or time smaller, and vice versa. The credibility weight of the observation points enables different importance to be given to data of different qualities during the alignment process. Observation points with high credibility will have a greater weight and thus occupy a larger proportion in the result.
[0071] After completing the spatio-temporal alignment of the data through the multi-source spatio-temporal alignment function, the data is dynamically standardized. In this embodiment, considering that the arsenic content in rice is comprehensively affected by various complex factors such as climate, soil, vegetation cover, and human activities, and these factors will show significant differences in different regions and different times, simply performing traditional standardization (such as Z-score standardization) on all data is often insufficient. These traditional standardization methods usually assume that the data follows a certain global and single distribution, which will lose accuracy and applicability when dealing with highly heterogeneous environmental data. Therefore, the core goal of the dynamic standardization in this embodiment is: by converting variables of different types and different dimensions (such as arsenic content, rainfall, temperature, NDVI, etc.) into dimensionless and comparable values, the dimension and scale differences are eliminated. Considering the influence of different climate zones, soil types, vegetation conditions, and human activity intensities on the data distribution, local and adaptive adjustments are made to achieve adaptation to regional specificity. And the standardized arsenic content data of rice still has comparability in different regions and different environmental backgrounds, so that risk assessment or model construction can be carried out more accurately, and the comparability of the data is improved.
[0072] Specifically, the dynamic standardization in this embodiment can be expressed by the following formula:
[0073] ,
[0074] wherein, is the dynamic standardization value, indicating the th type of variable at the spatial position Normalized value on is the original observation value, representing the class variable at the spatial location of the original observation value; is the dynamic mean of the climate zone, which is obtained by taking the weighted average of the original observation values at all spatial locations within the climate zone ; is the standard deviation of the climate zone , used to measure the variation range of the original observation values within this zone; is the vegetation cover correction factor, used to multiplicatively adjust the standardized result according to the vegetation cover, reflecting the influence of vegetation on the absorption, migration or bioaccumulation of nutrients (including arsenic); is the human activity compensation factor, used to directly compensate or correct the influence of human activities (mainly fertilization and irrigation) on arsenic content. These activities may directly introduce arsenic into the soil (such as arsenic in phosphate fertilizers), or change the soil environment, thus affecting the migration and transformation of arsenic.
[0075] Specifically, in the above formula:
[0076] The dynamic mean of the climate zone can be expressed by the following formula:
[0077] ,
[0078] where is the climate zone, is the number of data points contained within the climate zone, is the transformation function symbol, is the transformation function, which acts on the original observation value. Its purpose is to handle the non - linear relationship, skewed distribution or outliers of the data. Here, the transformation function symbol can be logarithmic transformation or Box - Cox transformation, making the data closer to a normal distribution, so that the mean is more representative. In short, the role of this transformation function is to adjust the weight of the contribution of the original observation value to the mean, making it more in line with the true situation of the data distribution when calculating the mean. is the climate weight, used to measure the representativeness or importance of the spatial location within its affiliated climate zone . For example, within a climate zone, the climate characteristics of a certain location may be more typical or more stable than those of other locations, so its weight is higher; this weight can be set according to climate similarity, climate stability index, or distance from the center of the climate zone, etc., ensuring that when calculating the zonal mean, it can better reflect the core climate characteristics of this climate zone.
[0079] Climate zone The standard deviation can be expressed by the following formula:
[0080] ,
[0081] where, is the soil weight, which is used to measure the influence of soil properties at spatial locations on its variability (different soil types (e.g., sandy soil, clay), pH value, organic matter content, etc. will all affect the adsorption, desorption, and migration of arsenic, thus affecting the variability of its spatial distribution). This weight can be set according to the influence degree of soil type, soil physical and chemical properties (pH, organic matter, clay content, etc.) on arsenic retention.
[0082] The vegetation coverage correction factor can be expressed by the following formula:
[0083] ,
[0084] where, is the normalized difference vegetation index, is the maximum value of the normalized difference vegetation index, which reflects the growth vigor and coverage rate of vegetation; is the symbol of the logarithmic function, is the vegetation growth status parameter of the th class variable; is the benchmark value of vegetation growth, representing the average growth level of vegetation.
[0085] It should be noted that the vegetation coverage correction factor is used as a multiplicative factor in dynamic standardization to adjust the standardization result to reflect the possible biological impact of vegetation coverage on arsenic content.
[0086] The human activity compensation factor can be expressed by the following formula:
[0087] ,
[0088] where, is the sensitivity coefficient of fertilization activity, which is used to measure the sensitivity or influence intensity of fertilization activity on the standardized arsenic content. The larger the sensitivity coefficient of fertilization activity, the greater the promotion effect of fertilization on the standardized value. The typical value of this coefficient is: 0.2 - 0.8, indicating that fertilization may have a medium to strong compensation effect on arsenic content; is the symbol of the hyperbolic tangent function, is the fertilization amount, is the critical fertilization level, is the hyperbolic tangent function, which is used for the saturation effect of the fertilization rate on the arsenic content. When the fertilization rate is low, the hyperbolic tangent function grows almost linearly, indicating that the arsenic content increases significantly with the increase of the fertilization rate. When the fertilization rate is very high, the hyperbolic tangent function approaches 1, indicating that the additional effect of fertilization on the arsenic content gradually saturates, that is, after reaching a certain fertilization rate, increasing the fertilization rate further has little effect on the increase of the arsenic content. is the irrigation level sensitivity coefficient, is the activation function symbol, is the irrigation level, which can be indicators such as irrigation intensity, frequency or total amount; is the activation function, which is used to simulate the non-linear effect of the irrigation level on the arsenic content. The activation function maps any real value to the interval (0,1), and is often used to represent the cumulative effect or probability from low to high and gradually saturating. In this formula, the activation function is used to represent that when the irrigation level is low, the effect is small; as the irrigation level increases, the effect increases rapidly; when reaching a high level, the effect tends to saturate, which is used to reflect the complex non-linear effect of irrigation on arsenic migration and bioavailability.
[0089] After the aforementioned multi-dimensional space-time alignment function and dynamic standardization, the data has been unified in dimension and corrected for the environmental background. To further improve the quality and reliability of the data, the quality and reliability of the data can also be further improved through the integration of the standardization process.
[0090] In the integration of the standardization process, through data type discrimination, pre-conversion is performed according to the different characteristics of variables to make them more in line with the requirements of statistical processing; at the same time, considering the inherent measurement errors of the data, uncertainty compensation and correction are performed on the data to improve the data accuracy; and using the spatial correlation of the data, local weighted smoothing is performed to achieve spatial dependence correction.
[0091] Specifically, data type discrimination can be achieved by constructing a standardization path selection function, and the standardization path selection function can be expressed by the following formula:
[0092] ,
[0093] where, is the standardization path selection function, which automatically selects the best preprocessing (transformation) method according to the type of the input variable; it ensures that each type of data is properly prepared before entering the dynamic standardization.
[0094] is the Box-Cox transformation, which is for continuous data (such as rainfall, temperature, humidity, soil pH value, etc.). By finding an optimal parameter, the original non-normal distribution data (such as skewed distribution) is transformed into an approximately normal distribution, or its variance is stabilized to make it more in line with the assumptions of many statistical models.
[0095] is the Fisher - z transform, which is used to transform remotely sensed related data to make its distribution closer to normal and stabilize its variance.
[0096] is the current observed value, is the current observed value of the is the minimum value of all current observed values, is the inter - quartile range, is the inter - quartile range normalization function. For categorical or ordinal data (such as fertilization intensity levels: low, medium, high; irrigation levels: none, small amount, moderate amount, large amount, etc.), it linearly scales them to a relatively unified scale and reduces the influence of outliers.
[0097] It should be noted that the normalization path selection function can be used as a pre - processing for dynamic normalization. After calculating , according to the type of variable , select the corresponding transformation function to process to obtain a pre - processed This pre - processed value will be used as the original observed value for actual calculation when.
[0098] Considering that the collected data is accompanied by uncertainties or errors, the influence of these errors on the normalization result can be quantified and corrected to make the data closer to the true value. In this embodiment, the measurement error function is used to achieve the compensation for uncertainties. The measurement error function can be expressed by the following formula:
[0099] ,
[0100] where, is the normalized value considering measurement errors, is the normalized value from the dynamic normalization equation, is the symbol of the natural exponential function, is the attenuation factor coefficient, which is used to control the intensity of the influence of measurement errors on the normalized value. Through it, the signal - to - noise ratio of the data is converted into a coefficient that can be used to measure error attenuation. is the data error parameter, which is used to quantify the inherent inaccuracy of the data of the category variable.
[0101] Finally, through geographically weighted normalization, the data is locally smoothed and corrected to make the final result more in line with the characteristics of geospatial data. In this embodiment, geographically weighted normalization can be expressed by the following formula:
[0102] ,
[0103] Among them, is the index of the adjacent observation points, is the total number of the near observation points, is the observation point after considering the measurement error; is the geographically weighted kernel function, which is used to measure the influence intensity of the observation point on the target point . This kernel function can adopt the common Gaussian attenuation function.
[0104] is the standardized value after the final spatial dependence correction. The set of all standardized values after the spatial dependence correction is a data set, which is the final clean data set obtained by processing through the hierarchical adaptive standardization model. The data in this data set are aligned in space-time, unified in dimension, corrected in environmental background, compensated in measurement error, and smoothed in space, thus providing a highly reliable data basis for subsequent advanced applications such as arsenic pollution risk assessment, spatial distribution mapping, and driving factor analysis.
[0105] In this embodiment, the geographically weighted kernel function can be expressed by the following formula:
[0106] ,
[0107] Among them, is the Euclidean distance between the spatial position and the i-th observation point, is the Euclidean distance between the spatial position and the j-th observation point, is the spatial influence radius, which is used to characterize the speed of distance attenuation. The larger the spatial influence radius, the wider the influence range of an observation point, indicating that even data points at a relatively far distance can have a significant impact on the target point. The smaller the spatial influence radius, the more limited the influence range, and only very close data points have an obvious impact. is the weight attenuation index, which is used to control the non-linearity degree of distance attenuation; when k = 1, the attenuation is exponential linear attenuation; when k > 1 (for example, k = 2 becomes Gaussian attenuation), the attenuation speed is faster, indicating that the weights assigned to data points at a relatively far distance decrease sharply. The typical value range of this weight attenuation index can be 1.5 - 2.5.
[0108] Step 3: Import the data set processed by the hierarchical adaptive standardization model into the GIS software. Combining with the spatial interpolation algorithm, generate the spatial distribution map of the arsenic content in rice in the GIS software.
[0109] Specifically, in this embodiment, the data in the dataset processed by the hierarchical adaptive normalization model usually exists in the form of point features, and each point is associated with longitude, latitude, time, and the normalized value after the corresponding spatial dependence correction.
[0110] Since the data in this dataset are actually discrete sampling points. In order to generate a continuous spatial distribution map, spatial interpolation is required, which estimates the values of unknown points by using the values of known points. In this embodiment, spatial interpolation can be achieved by using the Kriging method, thereby generating the final spatial distribution map of the arsenic content in rice.
[0111] Embodiment 2
[0112] In this embodiment, a high-resolution spatial distribution system for predicting the arsenic content in rice on a large scale is disclosed.
[0113] The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it can implement a method for the high-resolution spatial distribution of the arsenic content in rice on a large scale in Embodiment 1.
[0114] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A high-resolution spatial distribution method for arsenic content in rice over a large scale range, characterized in that, The high-resolution spatial distribution method includes the following: S100. Collect data related to the arsenic content in rice and integrate the collected data into a rice arsenic content database; S200. Standardize the data in the rice arsenic content database through a hierarchical adaptive normalization model. The hierarchical adaptive normalization model fuses data by considering the characteristics of different data. And through an interpolation method, align them to the same temporal and spatial resolutions to achieve data temporal synchronization and spatial matching; S300. Import the data set processed by the hierarchical adaptive normalization model into GIS software, and combine with a spatial interpolation algorithm to generate a spatial distribution map of the arsenic content in rice in the GIS software.
2. The high-resolution spatial distribution method of arsenic content in rice on a large scale range according to claim 1, characterized in that, The data related to the arsenic content in rice includes: climate change data, remote sensing data, environmental factor data, and historical rice arsenic content data.
3. The high-resolution spatial distribution method for arsenic content in rice over a large scale range according to claim 1, characterized in that, The hierarchical adaptive normalization model includes the following sub-steps: S210. By weighting data from different sources, align the data from different spatial positions and time points to a unified coordinate system. So that the data at each observation point is weighted according to the spatio-temporal distance, and the observation points closer to the target point will have a greater impact on the final result. Combine with a Gaussian kernel function to reduce the impact of observation points far from the target point on the result by attenuation. By setting credibility weights for the observation points, assign different degrees of importance to data of different qualities during the alignment process, thereby constructing a multi-source spatio-temporal alignment function; S220. After completing the spatio-temporal alignment of the data through the multi-source spatio-temporal alignment function, perform dynamic normalization on the data. The dynamic normalization comprehensively considers the significant differences in the rice arsenic content data affected by climate, soil, vegetation cover, and human activities in different regions and at different times. Convert the rice arsenic content data of different types and dimensions into dimensionless values to eliminate the differences in data dimension and scale.
4. The method for high-resolution spatial distribution of arsenic content in rice on a large scale according to claim 3, wherein The multi-source spatio-temporal alignment function is represented by the following formula: , Among them, is the multi - temporal - spatial alignment function for the type of data source, is the target spatio - temporal point, are the abscissa and ordinate in the spatial position, is the time point; is the data of the specific observation point, is the set of data of all observation points; is the data credibility weight of the th observation point; is the Gaussian kernel function symbol, is the weighted Euclidean distance, is the anisotropic Gaussian kernel function, is the original observation value matrix.
5. The method for high-resolution spatial distribution of arsenic content in rice on a large scale according to claim 3, wherein The dynamic normalization is represented by the following formula: , Among them, is the dynamic standardization value, indicating the standardization value of the nth type of variable at the spatial position; is the original observed value, indicating the original observed value of the nth type of variable at the spatial position;<l is the climate division, is the dynamic mean of the climate division, is the standard deviation of the climate division, is the vegetation coverage correction factor, is the human activity compensation factor.
6. The high-resolution spatial distribution method for arsenic content in rice on a large scale according to claim 3, characterized in that The hierarchical adaptive normalization model further includes: S230. A standardization process integration and optimization step for optimizing the data after passing through the multi-source spatio-temporal alignment function and dynamic normalization. Through data type discrimination, perform pre-conversion according to the different characteristics of the variables. At the same time, considering the inherent measurement errors in the data, perform uncertainty compensation and correction on the data to improve the accuracy of the data. And utilize the spatial correlation of the data to perform local weighted smoothing to achieve spatial dependence correction.
7. The method for high-resolution spatial distribution of arsenic content in rice on a large scale according to claim 6, characterized in that The data type discrimination is achieved by constructing a standardization path selection function. The standardization path selection function is represented by the following formula: , Among them, is a standardized path selection function, is the Box-Cox transformation, which is for continuous data and transforms the original non-normal distributed data into an approximately normal distribution by finding an optimal parameter; The Fisher-z transform is used to transform remote sensing data, making the distribution of the remote sensing data closer to a normal distribution and stabilizing the variance of the remote sensing data; is the current observation value, is the current observation value of the category variable, is the minimum value of all current observation values, is the interquartile range, is the interquartile range normalization function.
8. The high-resolution spatial distribution method for arsenic content in rice on a large scale range according to claim 6, wherein The uncertainty compensation and correction is achieved through a measurement error function. The measurement error function is represented by the following formula: , Among them, is the standardized value after considering measurement errors, is the standardized value from the dynamic standardization equation, is the natural base function symbol, is the attenuation factor coefficient, is the data error parameter.
9. The high-resolution spatial distribution method for arsenic content in rice on a large scale range according to claim 6, characterized in that The spatial dependence correction is represented by the following formula: , Among them, is the standardized value after spatial dependence correction, is the index of adjacent observation points, is the total number of near observation points, is the observation point 's standardized value after considering measurement errors; is the geographically weighted kernel function.
10. A high-resolution spatial distribution system for predicting the arsenic content in rice on a large-scale range, characterized in that, The high-resolution spatial distribution system includes: A processor; A memory storing a computer program, which when executed by the processor, implements the high-resolution spatial distribution method for the arsenic content in rice on a large-scale range as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Self-adaption spatial interpolation method and system based on spatial feature analysis
CN103353923A
Risk assessment method for predicting arsenic element content in rice
CN111047223A
Soil heavy metal spatial interpolation method and device and computer readable storage medium
CN113012771A
Coked soil pollution spatial distribution prediction optimization method and system
CN114118613A
Soil arsenic concentration spatial distribution inversion method and device and computer equipment
CN114384023A