A method and system for high-resolution spatial distribution of arsenic content in rice over a large scale

The rice arsenic content data was processed by a hierarchical adaptive normalization model and a multivariate spatiotemporal alignment function, which solved the problem of heterogeneous data in multi-source data fusion, achieved the generation of high-resolution spatial distribution maps of rice arsenic content, and improved the accuracy and reliability of the data.

CN120408544BActive Publication Date: 2025-09-26INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510927045.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-26
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

The existing multi-source data fusion methods have differences in spatiotemporal resolution, data format and quality when processing heterogeneous data, which makes data fusion difficult and lacks uncertainty assessment of prediction results, affecting the reliability of management decisions.

Method used

A hierarchical adaptive normalization model was used to standardize the data in the rice arsenic content database. Multivariate spatiotemporal alignment functions and dynamic normalization technology were used to align data from different sources to the same temporal and spatial resolution. GIS software was then used to generate a spatial distribution map of rice arsenic content.

Benefits of technology

It has achieved effective fusion and standardized processing of data from different sources, improved the integration and comparability of data, enhanced the accuracy and reliability of analysis results, and provided a highly reliable data basis for subsequent arsenic pollution risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408544B_ABST
    Figure CN120408544B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for high-resolution spatial distribution of arsenic content in rice over a large scale. This method, which belongs to the field of spatial distribution generation, includes collecting data related to rice arsenic content and integrating the collected data into a rice arsenic content database; standardizing the data in the rice arsenic content database using a hierarchical adaptive normalization model; importing the data set processed by the hierarchical adaptive normalization model into GIS software, and using a spatial interpolation algorithm to generate a spatial distribution map of arsenic content in rice within the GIS software. Through dynamic normalization, this invention effectively improves the accuracy and reliability of the analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of spatial distribution generation, and in particular to a method and system for high-resolution spatial distribution of arsenic content in rice over a large scale. Background Art

[0002] With global climate change and intensified human activities, the problem of arsenic content in rice is becoming increasingly prominent. With the development of remote sensing technology, geographic information systems, and data mining, multi-source data fusion methods have been widely used in soil nutrient analysis. These methods combine multi-source information such as satellite remote sensing imagery, topographic data, and meteorological data to achieve large-scale, high-frequency monitoring and assessment of soil nutrient status. However, existing multi-source data fusion methods still face many challenges when processing heterogeneous data. First, data from different sources have significant differences in spatiotemporal resolution, data format, and quality, making the data fusion process difficult. Second, traditional data processing methods cannot effectively capture the spatiotemporal variation characteristics and nonlinear relationships of soil nutrients. Third, existing analysis methods often lack uncertainty assessment of prediction results, which affects the reliability of management decisions. Summary of the Invention

[0003] One of the purposes of the present invention is to provide a high-resolution spatial distribution method for arsenic content in rice over a large scale, so as to solve the problem in the prior art that data from different sources have significant differences in spatiotemporal resolution, data format and quality, which makes the data fusion process difficult.

[0004] The present invention is achieved through the following technical solution: a method for high-resolution spatial distribution of arsenic content in rice over a large scale, comprising the following steps: S100, collecting data related to arsenic content in rice, and integrating the collected data into a rice arsenic content database; S200, standardizing the data in the rice arsenic content database using a hierarchical adaptive normalization model, wherein the hierarchical adaptive normalization model fuses the data by considering the characteristics between different data, and aligns the data to the same time and space resolution through an interpolation method to achieve temporal synchronization and spatial matching of the data; S300, importing the data set processed by the hierarchical adaptive normalization model into GIS software, and generating a spatial distribution map of arsenic content in rice in the GIS software in combination with a spatial interpolation algorithm.

[0005] Furthermore, data related to rice arsenic content include: climate change data, remote sensing data, environmental factor data and historical rice arsenic content data.

[0006] Furthermore, climate change data include: historical climate data, consisting of temperature, precipitation, humidity, CO2 concentration, frequency and intensity of extreme weather events; future climate scenario data, based on global / regional climate model prediction data from organizations such as the IPCC; surface hydrological data, consisting of surface runoff, soil moisture, and groundwater level data.

[0007] Furthermore, the remote sensing data includes: high-resolution optical remote sensing data, which is obtained by extracting rice physiological parameters such as vegetation index, chlorophyll content, and canopy temperature from relevant remote sensing data purchased from commercial satellite websites; synthetic aperture radar data, which is composed of synthetic aperture radar data purchased from commercial satellite websites and is used to monitor rice planting area, growth stage, and flooding conditions; hyperspectral remote sensing data, whose specific bands are more sensitive to plant physiological changes that may be caused by arsenic stress.

[0008] Furthermore, environmental factor data include: soil data, which consists of soil type, pH value, organic matter content, redox potential, and heavy metal background value; water body data, which consists of irrigation water quality (arsenic content, pH, etc.) and groundwater arsenic content; terrain data, including elevation, slope, water system distribution, etc.

[0009] Furthermore, historical rice arsenic content data include: field sampling data, which are rice grain arsenic content data with clear geographical coordinates and sampling time; literature data / regional research reports, which are used to integrate existing regional survey data.

[0010] Furthermore, the hierarchical adaptive normalization model includes the following sub-steps: S210, by weighting data from different sources, aligning data from different spatial locations and time points to a unified coordinate system, so that the data of each observation point is weighted according to the spatiotemporal distance, and the observation points close to the target point will have a greater impact on the final result. Combined with the Gaussian kernel function, the influence of observation points far from the target point on the result is reduced by attenuation. By setting credibility weights for the observation points, different qualities of data are given different importance in the alignment process, thereby constructing a multivariate spatiotemporal alignment function; S220, after completing the spatiotemporal alignment of the data through the multivariate spatiotemporal alignment function, the data is dynamically standardized. The dynamic standardization comprehensively considers the significant differences in rice arsenic content data due to climate, soil, vegetation cover and human activities in different regions and at different times, and converts the rice arsenic content data of variables of different types and dimensions into dimensionless values, thereby eliminating data dimension and scale differences.

[0011] Furthermore, the multivariate spatiotemporal alignment function is expressed as follows:

[0012] ,in, For the Class data source multivariate spatiotemporal alignment function, is the target space-time point, are the horizontal and vertical coordinates of the spatial position, For time point; is the data of a specific observation point, is the set of all observation point data; For the The data credibility weight of each observation point; is the symbol of Gaussian kernel function, is the weighted Euclidean distance, is the anisotropic Gaussian kernel function, is the original observation matrix.

[0013] Furthermore, the weighted Euclidean distance can be expressed as follows:

[0014] ,in, is the time dimension scaling factor; is the target space-time point; is the original space-time point.

[0015] Furthermore, dynamic normalization is expressed as follows:

[0016] ,in, is a dynamic normalized value, Indicates the Class variables in spatial position Normalized value on ; is the original observation value, Indicates the Class variables in spatial position The original observations on ; For climate zones, is the dynamic mean of the climate zone; is the standard deviation of the climate zone, is the vegetation cover correction factor, is the compensation factor for human activities.

[0017] Furthermore, the dynamic mean of the climate zone can be expressed as follows:

[0018] ,in, is the number of data points contained in the climate partition, is the symbol of the transformation function, is the transformation function, is the climate weight, which is used to measure the spatial location In its climate zone representativeness or importance in the

[0019] Furthermore, climate zoning The standard deviation can be expressed as follows:

[0020] ,in, is the soil weight, used to measure the spatial position The influence of soil characteristics on its variability.

[0021] Furthermore, the vegetation coverage correction factor can be expressed as follows:

[0022] ,in, It is the normalized vegetation index, which reflects the luxuriant degree and coverage of vegetation; is the maximum value of the normalized vegetation index, is the logarithmic function symbol, For the Vegetation growth condition parameters of class variables; It is the benchmark value for vegetation growth and represents the average growth level of vegetation.

[0023] Furthermore, the human activity compensation factor can be expressed as follows:

[0024] ,in, is the sensitivity coefficient of fertilization activity, which is used to measure the sensitivity or impact intensity of fertilization activity to standardized arsenic content. is the symbol of the hyperbolic tangent function, is the amount of fertilizer applied, is the critical fertilization level, is the hyperbolic tangent function, is the irrigation level sensitivity coefficient, is the activation function symbol, for irrigation levels; is the activation function used to simulate the nonlinear effect of irrigation level on arsenic content.

[0025] Furthermore, the hierarchical adaptive standardization model also includes: S230, a standardization process integrated optimization step, which is used to optimize the data after the multivariate spatiotemporal alignment function and dynamic standardization, and pre-convert the data according to the different characteristics of the variables through data type discrimination. At the same time, the inherent measurement errors of the data are taken into account, and the uncertainty compensation and correction of the data are performed to improve the accuracy of the data. The spatial correlation of the data is used to perform local weighted smoothing to achieve spatial dependency correction.

[0026] Furthermore, data type discrimination is achieved by constructing a standardized path selection function, which is represented by the following formula:

[0027] ,in, is the normalized path selection function, The Box-Cox transformation is for continuous data. It transforms the original non-normal distribution data into an approximate normal distribution by finding an optimal parameter. Fisher-z transformation is used to transform remote sensing data to make the distribution of remote sensing data closer to normal distribution and stabilize the variance of remote sensing data; is the current observation value, For the the currently observed value of the class variable, is the minimum value of all current observations, is the interquartile range, is the interquartile range standardization function.

[0028] Furthermore, the uncertainty compensation correction is achieved by a measurement error function, which is expressed by the following formula: ,in, is the standardized value after taking into account the measurement error, is the normalized value from the dynamic normalization equation, is the natural base function symbol, is the attenuation factor coefficient, is the data error parameter.

[0029] Furthermore, the spatial dependence correction is expressed as follows: ,in, is the standardized value after spatial dependence correction, is the index of the neighboring observation point, is the total number of near observation points, For observation points The standardized value after considering the measurement error; is the geographically weighted kernel function.

[0030] Furthermore, the geographically weighted kernel function can be expressed as follows:

[0031] ,in, For spatial location With the i-th observation point The Euclidean distance between The spatial influence radius is used to characterize the rate of distance attenuation. A larger spatial influence radius means a wider range of influence for an observation point, and even distant data points can have a significant impact on the target point. A smaller spatial influence radius means a more limited range of influence, and only very close data points have a significant impact. is the weight decay exponent, which is used to control the nonlinearity of distance decay. When k=1, the decay is exponential linear decay. When k>1 (for example, k=2 becomes Gaussian decay), the decay rate is faster, which means that the weight given to data points with a longer distance is sharply reduced. The typical value range of this weight decay exponent is 1.5~2.5.

[0032] Furthermore, the spatial interpolation algorithm is Kriging.

[0033] On the other hand, the present invention provides a high-resolution spatial distribution system for predicting the arsenic content of rice over a large scale, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the high-resolution spatial distribution method for predicting the arsenic content of rice over a large scale as described above is implemented.

[0034] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0035] 1. By collecting and integrating multi-source heterogeneous data such as climate change data, remote sensing data, environmental factor data, and historical rice arsenic content data, the present invention realizes the effective fusion and standardization of multimodal data from different sources, different times, and different spatial locations, overcoming the problems of data silos and data inconsistency in the existing technology, and improving the integration and comparability of data.

[0036] 2. By adopting multivariate spatiotemporal alignment functions and dynamic normalization technology, the present invention enables variables of different types and dimensions to be unified in terms of dimension, eliminating dimensional and scale differences. Through dynamic normalization, it effectively improves the accuracy and reliability of the analysis results, overcoming the limitation of traditional normalization methods that assume that data obey a certain global, single distribution.

[0037] 3. The standardized process integration of the present invention, including steps such as data type identification, pre-conversion, uncertainty compensation correction, and spatial dependency correction, significantly improves the quality and reliability of data and solves the error and uncertainty problems existing in the data processing process of the prior art.

[0038] 4. The present invention imports the processed dataset into GIS software and combines it with a spatial interpolation algorithm to generate a spatial distribution map of arsenic content in rice. This provides a highly reliable data foundation for subsequent advanced applications such as arsenic pollution risk assessment, spatial distribution mapping, and driving factor analysis, overcoming the problems of poor data timeliness and limited sampling points in existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0040] Figure 1 This is a flow chart of the method provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0042] Example 1

[0043] Currently, a common method for predicting large-scale arsenic levels in rice is to collect national-scale data on arsenic content in rice samples and soil arsenic content, as well as national and provincial-level rice yield data. The annual rice yield of each province is then evenly distributed across agricultural land in each province on a per square kilometer basis. Based on the rice yield and soil arsenic concentration per square kilometer of agricultural land, the rice yield on land with soil arsenic concentrations exceeding 20 mg kg⁻¹ is calculated. Based on the collected data on inorganic arsenic concentrations in rice and the proportion of rice with excessive arsenic content, average values ​​for rice samples collected across different soil arsenic concentration ranges are obtained. However, existing technologies have limited resolution, providing only a national overview of rice arsenic content and lacking regional and high-resolution predictions.

[0044] This embodiment discloses a method for high-resolution spatial distribution of arsenic content in rice over a large scale. Figure 1 The flowchart of the overall method in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:

[0045] Step 1: First, collect data related to rice arsenic content, including climate change data, remote sensing data, environmental factor data, and historical rice arsenic content data, and integrate the collected data into a complete rice arsenic content-related database to facilitate subsequent processing.

[0046] Specifically, climate change data can include:

[0047] Historical climate data: temperature, precipitation, humidity, CO2 concentration, frequency and intensity of extreme weather events (such as floods and droughts), etc.

[0048] Future climate scenario data: based on global / regional climate model prediction data from organizations such as the IPCC.

[0049] Surface hydrological data: surface runoff, soil moisture, groundwater level, etc.

[0050] Remote sensing data can include:

[0051] High-resolution optical remote sensing data: Relevant remote sensing data purchased from commercial satellite websites are used to extract rice physiological parameters such as vegetation index (NDVI, EVI), chlorophyll content, and canopy temperature.

[0052] Synthetic Aperture Radar (SAR) data: SAR data purchased from commercial satellite websites are used to monitor rice planting area, growth stage, flooding conditions, etc.

[0053] Hyperspectral remote sensing data: Specific bands are more sensitive to plant physiological changes that may be caused by arsenic stress and can be used for early warning and accurate diagnosis.

[0054] Environmental factor data may include:

[0055] Soil data: soil type, pH value, organic matter content, redox potential, heavy metal background value (especially arsenic), etc.

[0056] Water data: irrigation water quality (arsenic content, pH, etc.), groundwater arsenic content.

[0057] Topographic data: elevation, slope, water distribution, etc.

[0058] Historical rice arsenic data may include:

[0059] Field sampling data: arsenic content data of rice grains with clear geographical coordinates and sampling time.

[0060] Literature data / regional research reports: used to integrate existing regional survey data.

[0061] Step 2: Considering that the collected data on arsenic content in rice are multimodal data from different sources, different times, and different spatial locations, it is very important to standardize these data to make subsequent calculations more accurate. In this embodiment, the collected data are standardized through a hierarchical adaptive normalization model.

[0062] The hierarchical adaptive normalization model in this embodiment considers the specific characteristics of different data types and how to fuse them together. It also uses interpolation methods to align them to the same temporal and spatial resolutions, achieving temporal synchronization and spatial matching of the data. Furthermore, the model uses hierarchical processing, such as first normalizing each data type individually and then merging them into a unified spatial grid for fusion.

[0063] Specifically, the core of the hierarchical adaptive normalization model is the multivariate spatiotemporal alignment function. When dealing with complex data such as rice arsenic levels, a key challenge often arises: data from diverse sources, observation times, and sampling locations. For example, some arsenic levels may be obtained through satellite remote sensing in 2023, others through field sampling in 2024, and still others based on average values ​​for a specific region recorded in historical documents. These data are like fragments of information scattered across time, location, and described using different "languages" (measurement methods). The core function of the multivariate spatiotemporal alignment function in the hierarchical adaptive normalization model is to unify these fragments into a standard "spatiotemporal coordinate system," enabling meaningful comparison and integration of data from different sources. In other words, the multivariate spatiotemporal alignment function aims to map all this information onto a single, accurate, and unified global map.

[0064] In this embodiment, the multivariate spatiotemporal alignment function can be expressed by the following formula:

[0065] ,

[0066] in, For the Class data source multivariate spatiotemporal alignment function, is the target space-time point, are the horizontal and vertical coordinates of the spatial position, is a time point; it represents the time in a unified space-time coordinate On the first An estimate of a type of data source (for example, arsenic levels measured by a particular type of sensor, soil arsenic concentration derived from a particular analytical method, and so on) is what the multivariate spatiotemporal alignment function attempts to output: a standardized or interpolated data point at a specific point in space and time. is the data of a specific observation point, is the set of all observation point data; For the The data credibility weight of each observation point; is the symbol of Gaussian kernel function, is the weighted Euclidean distance, is the anisotropic Gaussian kernel function, is the original observation value matrix. The data in this matrix is ​​the input data of the model. The matrix is ​​used here to represent that each observation point is not just a single arsenic content value, but may also contain other related observation characteristics, such as soil type, moisture content, temperature, vegetation conditions, etc. These are raw data obtained directly without any processing.

[0067] The weighted Euclidean distance can be expressed as follows:

[0068] ,

[0069] in, is the time dimension scaling factor; is the target space-time point; is the original space-time point.

[0070] It should be noted that the above formula aligns data from different spatial locations and time points into a unified coordinate system by weighting. The data of each observation point is weighted according to its spatial and temporal distance. The observation points closer to the target point will have a greater impact on the final result. At the same time, the Gaussian kernel function is combined to reduce the influence of observation points far from the target point on the result by attenuation. In this kernel function, Determines the degree of spatial and temporal smoothness, large standard deviation ( A large distance (in the sense of distance) will reduce the impact of spatial or temporal distance, while a small distance will have a large impact. The credibility weight of the observation point gives different weights to data of different qualities during the alignment process. Observations with high credibility will have a greater weight and thus occupy a larger proportion in the results.

[0071] After completing spatiotemporal alignment of the data using a multivariate spatiotemporal alignment function, the data is dynamically normalized. In this embodiment, considering that rice arsenic content is influenced by a complex set of factors, including climate, soil, vegetation cover, and human activity, which can vary significantly across regions and over time, simply applying traditional normalization (such as Z-score normalization) to all data is often insufficient. These traditional normalization methods typically assume that the data adhere to a global, single distribution, which can compromise accuracy and applicability when dealing with highly heterogeneous environmental data. Therefore, the core goal of the dynamic normalization process in this embodiment is to eliminate dimension and scale differences by converting variables of different types and dimensions (such as arsenic content, rainfall, temperature, and NDVI) into dimensionless, comparable values. Taking into account the impact of different climate zones, soil types, vegetation conditions, and human activity intensity on data distribution, localized and adaptive adjustments are performed to achieve regional specificity. Furthermore, the standardized rice arsenic content data remains comparable across different regions and environmental contexts, enabling more accurate risk assessment or model building and improving data comparability.

[0072] Specifically, the dynamic normalization in this embodiment can be expressed by the following formula:

[0073] ,

[0074] in, is a dynamic normalized value, indicating the Class variables in spatial position Normalized value on ; is the original observation value, indicating the Class variables in spatial position The original observations on ; is the dynamic mean of the climate zone, which is obtained by All spatial locations The weighted average of the original observations on is obtained; For climate zones The standard deviation is used to measure the variation of the original observation value within the partition; The vegetation cover correction factor is used to make multiplicative adjustments to the standardized results according to the vegetation cover, reflecting the effect of vegetation on the absorption, migration or bioaccumulation of nutrients (including arsenic); It is a human activity compensation factor used to directly compensate or correct the impact of human activities (mainly fertilization and irrigation) on arsenic content. These activities may directly introduce arsenic into the soil (such as arsenic in phosphate fertilizers) or change the soil environment, thereby affecting the migration and transformation of arsenic.

[0075] Specifically, in the above formula:

[0076] The dynamic mean of the climate zone can be expressed as follows:

[0077] ,

[0078] in, For climate zones, is the number of data points contained in the climate partition, is the symbol of the transformation function, It is a transformation function that acts on the original observations. Its purpose is to deal with nonlinear relationships, skewed distributions or outliers in the data. The transformation function symbol here can be logarithmic transformation or Box-Cox transformation, which makes the data closer to the normal distribution, making the mean more representative. In short, the role of this transformation function is to adjust the weight of the original observations' contribution to the mean, so that it is more in line with the actual situation of the data distribution when calculating the mean. is the climate weight, which is used to measure the spatial location In its climate zone The representativeness or importance of a climate zone. For example, within a climate zone, the climate characteristics of a certain location may be more typical or more stable than those of other locations, so its weight will be higher. The weight can be set according to climate similarity, climate stability index, or distance from the center of the climate zone, ensuring that the core climate characteristics of the climate zone are better reflected when calculating the zone mean.

[0079] Climate zones The standard deviation can be expressed as follows:

[0080] ,

[0081] in, is the soil weight, used to measure the spatial position The weight can be set based on the influence of soil characteristics on its variability (different soil types (e.g., sandy soil, clay), pH value, organic matter content, etc. will affect the adsorption, desorption and migration of arsenic, thereby affecting the variability of its spatial distribution).

[0082] The vegetation coverage correction factor can be expressed as follows:

[0083] ,

[0084] in, is the normalized difference vegetation index, It is the maximum value of the normalized vegetation index, reflecting the luxuriant growth and coverage of vegetation; is the logarithmic function symbol, For the Vegetation growth condition parameters of class variables; It is the benchmark value for vegetation growth and represents the average growth level of vegetation.

[0085] It should be noted that the vegetation cover correction factor is used as a multiplicative factor in dynamic standardization to adjust the standardization results to reflect the possible biological impact of vegetation cover on arsenic content.

[0086] The human activity compensation factor can be expressed as follows:

[0087] ,

[0088] in, The fertilization activity sensitivity coefficient is used to measure the sensitivity or impact of fertilization on the standardized arsenic content. The larger the fertilization activity sensitivity coefficient, the greater the effect of fertilization on the standardized value. Typical values ​​of this coefficient are: 0.2-0.8, indicating that fertilization may have a moderate to strong compensatory effect on arsenic content. is the symbol of the hyperbolic tangent function, is the amount of fertilizer applied, is the critical fertilization level, is the hyperbolic tangent function, which is used to saturate the effect of fertilizer application on arsenic content. When the fertilizer application amount is low, the hyperbolic tangent function grows almost linearly, indicating that the arsenic content increases significantly with the increase of fertilizer application amount; when the fertilizer application amount is very high, the hyperbolic tangent function approaches 1, indicating that the additional effect of fertilizer application on arsenic content gradually saturates, that is, after reaching a certain fertilizer application amount, the effect of increasing the fertilizer application amount on the arsenic content is not obvious. is the irrigation level sensitivity coefficient, is the activation function symbol, The irrigation level can be indicators such as irrigation intensity, frequency or total amount; is an activation function used to simulate the nonlinear effect of irrigation level on arsenic content. The activation function maps any real value to the range (0,1). It is often used to represent the cumulative effect or probability from low to high and gradually saturated. In this formula, the activation function is used to indicate that when the irrigation level is low, the impact is small; as the irrigation level increases, the impact increases rapidly; when it reaches a high level, the impact tends to saturation, which is used to reflect the complex nonlinear effect of irrigation on arsenic migration and bioavailability.

[0089] The aforementioned multivariate spatiotemporal alignment function and dynamic normalization unify the data dimensions and correct for environmental backgrounds. To further enhance the data quality and reliability, standardized process integration can be used to further enhance the data quality and reliability.

[0090] In the integration of standardized processes, data types are identified and pre-converted according to the different characteristics of variables to make them more consistent with the requirements of statistical processing. At the same time, the inherent measurement errors of the data are taken into account and uncertainty compensation corrections are made to the data to improve data accuracy. The spatial correlation of the data is utilized to perform local weighted smoothing to achieve spatial dependency correction.

[0091] Specifically, data type discrimination can be achieved by constructing a standardized path selection function, which can be expressed as follows:

[0092] ,

[0093] in, A function is selected for the normalization path, which automatically chooses the best preprocessing (transformation) method depending on the type of input variables; it ensures that each type of data is most appropriately prepared before entering dynamic normalization.

[0094] The Box-Cox transformation is used for continuous data (such as rainfall, temperature, humidity, soil pH, etc.). By finding an optimal parameter, the original non-normal distribution data (such as skewed distribution) is transformed into an approximate normal distribution, or its variance is stabilized to make it more consistent with the assumptions of many statistical models.

[0095] Fisher-z transformation is used to transform remote sensing related data to make its distribution closer to normal and stabilize its variance.

[0096] is the current observation value, For the the currently observed value of the class variable, is the minimum value of all current observations, is the interquartile range, The interquartile range normalization function is used to linearly scale categorical or ordinal data (e.g., fertilization intensity levels: low, medium, high; irrigation levels: none, small, moderate, large, etc.) to a relatively uniform scale and reduce the impact of outliers.

[0097] It should be noted that the standardized path selection function can be used as a pre-processing for dynamic standardization. Afterwards, according to the variable Type, select the corresponding transformation function Processing is performed to obtain a pre-processed This preprocessed value will be used as the actual calculation The original observations used.

[0098] Considering that the collected data is accompanied by uncertainty or error, the impact of these errors on the normalization results can be quantified and corrected to make the data closer to the true value. In this embodiment, a measurement error function is used to compensate for the uncertainty. The measurement error function can be expressed as follows:

[0099] ,

[0100] in, is the standardized value after taking into account the measurement error, is the normalized value from the dynamic normalization equation, is the natural base function symbol, is the attenuation factor coefficient, which is used to control the intensity of the influence of measurement error on the standardized value. It converts the data signal-to-noise ratio into a coefficient that can be used to measure error attenuation. is the data error parameter, used to quantify the The data of class variables are inherently inaccurate.

[0101] Finally, through geographically weighted normalization, the data is locally smoothed and corrected so that the final result is more consistent with the characteristics of geographic spatial data. In this embodiment, geographically weighted normalization can be expressed by the following formula:

[0102] ,

[0103] in, is the index of the neighboring observation point, is the total number of near observation points, For observation points The standardized value after considering the measurement error; is a geographically weighted kernel function used to measure the observation points To the target point The kernel function can use the commonly used Gaussian attenuation function to determine the intensity of the influence.

[0104] The final standardized value after spatial dependency correction is obtained by combining all the standardized values ​​after spatial dependency correction into one data set, which is the final clean data set obtained through the hierarchical adaptive standardization model. The data in this data set are aligned in time and space, unified in dimension, corrected in environmental background, compensated in measurement error, and smoothed in space, thus providing a highly reliable data foundation for subsequent advanced applications such as arsenic pollution risk assessment, spatial distribution mapping, and driving factor analysis.

[0105] In this embodiment, the geographically weighted kernel function can be expressed by the following formula:

[0106] ,

[0107] in, For spatial location The Euclidean distance between the i-th observation point, For spatial location The Euclidean distance between the jth observation point and The spatial influence radius is used to characterize the rate of distance attenuation. A larger spatial influence radius means a wider range of influence for an observation point, and even distant data points can have a significant impact on the target point. A smaller spatial influence radius means a more limited range of influence, and only very close data points have a significant impact. is the weight decay exponent, which is used to control the nonlinearity of distance decay. When k=1, the decay is exponential linear decay. When k>1 (for example, k=2 becomes Gaussian decay), the decay rate is faster, which means that the weight given to data points with a longer distance is sharply reduced. The typical value range of this weight decay exponent is 1.5~2.5.

[0108] Step 3: Import the dataset processed by the hierarchical adaptive normalization model into the GIS software. Combined with a spatial interpolation algorithm, a spatial distribution map of arsenic content in rice is generated in the GIS software.

[0109] Specifically, in this embodiment, the data in the data set processed by the hierarchical adaptive normalization model usually exists in the form of point elements, and each point carries longitude, latitude, time and corresponding standardized values ​​after spatial dependency correction.

[0110] Since the data in this dataset are actually discrete sampling points, generating a continuous spatial distribution map requires spatial interpolation, which uses the values ​​of known points to estimate the values ​​of unknown points. In this example, spatial interpolation can be achieved using kriging to generate the final spatial distribution map of arsenic content in rice.

[0111] Example 2

[0112] This embodiment discloses a high-resolution spatial distribution system for predicting arsenic content in rice over a large scale.

[0113] The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for high-resolution spatial distribution of arsenic content in rice over a large scale in Example 1 can be implemented.

[0114] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for high-resolution spatial distribution of arsenic content in rice over a large scale, characterized by: The high-resolution spatial distribution method comprises: S100, collecting data related to arsenic content in rice, and integrating the collected data into a rice arsenic content database; S200, standardize the data in the rice arsenic content database using a hierarchical adaptive standardization model, The hierarchical adaptive normalization model fuses data by considering the characteristics of different data. The temporal synchronization and spatial matching of data are achieved by aligning them to the same temporal and spatial resolutions through interpolation methods; S300, importing the data set processed by the hierarchical adaptive normalization model into GIS software, and generating a spatial distribution map of arsenic content in rice in the GIS software by combining a spatial interpolation algorithm; The hierarchical adaptive normalization model includes the following sub-steps: S210, by weighting the data from different sources, aligning the data from different spatial locations and time points into a unified coordinate system, The data of each observation point is weighted according to the time and space distance, and the observation points closer to the target point will have a greater impact on the final result. Combined with the Gaussian kernel function, the influence of observation points far away from the target point on the result is reduced by attenuation. By setting credibility weights for observation points, different quality data are given different importance in the alignment process, thus constructing a multivariate spatiotemporal alignment function. S220, after completing the spatiotemporal alignment of the data through the multivariate spatiotemporal alignment function, dynamically standardize the data. The dynamic standardization comprehensively considers the significant differences in rice arsenic content data due to climate, soil, vegetation cover and human activities in different regions and at different times. The rice arsenic content data of different types and dimensions of variables were converted into dimensionless values ​​to eliminate the dimension and scale differences of the data; The multivariate spatiotemporal alignment function is expressed by the following formula: , in, For the Class data source multivariate spatiotemporal alignment function, is the target space-time point, are the horizontal and vertical coordinates of the spatial position, For time point; is the data of a specific observation point, is the set of all observation point data; For the The data credibility weight of each observation point; is the symbol of Gaussian kernel function, is the weighted Euclidean distance, is the anisotropic Gaussian kernel function, is the original observation matrix; The dynamic normalization is expressed by the following formula: , in, is a dynamic normalized value, Indicates the Class variables in spatial position Normalized value on ; is the original observation value, Indicates the Class variables in spatial position The original observations on ; For climate zones, is the dynamic mean of the climate zone, is the standard deviation of the climate zone, is the vegetation cover correction factor, is the compensation factor for human activities.

2. The method for high-resolution spatial distribution of arsenic content in rice over a large scale according to claim 1, characterized in that: The data related to rice arsenic content include: climate change data, remote sensing data, environmental factor data and historical rice arsenic content data.

3. The method for high-resolution spatial distribution of arsenic content in rice over a large scale according to claim 1, characterized in that: The hierarchical adaptive normalization model further includes: S230, standardization process integration optimization step, used to optimize the data after multivariate spatiotemporal alignment function and dynamic standardization, By distinguishing the data type, pre-conversion is performed according to the different characteristics of the variables. At the same time, the inherent measurement error of the data is taken into consideration, and the uncertainty compensation correction is performed on the data to improve the accuracy of the data. The spatial correlation of the data is used to perform local weighted smoothing to achieve spatial dependency correction.

4. The method for high-resolution spatial distribution of arsenic content in rice over a large scale according to claim 3, characterized in that: The data type discrimination is achieved by constructing a standardized path selection function. The normalized path selection function is expressed by the following formula: , in, is the normalized path selection function, is the original observation value; The Box-Cox transformation is for continuous data. It transforms the original non-normal distribution data into an approximate normal distribution by finding an optimal parameter. Fisher-z transformation is used to transform remote sensing data to make the distribution of remote sensing data closer to normal distribution and stabilize the variance of remote sensing data; is the current observation value, For the the currently observed value of the class variable, is the minimum value of all current observations, is the interquartile range, is the interquartile range standardization function.

5. The method for high-resolution spatial distribution of arsenic content in rice over a large scale according to claim 3, characterized in that: The uncertainty compensation correction is achieved through the measurement error function, The measurement error function is expressed by the following formula: , in, is the standardized value after taking into account the measurement error, is the normalized value from the dynamic normalization equation, is the natural base function symbol, is the attenuation factor coefficient, is the data error parameter.

6. The method for high-resolution spatial distribution of arsenic content in rice over a large scale according to claim 3, characterized in that: The spatial dependence correction is expressed by the following formula: , in, is the standardized value after spatial dependence correction, is the index of the neighboring observation point, is the total number of near observation points, For observation points The standardized value after considering the measurement error; is the geographically weighted kernel function.

7. A high-resolution spatial distribution system for predicting arsenic content in rice over a large scale, characterized by: The high-resolution spatial distribution system comprises: processor; A memory storing a computer program, which, when executed by a processor, implements the large-scale high-resolution spatial distribution method of arsenic content in rice as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Self-adaption spatial interpolation method and system based on spatial feature analysis

    CN103353923A

  • Risk assessment method for predicting arsenic element content in rice

    CN111047223A