A soil data simulation method integrating machine learning and spatial analysis collaboration
By integrating machine learning and spatial analysis, multi-source soil data are obtained, collaborative simulation maps are generated and consistent detection is carried out, the inefficiency and low-precision problems of traditional soil data acquisition and simulation methods are solved, and high-precision soil data simulation and real-time monitoring are realized, and precise agriculture and environmental protection are supported.
Patent Information
- Application Number
- CN202510739437.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Traditional soil data acquisition methods consume a lot of manpower and material resources, the data is fragmented and the simulation results are inaccurate, making it difficult to meet the needs of high accuracy. The existing models have not fully utilized spatial analysis technology and cannot reflect the spatial change law of soil properties.
Fusion machine learning and spatial analysis, by obtaining multi-source soil data, calculating the spatial characteristic parameters of soil data units, loading pre-trained machine learning models, generating collaborative simulation maps, and performing data consistency detection and parameter fusion, real-time monitoring and correction, generating high-precision soil data simulation results.
It realizes all-round collection of soil information, improves the accuracy and reliability of simulation results, provides intuitive spatial distribution information of soil attributes, and supports the decision-making of precise agriculture and environmental protection.
Smart Images

Figure CN120257853B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soil data processing, and in particular to a soil data simulation method that integrates machine learning and spatial analysis in a collaborative manner. Background Art
[0002] In many fields such as modern soil science research, agricultural production planning, and ecological environment monitoring, accurate soil data is an important basis for decision-making. However, traditional soil data acquisition and simulation methods face many difficulties and are difficult to meet the growing high-precision requirements.
[0003] From the perspective of data acquisition, on the one hand, soil data has a high degree of spatial heterogeneity. Soils in different geographical locations have significant differences in their physical properties (such as texture, porosity) and chemical properties (such as pH value, nutrient content). In large-scale regional studies, relying on traditional field sampling methods to obtain soil data not only consumes a large amount of human, material, and time costs, but also the number of sampling points is limited, making it difficult to comprehensively and accurately reflect the spatial variation of soil characteristics. For example, when conducting a soil fertility survey in a vast farmland area, sparse sampling points may miss local high or low soil fertility areas, resulting in misjudgment of the overall soil fertility status, which in turn affects the formulation of fertilization strategies, causing fertilizer waste or restricted crop growth. On the other hand, the data obtained by a single sensor has limitations. Although existing soil sensors can quickly obtain some soil information, each sensor can only monitor specific soil properties. For example, a soil moisture sensor can only measure soil water content and cannot simultaneously obtain other key information such as soil nutrient content. This makes the obtained soil data fragmented and difficult to form a comprehensive and integrated soil data set, bringing difficulties to subsequent data processing and analysis.
[0004] In terms of data simulation, traditional simulation methods also expose many deficiencies. Early soil data simulation models based on empirical formulas usually rely on simple linear relationship assumptions and ignore the complex non-linear interactions in the soil system. For example, when simulating the process of soil nutrient migration, these models do not fully consider the comprehensive effects of soil texture, microbial activities, and climate factors on nutrient migration, resulting in a large deviation between the simulation results and the actual situation. With the development of computer technology, some models based on physical processes have been developed, but these models often require a large number of input parameters and have extremely high requirements for the accuracy of the parameters. In practical applications, due to data acquisition limitations, many parameters are difficult to accurately measure, thus affecting the simulation accuracy of the models. In addition, most of these models do not fully utilize spatial analysis techniques and cannot effectively integrate the spatial characteristics of soil data, making the expression of the simulation results in terms of spatial distribution inaccurate and unable to intuitively reflect the variation law of soil properties at different geographical locations.
[0005] With the rapid development of emerging technologies such as big data and artificial intelligence, the field of soil data processing has ushered in new opportunities and challenges. Key challenges await: integrating multi-source soil data, fully exploring the spatial characteristics and underlying patterns within soil data, and leveraging advanced machine learning algorithms and spatial analysis techniques to achieve high-precision soil data simulation. This invention addresses this context, aiming to overcome the shortcomings of traditional methods and provide a more accurate and efficient soil data simulation solution for soil science research and related applications. Summary of the Invention
[0006] The purpose of the present invention is to provide a soil data simulation method that integrates machine learning and spatial analysis to solve the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a soil data simulation method that integrates machine learning and spatial analysis, the method comprising:
[0008] Acquire multi-source soil data of the target area and determine spatial characteristic parameters of each soil data unit from a preset spatial analysis model, wherein the spatial characteristic parameters include geographic location attributes, physical attributes, and chemical attributes of the soil data unit;
[0009] For each soil data unit, a machine learning characteristic index of the soil data unit is calculated based on the spatial characteristic parameters, and a simulated input set of the soil data unit is generated based on the multiple machine learning characteristic indexes; wherein the machine learning characteristic index is a comprehensive quantitative parameter that integrates the geographical location attribute and physical and chemical properties of each soil data unit; and the simulated input set is a standardized data set composed of the selected multiple machine learning characteristic indexes;
[0010] For each soil data unit, a pre-trained machine learning model is loaded, and based on the correlation between multiple characteristic indicators included in the simulation input set and the inference rules of the machine learning model, a dynamic evolution parameter of the soil data unit during the simulation process is determined;
[0011] Marking the dynamic evolution parameters for each soil data unit in the spatial analysis model to generate a collaborative simulation map;
[0012] For each soil data unit, performing data consistency detection according to the dynamic evolution parameter to obtain a detection result;
[0013] Soil data simulation is performed based on the detection results to generate soil data simulation results.
[0014] Preferably, simulating soil data based on the detection result to generate a soil data simulation result includes: when the detection result indicates that there are conflicts among multiple dynamic evolution parameters included in the soil data unit, determining the soil data unit as an abnormal unit; obtaining the conflicting parameters of the abnormal unit as the correction object, and determining the multi-source data priorities corresponding to the conflicting parameters; and performing parameter fusion on the soil data unit based on the data priorities to generate a soil data simulation result.
[0015] Preferably, the multi-source soil data includes first sensor data and second sensor data, the data priority of the first sensor data is the first priority, and the data priority of the second sensor data is the second priority; performing parameter fusion on the soil data unit based on the data priorities to generate a soil data simulation result includes:
[0016] Adjusting the simulation input set corresponding to the second sensor data according to the order of the first priority and the second priority;
[0017] Based on the collaborative simulation map, extracting the dynamic evolution parameters of the second sensor data in the adjusted simulation input set;
[0018] Combining the dynamic evolution parameters and the collaborative simulation map for parameter fusion to generate a soil data simulation result.
[0019] Preferably, for each soil data unit, calculating the machine learning feature index of the soil data unit based on the spatial feature parameter respectively, and generating the simulation input set according to multiple machine learning feature indexes, includes:
[0020] For each soil data unit, determining multiple adjacent units adjacent to it in space;
[0021] Calculating the correlation degree between the spatial feature parameter and the physical and chemical properties of each adjacent unit, and determining the first feature unit with the highest correlation degree from the adjacent units based on multiple correlation degrees;
[0022] Based on the first feature unit, determining multiple extended units adjacent to the feature unit, calculating the correlation degree between the spatial feature parameter and the physical and chemical properties of each extended unit, and determining the next feature unit with the highest correlation degree from the extended units based on multiple correlation degrees;
[0023] Repeatedly executing the steps of determining multiple extended units adjacent to the feature unit, calculating the correlation degree of each extended unit, and determining the next feature unit with the highest correlation degree until all soil data units in the target area are traversed to obtain the machine learning feature index set of the soil data unit;
[0024] Generate the simulated input set based on the set of machine learning feature metrics.
[0025] Preferably, calculating the correlation degree between the spatial feature parameters and the physical and chemical properties of each adjacent unit includes:
[0026] Calculate the first correlation degree between the geographical location attribute and the physical property of each adjacent unit, and the second correlation degree between the geographical location attribute and the chemical property of each adjacent unit respectively;
[0027] For each adjacent unit, use the weighted sum of the first correlation degree and the second correlation degree as the correlation degree of the adjacent unit.
[0028] Preferably, determining the first feature unit with the highest correlation degree from the adjacent units based on the multiple correlation degrees includes:
[0029] Store the multiple correlation degrees corresponding to the multiple adjacent units into a feature candidate list, and perform a descending order sorting on the correlation degrees in the feature candidate list to obtain a sorting result;
[0030] Based on the sorting result, determine the adjacent unit corresponding to the highest correlation degree as the first feature unit.
[0031] Preferably, after performing soil data simulation based on the detection result and generating a soil data simulation result, it further includes:
[0032] When the soil data simulation result is generated, for each soil data unit, monitor the attribute difference value between the soil data unit and the adjacent unit in real time;
[0033] When there is an attribute difference value exceeding a preset threshold, use the adjacent unit as an abnormal reference unit, and perform dynamic parameter correction on the soil data unit based on the abnormal reference unit to generate an updated soil data simulation result.
[0034] Preferably, the training process of the pre-trained machine learning model includes:
[0035] Obtain a historical soil data set, and extract the mapping relationship between the spatial feature parameters and the dynamic evolution parameters in the data set;
[0036] Construct an initial machine learning model according to the mapping relationship, and use the cross-validation method to iteratively optimize the model;
[0037] When the error between the predicted parameters output by the model and the measured parameters is lower than a preset threshold, determine that the model training is completed.
[0038] Preferably, the spatial analysis model further includes three-dimensional spatial layers of terrain undulation, soil type distribution, and hydrological conditions. When generating the collaborative simulation atlas, the dynamic evolution parameters are superimposed on the three-dimensional spatial layers for visual expression.
[0039] Preferably, the method for generating the three-dimensional spatial layers includes:
[0040] Collect elevation data, soil sampling data, and groundwater level data of the target area;
[0041] Perform interpolation processing on the elevation data to generate a terrain undulation layer, perform classification coding on the soil sampling data to generate a soil type distribution layer, and perform spatial interpolation on the groundwater level data to generate a hydrological conditions layer;
[0042] Perform spatial overlay analysis on the terrain undulation layer, soil type distribution layer, and hydrological conditions layer to form the three-dimensional spatial layer.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] 1. In the present invention, from the perspective of the comprehensiveness and accuracy of data processing, the method realizes the comprehensive collection of soil information by obtaining multi-source soil data of the target area and determining the spatial characteristic parameters of each soil data unit, covering geographical location attributes, physical attributes, and chemical attributes. When calculating the characteristic indicators of machine learning, the geographical location attributes and physical and chemical attributes of the soil data unit are fully integrated to generate a standardized simulation input set, which enables the subsequent simulation process to comprehensively consider the influence of various factors. Compared with the traditional method that only relies on a single attribute or a small amount of data for simulation, the present invention greatly improves the integrity and accuracy of the data and can more truly reflect the actual situation of the soil. For example, when analyzing the soil fertility of a certain area, not only the chemical attributes of soil nutrients are considered, but also its geographical location and the physical attributes of the surrounding soil are combined to accurately locate the fertility change area and avoid misjudgment caused by data one-sidedness.
[0045] 2. In the present invention, during the simulation process, a pre-trained machine learning model is loaded, and dynamic evolution parameters are determined based on the correlation degree between feature indicators in the simulation input set and the model inference rules. This process makes full use of the powerful learning and prediction capabilities of machine learning. The machine learning model can automatically mine complex non-linear relationships in the data and capture the changing rules of soil properties that are difficult to discover by traditional methods. Through learning and iterative optimization of a large number of historical soil data sets, the model can more accurately predict the dynamic changes of soil data units during the simulation process, improving the reliability of the simulation results. When simulating the migration of heavy metal elements in soil, the model can comprehensively consider the dynamic effects of various factors such as soil pH, texture, and microbial activities, providing a more scientific basis for soil pollution prevention and control.
[0046] 3. In the present invention, in the link of generating the collaborative simulation map, the dynamic evolution parameters are superimposed on a three-dimensional spatial layer containing terrain undulation, soil type distribution, and hydrological conditions for visual expression, providing researchers with intuitive and comprehensive spatial distribution information of soil data. This visualization method helps to quickly identify the spatial differences and changing trends of soil properties, facilitating spatial analysis and decision-making by researchers. When formulating an agricultural irrigation plan, by combining the collaborative simulation map, it is possible to clearly see the soil moisture distribution in different terrain and soil type areas, reasonably plan the irrigation area and water volume, and improve the water resource utilization efficiency.
[0047] 4. In the present invention, the data consistency detection mechanism is also a major highlight of the present invention. During the simulation process, data consistency detection is performed for each soil data unit. When a conflict in dynamic evolution parameters is found, the abnormal unit can be determined in a timely manner, and parameter fusion is performed according to the priority of multi-source data. This mechanism ensures the quality of the simulation data and avoids deviations in the simulation results caused by data conflicts. When obtaining soil data from multiple sensors, different sensors may produce data differences due to factors such as accuracy and measurement environment. Through data consistency detection and parameter fusion, these data can be effectively integrated, making the simulation results more in line with the actual situation.
[0048] 5. In the present invention, after the soil data simulation results are generated, the attribute difference values between the soil data unit and adjacent units are monitored in real time, and the dynamic parameters of the soil data unit are corrected based on the abnormal reference unit, further improving the timeliness and accuracy of the simulation results. During the agricultural production process, soil properties change with time and environmental factors. Through dynamic monitoring and correction, these changes can be reflected in a timely manner, providing real-time and accurate data support for agricultural production management and facilitating the development of precision agriculture. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the working principle diagram of the soil data simulation method that combines machine learning and spatial analysis collaboration of the present invention;
[0050] Figure 2 Flow chart for calculating machine learning feature metrics and generating a simulation input set
[0051] Figure 3 Flow chart for calculating the correlation degree between adjacent units
[0052] Figure 4 Flow chart for dynamically correcting the simulation results of soil data
[0053] Figure 5 Distribution map of soil sampling points
[0054] Figure 6 pH prediction map obtained by the method in the present invention
[0055] Figure 7 Total nitrogen prediction map obtained by the method in the present invention
[0056] Figure 8 Organic matter prediction map obtained by the method in the present invention
[0057] Figure 9 Available phosphorus prediction map obtained by the method in the present invention
[0058] Figure 10 Available potassium prediction map obtained by the method in the present invention
[0059] Figure 11 Soil conductivity prediction map obtained by the method in the present invention Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0061] Please refer to Figures 1-4 , the present invention provides a technical solution: The present invention relates to a soil data simulation method that combines machine learning and spatial analysis in a coordinated manner. The specific implementation solution is as follows:
[0062] S1. When conducting soil data simulation, it is necessary to obtain multi-source soil data of the target area first. The sources of these data are extensive. For example, data can be collected through a professional soil sensor network, which can be distributed at different locations in the target area to collect various types of soil information in real time; or through on-site sampling. Researchers select multiple sampling points in the target area, collect soil samples and bring them back to the laboratory for analysis and determination. After obtaining the data, use a preset spatial analysis model to determine the spatial characteristic parameters of each soil data unit. For the geographical location attribute, the global positioning system (GPS) can be used to accurately locate the location information of each soil data unit; in terms of physical properties, soil texture can be determined by sieving method, and bulk density can be measured by the core cutter method; for chemical properties such as soil pH value, a pH meter can be used for measurement, and nutrient content can be determined by chemical analysis methods.
[0063] S2. For each soil data unit, calculate the machine learning feature indicators based on the spatial characteristic parameters obtained previously. This requires comprehensively considering the internal relationship between geographical location attributes and physical and chemical properties. For example, analyze the relationship between the climate conditions of the soil's geographical location and the soil nutrient loss rate to construct a comprehensive quantification parameter. After obtaining multiple such machine learning feature indicators, select appropriate indicators from them to form a simulation input set. To ensure the standardization and comparability of the data, this set also needs to be standardized, such as performing normalization operations on the data of different indicators to make all data within the same order of magnitude range.
[0064] S3. Load a pre-trained machine learning model to process each soil data unit. This model is trained based on a large amount of historical soil data and has learned the complex patterns and rules in the data. According to the correlation degree between the feature indicators in the simulation input set and the inference rules set inside the machine learning model, predict the dynamic evolution parameters of the soil data unit during the simulation process. For example, based on the correlation between the current nutrient content, humidity of the soil and the feature indicators of the surrounding environmental factors, the model can predict parameters such as the change trend of soil nutrients and the fluctuation range of humidity in the future period.
[0065] S4. Mark the determined dynamic evolution parameters on each soil data unit in the spatial analysis model to generate a collaborative simulation map. The spatial analysis model can be constructed based on the geographical information system (GIS), which can intuitively display the spatial distribution of soil data. By combining the dynamic evolution parameters with the spatial location, the change trends of soil data in different regions and the mutual influence relationships between each unit can be clearly seen on the map, providing an intuitive basis for subsequent analysis.
[0066] S5. Perform data consistency detection on each soil data unit according to the dynamic evolution parameters. Check whether there are contradictions or irrationalities among these parameters, such as whether the change trend of soil humidity conforms to the precipitation and irrigation conditions in this area, and whether the change of soil pH is coordinated with the surrounding soil and environmental factors. If conflicts are found in the dynamic evolution parameters of a certain soil data unit, subsequent processing is required.
[0067] S6. Conduct soil data simulation based on the detection results. If the detection results show that the data is normal, the simulation can be directly carried out according to the dynamic evolution parameters; if there are abnormalities, they should be corrected according to the corresponding rules and then simulated, and finally an accurate and reliable soil data simulation result is obtained. This result can be applied to agricultural production planning to help farmers fertilize and irrigate reasonably; it can also be used in the field of environmental protection to evaluate the diffusion trend of soil pollution, etc.
[0068] Example 1: During the process of soil data simulation, when conducting soil data simulation based on the detection results, if it is detected that there are conflicts in multiple dynamic evolution parameters included in a soil data unit, a series of measures need to be taken. First, clarify where the conflict is, and identify the soil data unit with parameter conflicts as an abnormal unit. For example, when simulating the change of soil nutrients in a certain area, it is found that the change trend of the nitrogen element content in a soil data unit is contradictory to the change trends of phosphorus and potassium elements, which does not conform to the normal law of coordinated change of soil nutrients. At this time, this unit is an abnormal unit.
[0069] After determining the abnormal unit, find out its conflict parameters as the objects to be corrected. Continuing with the above example, if there are conflicts in the change trends of nitrogen, phosphorus, and potassium element contents, then the relevant parameters of these elements are the objects to be corrected. Next, determine the priority of multi-source data corresponding to the conflict parameters. Multi-source data may come from different measurement means or devices. For example, some data is obtained through high-precision professional sensors, while some is collected by ordinary monitoring devices. The data of high-precision professional sensors is more accurate, and its data priority can be set higher; the data priority of ordinary monitoring devices is relatively lower.
[0070] Perform parameter fusion on the soil data unit based on the data priority to generate a soil data simulation result. Taking soil nutrient data as an example, give priority to referring to the nutrient change trend reflected by high-priority data. If the high-precision sensor shows that the nitrogen element content is stable, while the ordinary device shows large fluctuations, then during parameter fusion, based on the data of the high-precision sensor, make appropriate adjustments in combination with the data of the ordinary device. For example, analyze the reasons for the fluctuations in the data of the ordinary device, whether it is affected by external interference, etc., and then reasonably correct its data, and then fuse it with the data of the high-precision sensor to finally generate a soil data simulation result that conforms to the actual situation, ensuring that the simulation result can accurately reflect the true state of soil nutrients.
[0071] Example 2: In practical applications, multi-source soil data often includes first sensor data and second sensor data, and the data priority of the first sensor data is the first priority, and the data priority of the second sensor data is the second priority. When generating a soil data simulation result by parameter fusion of soil data units based on the data priority, the specific operations are as follows.
[0072] According to the order of the first priority and the second priority, adjust the simulation input set corresponding to the second sensor data. Assume that the first sensor is a professional soil nutrient sensor with high precision and good stability; the second sensor is an ordinary multi-functional soil sensor. Since the first sensor data has a higher priority, the second sensor data is processed with reference to the nutrient data range and precision standard measured by it. For example, the soil nitrogen element content data measured by the second sensor has a wider range and lower precision. According to the measurement precision of the first sensor, the nitrogen element content data of the second sensor is screened and refined, and obvious unreasonable data points are removed, so that the second sensor data is more comparable with the first sensor data in terms of precision and range.
[0073] Based on the co-simulation map, extract the dynamic evolution parameters of the second sensor data in the adjusted simulation input set. The co-simulation map intuitively shows the spatial distribution and dynamic changes of soil data. Based on the adjusted simulation input set, find the soil data unit corresponding to the second sensor data from the map, and extract its dynamic evolution parameters during the simulation process, such as the change parameters of soil humidity over time, the change trend parameters of soil acidity and alkalinity, etc.
[0074] Combine the dynamic evolution parameters with the co-simulation map for parameter fusion to generate a soil data simulation result. During the fusion process, fully consider the mutual relationship between each soil data unit in the co-simulation map. For example, the soil humidity of other units around a certain soil data unit is high, and it shows a trend of water convergence in the co-simulation map. Then, when fusing the dynamic evolution parameters of the second sensor data, this surrounding environmental factor needs to be considered. If the second sensor data shows that the soil humidity of this unit is low, but through combined map analysis, it may be a measurement error or a temporary fluctuation. At this time, the dynamic evolution parameters of the second sensor data need to be adjusted according to the overall trend in the map and then fused with other data to generate a soil data simulation result that is more in line with the actual situation and improve the accuracy of the simulation.
[0075] Example 3: For each soil data unit, when respectively calculating the machine learning feature indicators based on the spatial feature parameters and generating a simulation input set, the specific process is as follows.
[0076] Identify multiple adjacent units that are spatially adjacent to it. This is achieved by using a spatial analysis algorithm and combining the geographical location information of the soil data units. For example, when simulating soil data in a farmland area, with a specific soil data unit as the center, a certain search radius is set , and through the spatial query function of the Geographic Information System (GIS), other soil data units within this radius are obtained as adjacent units. Here, the search radius is determined according to the size of the study area and the scale of soil property changes, with the unit of meters (m). These adjacent units are closely connected to the target unit in space, and their soil properties may have an impact on the target unit.
[0077] Calculate the correlation degree between the spatial characteristic parameters and the physical and chemical properties of each adjacent unit. Analyze the correlation between the geographical location attributes and the physical properties, and the correlation between the geographical location attributes and the chemical properties of each adjacent unit respectively. For example, it is found that there is a certain relationship between the bulk density (physical property) of the soil in a certain area and the altitude (geographical location attribute), and the bulk density of the soil is relatively low at higher altitudes; at the same time, the soil pH value (chemical property) is related to the topography and landform (geographical location attribute), and the soil pH value may be relatively high in valley areas. Through statistical analysis of a large amount of data, determine the tightness of this relationship, and obtain the first correlation degree (the correlation degree between the geographical location attribute and the physical property) and the second correlation degree (the correlation degree between the geographical location attribute and the chemical property) respectively. For each adjacent unit, the first correlation degree and the second correlation degree are synthesized according to certain rules. Here, the weighted sum method is used to calculate the correlation degree of this adjacent unit , and the formula is: , where and are the weights of the first correlation degree and the second correlation degree respectively, and . The values of and are set according to the actual research focus. If more attention is paid to the impact of physical properties on soil data, the value of can be appropriately increased; if more attention is paid to the impact of chemical properties, the value of is increased. Finally, the correlation degree of this adjacent unit is obtained.
[0078] Determine the first characteristic unit with the highest degree of association from adjacent units based on multiple degrees of association. Store the multiple degrees of association corresponding to multiple adjacent units in a characteristic candidate list, and use sorting algorithms in computer programs, such as bubble sort or quick sort, to sort the degrees of association in the list in descending order. After sorting, the adjacent unit corresponding to the degree of association at the top of the list is the unit with the highest degree of association, and it is determined as the first characteristic unit. For example, in the simulation of soil data in a certain orchard, after calculation and sorting, it is found that an adjacent unit that is closer to the target unit and has a high degree of similarity in soil texture, nutrient content, and other attributes with the target unit has the highest degree of association, and it is determined as the first characteristic unit.
[0079] Based on the first characteristic unit, determine multiple extended units adjacent to this characteristic unit, calculate the degree of association between the spatial characteristic parameters and physical and chemical properties of each extended unit again, and based on these degrees of association, determine the next characteristic unit with the highest degree of association from the extended units. Continuously repeat this process, that is, determine multiple extended units adjacent to the characteristic unit, calculate the degree of association and determine the next characteristic unit with the highest degree of association, until all soil data units in the target area are traversed, so as to obtain a set of machine learning characteristic indicators for soil data units.
[0080] Generate a simulation input set based on the set of machine learning characteristic indicators. Screen and organize the obtained set of machine learning characteristic indicators, remove obviously unreasonable or redundant indicators, and then combine the remaining indicators into a simulation input set in a certain order and format to provide high-quality data input for subsequent machine learning model processing.
[0081] Example 4: When calculating the degree of association between the spatial characteristic parameters and physical and chemical properties of each adjacent unit, the specific calculation method is as follows.
[0082] For each adjacent unit, first analyze the relationship between its geographical location attributes and physical properties to determine the first degree of association. Taking soil texture (physical property) and terrain slope (geographical location attribute) in geographical location as an example, it is found in soil research in mountainous areas that in areas with steeper slopes, the soil particles are relatively coarser and the texture is relatively loose; while in areas with gentler slopes, the soil particles are relatively finer and the texture is relatively compact. By collecting a large number of soil samples in multiple areas with different slopes, analyzing the relationship between soil texture and slope, and using statistical analysis methods, such as correlation analysis, to determine the quantitative relationship between the two, and then obtaining the first degree of association.
[0083] Meanwhile, calculate the second correlation degree between the geographical location attributes and chemical attributes of each adjacent unit. For example, study the relationship between soil pH (chemical attribute) and the groundwater level where the soil is located (a geographical location-related factor). In some plain areas, the soil in areas with a higher groundwater level is prone to salinization, resulting in an increase in soil pH. By long-term monitoring of the changes in soil pH in areas with different groundwater levels, collecting a large amount of data for analysis, and establishing a connection model between the groundwater level and soil pH, the second correlation degree can be determined.
[0084] For each adjacent unit, comprehensively calculate the first correlation degree and the second correlation degree to obtain the correlation degree of this adjacent unit. The comprehensive calculation here can adopt the weighted sum method, and the weights are set according to the actual research focus and data characteristics. If the current research mainly focuses on the impact of the physical properties of the soil on plant growth, then the weight of the first correlation degree can be appropriately increased; if more attention is paid to the impact of soil chemical properties on the ecological environment, then the weight of the second correlation degree is increased. By reasonably setting the weights, the comprehensive correlation degree between the adjacent unit and the target soil data unit in terms of spatial characteristic parameters and physical and chemical attributes can be more accurately reflected, providing more reliable data support for subsequent determination of characteristic units and generation of simulation input sets.
[0085] Example 5: When determining the first characteristic unit with the highest correlation degree from adjacent units based on multiple correlation degrees, the specific operation is as follows.
[0086] Store the multiple correlation degrees corresponding to multiple adjacent units into the characteristic candidate list. In actual operation, use data structures in computer programming languages, such as arrays or lists, to implement. For example, in the Python language, a list object can be created, and the correlation degrees calculated for each adjacent unit are sequentially added to this list. In this way, the correlation degrees of all adjacent units are centrally stored in a data structure, facilitating subsequent processing.
[0087] Sort the correlation degrees in the characteristic candidate list in descending order to obtain the sorting result. Various mature sorting algorithms can be used to implement, such as the bubble sort algorithm. Its basic principle is to gradually "bubble" the largest (or smallest) element to the end of the list by comparing adjacent elements multiple times and swapping positions. When sorting the correlation degrees, select descending order sorting, that is, arrange the larger correlation degrees in the front. After sorting, the elements in the list are arranged in descending order of correlation degree.
[0088] Based on the sorting result, the adjacent unit corresponding to the highest correlation degree is determined as the first feature unit. For example, when processing the soil data of a certain experimental field, the correlation degrees of multiple adjacent units with the target unit are calculated and stored in the feature candidate list. After bubble sorting, the adjacent unit corresponding to the first element in the list is the unit with the highest correlation degree, which is determined as the first feature unit. This first feature unit is most closely related to the target unit in terms of spatial feature parameters and physical and chemical properties, and can represent the main influencing factors of the surrounding soil data units on the target unit, providing important basic data for subsequent construction of the machine learning feature index set and generation of the simulation input set.
[0089] Example 6: After generating the soil data simulation result based on the detection result, subsequent monitoring and correction work need to be carried out. The specific steps are as follows.
[0090] When the soil data simulation result is generated, for each soil data unit, the attribute difference value between the soil data unit and the adjacent unit is monitored in real time. This can be achieved by arranging multiple sensors in the target area to form a real-time monitoring network. These sensors regularly collect the attribute information of the soil data unit and its adjacent unit, such as soil humidity, nutrient content, pH value, etc. Then, a data processing program is used to calculate the difference value between the two. For example, calculate the difference in soil humidity between a certain soil data unit and its adjacent unit. If the difference is large, it may mean that there is a special water distribution situation in this area.
[0091] When there is an attribute difference value exceeding the preset threshold, the adjacent unit is used as an abnormal reference unit, and the dynamic parameters of the soil data unit are corrected based on the abnormal reference unit to generate an updated soil data simulation result. The preset threshold can be determined according to the actual research accuracy requirements and the characteristics of the soil data. For example, when studying the soil data under the growth environment of a certain specific crop, according to the sensitivity of the crop to soil nutrients, the threshold for the difference in soil nutrient content is set. If it is found that the difference in nitrogen element content between a certain soil data unit and its adjacent unit exceeds the preset threshold, then the adjacent unit is used as an abnormal reference unit. Analyze the reasons for the difference. It may be that there are local fertilization differences in the area where the soil data unit is located, or it is affected by the surrounding irrigation water. According to the situation of the abnormal reference unit, adjust the dynamic parameters of the soil data unit, such as adjusting the change trend parameter of the soil nitrogen element in this unit, and re-perform the simulation to generate an updated soil data simulation result. Through this real-time monitoring and dynamic correction method, abnormal situations in the soil data can be discovered in a timely manner, and the soil data simulation result can be continuously optimized to make it more in line with the actual soil conditions, providing more accurate data support for related agricultural production, environmental protection and other applications.
[0092] Example 7: Refer to the appendix Figures 5-11, To better illustrate the soil data simulation method that integrates machine learning and spatial analysis as described in the present invention, the following takes the experimental farmland in a certain place as an example for detailed description.
[0093] The experimental farmland covers an area of approximately 209 hectares and is divided into 458 plots. The average area of each plot is 0.45 hectares. The soil is mainly composed of sandy soil and loam. The main crops planted are winter wheat and summer maize, and a small part of the area is used for planting soybeans and peanuts. Straw is regularly returned to the field after each harvest season.
[0094] Obtain multi-source soil data for the target area and determine spatial characteristic parameters: According to a grid of 100 meters × 100 meters, 208 samples were collected from 0.3 meters above the soil surface, as shown in Table 1. Apparent conductivity data of each sampling point was obtained using a proximal sensor EM38 - MK2. At the same time, parameters such as soil pH value, organic matter, total nitrogen, available phosphorus, and available potassium were tested for each soil sample. Among them, the geographical location attribute is represented by the coordinates of the sampling point (for example, the coordinates of the sampling point numbered 1 are X = 564042, Y = 3761324); the physical attribute includes apparent conductivity; the chemical attributes include soil pH value, organic matter, total nitrogen, available phosphorus, and available potassium, etc. These data constitute the spatial characteristic parameters of each soil data unit, as shown in Table 2.
[0095] Table 1 Soil Sampling Point Data (Partial)
[0096]
[0097] Table 2 Descriptive Statistical Analysis of Soil Attributes
[0098]
[0099] Calculate machine learning feature indicators to generate a simulation input set: For each soil data unit, determine multiple adjacent units that are spatially adjacent to it. Taking a certain soil data unit as an example, calculate the first correlation degree between the geographical location attribute and the physical attribute of each adjacent unit, and the second correlation degree between the geographical location attribute and the chemical attribute. Take the weighted sum of the first correlation degree and the second correlation degree as the correlation degree of the adjacent unit. Store the multiple correlation degrees corresponding to the multiple adjacent units in a feature candidate list, and sort the correlation degrees in the list in descending order. Determine the adjacent unit corresponding to the highest correlation degree as the first feature unit. Based on the first feature unit, determine multiple extended units adjacent to this feature unit, and repeat the steps of calculating the correlation degree and determining the next feature unit with the highest correlation degree until all soil data units in the target area are traversed, obtaining the machine learning feature indicator set of this soil data unit. After standardizing these indicators, generate a simulation input set.
[0100] Determine dynamic evolution parameters: Load a pre-trained machine learning model, and the training process of this model is as follows: Obtain a historical soil dataset, and extract the mapping relationship between the spatial feature parameters and the dynamic evolution parameters in the dataset; Construct an initial machine learning model according to the mapping relationship, and use the cross-validation method to iteratively optimize the model; When the error between the predicted parameters output by the model and the measured parameters is lower than the preset threshold, it is determined that the model training is completed. Based on the correlation degree between multiple feature indicators included in the simulation input set and the inference rules of the machine learning model, determine the dynamic evolution parameters of each soil data unit during the simulation process.
[0101] Generate a collaborative simulation map: The spatial analysis model used in this embodiment includes three-dimensional spatial layers of terrain undulation, soil type distribution, and hydrological conditions. By collecting elevation data, soil sampling data, and groundwater level data of the target area, perform interpolation processing on the elevation data to generate a terrain undulation layer, perform classification coding on the soil sampling data to generate a soil type distribution layer, perform spatial interpolation on the groundwater level data to generate a hydrological condition layer, and perform spatial overlay analysis on these three layers to form a three-dimensional spatial layer. Mark the above-determined dynamic evolution parameters for each soil data unit in this spatial analysis model, and superimpose the dynamic evolution parameters onto the three-dimensional spatial layer for visual expression to generate a collaborative simulation map.
[0102] Perform data consistency detection: For each soil data unit, perform data consistency detection according to the determined dynamic evolution parameters. If the detection result indicates that there are conflicts among multiple dynamic evolution parameters included in a soil data unit, determine this soil data unit as an abnormal unit.
[0103] Generate the soil data simulation result: If there are abnormal units, obtain the conflict parameters of the abnormal units as the correction objects, and determine the priority of multi-source data corresponding to the conflict parameters (assuming that in this embodiment, there is first sensor data and second sensor data, the data priority of the first sensor data is the first priority, and the data priority of the second sensor data is the second priority). According to the order of the first priority and the second priority, adjust the simulation input set corresponding to the second sensor data; Based on the collaborative simulation map, extract the dynamic evolution parameters of the second sensor data in the adjusted simulation input set; Combine this dynamic evolution parameter with the collaborative simulation map for parameter fusion to generate the soil data simulation result.
[0104] Dynamic parameter correction: After the soil data simulation result is generated, for each soil data unit, monitor the attribute difference value between this soil data unit and its adjacent units in real time. When there is an attribute difference value exceeding the preset threshold, use the adjacent unit as an abnormal reference unit, and perform dynamic parameter correction on the soil data unit based on the abnormal reference unit to generate an updated soil data simulation result.
[0105] Through the above steps, the soil data simulation of the target area of this farmland is completed, demonstrating the implementation process of the method described in the present invention in practical applications.
[0106] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0107] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A soil data simulation method integrating machine learning and spatial analysis collaboration, characterized in that, The method includes: S1. Obtain multi-source soil data of a target area, and determine the spatial characteristic parameters of each soil data unit from a preset spatial analysis model, where the spatial characteristic parameters include the geographical location attribute, physical attribute, and chemical attribute of the soil data unit; S2. For each soil data unit, calculate the machine learning feature indicators of the soil data unit respectively based on the spatial characteristic parameters, and generate a simulated input set according to the multiple machine learning feature indicators; where the machine learning feature indicators are comprehensive quantization parameters for each soil data unit, integrating its geographical location attribute and physical and chemical attributes; the simulated input set is a standardized data set composed of multiple selected machine learning feature indicators; S3. For each soil data unit, load a pre-trained machine learning model, and determine the dynamic evolution parameters of the soil data unit during the simulation process based on the correlation degree between the multiple feature indicators included in the simulated input set and the inference rules of the machine learning model; S4. Mark the dynamic evolution parameters for each soil data unit in the spatial analysis model to generate a collaborative simulation map; S5. For each soil data unit, perform data consistency detection according to the dynamic evolution parameters to obtain a detection result; S6. Perform soil data simulation based on the detection result to generate a soil data simulation result.
2. The soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 1, wherein, The performing soil data simulation based on the detection result to generate a soil data simulation result includes: when the detection result indicates that there are conflicts among the multiple dynamic evolution parameters included in the soil data unit, determine the soil data unit as an abnormal unit; obtain the conflicting parameters of the abnormal unit as the correction object, and determine the multi-source data priorities corresponding to the conflicting parameters; perform parameter fusion on the soil data unit based on the data priorities to generate a soil data simulation result.
3. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 1, characterized in that The multi-source soil data includes first sensor data and second sensor data, the data priority of the first sensor data is the first priority, and the data priority of the second sensor data is the second priority; the performing parameter fusion on the soil data unit based on the data priorities to generate a soil data simulation result includes: Adjust the simulated input set corresponding to the second sensor data according to the order of the first priority and the second priority; Based on the collaborative simulation map, extract the dynamic evolution parameters of the second sensor data in the adjusted simulated input set; Combine the dynamic evolution parameters and the collaborative simulation map to perform parameter fusion to generate a soil data simulation result.
4. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 1, characterized in that, The for each soil data unit, calculating the machine learning feature indicators of the soil data unit respectively based on the spatial characteristic parameters, and generating the simulated input set according to the multiple machine learning feature indicators includes: For each soil data unit, determine multiple adjacent units adjacent to it in space; Calculate the correlation degree between the spatial feature parameters and the physical and chemical properties of each adjacent unit, and determine the first feature unit with the highest correlation degree from the adjacent units based on multiple said correlation degrees; Based on the first feature unit, determine multiple extended units adjacent to the feature unit, calculate the correlation degree between the spatial feature parameters and the physical and chemical properties of each extended unit, and determine the next feature unit with the highest correlation degree from the extended units based on multiple said correlation degrees; Repeat the steps of determining multiple extended units adjacent to the feature unit, calculating the correlation degree of each extended unit, and determining the next feature unit with the highest correlation degree until all soil data units in the target area are traversed to obtain the machine learning feature index set of the soil data units; Generate the simulation input set based on the machine learning feature index set.
5. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 4, characterized in that The calculation of the correlation degree between the spatial feature parameters and the physical and chemical properties of each adjacent unit includes: Calculate the first correlation degree between the geographical location attribute and the physical property of each adjacent unit, and the second correlation degree between the geographical location attribute and the chemical property of each adjacent unit respectively; For each adjacent unit, use the weighted sum of the first correlation degree and the second correlation degree as the correlation degree of the adjacent unit.
6. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 4, characterized in that, The determination of the first feature unit with the highest correlation degree from the adjacent units based on multiple said correlation degrees includes: Store the multiple correlation degrees corresponding to the multiple adjacent units in a feature candidate list, and sort the correlation degrees in the feature candidate list in descending order to obtain a sorting result; Based on the sorting result, determine the adjacent unit corresponding to the highest correlation degree as the first feature unit.
7. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 1, characterized in that, After generating the soil data simulation result based on the detection result, it further includes: When the soil data simulation result is generated, for each soil data unit, real-time monitor the attribute difference value between the soil data unit and the adjacent unit; When there is an attribute difference value exceeding the preset threshold, use the adjacent unit as an abnormal reference unit, and perform dynamic parameter correction on the soil data unit based on the abnormal reference unit to generate an updated soil data simulation result.
8. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 1, characterized in that, The training process of the pre-trained machine learning model includes: Obtain a historical soil data set, and extract the mapping relationship between the spatial feature parameters and the dynamic evolution parameters in the data set; Construct an initial machine learning model according to the mapping relationship, and use the cross-validation method to iteratively optimize the model; When the error between the predicted parameter and the measured parameter output by the model is lower than the preset threshold, determine that the model training is completed.
9. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 1, characterized in that The spatial analysis model also includes three-dimensional spatial layers of terrain undulation, soil type distribution, and hydrological conditions. When generating the collaborative simulation atlas, overlay the dynamic evolution parameters on the three-dimensional spatial layers for visual expression.
10. A soil data simulation method integrating machine learning and spatial analysis collaboration according to claim 9, characterized in that, The generation method of the three-dimensional spatial layer includes: Collect elevation data, soil sampling data, and groundwater level data of the target area; Interpolate the elevation data to generate a terrain undulation layer, classify and code the soil sampling data to generate a soil type distribution layer, and perform spatial interpolation on the groundwater level data to generate a hydrological condition layer; Perform spatial overlay analysis on the terrain undulation layer, soil type distribution layer and hydrological condition layer to form the three-dimensional space layer.
Citation Information
Patent Citations
Intelligent construction method for water collection and corrosion prevention plough layer of slope cropland
CN118941955A
Farmland soil data digital analysis method and system
CN119904147A