Soil type map updating method fusing soil voluntary and soil map data

By constructing a soil-environmental text information model and clustering analysis, combining soil criterion text and soil map data, a representative sample set is generated and a soil prediction and mapping model is trained, which solves the problem that traditional soil maps and soil map texts are difficult to integrate, and improves the accuracy of soil prediction and mapping and the utilization value of historical data.

CN120508643APending Publication Date: 2025-08-19ZHENGZHOU NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535376.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

It is difficult to integrate and utilize traditional soil maps and soil map texts, resulting in low accuracy of soil prediction and mapping and insufficient historical soil information.

Method used

A soil-environmental text information model was constructed, and a representative sample point set was generated through cluster analysis and frequency sampling method, combining soil census text and soil map data, and a soil prediction and mapping model was trained.

Benefits of technology

The multi-source data fusion of soil prediction and mapping has been realized, the mapping accuracy and utilization value of historical data have been improved, and the spatial-text data alignment failure and multi-grained information fusion obstacles are broken.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508643A_ABST
    Figure CN120508643A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of soil prediction mapping. The invention provides a soil type map updating method fusing a soil voluntary and soil map data. According to the embodiment of the invention, by constructing the soil-environment text information model, the unstructured soil voluntary and the soil map vector polygon are uniformly converted into the sample point data set. Obtaining a partition clustering number based on a soil-environment text information model, and obtaining a class cluster corresponding to an environment factor combination in each parent material partition by adopting partition clustering analysis; based on the environment factor information of the environment factor combination class clusters and the quantified environment factor information in the soil-environment text information model, calculating the similarity between the environment factor combination and the text information framework, and further determining the soil type semantic information of each environment factor combination class cluster; and a representative point set corresponding to each soil type is screened out from each environment combination class cluster, and a quantitative soil-environment relationship is obtained by adopting a kernel density estimation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of soil prediction mapping, and in particular to a soil type map updating method that integrates soil chronicle text and soil map data. Background Art

[0002] Soil is the material foundation of human survival and an indispensable, non-renewable natural resource. Accurate understanding of soil spatial distribution is essential for its rational utilization. This information serves as fundamental data for precision agriculture, environmental change simulation, natural resource management and utilization, and global change monitoring, and holds significant scientific and application value. Soil prediction and mapping is the primary method for estimating soil spatial distribution information.

[0003] Traditional soil maps and textual materials such as soil annals are not only important data sources for predictive soil mapping but also core outputs of soil surveys. Many countries around the world have extensive historical soil maps, providing a crucial data foundation for predictive soil mapping. Compared with using field samples for predictive soil mapping, methods based on existing soil maps offer easier data acquisition, higher execution efficiency, and lower production costs. However, this approach is limited by the quality of existing soil maps. Currently, the majority of soil map resources are traditional soil maps, which generally lack high accuracy due to limitations in basic data quality and traditional mapping techniques. Through long-term research and practical activities, the field of soil science has accumulated a vast amount of textual materials containing soil information. my country's second soil survey alone produced over 2,000 volumes of soil annals at the county, municipal, provincial, and national levels. Due to the small number of representative sample points and ambiguous location information recorded in textual materials such as soil annals, they are currently primarily used as query tools, and the underlying soil information they contain has yet to be fully utilized and explored.

[0004] As two parallel outputs of the soil survey, soil maps and soil annals construct soil information systems through spatial representation and textual description, respectively, and are naturally complementary. However, due to their heterogeneous data forms, soil maps represent soil spatial distribution using vector geographic data, while soil annals record soil formation information using unstructured text, making their integration difficult. This technical barrier to multimodal data fusion not only restricts the further development of historical data but also poses challenges to the development of new soil prediction and mapping methods.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the Invention

[0007] The purpose of the embodiments of the present disclosure is to provide a soil type map updating method that integrates soil log text and soil map data, thereby overcoming one or more problems caused by the limitations and defects of related technologies, at least to a certain extent.

[0008] According to an embodiment of the present disclosure, a method for updating a soil type map by integrating soil log text and soil map data is provided, the method comprising: Determine soil types and their corresponding environmental factors based on soil annals text to build a soil-environment text information model, and use domain knowledge base and corpus to fill in the soil-environment text information model; The study area was divided into several different parent material zones using parent material factors. The number of zone clusters was determined by combining the soil-environment text information model. The environmental factor data of each parent material zone in the environmental factor geographic database were clustered using a clustering algorithm to generate environmental factor combination clusters. Based on the environmental factor type, the environmental factor text information of each instance in the soil-environment text information model is quantitatively processed. Combined with the environmental factor combination clusters, the comprehensive environmental similarity between each environmental factor combination cluster and the instance is calculated, and the optimal soil semantic information corresponding to each environmental factor combination is determined. Based on the optimal soil semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type. After spatially associating the soil type map with the environmental factor geographic database, the map was split into multiple soil type polygons according to geographic boundaries. Within each soil type polygon, representative sample points were selected using the frequency sampling method to generate a sample point set for each polygon. Merge the representative sample point set of each soil type and the sample point set of each polygon to generate the soil type sample point set; Soil prediction mapping models are trained and tested on soil type sample sets to generate and update soil type maps.

[0009] Furthermore, the steps of determining soil types and corresponding environmental factor information based on soil annals text to construct a soil-environment text information model, and using a domain knowledge base and a corpus to fill the soil-environment text information model include: Based on the soil annals text, determine the soil type and its corresponding environmental factor information; Analyze the language description characteristics of soil types and their corresponding environmental factors in soil chronicle texts, and construct a soil-environment text information model; Build the domain knowledge base and corpus required to extract soil-environment text information; When the expression pattern of text information is single and regular, a rule-based approach is used to extract soil-environment text information from the domain knowledge base and corpus; When the expression pattern of text information is complex, changeable and irregular, a method based on the BiLSTM-CRF model is used to extract soil-environment text information from the domain knowledge base and corpus.

[0010] Furthermore, the study area is divided into multiple parent material partitions using parent material factors. The number of partition clusters is determined by combining the soil-environment text information model. The environmental factor data of each parent material partition in the environmental factor geographic database are clustered using a clustering algorithm to generate environmental factor combination clusters, including the following steps: The study area is divided into several different parent material zones using parent material factors; Conduct statistical analysis on the extracted soil-environment text information to determine the parent material type of each parent material partition and the corresponding partition cluster number; Based on the environmental factor geographic database, select environmental factor data and perform standardized preprocessing on each environmental factor data; Based on the preprocessed environmental factor data, the K-means++ algorithm was used to cluster each parent material partition to generate environmental factor combination clusters.

[0011] Furthermore, the steps of quantitatively processing the environmental factor text information of each instance in the soil-environment text information model based on the environmental factor type, combining each environmental factor combination cluster, calculating the comprehensive environmental similarity between each environmental factor combination cluster and the instance, and determining the optimal soil semantic information corresponding to each environmental factor combination include: Based on the environmental factor type, the environmental factor text information of each instance in the soil-environment text information model is quantitatively processed; For each group of environmental factor combination clusters in each parent material partition, the comprehensive environmental similarity between the soil-environmental text information model and the environmental factor combination cluster is calculated based on the corresponding environmental factor information and the environmental factor text information in the soil-environmental text information model; The soil-environment text information model that best matches the environmental factor combination cluster is determined based on the comprehensive environmental similarity, and the soil type information in the model is passed to the environmental factor combination cluster to determine the optimal soil semantic information corresponding to each environmental factor combination.

[0012] Furthermore, based on the optimal soil semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type, including: Based on the optimal soil semantic information, a frequency histogram is created for each environmental factor in the environmental factor combination cluster; The pixel points whose environmental factor values fall into the highest frequency interval are determined as representative sample points corresponding to the environmental factor; if there are multiple pixels with the same highest frequency interval, all of them are selected; The representative sample points corresponding to each environmental factor were merged to form a representative sample point set corresponding to each soil type.

[0013] Furthermore, the step of merging the representative sample point set of each soil type and the sample point set of each polygon to generate the soil type sample point set includes: The representative sample point set of each soil type is merged with the sample point set of each polygon extracted from the soil type polygon, and duplicates are removed by geographic location to generate a non-redundant soil type sample point set.

[0014] Furthermore, the step of training and testing a soil prediction mapping model based on the soil type sample point set to generate and update a soil type map includes: The extreme gradient boosting algorithm was used to construct a soil prediction mapping model; The soil type sample point set is divided into a training sample point set and a verification sample point set according to a preset ratio; The soil prediction mapping model is trained and tested using the training sample set and the validation sample set to obtain a trained soil prediction mapping model; The trained soil prediction mapping model is used to predict the spatial distribution of soil types and generate and update soil type maps.

[0015] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: In the disclosed embodiments, the soil type map updating method, which integrates soil log text and soil map data, firstly, constructs a soil-environment text information model that parses soil log text and spatially interprets soil maps. This uniformly converts unstructured soil log text and soil map vector polygons into a sample dataset, thereby achieving multi-source data fusion modeling for soil prediction mapping. Based on the soil-environment text information model, the number of partition clusters is determined. Partition cluster analysis is then used to obtain clusters corresponding to environmental factor combinations within each parent material partition. Based on the environmental factor information of the environmental factor combination clusters and the quantified environmental factor information in the soil-environment text information model, the similarity between the environmental factor combination and the text information framework is calculated to determine the soil type semantic information for each environmental factor combination cluster. Representative point sets corresponding to each soil type are screened from each environmental combination cluster, and kernel density estimation is used to obtain quantitative soil-environment relationships. Furthermore, this method overcomes core issues in traditional technologies, such as the failure of spatial-text data alignment, barriers to multi-granular information fusion, and the lack of semantic parsing technology. This significantly improves the utilization value of historical data and provides a reliable data foundation and technical support for soil type map updates and improved mapping accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0017] Figure 1 A diagram showing the steps of a soil type map updating method that integrates soil chronicle text and soil map data in an exemplary embodiment of the present disclosure; Figure 2 A specific flow chart showing a soil type map updating method integrating soil log text and soil map data in an exemplary embodiment of the present disclosure is provided; Figure 3 An example diagram showing the structured organization of soil-environment text information in an exemplary embodiment of the present disclosure Figure 4 A soil type diagram inferred based on the method proposed in this application in an exemplary embodiment of the present disclosure is shown; Figure 5 A graph showing the estimated uncertainty of the mapping results of the method proposed in this application in an exemplary embodiment of the present disclosure; Figure 6 A soil type diagram inferred based on a direct screening method in an exemplary embodiment of the present disclosure is shown; Figure 7 A graph showing the inferred uncertainty of the direct screening method mapping results in an exemplary embodiment of the present disclosure; Figure 8 A diagram showing a traditional soil type in an exemplary embodiment of the present disclosure is shown; Figure 9 A diagram showing a conventional soil type in an exemplary embodiment of the present disclosure is shown; Figure 10 A soil type map corresponding to the mapping result of the method proposed in this application in an exemplary embodiment of the present disclosure is shown; Figure 11 A soil type map corresponding to the mapping result of the direct screening method in an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0018] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0019] In addition, the accompanying drawings are merely schematic illustrations of embodiments of the present disclosure and are not necessarily drawn to scale. Like reference numerals in the figures represent like or similar parts, and thus repeated descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically separate entities.

[0020] This example embodiment provides a soil type map updating method that integrates soil log text and soil map data. Figure 1 As shown in , the soil type map updating method integrating soil log text and soil map data may include: Step S101: determining soil types and their corresponding environmental factors based on soil annals text to construct a soil-environment text information model, and filling the soil-environment text information model with domain knowledge base and corpus; Step S102: Divide the study area into multiple parent material partitions using parent material factors, determine the number of partition clusters in combination with the soil-environment text information model, and use a clustering algorithm to perform cluster analysis on the environmental factor data of each parent material partition in the environmental factor geographic database to generate environmental factor combination clusters; Step S103: Quantifying the environmental factor text information of each instance in the soil-environment text information model based on the environmental factor type, calculating the comprehensive environmental similarity between each environmental factor combination cluster and the instance, and determining the optimal soil semantic information corresponding to each environmental factor combination; Step S104: Based on the optimal soil semantic information, a frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type; Step S105: After spatially associating the soil type map with the environmental factor geographic database, the map is split into multiple soil type polygons according to geographic boundaries. Within each soil type polygon, representative sample points are selected using a frequency sampling method to generate a sample point set for each polygon. Step S106: merging the representative sample point set of each soil type and the sample point set of each polygon to generate a soil type sample point set; Step S107: training and testing a soil prediction mapping model based on the soil type sample point set to generate and update a soil type map.

[0021] This soil type map updating method, which integrates soil text and soil map data, firstly constructs a soil-environment text information model that parses soil text and spatially interprets soil maps. This model converts unstructured soil text and soil map vector polygons into a unified point dataset, thereby enabling multi-source data fusion modeling for soil predictive mapping. Based on the soil-environment text information model, the number of partition clusters is determined. Partition cluster analysis is then used to identify clusters corresponding to environmental factor combinations within each parent material partition. Based on the environmental factor information of the environmental factor combination clusters and the quantitative environmental factor information in the soil-environment text information model, the similarity between the environmental factor combination and the text information framework is calculated to determine the semantic information of soil types for each environmental factor combination cluster. Representative point sets corresponding to each soil type are selected from each environmental combination cluster, and kernel density estimation is used to quantify soil-environment relationships. Furthermore, this method overcomes core issues inherent in traditional techniques, such as the failure of spatial-text data alignment, barriers to multi-granular information fusion, and the lack of semantic parsing technology. This method significantly enhances the value of historical data and provides a reliable data foundation and technical support for updating soil type maps and improving mapping accuracy.

[0022] Below, we will refer to Figures 1 to 11 The various steps of the soil type map updating method for fusing soil log text and soil map data in this example embodiment are described in more detail.

[0023] In step S101, soil-environment text information is extracted and structured: First, based on the county-level soil annals text data, the target soil type information and its corresponding environmental factor type are determined, the language description characteristics of the soil type and environmental factor information in the text data are analyzed, and a soil-environment text information model is designed; secondly, the domain knowledge base and corpus required for extracting text information are constructed, such as soil information dictionary, environmental factor dictionary, keyword dictionary, rule library, annotated corpus, etc.; then, two methods based on rules and Bi-directional Long Short-Term Memory Conditional Random Fields (BiLSTM-CRF) model are combined to extract different target text information. When the expression pattern of the target information is single and regular, the rule-based method is adopted; when the expression pattern of the target information is complex, changeable and irregular, the method based on the BiLSTM-CRF model is adopted; finally, the extraction results of the target information are summarized, and the extracted environmental factor information is filled into the soil-environment text information model framework (i.e., soil-environment text information model) in the format of "variable name: variable value", as shown in Figure 2. Figure 2As shown in the figure, (a) is the extraction result based on rules, (b) is the extraction result based on the BiLSTM-CRF model, and (c) is the result summary.

[0024] In step S102, soil zoning clustering is performed to obtain environmental factor combinations: First, the environmental factors used for clustering were selected and the data for each environmental factor was standardized and preprocessed. Then, the study area was divided into different parent material partitions based on the parent material factor information, and the number of soil types contained in each parent material partition was estimated based on the soil-environment text information set. For example, if it is known that the parent materials corresponding to soil types A, B, C, D, E, and F in a study area are P1, P2, P1, P2, P3, and P2, respectively, then the corresponding relationship between each parent material and soil type in the study area is: P1: A, C, P2: B, D, F, P3: E. It can be inferred that the number of soil types corresponding to parent material partitions P1, P2, and P3 should be 2, 3, and 1, respectively, which is the number of parent material partition clusters. Finally, based on the environmental factor geographic database, K-means++ and other algorithms were used to obtain the environmental factor combinations within each parent material partition through cluster analysis.

[0025] In step S103, soil environment factor combined semantic information is calculated: According to the environmental factor information of each combination and the environmental factor information in the soil-environment text information model, the comprehensive environmental similarity between the soil-environment text information model instance and the environmental factor combination cluster is calculated, the soil-environment text information model instance that best matches each environmental factor combination cluster is determined, and the soil type information in the instance is passed to the environmental factor combination cluster. Before semantic calculation, the environmental factor text description in the soil-environment text information model needs to be quantitatively converted based on the domain knowledge base and the environmental characteristics of the study area, such as converting "steeper slope" into "slope: 15-25°". The environmental factor combination cluster after semantic calculation has both spatial information and soil type information. Given a parent material information set , Quantified soil-environment text information framework set , the specific calculation process is as follows: 1) According to the parent material partition information, parent material partition The combined clusters of all environmental factors to be determined for soil semantic information are , the soil-environment text information framework set corresponding to the parent material condition is ; 2) For collections Any combination of environmental factors in , calculate the cluster center and the range of each environmental factor, and take The mode of the environmental factor values at each point is taken as the central value. The value range of each environmental factor in the current cluster is used as the value range of each environmental factor; 3) Traversal Corresponding soil-environment text frame set ,according to Calculation of cluster centers and ranges of environmental factors and The similarity of each frame in Similarity sequence . Represents a cluster of environmental factor combinations Frame for text with soil - environment The comprehensive environmental similarity between the two is measured using the Gower similarity coefficient, which is calculated as follows:

[0026] in, is the calculation function of comprehensive environment similarity, for and In the environmental factor set Gower similarity coefficient sequence on , is the number of environmental factors. The calculation formula is as follows:

[0027] in, and Environmental factors Corresponding cluster Central value and framework The text information value in Cluster Medium environmental factors The value range of and Environmental factors The segment value on and Clusters that can be covered The number of pixels in , It is a cluster The total number of all pixels contained in . After the calculation is completed, the minimum limiting factor method is used to select The minimum similarity is and The comprehensive similarity of geographical environment between 4) Traversal Repeat steps 2 and 3 for each environmental factor combination cluster in the The combination clusters of environmental factors are Similarity sequence ; 5) Calculate parent material conditions The semantic information corresponding to each combination of environmental factors is calculated as follows:

[0028]

[0029]

[0030] in, Combination of environmental factors The corresponding soil semantic information in the soil-environment text information framework, is the number of environmental factor combinations to determine the semantic information, is the number of semantic labels available for selection. During the semantic calculation process, the greater the similarity value, the higher the priority. The semantically determined environmental factor combination clusters and the selected semantic information labels do not participate in the semantic information calculation of other environmental factor combination clusters. 6) Traverse all parent material conditions , repeat steps 2-5 to obtain the soil semantic information corresponding to all environmental factor combination clusters (i.e., soil semantic information).

[0031] In step S104, soil type sample points are screened based on the combination of environmental factors: Based on the frequency sampling method, representative point sets corresponding to each soil type were selected from each environmental factor combination cluster. The main steps are as follows: First, a frequency histogram is created for each environmental factor in the environmental factor combination cluster; then, the pixel points whose environmental factor values fall into the highest frequency interval are determined as the representative sample points corresponding to the environmental factor. If a certain environmental factor has two or more groups with the same and the largest number of pixel points, then these sample points are all used as representative sample points; finally, the representative sample points corresponding to each environmental factor are combined to form a representative sample set corresponding to each soil type.

[0032] More specifically, for a given set of representative points , the number of sample points is , the environmental factors corresponding to each point The factor value of is: , then the continuous environmental factor The kernel density estimate of is as follows:

[0033] in, Soil type and environmental factors The probability density function of the relationship between Environmental factors bandwidth, Environmental factors The environmental factor value at location i, n is the number of representative sample points, is the kernel density function. The common Gaussian kernel function is used to estimate the kernel density curve, and the environmental factors are calculated based on the "rule of thumb" rule. Bandwidth , the Gaussian kernel function and environmental factor bandwidth expressions are as follows:

[0034]

[0035] in, Environmental factors The standard deviation of over n sample points.

[0036] By normalizing the obtained soil-environment relationship probability density function, we can obtain the similarity between each environmental factor and the typical environmental factor value corresponding to a certain soil information under the environmental factor. The value range is 0 to 1, and the calculation formula is as follows:

[0037] in, It is an environmental factor The probability density function between and soil type, Environmental factors The corresponding maximum value of the probability density function is It is an environmental factor The membership function between the soil type and the environmental factor The similarity between the typical environmental factor value corresponding to a soil type under this environmental factor.

[0038] For nominal or categorical variables, if the environmental factor value at a point is the same as the optimal environmental factor value corresponding to the soil type at that point, then the optimal environmental factor value for that point is 1. Otherwise, the optimal environmental factor value is 0. The environmental factor value corresponding to the soil type (such as parent material information) can be obtained from the description of the soil type or from the corresponding typical profile information.

[0039] In step S105, the soil type polygons are split: First, the soil type map was spatially associated with the environmental factor geographic database based on geographic coordinates; then, the soil type map was split into multiple soil type polygons according to the geographical boundaries of each soil type polygon.

[0040] In step S106, soil type sample points are screened based on the soil type polygon: Within each soil type polygon, representative points are screened from the environmental factor data corresponding to the soil type based on the frequency sampling method, and the representative sample points corresponding to each factor are merged as the sample point set corresponding to the soil type polygon.

[0041] The representative sample points of each soil type extracted based on the combination of environmental factors and soil type polygons are merged, and the geographical location is deduplicated to generate a non-redundant soil type sample point set, providing a high-quality data foundation for subsequent modeling.

[0042] In step S107, soil prediction mapping: A soil prediction mapping model is constructed based on the soil type sample point set, and the model is used to infer the spatial distribution of soil information to generate a high-precision soil type distribution map, thereby realizing the dynamic update of the soil type map and improving the mapping accuracy.

[0043] More specifically, for soil type information, by calculating the comprehensive environmental similarity of all soil types corresponding to a specified point, the soil type information and soil type estimation uncertainty of the point can be obtained. Their calculation formulas are as follows:

[0044]

[0045]

[0046] in, is the similarity vector between point p and each soil type, is the comprehensive geographical environment similarity of point p, n is the number of soil types, is the soil type corresponding to point p, is a calculation method for the comprehensive soil type similarity vector, is the uncertainty of soil type estimation at point p. Here, the maximum limiting factor method is used to estimate Conduct synthesis.

[0047] The comprehensive geographical environment similarity between a designated point and a soil type is calculated by combining the similarity between the environmental factor values of a designated point and the typical environmental factor values corresponding to a soil type. The formula is as follows:

[0048] in, is the similarity between the kth environmental factor at position i and the typical environmental factor value of a certain soil type, m is the number of environmental factors, It represents the comprehensive geographical environment similarity at location i. The environmental similarity synthesis method used here is the minimum limiting factor method. It can also be synthesized using the average method, maximum limiting factor method, weighted method, etc.

[0049] In a specific embodiment, combining Figure 3 The flowchart shown in the figure takes the update of the soil type map of a certain district in a certain province and city as an example to illustrate the specific implementation process of this application: Extraction and Structuring of Soil-Environmental Text Information The text data for the study area is titled "Soils of a Certain Prefecture." This document summarizes the second soil survey conducted in a certain district of a certain province. It describes the district's administration, agricultural production, soil-forming conditions, soil formation processes, soil classification and distribution patterns, soil type evaluation, soil fertility, and soil utilization. The document contains 73 soil types, classified into 40 genera, 16 subgenera, 9 soil classes, and 5 soil orders. Each soil type includes both soil type and a typical profile. Through statistical analysis of the completeness of soil and environmental information contained in the soil text data, the target information to be extracted for the experiment included soil type, elevation, slope, aspect, accumulated temperature, annual average temperature, annual precipitation, frost-free period, geomorphic location, and parent material. During the extraction process, a combination of rule-based and BiLSTM-CRFs models was used. The extraction results were then incorporated into various soil-environmental text information models named after the soil type.

[0050] Partition clustering to obtain environmental factor combinations Through statistical analysis of extracted soil-environmental textual information, the parent material types within the study area and the corresponding number of soil types (number of sub-region clusters) were determined. Based on a geographic database of environmental factors, slope, plan curvature, profile curvature, and terrain moisture index were selected as clustering factors. The spatial resolution was 30 m, and the study area contained 1,363,547 pixels. After standardizing each factor preprocessing, the K-means++ algorithm was used to cluster each parent material sub-region to obtain a specified number of environmental factor combinations. Parent material sub-regions containing only one soil genus type were excluded from clustering; instead, the soil type information was directly assigned to the parent material sub-region when calculating semantic information.

[0051] Computation of semantic information based on environmental factor combination Firstly, the text information of each environmental factor in the soil-environment text information model instance was quantitatively processed according to the environmental factor type. Then, based on the characteristics of the study area and the statistical analysis of the soil-environment text information extraction results, the environmental factors involved in the calculation of the semantic information of the environmental factor combination were selected, including parent material, elevation, slope, aspect, etc. Finally, the comprehensive environmental similarity between each environmental factor combination cluster in each parent material partition and its corresponding soil-environment text information framework was calculated, and then the optimal soil type information corresponding to each environmental factor combination was determined.

[0052] Soil type sample screening based on environmental factor combination First, draw the frequency distribution histogram of each environmental factor in the environmental factor set corresponding to each soil type. The environmental factors selected in this process include parent material, slope, plane curvature, profile curvature, terrain moisture index, annual average temperature, and annual average precipitation. Then, the pixel points whose environmental factor values fall into the highest frequency interval are determined as the representative sample points corresponding to the environmental factor. If a certain environmental factor has the same and largest number of pixel points in two or more group intervals, then these sample points are all regarded as representative sample points. Finally, merge the representative sample points corresponding to each environmental factor to form a representative sample point set corresponding to each soil type and save it to csv file 1 (including latitude and longitude coordinate information).

[0053] Soil type polygon splitting The soil map data for the study area was derived from a soil type map produced during the Second Soil Survey in 1987 at a scale of 1:250,000. After scanning, georeferencing, and vectorization, the soil type map was first spatially linked to a geographic database of environmental factors based on geographic coordinates. The soil type map was then split into multiple soil type polygons based on the geographic boundaries of each soil type (soil genus) polygon.

[0054] Soil type sample point screening based on soil type polygons Within each soil type polygon, representative points are selected from the environmental factor data corresponding to the soil type based on the frequency sampling method, and the representative sample points corresponding to each factor are merged as the sample point set corresponding to the soil type polygon and saved in CSV file 2 (including latitude and longitude coordinate information).

[0055] Generate soil type sample set The representative sample points of each soil type extracted based on the combination of environmental factors and soil type polygons are merged and deduplicated by geographic location to generate a non-redundant soil type sample point set.

[0056] Specifically, the CSV file 1 containing representative sample point data and the CSV file 2 containing representative sample point data of soil types are merged, and duplicate points with the same longitude and latitude coordinates are removed to generate a non-redundant soil type sample point set.

[0057] Soil prediction mapping First, the soil type sample set was divided into a training sample set and a validation sample set in a ratio of 8:2. Then, a soil prediction mapping model based on comprehensive environmental similarity was constructed based on the training sample set. The model was used to predict the spatial distribution of soil types and generate a high-precision soil type distribution map. Finally, the accuracy of the mapping results was evaluated based on the validation sample set.

[0058] Based on the quantitative soil-environment relationship obtained, the soil type distribution information of the study area was inferred, and the soil type distribution information was obtained based on the subordinate relationship between soil types and soil classes in the study area. The mapping results of this application method were compared with those of the direct screening method and the traditional soil map. The results are shown in Table 1. Figures 4 to 11 As shown. Among them, Figure 4 This is the soil type map inferred based on the method proposed in this application. Figure 5 The estimated uncertainty map of the mapping results of the method proposed in this application, Figure 6 This is the soil type map inferred based on the direct screening method. Figure 7 This is a speculative uncertainty map of the direct screening method mapping results. Figure 8 This is a traditional soil type map. Figure 9 This is a traditional soil type map. Figure 10 This is the soil type map corresponding to the mapping result of the method proposed in this application. Figure 11 This is the soil type map corresponding to the mapping results of the direct screening method.

[0059] Experimental results show that: 1) Based on field validation samples, the direct screening method achieved an inferred mapping accuracy of 48.84%, while the traditional soil map achieved an accuracy of 51.16%. The proposed method achieved an inferred mapping accuracy of 69.77%, significantly improving the accuracy of the previous two methods. These results demonstrate that the proposed method can effectively capture soil-environment relationships from soil textual data and that textual data can serve as a single data source for quantifying soil-environment relationships. 2) Traditional soil maps for the study area mapped 19 soil genera (6 soil classes), while the inferred soil map based on the proposed method mapped 31 soil genera (8 soil classes), consistent with the number recorded in the textual data. This demonstrates that the proposed method not only effectively captures soil-environment relationship knowledge from textual data but also supplements soil type information that is often overlooked during traditional soil mapping due to cartographic synthesis and other factors. The reason is that textual materials can record soil survey information in detail and are not affected by factors such as traditional cartographic synthesis. 3) By comparing the soil map inferred by the present application method with the traditional soil map, it was found that the overall spatial distribution pattern of soils between the two is basically the same, especially in the central and northern parts of the study area. Because traditional soil maps express soil type information in soil type polygons, while the present application method uses pixels, the soil map generated by the present application method is better able to express the subtle changes in soil type information in space and the spatial gradient of soil than the traditional soil map. It can also present soil information in local small areas that are ignored in the traditional cartographic synthesis process. 4) The mapping results of the direct screening method are significantly different from those of the present application method and the traditional soil map in the central and northern parts of the study area. The reason for this result may be that the completeness of the soil-environment text information is limited, resulting in the limited representativeness of the sample set screened by the direct screening method for the target soil type. Therefore, the accuracy of the soil-environment relationship obtained based on these sample points is limited, which leads to the limited accuracy of the inferred soil type map, which shows that the method of this application can effectively deal with the problem of incomplete text information; 5) The inference uncertainty of the method of this application in areas with slightly flat terrain is generally higher than that in areas with undulating terrain, which shows that the environmental factors selected for inferred mapping in this application can better reflect the changes in soil types in areas with larger undulating terrain.

[0060] Table 1 Comparison of soil mapping results accuracy

[0061] This soil type map updating method, which integrates soil text and soil map data, firstly constructs a soil-environment text information model that parses soil text and spatially interprets soil maps. This model converts unstructured soil text and soil map vector polygons into a unified point dataset, thereby enabling multi-source data fusion modeling for soil predictive mapping. Based on the soil-environment text information model, the number of partition clusters is determined. Partition cluster analysis is then used to identify clusters corresponding to environmental factor combinations within each parent material partition. Based on the environmental factor information of the environmental factor combination clusters and the quantitative environmental factor information in the soil-environment text information model, the similarity between the environmental factor combination and the text information framework is calculated to determine the semantic information of soil types for each environmental factor combination cluster. Representative point sets corresponding to each soil type are selected from each environmental combination cluster, and kernel density estimation is used to quantify soil-environment relationships. Furthermore, this method overcomes core issues inherent in traditional techniques, such as the failure of spatial-text data alignment, barriers to multi-granular information fusion, and the lack of semantic parsing technology. This method significantly enhances the value of historical data and provides a reliable data foundation and technical support for updating soil type maps and improving mapping accuracy.

[0062] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0063] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. A soil type map updating method integrating soil chronicle text and soil map data, characterized in that: The method includes: Determine soil types and their corresponding environmental factors based on soil annals text to build a soil-environment text information model, and use domain knowledge base and corpus to fill in the soil-environment text information model; The study area was divided into several different parent material zones using parent material factors. The number of zone clusters was determined by combining the soil-environment text information model. The environmental factor data of each parent material zone in the environmental factor geographic database were clustered using a clustering algorithm to generate environmental factor combination clusters. Based on the environmental factor type, the environmental factor text information of each instance in the soil-environment text information model is quantitatively processed. Combined with the environmental factor combination clusters, the comprehensive environmental similarity between each environmental factor combination cluster and the instance is calculated, and the optimal soil semantic information corresponding to each environmental factor combination is determined. Based on the optimal soil semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type. After spatially associating the soil type map with the environmental factor geographic database, the map was split into multiple soil type polygons according to geographic boundaries. Within each soil type polygon, representative sample points were selected using the frequency sampling method to generate a sample point set for each polygon. Merge the representative sample point set of each soil type and the sample point set of each polygon to generate the soil type sample point set; Soil prediction mapping models are trained and tested on soil type sample sets to generate and update soil type maps.

2. The soil type map updating method according to claim 1, wherein: The steps of determining soil types and their corresponding environmental factors based on soil annals text to construct a soil-environment text information model and filling the soil-environment text information model with a domain knowledge base and corpus include: Based on the soil annals text, determine the soil type and its corresponding environmental factor information; Analyze the language description characteristics of soil types and their corresponding environmental factors in soil chronicle texts, and construct a soil-environment text information model; Build the domain knowledge base and corpus required to extract soil-environment text information; When the expression pattern of text information is single and regular, a rule-based approach is used to extract soil-environment text information from the domain knowledge base and corpus; When the expression pattern of text information is complex, changeable and irregular, a method based on the BiLSTM-CRF model is used to extract soil-environment text information from the domain knowledge base and corpus.

3. The soil type map updating method according to claim 2, wherein the method comprises: The study area was divided into multiple parent material zones using parent material factors. The number of zone clusters was determined by combining the soil-environment text information model. The environmental factor data of each parent material zone in the environmental factor geographic database were clustered using a clustering algorithm. The steps of generating environmental factor combination clusters included: The study area is divided into several different parent material zones using parent material factors; Conduct statistical analysis on the extracted soil-environment text information to determine the parent material type of each parent material partition and the corresponding partition cluster number; Based on the environmental factor geographic database, select environmental factor data and perform standardized preprocessing on each environmental factor data; Based on the preprocessed environmental factor data, the K-means++ algorithm was used to cluster each parent material partition to generate environmental factor combination clusters.

4. The soil type map updating method according to claim 3, wherein: The steps of quantitatively processing the environmental factor text information of each instance in the soil-environment text information model based on the environmental factor type, combining each environmental factor combination cluster, calculating the comprehensive environmental similarity between each environmental factor combination cluster and the instance, and determining the optimal soil semantic information corresponding to each environmental factor combination include: Based on the environmental factor type, the environmental factor text information of each instance in the soil-environment text information model is quantitatively processed; For each group of environmental factor combination clusters in each parent material partition, the comprehensive environmental similarity between the soil-environmental text information model and the environmental factor combination cluster is calculated based on the corresponding environmental factor information and the environmental factor text information in the soil-environmental text information model; The soil-environment text information model that best matches the environmental factor combination cluster is determined based on the comprehensive environmental similarity, and the soil type information in the model is passed to the environmental factor combination cluster to determine the optimal soil semantic information corresponding to each environmental factor combination.

5. The soil type map updating method according to claim 4, wherein: Based on the optimal soil semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type, including the following steps: Based on the optimal soil semantic information, a frequency histogram is created for each environmental factor in the environmental factor combination cluster; The pixel points whose environmental factor values fall into the highest frequency interval are determined as representative sample points corresponding to the environmental factor; if there are multiple pixels with the same highest frequency interval, all of them are selected; The representative sample points corresponding to each environmental factor were merged to form a representative sample point set corresponding to each soil type.

6. The soil type map updating method according to claim 5, characterized in that: The step of merging the representative sample point set of each soil type and the sample point set of each polygon to generate the soil type sample point set includes: The representative sample point set of each soil type is merged with the sample point set of each polygon extracted from the soil type polygon, and duplicates are removed by geographic location to generate a non-redundant soil type sample point set.

7. The soil type map updating method according to claim 6, characterized in that: The steps of training and testing a soil prediction mapping model based on a soil type sample set to generate and update a soil type map include: The extreme gradient boosting algorithm was used to construct a soil prediction mapping model; The soil type sample point set is divided into a training sample point set and a verification sample point set according to a preset ratio; The soil prediction mapping model is trained and tested using the training sample set and the validation sample set to obtain a trained soil prediction mapping model; The trained soil prediction mapping model is used to predict the spatial distribution of soil types and generate and update soil type maps.