Sample point generation method based on typical soil type
By constructing a soil-environmental text information model and clustering analysis, a sample set of typical soil types was generated, which solved the problem of insufficient information accuracy and sample number in soil survey data, and achieved efficient soil information acquisition and mapping.
Patent Information
- Application Number
- CN202510535380.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-19
AI Technical Summary
The soil type distribution information and typical sample point location information in the existing soil survey data are relatively accurate, and the number of sample points is limited, making it difficult to meet the quantitative analysis needs of soil-environmental relationships.
The soil-environmental text information model is constructed based on soil journal text, the domain knowledge base and corpus fill model is used, and the parent material factor is divided and partitioned and clustered analysis is performed. The similarity is calculated by combining environmental factor clusters, representative sample points are screened, and a set of sample points are generated in typical soil type.
It breaks through the bottleneck of the use of text materials, obtains a large number of soil type sample points, reduces the cost of mapping, improves the spatial speculation efficiency of soil information, and has important theoretical and practical application value.
Smart Images

Figure CN120508644A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the technical field of soil type sample point generation, and in particular to a method for generating sample points based on typical soil types. Background Art
[0002] Soil is the material foundation for human survival and an indispensable, non-renewable natural resource. Accurate understanding of soil spatial distribution is essential for the rational utilization of soil resources. This information serves as fundamental data for precision agriculture, environmental change simulation, natural resource management and utilization, and global change monitoring, and holds significant scientific and application value. Predictive soil mapping is the primary method for inferring soil spatial distribution information. The soil-environment relationship is a core issue in this process, and soil samples provide the data foundation for capturing this relationship.
[0003] Due to the needs of scientific research and practical production, a vast amount of soil survey data, such as soil survey reports and historical documents, has been accumulated in the field of soil science. In particular, during my country's second soil census, over 2,000 volumes of soil species records and other textual materials were accumulated. These textual materials contain a wealth of reliable soil survey data and pedogenesis information, embodying the understanding of numerous soil experts on various soil characteristics and their changing patterns, and contain a wealth of information on soil-environment relationships. They not only provide solid data support for soil mapping but also lay a solid foundation for soil science research and the sustainable development of agricultural production. However, the accuracy of soil type distribution information and the location of typical sampling points recorded in these soil species records is low, and the number of sampling points is extremely limited, making them inadequate for quantitative analysis of soil-environment relationships. Consequently, this type of data has not been fully explored and utilized.
[0004] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.
[0005] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the Invention
[0006] The purpose of the embodiments of the present disclosure is to provide a method for generating sample points based on typical soil types, thereby overcoming one or more problems caused by the limitations and defects of related technologies, at least to a certain extent.
[0007] According to an embodiment of the present disclosure, a method for generating sample points based on typical soil types is provided, the method comprising: Determine soil types and their corresponding environmental factors based on soil annals text to build a soil-environment text information model, and use domain knowledge base and corpus to fill in the soil-environment text information model; The study area was divided into several different parent material zones using parent material factors. The number of zone clusters was determined by combining the soil-environment text information model. The environmental factor data of each parent material zone in the environmental factor geographic database were clustered using a clustering algorithm to generate environmental factor combination clusters. Converting the environmental factor text information of the instance in the soil-environment text information model into numerical data; For each environmental factor combination cluster, calculate its comprehensive environmental similarity with the soil-environment text information model to determine the matching soil type semantic information; Based on the matched soil type semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type. Representative sample points of all soil types were merged to generate the final set of representative soil type sample points.
[0008] Furthermore, the steps of determining soil types and corresponding environmental factor information based on soil annals text to construct a soil-environment text information model, and using a domain knowledge base and corpus to fill the soil-environment text information model include: Based on the soil annals text, determine the soil type and its corresponding environmental factor information; Analyze the language description characteristics of soil types and their corresponding environmental factors in soil chronicle texts, and construct a soil-environment text information model; Build the domain knowledge base and corpus required to extract soil-environment text information; When the expression pattern of text information is single and regular, a rule-based approach is used to extract soil-environment text information from the domain knowledge base and corpus; When the expression pattern of text information is complex, changeable and irregular, a method based on the BiLSTM-CRF model is used to extract soil-environment text information from the domain knowledge base and corpus.
[0009] Furthermore, the study area is divided into multiple parent material partitions using parent material factors. The number of partition clusters is determined by combining the soil-environment text information model. The environmental factor data of each parent material partition in the environmental factor geographic database are clustered using a clustering algorithm to generate environmental factor combination clusters, including the following steps: The study area is divided into several different parent material zones using parent material factors; Conduct statistical analysis on the extracted soil-environment text information to determine the parent material type of each parent material partition and the corresponding partition cluster number; Based on the environmental factor geographic database, select environmental factor data and perform standardized preprocessing on each environmental factor data; Based on the preprocessed environmental factor data, the K-means++ algorithm was used to cluster each parent material partition to generate environmental factor combination clusters.
[0010] Furthermore, the text-based environmental factor values in the soil-environment text information model are converted into numerical data, including direct conversion of single-point values, conversion of segmented values into a unified format, and quantification of qualitative descriptions based on field classification standards, including: When the environmental factor text information is a single point value; if it is a numerical single point value, no conversion is required; if it is a text single point value, the text is directly converted to a number; When the text information of environmental factors is a segmented value; if it is a numerical segmented value, the prefix, suffix and connector of the numerical body are converted into a unified expression format and then parsed using rules; if it is a textual segmented value, quantitative standards are set for the qualitative description of each environmental factor in combination with the existing field classification standards and the actual situation of the study area.
[0011] Furthermore, for each environmental factor combination cluster, calculating its comprehensive environmental similarity with the soil-environment text information model and determining the matching soil type semantic information includes: Set the parent material information set to , the quantified soil-environment text information model is ; According to parent material partition All environmental factors to be determined for soil semantic information are combined into clusters , the soil-environment text information model set corresponding to the parent material condition is ; Combining clusters for environmental factors Any combination of environmental factors in Calculate the cluster center and the range of each environmental factor, and take the environmental factor combination cluster The mode of the environmental factor values at each point is taken as the central value, and the environmental factor combination clusters The value range of each environmental factor in the current cluster is used as the value range of each environmental factor; Traversing the parent material partitions Corresponding soil-environment text information model set , combining clusters based on environmental factors The cluster center and the range of each environmental factor are used to calculate the environmental factor combination cluster Soil-Environment Text Information Model Collection The similarity of each frame in the , and the combination cluster of environmental factors are obtained Similarity sequence ;in, Represents a cluster of environmental factor combinations Soil-Environment Text Information Model The comprehensive environmental similarity between them is measured using the Gower similarity coefficient:
[0012] Where, is the calculation function of comprehensive environment similarity, for and In the environmental factor set Gower similarity coefficient sequence on , is the number of environmental factors; The calculation formula is as follows:
[0013] Where, and Environmental factors Corresponding cluster Central value and framework The text information value in Cluster Medium environmental factors The value range of and Environmental factors The segment value on and Clusters that can be covered The number of pixels in , It is a cluster The total number of all pixels contained in After the calculation is completed, the minimum limiting factor method is used to select The minimum similarity is and The comprehensive similarity of geographical environment between Traversing environmental factor combination clusters For each environmental factor combination cluster in the dataset, the geographical environment comprehensive similarity is repeatedly calculated to obtain the parent material conditions. The combination clusters of environmental factors are Similarity sequence ; Calculation of parent material conditions The semantic information corresponding to each combination of environmental factors is as follows:
[0014]
[0015]
[0016] Where, Combination of environmental factors The corresponding soil semantic information in the soil-environment text information model, is the number of environmental factor combinations to determine the semantic information, is the number of semantic tags available for selection; Traverse all parent material conditions , obtain the soil semantic information corresponding to all environmental factor combination clusters, and determine the optimal soil semantic information corresponding to each environmental factor combination.
[0017] Furthermore, based on the matched soil type semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type, including: Based on the matched soil type semantic information, a frequency histogram is created for each environmental factor in the environmental factor combination cluster; The pixel points whose environmental factor values fall into the highest frequency interval are determined as representative sample points corresponding to the environmental factor; if there are multiple pixels with the same highest frequency interval, all of them are selected; The representative sample points corresponding to each environmental factor were merged to form a representative sample point set corresponding to each soil type.
[0018] Furthermore, the steps of merging representative sample points of all soil types to generate a final set of typical soil type sample points include: The representative sample points of all soil types were merged and geographical location was removed to generate the final set of typical soil type sample points.
[0019] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: In the embodiments of the present disclosure, the method for generating typical soil type samples is described above. On the one hand, a soil-environment text information model is designed based on soil species text data to organize and store soil types and their corresponding environmental factor text information extracted from soil species text. Based on the geographical similarity theory, it is assumed that there is a corresponding relationship between the distribution of soil type information in geographic space and the combination of environmental factor sets in attribute space. The problem of obtaining soil type samples based on text data is transformed into the problem of obtaining specific environmental factor combinations and their soil type semantic information. The soil-environment text information is used to guide the clustering of environmental factor data, thereby obtaining environmental factor combinations and their soil type information. Based on the environmental factor information, representative soil type samples are selected to generate a typical sample set. On the other hand, using the common text data soil species as the data source for obtaining soil type samples, a large number of soil type samples can be obtained without relying on soil experts. This breaks through the current bottleneck of making full and effective use of text data, helps reduce mapping costs, and improves the efficiency of soil information spatial inference. It has important theoretical significance and practical application value, and has broad application prospects in digital soil mapping and land resources surveys. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0021] Figure 1 A diagram showing the steps of a method for generating sample points based on typical soil types in an exemplary embodiment of the present disclosure is shown; Figure 2 A schematic diagram showing the structure of a soil-environment text information model in an exemplary embodiment of the present disclosure is shown; Figure 3 The following is an overall flow chart of soil-environment text information extraction in an exemplary embodiment of the present disclosure; Figure 4 A flowchart of rule-based soil-environment text information extraction in an exemplary embodiment of the present disclosure is shown; Figure 5 A flowchart of soil-environment text information extraction based on a BiLSTM-CRF model in an exemplary embodiment of the present disclosure is shown; Figure 6 A specific flow chart of a method for generating sample points based on typical soil types in an exemplary embodiment of the present disclosure is shown; Figure 7 An example of the structured organization of soil-environment text information in an exemplary embodiment of the present disclosure is shown; Figure 8 The following figure shows the environmental factor combination acquisition result in the exemplary embodiment of the present disclosure; Figure 9 A graph showing the accuracy change of partition clustering and semantic calculation results in an exemplary embodiment of the present disclosure is shown; Figure 10 A histogram showing the accuracy distribution of partition clustering and semantic calculation results in an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0023] In addition, the accompanying drawings are merely schematic illustrations of embodiments of the present disclosure and are not necessarily drawn to scale. Like reference numerals in the figures represent like or similar parts, and thus repeated descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically separate entities.
[0024] This example implementation provides a method for generating sample points based on typical soil types. Figure 1 As shown in , the method for generating sample points based on typical soil types may include: Step S101: determining soil types and their corresponding environmental factors based on soil annals text to construct a soil-environment text information model, and filling the soil-environment text information model with domain knowledge base and corpus; Step S102: Divide the study area into multiple parent material partitions using parent material factors, determine the number of partition clusters in combination with the soil-environment text information model, and use a clustering algorithm to perform cluster analysis on the environmental factor data of each parent material partition in the environmental factor geographic database to generate environmental factor combination clusters; Step S103: converting the environmental factor text information of the instance in the soil-environment text information model into numerical data; Step S104: for each environmental factor combination cluster, calculating its comprehensive environmental similarity with the soil-environment text information model, and determining the matching soil type semantic information; Step S105: Based on the matched soil type semantic information, a frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type; Step S106: Merge representative sample points of all soil types to generate a final typical soil type sample point set.
[0025] The above-mentioned method for generating typical soil type samples, on the one hand, designs a soil-environment text information model based on soil species text data to organize and store soil types and their corresponding environmental factor text information extracted from soil species text. Based on geographic similarity theory, assuming that the distribution of soil type information in geographic space corresponds to the combination of environmental factor sets in attribute space, the problem of obtaining soil type samples from text data is transformed into the problem of obtaining semantic information about specific environmental factor combinations and their soil types. Soil-environment text information is used to guide the clustering of environmental factor data, thereby obtaining environmental factor combinations and their soil type information. Representative soil type samples are selected based on the environmental factor information to generate a typical sample set. Furthermore, by using common textual soil species text data as a data source for obtaining soil type samples, a large number of soil type samples can be obtained without relying on soil experts. This overcomes the current bottleneck of fully and effectively utilizing text data, helps reduce mapping costs, and improves the efficiency of spatial inference of soil information. This method has important theoretical significance and practical application value, and has broad application prospects in digital soil mapping and land resources surveys.
[0026] Below, we will refer to Figures 1 to 10 Each step of the above-mentioned method for generating sample points based on typical soil types in this exemplary embodiment is described in more detail.
[0027] In step S101, soil types and their corresponding environmental factor information are determined based on soil log text to construct a soil-environment text information model, and the soil-environment text information model is filled using a domain knowledge base and a corpus.
[0028] Specifically, based on the soil annals text, the soil type and its corresponding environmental factor information are determined; the linguistic description characteristics of the soil type and its corresponding environmental factor information in the soil annals text are analyzed, and a soil-environment text information model is constructed; the domain knowledge base and corpus required for extracting soil-environment text information are constructed; when the expression pattern of the text information is single and regular, a rule-based method is used to extract the soil-environment text information from the domain knowledge base and corpus; when the expression pattern of the text information is complex, changeable and irregular, a method based on the BiLSTM-CRF model is used to extract the soil-environment text information from the domain knowledge base and corpus.
[0029] More specifically, Figure 2As shown in the figure, the structure of the soil-environment text information model is as follows: each soil-environment text information framework consists of a soil type description and a typical profile description. The former is a description of the overall information of a soil type, and the latter is a description of the typical profile information of a soil type. One soil type corresponds to one or more typical profiles. The soil type description and the typical profile description are both composed of target information to be extracted, such as soil type information, soil attribute information, and environmental factor information. Each target information to be extracted consists of a variable name and a variable value. For example, the variable name and variable value corresponding to the elevation information "elevation 100 meters" are "elevation" and "100 meters" respectively; the variable name and variable value corresponding to the soil type information "upper scorched brown red soil" are "soil type" and "upper scorched brown red soil" respectively; the variable name and variable value corresponding to the soil organic matter content information "organic matter 2.24%" are "organic matter content" and "2.24%" respectively. The above soil-environment text information framework structure can be flexibly adjusted according to the description characteristics of the soil and environmental factor information in each text material to meet the needs of structured organization and expression of soil-environmental factor information. Among them, Figure 2 (a) is the structure of the soil-environment text information model. Figure 2 (b) is an example of soil-environment text information model.
[0030] Build the domain knowledge base and corpus required for extracting text information, such as soil information dictionary, environmental factor dictionary, keyword dictionary, rule base, annotated corpus, etc.
[0031] The overall process of extracting soil-environment text information from text materials is as follows: Figure 3As shown in the figure, it mainly includes three steps: first, determine the target information to be extracted; then, split the text materials to be extracted into different texts according to soil type or chapter structure, and each text corresponds to a soil-environment text information framework. For each given text, an appropriate extraction method is selected based on the descriptive characteristics of the target information. For variables with simple, regular expression patterns (e.g., "altitude 300m," "slope 25°," "accumulated temperature 4612°C," and "A layer thickness 10cm" represent four different types of target information, but their variable values can be expressed using the same "numeral + quantifier" pattern), a rule-based method is used for extraction. For variables with complex, irregular expression patterns (e.g., "slope deposits of limestone weathering," "limestone weathering deposits," and "Quaternary red clay," which all represent parent material information but are difficult to represent using a unified pattern), a statistical method is used for extraction. Finally, after both the rule-based and statistical methods have extracted all the target information to be extracted from a given text, the target information extracted by each method is combined and aggregated into a single soil-environment text information framework in the "variable name:variable value" format. The framework name is generally directly extracted based on the text structure or soil type dictionary.
[0032] Rule-based method: In information extraction, sentences are usually used as the basic text analysis granularity, and according to the co-occurrence characteristics of the variable names and variable values of the target information to be extracted, for target information with limited expression patterns, its variable names and variable values usually appear together in the same sentence. Therefore, before performing information extraction, the text to be extracted needs to be preprocessed according to punctuation marks to generate a set of sentences. This process is called sentence segmentation in the field of natural language processing. After obtaining the sentence set of the text material, a specific information extraction method can be used to perform the information extraction operation. When using a rule-based method to extract the target information to be extracted, it is mainly divided into two parts: (1) screening sentences containing the variable names of the target information, and (2) extracting the variable values of the target information. The detailed process of this method is as follows. Figure 4 shown.
[0033] Statistical-based method: Since the construction of a high-quality corpus requires extremely high time, manpower, and material costs, the statistical-based method in this article is used to supplement the extraction of some target information that is difficult to extract using the rule-based method within the determined target information framework.
[0034] The information extraction method based on the Bi-directional Long Short-Term Memory Conditional Random Fields (BiLSTM-CRF) model is a commonly used information extraction method based on deep learning and statistics. This method regards information extraction as a classification task. The process of extracting target information based on the BiLSTM-CRF model is as follows: Figure 5 The BiLSTM model is implemented in Python, while the CRF model can be constructed using the open-source software package CRF++ (http: / / taku910.github.io / crfpp). The label scores predicted by the BiLSTM model serve as the input to the CRF model. When training a CRF model using CRF++, you need to prepare a training corpus file and a feature template file.
[0035] In step S102, the parent material factors are used to divide the study area into multiple different parent material partitions, the number of partition clusters is determined in combination with the soil-environment text information model, and the clustering algorithm is used to perform cluster analysis on the environmental factor data of each parent material partition in the environmental factor geographic database to generate environmental factor combination clusters.
[0036] Specifically, the study area was divided into multiple parent material partitions using parent material factors. The extracted soil-environment text information was statistically analyzed to determine the parent material type of each parent material partition and its corresponding partition cluster number. Based on the environmental factor geographic database, environmental factor data were selected and standardized preprocessed for each environmental factor data. Based on the preprocessed environmental factor data, the K-means++ algorithm was used to cluster each parent material partition to generate environmental factor combination clusters.
[0037] More specifically, first, the environmental factors used for clustering are selected, and the data of each environmental factor is standardized and preprocessed. Then, based on the parent material factor information, the study area is divided into different parent material partitions, and the number of soil types contained in each parent material partition is estimated based on the soil-environment text information set. For example, if it is known that the parent materials corresponding to soil types A, B, C, D, E, and F in a study area are P1, P2, P1, P2, P3, and P2, respectively, then the corresponding relationship between each parent material and soil type in the study area is: P1: A, C, P2: B, D, F, P3: E. It can be inferred that the number of soil types corresponding to parent material partitions P1, P2, and P3 should be 2, 3, and 1, respectively, which is the number of parent material partition clusters. Finally, based on the environmental factor geographic database, K-means++ and other algorithms are used to obtain the environmental factor combinations within each parent material partition through cluster analysis.
[0038] The following conditions should be considered when selecting environmental factors: 1) they should be able to reflect the differences in soil types within the region, that is, factors that change synergistically with soil types should be selected; 2) they should be easily accessible and measurable to reduce acquisition costs and processing difficulties; 3) at different soil type classification levels, the main environmental factors affecting soil formation are different, and the data characteristics of each environmental factor are also different. Therefore, the environmental factors selected during clustering should be adapted to the classification level of the soil type.
[0039] After selecting the environmental factors used for clustering, the environmental factor data involved in clustering need to be preprocessed. The content of preprocessing mainly includes outlier processing, standardization or normalization processing. The processing of outliers is based on the exploratory analysis of factor data, understanding the data structure, statistical characteristics, spatial distribution characteristics, and visualization results of each factor, finding the outliers of each environmental factor, and then using the average value or weighted average value of the factor values in its adjacent areas to replace them, or directly setting them to null values and not participating in subsequent processing. Factor standardization is to scale the value of each environmental factor to a range of mean 0 and standard deviation 1, but will not change the original distribution structure of the factor data. The purpose of standardization is to eliminate or reduce the deviation caused by different dimensions of different environmental factors and reduce the impact on clustering results. The Z-score standardization formula that can be used is as follows:
[0040] in, is the original environmental factor value, is the standardized value, μ and δ are the mean and standard deviation of the environmental factors, respectively.
[0041] Factor normalization converts environmental factor values from dimensional form to dimensionless form and scales the factor values to a specified interval (such as [0,1]). This does not change the numerical order of the original data and can simplify calculations. The following normalization formulas can be used:
[0042] in, and Environmental factors The maximum and minimum factor values.
[0043] For a given A sample set consisting of , It's a bit Corresponding Environmental factor values, K-means clustering is to divide the sample set into Clusters , the goal is to minimize the sum of squares of cluster errors:
[0044] in, Cluster The cluster center of The smaller it is, the higher the similarity between samples in the cluster.
[0045] The K-means algorithm uses a random method to initialize cluster centers, which makes the clustering results unstable and takes a long time to reach a stable state. To improve clustering efficiency and effectiveness, the K-means++ algorithm improves the selection of initial cluster centers based on the K-means algorithm. The idea of the K-means++ algorithm in selecting initial centroids is to ensure that the distance between the initial cluster centers is as far as possible. The specific steps of the algorithm are as follows: 1) Randomly select a sample point as the initial cluster center ; 2) Calculate each sample point and The shortest distance , then calculate the point The probability of being selected as the next cluster center: ,Finally, the next cluster center is selected according to the roulette wheel method; 3) Repeat step 2 until you select Cluster centers ; 4) For each sample point in the data set , calculate it to The distance between the cluster centers is calculated and assigned to the class corresponding to the cluster center with the smallest distance; 5) For each category , recalculate the centroid (cluster center) of all sample points of this class; 6) Repeat steps 4 and 5 until the location of the cluster center no longer changes.
[0046] In step S103, the textual environmental factor values in the soil-environment text information model are converted into numerical data, including direct conversion of single-point values, unified format conversion of segmented values, and quantification of qualitative descriptions based on field classification standards.
[0047] Specifically, when the text information of the environmental factor is a single-point value; if it is a numerical single-point value, no conversion is required; if it is a text-type single-point value, the text is directly converted into a number; when the text information of the environmental factor is a segmented value; if it is a numerical segmented value, the prefix, suffix and connector of the numerical body are converted into a unified expression format, and then parsed using rules; if it is a text-type segmented value, quantitative standards are set for the qualitative description of each environmental factor in combination with the existing field classification standards and the actual situation of the study area.
[0048] More specifically, in order to use soil-environment text information for mathematical calculations, the environmental factor text information in the soil-environment text information model instance needs to be quantified before semantic calculation. The various types of information are processed as follows: 1) Single point value: ① Numeric type, for example, "Altitude: 80m", no conversion is required and can be used directly; ② Text type, for example, "slope: 20 degrees", the text needs to be directly converted into numbers; 2) Segment value: ① Numeric type, for example, "Altitude: >35m, 30 to 50m, less than 100m". The main body of the text environmental factor value is a numeric value. During conversion, the prefix, suffix, and connector of the numeric body are converted into a unified expression format, and then parsed using rules. For example, "35 to 50m, 35~50m, 35-50 meters" are all uniformly converted to "35-50m"; ② Text type, for example, "steep slope, southwest slope, shady slope". This type of text needs to combine existing field classification standards and the actual situation of the study area to set quantitative standards for the qualitative description of each environmental factor. Taking the quantification of slope direction and slope as an example, as shown in Tables 1 and 2.
[0049] Table 1 Slope aspect quantitative comparison table
[0050] Table 2 Slope Quantification Comparison Table
[0051] In step S104, for each environmental factor combination cluster, the comprehensive environmental similarity between the cluster and the soil-environment text information model is calculated to determine the matching soil type semantic information.
[0052] Specifically, for each set of environmental factor combinations (clusters) within each parent material partition, the comprehensive environmental similarity between the soil-environmental text information model and the environmental factor combination cluster is calculated based on the environmental factor information of each combination and the environmental factor information in the soil-environmental text information model, the soil-environmental text information model that best matches the environmental factor combination cluster is determined, and the soil type information in the model is passed to the environmental factor combination cluster, thus completing the semantic information calculation of the environmental factor combination. After semantic calculation, the environmental factor combination cluster has both location information and soil type information. Given a parent material information set , Quantified soil-environment text information framework set , the detailed process of the algorithm is as follows: 1) According to the parent material partition information, parent material partition The combined clusters of all environmental factors to be determined for soil semantic information are , the soil-environment text information framework set corresponding to the parent material condition is ; 2) For collections Any combination of environmental factors in , calculate the cluster center and the range of each environmental factor, and take The mode of the environmental factor values at each point is taken as the central value. The value range of each environmental factor in the current cluster is used as the value range of each environmental factor; 3) Traversal Corresponding soil-environment text frame set ,according to Calculation of cluster centers and ranges of environmental factors and The similarity of each frame in Similarity sequence . Represents a cluster of environmental factor combinations Frame for text with soil - environment The comprehensive environmental similarity between the two is measured using the Gower similarity coefficient, which is calculated as follows:
[0053] in, is the calculation function of comprehensive environment similarity, for and In the environmental factor set Gower similarity coefficient sequence on , is the number of environmental factors. The calculation formula is as follows:
[0054] in, and Environmental factors Corresponding cluster Central value and framework The text information value in Cluster Medium environmental factors The value range of and Environmental factors The segment value on and Clusters that can be covered The number of pixels in , It is a cluster The total number of all pixels contained in . After the calculation is completed, the minimum limiting factor method is used to select The minimum similarity is and The comprehensive similarity of geographical environment between 4) Traversal Repeat steps 2 and 3 for each environmental factor combination cluster in the The combination clusters of environmental factors are Similarity sequence ; 5) Calculate parent material conditions The semantic information corresponding to each combination of environmental factors is calculated as follows:
[0055]
[0056]
[0057] in, Combination of environmental factors The corresponding soil semantic information in the soil-environment text information framework, is the number of environmental factor combinations to determine the semantic information, is the number of semantic labels available for selection. During the semantic calculation process, the greater the similarity value, the higher the priority. The semantically determined environmental factor combination clusters and the selected semantic information labels do not participate in the semantic information calculation of other environmental factor combination clusters. 6) Traverse all parent material conditions , repeat steps 2-5 to obtain the soil semantic information corresponding to all environmental factor combination clusters.
[0058] In step S105, based on the matched soil type semantic information, a frequency sampling method is used to screen representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type.
[0059] Specifically, first, create a frequency histogram for each environmental factor in the environmental factor combination cluster. Since each environmental factor has a different type, range, and distribution pattern, it is necessary to set an appropriate histogram bin distance for each factor based on its actual situation. The bin distance setting method is as follows:
[0060] in, Environmental factors The histogram bin distance of , and Environmental factors The number of points covered and their interquartile range; Then, the pixel points whose environmental factor values fall into the highest frequency interval are determined as the representative positions corresponding to the environmental factors. If a certain environmental factor has the same and largest number of pixel points in two or more group intervals, then these sample points are all regarded as representative sample points.
[0061] In step S105 , representative sample points of all soil types are merged to generate a final typical soil type sample point set.
[0062] Specifically, the representative point sets corresponding to each soil type are merged and deduplicated based on geographical location to form the final typical soil type sample point set.
[0063] In a specific embodiment, the following Figure 6 The specific flow chart of the method for generating typical soil type samples is shown, taking the soil type data "Soil of a Certain State" of a certain province, city, and district as an example, to illustrate the specific implementation process of this application: (1) Extraction and structuring of soil-environmental text information "Soils of a Certain Prefecture" is a summary of the soil survey conducted in a certain prefecture between 1982 and 1986. It documents the prefecture's administration, agricultural production, soil-forming conditions, soil formation processes, soil classification and distribution patterns, soil type evaluation, soil fertility, and soil utilization. This data uses a genetic classification system to describe soil types, using soil species as the basic descriptive unit. A total of 73 soil species are described for the entire region, belonging to 40 genera, 16 subgenera, 9 soil classes, and 5 soil orders. Each soil species description is divided into two parts: a description of the soil type (hereinafter referred to as the "soil type description") and a description of a corresponding profile (hereinafter referred to as the "typical profile description"). By statistically analyzing the completeness of soil and environmental factor information contained in the soil text data, the target information selected for extraction in the experiment included soil type, elevation, slope, aspect, accumulated temperature, average annual temperature, annual precipitation, frost-free period, geomorphic location, and parent material. In the extraction process, we combined the rule-based and BiLSTM-CRF model-based methods, and filled the extraction results into the soil-environment text information models named after the soil type. Figure 7 As shown in the figure, (a) is the rule-based extraction result; (b) is the BiLSTM-CRF model-based extraction result; and (c) is the extraction result that integrates different target information. It can be seen that 1) the rule-based method is mainly used to extract target information with a single, regular expression pattern, such as elevation and slope; 2) the BiLSTM-CRF model-based method is mainly used to extract target information with complex, variable, and irregular expression patterns, such as landform location and parent material information; 3) the soil-environment text information framework can effectively organize and structure the extraction results of different target information.
[0064] (2) Partition clustering to obtain environmental factor combinations According to the soil-environment text information extracted from "Soil of a Certain State", statistics show the types of parent materials contained in the study area and the number of corresponding soil types (number of partition clusters). Based on the environmental factor geographic database, slope, plane curvature, profile curvature, and terrain moisture index were selected as clustering factors with a spatial resolution of 30m. There are a total of 1,363,547 pixels in the study area. On the basis of standardized preprocessing of each factor, the K-means++ algorithm is used to cluster each parent material partition to obtain a specified number of environmental factor combinations. For parent material partitions with only one soil type, they do not participate in clustering. When calculating semantic information, the soil type information is directly assigned to the parent material partition. Figure 8 As shown in Figure 1, the results are obtained by combining environmental factors; (a) is the result before semantic calculation, and (b) is the result after semantic calculation.
[0065] (3) Quantification of text information on environmental factors According to the type of environmental factors, the environmental factor text information in the soil-environment text information model instance is quantitatively processed.
[0066] (4) Calculation of semantic information of environmental factor combination Through statistics and analysis of the soil-environment text information extraction results, the types of environmental factor text information selected for the calculation of environmental factor combination semantic information include soil type, parent material, elevation, slope, and slope direction. On this basis, the comprehensive environmental similarity between each environmental factor combination cluster in each parent material partition and its corresponding soil-environment text information framework is calculated to determine the soil type information corresponding to each environmental factor combination. Figure 9 As shown in , it is a graph showing the change in accuracy of partition clustering and semantic calculation results. Figure 10 The figure shows the distribution histogram of the accuracy of partition clustering and semantic calculation results.
[0067] (5) Selection of representative sampling points First, frequency distribution histograms of each environmental factor set corresponding to each soil type were drawn. The environmental factors selected in this step included parent material, slope, plan curvature, profile curvature, terrain moisture index, annual mean temperature, and annual mean precipitation. Then, based on the histogram results, the pixels within the highest frequency interval of each environmental factor histogram were combined to form a representative sample set corresponding to each soil type.
[0068] (6) Generate a typical soil type sample set The representative point sets corresponding to each soil type are merged and deduplicated based on geographical location, which is the final typical soil type sample point set.
[0069] The above-mentioned method for generating typical soil type samples, on the one hand, designs a soil-environment text information model based on soil species text data to organize and store soil types and their corresponding environmental factor text information extracted from soil species text. Based on geographic similarity theory, assuming that the distribution of soil type information in geographic space corresponds to the combination of environmental factor sets in attribute space, the problem of obtaining soil type samples from text data is transformed into the problem of obtaining semantic information about specific environmental factor combinations and their soil types. Soil-environment text information is used to guide the clustering of environmental factor data, thereby obtaining environmental factor combinations and their soil type information. Representative soil type samples are selected based on the environmental factor information to generate a typical sample set. Furthermore, by using common textual soil species text data as a data source for obtaining soil type samples, a large number of soil type samples can be obtained without relying on soil experts. This overcomes the current bottleneck of fully and effectively utilizing text data, helps reduce mapping costs, and improves the efficiency of spatial inference of soil information. This method has important theoretical significance and practical application value, and has broad application prospects in digital soil mapping and land resources surveys.
[0070] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.
[0071] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A method for generating sample points based on typical soil types, characterized in that: The method includes: Determine soil types and their corresponding environmental factors based on soil annals text to build a soil-environment text information model, and use domain knowledge base and corpus to fill in the soil-environment text information model; The study area was divided into several different parent material zones using parent material factors. The number of zone clusters was determined by combining the soil-environment text information model. The environmental factor data of each parent material zone in the environmental factor geographic database were clustered using a clustering algorithm to generate environmental factor combination clusters. Converting the environmental factor text information of the instance in the soil-environment text information model into numerical data; For each environmental factor combination cluster, calculate its comprehensive environmental similarity with the soil-environment text information model to determine the matching soil type semantic information; Based on the matched soil type semantic information, the frequency sampling method is used to select representative sample points from each environmental factor combination cluster to generate a representative sample point set for each soil type. Representative sample points of all soil types were merged to generate the final set of representative soil type sample points.
2. The method for generating sample points based on typical soil types according to claim 1, characterized in that: The steps of determining soil types and their corresponding environmental factors based on soil annals text to construct a soil-environment text information model and filling the soil-environment text information model with a domain knowledge base and corpus include: Based on the soil annals text, determine the soil type and its corresponding environmental factor information; Analyze the language description characteristics of soil types and their corresponding environmental factors in soil chronicle texts, and construct a soil-environment text information model; Build the domain knowledge base and corpus required to extract soil-environment text information; When the expression pattern of text information is single and regular, a rule-based approach is used to extract soil-environment text information from the domain knowledge base and corpus; When the expression pattern of text information is complex, changeable and irregular, a method based on the BiLSTM-CRF model is used to extract soil-environment text information from the domain knowledge base and corpus.
3. The method for generating sample points based on typical soil types according to claim 2, characterized in that: The study area was divided into multiple parent material zones using parent material factors. The number of zone clusters was determined by combining the soil-environment text information model. The environmental factor data of each parent material zone in the environmental factor geographic database were clustered using a clustering algorithm. The steps of generating environmental factor combination clusters included: The study area is divided into several different parent material zones using parent material factors; Conduct statistical analysis on the extracted soil-environment text information to determine the parent material type of each parent material partition and the corresponding partition cluster number; Based on the environmental factor geographic database, select environmental factor data and perform standardized preprocessing on each environmental factor data; Based on the preprocessed environmental factor data, the K-means++ algorithm was used to cluster each parent material partition to generate environmental factor combination clusters.
4. The method for generating sample points based on typical soil types according to claim 3, characterized in that: The steps of converting the textual environmental factor values in the soil-environment text information model into numerical data include direct conversion of single-point values, conversion of segmented values into a unified format, and quantification of qualitative descriptions based on domain classification standards, including: When the environmental factor text information is a single point value; if it is a numerical single point value, no conversion is required; if it is a text single point value, the text is directly converted to a number; When the text information of environmental factors is a segmented value; if it is a numerical segmented value, the prefix, suffix and connector of the numerical body are converted into a unified expression format and then parsed using rules; if it is a textual segmented value, quantitative standards are set for the qualitative description of each environmental factor in combination with the existing field classification standards and the actual situation of the study area.
5. The method for generating sample points based on typical soil types according to claim 4, characterized in that: For each environmental factor combination cluster, the comprehensive environmental similarity between the cluster and the soil-environment text information model is calculated, and the steps of determining the matching soil type semantic information include: Set the parent material information set to , the quantified soil-environment text information model is ; According to parent material partition All environmental factors to be determined for soil semantic information are combined into clusters , the soil-environment text information model set corresponding to the parent material condition is ; Combining clusters for environmental factors Any combination of environmental factors in Calculate the cluster center and the range of each environmental factor, and take the environmental factor combination cluster The mode of the environmental factor values at each point is taken as the central value, and the environmental factor combination clusters The value range of each environmental factor in the current cluster is used as the value range of each environmental factor; Traversing the parent material partition Corresponding soil-environment text information model set , combining clusters based on environmental factors The cluster center and the range of each environmental factor are used to calculate the environmental factor combination cluster Soil-Environment Text Information Model Collection The similarity of each frame in the , and the combination cluster of environmental factors are obtained Similarity sequence ;in, Represents a cluster of environmental factor combinations Soil-Environment Text Information Model The comprehensive environmental similarity between them is measured using the Gower similarity coefficient: Where, is the calculation function of comprehensive environment similarity, for and In the environmental factor set Gower similarity coefficient sequence on , is the number of environmental factors; The calculation formula is as follows: Where, and Environmental factors Corresponding cluster Central value and framework The text information value in Cluster Medium environmental factors The value range of and Environmental factors The segment value on and Clusters that can be covered The number of pixels in , It is a cluster The total number of all pixels contained in After the calculation is completed, the minimum limiting factor method is used to select The minimum similarity is and The comprehensive similarity of geographical environment between Traversing environmental factor combination clusters For each environmental factor combination cluster in the dataset, the geographical environment comprehensive similarity is repeatedly calculated to obtain the parent material conditions. The combination clusters of environmental factors are Similarity sequence ; Calculation of parent material conditions The semantic information corresponding to each combination of environmental factors is as follows: Where, Combination of environmental factors The corresponding soil semantic information in the soil-environment text information model, is the number of environmental factor combinations to determine the semantic information, is the number of semantic tags available for selection; Traverse all parent material conditions , obtain the soil semantic information corresponding to all environmental factor combination clusters, and determine the optimal soil semantic information corresponding to each environmental factor combination.
6. The method for generating sample points based on typical soil types according to claim 5, characterized in that: Based on the matched soil type semantic information, the steps of selecting representative sample points from each environmental factor combination cluster using the frequency sampling method to generate a representative sample point set for each soil type include: Based on the matched soil type semantic information, a frequency histogram is created for each environmental factor in the environmental factor combination cluster; The pixel points whose environmental factor values fall into the highest frequency interval are determined as representative sample points corresponding to the environmental factor; if there are multiple pixels with the same highest frequency interval, all of them are selected; The representative sample points corresponding to each environmental factor were merged to form a representative sample point set corresponding to each soil type.
7. The method for generating sample points based on typical soil types according to claim 6, characterized in that: The steps of merging representative sample points of all soil types to generate the final representative soil type sample point set include: The representative sample points of all soil types were merged and geographical location was removed to generate the final set of typical soil type sample points.