Method for constructing soil bulk density prediction model, related device and equipment
By constructing a soil bulk density prediction model and utilizing the physicochemical properties, environmental characteristics, and spectral information of soil samples to screen key parameters, the model solves the problem of difficulty in measuring soil bulk density in karst landform areas, achieving accurate soil bulk density prediction and providing an important reference for agricultural ecology and ecological environmental protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2026-04-07
AI Technical Summary
In areas with complex geological conditions, such as karst landforms, it is difficult to directly measure soil bulk density. How can we establish an accurate soil bulk density prediction model for prediction?
By acquiring the physicochemical properties, environmental characteristics, and spectral information of soil samples, a soil bulk density prediction model was constructed. Using algorithms such as principal component analysis, correlation screening, and inverse stepwise regression analysis, key parameters were selected. Combined with visible/near-infrared spectral data, multiple soil bulk density prediction models were constructed, and the model with the smallest difference function was selected as the target model.
It enables more accurate and scientific prediction of soil bulk density, providing decision-making references for agricultural ecology and ecological environmental protection.
Smart Images

Figure CN117192081B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of soil bulk density prediction technology, and in particular to a method for constructing a soil bulk density prediction model, related apparatus and equipment. Background Technology
[0002] Soil bulk density refers to the mass of soil per unit volume. It is one of the fundamental physical properties of soil, reflecting its looseness and structure. It is an important indicator for measuring soil quality and productivity, providing valuable reference for analyzing soil permeability, infiltration capacity, water retention capacity, solute migration characteristics, and erosion resistance. It is also a crucial parameter for assessing soil organic carbon and nutrient reserves. Therefore, measuring soil bulk density within a certain range allows for the analysis of soil physicochemical properties, providing important decision-making references for agricultural ecology and environmental protection.
[0003] Under normal circumstances, the magnitude and spatial distribution of soil bulk density are influenced by factors such as soil texture and structure, land use, topography, and climate. For some geologically complex areas (e.g., karst landform areas), it is difficult to directly measure the soil bulk density. Therefore, how to establish accurate soil bulk density prediction models to predict soil bulk density is an increasingly important issue for technical personnel. Summary of the Invention
[0004] This application provides a method, related apparatus, and equipment for constructing a soil bulk density prediction model. The soil bulk density prediction model constructed by this method can more accurately predict the bulk density value of soil.
[0005] In a first aspect, embodiments of this application provide a method for constructing a soil bulk density prediction model, comprising: acquiring soil sample information of a target area, the soil sample information including first information, second information, and third information, the first information being used to characterize the physicochemical properties of the soil sample, the second information being used to characterize the environmental characteristics of the target area, and the third information being used to characterize the spectral information of the soil sample; calculating a first target parameter set of the soil sample based on the first information and the second information, the data in the first target parameter set being used to construct a soil bulk density prediction model, the data in the first target parameter set being the data in the first information and the second information; processing the third information in the soil sample to obtain a second target parameter set of the soil sample, the second target parameter set including L principal components of the third information of the soil sample; and constructing a target soil bulk density prediction model based on the first target parameter set and the second target parameter set.
[0006] In the above embodiments, by using the physicochemical properties of the soil, the spectral information of the soil, and the environmental parameters of the soil's environment as training samples for training the soil bulk density prediction model, the predicted soil bulk density value output by the training model is obtained through analysis of the soil's physicochemical properties, spectral information, and environmental parameters. This makes the predicted soil bulk density value output by the soil bulk density prediction model more accurate, comprehensive, and scientific.
[0007] In conjunction with the first aspect, in one possible implementation, the first information includes four data points: the percentage of gravel content in the soil rock, soil thickness, soil pH value, and particle content percentage; the second information includes three data points: rock exposure rate, elevation value, and slope aspect data; and the third information includes visible / near-infrared spectral data of the soil sample.
[0008] In conjunction with the first aspect, in one possible implementation, the first target parameter set of the soil sample is calculated based on the first information and the second information, including: calculating the first target parameter set based on the four data in the first information and the three data in the second information using the first algorithm, the second algorithm, and the third algorithm respectively, to obtain three first target parameter sets.
[0009] In conjunction with the first aspect, in one possible implementation, the first target parameter set is calculated using a first algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the target parameter set according to the formula... Calculate the score for each data point in the first and second information; The value of the j-th data point out of the seven data points in the first and second information is... Let h be the h-th principal component among the m principal components obtained after principal component processing of the first and second information, and r be the Pearson correlation coefficient between soil bulk density and the principal component. The j-th data point is the interpreted value of the h-th principal component after principal component processing; data points with scores greater than or equal to the first threshold are identified as data in the first target parameter set.
[0010] In conjunction with the first aspect, in one possible implementation, the first target parameter set is calculated using a second algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the first target parameter set according to the formula... = Calculate the correlation value between any two data points out of the seven data points in the first and second information; For the k-th correlation value, For the i-th data point out of 7 data points, For the j-th data point out of 7 data points, In all soil samples of the first soil type The mean, In all soil samples of the first soil type The mean value of i and j is used to determine the first soil type, which is the soil type corresponding to the current soil sample. The first soil type is not equal to i and j. The correlation values that are greater than or equal to the correlation threshold are determined as the first correlation values. The data corresponding to the first correlation values are determined as the first data. If there is a second data in the first data, the correlation values of the second data in the first correlation values are deleted. The second data is the first data with an importance value lower than the first importance threshold. The data in the first correlation values that are not deleted are determined as the data in the first target parameter set.
[0011] In conjunction with the first aspect, in one possible implementation, a first target parameter set is calculated using a third algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the significance index value of each data point in the first and second information using a reverse stepwise regression analysis algorithm. The significance index value is used to characterize the impact of the data in the first and second information on soil bulk density; and identifying data points with significance index values less than or equal to a first significance index threshold as data points in the first target parameter set.
[0012] In conjunction with the first aspect, in one possible implementation, the third information in each soil sample is processed to obtain a second target parameter set for that soil sample. Specifically, this includes: deleting a first edge band from the visible / near-infrared spectrum of the soil sample to obtain a first spectrum, wherein the wavelength range of the first edge band is 350nm-399nm and 2401nm-2500nm; resampling the first spectrum at a first spectral interval to obtain a sampled spectrum; smoothing and denoising the sampled spectrum to obtain a processed spectrum; and performing principal component analysis on the processed spectrum to obtain a second target parameter set, wherein the second target parameter set includes L principal components.
[0013] In conjunction with the first aspect, in one possible implementation, a target soil bulk density prediction model is constructed based on a first target parameter set and a second target parameter set. Specifically, this includes: combining the calculated three first target parameter sets with L principal components from the second target parameter set to obtain L first parameter sets, L second parameter sets, and L third parameter sets; using the L first parameter sets of each soil sample as training samples for the corresponding L first soil bulk density prediction models; using the L second parameter sets of each soil sample as training samples for the corresponding L second soil bulk density prediction models; and using the L third parameter sets of each soil sample as training samples for the corresponding L third soil bulk density prediction models. A third soil bulk density prediction model is established; among the 3L soil bulk density prediction models, the model with the smallest difference function is selected as the target soil bulk density prediction model; wherein, each first parameter set includes all data in the first target parameter set calculated by the first algorithm, each second parameter set includes all data in the first target parameter set calculated by the second algorithm, and each third parameter set includes all data in the first target parameter set calculated by the third algorithm; the i-th parameter set in the L first parameter sets includes the first i principal components of the second target parameter set, the i-th parameter set in the L second parameter sets includes the first i principal components of the second target parameter set, and the i-th parameter set in the L third parameter sets includes the first i principal components of the second target parameter set.
[0014] Thirdly, embodiments of this application provide a soil bulk density prediction model construction device, including a memory and a processor;
[0015] The memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the soil bulk density prediction model construction method in the first aspect and its various possible implementations.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the soil bulk density prediction model construction method described in the first aspect and its various possible implementations.
[0017] Fifthly, embodiments of this application provide a computer program including instructions that, when executed by a computer, cause a soil bulk density prediction model construction device to execute the processes performed by the soil bulk density prediction model construction device in the first aspect and its various possible implementations described above. Attached Figure Description
[0018] The accompanying drawings used in the embodiments of this application are described below.
[0019] Figure 1 This is a flowchart of a method for establishing a soil bulk density prediction model provided in an embodiment of this application;
[0020] Figure 2 This is a flowchart of a principal component acquisition method provided in an embodiment of this application;
[0021] Figure 3 This is a flowchart illustrating the construction of a soil bulk density prediction model for limestone soil, as provided in an embodiment of this application.
[0022] Figure 4 This is a schematic diagram of the structure of a soil bulk density prediction model construction device 40 provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the structure of a soil bulk density prediction model construction device 50 provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. The term "embodiment" as used herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The appearance of this phrase in different places in the specification does not necessarily indicate the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0025] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects and not to describe a particular order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, it may include a series of steps or units, or optionally, steps or units not listed, or other steps or units inherent to these processes, methods, products, or devices.
[0026] The accompanying drawings show only the portions relevant to this application, not all of them. Before discussing exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations may be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations may be rearranged. The process may be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process may correspond to a method, function, procedure, subroutine, subprogram, etc.
[0027] The terms “component,” “module,” “system,” “unit,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or distributed between two or more computers. Furthermore, these units can be executed from various computer-readable media on which various data structures are stored. Units can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from a second unit interacting with another unit between a local system, a distributed system, and / or a network; for example, the Internet interacting with other systems via signals).
[0028] Soil bulk density refers to the mass of soil per unit volume. It is one of the fundamental physical properties of soil, reflecting its looseness and structure. It is an important indicator for measuring soil quality and productivity, providing valuable reference for analyzing soil permeability, infiltration capacity, water retention capacity, solute migration characteristics, and erosion resistance. It is also a crucial parameter for assessing soil organic carbon and nutrient reserves. Therefore, measuring soil bulk density within a certain range allows for the analysis of soil physicochemical properties, providing important decision-making references for agricultural ecology and environmental protection.
[0029] Under normal circumstances, the magnitude and spatial distribution of soil bulk density are influenced by factors such as soil texture and structure, land use, topography, and climate. For some geologically complex areas (e.g., karst landform areas), it is difficult to directly measure the soil bulk density. Therefore, how to establish accurate soil bulk density prediction models to predict soil bulk density is an increasingly important issue for technical personnel.
[0030] Therefore, to solve the above problems, this application provides a method for establishing a soil bulk density prediction model. The process of establishing a soil bulk density prediction model provided by this application will be described below with reference to the accompanying drawings.
[0031] Please see Figure 1 , Figure 1 This is a flowchart of a method for establishing a soil bulk density prediction model provided in an embodiment of this application. The specific process is as follows:
[0032] S101: Obtain soil sample information for N soil types in the target area. The soil sample information includes first information, second information, and third information. The first information is used to characterize the physicochemical properties of the soil sample, the second information is used to characterize the environmental characteristics of the target area, and the third information is used to characterize the spectral information of the sample soil.
[0033] Specifically, the target area can be a karst landform area. Before establishing the soil bulk density prediction model, soil sample information for N soil types in the target area can be obtained. Multiple soil samples can be collected for each soil type. Soil samples of the same type can be collected in the same area or in different areas; this application embodiment does not impose any restrictions on this. Soil types can be classified according to the physicochemical properties of the soil or according to the composition of the soil; this application embodiment does not impose any restrictions on the classification of soil types.
[0034] For ease of description, this application example uses three soil types as an example, which can be divided into three types: limestone soil, paddy soil, and yellow soil. For example, the number of samples for each type of soil can be 100.
[0035] After obtaining soil samples, their physicochemical properties can be analyzed to obtain soil sample information for each sample. This soil sample information includes first information, second information, and third information.
[0036] The first piece of information is used to characterize the physicochemical properties of the soil. This information includes the percentage of gravel in the soil matrix, soil thickness, soil pH value, and particle content percentage. The percentage of gravel content indicates the proportion of gravel in the soil sample, while the particle content percentage includes the percentage of sand, silt, and clay particles. Since soil bulk density refers to the mass of soil per unit volume, the physicochemical properties of the soil have a certain influence on its bulk density.
[0037] For example, regarding the percentage of gravel content, gravel in soil is not considered part of the soil itself. Within a given volume, a higher percentage of gravel content results in a lower soil content and a lower soil bulk density; conversely, a lower percentage of gravel content results in a higher soil content and a higher soil bulk density. Regarding soil thickness, which characterizes the vertical distance between the sampled soil area and the surface, generally, for non-topsoil, due to the pressure from the upper soil layer, as depth increases, the soil porosity decreases, the soil density increases, and the soil content within a given volume increases, thus increasing the soil bulk density. Regarding soil pH, since soil acidity or alkalinity can affect changes in soil chemical properties (e.g., increasing / decreasing soluble substances in the soil), an increase or decrease in soluble substances in the soil can also affect the soil bulk density. In terms of particle content percentage, soil is composed of different particles, including sand, silt, organic matter, and water. Particles of different sizes have different physical and chemical properties, which affect soil bulk density. Sand and silt are the main solid particles in soil, accounting for most of the soil mass. Sand and silt have smaller particle sizes and higher densities. The higher the proportion of sand and silt, the higher the soil bulk density is usually, and the lower the proportion of sand and silt, the lower the soil bulk density is usually.
[0038] Therefore, the data listed in the first information above are parameters related to soil bulk density, and the bulk density value of the soil can be calculated using the data in the first information above.
[0039] The second piece of information characterizes the environmental conditions of the area where the soil sample is located. This information may include rock exposure rate, elevation, and slope aspect data. The rock exposure rate is the ratio of the area of surface rock in the target area to the total area of the target area. Slope aspect data characterizes the direction of the slope where the soil sample is located; for example, a numerical range can be used to represent the slope aspect. For instance, 1 represents a slope aspect of due east, 2 represents a slope aspect of due south, 3 represents a slope aspect of due west, and 4 represents a slope aspect of due north. Since soil bulk density refers to the mass of soil per unit volume, the environmental conditions of the area where the soil sample is located have a certain degree of influence on the soil bulk density.
[0040] For example, rock exposure rate is used to characterize soil development. A higher rock exposure rate indicates a lower degree of soil development, less soil content, and a lower bulk density; conversely, a lower rock exposure rate indicates a higher degree of soil development and a higher bulk density. Regarding the altitude of the soil sample location, different altitudes may have different vegetation types and growth conditions, which may affect the soil's physicochemical properties and its ability to retain nutrients, thus influencing soil quality and, consequently, bulk density. Regarding the slope aspect of the soil sample location, different slope aspects may result in different light intensity or duration, potentially affecting vegetation development and, consequently, bulk density. Therefore, the soil bulk density can be calculated using the data in the second piece of information mentioned above.
[0041] The third piece of information includes the visible / near-infrared spectral data of the soil sample. The visible / near-infrared spectral data is used to characterize the absorption, reflection, and transmission of visible / near-infrared light in different wavelength ranges by the soil. By analyzing the visible / near-infrared spectral data of the soil sample, the physicochemical properties of the soil can be analyzed, thereby calculating the soil bulk density value.
[0042] S102: Based on the first information and the second information, determine M first target parameter sets to participate in the construction of the soil bulk density prediction model. The data in the first target parameter sets are the data in the first information and the second information, and M is greater than or equal to 1.
[0043] Specifically, after obtaining soil sample information, due to the large amount of data in the soil sample information, the degree of influence of each data point on soil bulk density prediction varies. Therefore, to reduce the computational load in the soil bulk density prediction process, higher-priority data can be selected from the first and second information data as the first objective parameter set for constructing the soil bulk density prediction model, thus obtaining M first objective parameter sets.
[0044] The fractions F of the seven data points in the first information (gravel content percentage, soil thickness value, soil pH value, particle content percentage, rock exposure rate, altitude value, and slope aspect data in the second information) of each soil sample can be calculated using formula (1). Formula (1) is shown below:
[0045] (1)
[0046] in, Let represent the importance of the j-th data point to soil bulk density among these 7 data points; m represents the number of principal components retained after principal component analysis (PCA) of these 7 data points. It is the i-th principal component of these 7 data; r is the Pearson correlation coefficient between soil bulk density and the principal component; This refers to the explanatory value of the j-th data point among these 7 data points for the h-th principal component after PCA.
[0047] After calculating the scores of these seven data points, the data with scores greater than or equal to the first threshold can be identified as data in the first target parameter set. The first threshold can be obtained experimentally, empirically, or from experimental data, and can be set to 0.8 or 1; this embodiment does not impose any limitations on this.
[0048] In one possible implementation, the data in the first target parameter set can also be determined from the first and second information by the correlation screening method. Specifically, the correlation value of any two data in the above 7 data can be calculated by formula (2), thereby obtaining 21 correlation values. Formula (2) is shown below:
[0049] = (2)
[0050] Among them, the The k-th relevant value out of 21 relevant values. For the i-th data among the above 7 data, For the j-th data point among the above 7 data points, For data The mean of all soil samples of the corresponding soil type, where y is the data. In the mean values of all soil samples of the corresponding soil type, i and j are not equal.
[0051] For example, suppose The calculation focuses on the correlation between the third data point, "rock exposure rate," and the fourth data point, "soil thickness value," in a soil sample of 100 soil types, specifically calcareous soil. This can be represented as the "rock exposure rate" (the third data point) in the soil sample. It can be the average of the "rock exposure rate" of these 100 soil samples with the soil type of limestone soil; This can be represented as the "rock exposure rate" (4th data point) in the soil sample. It can be the average of the "soil thickness value" of these 100 soil samples with the soil type of limestone soil.
[0052] After calculating the 21 correlation values for these 7 data points, an initial correlation threshold can be set first. Then, multiple correlation thresholds can be obtained with a certain step size. For each correlation threshold, correlation values greater than or equal to that threshold are selected. Then, the data corresponding to the selected correlation values are used to calculate the first data point through a random forest model. The first data point is the data with relatively low importance in calculating soil bulk density in the random forest model.
[0053] Each of the above seven data points has a relative importance characteristic value for soil bulk density. The characteristic value of the relative importance of each data point can be calculated using the formula (1) above. The average of the relative importance features of these seven data points can be used as a threshold. Indicators above the threshold are considered important, while those below the threshold are considered unimportant.
[0054] In this way, from the data corresponding to the filtered correlation values, the data with correlation values greater than or equal to the first data and the first data can be deleted, thereby obtaining the second data, which is the data in the first target parameter set.
[0055] For example, the 21 correlation values are: correlation value 1 (data 1 and data 2, 0.9), correlation value 2 (data 1 and data 3, 0.5), correlation value 3 (data 1 and data 4, 0.8), correlation value 4 (data 1 and data 5, 0.6), correlation value 5 (data 1 and data 6, 0.4), correlation value 6 (data 1 and data 7, 0.7), correlation value 7 (data 2 and data 3, 0.1), correlation value 8 (data 2 and data 4, 0.2), correlation value 9 (data 2 and data 5, 0.3), correlation value 10 (data 2 and data 6, 0.6), and correlation value 11. (Data 2 and Data 7, 0.8), Correlation value 12 (Data 3 and Data 4, 0.5), Correlation value 13 (Data 3 and Data 5, 0.4), Correlation value 14 (Data 3 and Data 6, 0.3), Correlation value 15 (Data 3 and Data 7, 0.1), Correlation value 16 (Data 4 and Data 5, 0.55), Correlation value 17 (Data 4 and Data 6, 0.8), Correlation value 18 (Data 4 and Data 7, 0.6), Correlation value 19 (Data 5 and Data 6, 0.4), Correlation value 20 (Data 5 and Data 7, 0.7), Correlation value 21 (Data 6 and Data 7, 0.6).
[0056] Assuming an initial correlation threshold of 0.4 and a step size of 0.2, multiple correlation thresholds can be obtained: 0.6, 0.8, and 1 (the maximum correlation threshold is 1). For a correlation threshold of 0.4, correlation values greater than or equal to 0.4 can be retained, resulting in correlation values 1, 2, 3, 4, 5, 6, 10, 11, 12, 13, 16, 17, 18, 19, 20, and 21. Based on these correlation values, the corresponding data (the first target data) can be obtained as: data 1, data 2, data 3, data 4, data 5, data 6, and data 7. These 7 data points are then processed using a forest random model, resulting in data 6 as the first data point. The first set of data corresponds to 16 relevant values selected from the initial screening. These values include: 5 (data 1 and data 6), 10 (data 2 and data 6), 17 (data 4 and data 6), 19 (data 5 and data 6), and 21 (data 5 and data 6). Therefore, data 1, data 2, data 4, and data 5 have a correlation greater than or equal to 0.4 with data 6. From these seven target data points, data 1, data 2, data 3, data 4, and data 6 are removed, leaving data 7. Data 7 is then used as the data in the first target parameter set.
[0057] Similarly, the data in the first target parameter set can be determined using the above method for correlation thresholds of 0.6, 0.8, and 1, respectively. Then, the data in the first target parameter set determined under these five correlation thresholds are all included in the data in the first target parameter set, so that the final first target parameter set includes the data determined under these four correlation thresholds. The step size of the correlation threshold can be set arbitrarily; it can be obtained from experimental values, empirical values, or experimental data. This embodiment does not impose any restrictions on this.
[0058] The above method can identify multiple highly correlated data points. However, when using two highly correlated data points for soil bulk density prediction, these two data points may overlap significantly. Therefore, if a second data point is calculated using a random forest model, the data points with strong correlation to the second data point can be removed, thereby reducing unnecessary data calculations and computational load during soil bulk density prediction.
[0059] In one possible implementation, a significance value p can be calculated for each data point in the first and second information, and then data with significance values less than or equal to a first significance threshold can be identified as data in the first target parameter set.
[0060] The significance index value of each data point in the first information and the second information can be calculated using a reverse stepwise regression analysis algorithm. The significance index value is used to characterize the impact of the data in the first information and the second information on soil bulk density. Data points whose significance index values are less than or equal to a first significance index threshold are determined as data points in the first target parameter set.
[0061] S103: Obtain L principal components from the visible / near-infrared spectrum of the third information of the soil sample to obtain a second target parameter set, the second target parameter set including the L principal components.
[0062] Specifically, the visible / near-infrared spectrum of a soil sample represents the intensity of visible / near-infrared light across a specific wavelength. Since the wavelength range of light is continuous, the information from the visible / near-infrared spectrum of a soil sample is compressed into discrete information components, namely, L principal components. Different soil samples may exhibit different visible / near-infrared spectra due to variations in their physicochemical properties, resulting in differences in the corresponding principal components of their visible / near-infrared spectra. However, the number of principal components remains constant at L.
[0063] To illustrate this better, the following will be combined with... Figure 2 The procedure for obtaining the L principal components of the visible / near-infrared spectrum in the third information of the soil sample is described. Please refer to [link to documentation]. Figure 2 , Figure 2 This application provides a flowchart for obtaining principal components, and the specific process is as follows:
[0064] S201: Delete the first edge band in the visible / near-infrared spectrum of the soil sample to obtain a first spectrum, wherein the wavelength range of the first edge band is 350nm-399nm and 2401nm-2500nm.
[0065] S202: Resample the first spectrum at the first spectral interval to obtain the sampled spectrum.
[0066] For example, the first spectral interval can be 10 nm.
[0067] S203: Perform smoothing and noise reduction processing on the sampled spectrum to obtain the processed spectrum.
[0068] S204: Perform principal component analysis on the processed spectrum to obtain the second target parameter set, which includes L principal components.
[0069] Specifically, Principal Component Analysis (PCA) is a commonly used data dimensionality reduction method that can transform multiple indicators into a few principal components. These principal components are linear combinations of the original variables and are uncorrelated with each other, thus reflecting most of the information in the original data.
[0070] In PCA, the covariance matrix of the processed spectral data is first calculated, and then the eigenvectors and eigenvalues of this covariance matrix are solved. These eigenvectors form an orthogonal basis. Projecting each sample in the original dataset onto this orthogonal basis yields a new dataset with k dimensions, where the first k dimensions are the original eigenvectors, and the last dimension is the noise dimension. This new dataset can be represented as a matrix X', where each row of X' represents a sample. Each sample is a principal component.
[0071] It should be understood that S103 can be executed before S102, after S102, or simultaneously with S102. This application does not restrict the execution order of S103 and S102.
[0072] S104: Construct a soil bulk density prediction model for each soil type based on the first target parameter set and the second target parameter set.
[0073] Specifically, after determining the first set of target parameters and the second set of target parameters, a soil bulk density prediction model for each soil type can be constructed based on the first set of target parameters and the second set of target parameters.
[0074] For ease of description, this application uses the construction of a soil bulk density prediction model based on a first set of objective parameters and a second set of objective parameters for a soil sample of limestone soil type as an example for illustration.
[0075] Please see Figure 3 , Figure 3 This is a flowchart illustrating the construction of a soil bulk density prediction model for limestone soil, provided in an embodiment of this application. Assuming there are 100 soil samples of limestone soil, the specific process is as follows:
[0076] S301: Based on the number L of principal components in the first target parameter set and the second target parameter set of each lime soil sample, 3L third target parameter sets are obtained.
[0077] Specifically, in S102 above, three first objective parameter sets were obtained for each limestone soil sample using three methods, and in S103 above, L principal components were obtained for each limestone soil sample. To determine which method's first objective parameter set combined with the principal components is more accurate as training parameters for the soil bulk density prediction model, the three first objective parameter sets and the L principal components can be combined to obtain 3L third objective parameter sets. A portion of the 100 soil samples can be used as training samples for the soil bulk density prediction model; another portion can be used as validation samples to verify the accuracy of the soil bulk density prediction model.
[0078] For example, if the three target parameter sets are parameter set 1, parameter set 2, and parameter set 3, and the second target parameter set includes three principal components (principal component 1 to principal component 3), then nine third target parameter sets can be obtained for each limestone soil sample. These nine third target parameter sets are as follows:
[0079] Parameter set 1 + principal component 1, Parameter set 1 + principal component 1 + principal component 2, Parameter set 1 + principal component 1 + principal component 2 + principal component 3, Parameter set 2 + principal component 1, Parameter set 2 + principal component 1 + principal component 2, Parameter set 2 + principal component 1 + principal component 2 + principal component 3, Parameter set 3 + principal component 1, Parameter set 3 + principal component 1 + principal component 2, Parameter set 3 + principal component 1 + principal component 2 + principal component 3.
[0080] S302: Normalize the data in the 3L sets of third target parameters to obtain 3L sets of fourth target parameters for each soil sample of lime soil.
[0081] Specifically, before training the soil bulk density prediction model, each data point in the fourth objective parameter set can be normalized to obtain the normalized third objective parameter set, i.e., the fourth objective parameter set.
[0082] To facilitate understanding, we will take the normalization of the third objective parameter set of a single limestone soil sample as an example. Assume there are 100 limestone soil samples. The normalization of principal component 1 can be obtained using formula (3), as shown below:
[0083] =( - ) / ( - (3)
[0084] in, Let the normalized principal component 1 be the i-th soil sample from these 100 limestone soil samples. Let 1 be the principal component of the i-th soil sample out of these 100 limestone soil samples. The maximum principal component value among these 100 limestone soil samples. The minimum principal component value among these 100 limestone soil samples.
[0085] Similarly, by using the above method, each data point in each third objective parameter set of each soil sample can be normalized to obtain each normalized third objective parameter set, i.e., the fourth objective parameter set.
[0086] S303: Train the corresponding soil bulk density prediction model using the 3L fourth objective parameter sets of each soil training sample to obtain 3L soil bulk density training models.
[0087] Specifically, after calculating the 3L fourth objective parameter sets for each soil training sample, 3L soil bulk density prediction models can be constructed, and each of these 3L soil bulk density prediction models corresponds one-to-one with one of the 3L fourth objective parameter sets.
[0088] For example, each soil bulk density prediction model can be trained using formula (4), which is shown below:
[0089] = + + + +……+ + (4)
[0090] in, Let be the soil bulk density value of the i-th sample in the training soil sample of limestone soil. ~ It is the predictor of the j-th soil bulk density prediction model for limestone soil samples. This is the error term of the soil bulk density prediction model for the j-th limestone soil sample. ~ It is the data in the j-th fourth objective parameter set of the i-th sample in the limestone soil training samples.
[0091] Using the above method, the corresponding soil bulk density prediction model can be trained using data from the 3L fourth objective parameter sets of each limestone soil training sample, thereby obtaining the prediction factor of each soil bulk density prediction model. This allows the soil bulk density prediction model to output the predicted bulk density value of this type of soil when the fourth objective parameter set of limestone soil is input into the soil bulk density prediction model.
[0092] In some embodiments, the 3L fourth objective parameter sets of each lime soil validation sample can be used as validation samples to test the accuracy of the corresponding soil bulk density samples, so that each soil bulk density prediction model outputs the corresponding soil bulk density prediction value.
[0093] S304: Determine the target soil bulk density prediction model among the 3L soil bulk density prediction models.
[0094] Specifically, after validating each soil bulk density prediction model, the first evaluation parameter Q1, the second evaluation parameter Q2, and the third evaluation parameter Q3 of each soil bulk density prediction model can be calculated according to formulas (5), (6), and (7), respectively. Formulas (5), (6), and (7) are shown below:
[0095] = / [ (5)
[0096] = (6)
[0097] = (7)
[0098] Where w represents the number of limestone soil validation samples. The average bulk density value of each limestone soil validation sample. Let be the soil bulk density value of the i-th limestone soil validation sample. Let be the predicted soil bulk density value output by the j-th soil bulk density model based on the fourth objective parameter set corresponding to the i-th soil sample. The j-th model outputs the average value of the predicted soil bulk density for the fourth objective parameter set corresponding to all limestone soil samples.
[0099] Then, according to , and From these 3L soil bulk density prediction models, the target soil bulk density prediction model for limestone soil was selected. This can be... The top 5 soil bulk density prediction models were selected as the primary soil bulk density prediction model. Then, among these 5 models, [the following will be considered]... The soil bulk density prediction model with the smallest sum value was determined as the target soil bulk density prediction model.
[0100] Similarly, the target soil bulk density prediction model for yellow soil and paddy soil can also be obtained through the above method.
[0101] In this embodiment of the application, the physical and chemical properties of the soil, the spectral information of the soil, and the environmental parameters of the soil environment are used as training samples to train the soil bulk density prediction model. This makes the soil bulk density prediction value output by the soil bulk density prediction model at the training point more accurate, comprehensive and scientific.
[0102] It should be understood that the steps and the execution order of each step in the embodiments of this application are merely illustrative examples. The order of each step in the embodiments of this application can be adjusted and / or one or more steps can be deleted to obtain different embodiments. The obtained embodiments still fall within the protection scope of the embodiments of this application.
[0103] The methods of the embodiments of this application have been described in detail above. The related devices, equipment, computer-readable storage media, and computer programs of the embodiments of this application are described below.
[0104] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a soil bulk density prediction model construction device 40 provided in an embodiment of this application. The soil bulk density prediction model construction device 40 may include an acquisition unit 401, a first calculation unit 402, a first processing unit 403, and a model construction unit 404; wherein:
[0105] Acquisition unit 401 is used to acquire soil sample information of the target area;
[0106] The first calculation unit 402 is used to calculate the first target parameter set of the soil sample based on the first information and the second information.
[0107] The first processing unit 403 is used to process the third information in the soil sample to obtain the second target parameter set of the soil sample.
[0108] Model building unit 404 is used to build a target soil bulk density prediction model based on the first target parameter set and the second target parameter set.
[0109] In one possible implementation, the first target parameter set of the soil sample is calculated based on the first information and the second information, including: calculating the first target parameter set using the first algorithm, the second algorithm and the third algorithm respectively based on the four data in the first information and the three data in the second information, to obtain three first target parameter sets.
[0110] In one possible implementation, the first target parameter set is calculated using a first algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the target parameter set according to the formula... Calculate the score for each data point in the first and second information; The value of the j-th data point out of the seven data points in the first and second information is... Let h be the h-th principal component among the m principal components obtained after principal component processing of the first and second information, and r be the Pearson correlation coefficient between soil bulk density and the principal component. The j-th data point is the interpreted value of the h-th principal component after principal component processing; data points with scores greater than or equal to the first threshold are identified as data in the first target parameter set.
[0111] In one possible implementation, the first target parameter set is calculated using a second algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the target parameter set according to the formula... = Calculate the correlation value between any two data points out of the seven data points in the first and second information; For the k-th correlation value, For the i-th data point out of 7 data points, For the j-th data point out of 7 data points, In all soil samples of the first soil type The mean, In all soil samples of the first soil type The mean value of i and j is used to determine the first soil type, which is the soil type corresponding to the current soil sample. The first soil type is not equal to i and j. The correlation values that are greater than or equal to the correlation threshold are determined as the first correlation values. The data corresponding to the first correlation values are determined as the first data. If there is a second data in the first data, the correlation values of the second data in the first correlation values are deleted. The second data is the first data with an importance value lower than the first importance threshold. The data in the first correlation values that are not deleted are determined as the data in the first target parameter set.
[0112] In one possible implementation, a first target parameter set is calculated using a third algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the significance index value of each data point in the first and second information using a reverse stepwise regression analysis algorithm. The significance index value is used to characterize the impact of the data in the first and second information on soil bulk density; and determining the data whose significance index value is less than or equal to a first significance index threshold as data in the first target parameter set.
[0113] In one possible implementation, the third information in each soil sample is processed to obtain a second target parameter set for that soil sample. Specifically, this includes: deleting a first edge band from the visible / near-infrared spectrum of the soil sample to obtain a first spectrum, wherein the wavelength range of the first edge band is 350nm-399nm and 2401nm-2500nm; resampling the first spectrum at a first spectral interval to obtain a sampled spectrum; smoothing and denoising the sampled spectrum to obtain a processed spectrum; and performing principal component analysis on the processed spectrum to obtain a second target parameter set, wherein the second target parameter set includes L principal components.
[0114] In one possible implementation, a target soil bulk density prediction model is constructed based on a first target parameter set and a second target parameter set. Specifically, this includes: combining the calculated three first target parameter sets with L principal components from the second target parameter set to obtain L first parameter sets, L second parameter sets, and L third parameter sets; using the L first parameter sets of each soil sample as training samples for the corresponding L first soil bulk density prediction models; using the L second parameter sets of each soil sample as training samples for the corresponding L second soil bulk density prediction models; using the L third parameter sets of each soil sample as training samples for the corresponding L third soil bulk density prediction models; and selecting the model with the smallest difference function among the 3L soil bulk density prediction models as the target soil bulk density prediction model.
[0115] Please see Figure 5 , Figure 5 This is a schematic diagram of a soil bulk density prediction model construction device 50 provided in an embodiment of this application. The soil bulk density prediction model construction device 50 may include a memory 501, a communication module 502, and a processor 503; wherein, the detailed description of each unit is as follows:
[0116] Memory 501 is used to store program code.
[0117] Processor 503 is used to call program code stored in memory to perform the following steps:
[0118] The communication module 502 acquires soil sample information of the target area; calculates the first target parameter set of the soil sample based on the first and second information; processes the third information in the soil sample to obtain the second target parameter set of the soil sample; and constructs a target soil bulk density prediction model based on the first and second target parameter sets.
[0119] In conjunction with the first aspect, in one possible implementation, the first target parameter set of the soil sample is calculated based on the first information and the second information, including: calculating the first target parameter set based on the four data in the first information and the three data in the second information using the first algorithm, the second algorithm, and the third algorithm respectively, to obtain three first target parameter sets.
[0120] In conjunction with the first aspect, in one possible implementation, the first target parameter set is calculated using a first algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes calculating the score of each data point in the first and second information according to a formula.
[0121] In conjunction with the first aspect, in one possible implementation, the first target parameter set is calculated using a second algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the first target parameter set according to the formula... = Calculate the correlation value between any two data points out of the seven data points in the first and second information; For the k-th correlation value, For the i-th data point out of 7 data points, For the j-th data point out of 7 data points, In all soil samples of the first soil type The mean, In all soil samples of the first soil type The mean value of i and j is used to determine the first soil type, which is the soil type corresponding to the current soil sample. The first soil type is not equal to i and j. The correlation values that are greater than or equal to the correlation threshold are determined as the first correlation values. The data corresponding to the first correlation values are determined as the first data. If there is a second data in the first data, the correlation values of the second data in the first correlation values are deleted. The second data is the first data with an importance value lower than the first importance threshold. The data in the first correlation values that are not deleted are determined as the data in the first target parameter set.
[0122] In one possible implementation, a first target parameter set is calculated using a third algorithm based on four data points from the first information and three data points from the second information. Specifically, this includes: calculating the significance index value of each data point in the first and second information using a reverse stepwise regression analysis algorithm. The significance index value is used to characterize the impact of the data in the first and second information on soil bulk density; and determining the data whose significance index value is less than or equal to a first significance index threshold as data in the first target parameter set.
[0123] In one possible implementation, the third information in each soil sample is processed to obtain a second target parameter set for that soil sample. Specifically, this includes: deleting a first edge band from the visible / near-infrared spectrum of the soil sample to obtain a first spectrum, wherein the wavelength range of the first edge band is 350nm-399nm and 2401nm-2500nm; resampling the first spectrum at a first spectral interval to obtain a sampled spectrum; smoothing and denoising the sampled spectrum to obtain a processed spectrum; and performing principal component analysis on the processed spectrum to obtain a second target parameter set, wherein the second target parameter set includes L principal components.
[0124] In one possible implementation, a target soil bulk density prediction model is constructed based on a first target parameter set and a second target parameter set. Specifically, this includes: combining the calculated three first target parameter sets with L principal components from the second target parameter set to obtain L first parameter sets, L second parameter sets, and L third parameter sets; using the L first parameter sets of each soil sample as training samples for the corresponding L first soil bulk density prediction models; using the L second parameter sets of each soil sample as training samples for the corresponding L second soil bulk density prediction models; using the L third parameter sets of each soil sample as training samples for the corresponding L third soil bulk density prediction models; and selecting the model with the smallest difference function among the 3L soil bulk density prediction models as the target soil bulk density prediction model.
[0125] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for constructing a soil bulk density prediction model in the above embodiments and their various possible implementations.
[0126] This application provides a computer program that includes instructions that, when executed by a computer, enable a soil bulk density prediction model construction device to perform the processes executed by the soil bulk density prediction model construction device in the above embodiments and their various possible implementations.
[0127] It should be noted that the memory in the above embodiments can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor via a bus. The memory can be integrated with the processor.
[0128] The processor in the above embodiments may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the above scheme program.
[0129] For the foregoing method embodiments, in order to simplify the description, they are all expressed as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps may be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0130] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of the units described above is merely a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0131] The units described above as separate components may or may not be physically separate. Similarly, the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0132] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The aforementioned integrated unit can be implemented in hardware or as a software functional unit.
[0133] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in software form. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium may include various media capable of storing program code, such as a USB flash drive, portable hard drive, magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM).
[0134] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for constructing a soil bulk density prediction model, characterized in that, include: For each soil type, perform the following operations: Soil sample information from the target area is obtained. This soil sample information includes first information, second information, and third information. The first information characterizes the physicochemical properties of the soil sample; the second information characterizes the environmental characteristics of the target area; and the third information characterizes the spectral information of the soil sample. The first information includes four data points: percentage of gravel content in the soil rock, soil thickness, soil pH, and percentage of particle content. The second information includes three data points: rock exposure rate, altitude, and slope aspect. The third information includes visible / near-infrared spectral data of the soil sample. Based on the four data points in the first information and the three data points in the second information, the first target parameter set is calculated using the first algorithm, the second algorithm, and the third algorithm, respectively, resulting in three first target parameter sets. The data in the first target parameter sets are used to construct the soil bulk density prediction model. The data in the first target parameter sets are the data in the first information and the second information. The third information in the soil sample is processed to obtain a second target parameter set of the soil sample, which includes L principal components of the third information of the soil sample. The calculated three sets of first objective parameters are combined with the L principal components in the second objective parameter set to obtain L sets of first parameters, L sets of second parameters, and L sets of third parameters. The L first parameter sets of each soil sample are used as training samples for the corresponding L first soil bulk density prediction models to train the L first soil bulk density prediction models. The L sets of second parameters for each soil sample are used as training samples for the corresponding L second soil bulk density prediction models to train the L second soil bulk density prediction models. The L sets of third parameters for each soil sample are used as training samples for the corresponding L third soil bulk density prediction models to train the L third soil bulk density prediction models. Among the 3L soil bulk density prediction models, the model with the smallest difference function was selected as the target soil bulk density prediction model. Each first parameter set includes all data from the first target parameter set calculated by the first algorithm; each second parameter set includes all data from the first target parameter set calculated by the second algorithm; each third parameter set includes all data from the first target parameter set calculated by the third algorithm; the i-th parameter set in the L first parameter sets includes the first i principal components of the second target parameter set; the i-th parameter set in the L second parameter sets includes the first i principal components of the second target parameter set; and the i-th parameter set in the L third parameter sets includes the first i principal components of the second target parameter set.
2. The method as described in claim 1, characterized in that, Based on the four data points in the first information and the three data points in the second information, a first target parameter set is calculated using a first algorithm, specifically including: According to the formula Calculate the score for each data point in the first and second information; The value of the j-th data point among the seven data points of the first and second information is... The h-th principal component is one of the m principal components obtained after principal component processing of the first and second information, where r is the Pearson correlation coefficient between soil bulk density and the principal component. The j-th data point among the seven data points is the explanatory value of the h-th principal component after principal component processing, where y is the sample value of soil bulk density; Data with scores greater than or equal to the first threshold are identified as data in the first target parameter set.
3. The method as described in claim 1, characterized in that, Based on the four data points in the first information and the three data points in the second information, the first target parameter set is calculated using the second algorithm, specifically including: According to the formula = Calculate the correlation value between any two data points from the seven data points in the first and second information; For the k-th correlation value, the For the i-th data among the seven data points, For the j-th data among the 7 data, the In all soil samples of the first soil type The mean, the For all soil samples of the first soil type The average value, where the first soil type is the soil type corresponding to the current soil sample, and i and j are not equal; The correlation value that is greater than or equal to the correlation threshold is defined as the first correlation value: The data corresponding to the first correlation value is determined as the first data; If there is second data in the first data, the correlation value of the second data in the first correlation value is deleted. The second data is the first data whose importance value is lower than the first importance threshold. The data in the first correlation values that have not been deleted are identified as data in the first target parameter set.
4. The method as described in claim 1, characterized in that, Based on the four data points in the first information and the three data points in the second information, the first target parameter set is calculated using a third algorithm, specifically including: The significance index value of each data point in the first information and the second information is calculated by the reverse stepwise regression analysis algorithm. The significance index value is used to characterize the impact of the data in the first information and the second information on soil bulk density. Data whose significance index value is less than or equal to the first significance index threshold are identified as data in the first target parameter set.
5. The method as described in claim 1, characterized in that, The process of processing the third information in each soil sample to obtain the second target parameter set of the soil sample specifically includes: The first spectrum is obtained by deleting the first edge band in the visible / near-infrared spectrum of the soil sample, and the wavelength range of the first edge band is 350nm-399nm and 2401nm-2500nm. The first spectrum is resampled at a first spectral interval to obtain the sampled spectrum; The sampled spectrum is smoothed and denoised to obtain the processed spectrum; Principal component analysis is performed on the processed spectrum to obtain the second target parameter set, which includes L principal components.
6. A soil bulk density prediction model construction device, characterized in that, Includes a unit that performs the method as described in any one of claims 1-5.
7. A soil bulk density prediction model construction device, characterized in that, include: Memory and processor, wherein: The memory is used to store computer programs, the computer programs including program instructions; The processor is used to invoke the program instructions, causing the soil bulk density prediction model construction device to perform the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Method for estimating content of heavy metals in soil based on hyperspectral remote sensing technology
CN114018833A
Soil heavy metal content inversion method fusing multi-source environment variables and spectral information
CN114814167A
Soil organic carbon spectrum prediction method and device based on spectrum guided ensemble learning
CN116818687A