Soil nutrient prediction method based on high-dimensional spectral data compression pretreatment

Through adaptive segmentation compression and non-uniform segmentation pretreatment technology, the problems of low data processing efficiency and neglect of local features in soil nutrient prediction are solved, rapid dimensionality reduction and efficient prediction are achieved, and the accuracy of the prediction model is improved.

CN120148679AActive Publication Date: 2025-06-13CHANGCHUN UNIV OF SCI & TECH +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510631562.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The prior art faces the problems of low data processing efficiency, redundant information, complex preprocessing process and ignoring local characteristics of data in soil nutrient prediction, which affects the efficiency and accuracy of the prediction model.

Method used

Adaptive segmentation compression and non-uniform segmentation pretreatment methods based on high-dimensional spectral data are adopted to reduce the data dimensions through adaptive segmentation compression, and the local characteristics of the data are captured through non-uniform segmentation pretreatment, and finally a partial least squares model is established for soil nutrient prediction.

Benefits of technology

This method optimizes the data processing process, quickly reduces the spectral data dimension, improves prediction efficiency and accuracy, can better capture local features of the data, and improves the accuracy of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148679A_ABST
    Figure CN120148679A_ABST
Patent Text Reader

Abstract

The invention discloses a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, and belongs to the technical field of soil sample analysis. The method comprises the following steps: self-adaptive segmented compression: segmenting according to data characteristics and carrying out equal-proportion compression on each subsegment, and non-uniform segmented preprocessing: carrying out non-uniform segmented preprocessing on the compressed data, the method comprises the following steps of: determining segmentation points according to data characteristics, independently using a standard normal transformation preprocessing method for each sub-segment, finally preprocessing global data by using a smoothing method, and establishing a partial least square model: establishing the partial least square model by using the data subjected to adaptive segmentation compression and non-uniform segmentation preprocessing, according to the method, the data processing process can be optimized, the dimension of the spectral data can be quickly compressed, and the local features of the spectral data can be captured, so that the efficiency and accuracy of predicting the soil nutrients can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil sample analysis, and in particular to a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing. Background Art

[0002] As a non-destructive testing technology, near infrared spectroscopy has been widely used in soil nutrient prediction. Existing prediction methods are mainly divided into the following steps: data preprocessing: select a single or a combination of preprocessing methods (such as baseline correction, smoothing, etc.) to process the collected spectral data; feature extraction and selection: select feature extraction methods (such as principal component analysis, continuous projection algorithm, etc.) to reduce the dimension of preprocessed spectral data; establish a prediction model: select a suitable regression model (such as partial least squares regression, neural network, etc.) to model, analyze and predict the processed spectral data.

[0003] Near infrared spectroscopy technology faces problems such as low data processing efficiency, information redundancy, complex preprocessing process, and neglect of local data features, which affect the efficiency and accuracy of the prediction model. Specifically: 1. Low data processing efficiency: Existing technologies usually rely on high-dimensional features when processing near-infrared spectral data, which requires more computing resources. As the amount of data increases, the processing speed in practical applications will also decrease.

[0004] 2. Data redundancy and information loss: Feature extraction of near-infrared data can reduce the data dimension to a certain extent, but due to the high redundancy of the data, it will lead to a large information loss problem, which may lead to a decrease in prediction accuracy.

[0005] 3. Complex preprocessing process: Different preprocessing methods have uncertain effects on subsequent steps, and multiple experiments and adjustments are required based on data characteristics. At the same time, high-dimensional data will increase the complexity of the preprocessing process and affect prediction accuracy.

[0006] 4. Ignoring local data features: Data preprocessing and feature extraction are both aimed at global data, and fail to effectively capture subtle differences and changes in local areas of the data, which may reduce the accuracy of the prediction model. Summary of the invention

[0007] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0008] Therefore, the object of the present invention is to provide a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, which can optimize the data processing process, realize rapid compression of the spectral data dimension, capture the local characteristics of the spectral data, and thus improve the efficiency and accuracy of predicting soil nutrients.

[0009] To solve the above technical problems, according to one aspect of the present invention, the following technical solutions are provided: A soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, comprising the following steps: S1. Adaptive segmented compression: Segment according to data characteristics and compress each sub-segment proportionally; S2. Non-uniform segmented preprocessing: Perform non-uniform segmented preprocessing on the compressed data, determine the segmentation points according to data characteristics, use the standard normal transformation preprocessing method for each sub-segment separately, and finally use the smoothing method to preprocess the global data; S3. Establish a partial least squares model: Use the data after adaptive segmented compression and non-uniform segmented preprocessing to establish a partial least squares model for predicting soil nutrients.

[0010] As a preferred scheme of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing of the present invention, in step S1, the specific steps of segmenting according to data characteristics and compressing each sub-segment proportionally are as follows: Determine the adaptive spectral segmentation points: By calculating the average vector of the training samples at each wavelength point , and taking the extreme points of the vector as the segmentation points, the calculation formula is:

[0011] Among them, n is the number of training samples, is the value of the i-th sample at the j-th wavelength point; Calculate the adaptive segmented compression parameter: First, calculate the number of spectral points after compression through the formula , and then use the formula to calculate the number of spectral points after compression for each sub-segment ; Among them, is the compression parameter, which means compressing the spectral dimension to 1 / n of the original spectral dimension, rounding down the calculation result, setting the minimum value of the sub-segment to min = 2, and calculating the number of sub-segments with a spectral number of min as , and adjust the final value of each sub-segment according to the formula , and the adjustment rule is: ; Extract adaptive compression spectral points: Assume that the starting point of each sub-segment is , and the ending point is . According to , perform point compression. The points taken are , where the point-taking interval is and round down.

[0012] As a preferred embodiment of a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to the present invention, in step S2, perform non-uniform segmentation preprocessing on the compressed data, determine the segmentation points according to the data characteristics, use the standard normal transformation preprocessing method for each sub-segment separately, and finally use the smoothing method to preprocess the global data. The specific steps are as follows: Determine the non-uniform spectral preprocessing segmentation points: By calculating the average value vector of the samples at each wavelength point , and use its extreme points as the spectral segmentation points. The calculation formula is:

[0013] where n is the number of training samples, is the value of the i-th sample at the j-th wavelength point; Local spectral preprocessing: Perform standard normal transformation preprocessing on each sub-segment. The formula is:

[0014] where, is the reflectance of the sample, , are the mean and standard deviation of the spectral sub-segment of the sample; Then perform Savitzky-Golay smoothing preprocessing on the full spectrum. The formula is: .

[0015] where m is the window half-width, is the smoothing coefficient, is the original data point.

[0016] As a preferred embodiment of a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to the present invention, where the compression parameter has a value range of 1 < n_compression ≤ 5.

[0017] As a preferred embodiment of a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to the present invention, in the Savitzky-Golay smoothing preprocessing of the non-uniform segmentation preprocessing step, the parameter has a value range of ≤ m ≤ , where k is the order of the fitting polynomial, N is the total number of points, and the total window size 2m + 1 must be odd.

[0018] As a preferred embodiment of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to the present invention, wherein the method is applied to laboratory soil sample analysis and agricultural land soil sample analysis for predicting nutrient components such as nitrogen, phosphorus, and potassium in the soil.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: The method performs adaptive segmented compression before data preprocessing, segments according to data characteristics, and performs equal-proportion compression on each sub-segment, which can avoid the influence of the uncertainty of the preprocessing result on subsequent data dimensionality reduction. At the same time, it can simply and quickly reduce the data dimension and improve efficiency; perform non-uniform segmented preprocessing on the compressed data, determine the segmentation points according to data characteristics, use the standard normal transformation preprocessing method for each sub-segment separately, and finally use the smoothing method to preprocess the global data, which can reduce the calculation amount, save computing resources, and can better capture the local characteristics of the data, improve the sensitivity of the prediction model to details, and improve the accuracy of the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the drawings and specific embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them: Figure 1 is the flowchart of the adaptive segmented compression provided by the present invention; Figure 2 is the flowchart of the non-uniform segmented preprocessing provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings.

[0022] The present invention provides a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, which can optimize the data processing process, realize rapid compression of the spectral data dimension, capture the local characteristics of the spectral data, and thus improve the efficiency and accuracy of predicting soil nutrients.

[0023] The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing includes three steps: adaptive segmented compression, non-uniform segmented preprocessing, and establishing a partial least squares model.

[0024] Step 1, Adaptive Segmented Compression (ASC), as Figure 1 shown.

[0025] 1.1 Determine the adaptive spectral segmentation points: Calculate the average vector of the training samples at each wavelength point through formula (1) , and take the extreme points of this vector as the segmentation points.

[0026]

[0027] Among them, n is the number of training samples, is the value of the i-th sample at the j-th wavelength point.

[0028] 1.2 Calculate the adaptive segmented compression parameters: Calculate the number of spectral points after compression through formula (2) , and then use formula (3) to calculate the number of spectral points in each sub-segment .

[0029]

[0030]

[0031] Among them is the compression parameter, which means compressing the spectral dimension to 1 / n of the original spectral dimension. is the number of spectral points after compression in each sub-segment. The calculation result is rounded down, and the minimum value of the sub-segment is set to min = 2. Then calculate the number of sub-segments with a spectral number of min as .

[0032] Adjust the final value of each sub-segment according to formula (5).

[0033]

[0034]

[0035] 1.3 Extract the adaptive compressed spectral points: Assume that the starting point of each sub-segment is , and the ending point is . Take point compression according to formula (6).

[0036]

[0037] The points taken in formula (6) are , where the point-taking interval is and it is rounded down.

[0038] Step 2, Non-uniform segmented pretreatment (NUSP) is as follows Figure 2 shown.

[0039] 2.1 Determine the non-uniform spectral pretreatment segmentation points: Calculate the average vector of the samples at each wavelength point through formula (1) , and use its extreme points as the spectral segmentation points.

[0040] 2.2 Local spectral pretreatment: Perform standard normal transformation pretreatment on each sub-segment, as shown in formula (7), and then perform Savitzky-Golay (SG) smoothing pretreatment on the full spectrum, as shown in formula (8).

[0041]

[0042] where is the reflectivity of the sample, , are the mean and standard deviation of the spectral sub-segment of the sample.

[0043]

[0044] where m is the window half-width, is the smoothing coefficient, is the original data point.

[0045] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the disclosed embodiments of the present invention can be combined with each other in any way, and the exhaustive description of the situations of these combinations is omitted in this specification only for the consideration of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed in the text, but includes all technical solutions falling within the scope of the claims.

Claims

1. A soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, characterized in that: The steps include: S1, Adaptive segment compression: Segment the data according to its characteristics and compress each sub-segment in equal proportion; S2, non-uniform segmentation preprocessing: The compressed data is preprocessed by non-uniform segmentation, the segmentation points are determined according to the data characteristics, the standard normal transformation preprocessing method is used for each sub-segment separately, and finally the global data is preprocessed by smoothing method; S3. Establish partial least squares model: The partial least squares model was established using the data after adaptive segmentation compression and non-uniform segmentation preprocessing to predict soil nutrients.

2. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 1 is characterized in that: In step S1, the specific steps of segmenting data according to data features and compressing each sub-segment in equal proportion are as follows: Determine the adaptive spectral segmentation points: by calculating the average vector of the training samples at each wavelength point , and the extreme point of the vector is used as the segmentation point. The calculation formula is: ; Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength; Calculate the adaptive segment compression parameters: First use the formula Calculate the number of spectral points after compression, and then use the formula Calculate the number of spectral points after compression of each sub-segment ; in, is the compression parameter, which means compressing the spectrum dimension to 1 / n of the original spectrum dimension, rounding down the calculation result, setting the minimum value of the sub-segment to min=2, and calculating the number of sub-segments with the spectrum number min as , according to the formula Adjust the final value of each sub-segment. The adjustment rules are: ; Extract adaptive compressed spectrum points: Assume that the starting point of each sub-segment is , the end point is ,according to Point compression, the points taken , where the point interval is And round down.

3. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 1 is characterized in that: In step S2, the compressed data is preprocessed by non-uniform segmentation, the segmentation points are determined according to the data characteristics, the standard normal transformation preprocessing method is used for each sub-segment separately, and finally the global data is preprocessed by the smoothing method. The specific steps are as follows: Determine the non-uniform spectrum preprocessing segmentation points: by calculating the average vector of the sample at each wavelength point , and use its extreme point as the spectrum segmentation point. The calculation formula is: ; Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength; Local spectrum preprocessing: Perform standard normal transformation preprocessing on each sub-segment, the formula is: ; in, is the reflectance of the sample, , is the mean and standard deviation of the sample sub-segment spectrum; Then the full spectrum is preprocessed by Savitzky-Golay smoothing, the formula is: ; Where m is the half-width of the window, is the smoothing coefficient, is the original data point.

4. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 2 is characterized in that: The compression parameters The value range is 1< ≤5.

5. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 3 is characterized in that: In the Savitzky-Golay smoothing preprocessing of the non-uniform segmentation preprocessing step, the parameter The value range is ≤m≤ , where k is the order of the fitting polynomial, N is the total number of points, and the total window size 2m+1 must be an odd number.

6. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 1 is characterized in that: The method is applied to laboratory soil sample analysis and agricultural land soil sample analysis to predict nutrients such as nitrogen, phosphorus, potassium, etc. in the soil.

Citation Information

Patent Citations

  • Method for predicting soil properties based on multi-sensor fusion

    CN109669023A

  • Spectrometric analysis

    CN112557491A

  • Segmented preprocessing method for near infrared spectral data

    CN114062306A

  • Guanxinning quality detection method based on hyperspectral imaging technology and application

    CN116840110A

  • Estimation method of soil organic matter content

    CN117929325A