A soil nutrient prediction method based on high-dimensional spectral data compression preprocessing

The high-dimensional spectral data is processed through adaptive segmentation compression and non-uniform segmentation pretreatment, which solves the problems of low data processing efficiency and neglect of local features in near-infrared spectroscopy technology, and achieves high efficiency and high accuracy of soil nutrient prediction.

CN120148679BActive Publication Date: 2025-08-19CHANGCHUN UNIV OF SCI & TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510631562.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-19
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing near-infrared spectroscopy technology has problems such as low data processing efficiency, redundant information, complex preprocessing process and ignoring local characteristics of the data in soil nutrient prediction, which affects the efficiency and accuracy of the prediction model.

Method used

Adaptive segmentation compression and non-uniform segmentation pretreatment methods are used to process high-dimensional spectral data, and combined with standard normal transformation and smoothing methods, a partial least squares model is established for soil nutrient prediction.

Benefits of technology

Optimize the data processing process, quickly reduce the spectral data dimension, capture local data characteristics, and improve the efficiency and accuracy of soil nutrient prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148679B_ABST
    Figure CN120148679B_ABST
Patent Text Reader

Abstract

The present invention discloses a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing in the technical field of soil sample analysis, including adaptive segmentation compression: segmenting according to data characteristics and compressing each sub-segment proportionally, non-uniform segmentation preprocessing: performing non-uniform segmentation preprocessing on the compressed data, determining the segmentation points according to the data characteristics, using the standard normal transformation preprocessing method for each sub-segment separately, and finally using the smoothing method to preprocess the global data, and establishing a partial least squares model: using the data after adaptive segmentation compression and non-uniform segmentation preprocessing to establish a partial least squares model for soil nutrient prediction. The present invention can optimize the data processing process, realize rapid compression of spectral data dimensions, capture local characteristics of spectral data, and thus improve the efficiency and accuracy of predicting soil nutrients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil sample analysis, and in particular to a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing. Background Art

[0002] Near-infrared spectroscopy, a nondestructive testing technique, has been widely used in soil nutrient prediction. Existing prediction methods are primarily divided into the following steps: data preprocessing: selecting a single or a combination of preprocessing methods (such as baseline correction and smoothing) to process the collected spectral data; feature extraction and selection: selecting a feature extraction method (such as principal component analysis or successive projections) to perform dimensionality reduction on the preprocessed spectral data; and prediction model development: selecting an appropriate regression model (such as partial least squares regression or neural networks) to perform modeling, analysis, and prediction on the processed spectral data.

[0003] Near-infrared spectroscopy technology faces problems such as low data processing efficiency, information redundancy, complex preprocessing, and neglect of local data features, which affect the efficiency and accuracy of the prediction model. Specifically:

[0004] 1. Low data processing efficiency: Existing technologies for processing near-infrared spectral data typically rely on high-dimensional features, which require more computing resources. As the amount of data increases, the processing speed in practical applications also decreases.

[0005] 2. Data redundancy and information loss: Feature extraction of near-infrared data can reduce the data dimension to a certain extent, but due to the high redundancy of the data, it will lead to a large information loss problem, which may lead to a decrease in prediction accuracy.

[0006] 3. Complex preprocessing process: Different preprocessing methods have uncertain effects on subsequent steps, requiring multiple experiments and adjustments based on data characteristics. At the same time, high-dimensional data will also increase the complexity of the preprocessing process and affect prediction accuracy.

[0007] 4. Ignoring local data features: Data preprocessing and feature extraction are both targeted at global data, and fail to effectively capture subtle differences and changes in local areas of the data, which may reduce the accuracy of the prediction model. Summary of the Invention

[0008] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid blurring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.

[0009] Therefore, the purpose of the present invention is to provide a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, which can optimize the data processing process, achieve rapid compression of spectral data dimensions, capture local characteristics of spectral data, and thus improve the efficiency and accuracy of soil nutrient prediction.

[0010] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:

[0011] A soil nutrient prediction method based on high-dimensional spectral data compression preprocessing includes the following steps:

[0012] S1, Adaptive segment compression:

[0013] Segment the data according to its characteristics and compress each sub-segment proportionally;

[0014] S2. Non-uniform segmentation preprocessing:

[0015] The compressed data is preprocessed by non-uniform segmentation, the segmentation points are determined according to the data characteristics, each sub-segment is preprocessed using the standard normal transformation method, and finally the global data is preprocessed using the smoothing method;

[0016] S3. Establish a partial least squares model:

[0017] A partial least squares model was established using data that had undergone adaptive segmentation compression and non-uniform segmentation preprocessing to predict soil nutrients.

[0018] As a preferred embodiment of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing of the present invention, the specific steps of segmenting the data according to its characteristics and compressing each sub-segment proportionally in step S1 are as follows:

[0019] Determine the adaptive spectral segmentation points: by calculating the average vector of the training samples at each wavelength point , and use the extreme point of the vector as the segmentation point. The calculation formula is:

[0020]

[0021] Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength;

[0022] Calculate the adaptive segment compression parameters: first use the formula Calculate the number of spectral points after compression, and then use the formula Calculate the number of spectral points after compression of each sub-segment ;

[0023] in, is the compression parameter, which means compressing the spectrum dimension to 1 / n of the original spectrum dimension, rounding the calculation result down, setting the minimum value of the sub-segment to min=2, and calculating the number of sub-segments with the minimum spectrum number as , according to the formula Adjust the final value of each sub-segment. The adjustment rules are:

[0024] ;

[0025] Extract adaptive compressed spectrum points: Assume that the starting point of each sub-segment is , the end point is ,according to Point compression, the points , where the point interval is And round down.

[0026] As a preferred embodiment of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing described in the present invention, in step S2, the compressed data is subjected to non-uniform segmentation preprocessing, segmentation points are determined according to data characteristics, a standard normal transformation preprocessing method is applied to each sub-segment separately, and finally a smoothing method is used to preprocess the global data. The specific steps are as follows:

[0027] Determine the segmentation points for non-uniform spectrum preprocessing: by calculating the average vector of the sample at each wavelength point , and use its extreme value points as spectrum segmentation points. The calculation formula is:

[0028]

[0029] Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength;

[0030] Local spectrum preprocessing: Perform standard normal transformation preprocessing on each sub-segment. The formula is:

[0031]

[0032] in, is the reflectivity of the sample, 、 is the mean and standard deviation of the sample sub-segment spectrum;

[0033] Then the full spectrum is preprocessed with Savitzky-Golay smoothing, the formula is:

[0034] .

[0035] Where m is the half-width of the window, is the smoothing coefficient, is the original data point.

[0036] As a preferred solution of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing described in the present invention, the compression parameter The value range is 1 <n_compression≤5。

[0037] As a preferred embodiment of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing of the present invention, in the Savitzky-Golay smoothing preprocessing of the non-uniform segmentation preprocessing step, the parameter The value range is ≤m≤ , where k is the order of the fitting polynomial, N is the total number of points, and the total window size 2m+1 must be an odd number.

[0038] As a preferred embodiment of the soil nutrient prediction method based on high-dimensional spectral data compression preprocessing described in the present invention, the method is applied to laboratory soil sample analysis and agricultural land soil sample analysis to predict nutrients such as nitrogen, phosphorus, and potassium in the soil.

[0039] Compared with the prior art, the present invention has the following beneficial effects: the method performs adaptive segmentation compression before data preprocessing, segments the data according to data characteristics and compresses each sub-segment proportionally, which can avoid the impact of the uncertainty of the preprocessing results on subsequent data dimensionality reduction, and at the same time simply and quickly reduce the data dimension and improve efficiency; the compressed data is preprocessed in a non-uniform segmentation manner, the segmentation points are determined according to the data characteristics, and the standard normal transformation preprocessing method is used separately for each sub-segment. Finally, the global data is preprocessed using a smoothing method, which can reduce the amount of calculation, save computing resources, and better capture the local characteristics of the data, thereby improving the sensitivity of the prediction model to details and improving the accuracy of the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort. Among them:

[0041] Figure 1 A flowchart of the adaptive segmented compression provided by the present invention;

[0042] Figure 2 This is a flow chart of the non-uniform segmentation preprocessing provided by the present invention. DETAILED DESCRIPTION

[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0044] The present invention provides a soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, which can optimize the data processing process, realize rapid compression of spectral data dimensions, and capture local characteristics of spectral data, thereby improving the efficiency and accuracy of soil nutrient prediction.

[0045] The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing includes three steps: adaptive segmentation compression, non-uniform segmentation preprocessing, and establishing a partial least squares model.

[0046] Step 1: Adaptive segmented compression (ASC), such as Figure 1 shown.

[0047] 1.1 Determine the adaptive spectrum segmentation point: Calculate the average value vector of the training sample at each wavelength point using formula (1) , and use the extreme points of the vector as segmentation points.

[0048]

[0049] Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength.

[0050] 1.2 Calculation of adaptive segment compression parameters: Calculate the number of spectral points after compression using formula (2): , and then use formula (3) to calculate the number of spectral points in each sub-segment .

[0051]

[0052]

[0053] in is the compression parameter, which means compressing the spectral dimension to 1 / n of the original spectral dimension. The number of spectral points after compression for each sub-segment is rounded down, and the minimum value of the sub-segment is set to min=2. Then the number of sub-segments with the minimum spectral number is calculated as .

[0054] According to formula (5), adjust the final The value of .

[0055]

[0056]

[0057] 1.3 Extracting adaptive compressed spectrum points: Assume that the starting point of each sub-segment is , the end point is , take point compression according to formula (6).

[0058]

[0059] The point in formula (6) is , where the point interval is And round down.

[0060] Step 2: Non-uniform segmented pretreatment (NUSP) Figure 2 shown.

[0061] 2.1 Determine the segmentation points for non-uniform spectrum preprocessing: Calculate the average value vector of the sample at each wavelength point using formula (1): , and use its extreme points as spectrum segmentation points.

[0062] 2.2 Local spectrum preprocessing: Perform standard normal transformation preprocessing on each sub-segment, as shown in formula (7), and then perform Savitzky-Golay (SG) smoothing preprocessing on the full spectrum, as shown in formula (8).

[0063]

[0064] in is the reflectivity of the sample, 、 are the mean and standard deviation of the sample sub-segment spectrum.

[0065]

[0066] Where m is the half-width of the window, is the smoothing coefficient, is the original data point.

[0067] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A soil nutrient prediction method based on high-dimensional spectral data compression preprocessing, characterized in that: The steps include: S1, Adaptive segment compression: Segment the data according to its characteristics and compress each sub-segment proportionally; S2. Non-uniform segmentation preprocessing: The compressed data is preprocessed by non-uniform segmentation, the segmentation points are determined according to the data characteristics, each sub-segment is preprocessed using the standard normal transformation method, and finally the global data is preprocessed using the smoothing method; S3. Establish a partial least squares model: A partial least squares model is established using the data after adaptive segmentation compression and non-uniform segmentation preprocessing to predict soil nutrients. In step S1, the specific steps of segmenting according to data characteristics and compressing each sub-segment proportionally are as follows: Determine the adaptive spectral segmentation points: by calculating the average vector of the training samples at each wavelength point , and use the extreme point of the vector as the segmentation point. The calculation formula is: ; Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength; Calculate the adaptive segment compression parameters: first use the formula Calculate the number of spectral points after compression, and then use the formula Calculate the number of spectral points after compression of each sub-segment ; in, is the compression parameter, which means compressing the spectrum dimension to 1 / n of the original spectrum dimension, rounding the calculation result down, setting the minimum value of the sub-segment to min=2, and calculating the number of sub-segments with the minimum spectrum number as , according to the formula Adjust the final value of each sub-segment. The adjustment rules are: ; Extract adaptive compressed spectrum points: Assume that the starting point of each sub-segment is , the end point is ,according to Point compression, the points , where the point interval is And round down.

2. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 1, characterized in that: In step S2, the compressed data is preprocessed by non-uniform segmentation, the segmentation points are determined according to the data characteristics, the standard normal transformation preprocessing method is used for each sub-segment separately, and finally the global data is preprocessed by the smoothing method. The specific steps are as follows: Determine the segmentation points for non-uniform spectrum preprocessing: by calculating the average vector of the sample at each wavelength point , and use its extreme value points as spectrum segmentation points. The calculation formula is: ; Where n is the number of training samples, is the value of the i-th sample at the j-th wavelength; Local spectrum preprocessing: Perform standard normal transformation preprocessing on each sub-segment. The formula is: ; in, is the reflectivity of the sample, is the mean and standard deviation of the sample sub-segment spectrum; Then the full spectrum is preprocessed with Savitzky-Golay smoothing, the formula is: ; Where m is the half-width of the window, is the smoothing coefficient, is the original data point.

3. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 1, characterized in that: The compression parameters The value range is 1< ≤5.

4. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 2, characterized in that: In the Savitzky-Golay smoothing preprocessing of the non-uniform segmentation preprocessing step, the parameter The value range is , where k is the order of the fitting polynomial, N is the total number of points, and the total window size 2m+1 must be an odd number.

5. The soil nutrient prediction method based on high-dimensional spectral data compression preprocessing according to claim 1, characterized in that: The method is applied to the analysis of laboratory soil samples and agricultural land soil samples to predict the nitrogen, phosphorus and potassium nutrients in the soil.

Citation Information

Patent Citations

  • Segmented preprocessing method for near infrared spectral data

    CN114062306A