Sugar beet multispectral data dimension reduction processing method based on principal component analysis

Through the dual-path parallel mechanism and orthogonality verification of principal component analysis, the high-dimensional problem of beet multi-spectral data is solved, efficient dimensionality reduction of the data and the retention of key features is achieved, the efficiency and accuracy of beet growth monitoring is improved, and the rapid and reliable crop analysis is supported.

CN120296404AInactive Publication Date: 2025-07-11ZHANGYE ACAD OF AGRI SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510796272.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The high-dimensional characteristics of beet multispectral data lead to serious data redundancy and increased computational burden, which affects the feasibility of real-time analysis and the accuracy and robustness of subsequent machine learning models. Traditional dimensionality reduction methods cannot fully retain discriminant information, affecting the efficiency and accuracy of beet growth monitoring.

Method used

The dual-path parallel mechanism based on principal component analysis is adopted to extract principal components through standard principal component analysis and discriminant enhancement principal component analysis, and orthogonality verification and fusion are carried out. Combined with the dual screening of reconstruction error and classification accuracy, the kernel function parameter weight factor is dynamically adjusted to ensure the high-fidelity retention and dimensionality reduction effect of key features.

Benefits of technology

While reducing the data dimensions, the core structure of beet multi-spectral data is effectively retained, the analysis efficiency and accuracy are improved, the early disease detection rate and the classification accuracy of defective states are improved, and the reliable basis for precise fertilization and drug application is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296404A_ABST
    Figure CN120296404A_ABST
Patent Text Reader

Abstract

The invention discloses a beet multispectral data dimension reduction processing method based on principal component analysis, and relates to the technical field of agricultural information processing, and the method comprises the following steps: (a) carrying out centralized preprocessing on original beet multispectral data; (b) extracting principal components by adopting a double-path parallel mechanism: executing standard principal component analysis on a path I, and screening the principal components according to a variance contribution rate on the basis of a sample covariance matrix; according to the beet multispectral data dimension reduction processing method based on principal component analysis, through a double-path principal component fusion mechanism and orthogonal verification design, the inherent contradiction that global structure reservation and discriminant feature enhancement are difficult to consider in a traditional dimension reduction method is overcome while the dimension of beet multispectral data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural information processing, and in particular to a method for reducing the dimensionality of beet multispectral data based on principal component analysis. Background Art

[0002] In the practice of precision agriculture for beet cultivation, multispectral imaging technology has become a key tool for monitoring the growth status of beets. By capturing reflection information in multiple bands, it helps to evaluate the health level, nutritional status, and potential disease threats of beets. However, the multispectral data generated by this technology often has the characteristics of high dimensionality, containing a large amount of band information, resulting in serious data redundancy and high correlation, significantly increasing the computational burden and storage requirements. High-dimensional data not only makes the processing process slow, affecting the feasibility of real-time field analysis, but also reduces the accuracy and robustness of subsequent machine learning models such as classification or regression, because redundant information will obscure key features, making feature extraction and pattern recognition difficult. Traditional dimensionality reduction methods such as band screening or simple linear transformation can partially alleviate the dimensionality problem, but often cannot fully retain the discriminative information of the data, resulting in a decline in model performance and affecting the improvement of the efficiency and accuracy of beet growth monitoring. Therefore, there is an urgent need for an efficient data processing method that can reduce the dimensionality while maximizing the retention of the core structure of the original data to support fast and reliable beet growth analysis. Summary of the Invention

[0003] (I) Technical Problems to be Solved Aiming at the deficiencies of the prior art, the present invention provides a method for reducing the dimensionality of beet multispectral data based on principal component analysis, which solves the problem of how to effectively reduce the data dimensionality to eliminate redundancy while ensuring the high-fidelity retention of key features, thereby improving the analysis efficiency and accuracy of beet multispectral data processing.

[0004] (II) Technical Solutions To achieve the above objectives, the present invention is realized through the following technical solutions: A method for reducing the dimensionality of beet multispectral data based on principal component analysis, including the following steps: (a) Perform centering preprocessing on the original beet multispectral data; (b) Use a dual-path parallel mechanism to extract principal components: Path one performs standard principal component analysis, and screens principal components based on the variance contribution rate of the sample covariance matrix; Path two performs discriminant-enhanced principal component analysis, and solves the generalized eigenvalue problem in the non-linear feature space to extract discriminant principal components; (c) Perform orthogonality verification on the dual-path principal components, and perform principal component fusion when the absolute value of the dot product of the principal component vectors of path one and path two is less than the threshold; (d) Reconstruct the data based on the fused principal components and calculate the average reconstruction error at the band level. At the same time, input it into the classifier to verify the classification accuracy of the growth state; (e) When the error exceeds the noise threshold or the accuracy is lower than the reference value, adjust the kernel function parameter weight factor of the discriminative enhancement path and then re - execute steps (b) - (d).

[0005] In the entire technical process of dimensionality reduction processing of sugar beet multi - spectral data, first, perform centering pre - processing on the original spectral data. In this step, calculate the arithmetic mean of the reflectivities of all samples band - by - band, and subtract the average value of the corresponding band from the band reflectivity of each sample to eliminate the baseline drift of the device. When the absolute difference between the reflectivity of a specific band of a sample and the average value exceeds three times the standard deviation of that band, it is determined as an outlier and the root cause disposal is initiated: if it is caused by instantaneous cloud occlusion, use the median of the same - batch samples at adjacent spatial positions to replace it; if it is caused by sensor failure, eliminate the sample and record the missing mark. Implement differential compensation for different soil types: for high - reflective black soil plots, additionally subtract the moving average of the soil background reflectivity after centering; for low - reflective red soil plots, retain the original centering result. For data sets collected continuously for multiple days, calculate the band offset coefficient for subsequent dates based on the average value of each band of the first - day data, and perform time - series drift correction by multiplying the reciprocal of the offset coefficient and then execute standard centering.

[0006] After completing the preprocessing, it enters the dual-path principal component extraction stage. The standard principal component analysis path calculates the sample covariance matrix and solves the eigenvalue problem, and screens the principal components according to the variance contribution rate: starting from the component with the largest eigenvalue, the contribution rate is accumulated sequentially. When the preset target value is reached for the first time, the principal component serial number is recorded. If the difference in eigenvalues between adjacent components is less than a specific proportion of the median, continue to include them until the difference condition is met. Adjust the screening strategy dynamically according to the agricultural scenario: in the early disease monitoring scenario, increase the cumulative contribution rate threshold to retain weak abnormal features; in the nutrient assessment scenario, forcefully retain the principal components with the load of fertilizer-sensitive bands meeting the standard. The discriminative enhancement path calculates the cluster centers of each category based on the growth state labels, and dynamically assigns the kernel function weights according to the mutual information entropy value between bands: for strongly correlated band groups with an entropy value higher than 0.8, such as the red edge bands in the 780-850nm range, reduce their kernel parameter weights to 60% of the reference value; for weakly correlated bands with an entropy value lower than 0.3, such as the short-wave infrared band at 1550nm, increase the weight to 140%. When performing non-linear mapping using the radial basis kernel function, adjust the kernel range according to the growth period of sugar beet: use a wide kernel radius to tolerate large variations during the low leaf area index period, and switch to a narrow kernel radius to focus on fine changes during the high leaf area index period. Synchronously optimize the between-class and within-class dispersions in the mapped space: when the distance between the healthy and diseased class centers is insufficient, apply a gain to the reflectance of the diseased samples; when the number of samples lacking elements is scarce, generate simulation data based on the proximity method; split subgroups for high-dispersion groups; temporarily do not participate in the within-class dispersion calculation for abnormal groups caused by water stress. The discriminant principal component extraction preferentially selects the projection direction with the largest between-class and within-class dispersion ratio. When multiple ratios are similar, select the direction with a higher correlation with the key physiological bands.

[0007] Preferably, the discriminative enhanced principal component analysis is implemented through the following steps: Calculate the kernel function parameter weight factor according to the mutual information entropy between bands; Map the data to the non-linear space using the radial basis kernel function with the weight factor; Solve the discriminant principal component with the goal of maximizing the ratio of between-class dispersion to within-class dispersion.

[0008] Preferably, the orthogonality verification specifically includes: Perform Gram-Schmidt orthonormalization on the principal component sequence of path one; Perform Gram-Schmidt orthonormalization on the principal component sequence of path two; Calculate the absolute value of the dot product of the i-th principal component of path one and the j-th principal component of path two pairwise; When the absolute value of all dot products is less than the preset orthogonality threshold, it is determined that the principal component information of the two paths is complementary; among them, the preset orthogonality threshold is determined based on the statistical value of the correlation between bands of sugar beet multispectral data. That is: the preset orthogonality threshold is set by calculating the average value of the correlation coefficients of adjacent bands of historical sugar beet multispectral data.

[0009] Orthogonality verification is performed after the main components are output through two paths. The main component sequences of the two paths are respectively orthogonalized to eliminate internal redundancy, and then the absolute value of the dot product of the cross-path main components is calculated pair by pair. When all dot product values are less than the preset orthogonality threshold, it is determined that the information is complementary, and this threshold is determined based on the statistical value of the correlation between historical data bands. For special growth periods such as the tuber maturity period, due to the spectral convergence caused by leaf aging, the threshold is relaxed to a specific multiple of the standard value. When the verification fails, a hierarchical disposal is initiated: if some dot products exceed the standard, resampling is performed to retain the core physiological bands; if most dot products exceed the standard, it is determined that there is soil background interference and spectral unmixing is performed. When encountering bad weather, a compensation mechanism is enabled: for continuous rainy days, ratio correction using historical data of the same period is adopted; when the red edge slope abnormally decays under high temperature stress, the verification is suspended.

[0010] Preferably, the main component fusion is achieved through the following steps: Merge the main component sets after orthogonalization of the two paths; Normalize the variance contribution rate of the main components of path one and the discriminant contribution rate of the main components of path two into a comprehensive eigenvalue index; Sort in descending order according to the comprehensive eigenvalue index, and select the top M main components as the output.

[0011] After the orthogonality verification is passed, the main component fusion is implemented. Merge the orthogonalized component sets of the two paths and assign importance indicators: the components of the standard path are normalized according to the variance contribution rate, and the components of the discriminant path are normalized according to the separation significance of categories. During the regular growth period, weighted synthesis is performed using a fixed weight ratio, and the weight ratio is switched during the key turning period to strengthen the discriminant contribution. After sorting according to the comprehensive index, the number of basic components is determined, and the correlation coefficient between each of them and the core physiological parameters is checked one by one: the unqualified components are replaced by the first qualified substitute component; when there is no qualified substitute, the total number is reduced and rescreened. For plots with high soil reflectance, the weight of the soil characteristic components is forced to be reduced, and at the same time, the importance of the components related to the vegetation index is enhanced. Finally, the fused components are output and a feature traceability report is generated, marking the path sources and agronomic meanings of each component.

[0012] The dimensionality reduction results need to be verified by double screening. Reconstruction error verification stage: Reconstruct the original data using the fusion components and calculate the average absolute error of each band, and compare it with the dynamic threshold set based on the historical noise standard deviation. The threshold is relaxed for sandy arid plots; the standard deviation is calculated separately using the shadow area data for the greenhouse environment. Root cause diagnosis is initiated when the error exceeds the standard: continuous multi-band exceeding the standard triggers the sensor whiteboard correction; single-band exceeding the standard is marked as a potential agricultural event. In strong wind weather, the anti-wind noise mode is enabled, and only the leaf flip insensitive band is used for reconstruction; after rain, the leaf droplet samples are removed and re-verified. Classification accuracy verification stage: The dimensionality reduction data is input into the support vector machine classifier, and the overall accuracy is required to be no less than the preset margin value of the original data accuracy. The margin value is dynamically set according to the confidence probability of the classification decision. The disease category implements a zero missed judgment strategy; the deficiency category allows limited misjudgment but records the soil conductivity value. When the sample is unbalanced, the minority class simulation data is synthesized; the saline-alkali plot is forcibly included in the historical salt stress samples. Initiate low-light compensation migration verification under continuous rainy environment; focus on heat shock response band and relax health standards during high temperature and drought period.

[0013] Preferably, the reconstruction error verification in step (d) must satisfy: The band-level average reconstruction error does not exceed a predefined multiple of the standard deviation of the noise level of the historical data; wherein the predefined multiple is determined by the noise distribution confidence interval of the sugar beet multi-spectral data.

[0014] The predefined multiples are determined by the following steps: a) Calculate the standard deviation σ of noise in each band of historical data; b) The probability P that the statistical noise value falls within the interval [μ-1.5σ, μ+1.5σ]; c) When P ≥ 0.85, the multiple k is taken as 1.5, and k is preferably 1.2 during implementation.

[0015] Preferably, the classification accuracy verification in step (d) needs to meet the following requirements: the classification accuracy is not lower than a preset accuracy margin value of the original data classification accuracy; wherein the preset accuracy margin value is determined by the decision confidence probability of the beet growth status classification model.

[0016] The classification accuracy verification in step (d) adopts a support vector machine, and the verification objects include three types of growth states: healthy, nutrient-deficient, and diseased.

[0017] Preferably, the adjustment of the kernel function parameter weight factor in step (e) adopts a gradient descent method, and the optimization goal is to minimize the weighted sum of the reconstruction error and the classification accuracy deviation.

[0018] If any verification fails to meet the standard, adaptive optimization will be triggered. The gradient descent method is used to adjust the kernel function parameter weight factor of the discrimination path, and the initial step size is dynamically constrained according to the growth period: large adjustments are allowed in the early growth period, and fine adjustments are limited in the late growth period. The optimization goal is the weighted sum of the reconstruction error and the precision deviation, and a directional response is implemented: only when the reconstruction error exceeds the standard, the weight of the high mutual information entropy band is adjusted; only when the precision is insufficient, the discrete weight of the adjacent category is strengthened. The parameters of the red edge band are frozen in high temperature weather; only the weight of the water-insensitive band is allowed to be adjusted in rainy environment. When soil characteristic pollution is detected in high organic matter black soil plots, the relevant weights are forced to return to zero; a special weight channel for ion poisoning is added to salinized plots. The dual indicators are re-tested immediately after each optimization, and manual consultation is initiated if the two iterations fail to meet the standard. The optimization process is embedded with the principle of agronomic indicator priority termination: stop immediately when the prediction error of physiological parameters is significantly improved; and roll back parameters when the leaf spectrum is distorted. If the moisture band error is still high after the first iteration of sandy dry land, the second iteration is prohibited; the special discrimination model is switched for heavy clay and waterlogged plots. High temperature or rainfall interference triggers the meteorological fuse, delaying restart until the environment stabilizes.

[0019] Preferably, the number of iterations of the gradient descent method does not exceed two.

[0020] Preferably, the classifier is a linear kernel support vector machine. The classifier is constructed using a linear kernel support vector machine, and the classification boundary parameters are dynamically configured according to the growth period. High weights are assigned to chlorophyll-sensitive bands, and the weights of soil background bands are reset to zero. When the sample is unbalanced, cost-sensitive learning is implemented and weak-label simulation samples are generated. A special feature channel for sodium ion poisoning is added to monitor saline-alkali land; a moisture-resistant classifier is constructed in a rainy environment and supplemented by a historical sunny day model reference; the high temperature period relies on the coordinated discrimination of visible light and canopy temperature. When the model is running, the classification confidence is monitored in real time, and continuous low-confidence samples trigger spectral re-collection; saline-alkali associated misjudged samples are corrected in combination with soil available nutrient data. During monthly incremental training, the core feature weights are frozen and only the newly added boundaries are adjusted.

[0021] Preferably, in the discriminant enhanced principal component analysis, the between-class scatter is calculated using the Mahalanobis distance, and the within-class scatter is calculated using the trace of the covariance matrix. In the scatter calculation process, the between-class scatter is measured using the Mahalanobis distance: the combined covariance matrix of the entire plot is used in the early growth stage, and independent models are built for different soil types in the late growth stage. When the class boundary is fuzzy, virtual samples are introduced to forcibly expand the distance. The within-class scatter is measured by the trace of the covariance matrix: for saline-alkali plots, samples with sodium interference are isolated for separate calculation; for abnormal canopy samples in high-temperature weather, they are frozen; after rain, a short-wave infrared correction factor is enabled to compress the scatter value. When the within-class scatter suddenly increases, root cause diagnosis is performed: soil moisture fluctuations trigger water compensation; red edge shift is regarded as a warning of physiological disorders. For severely saline-alkali plots, an ion toxicity scatter channel is added, and its weight increases with the increase of the soil sodium-potassium ratio. Before the scatter data is output, verify its correlation coefficient with agronomic parameters, and re-sampling is triggered when it exceeds the standard; check the consistency of field phenotypes and associate with UAV images for re-verification. When updating the scatter reference library monthly, a sliding window mechanism is used to only retain recent valid data.

[0022] (III) Beneficial Effects The present invention provides a method for reducing the dimensionality of sugar beet multispectral data based on principal component analysis. It has the following beneficial effects: (I) The method for reducing the dimensionality of sugar beet multispectral data based on principal component analysis, through a dual-path principal component fusion mechanism and orthogonal verification design, while reducing the dimensionality of sugar beet multispectral data, overcomes the inherent contradiction in traditional dimensionality reduction methods that it is difficult to balance the retention of global structure and the enhancement of discriminant features. Among them, the standard principal component analysis path ensures the high-fidelity compression of the main information of spectral variation, the discriminant enhancement path specifically extracts weak agronomic features such as diseases and nutrient deficiencies, and the orthogonality verification proves the complementarity of the dual-path information from a mathematical level, making the feature set after dimensionality reduction have both comprehensiveness and specificity. The double-screening closed-loop of reconstruction error and classification accuracy further guarantees the output quality: the dynamic noise threshold combined with the agricultural situation adaptive mechanism reduces the reconstruction error; the accuracy margin control based on decision confidence maintains a small loss of classification accuracy under complex working conditions such as saline-alkali interference and rainy environments.

[0023] (II) The method for reducing the dimensionality of sugar beet multispectral data based on principal component analysis, by deeply embedding the agronomic logic of the entire growth period of sugar beet, transforms the abstract mathematical process into an executable agricultural situation operation chain. The parameter adjustment rules adaptive to the growth period make the method have field universality and break through the requirements of traditional algorithms for a stable environment; the abnormal handling mechanism of soil-meteorological coupling improves the robustness under harsh farmland conditions, and the data availability rate is improved compared with traditional methods. Finally, a qualitative change in agricultural analysis efficiency is achieved, the early disease detection rate is also improved, the classification accuracy of nutrient deficiency status is improved, and the prediction error of root sugar content is controlled, providing a reliable basis for precise fertilization and spraying. Description of the Drawings

[0024] Figure 1 It is a schematic flow diagram of the whole invention; Figure 2 It is a schematic framework diagram of the invention. Specific embodiments

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0026] Please refer to Figure 1 and Figure 2 , the present invention provides a technical solution: a method for reducing the dimensionality of beet multispectral data based on principal component analysis, including the following steps: (a) Perform centering preprocessing on the original beet multispectral data; it should be further noted that in the specific implementation process, it includes data standardization correction, outlier adaptive processing, plot environment compensation, and time dimension normalization. Among them, data standardization correction includes: for the original band reflectance data collected by the beet multispectral imaging device, calculate the arithmetic mean of the reflectance of all sample points band by band; subtract the average value corresponding to each band from the reflectance value of each sample point in each band to eliminate the systematic deviation caused by the difference in light intensity or sensor baseline drift, and align the data distribution center to the zero value reference.

[0027] Outlier adaptive processing includes: performing outlier adaptive processing based on a judgment logic. Among them, the judgment logic includes: if the absolute value of the difference between the reflectance value of a specific band of a sample point and the average value of that band exceeds three times the standard deviation of that band, it is determined as an outlier.

[0028] When the outlier is caused by instantaneous cloud occlusion, replace it with the median of the corresponding band reflectance of the same batch of beet samples at adjacent spatial positions; when the outlier is caused by a sensor failure, directly remove the sample and record the missing mark, where sensor failure includes more than 10 consecutive sample points with the same band being abnormal.

[0029] Plot environment compensation includes: for the soil background reflection interference of different beet planting plots, that is: if the soil type is highly reflective black soil, that is: the reflectance is greater than 0.3, subtract the moving average of the soil background reflectance additionally after centering; if the soil type is low-reflective red soil, that is: the reflectance does not exceed 0.3, retain the original centering result without compensation to avoid losing effective signals due to overcorrection.

[0030] The time - dimension normalization process includes: for the dataset collected for consecutive days of the same plot, that is, taking the average value of each band of the data on the first day as the benchmark, calculating the band offset coefficients of the data for subsequent days; for the data on non - first days, first multiplying by the reciprocal of the offset coefficient for rough correction, and then performing the standard centering operation to eliminate the temporal drift error introduced by the daily environmental changes.

[0031] (b)Adopt a dual - path parallel mechanism to extract the principal components: Path 1 performs standard principal component analysis, and screens the principal components based on the variance contribution rate of the sample covariance matrix; Path 2 performs discriminant - enhanced principal component analysis, and solves the generalized eigenvalue problem in the non - linear feature space to extract the discriminant principal components; In step (b), the process of Path 1 performing standard principal component analysis and screening the principal components based on the variance contribution rate of the sample covariance matrix includes: The implementation process of the standard principal component analysis path includes: First, based on the pre - processed beet multi - spectral dataset after centering, calculate the reflectance covariance matrix of all sample points in each spectral band, which characterizes the variation correlation strength between different bands; by solving the eigenvalue problem of this covariance matrix, obtain the sequence of eigenvectors arranged in descending order of eigenvalues, each eigenvector corresponding to a principal component direction, and the corresponding eigenvalue reflecting the data variation amount in this direction.

[0032] The implementation process of screening the principal components for the variance contribution rate includes the following logic: Cumulative contribution rate threshold determination: Start accumulating the variance contribution rate sequentially from the principal component corresponding to the largest eigenvalue; when the cumulative contribution rate first reaches the preset target value, record the current principal component serial number K; if the difference between the eigenvalues of adjacent principal components is less than 5% of the median eigenvalue, then continue to include the next principal component until the difference condition is met to avoid losing weakly - related discriminant information.

[0033] Adaptive adjustment for agricultural scenarios: Early disease monitoring scenario: When the proportion of disease samples in historical data is less than 10%, raise the cumulative contribution rate threshold to 98% to ensure that weak abnormal spectral features are not filtered; Nutritional status assessment scenario: If the goal is nutrient - deficiency classification, and the nutrient - deficiency classification includes nitrogen, phosphorus, and potassium deficiencies, only retain the principal components whose absolute value of the load of the fertilizer - sensitive bands in the eigenvector is greater than 0.3 to suppress the interference of irrelevant bands.

[0034] Verification of principal component effectiveness: For the first K principal components selected, check the correlation coefficients with the key physiological indicators of beets item by item: Case 1: If the correlation coefficient of the principal component with the SPAD value of chlorophyll content or biomass is less than 0.4, it is determined as an environmental noise component and excluded; Case 2: If the cumulative contribution rate is less than 92% after exclusion, then supplement the subsequent principal components according to the eigenvalue size until the requirement is met.

[0035] (c) Verify the orthogonality of the dual-path principal components. When the absolute value of the dot product of the principal component vectors of Path 1 and Path 2 is less than the threshold, perform principal component fusion; (d) Reconstruct the data based on the fused principal components and calculate the average reconstruction error at the band level. At the same time, input the classifier to verify the classification accuracy of the growth state; (e) When the error exceeds the noise threshold or the accuracy is lower than the reference value, adjust the kernel function parameter weight factor of the discriminative enhancement path and then re-execute steps (b)-(d).

[0036] Discriminative enhanced principal component analysis is implemented through the following steps: Calculate the kernel function parameter weight factor according to the mutual information entropy between bands; Map the data to the non-linear space using the radial basis kernel function with weight factors; Solve the discriminant principal components with the goal of maximizing the ratio of between-class scatter to within-class scatter.

[0037] It should be further noted that in the specific implementation process, the implementation content of discriminative enhanced principal component analysis includes: First, based on the prior labels of the sugar beet growth state, the prior labels include healthy, nutrient deficiency, and disease, calculate the clustering centers of each category of samples in the multi-spectral band space; By dynamically analyzing the mutual information entropy values between bands, determine the weight allocation ratio of the kernel function parameters, specifically: for strongly correlated band groups with entropy values higher than 0.8, such as the red edge bands of 780-850 nm, reduce their kernel parameter weights to 60% of the reference value to suppress the amplification of redundant information; for weakly correlated bands with entropy values lower than 0.3, such as the short-wave infrared of 1550 nm, increase the weight to 140% to enhance the independent discriminant signal.

[0038] When performing non-linear mapping using the radial basis kernel function, adjust the kernel action range according to the growth period of the sugar beet: From the seedling stage to the stage of closing the ridges, that is, the leaf area index is less than 3: adopt a wide kernel radius, scaling factor σ = 2.5, and tolerate large spectral variations; During the tuber bulking stage, that is, the leaf area index is not less than 3: switch to a narrow kernel radius, scaling factor σ = 1.0, and focus on subtle physiological changes.

[0039] In the mapped high-dimensional space, synchronously optimize the between-class scatter and within-class scatter: Strengthen between-class scatter: Calculate the Mahalanobis distance between the healthy and diseased class centers. If the distance is less than 1.3 times the historical minimum value, apply a 1.05-fold gain to the spectral reflectance of the diseased samples to artificially expand the class differences; When the number of samples in the nutrient deficiency category is less than 15% of the total number, enable the synthetic minority over-sampling technique to generate simulated nutrient deficiency spectral data based on the nearest neighbor method.

[0040] Intra-class discrete suppression: Dynamically group the spectral curves of samples within the same class. If the standard deviation within the group exceeds 1.8 times the standard deviation of the entire class, it is split into sub-groups. For the abnormally high discrete group caused by water stress, temporarily remove the data of this group from participating in the intra-class discrete calculation, where the abnormally high discrete group is the soil water content less than 40%.

[0041] The extraction of discriminant principal components follows the generalized eigenvalue solution rule: preferentially select the projection direction that maximizes the between-class and within-class discrete ratio. When there are multiple ratios that are close, that is, the difference is less than 5%, select the direction with a higher correlation with the chlorophyll-sensitive band of 720nm.

[0042] This implementation content transforms the complex discriminant enhancement process into a stable intelligent analysis process through a multi-level agricultural situation adaptation strategy, providing high-value discriminant feature inputs for the dual-path fusion.

[0043] The orthogonality verification specifically includes: Perform Gram-Schmidt orthogonalization on the principal component sequence of Path 1; Perform Gram-Schmidt orthogonalization on the principal component sequence of Path 2; Calculate the absolute value of the dot product of the i-th principal component of Path 1 and the j-th principal component of Path 2 pair by pair; When all absolute values of dot products are less than the preset orthogonality threshold, it is determined that the principal component information of the two paths is complementary; among them, the preset orthogonality threshold is determined based on the inter-band correlation statistical value of the sugar beet multi-spectral data.

[0044] Among them, the preset orthogonality threshold is set by calculating the average value of the correlation coefficients of adjacent bands of historical sugar beet multi-spectral data, and 0.05 is taken in the specific implementation, that is: when all absolute values of dot products are less than 0.05, it is determined that the information is complementary.

[0045] It should be further noted that in the specific implementation process, the content of the orthogonality verification includes the following: for the principal component sequence output by the standard principal component analysis path, perform Gram-Schmidt orthogonalization processing vector by vector to ensure that the directions of the principal components within Path 1 are perpendicular to each other, eliminating component redundancy caused by the high correlation of sugar beet spectral bands; synchronously perform the same operation on the principal component sequence output by the discriminant enhancement path to force the components to satisfy the orthogonality constraint and avoid self-overlap of category discriminant information.

[0046] Based on the orthogonality of the dual-path principal component space, carry out cross-path dot product verification: Complementary quantification judgment: successively select the i-th principal component vector of Path 1 and the j-th principal component vector of Path 2, and calculate the absolute value of their dot product; when the absolute values of the dot products of all combinations are less than the preset orthogonal threshold (determined based on the statistical value of the correlation between historical data bands), it is determined that there is a strict orthogonal relationship between the principal component spaces of the two paths; Typical agricultural situation special case: If the sugar beet is in the root maturity stage (growth days > 120 days), due to the aging of the leaves resulting in similar spectral characteristics, the orthogonal threshold is relaxed to 1.3 times the standard value to avoid information loss of the path caused by overly strict verification.

[0047] Verification failure handling mechanism: Case 1: When 30% - 50% of the dot product values exceed the standard, start band resampling: only retain the core bands with a correlation coefficient greater than 0.6 with chlorophyll content, and re-execute the dual-path analysis; Case 2: When more than 50% of the dot product values exceed the standard, it is determined as soil background interference, that is: the proportion of soil reflectivity is greater than 40%, and trigger the background subtraction process: collect the spectrum of the bare soil area as the background; perform soil linear spectral unmixing on the reflectivity of each band in the sugar beet area; recalculate the principal components with the unmixed pure vegetation spectrum.

[0048] Adaptability to extreme weather: In case of continuous rainy weather, that is: when the effective light intensity is less than 200 W / m²: If the fluctuation range of the dot product verification exceeds plus or minus 15%, enable the compensation of the cloudy mode spectral library; based on the healthy sugar beet spectrum in the same historical period, perform reflectivity ratio correction on the current data.

[0049] When high temperature stress above 35°C causes spectral distortion: detect the change rate of the slope of the red edge band; when the slope abnormally decreases by more than 20%, suspend the orthogonal verification and mark it as an invalid data batch.

[0050] After the orthogonal verification passes, generate a dual-path complementarity report, record the global variation characteristics retained by Path 1 and the local discriminant characteristics enhanced by Path 2, and provide a quantitative basis for subsequent principal component fusion.

[0051] This implementation content transforms the abstract orthogonal determination into an executable agricultural situation operation specification through an environment-aware adaptive verification strategy, providing technical support for the core innovation point of "dual-path complementary fusion".

[0052] Principal component fusion is achieved through the following steps: Merge the principal component sets after orthogonalization of the two paths; Normalize the variance contribution rate of the principal components of Path 1 and the discriminant contribution rate of the principal components of Path 2 into a comprehensive eigenvalue index; Sort in descending order according to the comprehensive eigenvalue index, and select the top M principal components as the output.

[0053] It should be further noted that in the specific implementation process, the implementation content of principal component fusion includes the following processes: After the orthogonality verification passes, the set of orthogonal principal components output by the standard principal component analysis path and the set of orthogonal principal components output by the discriminant enhancement path are merged to form a comprehensive principal component pool. A quantitative importance index is assigned to each principal component: For the principal components of Path 1, its index value is the normalized variance contribution rate, that is, the percentage of the eigenvalue of this component in the sum of the eigenvalues of all components; for the principal components of Path 2, its index value is the normalized discriminant contribution rate, that is, by calculating the significance score of this component in class separation and converting it into a relative value of 0% to 100%.

[0054] Perform weighted synthesis of importance indicators: During the regular growth period, that is, from the seedling stage to the sugar accumulation stage, during this period, a fixed weight ratio of 7:3 is adopted, that is, the comprehensive eigenvalue index = 70% × variance contribution rate + 30% × discriminant contribution rate.

[0055] During the key transition period, that is, the initiation stage of root tuber swelling and the latent stage of diseases, during this period, the weight ratio is dynamically switched to 3:7 to strengthen the weight of the discriminant contribution rate and ensure that weak abnormal features are not submerged.

[0056] After sorting in descending order according to the comprehensive eigenvalue index, principal component optimization is implemented. Among them, the determination of the basic quantity includes: According to the total number of original data bands N, set the initial number of fused components M = ceil(0.3×N).

[0057] Perform agronomic effectiveness screening, including: checking the correlation coefficients between the first M components and the core physiological parameters of sugar beet one by one, where the physiological parameters include chlorophyll and canopy nitrogen content; The process of checking the correlation coefficients between the first M components and the core physiological parameters of sugar beet one by one includes: If the absolute value of the component correlation coefficient < 0.4: Mark it as an inefficient component and trigger the substitution mechanism, that is, select the first component with a correlation coefficient ≥ 0.5 from the subsequent components to replace it; If all substitute components do not meet the standards: Decrease the value of M by 1 and rescreen until the requirements are met.

[0058] When it is detected that the proportion of soil reflectance > 40%, for the components with soil characteristic band load > 0.6 in the principal components of Path 1, the weight is forced to be reduced by 50%; increase the discriminant contribution rate of the vegetation index-related components in Path 2 by 20%. Finally, the fused M principal components are output, and a feature traceability report is generated synchronously: Mark the source path of each component and its main agronomic significance represented for subsequent growth analysis reference.

[0059] This implementation content transforms mathematical features into operable agronomic knowledge through a dynamic fusion strategy of agricultural perception, supporting the efficient execution of dual screening.

[0060] The reconstruction error verification in step (d) shall meet the following requirements: The average reconstruction error at the band level does not exceed a predefined multiple of the standard deviation of the historical data noise level; where the predefined multiple is determined by the confidence interval of the noise distribution of the sugar beet multispectral data.

[0061] The average reconstruction error at the band level does not exceed 1.2 times the standard deviation of the historical data noise level.

[0062] The predefined multiple is determined through the following steps: a) Calculate the standard deviation σ of the noise of each band in the historical data; b) Statistically calculate the probability P that the noise value falls within the interval [μ - 1.5σ, μ + 1.5σ]; c) When P ≥ 0.85, take the multiple k = 1.5, and preferably k = 1.2 during implementation; By dynamically setting the error threshold through the noise distribution confidence interval, compared with the fixed multiple scheme, the tolerance of the reconstruction error is improved and the classification false positive rate is reduced in the detection of sugar beet nutrient deficiency status.

[0063] It should be further noted that in the specific implementation process, the content of the reconstruction error verification includes: performing reconstruction calculations on the original sugar beet multispectral data using the principal components after fusion and dimensionality reduction, calculating the absolute error between the reflectance reconstruction value and the original value band by band; taking the arithmetic mean of the errors of all samples in the same band to obtain the average reconstruction error of that band. Compare the average error of each band with the preset noise threshold: The setting of the preset noise threshold includes: calling the standard deviation σ of the noise in the historical data of the same plot and the same growth period, and taking the predefined multiple k·σ as the threshold, where k is determined by the confidence interval of the noise distribution; When the soil water content is lower than 40%, due to surface fissures resulting in abnormal spectral dispersion, the threshold is relaxed to 1.5 times the standard value; If it is detected that there is interference from the supplementary light, that is, when the reflectance mutation in a specific band > 30%, calculate σ separately using the data in the shadow area.

[0064] Subsequently, perform root cause diagnosis of the error exceeding the standard. First, identify systematic biases, including: if the errors of more than 3 consecutive bands exceed the standard, it is determined that the sensor calibration has drifted, and trigger the whiteboard calibration process. The whiteboard calibration process includes: collecting the reflectance data of the standard whiteboard on the same day; performing gain compensation on the original data. The process includes: correction value = original value × theoretical whiteboard value, correction value = original value × measured whiteboard value; when only a single band exceeds the standard, it is diagnosed as a crop physiological abnormality, retain the error and mark it as a potential agricultural situation event.

[0065] During the response to meteorological interference, when encountering strong wind weather with a wind speed greater than 8 m / s, it may cause the variation of blade attitude. During this period, calculate the canopy texture index CTI. If the CTI fluctuation is greater than 15%, enable the anti-wind noise mode, that is: replace the original value of the day with the average value under the same lighting conditions in the past three days; only use the bands that are insensitive to blade flipping during reconstruction.

[0066] During the period when water droplets adhere to the blades after rainfall, detect the samples with a reduction in short-wave infrared reflectance greater than 20%; after removing the affected samples, re-verify the error.

[0067] Finally, generate an error traceability report: mark the exceeded bands, meteorological correlation degree and equipment status. When the overall compliance rate is not less than 95%, it is determined that the verification is passed. For the marked physiological abnormal samples, automatically push them to the agronomist terminal for review.

[0068] This implementation content transforms the abstract reconstruction verification into a technical gate for ensuring data fidelity through a multi-factor coupled error control strategy, providing reliable input for subsequent classification accuracy verification.

[0069] The classification accuracy verification in step (d) needs to meet: the classification accuracy is not less than the preset accuracy margin value of the classification accuracy of the original data; among them, the preset accuracy margin value is determined by the decision confidence probability of the sugar beet growth state classification model.

[0070] The preset accuracy margin value is dynamically set according to the following rules: a) Calculate the average decision confidence probability P of the original data in the SVM classifier avg ; b) When P avg ≥0.95, set the preset accuracy margin value δ = 2%, corresponding to the 98% benchmark during implementation; c) When 0.90 ≤ P avg <0.95, set the preset accuracy margin value δ = 5%.

[0071] The classification accuracy verification in step (d) uses a support vector machine, and the verification objects include three growth states: healthy, nutrient deficiency, and disease.

[0072] It should be further noted that in the specific implementation process, the implementation content of the classification accuracy verification includes: inputting the dimension-reduced data set into the preset support vector machine classifier, and performing the three-category discrimination of the sugar beet growth state: healthy, nutrient deficiency, and disease. The accuracy verification needs to meet: the overall accuracy rate of the classification results is not less than the preset accuracy margin value of the classification accuracy of the original high-dimensional data. This margin value is dynamically determined by analyzing the decision confidence probability of the classification model in historical data. When the confidence probability is greater than 0.95, the margin value is set to 2%, and when the confidence probability is in the range of 0.90 - 0.95, the margin value is set to 5%.

[0073] The implementation process embeds a three - layer agricultural condition logic control, including the following: Category - specific fault - tolerance mechanism: Implement a zero - tolerance policy for misjudgments of disease categories, that is, the misjudgment rate must be ≤ 1%. Because the spectral characteristics of early - stage diseases are weak, once the recognition confidence is lower than 0.85, an artificial review process is triggered; For nutrient - deficiency categories, a maximum misjudgment tolerance of 10% is allowed, but the soil EC value of misjudged samples needs to be recorded. When EC > 4 dS / m, it is automatically attributed to saline - alkali interference.

[0074] Sample representativeness verification: If the proportion of healthy samples in the plot exceeds 90%, the synthetic minority over - sampling technique (SMOTE) is enabled: Generate simulated nutrient - deficiency and disease spectral data in the feature space based on Mahalanobis distance to ensure class balance in the validation set; For saline - alkali plots with pH > 8.5, 10% of historical salt - stress samples are compulsorily included to prevent the classifier from ignoring ion - toxicity characteristics.

[0075] Meteorological coupling verification: When there are consecutive rainy days with cumulative sunshine < 5 hours / day, the low - illumination compensation mode is activated, that is: Use data from the same growth period on sunny days to train an auxiliary classifier and perform transfer verification on rainy - day data; When the transfer - verification accuracy is lower than the benchmark value by 3%, the current verification result is postponed and marked as "meteorologically invalid".

[0076] When the temperature is greater than 35°C for 3 consecutive days, that is, during the high - temperature and drought period, during which: Detect the transpiration stress index (TSI) of sugar beets. If TSI > 0.6: The verification of disease categories only focuses on the heat - shock response band; The health standard is relaxed to include the mild wilting state.

[0077] Finally, a classification confidence report is generated, including the recall rate of each category, soil interference factors, and meteorological impact coefficients. When the overall accuracy margin meets the standard and the disease misjudgment rate is qualified, the verification is determined to pass. For saline - alkali - related misjudged samples, soil improvement suggestions are automatically pushed to the farm management system.

[0078] This implementation content uses an intelligent verification strategy of multi - factor coupling to transform the classification - accuracy requirements into executable agricultural - condition operation standards, providing a reliable agronomic application guarantee for the dimensionality - reduction results.

[0079] In step (e), the gradient - descent method is used to adjust the weight factors of the kernel - function parameters, and the optimization objective is to minimize the weighted sum of the reconstruction error and the classification - accuracy deviation.

[0080] It should be further noted that in the specific implementation process, the adaptive optimization implementation content includes: when the reconstruction error verification or classification accuracy verification fails to meet the standard, start the parameter adjustment process driven by the gradient descent method, including: taking the weighted sum of the reconstruction error and the classification accuracy deviation as the optimization goal, and iteratively correcting the kernel function parameter weight factor of the discriminative enhancement path. The initial weight adjustment step size is set to 0.1, and the adjustment amplitude is dynamically constrained according to the growth period of sugar beet. A large adjustment of plus or minus 0.3 is allowed during the seedling stage, and the adjustment is restricted to a fine-tuning of plus or minus 0.05 during the mature stage to avoid oscillation.

[0081] Implement a three-layer optimization logic control, including the following processes: Error source directional response: If only the reconstruction error exceeds the standard, focus on the mutual information entropy weight factor of the bands, and adjust it inversely according to the entropy value of the bands with exceeded error, that is: for the bands with entropy value greater than 0.8, reduce the weight step size by multiplying by 2, and for the bands with entropy value less than 0.3, increase the step size by multiplying by 1.5; If only the classification accuracy is insufficient, strengthen the between-class scatter weight, and apply an additional 0.15-fold gain compensation to the class pairs with the center distance between healthy and diseased classes less than 1.2 times the historical average.

[0082] Meteorological locking strategy: When encountering continuous high temperature with temperature greater than 35°C, freeze the parameter adjustment of the red edge band because high temperature may cause red edge displacement artifacts; concentrate the optimization resources on the visible light band, which is less affected by temperature.

[0083] When encountering a rainy period with cumulative sunshine less than 100 W / m²·d, enable the anti-humidity noise mode: only allow the adjustment of the weights of the water-insensitive bands; impose an adjustment ban on the short-wave infrared band to avoid the amplification of water droplet interference.

[0084] Soil coupling intervention, including: when the OM of the high-organic-matter black soil is greater than 5% and the soil carbon characteristics are detected to be mixed into the discrimination path, force the weights of the soil characteristic bands to zero; Trigger the organic matter compensation algorithm, that is: correct the spectral reflectance based on the measured value of soil carbon content.

[0085] For saline-alkali land EC, that is: increase the special weight channel for ion toxicity: set an independent adjustment factor for the sodium-sensitive band; when the sodium ion concentration > 1000 ppm, the step size of this channel is amplified to 3 times the standard value.

[0086] Immediately re-verify the two indicators after each optimization iteration. If the standard is still not met after two iterations, start the agricultural society consultation mechanism, including: pause the automated process, push the spectral feature anomaly report to the agronomy experts, and determine whether it belongs to real physiological disorders in combination with field sampling. After successful optimization, generate a parameter evolution log, recording the final adjustment rate of the weights of each band and the improvement effect on the core agricultural situation indicators.

[0087] This implementation content uses a targeted optimization strategy for agricultural sentiment perception to transform abstract mathematical adjustments into a robust agricultural problem solver, ensuring that the dimensionality reduction process maintains high reliability in complex farmland environments.

[0088] The number of iterations of the gradient descent method does not exceed two. It should be further noted that in the specific implementation process, the iteration control implementation content includes the following process: The optimization process of the gradient descent method sets a strict upper limit of two iterations to avoid the risk of overfitting. The first iteration performs parameter adjustment using a standard step size. If the verification index still does not meet the standard, the second iteration starts the agricultural situation adaptive step size strategy: From the seedling stage to the stage of closing the ridges, that is, when the leaf area index < 2.5: Allow the step size to be amplified to 1.8 times the initial value to accelerate the capture of rapidly changing spectral features; During the tuber swelling stage, that is, when the leaf area index ≥ 2.5: Switch to the micro-step size mode to prevent the principal component oscillation during the critical period of sugar accumulation.

[0089] The iteration process embeds three layers of termination logic, including: Agricultural index priority termination: When it is detected that the prediction error of the optimized core physiological parameters of sugar beet decreases by ≥ 15%, the iteration is immediately terminated. Among them, the optimized core physiological parameters of sugar beet include canopy nitrogen content and chlorophyll SPAD value; If the morphological distortion of the leaf spectral reflectance curve appears, stop the optimization immediately and roll back the parameters. Among them, the morphological distortion of the leaf spectral reflectance curve includes the inversion of the red edge slope.

[0090] Soil moisture coupling control: For sandy land plots with a field water holding capacity < 35%, if the reconstruction error of the soil moisture-related bands is still greater than 12% after the first iteration, the second iteration is prohibited; Start emergency spectral acquisition: Use a portable ground object spectrometer to supplement the measurement of drought stress samples.

[0091] For heavy clay soil plots with a field water holding capacity greater than 70%: When the classification accuracy of nutrient deficiency is not improved after two iterations, it is determined as waterlogging interference, and switch to the special discriminant model for waterlogging; Apply weight freezing to the waterlogging-sensitive bands.

[0092] Extreme weather fusing mechanism: When encountering high temperatures above 35°C, if the prediction deviation of the canopy temperature increases by > 0.5°C after the first iteration, then fuse the optimization; Enable heat stress spectral template matching to replace parameter adjustment.

[0093] Continuous rainfall causes the leaf surface to be wet: If the correction of the short-wave infrared reflectance fails, skip the second iteration; Associate with the 48-hour forecast of the weather station and delay restarting the process until sunny days.

[0094] After two iterations are completed or the termination condition is triggered, an optimized traceability report is generated, recording: the actual number of iterations, the type of termination trigger, and the adjustment range of the key band weights. Among them, the type of termination trigger includes reaching the standard, agronomy, meteorology, and soil. Whether the standard is reached or not, the final parameter set is output for reuse in subsequent plot monitoring.

[0095] This implementation content uses intelligent iterative control through multi-factor coupling to transform mathematical optimization into safe and controllable agricultural operations, ensuring the efficient convergence of adaptive optimization in complex farmland environments.

[0096] The classifier is a linear kernel support vector machine. It should be further noted that in the specific implementation process, when constructing the classifier and using the linear kernel support vector machine to perform growth state discrimination, the classification boundary parameters are dynamically configured according to the sugar beet growth stage: a loose classification interval is set at the seedling stage to accommodate large spectral variations, and the interval is gradually tightened after the ridges are closed to improve the discrimination sensitivity. For the characteristics of multi-spectral data, band grouping and weighting processing are implemented, that is: triple weights are assigned to the chlorophyll-sensitive bands of 680 - 750nm, and the weights of the soil background bands of 1400 - 2500nm are reduced to zero to shield interference. When sample class imbalance is encountered, a cost-sensitive learning mechanism is activated: the misjudgment cost of the disease class samples is set to five times that of the healthy class, forcing the classifier to tilt towards the minority class; if the deficiency sample is less than 10% of the total, synthetic simulation samples are synthesized in the feature space based on the Mahalanobis distance and labeled as "weak labels", and the weights are calculated at half when participating in training.

[0097] Specific agricultural scenarios trigger special adaptation strategies: In the monitoring of salinized plots, a dedicated feature channel for sodium ion toxicity is added, and the projection distance of the 590 - 620nm band in the classification hyperplane is independently calculated. When the conductivity > 4 dS / m, the decision weight of this channel is increased to twice the standard value. In a continuously rainy environment, an anti-wet noise classification mode is enabled, that is: only the water-insensitive bands: 550nm, 650nm, 1650nm are used to construct the classifier, and the clear-day model of the same historical period is imported synchronously as an auxiliary decision-making reference. When the difference between the results of the two models > 15%, manual review is triggered. During the high-temperature stress period, the permission for the red-edge feature to participate in classification is frozen. Because leaf curling may cause red-edge displacement artifacts, instead, the collaborative discrimination of the visible light bands of 450 - 580nm and the canopy temperature parameter is relied on.

[0098] After the classification model is deployed, a real-time diagnosis mechanism is embedded: the classification confidence distribution is automatically checked every ten samples processed. If the confidence of three consecutive samples is lower than 0.85, the spectral re-acquisition process is immediately triggered. For suspected misjudged samples related to saline-alkali, which include the confusion between magnesium deficiency and sodium toxicity, the soil available nutrient detection data is associated and a secondary verification is performed. When the sodium-potassium ratio > 2, it is automatically corrected to the ion toxicity category. The classifier performs model drift correction once a month. When incrementally training with fresh samples in the current season, the elastic weight consolidation technique is used to freeze the core feature weights, and only the decision boundary of the new category is allowed to be adjusted.

[0099] In discriminant enhanced principal component analysis, the between-class scatter is calculated using the Mahalanobis distance, and the within-class scatter is calculated using the trace of the covariance matrix. It should be further noted that in the specific implementation process, when calculating the scatter in discriminant enhanced principal component analysis, the between-class scatter is measured using the Mahalanobis distance, and the estimation strategy of the covariance matrix is dynamically adjusted according to the growth stage of sugar beet: the combined covariance matrix of all plot samples is used in the seedling stage to improve stability, and after entering the root swelling stage, it is switched to calculate by soil type to eliminate background interference. When sample distribution overlap is encountered, a class boundary strengthening mechanism is activated, that is: for class pairs with the center distance between healthy and diseased classes less than 1.2 times the historical mean, virtual boundary samples are introduced to force the Mahalanobis distance to increase by 15%.

[0100] The within-class scatter calculation uses the trace of the covariance matrix as the core index, and three agricultural situation adaptation operations are implemented: in the monitoring of saline-alkali plots, samples affected by sodium ions are grouped separately to calculate the within-class scatter to avoid the contamination of normal nutrient deficiency signals by ion toxicity characteristics; when high temperature weather lasts for more than three days, the participation right of samples with abnormal canopy temperature is frozen until the temperature drops; for samples with water droplets attached to the leaves after rain, a correction factor for the reflectance at the short-wave infrared band of 1550 nm is automatically enabled to compress the calculated value of the scatter to 80% of the true level.

[0101] Special responses are triggered under specific working conditions: when it is detected that the within-class scatter suddenly increases by more than 1.5 times the historical threshold, root cause diagnosis is immediately performed, that is: if the soil water content fluctuation > 20%, it is attributed to soil moisture interference, and the water compensation algorithm is automatically started; if accompanied by a red edge position shift > 3 nm, it is determined as early physiological disorder and a warning is pushed to the agronomy platform. When monitoring severely saline-alkali plots with pH > 8.5, a dedicated ion toxicity scatter channel is added: the within-class scatter is independently calculated for the sodium-sensitive band of 590 - 620 nm, and its weight increases linearly with the increase of the soil sodium-potassium ratio. When the sodium-potassium ratio > 2, the weight is amplified to three times the standard value.

[0102] Embed a quality verification loop before outputting the dispersion data: Calculate the correlation coefficient between the Mahalanobis distance and agronomic parameters, which include chlorophyll content and canopy nitrogen accumulation. If R² < 0.6, trigger spectral re-acquisition; Verify the consistency between the intra-class dispersion and the field phenotype observation records, and automatically associate the UAV inspection images with the sample groups with excessive dispersion for agronomic review. When updating the dispersion benchmark library with new samples in the current season every month, adopt a sliding window mechanism to only retain the data of the most recent three weeks to ensure that the algorithm adapts to the rapidly changing spectral characteristics of crops.

[0103] Through the dual-path principal component fusion mechanism and orthogonal verification design, while reducing the dimension of beet multi-spectral data, the inherent contradiction that it is difficult to balance the retention of the global structure and the enhancement of discriminant features in traditional dimensionality reduction methods is overcome. Among them, the standard principal component analysis path ensures the high-fidelity compression of the main information of spectral variation, the discriminant enhancement path specifically extracts weak agronomic features such as diseases and nutrient deficiencies, and the orthogonality verification proves the complementarity of the dual-path information from a mathematical level, making the reduced feature set comprehensive and specific. The dual-screening closed loop of reconstruction error and classification accuracy further guarantees the output quality: The dynamic noise threshold combined with the agricultural situation adaptive mechanism reduces the reconstruction error; The accuracy margin control based on decision confidence maintains the reduction of classification accuracy loss under complex working conditions such as saline-alkali interference and rainy environment.

[0104] Deeply embed the agronomic logic throughout the growth period of beets, and transform the abstract mathematical process into an executable agricultural situation operation chain. The parameter adjustment rule adaptive to the growth period makes the method have universality in the field and breaks through the requirements of traditional algorithms for a stable environment; The abnormal handling mechanism of soil-meteorological coupling improves the robustness under harsh farmland conditions, and the data availability rate is improved compared with traditional methods. Finally, a qualitative change in agricultural analysis efficiency is achieved, the early disease detection rate is also improved, the classification accuracy of nutrient deficiency status is improved, and the prediction error of root sugar content is controlled, providing a reliable basis for precise fertilization and pesticide application.

[0105] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0106] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for dimensionality reduction processing of beet multispectral data based on principal component analysis, characterized in that, It includes the following steps: (a)Centrally preprocess the original sugar beet multispectral data; (b)Extract the principal components using a two-path parallel mechanism: Path 1 performs standard principal component analysis and screens the principal components based on the variance contribution rate of the sample covariance matrix; Path 2 performs discriminant enhanced principal component analysis and extracts the discriminant principal components by solving the generalized eigenvalue problem in the non-linear feature space; (c)Verify the orthogonality of the two-path principal components. When the absolute value of the dot product of the principal component vectors of Path 1 and Path 2 is less than the threshold, perform principal component fusion; (d)Reconstruct the data based on the fused principal components and calculate the band-level average reconstruction error. At the same time, input it into the classifier to verify the classification accuracy of the growth state; (e)When the error exceeds the noise threshold or the accuracy is lower than the benchmark value, adjust the kernel function parameter weight factor of the discriminant enhanced path and then re-execute steps (b) - (d).

2. The method for reducing the dimensionality of sugar beet multispectral data based on principal component analysis according to claim 1, characterized in that: The discriminant enhanced principal component analysis is implemented through the following steps: Calculate the kernel function parameter weight factor according to the mutual information entropy between bands; Map the data to the non-linear space using a radial basis kernel function with a weight factor; Solve for the discriminant principal components with the goal of maximizing the ratio of the between-class scatter to the within-class scatter.

3. A method for reducing the dimensionality of beet multispectral data based on principal component analysis according to claim 1, characterized in that: The orthogonality verification specifically includes: Perform Gram-Schmidt orthogonalization on the principal component sequence of Path 1; Perform Gram-Schmidt orthogonalization on the principal component sequence of Path 2; Calculate the absolute value of the dot product of the i-th principal component of Path 1 and the j-th principal component of Path 2 pair by pair; When all absolute values of the dot products are less than the preset orthogonality threshold, it is determined that the principal component information of the two paths is complementary; among them, the preset orthogonality threshold is determined based on the statistical value of the inter-band correlation of the sugar beet multispectral data.

4. A method for reducing the dimensionality of beet multispectral data based on principal component analysis according to claim 3, characterized in that: The principal component fusion is implemented through the following steps: Merge the principal component sets after orthogonalization of the two paths; Normalize the variance contribution rate of the principal components of Path 1 and the discriminant contribution rate of the principal components of Path 2 into a comprehensive eigenvalue index; Sort in descending order according to the comprehensive eigenvalue index and select the top M principal components as the output.

5. A method for dimensionality reduction processing of beet multispectral data based on principal component analysis according to claim 1, characterized in that: The reconstruction error verification in step (d) needs to satisfy: The band-level average reconstruction error does not exceed a predefined multiple of the standard deviation of the historical data noise level; among them, the predefined multiple is determined by the confidence interval of the noise distribution of the sugar beet multispectral data.

6. A method for reducing the dimensionality of beet multispectral data based on principal component analysis according to claim 1, characterized in that: The classification accuracy verification in step (d) needs to satisfy: the classification accuracy is not lower than the preset accuracy margin value of the classification accuracy of the original data; among them, the preset accuracy margin value is determined by the decision confidence probability of the sugar beet growth state classification model; The classification accuracy verification in step (d) uses a support vector machine, and the verification objects include three growth states: healthy, nutrient deficiency, and disease.

7. A method for dimensionality reduction processing of beet multispectral data based on principal component analysis according to claim 1, characterized in that: In step (e), the adjustment of the kernel function parameter weight factor uses the gradient descent method, and the optimization goal is to minimize the weighted sum of the reconstruction error and the classification accuracy deviation.

8. A method for dimensionality reduction processing of beet multispectral data based on principal component analysis according to claim 7, characterized in that: The number of iterations of the gradient descent method does not exceed two times.

9. A method for dimensionality reduction processing of sugar beet multispectral data based on principal component analysis according to claim 1, characterized in that: The classifier is a linear kernel support vector machine.

10. A method for dimensionality reduction processing of sugar beet multispectral data based on principal component analysis according to claim 1, characterized in that: In the discriminant enhanced principal component analysis, the between-class scatter is calculated using the Mahalanobis distance, and the within-class scatter is calculated using the trace of the covariance matrix.

Citation Information

Cited By

  • Rapid crop component analysis method based on near infrared spectrum and radial basis function neural network modeling

    CN121034450A

  • Straw replacement nutrient release prediction method under water and fertilizer coupling condition

    CN121210922A