Method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy
By employing a local similarity matching and dynamic weighted fusion strategy, combined with near-infrared spectral characteristic wavelengths and depth features, the problems of poor adaptability and model instability in sugar beet seed detection were solved, achieving high-precision and robust seed quality detection.
Patent Information
- Application Number
- CN202511911552.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-12-17
AI Technical Summary
Existing near-infrared spectroscopy techniques based on global modeling face challenges in detecting sugar beet seed quality, including poor adaptability, model instability, and difficulty in overcoming the problems of low prediction accuracy and insufficient adaptability caused by the large individual biological variation of sugar beet seeds.
A robust detection method is established by employing a local similarity matching and dynamic weighted fusion strategy. By acquiring the near-infrared spectral wavelength and depth features of sugar beet seeds and combining them with a convolutional neural network, the most similar historical samples are dynamically selected for local interpolation scoring. The local density information of reference samples is introduced for credibility weighting.
It significantly improves the predictive accuracy and adaptability of sugar beet seed quality testing, enabling accurate and robust testing of different batches and varieties of seeds, and enhancing the robustness of evaluation results and the level of intelligent sorting decision-making in complex real-world scenarios.
Smart Images

Figure CN121324300B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of nondestructive testing of seed quality, in particular to a method for rapidly and nondestructively testing sugar beet seed quality based on near-infrared spectroscopy. BACKGROUND
[0002] Sugar beet is an important sugar crop and economic crop, and the quality of its seeds is the basis for ensuring high yield and high quality. Internal insect damage and mold are two common and hidden quality problems that are often difficult to detect from the appearance of the seeds but can directly lead to loss of seed vigor, a sharp drop in germination rate, and even the spread of field diseases, resulting in direct economic losses to agricultural production. Therefore, there is an urgent need to develop a technology that can rapidly, accurately, and nondestructively detect internal defects in sugar beet seeds to achieve precise seed sorting, ensure seed quality, and improve agricultural efficiency.
[0003] Traditional seed quality detection methods, such as standard germination tests, are accurate but time-consuming, taking up to 7-10 days, and are destructive, making them unsuitable for rapid sorting in production. Near-infrared spectroscopy technology has been explored for seed quality detection due to its speed and nondestructive nature. Existing technologies typically use the following process: collecting a large number of sample spectra, establishing a global prediction model (such as PLS, SVM) between the spectra and standard germination rate indicators, and using the model to predict new samples.
[0004] However, this method faces significant challenges when applied to individual sugar beet seeds: individual sugar beet seeds are small, with large biological variations, and individual spectra are severely affected by irrelevant factors such as shape, position, and seed coat texture; establishing a globally applicable and highly accurate global model requires a large number of samples covering all variation types, which is costly and difficult to maintain; the model is prone to overfitting or underfitting, and the prediction stability for new batches and new varieties of seeds is poor. Therefore, how to overcome the limitations of global modeling and develop a more robust and accurate nondestructive detection method that is more suitable for individual differences in sugar beet seeds has become a technical problem that needs to be solved in this field.
[0005] The above information disclosed in the BACKGROUND section is only intended to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0006] The present application aims to provide an optimized configuration method and device for a magnetic resistance type dynamic voltage restorer to solve the problems raised in the background.
[0007] To achieve the above object, the present application provides the following technical solutions:
[0008] A method for rapidly and nondestructively detecting the quality of sugar beet seeds based on near-infrared spectroscopy technology, and the specific steps include:
[0009] Step 1: Obtain a sample set of sugar beet seeds, collect the near-infrared spectrum of each seed in the sample set and the sugar beet seeds to be detected, form the near-infrared spectrum vector of each seed, and extract the characteristic wavelength absorbance and depth feature of the corresponding seed from the near-infrared spectrum vector of each seed as the identification feature;
[0010] Step 2: Determine the quality score of each sugar beet seed in the sample set, associate the quality score with the corresponding identification feature to form a control group, integrate the control groups of all sugar beet seeds in the sample set to form a control set, and determine the correlation degree between the sugar beet seed to be detected and each control group by traversing the identification feature of the sugar beet seed to be detected in the control set;
[0011] Step 3: Screen out the control groups that meet the preset condition in the correlation degree with the seed to be detected, mark them as reference control units, obtain the quality score of each reference control unit, and integrate all reference control units to form a reference control group;
[0012] Step 4: Take the deviation degree of the reference control unit and the seed to be detected as a reference, and take the quality score of the reference control unit as a reference, to obtain the preliminary quality score estimate of the sugar beet seed to be detected relative to each reference control unit;
[0013] Step 5: According to the correlation degree of the reference control unit to the sugar beet seed to be detected and the density degree of the reference control unit in the reference control group, obtain the credibility weight of each reference control unit, and combine the credibility weight and the preliminary quality score estimate to perform comprehensive evaluation, so as to determine the final quality score of the seed to be detected and judge whether the quality of the seed to be detected meets the requirements.
[0014] Further, the extraction of the characteristic wavelength and the depth feature specifically includes:
[0015] Integrate the near-infrared spectrum vectors of all seeds in the sample set to form an original spectrum matrix, and perform the following three steps in sequence for the pretreatment of the original spectrum matrix:
[0016] The original spectrum matrix is first subjected to multivariate scatter correction, the average value of the near-infrared spectrum vectors of all seeds in the sample set is taken as the average spectrum, the linear relationship between the near-infrared spectrum vector of each seed and the average spectrum is corrected, the scattering effect caused by the irregular shape and surface roughness difference of the sugar beet seeds is eliminated, and only the absorbance information reflecting the chemical composition difference is obtained;
[0017] On the basis of multiplicative scatter correction, the near-infrared spectrum vector of each seed is standardized to eliminate the baseline offset and intensity variation caused by the slight fluctuation of seed placement and measurement conditions, so that the spectral curve shape is more stable.
[0018] The near-infrared spectrum vector is smoothed by using a moving window with a specific window width and a polynomial order, thereby filtering out high-frequency noise and improving the signal-to-noise ratio of the spectrum while preserving the shape of the characteristic absorption peaks of the sugar beet seeds; the characteristic absorption peaks include at least three key peaks at 960 nm, 1200 nm and 1450 nm;
[0019] The characteristic wavelengths are extracted by using a competitive adaptive reweighted sampling method improved for the spectral characteristics of sugar beet seeds, specifically as follows:
[0020] The Monte Carlo sampling strategy is used, and each time sampling randomly selects a preset proportion of sugar beet seeds from the sample set as a modeling subset;
[0021] In each sampling cycle, a regression model of the near-infrared spectrum vector of the sugar beet seeds and the quality score is established by using the partial least squares method, and the regression coefficients of the wavelength variables in the near-infrared spectrum vector are recorded;
[0022] A double-factor weighted wavelength importance evaluation mechanism is designed, which considers both the absolute value of the regression coefficient of the wavelength variable and the spectral information entropy, and preferentially retains the wavelengths with large contribution to the prediction of the quality of the sugar beet seeds and rich information content;
[0023] Based on the weighted importance, adaptive resampling is performed to iteratively eliminate the wavelength variables that are not sensitive to the prediction of the quality of the sugar beet seeds;
[0024] The frequency of each wavelength in all sampling cycles is counted, and the wavelengths with a frequency exceeding a preset threshold are selected as the characteristic wavelengths, and the absorbance of the characteristic wavelengths is used as part of the identification features; the spectral information entropy covers the distribution characteristics of the absorbance values of the sugar beet seeds, and the absorbance values of the sugar beet seeds at each wavelength are divided into multiple quantization levels, the probability of belonging to each level is counted, and the information entropy of the wavelength is calculated based on these probabilities, and the specific calculation formula is as follows:
[0025] ;
[0026] wherein, is the information entropy of the jth wavelength, and j is the wavelength index; represents the probability that the absorbance value at the jth wavelength belongs to the cth level, c is the level index, and J is the total number of quantization levels;
[0027] Each probability value is obtained by kernel density estimation, and the formula is as follows:
[0028] ;
[0029] in, Let be the absorbance of the i-th beet seed at the j-th wavelength. Let be the center absorbance value of grade c at the j-th wavelength, and h be the bandwidth parameter. Let i represent the Gaussian kernel function, i be the index of the beet seed in the sample set, and n be the number of beet seeds in the sample set.
[0030] The deep feature is constructed using a convolutional neural network architecture specifically designed for the one-dimensional near-infrared spectroscopy of beet seeds. The input data to the convolutional neural network is the enhanced near-infrared spectral vector of each beet seed, the dimension of which corresponds to the number of wave points acquired by the spectrometer. The output data is a low-dimensional deep feature vector automatically learned and extracted from the input data. This network architecture includes at least:
[0031] Multiple one-dimensional convolution kernels are used to perform sliding convolution operations on the input near-infrared spectral vector to extract the feature patterns between local adjacent wave points in the spectrum.
[0032] A non-linear activation layer performs a non-linear transformation on the output of the convolutional layer;
[0033] Pooling layers downsample features, enhancing feature translation invariance and reducing data dimensionality;
[0034] Fully connected layers or global pooling layers map the extracted distributed features into a final fixed-length deep feature vector.
[0035] Furthermore, the enhanced near-infrared spectral vector refers to the enhancement of the near-infrared spectral vector of each beet seed by spectral translation, noise injection, and perturbation of key regions.
[0036] Spectral shift enhancement specifically refers to randomly shifting the input near-infrared spectral vector along the wave point dimension to simulate the system wavelength shift between different spectrometers;
[0037] Noise injection enhancement specifically refers to adding random noise to the input near-infrared spectral vector, the intensity of which matches the noise statistical characteristics in the actual measurement environment of beet seeds;
[0038] Key region perturbation enhancement specifically refers to applying a directed random perturbation to the absorbance value of the input near-infrared spectral vector within the spectral point range corresponding to the absorption bands of known key chemical components in beet seeds, in order to enhance the network's ability to identify this feature region.
[0039] Furthermore, the control set is specifically defined as including:
[0040] each beet seed in the sample set is single-seeded in a standardized greenhouse or field test plot; key traits of each beet seed are measured for the corresponding single plant at key growth stages of the beet seed; the key traits include but are not limited to: seedling plant height, rhizome at the tuber bulking stage, single plant fresh weight of tubers at the harvest stage, and sugar content of tubers;
[0041] each key trait of each beet seed is normalized, and the normalized values of each key trait are weighted and summed to obtain a quality score of each beet seed; in the normalization process, the minimum and maximum values of each key trait are the minimum and maximum values of all corresponding key traits in the sample set;
[0042] The quality score of each beet seed is associated with its identified features, and the control groups of all beet seeds are integrated to form a control set.
[0043] Further, the feature wavelength and the depth feature are subjected to Z-score standardization; the standardized traditional feature vector and the standardized depth feature vector are spliced to form a fusion feature vector;
[0044] The degree of association is determined by calculating the similarity between the fusion feature vector of the beet seed to be detected and the fusion feature vector of each control group in the control set; the similarity is calculated in the form of the inverse of the distance measure, and the formula is:
[0045] ;
[0046] wherein, represents the similarity between the xth control group and the beet seed to be detected; represents the distance function between the xth control group and the beet seed to be detected; represents the fusion feature vector of the beet seed to be detected, represents the fusion feature vector of the xth control group; x is the index of the control group in the control set;
[0047] The similarity between the beet seed to be detected and each control group is preliminarily screened with a dynamic threshold, which is adaptively determined according to the feature distribution density of the control set, and the calculation formula is:
[0048] ;
[0049] wherein, is the dynamic threshold; is the mean of the similarity calculated between the beet seed to be detected and all control groups in the control set; is the standard deviation of the similarity calculated between the beet seed to be detected and all control groups in the control set, is an adjustment factor.
[0050] Further, the distance function adopts a weighted Euclidean distance, and the specific formula is as follows:
[0051]
[0052] M is the total dimension number of the fusion feature vector; respectively represent and the feature value in the kth dimension; is the weight coefficient of the kth dimension feature, and k is the dimension index;
[0053] The weight coefficient is determined by analyzing the correlation between the feature value of the sugar beet seed in different dimensions in the sample set and the corresponding quality score, and the specific logic is as follows:
[0054] The correlation between the feature value in the kth dimension of all sugar beet seeds in the sample set and the corresponding quality score is analyzed, and the Spearman rank correlation coefficient is calculated;
[0055] The absolute values of the correlation coefficients of all dimensions are normalized to obtain the initial weight of each dimension;
[0056] The initial weight is smoothed and amplified to enhance the distinction between important features and secondary features, and the final weight calculation formula is as follows:
[0057]
[0058] wherein, indicates the initial weight.
[0059] Further, the reference control group specifically includes:
[0060] The correlation degree between all control groups in the control set and the sugar beet seed to be detected is sorted, that is, sorted from high to low according to the similarity value; select the control groups with the top several correlation degrees to form a primary candidate set;
[0061] The primary candidate set is subjected to feature space consistency verification, and the specific steps are as follows:
[0062] For each dimension in the primary candidate set, the absolute deviation of the feature value of all control groups in the primary candidate set in the dimension and the feature value of the sugar beet seed to be detected in the dimension is calculated; based on all the calculated absolute deviations, the standard deviation is calculated, which represents the difference degree in the dimension;
[0063] A difference degree threshold is preset, and for each control group in the primary candidate set, the number of dimensions with a difference degree greater than the difference degree threshold in the control group is counted, and if the number of dimensions exceeds the preset number threshold, the control group is marked as abnormal.
[0064] The average value and the standard deviation of all similarities in the primary candidate set are calculated to determine a dynamic correlation threshold; the dynamic correlation threshold is also adaptively determined according to the characteristic distribution density of the primary candidate set;
[0065] From the primary candidate set, a control group with a similarity not less than the dynamic correlation threshold is screened out, and the screened control group is labeled as a reference control unit; wherein, when compared with the dynamic correlation threshold, the similarity value of the control group marked as an anomaly is halved.
[0066] Further, the quality score of the to-be-detected sugar beet seed relative to each reference control unit is obtained by a local weighted regression interpolation algorithm, and the specific process is as follows:
[0067] The feature space distance between the fusion feature vector of the to-be-detected sugar beet seed and the fusion feature vector of each reference control unit is calculated, and the feature space distance is used as the deviation degree between the two; the feature space distance adopts Mahalanobis distance, and the calculation formula is as follows:
[0068] ;
[0069] Wherein, denotes the feature space distance between the to-be-detected sugar beet seed and the gth reference control unit, and g is the index of the reference control unit; denotes the fusion feature vector of the gth reference control unit; denotes the inverse matrix of the covariance matrix calculated by the fusion feature vectors of all screened reference control units;
[0070] According to the deviation degree, a monotonically decreasing kernel function is used to calculate the influence weight of each reference control unit relative to the quality score of the to-be-detected sugar beet seed; the monotonically decreasing kernel function is a Gaussian kernel function or a double square function; when the Gaussian kernel function is adopted, the calculation formula of the influence weight is as follows:
[0071] ;
[0072] Wherein, denotes the influence weight of the gth reference control unit relative to the to-be-detected sugar beet seed; is a bandwidth parameter, which is used to control the rate of weight decay with distance, and its value is dynamically determined according to the median or quantile of all reference control units;
[0073] The influence weight of the to-be-detected sugar beet seed and each reference control unit is used to perform weighted calculation on the quality score of the reference control unit, so as to obtain a preliminary quality score estimation value of the to-be-detected sugar beet seed relative to each reference control unit; for each reference control unit, a local linear model with the reference control unit as the center is used to calculate the preliminary quality score estimation value of the to-be-detected sugar beet seed, and the calculation is as follows:
[0074] All reference control units are taken as samples, the fusion feature vectors thereof are taken as independent variables, and corresponding quality scores are taken as dependent variables, a local weighted least square model is constructed, for a current reference control unit, a fitting target is a minimum weighted error sum of squares; an optimal local regression parameter, that is, an intercept term and a regression coefficient vector, is solved; the fusion feature vector of the to-be-detected sugar beet seed is substituted into the local weighted least square model, and a preliminary quality score estimation value relative to the reference control unit is calculated; that is, a product of the fusion feature vector of the to-be-detected sugar beet seed and a transpose of the regression coefficient vector is calculated first, and a sum of the product and the intercept term is the preliminary quality score estimation value of the to-be-detected sugar beet seed.
[0075] Further, the judgment of whether the quality of the to-be-detected sugar beet seed meets the requirement specifically includes:
[0076] The confidence weight of each reference control unit is calculated, the confidence weight includes two parts, the first part is the correlation degree of the reference control unit and the to-be-detected sugar beet seed, and the second part is the density degree of the reference control unit in the reference control group, which is represented by calculating the number of adjacent units within a preset neighborhood radius with the reference control unit as the center; the calculation formula of the confidence weight is as follows:
[0077] ;
[0078] wherein, the confidence weight of the gth reference control unit is represented by g; the similarity of the gth reference control unit is represented by g, which is used to represent the correlation degree of the gth reference control unit and the to-be-detected sugar beet seed; the density weight of the gth reference control unit is represented by g, which is used to represent the density degree of the reference control unit, for the gth reference control unit, the number of reference control units in the reference control group with a Euclidean distance less than a preset neighborhood radius is calculated, and the number is the density weight; , the minimum value and the maximum value of the density weight of all reference control units are represented by min and max respectively; , the preset harmonic coefficient is represented by h; ;
[0079] According to the credibility weight corresponding to each reference control unit, the preliminary quality score estimation values of all reference control units are weighted and summed, and the credibility weights of each reference control unit are summed, and the ratio of the two is the final quality score of the corresponding reference control unit;
[0080] The final quality score of the to-be-detected sugar beet seed is compared with the preset quality threshold, if the final quality score of the to-be-detected sugar beet seed is not less than the preset quality threshold, it is determined that the seed quality meets the requirements, otherwise it is determined that it does not meet the requirements.
[0081] In the above technical solution, the technical effects and advantages provided by the present application are:
[0082] The present application solves the core problems of poor adaptability and unstable model of the existing near-infrared spectrum detection method based on global modeling when applied to sugar beet single seed. Through the innovative use of local similarity matching and dynamic weighted fusion strategy, accurate, robust and highly practical detection is achieved. Specifically, by dynamically selecting the most similar multiple historical samples for each to-be-detected seed for local interpolation scoring, the defect of insufficient global model generalization ability caused by large individual biological variation of sugar beet seeds is effectively overcome, and the prediction accuracy and adaptability to different batches and varieties of seeds are significantly improved. By introducing the local density information of reference samples for credibility weighting, and dynamically setting the quality threshold based on batch statistics, the robustness of the evaluation results and the intelligence level of the sorting decision in complex real scenarios are greatly enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0083] Fig. 1 The present application provides a whole method flowchart;
[0084] Fig. 2 The present application provides a similarity fitting curve diagram. DETAILED DESCRIPTION
[0085] In order to make the purpose, technical scheme and advantages of the present application more clear and obvious, the present application is further described in detail below in combination with specific embodiments.
[0086] It should be noted that the technical terms or scientific terms used in the present application should be understood as the general meaning understood by those skilled in the art to which the present application belongs, unless otherwise defined. The terms "first", "second" and the like used in the present application do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and the like mean that the elements or objects before the terms cover the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and the like are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0087] Embodiment:
[0088] Please refer to Figs. 1-2 The present application provides a technical solution:
[0089] A method for rapidly and non-destructively detecting the quality of sugar beet seeds based on near-infrared spectroscopy technology, the specific steps comprising:
[0090] Step 1: Obtain a sample set of sugar beet seeds, collect the near-infrared spectrum of each seed in the sample set and the sugar beet seeds to be detected, form the near-infrared spectrum vector of each seed, and extract the characteristic wavelength absorbance and depth feature of the corresponding seed from the near-infrared spectrum vector of each seed as the identification feature.
[0091] In this embodiment, a statistically representative sample set of sugar beet seeds is obtained. The sample set should cover different origins, different years, different vigor grades (for example, high vigor, normal, aging, damaged, etc.) of the varieties to be detected, so as to ensure that the model established has wide applicability and robustness. The number of samples is recommended to be not less than 200, and preferably 500-1000, to meet the basic requirements of subsequent modeling and statistical analysis.
[0092] A Fourier transform near-infrared spectrometer or a grating scanning spectrometer is used for collection. Preferably, the wavelength range collected should cover 900 nm to 1700 nm, which contains the characteristic absorption bands of water, oil, protein, carbohydrate and other key components in sugar beet seeds. The spectral resolution should be set to 8 or higher to ensure sufficient acquisition of spectral details. To simulate the single kernel detection scenario, an accessory with a small aperture sample cup or single kernel seed stage should be used to ensure that the spectrum collected each time is only from one seed. To reduce measurement errors, each seed should be measured at least three times in different directions, and the average spectrum is taken as the original near infrared spectrum vector of the seed. The spectrum vector is a one-dimensional array, and its length is equal to the number of wave points collected.
[0093] In this embodiment, extracting the characteristic wavelength and the depth feature specifically includes:
[0094] The original spectrum signal usually contains a variety of interference information from instrument noise, sample physical properties (such as size, shape, surface roughness), and environmental fluctuations, and must be preprocessed to enhance the spectral features related to chemical composition. The near infrared spectrum vectors of all seeds in the sample set are integrated to form an original spectrum matrix, and the original spectrum matrix is preprocessed by sequentially performing the following three steps:
[0095] The original spectrum matrix is first subjected to multivariate scatter correction, the average spectrum of the near infrared spectrum vectors of all seeds in the sample set is taken as the average spectrum, the linear relationship between the near infrared spectrum vector of each seed and the average spectrum is corrected, the scattering effect caused by the irregular shape and surface roughness difference of the sugar beet seeds is eliminated, and the absorbance information reflecting only the chemical composition difference is obtained. The specific operation is: first, calculate the average spectrum of the original spectrum vectors of all seeds in the sample set; then, perform a linear regression of the original spectrum vector of each seed with the average spectrum to obtain the correction coefficient (slope and intercept); finally, correct the original spectrum using the correction coefficient to align the corrected spectrum with the average spectrum in shape, thereby obtaining the absorbance information mainly reflecting the chemical composition difference.
[0096] On the basis of multivariate scatter correction, the near infrared spectrum vector of each seed is subjected to standardization processing to eliminate the baseline shift and intensity variation caused by the slight fluctuation of seed placement position and measurement conditions, and make the spectral curve shape more stable. The calculation method is: for the spectrum vector of the seed after multivariate scatter correction, first calculate the mean and standard deviation of the absorbance of all wave points of the vector, then subtract the mean from the absorbance value of each wave point and divide by the standard deviation.
[0097] To filter out high-frequency random noise while preserving important spectral details (especially characteristic absorption peaks), Savitzky-Golay convolution smoothing method is adopted. Two key parameters need to be preset: window width (such as 5, 7, 9, 11, etc. odd points) and polynomial fitting order (usually 2 or 3). For example, choose window width of 9 and polynomial order of 2 for smoothing. This algorithm performs least squares fitting of a second-order polynomial within each point of the spectrum vector and its adjacent window, and takes the value of the fitted polynomial at the center point as the output of smoothing. The characteristic absorption peaks at least include the absorption peaks near 960 nm, 1200 nm and 1450 nm which are closely related to key chemical components such as moisture, sugar and protein in sugar beet seeds. Smoothing must be carried out on the premise of preserving the peak position, peak height and peak shape characteristics of these absorption peaks. Excessive smoothing will lead to loss of characteristic information.
[0098] The competitive adaptive reweighted sampling method improved for the spectral characteristics of sugar beet seeds is used to extract feature wavelengths, as follows:
[0099] The Monte Carlo sampling strategy is used, and each time sampling randomly selects a preset proportion of sugar beet seeds from the sample set as the modeling subset. Set the total sampling cycle number (such as 50 times). In each cycle, randomly select a preset proportion (such as 70%-80%) of seeds from the sample set as the modeling subset, and the remaining seeds as the prediction subset for internal validation. This random sampling strategy simulates the uncertainty in model construction, which helps to evaluate the stability of the wavelengths.
[0100] In each sampling cycle, a regression model of sugar beet seed near-infrared spectrum vector and quality score is established by partial least squares (PLS), and the regression coefficients of each wavelength variable in the near-infrared spectrum vector are recorded.
[0101] For example, the near infrared spectrum vectors (pre-processed) of all sugar beet seeds in the modeling subset are constructed into an independent variable matrix X, and their corresponding known quality scores are constructed into a dependent variable vector Y. The PLS algorithm is used to fit the matrix X and Y. PLS establishes a linear regression model by finding the latent variables (principal components) of X and Y, which can simultaneously explain the maximum variation of X and the correlation with Y. In the fitting process, a key hyperparameter, the number of latent variables, needs to be determined. This number can be determined by cross-validation or other methods, for example, in each sampling cycle, using a prediction subset or internal cross-validation, the number of latent variables that minimizes the prediction error (such as root mean square error RMSE) is selected. When the PLS model is trained, a set of regression coefficients is obtained, and the coefficient vector defines the linear mapping relationship from the spectrum data X to the quality score Y. Record this regression coefficient vector, and each element in it is the contribution weight (regression coefficient) of the corresponding wavelength variable to the quality score. The greater the absolute value of the coefficient, the more important the wavelength is in predicting the quality score in this modeling.
[0102] The number of latent variables in the modeling process is a key model parameter, and its specific value is not fixed, but is dynamically determined through the verification step in the model training process. A typical determination method is: after establishing the modeling subset in each Monte Carlo sampling, reserve a part of the samples as the internal validation set, or perform K-fold cross-validation on the modeling subset. Try different numbers of latent variables (for example, from 1 to 20), respectively establish PLS models and calculate the prediction error (such as prediction residual sum of squares PRESS or root mean square error RMSE) on the validation set. Select the number of latent variables that minimizes the prediction error or reaches the platform as the optimal parameter of the PLS model in this cycle. The reason for this is that too few latent variables, the model underfits, and cannot fully extract spectral information; too many, the model may overfit, introducing noise. By dynamically determining through verification, it can be ensured that the PLS model established each time is the optimal or nearly optimal linear model for the current sampling subset, so that the extracted regression coefficients are more representative and reliable.
[0103] A double-factor weighted wavelength importance evaluation mechanism is designed, considering both the absolute value of the regression coefficient of the wavelength variable and the spectral information entropy, and the wavelengths that contribute greatly to the prediction of the quality of sugar beet seeds and have rich information content are preferentially retained. The absolute value of the regression coefficient directly reflects the contribution size of the wavelength to the prediction of the quality score. The spectral information entropy measures the degree of confusion or information content of the absorbance value distribution at the wavelength.
[0104] Based on the weighted importance, iteratively eliminate the wavelength variables that are not sensitive to the prediction of the quality of sugar beet seeds; combine the two factors, for example, define the wavelength importance index in the form of weighted sum or product, and the expression is wherein, for regression coefficients, for adjustment factors (e.g. ) to balance the scales of the two terms. Wavelengths with high information entropy imply that they can provide more diverse and non-redundant information, which helps to improve the discrimination ability and robustness of the model.
[0105] The adaptive re-weighted sampling uses an exponential decay function to dynamically determine a retention ratio based on the weighted importance metric, and eliminates the wavelengths with the lowest importance. The PLS modeling and re-sampling are performed in a loop until the number of retained wavelengths reaches a pre-set minimum value or the loop ends.
[0106] The frequency of selection of each wavelength in all sampling loops is counted, and the wavelengths with a frequency of selection exceeding a pre-set threshold are selected as characteristic wavelengths, and the absorbance of the characteristic wavelengths is used as part of the identification characteristics. For example, the threshold is set to 50% or 60% here. These wavelengths are variables that are continuously proven to be important and informative for sugar beet seed quality prediction in multiple random modeling. The spectral information entropy encompasses the distribution characteristics of the absorbance values of the sugar beet seeds, divides the absorbance values of the sugar beet seeds at each wavelength into multiple quantization levels (e.g. 10 intervals divided equally or equally in frequency), counts the probability of belonging to each level, and calculates the information entropy of the wavelength based on the probabilities. The specific calculation formula is as follows:
[0107] ;
[0108] wherein, is the information entropy of the jth wavelength, and j is the wavelength index; represents the probability that the absorbance value at the jth wavelength belongs to the cth level, and c is the level index, and J is the total number of quantization levels.
[0109] The information entropy calculation formula is used to measure the information content or uncertainty degree of the absorbance data at a specific wavelength in the near-infrared spectrum. The continuous spectral absorbance values are converted into a probability distribution through kernel density estimation, and then the discrete characteristics of the distribution are quantified based on the Shannon information entropy theory. The above calculation enables the information entropy value to sensitively reflect the statistical distribution pattern of the spectral data at the wavelength: if the absorbance of all seeds is highly concentrated in a few levels (distribution concentration), the probability distribution peak is sharp, and the calculated information entropy value is small; on the contrary, if the absorbance values are uniformly distributed in each level (distribution dispersion), the probability distribution is flat, and the information entropy value is larger. From the technical effect, the information entropy index provides a new dimension judgment standard for wavelength screening that is independent of the regression model. The larger the information entropy value, the more significant the difference in absorbance between different sugar beet seeds at the wavelength, the more diverse the distribution, and the more likely it contains more discriminative information; the smaller the information entropy value, the more similar the response of all seeds at the wavelength, and the weaker the discrimination ability.
[0110] Each probability value is obtained through kernel density estimation, as shown in the following formula:
[0111] ;
[0112] in, Let be the absorbance of the i-th beet seed at the j-th wavelength; is the center absorbance value of grade c at the j-th wavelength (e.g., the seed for this grade range); h is the bandwidth parameter (usually determined using the Silverman rule of thumb); denoted as Gaussian kernel function; i is the index of beet seeds in the sample set, and n is the number of beet seeds in the sample set.
[0113] Compared to simple histogram statistics, this method overcomes the problem of inaccurate probability estimation caused by limited sample size or discontinuous data distribution. By introducing a Gaussian kernel function to smooth the contribution of each data point, it obtains a more robust probability value that better reflects the true continuous distribution characteristics of absorbance values. This provides a high-quality probabilistic input for subsequent calculations of information entropy, ensuring the reliability of information entropy as an evaluation index of wavelength importance.
[0114] When the calculated probability value A higher value indicates that at wavelength j, the absorbance of the beet seed sample is highly concentrated within the range represented by the c-th quantization level. This suggests that the spectral response at that wavelength may be relatively uniform or singular, carrying less differential information. Conversely, when... The smaller the value, the less likely the sample absorbance value appears in that specific level range. If we consider the probability distribution of all levels, and the probability values of all levels are small and evenly distributed at a wavelength, it means that the absorbance value is widely distributed and varied at that wavelength, indicating that the wavelength may capture more significant chemical or physical differences between samples.
[0115] The deep features are constructed using a convolutional neural network architecture specifically designed for the one-dimensional near-infrared spectrum of beet seeds. The input data of the convolutional neural network is the enhanced near-infrared spectral vector of each beet seed, the dimension of which corresponds to the number of wave points collected by the spectrometer. The output data is a low-dimensional deep feature vector that is automatically learned and extracted from the input data.
[0116] For example, the network architecture includes:
[0117] Input layer: Receives a spectral vector with dimension L, where L is the number of wave points.
[0118] Convolutional layer: Multiple one-dimensional convolutional kernels (e.g., the first layer uses 64 convolutional kernels with a width of 5 and a stride of 1) are used to perform sliding convolution on the spectral sequence to extract spectral feature patterns (such as absorption peak shoulders, inflection points, etc.) between local adjacent wave points.
[0119] Activation layer: Introduce a nonlinear activation function (e.g., ReLU) after the convolution layer to increase the network's nonlinear representation ability.
[0120] Pooling layer: Use one-dimensional max-pooling or average-pooling (e.g., pool window size of 2, step size of 2) after the activation layer to downsample the features, enhance the feature's invariance to slight translation, and reduce the data dimension.
[0121] Repeated stacking: Repeat the above convolution-activation-pooling structure 2-3 times to extract higher-level, more abstract features.
[0122] Flattening and fully connected / global pooling layer: Flatten the final multi-channel one-dimensional feature map or directly use global average pooling to convert it into a fixed-length vector. Then connect one or two fully connected layers (e.g., 128, 64 neurons), and the output of the last fully connected layer is the final extracted deep feature vector, which can be set according to needs (e.g., 32 or 64 dimensions). The network learns the mapping from the original spectrum to the high-quality representation through end-to-end training.
[0123] In this embodiment, the enhanced near-infrared spectrum vector refers to the near-infrared spectrum vector of each sugar beet seed being subjected to spectral translation enhancement, noise injection enhancement, and key region disturbance enhancement.
[0124] Spectral translation enhancement specifically refers to randomly shifting the input spectrum vector left or right by 1-3 wave points in the wave point dimension to simulate the system wavelength shift caused by slight differences in calibration between different spectrometers or the same spectrometer.
[0125] Noise injection enhancement specifically refers to adding random noise conforming to a Gaussian distribution to the input spectrum vector, and the strength (standard deviation) of the noise is determined according to statistical analysis of the actual measurement environment noise, for example, the standard deviation is set to 10%-20% of the standard deviation of the spectrum signal in the smooth region.
[0126] Key region disturbance enhancement specifically refers to identifying the absorption bands (e.g., near 960nm, 1200nm, 1450nm) corresponding to key chemical components (e.g., sugar, water) of the sugar beet seed in the spectrum, and in these wave point intervals, a guided random disturbance (e.g., slightly changing the peak height or peak width) based on the typical peak shape change in the region is applied to the absorbance value, forcing the network to pay more attention to and learn the feature patterns in these key regions.
[0127] Step 2: Determine the quality score of each sugar beet seed in the sample set, and associate the quality score with the corresponding identification feature to form a control group. Integrate the control groups of all sugar beet seeds in the sample set to form a control set. Traverse the identification feature of the to-be-detected sugar beet seed in the control set to determine the correlation degree between the to-be-detected sugar beet seed and each control group.
[0128] In this embodiment, the control set specifically includes:
[0129] Each sugar beet seed in the sample set is single-seeded in a standardized greenhouse with controllable environment or a field test plot with uniform soil, irrigation, and fertilization conditions. This is done to eliminate the interference of planting density, competition, and other factors on the traits of individual plants, and to ensure that the observed trait differences are mainly due to the quality of the seeds themselves.
[0130] During the key growth periods of sugar beet seed production, the key traits of the corresponding individual plant of each sugar beet seed are measured. The key traits include but are not limited to: seedling plant height, rootstock diameter during root enlargement period, single plant root fresh weight at harvest period, and sugar content of root.
[0131] Seedling plant height refers to the measurement at a certain number of days after emergence (e.g., 30 days after emergence), reflecting the germination vigor and early growth potential of the seed. Rootstock diameter during root enlargement period refers to the measurement at the middle growth stage (e.g., 80-100 days after sowing), reflecting the basis of vegetative growth and root development. Single plant root fresh weight at harvest period refers to the measurement at the harvest time of the technological maturity period, which is a direct reflection of yield. The sugar content of the root is determined by a refractometer or near-infrared spectrometer to measure the juice of the harvested root, which is a core indicator of sugar beet quality.
[0132] The selection of the above key traits is based on the breeding and cultivation goals of sugar beet production, but the present application is not limited thereto. Depending on different detection purposes (such as focusing on yield or sugar content), other traits such as leaf number, disease resistance, etc. can be added.
[0133] Each key trait of each sugar beet seed is normalized to map the original value to the [0, 1] interval. This normalization based on the global range of the sample set ensures that the scores of all seeds are comparable on the same scale. In the normalization process, the minimum and maximum values of each key trait are the minimum and maximum values of all corresponding key traits in the sample set.
[0134] For example, assuming that the sample set has 100 seeds, the maximum root fresh weight is 800 grams and the minimum is 300 grams. The fresh weight of a certain seed is 650 grams, then its normalized fresh weight value is: (650-300) / (800-300)=0.7.
[0135] The normalized key trait values are then weighted and summed to obtain a quality score of each sugar beet seed. The weight coefficients reflect the importance of different traits in the comprehensive quality evaluation. The specific values depend on the specific breeding or production goals and can be determined by domain expert experience or mathematical methods.
[0136] Example 1 (balanced type): If both yield and quality are pursued, the weight coefficient of plant height at seedling stage can be set to 0.2, the weight coefficient of root diameter at root enlargement stage can be set to 0.2, the weight coefficient of fresh weight of single root at harvest stage can be set to 0.3, and the weight coefficient of sugar content of root can be set to 0.3.
[0137] Example 2 (high yield-oriented type): If high yield is emphasized, the weight coefficient of plant height at seedling stage can be set to 0.1, the weight coefficient of root diameter at root enlargement stage can be set to 0.2, the weight coefficient of fresh weight of single root at harvest stage can be set to 0.5, and the weight coefficient of sugar content of root can be set to 0.2.
[0138] The quality score of each sugar beet seed is associated with its identified characteristics, and all control groups of sugar beet seeds are integrated to form a control set.
[0139] In this embodiment, the characteristic wavelengths and the depth features are subjected to Z-score standardization. The mean value in the standardization is the mean value of all seeds in the sample set at the feature, and the standard deviation is also the standard deviation of all seeds at the feature. After standardization, the data mean value in each feature dimension is 0, and the standard deviation is 1, which are in the same order of magnitude. The standardized traditional feature vector and the standardized depth feature vector are then spliced to form a fusion feature vector.
[0140] The degree of association is determined by calculating the similarity between the fusion feature vector of the sugar beet seed to be detected and the fusion feature vectors of each control group in the control set.
[0141] The distance function adopts a weighted Euclidean distance, and the specific formula is as follows:
[0142] ;
[0143] wherein M is the total dimension number of the fusion feature vector; respectively represent and the feature value at the kth dimension; is the weight coefficient of the kth dimension, and k is the dimension index.
[0144] The formula gives different weight coefficients to each dimension of the feature vector based on the standard Euclidean distance, so that the feature dimension that contributes more to the quality prediction occupies a more important position in the distance calculation, thus more accurately reflecting the real difference between two samples in key quality characterization. The greater the value, the greater the difference between the detected seed and the reference seed in the core features affecting the quality score, i.e., the more dissimilar the quality characteristics of the two; on the contrary, the smaller the distance value, the closer the position of the two in the weighted feature space, the more matching the key features, i.e., the more similar the quality characteristics.
[0145] The weight coefficient is determined by analyzing the correlation between the feature values of the beet seeds in the sample set in different dimensions and their corresponding quality scores, and the specific logic is as follows:
[0146] The correlation between the kth feature value of all beet seeds in the sample set and its corresponding quality score is analyzed, and the Spearman rank correlation coefficient is calculated. The Spearman correlation coefficient does not assume that the data follows a normal distribution and is not sensitive to outliers, and is more suitable for this example.
[0147] The absolute values of the correlation coefficients of all dimensions are normalized to obtain the initial weight of each dimension.
[0148] The initial weight is smoothed and amplified to enhance the distinction between important features and secondary features, and the final weight calculation formula is:
[0149] ;
[0150] wherein, represents the initial weight.
[0151] For example, assuming that the fusion feature vector has 3 dimensions, and the absolute values of the Spearman correlation coefficients between them and the quality score are |r1|=0.8, |r2|=0.3, |r3|=0.5, respectively.
[0152] The initial weight is calculated as , , .
[0153] The final weight , , .
[0154] This method allows features highly related to seed quality (such as wavelengths corresponding to key chemical absorption bands) to occupy a larger proportion in distance calculation, thus making the similarity matching more focused on core spectral information affecting quality, improving the accuracy of matching.
[0155] The similarity is calculated in the form of the reciprocal of the distance metric, and its formula is:
[0156] ;
[0157] wherein, represents the similarity between the xth control group and the sugar beet seed to be detected; represents the distance function between the xth control group and the sugar beet seed to be detected; represents the fusion feature vector of the sugar beet seed to be detected, represents the fusion feature vector of the xth control group; x is the index of the control group in the control set.
[0158] This function maps the distance to the interval (0, 1], the closer the distance (the smaller D), the closer to 1 the similarity; the farther the distance, the closer to 0 the similarity. Adding a constant 1 is to avoid the denominator being zero and to control the range of similarity. Specifically, the greater the similarity value , the smaller the distance between the fusion feature vector of the seed to be detected and the control group in the feature space, that is, the closer the two are in the internal attributes such as chemical composition and physical structure reflected by the spectrum, indicating that they may have similar genetic background, physiological state or quality potential. On the contrary, the smaller the similarity value, the greater the distance between the two in the feature space, the more significant the difference in the internal attributes, and the higher the possibility of belonging to different categories or quality levels.
[0159] The similarity between the sugar beet seed to be detected and each control group is preliminarily screened with a dynamic threshold, which is adaptively determined according to the feature distribution density of the control set, and its calculation formula is:
[0160] ;
[0161] wherein, is the dynamic threshold; is the mean value of the similarity calculated between the sugar beet seed to be detected and all control groups in the control set; is the standard deviation of the similarity calculated between the sugar beet seed to be detected and all control groups in the control set, is an adjustment factor.
[0162] controls the strictness of screening. The greater the value is, the lower the dynamic threshold is, and the more candidate seeds are retained; The smaller the value is, the more stringent the screening is. The value of can be determined according to experience or optimized on the sample set through cross-validation. A typical value range is 1.0 to 3.0. For example, if the value of is taken as 2.0, The threshold is set at a position of two standard deviations below the mean of the similarity distribution. According to the normal distribution characteristics, this will filter out about 95% of the extremely low similarity samples, and retain about 5% of the most similar samples to the tested seed as candidates, which significantly improves the computing efficiency while ensuring the screening effect. The specific value of the threshold can be fine-tuned according to the size of the control set and the density of the data distribution.
[0163] In this embodiment, 30 control groups and sample data of the beet seeds to be tested are collected, and the similarity of each control group is calculated. The specific data is shown in the following table:
[0164] Table 1: Similarity and distance function value table
[0165]
[0166] In combination with the above table data and Fig. 2 It can be seen that the calculation of the similarity reveals a typical nonlinear decay relationship. From the data, it can be seen that as the distance increases, the similarity gradually decreases, but this decrease is not linear and uniform, but presents a smooth decay trend that gradually approaches zero. When the distance is small, the similarity is very sensitive to the change of the distance, for example, the distance increases from 0.1 to 0.5, and the similarity decreases from 0.909 to 0.667, with a significant change. When the distance is large, the similarity decreases gradually with the increase of the distance, for example, the distance increases from 4.0 to 6.0, and the similarity decreases from 0.200 to 0.143. This shows that the similarity function has strong discrimination ability in the near distance interval, and can effectively highlight the reference samples highly similar to the tested seed. In the far distance interval, it plays a smooth suppression role, avoiding excessive interference of individual extreme samples on the overall evaluation result.
[0167] From the perspective of pattern recognition and similarity measurement, the design of this function skillfully balances local sensitivity and global robustness. In the specific application of sugar beet seed quality detection, the spectral feature differences between seeds often focus on a few key wavelengths or feature combinations, and slight feature shifts can mean significant quality differences. Therefore, in areas where the feature space distance is close, the rapid decay of the similarity function helps to filter out truly valuable neighbor samples, providing a high-quality data foundation for subsequent local weighted regression. At the same time, the gentle decay of the function in the distant area also avoids complete rejection of non-similar samples, and when the data distribution is sparse or abnormal samples exist, it still maintains a certain numerical stability, enhancing the model's adaptability in complex real-world scenarios. In addition, this similarity calculation method is essentially a distance-based confidence measure, which not only reflects the closeness of samples in the feature space, but also provides a quantitative basis for subsequent credibility weighting and double dynamic threshold screening. In actual detection, the similarity value can be directly used as a quantitative indicator of sample correlation, to determine whether a historical sample should be included in the reference set, or to give it a corresponding weight in integrated prediction. Combined with the setting of dynamic threshold, the system can adaptively adjust the screening criteria according to the distribution density of the current sample set, further improving the intelligence and practicality of the detection method.
[0168] Step 3: Selecting the control group that meets the preset conditions in terms of correlation degree with the seed to be detected, labeling it as a reference control unit, obtaining the quality score of each reference control unit, and integrating all reference control units to form a reference control group.
[0169] In this embodiment, the selection of the reference control group specifically includes:
[0170] The correlation degree (similarity Sim(x)) between the sugar beet seed to be detected and each control group in the control set calculated in step 2 is arranged in descending order from high to low. After sorting, the top N control groups are selected to form a preliminary candidate set. The value of parameter N can be dynamically determined according to the actual application scenario and the size of the control set. A recommended determination method is: N = max(20, 0.1 * total number of control set samples), that is, at least 20 most similar samples are selected, or 10% (rounded up) of the total number of control set samples are selected. The reason for this value is that on the one hand, it is necessary to ensure a sufficient number of candidate samples for subsequent analysis to avoid statistical instability due to too few samples; on the other hand, it is necessary to control the computational complexity to avoid including too many irrelevant samples in the candidate set. For example, if the control set has a total of 500 seed data, N = max(20, 0.1 * 500) = max(20, 50) = 50, that is, the top 50 control groups with the highest similarity are selected to form the preliminary candidate set.
[0171] The primary candidate set is generated based on the overall similarity ranking, and there may be individual control groups that differ greatly from the seed to be detected in a few key feature dimensions, that is, the overall similarity and local abnormality. Therefore, the primary candidate set needs to be checked for feature space consistency, and the specific steps are as follows:
[0172] For each dimension in the primary candidate set, the absolute deviation of the feature value of all control groups in the primary candidate set on the dimension from the feature value of the sugar beet seed to be detected on the dimension is calculated respectively. Based on all the calculated absolute deviations, the standard deviation is calculated, which represents the difference degree on the dimension. The standard deviation reflects the dispersion degree of the difference between the samples in the candidate set and the seed to be detected on the dimension. The larger the standard deviation value, the greater the fluctuation of the difference between the candidate samples and the seed to be detected on the dimension, and the poorer the consistency.
[0173] A difference degree threshold is preset, and for each control group in the primary candidate set, the number of dimensions with a difference degree greater than the difference degree threshold in the control group is counted. If the number of dimensions exceeds the preset number threshold, the control group is marked as abnormal.
[0174] The difference degree threshold is used to judge whether the absolute deviation on a certain dimension is significantly deviated from the overall difference level of the candidate set. The difference degree threshold can be set as a certain multiple of the median of the standard deviation of all dimensions. For example, the difference degree threshold is set as 2 times the median of the standard deviation of each dimension. The rationality of such value is that, based on the common experience of identifying mild outliers in statistics (such as 1.5 times), the threshold is linked to the median of the data distribution, which can adapt to the difference scale of different feature dimensions and avoid misjudgment or omission of judgment caused by fixed threshold for certain dimension-sensitive features.
[0175] For example, assuming that the fusion feature vector dimension is equal to 50 and the number threshold is set to 5. For a control group, if it is found that the absolute deviation of more than 5 feature dimensions is significantly greater than the difference degree threshold of the corresponding dimension, the control group is marked as abnormal. The similarity value of the control group marked as abnormal is halved when compared with the dynamic correlation degree threshold, so as to weaken its influence.
[0176] After consistency checking, the primary candidate set needs to be further selected to determine the final reference control unit. This step introduces a dynamic correlation degree threshold, which is adaptively determined according to the current similarity distribution of the primary candidate set.
[0177] The average value and standard deviation of all similarities in the primary candidate set are calculated to determine the dynamic correlation degree threshold. The dynamic correlation degree threshold is also adaptively determined according to the feature distribution density of the primary candidate set.
[0178] First, for all the controls in the primary candidate set, their adjusted similarity is calculated (abnormal controls are halved, normal controls are unchanged). The mean and standard deviation of these similarities are calculated; the formula for the dynamic correlation threshold is: wherein is the dynamic correlation threshold, is the mean and standard deviation of all the similarities in the primary candidate set, respectively; is an adjustment factor, used to control the strictness of the screening. Its value range is usually 0.5 to 1.5. For example, take = 1.0, i.e. the threshold is set to the mean minus one standard deviation. The rationality of this value is that it sets the threshold in the lower region of the distribution based on the similarity distribution of the candidate set itself. When the internal similarity of the candidate set is highly concentrated (small ), the threshold is close to the mean, and the screening is relatively strict; when the similarity is dispersed (large ), the threshold is lowered, ensuring that a sufficient number of samples are selected, enhancing the adaptability of the algorithm.
[0179] From the primary candidate set, the controls with a similarity not less than the dynamic correlation threshold are screened out, and these screened controls are designated as the reference control unit.
[0180] For example, assume that there are 500 samples in the control set in this embodiment, and after the seed to be detected is calculated for similarity, the top N = 50 are taken as the primary candidate set.
[0181] Consistency check: the standard deviation of the difference degree of the 50 samples in M = 50 dimensions is calculated, and the median is 0.05. Set the multiple to 2, then the difference degree threshold is 0.1. Check each candidate sample, find that sample 3 has a deviation of more than 0.1 in 8 dimensions, exceeding the quantity threshold (value is 5), so it is marked as abnormal, and its similarity is adjusted from 0.92 to 0.46.
[0182] Dynamic threshold screening: the mean and standard deviation of the adjusted similarities of the 50 candidate samples are calculated. Set , then the dynamic correlation threshold is 0.78. From the 50 samples, the samples with an adjusted similarity not less than 0.78 are screened out, a total of 35. These 35 samples are the final reference control unit.
[0183] In this step, the core input data is the fusion feature vector generated in step 2. In the fusion feature vector, different features (such as traditional wavelength absorbance, CNN deep features) may have completely different numerical ranges and dimensions, and direct calculation of distance or similarity will lead to the dominance of features with large dimensions. Therefore, before performing the correlation degree calculation (step 2) and the feature space distance calculation of this step (such as Mahalanobis distance in subsequent step 4), it is necessary to ensure that these feature vectors have been standardized. Specifically, the Z-score standardization method should be used, that is, for each feature dimension, the mean and standard deviation are calculated based on the entire control set, and then all feature values (including the to-be-detected seed) on this dimension are converted to standardized values. Z-score standardization makes each feature dimension have a mean of 0 and a standard deviation of 1, ensuring that each dimension has equal importance when calculating Euclidean distance or Mahalanobis distance. And many machine learning algorithms and distance measures (especially those involving gradient descent or covariance calculation) perform more stably and converge faster after data standardization.
[0184] Step 4: Based on the deviation degree of the reference control unit and the to-be-detected seed, and the quality score of the reference control unit as a reference, obtain the preliminary quality score estimate of the to-be-detected sugar beet seed relative to each reference control unit;
[0185] In this embodiment, the quality score of the to-be-detected sugar beet seed relative to each reference control unit is obtained by a local weighted regression interpolation algorithm, and the specific process is as follows:
[0186] The feature space distance between the fusion feature vector of the to-be-detected sugar beet seed and the fusion feature vector of each reference control unit is calculated, and this distance is used as the deviation degree between them; the feature space distance uses Mahalanobis distance, and the calculation formula is as follows:
[0187] ;
[0188] Wherein, represents the feature space distance (deviation degree) between the to-be-detected sugar beet seed and the gth reference control unit, and g is the index of the reference control unit; represents the fusion feature vector of the gth reference control unit; represents the inverse matrix of the covariance matrix calculated by the fusion feature vectors of all reference control units selected. The covariance matrix is calculated using a subset of reference control units rather than the entire control set in order to make the distance measure more consistent with the local feature space distribution where the current to-be-detected seed is located, thereby improving the accuracy of distance calculation.
[0189] For example, assuming that the fusion feature vector is 3-dimensional, the to-be-detected seed feature vector is , the jth reference unit feature vector F(g) = [1.0, 1.0, 0.0], the inverse matrix of the covariance matrix calculated by the reference control unit subset is a 3x3 matrix. Then the difference vector is [0.2, -0.2, -0.5], the product of the difference vector and is calculated, and then the dot product of the transpose of the difference vector is multiplied, and finally the square root is taken, that is, the Mahalanobis distance .
[0190] According to the degree of deviation, the influence weight of each reference control unit relative to the quality score of the beet seed to be detected is calculated by using a monotonically decreasing kernel function; the monotonically decreasing kernel function is a Gaussian kernel function or a double square function; when the Gaussian kernel function is used, the calculation formula of the influence weight is as follows:
[0191] ;
[0192] wherein, represents the influence weight of the gth reference control unit relative to the beet seed to be detected; is a bandwidth parameter for controlling the rate of weight decay with distance, and the value is dynamically determined according to the median or quantile of all reference control units; for example, the median is used to determine , which represents the median of the feature space distance of all reference control units.
[0193] When the double square function is used, the calculation formula of the influence weight is as follows:
[0194] ;
[0195] wherein, the determination of the bandwidth parameter is crucial, which controls the rate of weight decay with distance. If is too large, the points with a relatively large distance still have a relatively large weight, the model tends to be a global average, and the locality is weakened; if is too small, only the points with a very small distance contribute, and the model may be sensitive to noise. The present application proposes an adaptive determination method: the value of the bandwidth parameter is dynamically determined according to the distribution of the distance between all reference control units and the seed to be detected. Preferably, the median is used to determine. The rationality of such value lies in that the median is not sensitive to outliers and can represent the central tendency of the sample distance; dividing by makes the sample weight at the median distance about , which ensures that the weight function has good decay characteristics, which can focus on the local neighborhood, and also will not lose enough data support due to too narrow neighborhood.
[0196] The quality score of the reference control unit is weighted and calculated by using the influence weight of the detected sugar beet seed and each reference control unit, so as to obtain a preliminary quality score estimation value of the detected sugar beet seed relative to a certain reference control unit; the preliminary quality score estimation value is calculated by a local linear model with a certain reference control unit as the center, and the calculation is as follows:
[0197] The local model is defined, and for the local area with the jth reference control unit as the center, a linear relationship is assumed , wherein is the quality score, F is the fusion feature vector, is the intercept, is the regression coefficient vector.
[0198] The weighted objective function is constructed: the weighted error sum of squares of all reference control units is minimized as the target, and the optimal local regression parameter and are solved. The objective function is , wherein is the total number of reference control units, q is also the reference control unit index, q≠g. is the weight of the qth reference control unit in the local model with the gth reference control unit as the center. It is the local distance weight for constructing the local regression model with the gth unit as the center. , wherein indicates the Euclidean distance or Mahalanobis distance between the qth reference control unit and the gth reference control unit; is the bandwidth parameter of the local model, which can be set as a fixed value or adaptively determined according to the median of . Preferably, a proportion (such as 0.5 times) of the median of the distance between all reference control units is taken. In this way, the size of the local model can be automatically adjusted according to the density of the current reference control group, a more local model is used in the data-intensive place, and a smoother model is used in the data-sparse place, thereby enhancing the robustness of the algorithm. is the known quality score of the qth reference control unit; is the fusion feature vector of the qth reference control unit.
[0199] The regression parameter is solved: the weighted least squares method is used to solve the above objective function, so as to obtain the optimal local regression parameter with the jth reference control unit as the center.
[0200] The preliminary estimation value is calculated: the fusion feature vector of the detected sugar beet seed is substituted into the local model, so as to calculate the preliminary quality score estimation value relative to the gth reference control unit, that is, , denotes the optimal local regression parameter of the gth reference control unit.
[0201] For example, assume there are 3 reference control units (g = 3) with the following data: Unit 1: F(1) = [1.0, 1.5], S(1) = 85; Unit 2: F(2) = [1.2, 1.3], S(2) = 88 (as the current center j = 2); Unit 3: F(3) = [0.8, 1.7], S(3) = 82; the seed to be tested = [1.1, 1.4].
[0202] First, with Unit 2 as the center, the distance weights of other units to it are calculated (for example, using a Gaussian kernel, = 0.5): , , .
[0203] Then, a weighted least squares model is established, aiming to minimize .
[0204] Solving the local parameters and .
[0205] Finally, calculate to obtain the preliminary estimate value based on Unit 2, for example, 86.5.
[0206] By iterating the above process through each reference control unit as the center, a set of preliminary quality score estimates corresponding to each reference control unit is obtained, providing the basis for the final integrated decision in Step 5.
[0207] Step 5: According to the relevance of the reference control unit to the sugar beet seed to be tested and the density of the reference control unit in the reference control group, obtain the credibility weight of each reference control unit, and conduct comprehensive evaluation combining the credibility weight and the preliminary quality score estimate to determine the final quality score of the seed to be tested, and judge whether the quality of the seed to be tested meets the requirements.
[0208] In this embodiment, judging whether the quality of the sugar beet seed to be tested meets the requirements specifically includes:
[0209] Calculate the credibility weight of each reference control unit, which includes two parts: the first part is the relevance of the reference control unit to the sugar beet seed to be tested, and the second part is the density of the reference control unit in the reference control group, which is represented by calculating the number of adjacent units within a preset neighborhood radius centered on the reference control unit; the calculation formula of the credibility weight is as follows:
[0210] ;
[0211] in, This represents the confidence weight of the g-th reference control unit; Sim(g) represents the similarity of the g-th reference control unit, which is used to characterize the degree of association between the g-th reference control unit and the beet seed to be tested. Before being used to calculate the confidence weight, all Sim(g) have been normalized so that their values are in the range of [0,1]. The density weight represents the density of the g-th reference control unit, which characterizes the density of the reference control unit. For the g-th reference control unit, the number of reference control units in the reference control group whose Euclidean distance to it is less than the preset neighborhood radius is calculated, and this number is the density weight. , These are the minimum and maximum values of the density weights for all reference control units, respectively. , These are the preset harmonic coefficients. .
[0212] Density weight This is characterized by calculating the number of neighboring units within a predetermined neighborhood radius centered on the g-th reference unit. The specific calculation steps are as follows:
[0213] Calculate the characteristic spatial distances between all pairs of reference control units within the reference control group. Euclidean distance is preferred because it is more intuitive and efficient in calculating local density.
[0214] Determining the neighborhood radius R is a key preset parameter. It is determined by calculating the median or 30th percentile of the Euclidean distances between all pairwise reference control units within the reference control group, and using this as the neighborhood radius. The reason for choosing the median or a lower percentile is that it can adapt to the density of the dataset. In dense regions of the feature space, this radius can include a suitable number of neighboring points; in sparse regions, it can also ensure that a certain number of neighboring points are included in the statistics, avoiding the failure of density estimation due to a radius that is too large or too small.
[0215] Assuming the reference control group has 100 units, 4950 pairwise distances are calculated. These distances are sorted in ascending order, and the median (the 2475th distance value) is taken; for example, if this value is 1.85, then the neighborhood radius is set. .
[0216] Counting the number of neighboring units: For each reference control unit g, count the units in the reference control group that satisfy the characteristic space distance being less than R. The quantity. This quantity is the density weight of the cell. .
[0217] harmonization coefficient and determines whether the model trusts more the individuals with high similarity to the seed to be tested or the representative individuals located in the dense region of the feature space. The optimal combination of can be determined by grid search or cross-validation on an independent validation set. The validation set consists of a batch of sugar beet seed samples with known quality scores but not involved in the modeling.
[0218] Specifically, a series of values (such as 0.1, 0.2, …, 0.9) are set, and the corresponding .
[0219] For each set of , steps 1 to 5 of the method of this embodiment are run on the validation set to obtain the predicted quality score of each validation seed.
[0220] Calculate the evaluation index between the predicted score and the true quality score (from field test), such as root mean square error or coefficient of determination.
[0221] Select the set of that minimizes the root mean square error or maximizes the coefficient of determination as the final preset value.
[0222] Through optimization based on the validation set, it can be ensured that the value of the harmonization coefficient is data-driven and can adapt to the complex mapping relationship between the spectral characteristics of sugar beet seeds and quality to the greatest extent, thereby improving the accuracy of integrated prediction. For example, on a validation set of a certain sugar beet seed variety, the optimal combination is , which indicates that in this data set, the direct similarity of the reference unit to the seed to be tested is slightly more important than its representativeness in the population.
[0223] According to the reliability weight corresponding to each reference control unit, the preliminary quality score estimate values of all reference control units are weighted and summed, and the reliability weights of each reference control unit are summed, and the ratio of the two is the final quality score of the corresponding reference control unit.
[0224] For example, there are 3 reference control units. The first reference control unit has a preliminary quality score estimate value of 85, a similarity of 0.92, and a density weight of 8; the second reference control unit has a preliminary quality score estimate value of 78, a similarity of 0.85, and a density weight of 15; and the third reference control unit has a preliminary quality score estimate value of 90, a similarity of 0.95, and a density weight of 5.
[0225] Let , The normalized density factor is calculated, with the first reference control unit being 0.3, the second reference control unit being 1, and the third reference control unit being 0.
[0226] The reliability weight is calculated, with the first reference control unit being 0.734, the second reference control unit being 0.895, and the third reference control unit being 0.665.
[0227] The final quality score is calculated to be about 83.7. It can be seen that although the third reference control unit has the highest preliminary score (90), its density is low, and the reliability weight is pulled down; although the second reference control unit has a lower preliminary score (78), it is located in the dense area, and the reliability is the highest, which has a significant impact on the final result. The final score 83.7 is closer to the weighted comprehensive judgment.
[0228] The final quality score of the detected sugar beet seed is compared with the preset quality threshold value. If the final quality score of the detected sugar beet seed is not less than the preset quality threshold value, it is determined that the seed quality meets the requirements; otherwise, it is determined that it does not meet the requirements.
[0229] The quality threshold value should be determined based on the quality score distribution of the sample set (or a larger, representative historical seed library) and the actual production needs. The specific determination method includes:
[0230] Determination based on historical data distribution; calculate the mean and standard deviation of the quality scores of all sugar beet seeds in the sample set; set the quality threshold value as the mean minus n times the standard deviation. n is a positive integer, which can be 1 or 1.5. For example, the mean is 75, the standard deviation is 10, and n = 1, then the quality threshold value is 65. This method determines that seeds with a score below one standard deviation below the mean are unqualified, which is a division based on statistical significance.
[0231] Determination based on actual planting performance; sort the sample set seeds by quality score and correlate them with their corresponding field actual performance (such as whether to germinate, whether the final yield meets the standard). Through receiver operating characteristic curve (ROC curve) analysis, find the score critical point that can most effectively distinguish between performance meeting the standard and performance not meeting the standard, and set this critical point as the quality threshold value. This method directly links the threshold value to the final economic traits, and is the most practical.
[0232] Determination in combination with industry standards or expert experience; if the industry has clear seed grading standards (such as germination rate, vigor index), the final quality score predicted by this method can be correlated with these standards through a regression model, and then the predicted score value corresponding to the industry standard is set as the quality threshold value. Experts can also determine a batch of seeds based on experience, and set the average predicted score corresponding to the dividing line between qualified and unqualified determined by experts as the quality threshold value.
[0233] The above formulas are all dimensionless values calculated, the formula is obtained by collecting a large amount of data to simulate the most recent real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.
[0234] The above embodiments can be implemented wholly or partially by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented wholly or partially in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solutions.
[0235] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, which can be located in one place or distributed on multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiments according to actual needs.
[0236] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application.
Claims
1. A method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy technology, characterized in that, The specific steps include: Step 1: Obtain a sample set of sugar beet seeds, collect the near-infrared spectrum of each seed in the sample set, and form a near-infrared spectrum vector of each seed; extract the characteristic wavelength absorbance and depth feature of each seed from the near-infrared spectrum vector of each seed as the identification feature; Step 2: Determine the quality score of each sugar beet seed in the sample set, correlate the quality score with the corresponding identification feature to form a control group, integrate the control group of all sugar beet seeds in the sample set to form a control set; traverse the identification feature of the sugar beet seed to be detected in the control set to determine the correlation degree between the sugar beet seed to be detected and each control group; Step 3: Screen out the control group that meets the preset condition in the correlation degree between the detected seed and the detected seed, and mark it as a reference control unit, and obtain the quality score of each reference control unit, and integrate all reference control units to form a reference control group; Step 4: Take the deviation degree of the reference control unit and the detected seed as a reference, and take the quality score of the reference control unit as a reference to obtain the preliminary quality score estimate of the detected sugar beet seed relative to each reference control unit; Step 5: According to the correlation degree of the reference control unit to the detected sugar beet seed and the density of the reference control unit in the reference control group, obtain the credibility weight of each reference control unit, and combine the credibility weight and the preliminary quality score estimate to make a comprehensive evaluation to determine the final quality score of the detected seed, and judge whether the quality of the detected seed meets the requirements; The characteristic wavelength and the depth feature are subjected to Z-score standardization processing; then the standardized traditional feature vector and the standardized depth feature vector are spliced to form a fusion feature vector; The correlation degree is determined by calculating the similarity between the fusion feature vector of the detected sugar beet seed and the fusion feature vector of each control group in the control set; the similarity is calculated in the form of the inverse of the distance metric, and the formula is: wherein, represents the similarity between the xth control group and the sugar beet seeds to be detected; represents the distance function between the xth control group and the sugar beet seeds to be detected; represents the fusion feature vector of the sugar beet seeds to be detected, represents the fusion feature vector of the xth control group; x is the index of the control group in the control group. The similarity of the detected sugar beet seed and each control group is preliminarily screened with a dynamic threshold, and the dynamic threshold is adaptively determined according to the feature distribution density of the control set, and the calculation formula is: wherein, is a dynamic threshold value; is the mean of the similarity values calculated for all control groups in the control set; is the standard deviation of the similarity values calculated for all control groups in the control set, is an adjustment factor; The distance function adopts weighted Euclidean distance, and the specific formula is as follows: Wherein, M is the total dimension number of the fusion feature vector; respectively represent and the eigenvalue of the kth dimension; is the weight coefficient of the kth dimension feature, and k is the dimension index. The weight coefficient is determined by analyzing the correlation between the feature value of the sugar beet seed in the sample set in different dimensions and its corresponding quality score, and the specific logic is: The correlation between the kth feature value of all sugar beet seeds in the sample set and its corresponding quality score is analyzed, and the Spearman rank correlation coefficient is calculated; The absolute values of all dimensions of the correlation coefficient are normalized to obtain the initial weight of each dimension; The initial weight is smoothed and amplified to enhance the distinction between important features and secondary features, and the final weight calculation formula is: wherein, denotes the initial weight.
2. The method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy according to claim 1, characterized in that, The extraction of the characteristic wavelength and the depth feature specifically includes: Integrate the near-infrared spectrum vector of all seeds in the sample set to form an original spectrum matrix, and perform the following three steps in sequence for pretreatment of the original spectrum matrix: The original spectral matrix is subjected to multivariate scattering correction, the average spectrum of all seeds in the sample set is taken as the average spectrum, the linear relationship between the near-infrared spectrum vector of each seed and the average spectrum is corrected, the scattering effect caused by the irregular shape and surface roughness difference of the sugar beet seeds is eliminated, and the absorbance information reflecting only the chemical composition difference is obtained; On the basis of multivariate scattering correction, the near-infrared spectrum vector of each seed is subjected to standardization processing, the baseline shift and intensity change caused by the slight fluctuation of the seed placement position and the measurement condition are eliminated, and the spectral curve form is more stable; The moving window with a specific window width and a polynomial order is used for smoothing the near-infrared spectrum vector, the high-frequency noise is filtered out on the premise of retaining the shape of the characteristic absorption peak of the sugar beet seed, and the signal-to-noise ratio of the spectrum is improved; the characteristic absorption peak at least includes three key peaks at 960 nm, 1200 nm and 1450 nm; The competitive adaptive reweighted sampling method improved for the spectral characteristics of the sugar beet seed is used to extract the characteristic wavelength, and the method is as follows: The Monte Carlo sampling strategy is used, and a preset proportion of the sugar beet seeds in the sample set is randomly selected as a modeling subset each time; In each sampling cycle, a regression model of the near-infrared spectrum vector of the sugar beet seed and the quality score is established by the partial least squares method, and the regression coefficients of the wavelength variables in the near-infrared spectrum vector are recorded; A double-factor weighted wavelength importance evaluation mechanism is designed, the absolute value of the regression coefficient of the wavelength variable and the spectral information entropy are considered at the same time, the wavelength with rich information content and great contribution to the quality prediction of the sugar beet seed is preferentially retained; Based on the weighted importance, adaptive resampling is carried out, and the wavelength variable insensitive to the quality prediction of the sugar beet seed is iteratively eliminated; The frequency of each wavelength in all sampling cycles is counted, the wavelength with a frequency higher than a preset threshold is selected as the characteristic wavelength, and the absorbance of the characteristic wavelength is taken as part of the identification feature; the spectral information entropy covers the distribution characteristics of the absorbance value of the sugar beet seed, the absorbance value of the sugar beet seed at each wavelength is divided into multiple quantization levels, the probability of belonging to each level is counted, and the information entropy of the wavelength is calculated based on the probabilities; the specific calculation formula is as follows: wherein, Hj is the information entropy for the jth wavelength, j being the wavelength index; Pj,c represents the probability that the absorbance value at the jth wavelength belongs to the cth class, c being the class index, J being the total number of quantization classes. The probability values are obtained by kernel density estimation, and the formula is as follows: wherein, is the absorbance of the i-th sugar beet seed at the j-th wavelength, is the absorbance central value of the c-class at the j-th wavelength, h is a bandwidth parameter, denotes a Gaussian kernel function, i is the index of a sugar beet seed in the sample set, n is the number of sugar beet seeds in the sample set; The deep feature is obtained by using a convolutional neural network architecture specially designed for the one-dimensional near-infrared spectrum of the sugar beet seed; the input data of the convolutional neural network is the enhanced near-infrared spectrum vector of each sugar beet seed, and the dimension corresponds to the number of wave points collected by the spectrometer; the output data is a low-dimensional deep feature vector automatically learned and extracted from the input data; the network architecture at least includes: A plurality of one-dimensional convolution kernels are used to perform sliding convolution operation on the input near-infrared spectrum vector to extract the feature mode between the local adjacent wave points of the spectrum; A nonlinear activation layer performs nonlinear transformation on the output of the convolution layer; The pooling layer down-samples the features to enhance the feature translation invariance and reduce the data dimension; The fully connected layer or the global pooling layer maps the extracted distributed features to the final fixed-length deep feature vector.
3. The method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy according to claim 2, characterized in that, The enhanced near-infrared spectrum vector refers to: performing spectrum shift enhancement, noise injection enhancement and key region disturbance enhancement on the near-infrared spectrum vector of each sugar beet seed; The spectrum shift enhancement specifically refers to randomly shifting the input near-infrared spectrum vector in the wave point dimension to simulate the system wavelength offset between different spectrometers; The noise injection enhancement specifically refers to adding random noise to the input near-infrared spectrum vector, and the intensity of the noise matches the statistical characteristics of the noise in the actual measurement environment of the sugar beet seed; The key region disturbance enhancement specifically refers to applying a guided random disturbance to the absorbance value of the input near-infrared spectrum vector in the wave point interval corresponding to the absorption band of the known key chemical component of the sugar beet seed, to enhance the identification ability of the network to this feature region.
4. The method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy according to claim 1, characterized in that, Determining the control set specifically includes: Each sugar beet seed in the sample set is sown individually in a standardized greenhouse or field test area; during the key growth period of the sugar beet seed, the key traits of each sugar beet seed corresponding to the single plant are measured; the key traits include but are not limited to: seedling plant height, rhizome during root enlargement period, single plant fresh weight of root at harvest period, and sugar content of root; Each key trait of each sugar beet seed is normalized, and the normalized key trait values are weighted and summed to obtain the quality score of each sugar beet seed; in the normalization process, the minimum and maximum values of each key trait are the minimum and maximum values of all corresponding key traits in the sample set; The quality score of each sugar beet seed is associated with its identification feature, and all control groups of sugar beet seeds are integrated to form a control set.
5. The method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy according to claim 1, characterized in that, Screening the reference control group specifically includes: Sort the correlation degree of all control groups in the control set with the sugar beet seed to be detected, that is, sort them from high to low according to the similarity value; select the control groups with the top several correlation degrees to form a preliminary candidate set; The feature space consistency verification of the preliminary candidate set includes the following steps: For each dimension in the preliminary candidate set, calculate the absolute deviation of the feature value of each control group in the preliminary candidate set in that dimension and the feature value of the sugar beet seed to be detected in that dimension; based on all the calculated absolute deviations, calculate the standard deviation, which represents the difference degree in that dimension; A difference degree threshold is preset, and for each control group in the preliminary candidate set, the number of dimensions with a difference degree greater than the difference degree threshold is counted, and if the number of dimensions exceeds the preset number threshold, the control group is marked as abnormal; Calculate the average value and standard deviation of all similarity degrees in the preliminary candidate set to determine the dynamic correlation degree threshold; the dynamic correlation degree threshold is also determined adaptively according to the feature distribution density of the preliminary candidate set; From the preliminary candidate set, filter out the control groups with a similarity degree not less than the dynamic correlation degree threshold, and mark these filtered control groups as reference control units; wherein, when compared with the dynamic correlation degree threshold, the similarity value of the control group marked as abnormal is halved.
6. The method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy according to claim 5, characterized in that, The quality score of the sugar beet seed to be detected relative to each reference control unit is specifically obtained through a local weighted regression interpolation algorithm, and the specific process is as follows: The feature space distance between the fusion feature vector of the to-be-detected sugar beet seed and the fusion feature vector of each reference control unit is calculated, and the feature space distance is taken as the deviation degree between the two; the feature space distance adopts Mahalanobis distance, and the calculation formula is as follows: wherein, denotes the feature space distance between the sugar beet seed to be detected and the gth reference control unit, g being the reference control unit index; denotes the fused feature vector of the gth reference control unit; denotes the inverse of the covariance matrix calculated from the fused feature vectors of all reference control units selected. According to the deviation degree, an influence weight of each reference control unit relative to the quality score of the to-be-detected sugar beet seed is calculated by using a monotonically decreasing kernel function; the monotonically decreasing kernel function is a Gaussian kernel function or a double square function; when the Gaussian kernel function is adopted, the calculation formula of the influence weight is as follows: wherein, represents the influence weight of the gth reference control unit with respect to the sugar beet seeds to be detected; is a bandwidth parameter for controlling the rate of decay of the weight with distance, the value of which is dynamically determined according to the median or quantile of all reference control units; The quality score of the reference control unit is calculated by using the influence weight of the to-be-detected sugar beet seed and each reference control unit, so as to obtain a preliminary quality score estimate value of the to-be-detected sugar beet seed relative to each reference control unit; for each reference control unit, the preliminary quality score estimate value of the to-be-detected sugar beet seed is calculated by using a local linear model with the reference control unit as the center, and the calculation is as follows: All reference control units are taken as samples, the fusion feature vectors thereof are taken as independent variables, and the corresponding quality scores are taken as dependent variables, so as to construct a local weighted least square model; for the current reference control unit, the fitting target is the minimum weighted error sum of squares; the optimal local regression parameter, i.e., the intercept term and the regression coefficient vector, is solved; the fusion feature vector of the to-be-detected sugar beet seed is substituted into the local weighted least square model, so as to calculate the preliminary quality score estimate value relative to the reference control unit; that is, the product of the fusion feature vector of the to-be-detected sugar beet seed and the transpose of the regression coefficient vector is calculated first, and the sum of the product and the intercept term is the preliminary quality score estimate value of the to-be-detected sugar beet seed.
7. The method for rapid non-destructive detection of sugar beet seed quality based on near infrared spectroscopy according to claim 1, characterized in that, The quality of the to-be-detected sugar beet seed is determined to meet the requirements, which specifically includes: The reliability weight of each reference control unit is calculated, and the reliability weight includes two parts, the first part is the correlation degree of the reference control unit and the to-be-detected sugar beet seed, and the second part is the dense degree of the reference control unit in the reference control group, which is represented by calculating the number of adjacent units in a preset neighborhood radius centered on the reference control unit; the calculation formula of the reliability weight is as follows: wherein, represents the reliability weight of the gth reference control unit; represents the similarity of the gth reference control unit, used to represent the correlation degree between the gth reference control unit and the sugar beet seeds to be detected; represents the density weight of the gth reference control unit, used to represent the density of the reference control unit, and for the gth reference control unit, the number of reference control units in the reference control group with the Euclidean distance less than the preset neighborhood radius is calculated, which is the density weight; , respectively represent the minimum value and the maximum value of the density weight of all reference control units; , respectively represent the preset harmonic coefficients, ; The preliminary quality score estimate values of all reference control units are weighted and summed according to the reliability weights of the reference control units, and the reliability weights of the reference control units are summed, and the ratio of the two is the final quality score of the corresponding reference control unit; The final quality score of the to-be-detected sugar beet seed is compared with the preset quality threshold, if the final quality score of the to-be-detected sugar beet seed is not less than the preset quality threshold, it is determined that the seed quality meets the requirements; otherwise, it is determined that it does not meet the requirements.
Citation Information
Patent Citations
Method for selecting cypress tree superior seed using hyperspectral image technology
KR102112088B1
Classification of seeds
WO2010000266A1