Spectral determination method and system for iNDF content of whole-plant corn silage
By performing neighborhood processing and feature screening on the corn sample set, a more representative model was constructed, which solved the problem of inaccurate measurement of iNDF content of whole corn in the prior art, and achieved more efficient and accurate spectral measurement.
Patent Information
- Application Number
- CN202411446619.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The prior art uses near-infrared spectroscopy (NIRS) combined with partial least squares regression (PLS) algorithm to measure the content of undigestable neutral detergent fiber (iNDF) in the whole corn in silage. It is mainly due to improper selection of principal components and failing to best reflect the characteristics of iNDF, affecting the prediction accuracy.
By performing neighborhood processing on the corn sample set, building a neighborhood set, extracting the first feature corresponding to the neighborhood set of each corn sample, performing data matching and feature screening, obtaining training features, which are used to train the target model and optimize the model training process.
It improves the consistency of data analysis and the generalization ability of the model, reduces the risk of overfitting, and improves the prediction efficiency and accuracy of iNDF content spectroscopy of the whole plant corn in silage.
Smart Images

Figure CN119104503B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spectral data analysis, and in particular to a spectral determination method and system for the iNDF content of whole-plant silage corn. Background Art
[0002] In the livestock industry, whole-plant corn silage is an important feed resource. The evaluation of its nutritional value and digestibility is crucial for feed formulation and animal husbandry management. Whole-plant corn silage refers to the harvesting of the entire above-ground portion of the corn plant, including the cobs, at an appropriate harvest time. This silage is then chopped, processed, and fermented to produce a feed suitable for herbivorous livestock such as cattle and sheep. This feed boasts high biomass yield, excellent fiber quality, long-lasting greenness, and optimal dry matter and moisture contents, making it ideal for closed silage storage with anaerobic fermentation.
[0003] When evaluating the nutritional value of whole-plant corn silage, the indigestible neutral detergent fiber (iNDF) content is a key indicator. Determination of iNDF content helps estimate the digestibility and range of organic matter, playing an important role in regulating diet digestibility and ruminant feed intake. Therefore, accurate determination of iNDF content in whole-plant corn silage is crucial for improving feed utilization and animal performance.
[0004] Near-infrared spectroscopy (NIRS) is widely used for rapid determination of feed composition, including iNDF content, due to its rapid, non-destructive, and reproducible nature. NIRS rapidly analyzes the spectral absorption characteristics of a sample and can rapidly complete spectral measurements. Within NIRS, the partial least squares regression (PLS) algorithm is a common method for constructing predictive models.
[0005] However, existing technologies for measuring corn iNDF using NIRS combined with PLS algorithms have limitations. During traditional PLS model construction, the principal component with the greatest similarity to both the explanatory and explained variables is typically selected as the target component. However, this approach is inaccurate because it fails to account for differences in iNDF content between corn samples, resulting in the selected principal component not optimally reflecting iNDF characteristics. Furthermore, the selected principal component may sometimes represent a general characteristic of corn rather than a key indicator of iNDF content. This inaccurate principal component selection can affect subsequent model construction, leading to a decrease in the accuracy of the NIRS algorithm in predicting corn iNDF content. Summary of the Invention
[0006] In order to solve the technical problem of inaccurate measurement of the iNDF content of whole-plant corn silage, the present invention aims to provide a spectral measurement method and system for the iNDF content of whole-plant corn silage. The technical solution adopted is as follows:
[0007] In a first aspect, an embodiment of the present invention provides a spectral method for measuring the iNDF content of whole-plant corn silage, which is applied to a spectral system for measuring the iNDF content of whole-plant corn silage. The method comprises:
[0008] Performing data processing on the corn sample set to be tested to obtain iNDF content data and corresponding infrared spectrum data of each corn sample in the corn sample set;
[0009] performing neighborhood processing on the iNDF content data of each corn sample of a training sample set in the corn sample set to obtain a neighborhood set of each corn sample, wherein the neighborhood set of each corn sample includes a plurality of corn samples having the same iNDF content data as that of each corn sample;
[0010] Performing data analysis on the neighborhood set of each corn sample to obtain a first feature corresponding to the neighborhood set of each corn sample, wherein the first feature represents infrared spectrum data common to the neighborhood set of each corn sample;
[0011] performing data matching on each corn sample in the corn sample set according to the first feature corresponding to the neighborhood set of each corn sample to obtain a plurality of feature sequences, wherein the plurality of feature sequences include each first feature and a feature sequence corresponding to each first feature;
[0012] Determining, based on the multiple feature sequences, whether a first feature corresponding to the neighborhood set of each corn sample is a target feature, wherein the target feature is used to determine the iNDF content data of each corn sample;
[0013] Feature screening is performed based on the target features to obtain training features, and the training features are used to train a target model, and the target model is used for spectral determination of iNDF content of whole-plant corn silage.
[0014] In a second aspect, an embodiment of the present invention provides a spectral measurement system for the iNDF content of whole-plant corn silage, comprising a memory and a processor, wherein the memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the spectral measurement system for the iNDF content of whole-plant corn silage implements the spectral measurement method for the iNDF content of whole-plant corn silage as described in the first aspect.
[0015] In a third aspect, an embodiment of the present invention provides a spectral measurement device for the iNDF content of whole-plant corn silage, which is applied to a spectral measurement system for the iNDF content of whole-plant corn silage. The device comprises:
[0016] a processing unit, configured to perform data processing on the corn sample set to be tested, and obtain iNDF content data and corresponding infrared spectrum data of each corn sample in the corn sample set;
[0017] The processing unit is further configured to perform neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain a neighborhood set of each corn sample, wherein the neighborhood set of each corn sample includes a plurality of corn samples having the same iNDF content data as that of each corn sample;
[0018] an analyzing unit, configured to perform data analysis based on the neighborhood set of each corn sample to obtain a first feature corresponding to the neighborhood set of each corn sample, wherein the first feature represents infrared spectrum data common to the neighborhood set of each corn sample;
[0019] a matching unit, configured to perform data matching on each corn sample in the corn sample set according to the first feature corresponding to the neighborhood set of each corn sample, to obtain a plurality of feature sequences, wherein the plurality of feature sequences include each first feature and a feature sequence corresponding to each first feature;
[0020] The processing unit is further configured to determine, based on the multiple feature sequences, whether a first feature corresponding to the neighborhood set of each corn sample is a target feature, wherein the target feature is used to determine the iNDF content data of each corn sample;
[0021] A screening unit is used to perform feature screening based on the target feature to obtain training features, wherein the training features are used to train a target model, and the target model is used for spectral determination of iNDF content of whole-plant corn silage.
[0022] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the spectral determination method for the iNDF content of whole-plant corn silage as described in the first aspect.
[0023] The present invention has the following beneficial effects: by performing neighborhood processing on a corn sample set, the present invention ensures that samples in a neighborhood set of each corn sample have similar iNDF contents, thereby improving data consistency during data analysis; the first feature of the obtained neighborhood set of each corn sample represents shared infrared spectral data, which helps to identify and extract the most valuable spectral features for iNDF content prediction; by performing data analysis on the neighborhood set, the feature sequences obtained help to build a model with greater generalization capability because they are based on the spectral features of multiple samples rather than a single sample; by performing data matching on each corn sample to obtain multiple feature sequences, this helps to find the similarity of spectral features between different samples, thereby optimizing subsequent data analysis and model training processes; by obtaining training features through feature screening, the risk of overfitting in the model training process can be reduced because the screened features are more likely to reflect the actual changes in iNDF content, that is, using the screened training features to train the target model can improve the model's prediction efficiency and accuracy for spectral determination of iNDF content in whole-plant silage corn. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 A schematic flow chart of a method for spectroscopic determination of iNDF content in whole-plant silage corn provided by one embodiment of the present invention;
[0026] Figure 2 A schematic structural diagram of a spectroscopic measurement device for the iNDF content of whole-plant corn silage provided by one embodiment of the present invention;
[0027] Figure 3 A schematic diagram of the structure of a spectral measurement system for the iNDF content of whole-plant corn silage provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0028] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effectiveness of a spectroscopic method and system for measuring iNDF content in whole-plant corn silage according to the present invention. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0029] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0030] The following describes in detail a specific scheme of a spectral measurement method and system for the iNDF content of whole-plant corn silage provided by the present invention with reference to the accompanying drawings.
[0031] See also Figure 1 , which shows a flow chart of a method for spectroscopic determination of iNDF content of whole-plant corn silage provided by one embodiment of the present invention, the method comprising the following steps:
[0032] S10. Perform data processing on the corn sample set to be tested to obtain iNDF content data and corresponding infrared spectrum data of each corn sample in the corn sample set.
[0033] Among them, the corn sample set includes corn samples of different varieties, growth conditions, maturity, etc., and each corn sample is assigned a unique identifier.
[0034] Among them, iNDF content data can be determined by wet chemical methods (such as nylon bag method) to determine the iNDF content of each sample.
[0035] Among them, the infrared spectrum data is obtained by using a near-infrared spectrometer to perform spectral scanning on each sample and collect the spectral absorption data of each sample within a specific wavelength range.
[0036] Specifically, before determining the iNDF content and collecting spectral data, the corn samples may need to be properly processed, such as drying and crushing; ensure the data quality during the measurement and collection process, and eliminate outliers and errors; record the iNDF content and spectral data of each sample in a database or data table for subsequent analysis; convert the spectral data into digital vectors, each vector containing absorbance values at a series of wavelengths; preprocess the spectral data, such as standardization or normalization, to eliminate the impact of differences in measurement conditions; use the iNDF content data as a label for the spectral data for subsequent machine learning model training.
[0037] As can be seen, this implementation yields a dataset containing the iNDF content of each corn sample and its corresponding infrared spectral data. This dataset serves as the foundation for subsequent prediction models and analysis of corn samples, and can be used for a variety of purposes, such as variety identification, quality assessment, and optimizing growing conditions.
[0038] S20. Performing neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain a neighborhood set of each corn sample, wherein the neighborhood set of each corn sample includes multiple corn samples having the same iNDF content data as that of each corn sample.
[0039] The process of neighborhood processing is to construct a neighborhood set for each corn sample in the training sample set based on its iNDF content data using the KNN method (k=5). That is, each sample will be grouped together with the other four samples with the closest iNDF content (plus the sample itself, a total of 5 samples) to form a neighborhood set. The purpose of this process is to cluster samples with similar iNDF contents, as these samples are likely to also have similar spectral characteristics.
[0040] Specifically, since different corn samples have different iNDF contents, it is difficult to extract the spectral features of iNDF. Therefore, a neighborhood is constructed for each sample point to obtain corn samples with consistent iNDF content in the sample, and the reference value of the corn sample in the PLS algorithm is obtained. The iNDF neighborhood of the corn sample is constructed using the KNN method Select a corn sample , get the neighborhood set containing the sample , to collect The sample corn in the dataset was used as the analysis object, and the PCA algorithm was used to obtain the corresponding principal components. The obtained principal components reflect the common infrared spectral characteristics of corn in the neighborhood set.
[0041] It can be seen that in this embodiment, through the above process, the neighborhood sets of each corn sample are constructed, and these sets can then be used for further data analysis and model training.
[0042] S30. Perform data analysis based on the neighborhood set of each corn sample to obtain a first feature corresponding to the neighborhood set of each corn sample, where the first feature represents infrared spectrum data common to the neighborhood set of each corn sample.
[0043] After constructing the neighborhood sets, data analysis is performed on each set to extract common infrared spectral data features. These features can be represented as the main change trend in the infrared spectral data of each sample, namely the first feature.
[0044] The extraction of the first feature is achieved through statistical methods such as principal component analysis (PCA), which can help identify the most important sources of variation in the data. In this process, the first feature of each neighborhood set is extracted, which can represent the common spectral characteristics of all samples in the set.
[0045] Before data analysis, the infrared spectral data for all samples in the neighborhood of each corn sample are aggregated. This data is typically presented as spectral curves, with each curve representing the infrared spectrum of a sample. Before analysis, the spectral data may require preprocessing, such as baseline correction, normalization, and smoothing, to eliminate interference from the instrument or sample processing.
[0046] It can be seen that in this embodiment, the first feature of the neighborhood set of each corn sample is extracted, and this feature can be used for further analysis, such as establishing a prediction model to associate infrared spectral data with iNDF content.
[0047] S40. Perform data matching on each corn sample in the corn sample set according to the first feature corresponding to the neighborhood set of each corn sample to obtain a plurality of feature sequences, where the plurality of feature sequences include each first feature and a feature sequence corresponding to each first feature.
[0048] The purpose of data matching is to find the sequence that best represents the spectral features of each corn sample. Through matching, the spectral data of each sample can be compared with the extracted first features to identify and construct the feature sequences corresponding to these features.
[0049] Specifically, for each corn sample, its infrared spectrum data is compared with the first feature of the neighborhood set; by calculating the similarity or correlation between the spectrum data and the first feature, the sample feature that matches the first feature can be identified.
[0050] Among them, the feature sequence corresponding to each first feature is composed of sample features that match the first feature. The feature sequence includes each first feature and its corresponding sample feature. These sequences will be used for subsequent analysis and model construction.
[0051] Since each corn sample may be associated with multiple neighborhood sets, multiple feature sequences matching different first features may be obtained. These feature sequences may contain different spectral bands and absorption intensity information, reflecting the diverse characteristics of the corn samples.
[0052] After constructing the feature sequence, further optimization and screening can be performed to ensure that the selected features are the most representative and informative for predicting iNDF content. Various feature selection and extraction techniques can be used to optimize the feature sequence, such as recursive feature elimination (RFE) and model-based feature selection.
[0053] It can be seen that, through this process in this embodiment, it is possible to ensure that the most representative and informative features are extracted from the spectral data of the corn sample, providing a basis for building an accurate prediction model.
[0054] S50. Determine, based on the multiple feature sequences, whether the first feature corresponding to the neighborhood set of each corn sample is a target feature, where the target feature is used to determine the iNDF content data of each corn sample.
[0055] Target features are those that can effectively distinguish corn samples with different iNDF levels. These features have high predictive value in the model and can help accurately determine the iNDF content of corn samples.
[0056] The feature sequence contains the first feature of each corn sample's neighborhood set and its corresponding relationship in the entire sample set. By analyzing these sequences, we can identify which features appear repeatedly in multiple samples, which may be the target features.
[0057] Among them, after the target features are determined, these feature sequences will be used to train prediction models, such as partial least squares regression (PLS) models, for spectral determination of iNDF content in whole-plant silage corn.
[0058] The steps for determining the target characteristic include: characteristic sequence evaluation, which evaluates the characteristic sequence of each corn sample to examine the frequency and consistency of the first characteristic across multiple samples; correlation analysis, which analyzes the correlation between the first characteristic and iNDF content data. The target characteristic should have a high correlation with iNDF content; and statistical testing, which may require statistical testing (such as t-tests, ANOVA, etc.) to determine the statistical significance of the first characteristic at different iNDF levels.
[0059] As can be seen, this example ensures that the most representative and informative features are extracted from the spectral data of corn samples, providing a foundation for building an accurate prediction model. This approach not only improves the accuracy and reliability of the model's predictions of iNDF content, but also enhances the model's interpretability, making the model's predictions easier to understand and trust.
[0060] S60: Perform feature screening based on the target features to obtain training features. The training features are used to train a target model. The target model is used for spectral determination of iNDF content of whole-plant silage corn.
[0061] The goal of target feature screening is to select the most effective and relevant features from all potential spectral features for predicting the iNDF content of whole-plant corn silage. These features are able to best reflect the changes in iNDF content. These feature sequences include the first feature extracted from the neighborhood set and its corresponding feature sequence.
[0062] Among them, the features retained after screening are called training features. These features will be used as input variables to train the model.
[0063] The steps for training the target model include: splitting the dataset containing training features and iNDF content into a training set and a test set; selecting an appropriate algorithm to build the model, such as partial least squares (PLS) regression, support vector machine (SVM), neural network, etc.; using the training set data to train the model and optimizing the model performance by adjusting the model parameters; and using the test set data to verify the model's predictive ability to ensure that the model has good generalization capabilities.
[0064] The trained target model can be used to determine the iNDF content of an unknown whole-plant corn silage sample. This typically involves the following steps: collecting spectral data from the unknown sample; processing the spectral data using the trained model; and outputting the model-predicted iNDF content.
[0065] It can be seen that through the above process in this embodiment, the model finally obtained can quickly and accurately predict the iNDF content of whole-plant corn silage, which has important application value in the fields of agricultural production, feed quality control and scientific research.
[0066] The present invention performs neighborhood processing on a corn sample set, thereby ensuring that samples in a neighborhood set of each corn sample have similar iNDF contents, thereby improving data consistency during data analysis; the first feature of each corn sample neighborhood set obtained represents shared infrared spectral data, which helps to identify and extract the most valuable spectral features for iNDF content prediction; the feature sequences obtained by performing data analysis on the neighborhood set help to build a model with greater generalization capability because they are based on the spectral features of multiple samples rather than a single sample; data matching is performed on each corn sample to obtain multiple feature sequences, which helps to find the similarity of spectral features between different samples, thereby optimizing subsequent data analysis and model training processes; and training features obtained by feature screening can reduce the risk of overfitting during model training because the screened features are more likely to reflect actual changes in iNDF content. That is, using the screened training features to train the target model can improve the model's prediction efficiency and accuracy for spectral determination of iNDF content in whole-plant silage corn.
[0067] In one embodiment, data matching is performed on each corn sample in the corn sample set according to the first feature corresponding to the neighborhood set of each corn sample to obtain multiple feature sequences, including: obtaining a target corn sample to be matched in the corn sample set and the infrared spectrum data to be matched in the target corn sample; calculating the infrared spectrum data to be matched in the target corn sample and the first feature corresponding to the neighborhood set of the target corn sample according to a first preset calculation formula to obtain the similarity between the infrared spectrum data to be matched in the target corn sample and the first feature corresponding to the neighborhood set of the target corn sample; traversing each corn sample in the corn sample set to obtain multiple similarities; and obtaining multiple feature sequences based on the multiple similarities.
[0068] The purpose of data matching is to find sequences that can represent the spectral features of each corn sample. These feature sequences will be used to build a prediction model.
[0069] Among them, the first preset calculation formula is:
[0070]
[0071] in, Indicates the infrared spectrum data to be matched in the selected corn sample The similarity between the kth first feature of the i-th corn sample and the first feature with the largest similarity is selected as the infrared spectrum data to be matched. The first feature to match. That is: .
[0072] in:
[0073]
[0074]
[0075] in, Indicates the infrared spectrum data to be matched in the selected sample set The corresponding spectral band set, represents the kth eigenvector in the i-th corn sample The corresponding spectral band set; Indicates the Jaccard similarity between two sets. It measures the similarity of the first feature by the overlap of the spectral wavelength features in the first feature. The higher the similarity, the more the principal components exhibit the same spectral feature. The larger the value, the greater the similarity between the spectral features. It represents the DTW distance between two sequences, reflecting the distance similarity between the sequences. Jaccard requires the wavelength of the spectrum to be consistent when calculating similarity. However, in reality, there will be certain deviations in the detection spectra of different corns, resulting in fewer overlapping spectral bands in the two sets, which in turn makes the similarity smaller. Therefore, it is necessary to calculate the DTW distance between the two sequences to correct the similarity. The smaller the DTW distance, the higher the similarity between the spectral band sequences, and thus the greater the similarity between the corresponding first features.
[0076] thus, Represents two sets Jaccard similarity between; Represents two sets The DTW distance between Indicates the index of the spectral band. In spectral analysis, each band corresponds to a specific wavelength. The index " ” is used to identify these bands. Specifically, select one first feature of the corn sample, select the first features of other corn samples, and complete the first feature matching of the two corn samples; repeat this process to complete the first feature matching of all corn samples and obtain corresponding multiple matching feature sequences.
[0077] Among them, by calculating the similarity between multiple first features of each sample and the infrared spectrum data to be matched, the first feature with the largest similarity is selected as the matching feature. This process involves traversing each corn sample and comparing the k first features of each sample to select the best matching feature, that is, comparing the infrared spectrum data of each target corn sample with each first feature in its neighborhood set. calculate . Traverse all first features and find the first feature with the greatest similarity.
[0078] Specifically, a first feature is selected from a corn sample as a reference; the reference feature is compared with the first features of other corn samples; the above matching process is repeated until the first features of all corn samples are matched; through the matching process, multiple feature sequences are obtained, each sequence represents the similarity ranking between the first feature of a sample and the first features of other samples in the corn sample set.
[0079] It can be seen that multiple feature sequences can be generated in this embodiment, and each sequence provides an important data basis for subsequent model training and iNDF content prediction.
[0080] In one embodiment, judging whether the first feature corresponding to the neighborhood set of each corn sample is a target feature based on the multiple feature sequences includes: obtaining a target corn sample to be matched in the corn sample set and infrared spectrum data to be judged in the target corn sample; obtaining a target feature sequence based on the infrared spectrum data to be judged in the target corn sample and the multiple feature sequences, wherein the target feature sequence is consistent with the infrared spectrum data to be judged in the target corn sample; obtaining a target neighborhood corn sample corresponding to the target corn sample based on the target corn sample, wherein the target neighborhood corn sample is in the neighborhood set of the target corn sample; and The target feature sequence and the target neighborhood corn sample are compared to obtain a target difference value and a target association value, wherein the target difference value is a difference value between the target neighborhood corn sample corresponding to the target corn sample and the target corn sample under the infrared spectrum data to be judged, and the target association value is an association value between the target neighborhood corn sample corresponding to the target corn sample and the target corn sample under the infrared spectrum data to be judged; according to the target difference value and the target association value, a target accurate value corresponding to the infrared spectrum data to be judged in the target corn sample is obtained, and the target accurate value is used to judge whether the first feature corresponding to the neighborhood set of each corn sample is a target feature.
[0081] Among them, the target corn sample is any corn sample selected from the corn sample set, and this sample will be the object of subsequent analysis and judgment.
[0082] The target feature sequence is determined from multiple feature sequences based on the infrared spectrum data of the target corn sample to be judged, and a target feature sequence that is consistent with the spectrum data is determined. This sequence will serve as the standard for judging the target feature.
[0083] For the target corn sample, find the corn samples in its neighborhood set. These samples will be used to calculate the target difference value and target association value.
[0084] Among them, an infrared spectrum data to be judged in a corn sample is selected, and the target feature sequence corresponding to the infrared spectrum data to be judged is obtained, and the features corresponding to different corn samples are determined; if there is similarity between the features corresponding to corn samples with similar iNDF contents; the features corresponding to different iNDF contents are different, and there is a strong correlation between the change in iNDF content and the change in the feature; then it means that the feature is more likely to be the feature required in the PLS model. The obtained neighboring corn samples are used as the feature change similarity judgment; a feature sequence is selected to obtain the features corresponding to the remaining corn samples in the neighborhood;
[0085]
[0086] Indicates the infrared spectrum data to be matched, Indicates the neighboring corn samples within this corn; Indicates the features in the neighborhood that match the features of the corn; Indicates the similarity between the eigenvectors in the neighborhood, reflecting the difference in absorption between different bands in the spectrum. It represents a summation operation of relevant parameters for all neighboring corn samples within the corn, which is a well-known technology.
[0087] However, there may be normal deviations in the actual spectral wavelength correspondence, so As a weight, the similarity is corrected; Indicates the infrared spectrum data to be matched Features that match the features of the corn in the neighborhood Spectral matching similarity between them; The difference value between the characteristics of this corn and the corn in the neighborhood; The larger the value, the greater the similarity of the features in the neighborhood, and thus the greater the possibility that the matching principal component sequence will express the iNDF content feature.
[0088] Among them, select a corn sample, determine the feature sequence to be judged, and calculate the remaining corn samples The characteristic differences corresponding to this corn sample , the corresponding iNDF content difference is ; Get the corresponding difference sequence , according to the difference sequence, the Pearson correlation coefficient of the corn sample characteristics is obtained ;Will As the correlation value between the changes in principal component characteristics and the changes in iNDF content.
[0089] Through the above process, the correlation value corresponding to each corn sample feature is obtained , feature difference value The stronger the correlation and the greater the similarity, the greater the target accuracy of the principal component as the iNDF content feature. ; .
[0090] It can be seen that in this embodiment, it is possible to determine whether the first feature in the neighborhood set of each corn sample has a strong correlation with the iNDF content, thereby providing accurate feature selection for establishing a spectral measurement model.
[0091] In one embodiment, after obtaining the target accuracy value corresponding to the infrared spectrum data to be judged in the target corn sample based on the target difference value and the target association value, the method further includes: judging whether the target accuracy value is greater than a preset threshold; if the target accuracy value is greater than or equal to the preset threshold, determining that the first feature corresponding to the neighborhood set of each corn sample is the target feature; or, if the target accuracy value is less than the preset threshold, determining that the first feature corresponding to the neighborhood set of each corn sample is not the target feature.
[0092] Among them, the preset threshold is determined in advance based on experimental data, experience or statistical methods, and is used to evaluate whether the target accuracy value is large enough to determine whether the feature has a significant correlation with the iNDF content.
[0093] Optionally, the preset threshold can be 0.7-0.9, that is, if the target accuracy value is a similarity metric (such as cosine similarity or correlation coefficient), it is generally believed that values between 0.7 and 0.9 indicate a higher correlation; or 0.5-0.7, values between 0.5 and 0.7 can indicate a moderate correlation; or 0.3-0.5, which indicates a lower correlation.
[0094] If the target accuracy value is greater than or equal to the preset threshold, this indicates a strong correlation between the feature and the iNDF content. Therefore, the feature can be determined to be a target feature, meaning it is useful for predicting iNDF content. If the target accuracy value is less than the preset threshold, this indicates a weak correlation between the feature and the iNDF content. Therefore, the feature is determined not to be a target feature, meaning it may not be very useful for predicting iNDF content.
[0095] This example ensures that only features highly correlated with iNDF content are incorporated into the final prediction model, thereby improving the model's predictive performance and reliability. This approach effectively filters the most useful information from a large amount of spectral data for spectral determination of iNDF content in whole-plant corn silage.
[0096] In one embodiment, before performing neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain the neighborhood set of each corn sample, the method further includes: performing random selection in the corn sample set to obtain a training sample set and a test sample set; and performing neighborhood processing on the iNDF content data of each corn sample in the training sample set to obtain the neighborhood set of each corn sample.
[0097] Among them, 75% of the samples are randomly selected as training samples, and the remaining 25% of the samples are used as test samples.
[0098] The random selection of samples from the corn sample set is to ensure the representativeness and randomness of the training sample set and the test sample set, and to avoid any potential bias. Random selection can be achieved through a random number generator or random sampling method.
[0099] The training set is used for model training, that is, to establish model parameters and select features. In most cases, the training set accounts for a larger proportion of the total sample size because more data helps the model generalize. The test set is used for model testing, that is, to evaluate the model's predictive performance. The test set does not participate in model training to ensure the objectivity of the evaluation results.
[0100] It can be seen that in this embodiment, by retaining a portion of samples as a test set, the performance of the model on unseen data can be evaluated, thereby avoiding overfitting of the model on the training data, and further ensuring that the model fully utilizes the available data during the training process, while maintaining a certain degree of independence and fairness in the model evaluation.
[0101] In one embodiment, the feature screening according to the target feature to obtain the training feature includes: constructing a plurality of second features of a preset model according to the iNDF content data of each corn sample in the training sample set, the second feature being used to represent the iNDF content data of each corn sample; performing difference calculation on the plurality of second features according to the plurality of feature sequences to obtain an error value between each second feature and the plurality of feature sequences; traversing the plurality of second features to obtain a plurality of error values; confirming that a third feature corresponding to a minimum error value among the plurality of error values is a training feature to be determined, the third feature being in the target feature; and correcting the training feature to be determined to obtain the training feature.
[0102] The second feature is constructed based on the iNDF content data of each corn sample in the training sample set. These features are usually obtained through the partial least squares (PLS) algorithm, which represents the relationship between the iNDF content data and the spectral data. In the specific implementation, the PLS algorithm is used to train all corn samples to build a PLS model and obtain the second feature of the whole set. The second feature can be recorded as , represents the set of all second features obtained by the partial least squares (PLS) algorithm, etc. are all specific elements in the set V_pls, each element represents a second feature obtained by the PLS algorithm, represents the first and second characteristics, Represents the second second feature.
[0103] The difference between multiple feature sequences (based on the first features of the neighborhood set) and multiple second features is calculated to obtain the error value between each second feature and the feature sequence. This step is to quantify the difference between the second features extracted from the PLS model and the feature sequence obtained by the neighborhood method.
[0104] Among them, for the set For each second feature in the dataset, the error between it and each feature sequence is calculated. By traversing the dataset, the error between each second feature and the feature sequence is obtained. During the traversal process, the second feature with the smallest error is found. This feature is considered the training feature to be determined, i.e., the third feature. The third feature is located in the target feature and represents the spectral information most relevant to the iNDF content.
[0105] The determined training features are modified to obtain the final training features. The modification may include noise removal, standardization, or feature weight adjustment.
[0106] In the specific implementation, the PLS model is constructed for all the trained corn samples through the PLS algorithm to obtain the second feature of the PLS model, which is recorded as ; Calculate the difference between the second feature constructed by all corn samples and the second feature sequence. The second feature sequence is a feature sequence constructed by neighborhood samples, that is, it is constructed based on the feature sequences of multiple samples with similar iNDF content to the target sample. Fix a second feature sequence and traverse The error between the kth PLS second feature and the second feature sequence is obtained; the error is usually calculated using the following methods: Euclidean distance, Manhattan distance or mean square error; the PLS with the smallest error is selected as the third feature to complete the matching of the global second feature and the local second feature. The matching error of all PLS third features in is If the error The smaller the error, the higher the similarity between the overall PLS second feature and the second feature of the local corn sample. However, it does not mean that the smaller the error, the PLS component is considered to be the second feature that needs to be selected in the PLS model, because there is normal fluctuation between the global PLS second feature and the local second feature. Selecting the smallest matching error may lead to overfitting, causing the second feature of the PLS model to over-learn the features of the trained corn sample, and the model error is large when facing the iNDF content determination of untrained corn.
[0107] Among them, there is a positive correlation between the global PLS second feature and the local second feature, which means that they should reflect the same spectral information to a certain extent, but they may also have differences, which is due to the difference in feature extraction between the global model and the local neighborhood model.
[0108] It can be seen that in this embodiment, the PLS model is trained using the most relevant and representative features, thereby improving the accuracy and reliability of the model in predicting the iNDF content of corn samples.
[0109] In one embodiment, the correcting the training feature to be determined to obtain the training feature includes: obtaining a second feature to be analyzed and a corresponding feature sequence; determining, based on the second feature to be analyzed and the corresponding feature sequence, a mean iNDF content in the feature sequence corresponding to the second feature to be analyzed; calculating, according to a second preset calculation formula, the mean iNDF content and a corresponding error value in the feature sequence corresponding to the second feature to be analyzed to obtain a determined value to be determined; screening the determined values to be determined, selecting the largest determined value as a target determined value, and the feature corresponding to the target determined value being the training feature.
[0110] Among them, it is necessary to combine the possibility of local principal components on iNDF content and thus modify the selection conditions of principal components in the PLS algorithm.
[0111] Among them, the second feature to be analyzed is extracted from the PLS model , which is the spectral feature related to iNDF content obtained by the PLS algorithm. Determine the feature sequence corresponding to each second feature to be analyzed ( ), this sequence is constructed based on a neighborhood set and contains the features of multiple samples with similar iNDF content to the target sample.
[0112] For each feature sequence corresponding to the second feature to be analyzed, the average value (EP) of the iNDF content of all samples in the sequence is calculated. This average value represents the central trend of the feature sequence in iNDF content.
[0113] Among them, all calculated determination values are screened and the largest determination value is selected as the target determination value. The feature corresponding to the target determination value will be selected as the training feature because it has the highest probability in predicting iNDF content.
[0114] When selecting principal components in the PLS model, the likelihood of local principal components contributing to iNDF content needs to be considered. If the local secondary feature sequence corresponding to a global secondary feature has a stronger performance in iNDF content (i.e., a higher likelihood of iNDF content), and the difference between the global and local features is smaller, then the likelihood of this global secondary feature serving as an iNDF content feature is greater.
[0115] Specifically, select a global second feature to be analyzed , the matching local second feature sequence is , the greater the possibility that the second feature sequence as a whole is iNDF content, and the smaller the difference between the global second feature and the local second feature, the greater the possibility that the global second feature is the iNDF content feature, and thus the greater the possibility that it is the second feature to be selected for the PLS model;
[0116]
[0117] Indicates the error value, represents the mean value of iNDF content in the second characteristic sequence; represents the determined value of the second feature of the model, The number of the th is obtained by the partial least squares (PLS) algorithm. The second feature, Is an integer indicating the A specific location in .
[0118] From this we get The determined values of all second features in ; Select the second feature with the largest determination value as the second feature of the PLS model.
[0119] It can be seen that in this example, the most useful features for predicting iNDF content were accurately selected from the PLS model, thereby improving the prediction performance of the model. That is, the predictive ability of the features and their performance in the local dataset were taken into account, which helped to build a more robust and effective prediction model.
[0120] In one embodiment, after screening the determination values to be determined and selecting the largest determination value as the target determination value, the method further includes: judging whether the target determination value meets the stop condition of feature screening; if the target determination value meets the stop condition of feature screening, stopping the feature screening of the target feature.
[0121] After selecting the maximum deterministic value as the target deterministic value, it is necessary to evaluate whether this value has reached the preset stopping condition. The stopping condition is used to determine when to stop adding new features to avoid overfitting and improve the generalization ability of the model.
[0122] Among them, when using the PLS algorithm for feature selection, the determination value Q corresponding to each selected feature is recorded; the determination value Q reflects the contribution of the feature to the model's prediction ability.
[0123] Among them, the PLS algorithm is used to complete the subsequent feature selection, and the determined value of each global feature selection is recorded ,when If the difference between the current feature and the first feature exceeds 20%, the iteration is stopped, thus completing the feature selection.
[0124]
[0125] in, represents the determined value of the first global feature selection, Indicates the Determined value for sub-global feature selection.
[0126] When the kth selected feature decreases If it is greater than 20%, the feature selection of the PLS algorithm is stopped.
[0127] Specifically, the stopping condition is that the ratio between the target determination value and the initial target determination value decreases to reach a preset ratio threshold.
[0128] Among them, when the degree of decrease of the feature selected for the kth time If the decrease in the determined value exceeds 20%, the feature selection of the PLS algorithm is stopped. That is, if the newly added features no longer significantly improve the predictive ability of the model (that is, the decrease in the determined value exceeds 20%), there is no need to continue adding more features.
[0129] The stopping condition is that the ratio of the target value to the initial target value decreases by a preset ratio threshold, such as 20%. If the target value decreases by more than 20% relative to the initial value, it is considered that further feature addition will no longer improve the model, so feature screening is stopped.
[0130] It can be seen that the feature screening process in this embodiment is effectively controlled, ensuring that the final model only contains the features that contribute most to the prediction of iNDF content, while reducing the risk of overfitting, which helps to build a simple and efficient prediction model.
[0131] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0132] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0133] As another aspect of the present invention, embodiments of the present invention provide a spectral device for measuring the iNDF content of whole-plant corn silage. The spectral device for measuring the iNDF content of whole-plant corn silage can be a software module comprising several instructions stored in a memory. A processor can access the memory and execute the instructions to perform the spectral method for measuring the iNDF content of whole-plant corn silage described in the various embodiments described above.
[0134] See also Figure 2 , Figure 2 Schematic diagram of a spectroscopic device for measuring the iNDF content of whole-plant corn silage provided in an embodiment of the present application. Figure 2 As shown, the spectrometer 200 for measuring the iNDF content of whole-plant corn silage includes:
[0135] The processing unit 201 is used to perform data processing on the corn sample set to be tested, and obtain the iNDF content data and corresponding infrared spectrum data of each corn sample in the corn sample set;
[0136] The processing unit 201 is further configured to perform neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain a neighborhood set of each corn sample, wherein the neighborhood set of each corn sample includes a plurality of corn samples having the same iNDF content data as that of each corn sample;
[0137] An analyzing unit 202 is configured to perform data analysis based on the neighborhood set of each corn sample to obtain a first feature corresponding to the neighborhood set of each corn sample, where the first feature represents infrared spectrum data common to the neighborhood set of each corn sample;
[0138] A matching unit 203 is configured to perform data matching on each corn sample in the corn sample set according to the first feature corresponding to the neighborhood set of each corn sample, to obtain a plurality of feature sequences, wherein the plurality of feature sequences include each first feature and a feature sequence corresponding to each first feature;
[0139] The processing unit 201 is further configured to determine, based on the multiple feature sequences, whether a first feature corresponding to the neighborhood set of each corn sample is a target feature, wherein the target feature is used to determine the iNDF content data of each corn sample;
[0140] The screening unit 204 is configured to perform feature screening based on the target features to obtain training features. The training features are used to train a target model, and the target model is used for spectral determination of iNDF content in whole-plant silage corn.
[0141] The present invention has the following beneficial effects: by performing neighborhood processing on a corn sample set, the present invention ensures that samples in a neighborhood set of each corn sample have similar iNDF contents, thereby improving data consistency during data analysis; the first feature of the obtained neighborhood set of each corn sample represents shared infrared spectral data, which helps to identify and extract the most valuable spectral features for iNDF content prediction; by performing data analysis on the neighborhood set, the feature sequences obtained help to build a model with greater generalization capability because they are based on the spectral features of multiple samples rather than a single sample; by performing data matching on each corn sample to obtain multiple feature sequences, this helps to find the similarity of spectral features between different samples, thereby optimizing subsequent data analysis and model training processes; by obtaining training features through feature screening, the risk of overfitting in the model training process can be reduced because the screened features are more likely to reflect the actual changes in iNDF content, that is, using the screened training features to train the target model can improve the model's prediction efficiency and accuracy for spectral determination of iNDF content in whole-plant silage corn.
[0142] In one embodiment, in performing data matching on each corn sample in the corn sample set according to the first feature corresponding to the neighborhood set of each corn sample to obtain multiple feature sequences, the matching unit 203 is further used to: obtain a target corn sample to be matched in the corn sample set and the infrared spectrum data to be matched in the target corn sample; calculate the infrared spectrum data to be matched in the target corn sample and the first feature corresponding to the neighborhood set of the target corn sample according to a first preset calculation formula to obtain the similarity between the infrared spectrum data to be matched in the target corn sample and the first feature corresponding to the neighborhood set of the target corn sample; traverse each corn sample in the corn sample set to obtain multiple similarities; and obtain multiple feature sequences based on the multiple similarities.
[0143] In one embodiment, in the process of determining whether the first feature corresponding to the neighborhood set of each corn sample is a target feature based on the multiple feature sequences, the processing unit 201 is further configured to: obtain a target corn sample to be matched in the corn sample set and infrared spectrum data to be determined in the target corn sample; obtain a target feature sequence based on the infrared spectrum data to be determined in the target corn sample and the multiple feature sequences, wherein the target feature sequence is consistent with the infrared spectrum data to be determined in the target corn sample; obtain a target neighborhood corn sample corresponding to the target corn sample based on the target corn sample, wherein the target neighborhood corn sample is in the neighborhood of the target corn sample. set; according to the target feature sequence and the target neighborhood corn sample, a target difference value and a target association value are obtained, the target difference value is the difference value between the target neighborhood corn sample corresponding to the target corn sample and the target corn sample under the infrared spectrum data to be judged, and the target association value is the association value between the target neighborhood corn sample corresponding to the target corn sample and the target corn sample under the infrared spectrum data to be judged; according to the target difference value and the target association value, a target accurate value corresponding to the infrared spectrum data to be judged in the target corn sample is obtained, and the target accurate value is used to judge whether the first feature corresponding to the neighborhood set of each corn sample is the target feature.
[0144] In one embodiment, after obtaining the target accurate value corresponding to the infrared spectrum data to be judged in the target corn sample based on the target difference value and the target correlation value, the processing unit 201 is further used to: determine whether the target accurate value is greater than a preset threshold; if the target accurate value is greater than or equal to the preset threshold, determine that the first feature corresponding to the neighborhood set of each corn sample is the target feature; or, if the target accurate value is less than the preset threshold, determine that the first feature corresponding to the neighborhood set of each corn sample is not the target feature.
[0145] In one embodiment, before performing neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain the neighborhood set of each corn sample, the processing unit 201 is further used to: perform random selection in the corn sample set to obtain a training sample set and a test sample set; and perform neighborhood processing on the iNDF content data of each corn sample in the training sample set to obtain the neighborhood set of each corn sample.
[0146] In one embodiment, in the feature screening according to the target feature to obtain the training feature, the screening unit 204 is used to: construct multiple second features of a preset model based on the iNDF content data of each corn sample in the training sample set, where the second feature is used to represent the iNDF content data of each corn sample; perform difference calculation on the multiple second features according to the multiple feature sequences to obtain an error value between each second feature and the multiple feature sequences; traverse the multiple second features to obtain multiple error values; confirm that a third feature corresponding to the minimum error value among the multiple error values is the training feature to be determined, and the third feature is in the target feature; and correct the training feature to be determined to obtain the training feature.
[0147] In one embodiment, in the process of correcting the training feature to be determined to obtain the training feature, the screening unit 204 is configured to: obtain a second feature to be analyzed and a corresponding feature sequence; determine, based on the second feature to be analyzed and the corresponding feature sequence, a mean iNDF content in the feature sequence corresponding to the second feature to be analyzed; calculate, based on a second preset calculation formula, the mean iNDF content and a corresponding error value in the feature sequence corresponding to the second feature to be analyzed to obtain a determined value to be determined; and screen the determined values to be determined, selecting the largest determined value as a target determined value, wherein the feature corresponding to the target determined value is the training feature.
[0148] In one embodiment, after screening the determination values to be determined and selecting the largest determination value as the target determination value, the screening unit 204 is used to: determine whether the target determination value meets the stop condition of feature screening; if the target determination value meets the stop condition of feature screening, stop the feature screening of the target feature.
[0149] In one embodiment, the stopping condition is that the ratio between the target determination value and the initial target determination value decreases to a preset ratio threshold.
[0150] It should be noted that the aforementioned spectral device for measuring the iNDF content of whole-plant corn silage can implement the spectral method for measuring the iNDF content of whole-plant corn silage provided in the embodiments of this application, and possesses the corresponding functional modules and beneficial effects of the method. For technical details not fully described in the embodiments of the spectral device for measuring the iNDF content of whole-plant corn silage, please refer to the spectral method for measuring the iNDF content of whole-plant corn silage provided in the embodiments of this application.
[0151] See also Figure 3 , Figure 3 Schematic diagram of a spectral measurement system for iNDF content of whole-plant corn silage provided in an embodiment of the present application. Figure 3 As shown, the system 300 includes a processor 301 and a memory 302. The processor 301 is in communication with the memory 302.
[0152] Processor 301 is configured to support the system in executing the corresponding functions of the spectroscopic method for determining the iNDF content of whole-plant corn silage in the above-described method embodiment. Processor 301 can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or any combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0153] Specifically, the processor 301 may include a sending card, a receiving card, and a driver chip.
[0154] Memory 302 is used to store program code, etc. Memory 302 may include volatile memory (VM), such as random access memory (RAM); non-volatile memory (NVM), such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or a combination of these types of memory.
[0155] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program includes program instructions, which, when executed by a computer, enable the computer to execute the spectral determination method for the iNDF content of whole-plant corn silage as described in the above embodiment.
[0156] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0157] The above disclosure is merely a preferred embodiment of the present application and certainly cannot be used to limit the scope of rights of the present application.
Claims
1. A spectral method for determining the iNDF content of whole-plant corn silage, characterized in that: A spectral determination system for the iNDF content of whole-plant corn silage is provided, the method comprising: Performing data processing on the corn sample set to be tested to obtain the iNDF content data and corresponding infrared spectrum data of each corn sample in the corn sample set; performing neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain a neighborhood set of each corn sample, wherein the neighborhood set of each corn sample includes multiple corn samples having the same iNDF content data as that of each corn sample; Performing data analysis on the neighborhood set of each corn sample to obtain a first feature corresponding to the neighborhood set of each corn sample, wherein the first feature represents infrared spectrum data common to the neighborhood set of each corn sample; Obtain a target corn sample to be matched and infrared spectral data to be matched in the target corn sample from the corn sample set; calculate the infrared spectral data to be matched in the target corn sample and the first feature corresponding to the neighborhood set of the target corn sample according to a first preset calculation formula to obtain a similarity between the infrared spectral data to be matched in the target corn sample and the first feature corresponding to the neighborhood set of the target corn sample; traverse each corn sample in the corn sample set to obtain multiple similarities; and obtain multiple feature sequences based on the multiple similarities, the multiple feature sequences including each first feature and a feature sequence corresponding to each first feature; Obtain a target corn sample to be matched and infrared spectrum data to be judged in the target corn sample in a corn sample set; obtain a target feature sequence based on the infrared spectrum data to be judged in the target corn sample and multiple feature sequences, and the target feature sequence is consistent with the infrared spectrum data to be judged in the target corn sample; obtain a target neighborhood corn sample corresponding to the target corn sample based on the target corn sample, and the target neighborhood corn sample is in a neighborhood set of the target corn sample; obtain a target difference value and a target correlation value based on the target feature sequence and the target neighborhood corn sample, the target difference value being the difference value between the target neighborhood corn sample corresponding to the target corn sample and the target corn sample under the infrared spectrum data to be judged, and the target correlation value being the correlation value between the target neighborhood corn sample corresponding to the target corn sample and the target corn sample under the infrared spectrum data to be judged; obtain a target accurate value corresponding to the infrared spectrum data to be judged in the target corn sample based on the target difference value and the target correlation value, the target accurate value being used to judge whether the first feature corresponding to the neighborhood set of each corn sample is the target feature, and the target feature being used to judge the iNDF content data of each corn sample; Feature screening is performed based on the target features to obtain training features, which are used to train the target model, which is used for spectral determination of iNDF content in whole-plant corn silage; Among them, the first preset calculation formula is: ; The feature vector representing the feature sequence of the infrared spectrum data to be matched in the selected corn sample The similarity between the kth first feature of the i-th corn sample, Represents two sets The Jaccard similarity between Represents two sets The DTW distance between Indicates the infrared spectrum data to be matched in the selected sample set The corresponding spectral band set, represents the kth eigenvector in the i-th corn sample The corresponding spectral band set; The target difference value is calculated as: ; represents the target difference value, The feature vector representing the feature sequence of the infrared spectrum data to be matched, Indicates the neighboring corn samples of this corn; The feature vector representing the feature sequence in the neighborhood that matches the features of the corn; represents the similarity between the feature vectors of the corn and the corns in the neighborhood, Indicates the summation of relevant parameters for all neighboring corn samples within the corn. Represents the feature vector of the feature sequence of the infrared spectrum data to be matched The feature vector of the feature sequence that matches the feature of the corn in the neighborhood The spectral matching similarity between them; The calculation process of the target correlation value is as follows: select a corn sample, determine the feature sequence to be judged, and calculate the target correlation value of the remaining corn samples. The characteristic differences corresponding to this corn sample , the corresponding iNDF content difference is ; Get the corresponding difference sequence , according to the difference sequence, the Pearson correlation coefficient of the corn sample characteristics is obtained ;Will as the target correlation value between the changes in principal component characteristics and the changes in iNDF content; The calculation formula for the target confirmation value is: .
2. A spectral determination method for the iNDF content of whole-plant corn silage according to claim 1, characterized in that: After obtaining the target accurate value corresponding to the infrared spectrum data to be judged in the target corn sample according to the target difference value and the target correlation value, the method further includes: Determine whether the target accuracy value is greater than a preset threshold; If the target accuracy value is greater than or equal to the preset threshold, the first feature corresponding to the neighborhood set of each corn sample is determined to be the target feature; or, If the target accuracy value is less than the preset threshold, it is determined that the first feature corresponding to the neighborhood set of each corn sample is not the target feature.
3. The spectral determination method for the iNDF content of whole-plant corn silage according to claim 1, characterized in that: Before performing neighborhood processing on the iNDF content data of each corn sample in the training sample set in the corn sample set to obtain a neighborhood set of each corn sample, the method further includes: Randomly select from the corn sample set to obtain a training sample set and a test sample set; Neighborhood processing is performed on the iNDF content data of each corn sample in the training sample set to obtain the neighborhood set of each corn sample.
4. The spectral determination method for the iNDF content of whole-plant corn silage according to claim 1, characterized in that: Feature screening is performed based on the target features to obtain training features, including: Constructing a plurality of second features of a preset model according to the iNDF content data of each corn sample in the training sample set, where the second features are used to represent the iNDF content data of each corn sample; Calculating differences between the plurality of second features according to the plurality of feature sequences to obtain an error value between each second feature and the plurality of feature sequences; Traversing multiple second features to obtain multiple error values; Confirming that a third feature corresponding to a minimum error value among the multiple error values is a training feature to be determined, and the third feature is in the target feature; The determined training features are modified to obtain training features.
5. The spectral determination method for the iNDF content of whole-plant corn silage according to claim 4, characterized in that: The training features to be determined are modified to obtain training features, including: Obtaining a second feature to be analyzed and a corresponding feature sequence; Determining, according to the second feature to be analyzed and the corresponding feature sequence, the mean iNDF content in the feature sequence corresponding to the second feature to be analyzed; Calculate the iNDF content mean and the corresponding error value in the feature sequence corresponding to the second feature to be analyzed according to the second preset calculation formula to obtain a determined value to be determined; The determined values to be determined are screened, and the largest determined value is selected as the target determined value, and the feature corresponding to the target determined value is the training feature.
6. The spectral determination method for the iNDF content of whole-plant corn silage according to claim 5, characterized in that: After screening the determined values to be determined and selecting the largest determined value as the target determined value, the method further includes: Determine whether the target determination value meets the stop condition of feature screening; If the target determination value meets the stop condition of feature screening, the feature screening of the target feature is stopped.
7. The spectral determination method for iNDF content of whole-plant corn silage according to claim 5, characterized in that: The stopping condition is that the ratio between the target determination value and the initial target determination value decreases to a preset ratio threshold.
8. A spectral measurement system for the iNDF content of whole-plant corn silage, characterized in that: The system comprises a memory and a processor, wherein the memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the system for spectral determination of the iNDF content of whole-plant corn silage implements the spectral determination method for the iNDF content of whole-plant corn silage as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Straw comprehensive utilization classification method and system based on machine learning and medium
CN117435986A
Generalized local adaptive fusion regression process based on physicochemical and physiochemical underlying hidden properties for quantitative analysis of molecular based spectroscopic data
US20230267369A1