A method and system for identifying the types of young crops
By collecting and analyzing the odor and environmental data of young seedlings, combining advanced analysis and machine learning technology, a dynamic prediction model is established, and the stability and accuracy of odor sensors in complex environments is solved, fine division and species identification of plant growth stages are realized, and the intelligent development of agriculture is promoted.
Patent Information
- Application Number
- CN202410873044.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-07-01
AI Technical Summary
Under complex environmental conditions, the stability and accuracy of high-sensitivity odor sensors are affected, resulting in inconsistent quality of plant growth data, affecting the accuracy of subsequent data analysis and model training, making it difficult to effectively extract indicator compounds in key growth stages, and the model generalization ability is insufficient.
By collecting odor samples and environmental parameter data of green seedlings, using multi-point synchronous acquisition method and advanced analysis technology, combining machine learning and statistical modeling, a dynamic prediction model is established, and compounds are separated by high-performance liquid chromatography and mass spectrometry technology are used to separate key compound characteristics, and the characteristics of key compounds are extracted, and computer vision data is used to identify and adjust them to build a plant species recognition system.
It realizes fine division and prediction of plant growth stages, improves the accuracy and efficiency of plant identification, provides decision-making support for precise agriculture, promotes the automation and intelligent development of agriculture, and reduces the uncertainty of manual intervention and subjective judgment.
Smart Images

Figure CN118861762B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and particularly to a method and system for identifying the types of young crops. Background Art
[0002] With the rapid development of modern agricultural technology, precision agriculture has gradually become a key means to improve crop yield and quality. In particular, the growth data of young crops regularly collected by highly sensitive odor sensors and environmental monitoring equipment can be used to monitor the growth status of crops in real time, and by analyzing the chemical substances in odor samples, the growth stage and health status of crops can be predicted. Although highly sensitive odor sensors can accurately capture the subtle changes in plant odors, under different environmental conditions (such as outdoor environments with significant temperature and humidity changes), the stability and accuracy of the sensors may be affected, which leads to inconsistent data quality, thus affecting the accuracy of subsequent data analysis and model training. When using the technology of high performance liquid chromatography-mass spectrometry for chemical substance analysis, how to effectively separate key growth stage indicator compounds from complex plant odor samples is a technical challenge. These technical contradictions reflect the problems that may be encountered in practical applications, such as the data consistency problem of sensors in complex environments, the effective extraction problem of key compounds, and the lack of model generalization ability. Solving these problems is crucial for improving the accuracy and practicality of plant growth monitoring. Summary of the Invention
[0003] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method and system for identifying the types of young crops. The method for identifying the types of young crops collects odor samples and environmental parameter data of young crops at different growth stages, and uses a multi-point synchronous collection method and an advanced analysis method to achieve non-invasive fine division and prediction of plant growth stages; combined with machine learning and statistical modeling techniques, a dynamic prediction model is established, which can accurately simulate and predict plant growth changes, provide decision-making support for precision agriculture; it can also evaluate the environmental adaptability of plants, recommend the most suitable growth parameters, and promote the sustainable development of agriculture; combined with odor characteristics and computer vision data, it improves the accuracy of plant identification, which is of great significance for wild plant investigation, species protection and resource utilization, promotes the development of agriculture towards automation and intelligence, helps to build a smart agricultural ecosystem, and realizes the efficient utilization of resources and the development of environment-friendly agriculture.
[0004] A method for identifying the types of young crops according to the present invention includes:
[0005] S1. Collect odor samples and environmental parameter data corresponding to different growth stages of young crops of different plant species, and integrate data streams through a multi-point synchronous collection method to form a sample data set with time stamps;
[0006] S2. Separate the collected odor samples, analyze the components in each separated sample to generate corresponding data streams, and perform qualitative identification using the isotope labeling method;
[0007] S3. Extract the characteristic features of key compounds at each growth stage of the plant, and use the K-means clustering algorithm to classify the identified compounds to form the odor characteristics corresponding to each growth stage;
[0008] S4. Combine the growth stage corresponding to the odor characteristics and the environmental parameter data, establish a first dynamic prediction model and train it, and optimize the model parameters through Bayesian optimization;
[0009] S5. Conduct trend analysis on the odor characteristics, growth stage, and environmental parameter data corresponding to the time stamp, and establish a second dynamic prediction model;
[0010] S6. Combine the odor characteristics corresponding to each growth stage of each plant species and the environmental parameter data corresponding to the time stamp, analyze the standard growth trends of each plant species under different environmental parameter data, establish a plant species recognition model and train it, and perform plant species recognition on a new sample data set without plant species association;
[0011] S7. Analyze the morphological characteristics of the plant in combination with the computer vision data, fuse the odor characteristics with the computer vision data, and adjust the output result of the plant species recognition model.
[0012] Preferably, step S1 further includes:
[0013] The collection contents of the odor samples and the environmental parameter data include temperature, humidity, light intensity, and wind speed;
[0014] The formation of the sample data set includes the following steps:
[0015] According to the seedling growth stage of different plant species, determine the time points for collecting the odor samples and the environmental parameter data, and integrate the odor samples and the environmental parameter data through a multi-point synchronous collection method to obtain a data stream with a time stamp;
[0016] Adopt data preprocessing technology to clean, denoise, and standardize the collected odor samples and environmental parameter data, arrange the data in chronological order using the time stamp information, construct a time series data set, and perform data interpolation operations according to the time stamp to obtain the processed sample data set.
[0017] Preferably, step S2 further includes:
[0018] According to the characteristics of plant species, odor samples released by the collected plants are separated and their components are analyzed by high performance liquid chromatography-mass spectrometry (HPLC-MS) technology to obtain corresponding data streams.
[0019] A high performance liquid chromatography column is used to separate the compounds in the odor samples. The conditions of mobile phase composition, pH value, and column temperature are optimized through orthogonal experimental design, and solvent gradient elution is adopted to achieve separation according to the differences in polarity and hydrophobicity of the compounds.
[0020] The separated samples are introduced into a mass spectrometer, the compounds are ionized by an electron impact ionization source, and the mass-to-charge ratio is measured in a quadrupole mass analyzer to obtain the mass spectrum of the compounds. According to the fragment ion peaks and molecular ion peaks in the mass spectrum, combined with the standard mass spectra in the compound database, the compounds are qualitatively identified by isotope labeling method.
[0021] Preferably, step S3 further includes:
[0022] Mass spectrometry identification data of odor samples at different growth stages of plants are obtained by high performance liquid chromatography-mass spectrometry technology. The principal component analysis method is used to reduce the dimension of the mass spectrometry data. Through orthogonal transformation, the original high-dimensional data are mapped into a set of linearly independent new feature spaces to extract the main variation information of the samples. The new features are sorted according to the variance size, and the first few principal components with the largest variance can explain most of the variation of the original data, forming a low-dimensional feature matrix.
[0023] The linear discriminant analysis method is used to further optimize the principal component features. Through linear transformation, the data are projected into a low-dimensional space, so that the samples at different growth stages are separated as much as possible after projection, while the samples within the same growth stage are aggregated as much as possible, maximizing the between-class difference between samples at different growth stages and minimizing the within-class difference of samples within the same growth stage to obtain the most discriminative feature combination.
[0024] Combined with the recursive feature elimination method, by iteratively removing the features that contribute less to the discrimination result, the key compound marker features highly related to the plant growth stage are screened out to construct an optimal feature subset.
[0025] According to the key compound characteristic signs screened out, the K-means clustering algorithm is used to perform unsupervised classification on the odor samples at different growth stages. First, randomly select K samples as the initial clustering centers, then calculate the distance from each sample to each clustering center, divide the samples into the category with the closest distance, and then recalculate the centroid of each category as the new clustering center. Repeat the above process until the clustering centers no longer change or reach the maximum number of iterations. By optimizing the clustering centers and sample allocation, the samples with similar odor characteristics are divided into the same category, forming the odor characteristic clustering results representing each growth stage of the plant.
[0026] Preferably, step S4 further includes:
[0027] Optimizing the model parameters includes the penalty coefficient and kernel function parameters, so that the support vector machine model associates the sample data set with the corresponding growth stage, and classifies the new sample data set without growth stage association according to the growth stage;
[0028] According to the corresponding relationship between the odor characteristics and the growth stage obtained in step S3, combined with the environmental parameter data, construct a multi-dimensional sample data set as the training input of the first dynamic prediction model;
[0029] Use the support vector machine model as the basis of the first dynamic prediction model, map the sample data to a high-dimensional space through the kernel function, and find the optimal classification hyperplane in the high-dimensional space to realize the non-linear association between the odor characteristics and the growth stage;
[0030] Use the Bayesian optimization method to automatically optimize the key parameters of the support vector machine model.
[0031] Preferably, step S5 further includes:
[0032] The second dynamic prediction model includes an autoregressive integrated moving average model or a long short-term memory neural network to predict the growth stage corresponding to future time points;
[0033] Preprocess the collected odor characteristic data, remove outliers and noise data, extract the key parameters of the odor characteristics, and obtain a standardized odor characteristic data set;
[0034] Clean and standardize the environmental parameter data, remove invalid data, and normalize the data to make the environmental parameter data of different dimensions comparable, forming a standardized environmental parameter data set;
[0035] Align the preprocessed odor characteristic data and the standardized environmental parameter data according to the time stamp, and construct a multi-dimensional time series data set of odor-environment-growth stage as the input of the trend analysis and prediction model;
[0036] Use the autoregressive integrated moving average model to perform trend analysis on multi-dimensional time series data. Through the autocorrelation and trend changes of historical data, establish a second dynamic prediction model between odor characteristics, environmental parameters, and growth stages.
[0037] Preferably, the step S6 further includes:
[0038] According to the odor sample data of each plant species at different growth stages, extract the odor feature vectors of each sample, including the types and contents of key compounds, and associate and store the odor feature vectors with the corresponding plant species, growth stages, and collection timestamps to form a structured odor feature dataset;
[0039] Adopt data preprocessing techniques to clean, denoise, and standardize the odor feature data, eliminate the influence of outliers and missing values, and align and synchronize the odor feature data with the environmental parameter data according to the timestamps to form a complete time series dataset. Each time point contains multi-dimensional information of odor characteristics, environmental parameters, and plant species;
[0040] Use the dynamic time warping algorithm DTW to calculate the distance matrix between the odor feature sequences of different plant species. Then, take the distance matrix as the input, apply the K-means clustering algorithm to perform clustering analysis on the odor feature time series, evaluate the clustering quality of different clustering numbers K through the silhouette coefficient index, select the optimal clustering result, and obtain the standard time series templates representing the growth patterns of different plant species;
[0041] For the standard time series templates in each cluster, use the method of multiple linear regression, with environmental parameters as independent variables and odor characteristics as dependent variables, and establish the regression equation as follows:
[0042] y = β0 + β1X1 + β2X2 + … + β P X P
[0043] Among them, y represents the odor characteristic and is the dependent variable; β i (i = 0, 1…P) are the regression coefficients, used to reflect the influence degree and direction of each environmental parameter on the odor characteristic; X i (i = 1, 2…P) represents the value of the i-th environmental parameter and is the independent variable;
[0044] Solve the regression coefficients by the least squares method, and conduct significance tests and model diagnostics to ensure the effectiveness and interpretability of the regression model, explore the growth laws and trend changes of different plant species under different environmental conditions, construct a plant species recognition model, with odor features and environmental parameters as inputs and plant species as outputs. First, perform feature selection and dimensionality reduction on the two types of features respectively, extract the key feature subsets that contribute the most to species discrimination, then use machine learning algorithms, and achieve non-linear transformation and fusion of features through kernel function mapping or decision tree combination. For multi-classification problems, adopt the "one-versus-rest" or "one-versus-one" strategy to transform multi-classification into multiple binary classification problems for solution, and then obtain the final species label through voting or probability combination. Finally, optimize the hyperparameters of the model through grid search and cross-validation, and train to obtain the optimal-performing plant species discrimination model;
[0045] When a newly collected odor sample dataset is input, first extract its odor feature vector and environmental parameters, and standardize them in the same preprocessing manner as the training data;
[0046] Input the preprocessed feature vector into the trained plant species recognition model, calculate the similarity scores between the sample and each plant species through the decision function of the model, and select the plant species with the highest score as the recognition result.
[0047] Preferably, step S6 further includes:
[0048] Combine the standard growth trends of each plant species corresponding to the new environmental parameter data in the new sample dataset, and use the nearest neighbor algorithm to match the corresponding plant species.
[0049] Preferably, step S7 further includes:
[0050] Obtain the morphological feature data of the plant through computer vision technology, and use image segmentation, object detection, and feature extraction algorithms to convert the plant image into a structured morphological feature vector;
[0051] Utilize the odor data released by the plant to be collected, extract the types and concentration characteristics of the odor components, and form an odor feature vector;
[0052] Input the odor feature sequence and morphological feature sequence into two long short-term memory networks (LSTMs) respectively, and achieve the alignment of the feature sequences in the time dimension through the transmission and update of the hidden state;
[0053] In the spatial dimension, through interpolation and mapping methods, unify the feature data collected by different sensors into the same coordinate system to construct a multi-modal plant feature dataset;
[0054] Using the methods of Relief, Lasso, principal component analysis and linear discriminant analysis, the importance and correlation of odor features and morphological features are evaluated from multiple perspectives, and the feature subset with the greatest contribution to plant species discrimination is selected, and feature dimensionality reduction is achieved;
[0055] Adopt a multi-kernel learning strategy. By defining different kernel functions, odor features and morphological features are processed separately, and selection is made according to the distribution characteristics of the features;
[0056] The outputs of each kernel function are weighted and combined to obtain a unified sample similarity metric;
[0057] The weights of the kernel functions are optimized through the least squares loss with L2 regularization and the Hinge loss multi-kernel learning objective function, and solved by quadratic programming or gradient descent algorithms. By alternately iterating the parameters of the kernel functions and the combined weights until convergence, the deep fusion of odor features and morphological features is achieved.
[0058] The present invention also proposes a young plant species recognition system, which adopts the young plant species recognition method as described above, including:
[0059] A collection module, the collection module includes an odor sensor and an environmental monitoring device, and is used to collect plant odor samples and environmental parameter data;
[0060] A data analysis module, the data analysis module is used to collect plant odor samples and environmental parameter data;
[0061] A feature analysis module, the feature analysis module is used to identify key compounds and classify them to form odor features;
[0062] A model processing module, the model processing module is used to establish and optimize a support vector machine model for growth stage prediction;
[0063] A dynamic analysis module, the dynamic analysis module is used to predict the growth stage and analyze the growth trend of plants;
[0064] A fusion module, the fusion module is used to perform plant species recognition and adjust the output in combination with visual data.
[0065] The young plant species recognition method and system described in the present invention have the following advantages:
[0066] A method for identifying the types of young crops provided by the present invention can finely divide and predict the growth stages of plants by collecting odor samples and environmental parameter data corresponding to different growth stages of young crops of different plant species, integrating data streams through a multi-point synchronous collection method, and adopting advanced analysis methods. It provides a non-invasive growth monitoring means based on plant physiological activities. This monitoring method helps to understand the growth status of crops, optimize planting conditions, and prevent pests and diseases in advance, thereby improving agricultural production efficiency and crop quality. By using machine learning and statistical modeling techniques, combining odor characteristics and environmental parameters, a highly accurate dynamic prediction model is established, which can more accurately simulate and predict the changes in plant growth stages and growth trends under different environmental conditions, providing a powerful decision-making support tool for precision agriculture practice. Through the analysis of the standard growth trends of different plant species under different environmental parameters, the adaptability of plants to environmental changes can be evaluated, and the most suitable growth environmental parameters can be recommended for specific plant species, which helps to screen crop varieties and plan planting areas under the background of climate change, and realize the sustainable development of agricultural production. By combining odor characteristics and computer vision data, not only can the growth stages of plants be identified, but also the plant species can be further identified. This identification model integrating multi-source information improves the accuracy of plant identification, which is of great significance for wild plant surveys, species protection, and the rational utilization of plant resources, promotes the development of agriculture towards automation and intelligence, reduces the uncertainty of manual intervention and subjective judgment, improves the efficiency and scientific nature of agricultural management, helps to build a smart agricultural ecosystem, and realizes the efficient utilization of resources and the development of environment-friendly agriculture. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a flowchart of the method for identifying the types of young crops according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] As Figure 1 shown, the method for identifying the types of young crops according to the present invention includes:
[0069] S1. Collect odor samples and environmental parameter data corresponding to different growth stages of young crops of different plant species, and integrate the data streams through a multi-point synchronous collection method to form a sample data set with time stamps. Specifically, use odor sensors and environmental monitoring devices to collect odor samples and environmental parameter data corresponding to different growth stages of young crops of different plant species, and integrate the data streams of the odor sensors and environmental monitoring devices through a multi-point synchronous collection method to form a sample data set with time stamps.
[0070] S2. Separate the collected odor samples, analyze the components in each separated sample to generate corresponding data streams, and perform qualitative identification using isotope labeling method; As can be seen from step S1 above, for the same plant species among different plant species, the corresponding odor sensors are used to separate the collected odor samples through high performance liquid chromatography - mass spectrometry technology, analyze the components in each separated sample to generate the data stream output by the odor sensors. Among them, identifying compounds is achieved by measuring the mass - to - charge ratio using a mass spectrometer, specifically including ester and phenolic compounds. Using the isotope labeling method for qualitative identification is based on a professional compound database.
[0071] S3. Extract the signature features of key compounds at each growth stage of the plant, and use the K - means clustering algorithm to classify the identified compounds to form odor features corresponding to each growth stage; Specifically, use the principal component analysis method and the linear discriminant analysis method to process the identification data of the mass spectrometer, combined with the recursive feature elimination method, extract the signature features of key compounds at each growth stage of the plant, and use the K - means clustering algorithm to classify the identified compounds to form odor features corresponding to each growth stage.
[0072] S4. Combine the growth stage corresponding to the odor features and the environmental parameter data, establish and train the first dynamic prediction model, and optimize the model parameters through the Bayesian method.
[0073] S5. Perform trend analysis on the odor features, growth stages, and environmental parameter data corresponding to the time stamps, and establish the second dynamic prediction model.
[0074] S6. Combine the odor features corresponding to each growth stage of each plant species and the environmental parameter data corresponding to the time stamps, analyze the standard growth trends of each plant species under different environmental parameter data, establish and train a plant species recognition model, and perform plant species recognition on a new sample data set that does not contain plant species associations.
[0075] S7. Analyze the morphological characteristics of the plant in combination with computer vision data, fuse the odor features with the computer vision data, and adjust the output results of the plant species recognition model.
[0076] Furthermore, in this embodiment, step S1 further includes:
[0077] The collection content of odor samples and environmental parameter data includes temperature, humidity, light intensity, and wind speed.
[0078] The formation of the sample data set includes the following steps:
[0079] According to the growth stages of young plants of different plant species, determine the time points for collecting odor samples and environmental parameter data, and integrate the odor samples and environmental parameter data through multi-point synchronous collection to obtain a data stream with timestamps. As mentioned above, the odor samples and environmental parameter data corresponding to different growth stages of young plants of different plant species are collected using odor sensors and environmental monitoring devices. Therefore, determine the time points for collecting odor samples and environmental parameter data, and integrate the odor sensors and environmental monitoring devices through multi-point synchronous collection to obtain a data stream with timestamps;
[0080] Adopt data preprocessing techniques to clean, denoise, and standardize the collected odor samples and environmental parameter data. Use the timestamp information to arrange the data in chronological order, construct a time series dataset, and calculate the sampling interval based on the timestamps to perform data interpolation operations to obtain a processed sample dataset;
[0081] After obtaining a high-quality sample dataset, through feature extraction algorithms, principal component analysis (PCA) or independent component analysis (ICA) can be selected to extract the key components that can reflect the plant odor characteristics from the sample dataset, and select a suitable algorithm in combination with the characteristics of the plant odor data to form a feature vector;
[0082] Adopt machine learning algorithms, and the machine learning algorithms can be selected as support vector machine (SVM), random forest (RandomForest), or neural network to train the feature vector, establish an odor sample classification model. When selecting the algorithm, factors such as the data volume, the number of features, and the computational complexity need to be considered, and the model is optimized and evaluated through cross-validation methods;
[0083] Deploy the optimized odor sample classification model to edge computing devices. The edge computing devices can be selected as Raspberry Pi. By converting the trained model into a format suitable for edge devices (such as TensorFlowLite) and combining with sensor data for inference, real-time classification of odor samples is achieved;
[0084] Through the Internet of Things platform, upload the data and classification results collected by the edge computing devices to the cloud for big data analysis and visualization display, providing data support for the research on plant odor characteristics. Based on the data analysis results on the cloud, further optimize the odor sample dataset, including removing abnormal samples and balancing the data category distribution, improving the quality and representativeness of the dataset. At the same time, continuously update and improve the odor sample classification model to improve the accuracy and generalization ability of the model, laying a foundation for subsequent plant odor feature analysis and applications;
[0085] Examples are as follows:
[0086] During the seedling, growth, and maturity stages of moss plants, odor samples and environmental parameter data were collected every two days, with each collection lasting 30 minutes. Through a multi-point synchronous collection method integrating four TGS2602 gas sensors and DHT11 temperature and humidity sensors, data streams were obtained at a frequency of once per second, and each data point was timestamped by an NTP server. The collected data was denoised, with median filtering algorithm used to remove spike noise, and linear interpolation algorithm used to fill in missing data. Then, the data points were sorted according to the timestamps to construct a time series dataset;
[0087] The PCA algorithm was used to extract features from the odor sensor data. By calculating the eigenvalues and eigenvectors of the covariance matrix, the top 3 principal components with a proportion of over 95% were selected as the odor feature vectors;
[0088] The SVM algorithm was used to classify the odor feature vectors. The RBF kernel function was adopted, and the optimal penalty coefficient C and kernel function parameter gamma were selected through 5-fold cross-validation, achieving a classification accuracy of 92% on the test set. The trained SVM model was converted to the TensorFlow Lite format and deployed on a Raspberry Pi 4 device. The odor sensor was connected through the I2C interface to collect and classify odor data in real-time, and the results were uploaded to the Alibaba Cloud IoT platform via the MQTT protocol;
[0089] In the cloud, Hive was used to perform ETL processing on the odor sample data. By using SQL statements to count the sample quantity distribution of each category, it was found that the proportion of odor samples of moss plants during the growth stage was relatively low, only 18%. Therefore, it was necessary to increase the data collection frequency during the growth stage to balance the dataset. At the same time, the stochastic gradient descent algorithm was used to perform online learning on the SVM model, continuously updating the model parameters, and the classification accuracy on the newly collected odor samples was improved to 95%. Finally, a high-quality time series dataset containing 1000 odor samples, covering the entire growth process of moss plants, was constructed, laying a data foundation for subsequent research on the variation law of the odor characteristics of moss plants.
[0090] Furthermore, in this embodiment, step S2 further includes:
[0091] According to the characteristics of plant species, the odor samples released by the plants were collected, and the odor samples were separated and analyzed for components through high-performance liquid chromatography-mass spectrometry technology to obtain the corresponding data streams; that is, the odor samples were separated and analyzed for components through high-performance liquid chromatography-mass spectrometry technology to obtain the data streams of the odor sensors;
[0092] The compounds in the odor sample are separated using a high-performance liquid chromatography column. The conditions of the mobile phase composition, pH value, and column temperature are optimized through orthogonal experimental design, and solvent gradient elution is adopted to achieve separation based on the differences in the polarity and hydrophobicity of the compounds;
[0093] The separated sample is introduced into a mass spectrometer. The compounds are ionized by an electron impact ionization source and the mass-to-charge ratio is measured in a quadrupole mass analyzer to obtain the mass spectrum of the compounds. Based on the fragment ion peaks and molecular ion peaks in the mass spectrum, combined with the standard mass spectra in the compound database, the compounds are qualitatively identified using the isotope labeling method;
[0094] Specifically, first, according to the structural characteristics of the compounds, appropriate isotope labeling reagents are selected, such as deuterated reagents or 13C-labeled reagents; then, the labeling reagent is mixed with the odor sample and a derivatization reaction is carried out under certain conditions; finally, the derivatized product is analyzed by mass spectrometry, and based on the mass-to-charge ratio differences of the isotope peaks, the molecular weights and structural information of the key compounds of esters and phenols are determined;
[0095] In the process of mass spectrometry data processing, the isotope peak area normalization method is used for quantitative analysis of the compounds. The isotope peak areas corresponding to each compound are extracted, and the sum of the isotope peak areas of each compound is calculated; then, the isotope peak area of each compound is divided by the total peak area to obtain the normalized relative content, generating a data matrix reflecting the component characteristics of the odor sample;
[0096] By comparing the differences in the compound compositions of odor samples of different plant species and combining the information on the biosynthetic pathways of the compounds, the species-specific markers of the odors are judged, and a corresponding relationship database between odor compounds and plant species is constructed. To further explore the relationship between odor components and plant growth conditions, plant growth environment parameters can be collected simultaneously during odor sampling. The plant growth environment parameters include temperature, humidity, and light, and the morphological indexes of the plants are measured. The morphological indexes include plant height, leaf area, and chlorophyll content;
[0097] The environmental parameters and morphological indexes are correlated with the odor component data to explore the variation rules of odor components under different growth conditions, laying a foundation for the subsequent establishment of a prediction model for the relationship between odor components and plant growth conditions;
[0098] Examples are as follows:
[0099] During the blooming period of roses, odor samples of 5 flowers were collected at 10 am every day. The SPME fiber was used to adsorb for 30 minutes, and then the fiber was inserted into the injection port of the GC-MS system. Helium was used as the carrier gas with a flow rate of 1.0 mL / min. The chromatographic column was an HP-5MS capillary column (30 m × 0.25 mm × 0.25 μm). The programmed temperature conditions were as follows: maintained at 50 °C for 2 min, heated to 150 °C at a rate of 10 °C / min, then heated to 300 °C at a rate of 5 °C / min, and maintained for 10 min. The EI ionization source was used with an ionization energy of 70 eV, and the mass scanning range was m / z 50 - 500;
[0100] By retrieving through the NIST mass spectrometry library and comparing with standards, the main components in the rose odor were identified as β-ionone, nerol, and eugenol. Using 13C-labeled β-ionone as the internal standard, isotope dilution method was used for its quantitative analysis, and its absolute content was determined to be 5.2 μg / g. At the same time of odor sampling, a temperature and humidity recorder and a quantum photometer were used to measure the flower growth environment parameters. It was found that when the temperature was 25 °C, the relative humidity was 60%, and the light intensity was 10000 lux, the content of β-ionone was the highest;
[0101] The partial least squares regression (PLSR) algorithm was used to establish a prediction model of environmental parameters and odor component contents. The optimal number of principal components was determined to be 3 through cross-validation, and the prediction correlation coefficient R2 was 0.85;
[0102] Using this model, the contents of key components of rose odor can be predicted according to environmental parameters, providing theoretical guidance for rose cultivation and harvesting.
[0103] Furthermore, in this embodiment, step S3 further includes:
[0104] The mass spectrometry identification data of odor samples at different growth stages of plants were obtained by high performance liquid chromatography-mass spectrometry (HPLC-MS) technology. The principal component analysis method was used to reduce the dimension of the mass spectrometry data. Through orthogonal transformation, the original high-dimensional data was mapped into a set of linearly independent new feature spaces, and the main variation information of the samples was extracted. The new features were sorted according to the variance size. The first few principal components with the largest variance could explain most of the variation of the original data, forming a low-dimensional feature matrix;
[0105] The linear discriminant analysis method was used to further optimize the principal component features. Through linear transformation, the data was projected into a low-dimensional space, so that the samples at different growth stages were separated as much as possible after projection, while the samples within the same growth stage were aggregated as much as possible, maximizing the between-class difference between samples at different growth stages and minimizing the within-class difference of samples within the same growth stage, obtaining the feature combination with the strongest discriminative power;
[0106] Combined with the recursive feature elimination method, by iteratively removing features that contribute less to the discrimination result, the key compound signature features highly related to the plant growth stage are screened out to construct an optimal feature subset;
[0107] According to the screened key compound signature features, the K-means clustering algorithm is used to perform unsupervised classification on the odor samples at different growth stages. First, randomly select K samples as the initial clustering centers, then calculate the distance from each sample to each clustering center, and divide the samples into the nearest category. Then recalculate the centroid of each category as the new clustering center, and repeat the above process until the clustering centers no longer change or reach the maximum number of iterations. By optimizing the clustering centers and sample allocation, the samples with similar odor features are divided into the same category to form the odor feature clustering results representing each growth stage of the plant;
[0108] After forming the odor feature clustering results representing each growth stage of the plant, use the decision tree model to analyze each clustering result. By recursively selecting the best splitting features, divide the samples into different leaf nodes until the samples in the leaf nodes belong to the same growth stage or reach the predetermined stopping condition. By analyzing the structure and key nodes of the decision tree, extract the mass spectrometry feature patterns of each growth stage, determine the key compound composition and the change law of relative content that can reflect different growth stages of the plant, and establish the corresponding relationship between the odor features and the growth stage;
[0109] Using the clustering results and decision tree feature rules, identify and predict the growth stage of unknown odor samples. For a newly collected odor sample, extract its mass spectrometry features, judge the growth stage category it belongs to by comparing with the clustering centers, or input its features into the trained decision tree model, and predict the growth stage according to the key compound combination rule to achieve real-time classification of odor features and rapid judgment of the plant growth state;
[0110] The example is as follows:
[0111] Taking the rose as an example, collect 30 odor samples for each of the three stages of the bud stage, full bloom stage, and withering stage, and perform mass spectrometry analysis using a GC-MS system to obtain an 180-dimensional mass spectrometry data matrix. Use the principal component analysis method to reduce the dimension of the data. By eigenvalue decomposition and eigenvector sorting, select the first 5 principal components, which cumulatively explain more than 85% of the original variation. Then use the linear discriminant analysis method to optimize the principal components. By maximizing the Fisher discriminant criterion, 3 discriminant functions are obtained, and the discriminant accuracy rate on the training set reaches 95%;
[0112] Combined with the recursive feature elimination method, the optimal feature subset is determined through cross-validation, and 20 key compounds highly related to the growth stage of roses are screened out, including β-ionone, nerol, and benzyl alcohol. The K-means clustering algorithm is used to classify the key compounds. The optimal number of clusters is determined to be 3 by the elbow method, and the silhouette coefficient is used to evaluate the clustering quality. The average silhouette coefficient is 0.82, indicating a stable clustering result. The CART decision tree algorithm is used to extract rules for each cluster. The optimal splitting feature is selected through the Gini index to generate a decision tree with a tree depth of 4. The mass spectrometry feature laws of each growth stage are extracted from the key nodes of the tree. For example, the bud stage is characterized by high β-ionone content and low nerol content, and the full-bloom stage is characterized by high benzyl alcohol content;
[0113] Ten unknown odor samples are predicted. First, the contents of their 20 key compounds are extracted, and their belonging clusters are determined by calculating the Euclidean distance. Then, the compound contents are input into the decision tree model for growth stage judgment, and the prediction accuracy reaches 90%, verifying the effectiveness of this method.
[0114] Furthermore, in this embodiment, step S4 further includes:
[0115] Optimizing the model parameters including the penalty coefficient and kernel function parameters to associate the support vector machine model with the sample data set and the corresponding growth stage, and classifying the new sample data set without growth stage association according to the growth stage;
[0116] According to the corresponding relationship between the odor characteristics and the growth stage obtained in step S3, combined with the environmental parameter data, a multi-dimensional sample data set is constructed as the training input of the first dynamic prediction model;
[0117] Using the support vector machine model as the basis of the first dynamic prediction model, the sample data is mapped to a high-dimensional space through the kernel function, and the optimal classification hyperplane is searched in the high-dimensional space to achieve the non-linear association between the odor characteristics and the growth stage;
[0118] Using the Bayesian optimization method to automatically optimize the key parameters of the support vector machine model;
[0119] During the parameter optimization process, the sample data set is divided into a training set and a validation set by 5-fold cross-validation, the search ranges of the penalty coefficient and kernel function parameters are determined, the performance of the candidate parameters is evaluated by maximizing the marginal likelihood estimation, and the parameter combination with the highest accuracy on the validation set is selected as the optimal solution;
[0120] Based on the optimized support vector machine model, all the sample data sets are used for training to obtain the final first dynamic prediction model for classifying and predicting the growth stage of newly collected odor samples;
[0121] When a new sample data set is input, first extract its odor features and environmental parameters, and perform standardization and normalization in the same preprocessing manner as the training data. Then, input the preprocessed new sample data into the trained support vector machine model, and convert the output of the decision function of the model into a probability value through the Platt scaling method. The calculation of Platt scaling is to convert the output z of the decision function into a probability value p through the Sigmoid function. The formula is as follows:
[0122] p = 1 / [1 + e (-Az-B)
[0123] where A and B are scaling parameters obtained by maximum likelihood estimation;
[0124] According to the KL divergence between the probability value of the sample and the true probability distribution of each category, obtain the similarity score. The smaller the KL divergence, the more similar the sample is to the category. Select the growth stage with the highest similarity score as the prediction result;
[0125] To further improve the reliability of prediction and the adaptability of the model, incremental learning can be adopted. When the new sample data set reaches a certain magnitude, trigger model retraining; when the prediction accuracy is lower than the preset threshold, start hyperparameter tuning; when new plant species or growth stages appear, make adaptive modifications to the model structure, add new category nodes, and on the basis of retaining the original knowledge, only train the newly added data to achieve the continuous evolution of the model;
[0126] At the same time, the temporal law of plant growth can also be used to check the rationality of the prediction results, filter out obvious misclassifications, and for prediction results with low confidence, manual annotation or intervention can be requested, and the feedback results are added to the training data set to further improve the robustness and interpretability of the model, forming an adaptive and continuously learning intelligent prediction system to provide accurate and real-time decision support for plant growth monitoring and agricultural production management;
[0127] The example is as follows:
[0128] Taking roses as an example, collect 100 odor samples at different growth stages. Each sample contains the contents of 10 main odor components and 5 environmental parameters, and construct a 1000-dimensional sample feature vector. Randomly divide the sample data set into a training set and a test set according to a ratio of 8:2. Use a Gaussian kernel function to construct a support vector machine model, and determine the search range of the penalty coefficient C as [0.1, 100] and the search range of the kernel function parameter γ as [0.001, 1] through 5-fold cross-validation;
[0129] Using the Bayesian optimization algorithm, set the initial number of iterations to 10, the acquisition function to Gaussian process, and the optimization objective to the average accuracy of cross-validation;
[0130] After 20 rounds of iteration, the optimal parameter combination is obtained as C = 15.8 and γ = 0.04. At this time, the accuracy of the model on the test set is 95.2%;
[0131] For a newly collected rose scent sample, extract the contents of 10 scent components and 5 environmental parameters of it, construct a feature vector, convert the output distance of the support vector machine model into a posterior probability through Platt scaling, and calculate the KL divergence between the sample and 5 growth stages. The 3rd growth stage with the smallest divergence value, that is, the bud stage, is obtained as the prediction result;
[0132] After continuously collecting 50 new samples, retrain the support vector machine model and update the parameters of Platt scaling. It is found that the accuracy of the model on the new samples is improved to 97.1%;
[0133] At the same time, using the prior knowledge of plant growth, perform a secondary discrimination on the samples predicted to be in the withering stage. If the collection time is earlier than the bud stage, then correct its prediction result to the bud stage, further improving the accuracy and interpretability of the model.
[0134] Furthermore, in this embodiment, step S5 further includes:
[0135] The second dynamic prediction model includes an autoregressive integrated moving average model or a long short-term memory neural network to predict the growth stage corresponding to a future time point;
[0136] Preprocess the collected scent feature data, remove outliers and noise data, extract the key parameters of the scent features, and obtain a standardized scent feature data set; As can be seen from the above, by using a scent sensor to obtain the scent feature data in the plant growth process, and collecting the environmental parameter data corresponding to the plant growth stage through an environmental monitoring device, including temperature, humidity, and light intensity, synchronize the scent features and environmental parameter data according to the collected timestamps;
[0137] Clean and standardize the environmental parameter data, remove invalid data, and normalize the data to make the environmental parameter data of different dimensions comparable, forming a standardized environmental parameter data set;
[0138] Align the preprocessed scent feature data and the standardized environmental parameter data according to the timestamp, and construct a multi-dimensional time series data set of scent - environment - growth stage as the input of the trend analysis and prediction model;
[0139] Use the autoregressive integrated moving average model to perform trend analysis on multi-dimensional time series data, and establish a second dynamic prediction model between odor characteristics, environmental parameters, and growth stages through the autocorrelation and trend changes of historical data;
[0140] The long short-term memory neural network can also be used to construct a growth stage prediction model. Taking multi-dimensional time series data as input, through the learning and memory ability of the network, capture the long-term and short-term dependencies between odor characteristics, environmental parameters, and growth stages;
[0141] During the training process of the prediction model, the cross-validation method is used to evaluate and optimize the model. By adjusting the hyperparameters and network structure of the model, improve the prediction accuracy and generalization ability of the model;
[0142] According to the trained prediction model, input the odor characteristics and environmental parameter data at future time points, and through the inference and prediction of the model, obtain the growth stage prediction results at the corresponding time points;
[0143] Compare the prediction results with the actual growth stages to evaluate the accuracy and reliability of the prediction model;
[0144] If there is a large deviation between the prediction results and the actual situation, the model needs to be further optimized and improved, such as increasing training data and adjusting the model structure to improve the prediction accuracy;
[0145] The example is as follows:
[0146] When the odor sensor obtains the odor characteristic data during the plant growth process, a gas sensor array based on metal oxide semiconductors can be used. Through the combined response of multiple sensors, selective detection of different odor molecules can be achieved. For example, an array composed of 6 TGS sensors can be used to detect methane, ethanol, acetone, ammonia, hydrogen sulfide, and carbon monoxide odor molecules respectively;
[0147] Collect the environmental parameter data corresponding to the plant growth stage through environmental monitoring equipment. For example, use the DHT11 temperature and humidity sensor to collect temperature and humidity data, use the BH1750 light intensity sensor to collect light intensity data, and synchronize the odor characteristics and environmental parameter data through timestamps;
[0148] When preprocessing the collected odor feature data, the median filtering algorithm can be used to remove outliers and noise data. By setting a threshold, data points outside the normal range are regarded as outliers and excluded. Then, key parameters of the odor features are extracted. For example, the odor concentration can be calculated through the correspondence between the sensor response value and the standard gas concentration, and the odor type can be discriminated through feature extraction algorithms such as the principal component analysis method to obtain a normalized odor feature data set. The environmental parameter data is cleaned and standardized to remove invalid data, such as values of temperature data that exceed the measurement range.
[0149] The data is normalized to map environmental parameter data in different dimensions into the interval [0, 1] to make it comparable. The preprocessed odor feature data and the standardized environmental parameter data are aligned according to the time stamp to construct a multi-dimensional time series data set of odor-environment-growth stage.
[0150] The autoregressive integrated moving average model is used to perform trend analysis on the multi-dimensional time series data. Through the autocorrelation and trend changes of historical data, a dynamic association model between odor features, environmental parameters, and growth stages is established.
[0151] The ARIMA(1, 1, 1) model is used to capture the short-term correlation and long-term trend of the time series data through first-order autoregression, first-order differencing, and first-order moving average, and predict the growth stage at future time points.
[0152] When using a long short-term memory neural network to construct a growth stage prediction model, a three-layer LSTM network structure can be adopted. The input layer receives multi-dimensional time series data, the hidden layer learns the long-term and short-term dependencies of the data through 128 LSTM units, and the output layer outputs the probability distribution of the growth stage through the Softmax function.
[0153] During the model training process, the 5-fold cross-validation method is adopted. The data set is divided into a training set and a validation set, and the hyperparameters of the model, including the learning rate and batch size, are optimized through grid search to improve the prediction accuracy and generalization ability of the model.
[0154] According to the trained prediction model, by inputting the odor features and environmental parameter data at future time points, through the inference and prediction of the model, the growth stage prediction results at the corresponding time points are obtained. The prediction results are compared with the actual growth stages, and evaluation indicators such as prediction accuracy, precision, and recall are calculated.
[0155] If the prediction accuracy rate is lower than 90%, it is necessary to further optimize and improve the model, such as increasing the amount of training data, adjusting the number of network layers and neurons, and introducing regularization methods to improve the accuracy and reliability of the prediction; through continuous iterative optimization, a stable and efficient plant growth stage prediction model is finally obtained to provide decision-making support for smart agriculture.
[0156] Further, in this embodiment, step S6 further includes:
[0157] According to the odor sample data of each plant species at different growth stages, extract the odor feature vector of each sample, including the types and contents of key compounds, and associate and store the odor feature vector with the corresponding plant species, growth stage, and collection timestamp to form a structured odor feature dataset;
[0158] Adopt data preprocessing techniques to clean, denoise, and standardize the odor feature data, eliminate the influence of outliers and missing values, and align and synchronize the odor feature data with the environmental parameter data according to the timestamp to form a complete time series dataset, where each time point contains multi-dimensional information such as odor features, environmental parameters, and plant species;
[0159] Use the dynamic time warping algorithm DTW to calculate the distance matrix between the odor feature sequences of different plant species, and then use the distance matrix as the input, apply the K-means clustering algorithm to perform clustering analysis on the odor feature time series, evaluate the clustering quality of different clustering numbers K through the silhouette coefficient index, select the optimal clustering result, and obtain the standard time series template representing the growth patterns of different plant species;
[0160] For the standard time series template in each cluster, use the method of multiple linear regression, with environmental parameters as independent variables and odor features as dependent variables, to establish the following regression equation:
[0161] y = β0 + β1X1 + β2X2 + … + β P X P
[0162] where y represents the odor feature, which is the dependent variable; β i (i = 0, 1…P) are the regression coefficients, used to reflect the influence degree and direction of each environmental parameter on the odor feature; X i (i = 1, 2…P) represents the value of the i-th environmental parameter, which is the independent variable;
[0163] The regression coefficients are solved by the least squares method, and significance tests and model diagnostics are carried out to ensure the effectiveness and interpretability of the regression model, explore the growth laws and trend changes of different plant species under different environmental conditions, construct a plant species recognition model, with odor characteristics and environmental parameters as inputs and plant species as outputs. First, feature selection and dimensionality reduction are respectively carried out on the two types of features to extract the key feature subsets that contribute the most to species discrimination. Then, machine learning algorithms are used, and non-linear transformation and fusion of features are achieved through kernel function mapping or decision tree combination. For multi-classification problems, the "one-versus-rest" or "one-versus-one" strategy is adopted to transform multi-classification into multiple binary classification problems for solution, and the final species label is obtained through voting or probability combination. Finally, the hyperparameters of the model are optimized through grid search and cross-validation, and the plant species discrimination model with the optimal performance is trained;
[0164] When a newly collected odor sample dataset is input, first extract its odor feature vector and environmental parameters, and standardize them in the same preprocessing manner as the training data;
[0165] Input the preprocessed feature vector into the trained plant species recognition model, calculate the similarity scores between the sample and each plant species through the decision function of the model, and select the plant species with the highest score as the recognition result;
[0166] Considering the diversity and timeliness of plant species recognition samples, a time series database (such as InfluxDB) or a NoSQL database (such as MongoDB) is used to store sample data. The former is suitable for storing structured data indexed by timestamps and supports efficient range queries and aggregation analysis. The latter is suitable for storing semi-structured and unstructured data and supports flexible data schemas and dynamic expansion;
[0167] When inserting and querying data, indexes are established according to the timestamp and species label key fields to improve the efficiency of data retrieval and analysis, and the model is retrained and optimized regularly to continuously update and expand the sample library for plant species recognition and improve the recognition accuracy and generalization ability;
[0168] The example is as follows:
[0169] Taking three kinds of flowers, rose, jasmine, and osmanthus, as examples, 100 odor samples are respectively collected at three stages: bud stage, full bloom stage, and decline stage. The content of 20 key volatile components is extracted from each sample, and at the same time, 10 environmental parameters such as temperature, humidity, and light intensity are recorded during collection, forming a structured odor feature dataset containing 30,000 records;
[0170] The Z-score normalization method is used to dimensionless process the odor characteristics and environmental parameters, and the Gaussian smoothing kernel function is used to supplement the missing values. Then, the FastDTW algorithm is used to calculate the distance matrix between the odor characteristic sequences of different plant species at different growth stages, and the parameters are set as w = 50 and δ = 10. The distance matrix is input into the K-means clustering algorithm, where K ranges from 2 to 10, and the clustering quality is evaluated by the Silhouette Coefficient. Finally, the optimal number of clusters is determined to be 3, corresponding to the typical odor characteristic templates of three plants: rose, jasmine, and osmanthus;
[0171] For the odor characteristic template of each plant, a multiple linear regression model is constructed between temperature, humidity, light intensity, and the content of key volatile components. The least squares method is used to estimate the regression coefficients, and the significance of the model and the contribution degree of each environmental parameter are judged through F-test and t-test. The results show that temperature and light intensity have the most significant impact on odor characteristics;
[0172] In the training of the plant species recognition model, the Pearson correlation coefficient method is first used to screen the odor characteristics, and the 10 characteristics with the highest correlation with the species label are selected. Then, the principal component analysis (PCA) method is used to reduce the dimension of the environmental parameters, and 3 principal components are extracted;
[0173] The two types of characteristics are input into the SVM and RF models. Among them, the SVM uses the Gaussian radial basis kernel function, and the RF uses 500 decision trees. The hyperparameters such as C, γ, and the maximum depth of the tree are optimized by 10-fold cross-validation. Finally, the recognition accuracy of the SVM model on the test set reaches 95%, which is better than 92% of the RF;
[0174] When a new odor sample is collected, first extract the contents of its 20 key components and 10 environmental parameter values, and perform the same standardization process. Then, input the standardized characteristics into the trained SVM model, and through the "one-versus-one" multi-classification strategy, obtain the probability values of the sample belonging to the three flower species respectively, and take the category with the highest probability as the final recognition result;
[0175] At the same time, information such as the odor characteristics, environmental parameters, collection time, and recognition results of the sample are stored in the MongoDB database in JSON format, and a composite index based on the timestamp and species label is created. The recognition model is retrained every 7 days to continuously optimize the model performance.
[0176] Furthermore, in this embodiment, step S6 further includes:
[0177] Combined with the standard growth trends of each plant species corresponding to the new environmental parameter data in the new sample dataset, the nearest neighbor algorithm is used to match the corresponding plant species;
[0178] Specifically, according to the newly collected sample data set, obtain the environmental parameter data and the corresponding plant species information therein, and use this data as the training set and the test set;
[0179] By analyzing the growth trends of each plant species under different environmental parameters, obtain the standard growth trend curve of each plant species as the reference standard for the nearest neighbor algorithm;
[0180] Obtain the environmental parameter data of the plant sample to be judged. According to this environmental parameter data, use the nearest neighbor algorithm to calculate the distance between the growth trend curve of this plant sample and each standard growth trend curve;
[0181] Based on the calculated distance values, determine which standard growth trend curve the growth trend curve of the plant sample to be judged is closest to, and obtain the plant species to which this plant sample most likely belongs;
[0182] According to the results of the nearest neighbor algorithm, judge the specific plant species of the plant sample to be judged to obtain the classification result of this plant sample;
[0183] Adopt the method of cross-validation, divide the new sample data set into multiple subsets, each time select a different subset as the test set, and the remaining subsets as the training set. Through multiple iterative trainings and tests, obtain the average accuracy rate of the nearest neighbor algorithm on this data set to evaluate the performance of the algorithm;
[0184] According to the classification results and performance evaluation results of the nearest neighbor algorithm, determine the practical application value of this algorithm in the plant species judgment task, and provide a reference for subsequent plant classification and recognition work;
[0185] The example is as follows:
[0186] Extract the environmental parameter data and the corresponding plant species information from the newly collected sample data set. The environmental parameters include temperature, humidity, light intensity, etc., a total of 15 parameters;
[0187] Randomly divide the extracted data into a training set and a test set according to a ratio of 8:2. For the training set data, group them according to plant species, and the number of plant samples in each group is not less than 1000;
[0188] Using the regression analysis algorithm, with the environmental parameters as independent variables and the plant growth state parameters (such as plant height, number of leaves, etc.) as dependent variables, model the growth trends of each plant species to obtain the standard growth trend curve equation of each plant species;
[0189] For the plant sample to be judged, obtain its growth environment parameters, substitute them into the standard growth trend curve equation of each plant species, calculate the theoretical growth state parameters, and then use the Euclidean distance formula to calculate the distance between the actual growth state parameters of the plant sample and the theoretical growth state parameters of each plant species. The one with the smallest distance is the most likely species to which the plant sample to be judged belongs;
[0190] To evaluate the accuracy of the model, a 5-fold cross-validation method is adopted. The test set data is divided into 5 subsets. Each time, 1 of the subsets is selected as the validation set, and the remaining 4 are used as the training set for 5 iterative trainings and validations;
[0191] By calculating the average accuracy of the 5 judgment results, the overall accuracy of this nearest neighbor algorithm on this data set is 97%, indicating that this method has high practical value in the task of plant species judgment;
[0192] To further improve the judgment accuracy, a deep learning algorithm can be introduced. The convolutional neural network model is used to extract features and classify plant images, and combined with the nearest neighbor algorithm to form a complete plant intelligent recognition system.
[0193] Furthermore, in this embodiment, step S7 further includes:
[0194] Obtain the morphological feature data of the plant through computer vision technology, and use image segmentation, object detection and feature extraction algorithms to convert the plant image into a structured morphological feature vector; the morphological features of the plant include plant height, crown width, leaf shape and flower color;
[0195] Utilize the collected odor data of the plant, extract the types and concentration characteristics of the odor components to form an odor feature vector; the odor data released by the plant is collected by an odor sensor;
[0196] Input the odor feature sequence and the morphological feature sequence into two long short-term memory networks (LSTM) respectively. Through the transfer and update of the hidden state, the alignment of the feature sequence in the time dimension is realized;
[0197] In the spatial dimension, through interpolation and mapping methods, the feature data collected by different sensors are unified into the same coordinate system to construct a multi-modal plant feature data set;
[0198] Adopt methods such as Relief, Lasso, principal component analysis and linear discriminant analysis to evaluate the importance and correlation of odor features and morphological features from multiple perspectives, select the feature subset with the greatest contribution to plant species discrimination, and achieve feature dimensionality reduction;
[0199] Adopt a multi - core learning strategy. By defining different kernel functions, deal with odor features and morphological features respectively, and make a selection according to the distribution characteristics of the features. Commonly used kernel functions include linear kernel, Gaussian kernel, and polynomial kernel.
[0200] Perform weighted combination on the outputs of each kernel function to obtain a unified sample similarity measure.
[0201] The weights of the kernel functions are optimized through the least - square loss with L2 - norm regularization and the multi - core learning objective function of Hinge loss, and solved using quadratic programming or gradient descent algorithms. By alternately iterating the parameters of the kernel functions and the combined weights until convergence, the deep fusion of odor features and morphological features is achieved.
[0202] Input the fused plant features into a pre - trained plant species recognition model, which can be softmax regression, support vector machine (SVM), or convolutional neural network (CNN).
[0203] In the test stage, select the category with the highest output probability as the preliminary recognition result. According to the association rule base of morphology - odor - species, such as "IF flower color = red AND number of petals = 5 AND odor = sweet THEN species = rose".
[0204] Match the preliminary recognition result with the association rule base. If there are contradictions or inconsistencies, trigger the rule reasoning process. The reasoning process can adopt forward reasoning or backward reasoning. Through the combination of rules and conflict resolution, obtain the corrected recognition result. At the same time, it is also possible to use past recognition cases. Through similarity calculation and priority ranking, find the historical case most similar to the current sample, and use its species label as a reference to verify and correct the preliminary recognition result.
[0205] In the process of reasoning and decision - making, use Dempster - Shafer evidence theory to model the support or confidence of different evidence sources such as model output, rule reasoning, and case analysis, represented by the mass function. Through Dempster's combination rule, fuse multiple evidences to obtain a comprehensive final recognition result. Use Dempster - Shafer entropy to measure the conflict degree between evidences as a reference for confidence evaluation.
[0206] Associate the finally identified plant species labels with the original data such as images and odors, and store them in the sample library. For structured feature data, use a relational database such as MySQL for storage, and implement efficient conditional queries and statistical analyses through SQL statements; for unstructured original data, including image files, odor waveforms, etc., use a non-relational database such as MongoDB for storage, which supports diverse data formats and dynamic schema expansion. Define metadata fields for collection time, location, and equipment to manage and trace the samples;
[0207] Feed the recognition results back to users or other application systems, including plant encyclopedias, ecological monitoring, etc., to provide real-time plant species information services. Regularly extract a portion of data from the sample library to fine-tune and evaluate the plant species recognition model. Through gradient descent and backpropagation algorithms, make the model parameters adapt to new data distributions and feature changes, and use methods such as cross-validation and leave-one-out to evaluate performance metrics such as the accuracy, recall rate, and F1 value of the model. Take corresponding regularization and early stopping measures to prevent overfitting or underfitting;
[0208] Regularly perform K-means, DBSCAN clustering analyses, as well as principal component analysis, factor analysis, and independent component analysis statistical modeling on the sample library to mine the commonalities and differences in odor and morphological characteristics among different plant species, update the association rule library and case library of morphology-odor-species, continuously optimize and improve the recognition model, and improve the accuracy, robustness, and generalization ability of plant species recognition, providing important data support and decision-making basis for smart agriculture and ecological environment protection;
[0209] Examples are as follows:
[0210] Taking three kinds of flowers, namely roses, jasmines, and osmanthus, as examples, collect 1000 flower images at different growth stages, and use the YOLOv5 object detection algorithm to extract the position and size information of key parts such as flowers, leaves, and branches. Then, extract 2048-dimensional visual feature vectors through the ResNet-50 convolutional neural network; at the same time, use an electronic nose to collect odor data released by the three kinds of flowers at different growth stages, collect each sample 10 times, each time lasting 20 seconds, and extract 32-dimensional odor features such as the types and concentrations of odor molecules;
[0211] Then, input the visual features and odor features into two neural networks composed of 128 LSTM units respectively. Taking 7 days as the time step, train the network parameters through error backpropagation to achieve synchronous alignment of odor and visual features in the time dimension;
[0212] Next, the Relief-F algorithm is used to calculate the contribution of each feature to the class discrimination between neighboring samples, and the 50 most important visual features and 10 olfactory features are selected. Then, Lasso regression is used for feature screening to obtain 30 visual features and 5 olfactory features. Finally, the PCA algorithm is used to reduce the dimensionality of the fused features to 20 dimensions;
[0213] The multi-kernel learning method is adopted to fuse features of different modalities. The linear kernel is used to process visual features, and the Gaussian kernel is used to process olfactory features. The weights of the kernel functions are optimized through the least squares loss function to obtain the fused multi-modal feature representation. The fused features are input into the softmax classifier for class recognition, and the 5-fold cross-validation method is used to evaluate the recognition accuracy, and the predicted probability values of different classes are recorded at the same time;
[0214] According to 20 morphological-olfactory-class association rules in the morphological-olfactory-class association rule library, such as "bright flower color, many petals, strong smell → rose; elegant flower color, slender petals, fresh smell → jasmine; yellow-green flower color, round petals, rich smell → osmanthus", the Drools rule engine is used to reason and verify the recognition results. When the recognition result is inconsistent with the rule inference result, the top-5 samples with the most similar morphological and olfactory features are retrieved from the historical sample library, and the Dempster-Shafer evidence support degree of their class labels is calculated, and Dempster evidence combination is performed with the predicted probability of the softmax classifier to obtain the comprehensive recognition result. The final recognition result, together with the original image, odor signal, acquisition metadata, etc., is stored in the MongoDB database, and the sample library is updated in real time;
[0215] Every 30 days, 100 samples are drawn from the sample library to form a fine-tuning dataset, and the Adam optimizer and the cross-entropy loss function are used to fine-tune the recognition model, and performance indicators such as precision, recall, and F1-score of each class are statistically calculated;
[0216] Every 60 days, the K-means++ algorithm is used to perform clustering analysis on the visual and olfactory features in the sample library. According to the silhouette coefficient, the optimal number of clusters is determined to be 5. It is found that different classes show obvious intra-class aggregation and inter-class separation trends in feature combinations such as "flower color-flower shape" and "odor-intensity". Based on this, the morphological-olfactory-class association rule library is optimized to improve the coverage and accuracy of rule inference, providing more comprehensive and refined knowledge support for flower recognition.
[0217] The present invention also proposes a seedling class recognition system, which adopts the above-mentioned seedling class recognition method, including:
[0218] The acquisition module, which includes an odor sensor and an environmental monitoring device, is used to collect plant odor samples and environmental parameter data;
[0219] The data analysis module is used to collect plant odor samples and environmental parameter data;
[0220] The feature analysis module is used to identify key compounds and classify them to form odor features;
[0221] The model processing module is used to establish and optimize a support vector machine model for growth stage prediction;
[0222] The dynamic analysis module is used to predict the growth stage and analyze the plant growth trend;
[0223] The fusion module is used to identify plant species and adjust the output in combination with visual data.
[0224] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by orientation words such as "front, back, up, down, left, right", "horizontal, vertical, level" and "top, bottom" is usually based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description. Without contrary description, these orientation words do not indicate and imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, it should not be construed as limiting the protection scope of the present invention.
[0225] For those skilled in the art, various corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all these changes and deformations should fall within the protection scope of the claims of the present invention.
Claims
1. A method for identifying the types of young seedlings, characterized in that, Including: S1. Collect odor samples and environmental parameter data corresponding to different growth stages of young seedlings of different plant species, and integrate data streams through multi-point synchronous collection to form a sample data set with timestamps. S2. Separate the collected odor samples, analyze the components in each separated sample to generate corresponding data streams, and use isotope labeling method for qualitative identification. S3. Extract the characteristic features of key compounds at each growth stage of the plant, and use the K-means clustering algorithm to classify the identified compounds to form odor characteristics corresponding to each growth stage. S4. Combine the growth stage corresponding to the odor characteristics and the environmental parameter data, establish a first dynamic prediction model and train it, and optimize the model parameters through the Bayesian method. S5. Analyze the trends of odor characteristics, growth stages, and environmental parameter data corresponding to timestamps, and establish a second dynamic prediction model. S6. Combine the odor characteristics corresponding to each growth stage of each plant species and the environmental parameter data corresponding to timestamps, analyze the standard growth trends of each plant species under different environmental parameter data, establish a plant species identification model and train it, and identify plant species for a new sample data set without plant species association. S7. Analyze the morphological characteristics of plants by combining computer vision data, fuse odor characteristics with computer vision data, and adjust the output results of the plant species identification model.
2. The method for identifying the variety of young seedlings according to claim 1, wherein The step S1 further includes: The collection content of the odor samples and the environmental parameter data includes temperature, humidity, light intensity, and wind speed. The formation of the sample data set includes the following steps: Determine the time points for collecting the odor samples and the environmental parameter data according to the growth stages of young seedlings of different plant species, and integrate the odor samples and the environmental parameter data through multi-point synchronous collection to obtain a data stream with timestamps. Adopt data preprocessing technology to clean, denoise, and standardize the collected odor samples and environmental parameter data, arrange the data in chronological order using timestamp information, construct a time series data set, and calculate the sampling interval according to the timestamp for data interpolation operation to obtain the processed sample data set.
3. The method for identifying the types of young crops according to claim 1, wherein The step S2 further includes: According to the characteristics of the plant species, collect the odor samples released by the plant, separate and analyze the components of the odor samples by high performance liquid chromatography-mass spectrometry (HPLC-MS) technology to obtain corresponding data streams. Separate the compounds in the odor samples using a high performance liquid chromatography column, optimize the conditions of mobile phase composition, pH value, and column temperature through orthogonal experimental design, and use the solvent gradient elution method to achieve separation according to the differences in polarity and hydrophobicity of the compounds. Import the separated samples into a mass spectrometer, ionize the compounds by electron impact ionization source, measure the mass-to-charge ratio in a quadrupole mass analyzer to obtain the mass spectrum of the compounds, and use isotope labeling method to qualitatively identify the compounds based on the fragment ion peaks and molecular ion peaks in the mass spectrum and the standard mass spectrum in the compound database.
4. The method for identifying the type of young seedlings according to claim 1, wherein The step S3 further includes: Obtain the mass spectrometry identification data of the odor samples at different growth stages of plants through high performance liquid chromatography-mass spectrometry (HPLC-MS) technology. Use the principal component analysis method to reduce the dimension of the mass spectrometry data. Through orthogonal transformation, map the original high-dimensional data to a set of linearly independent new feature spaces, extract the main variation information of the samples. The new features are sorted according to the variance size, and the first few principal components with the largest variance can explain most of the variation of the original data, forming a low-dimensional feature matrix; Adopt the linear discriminant analysis method to further optimize the principal component features. Project the data into a low-dimensional space through linear transformation, so that the samples at different growth stages are separated as much as possible after projection, while the samples within the same growth stage are clustered as much as possible, maximizing the between-class difference between samples at different growth stages, and at the same time minimizing the within-class difference of samples within the same growth stage, to obtain the feature combination with the strongest discriminative power; Combine the recursive feature elimination method, and iteratively remove the features that contribute less to the discrimination result, screen out the key compound marker features highly related to the plant growth stage, and construct an optimal feature subset; According to the screened key compound marker features, use the K-means clustering algorithm to perform unsupervised classification on the odor samples at different growth stages. First, randomly select K samples as the initial clustering centers, then calculate the distances from each sample to each clustering center, divide the samples into the closest categories, and then recalculate the centroids of each category as the new clustering centers. Repeat the above process until the clustering centers no longer change or reach the maximum number of iterations. By optimizing the clustering centers and sample assignments, the samples with similar odor features are divided into the same category, forming the odor feature clustering results representing each growth stage of the plant.
5. The method for identifying the types of young seedlings according to claim 1, wherein, The step S4 further includes: Optimize the model parameters including the penalty coefficient and kernel function parameters, so that the support vector machine model associates the sample data set with the corresponding growth stage, and classifies the new sample data set without growth stage association according to the growth stage; According to the corresponding relationship between the odor features and the growth stages obtained in the step S3, combine the environmental parameter data to construct a multi-dimensional sample data set as the training input of the first dynamic prediction model; Adopt the support vector machine model as the basis of the first dynamic prediction model, map the sample data to a high-dimensional space through the kernel function, and find the optimal classification hyperplane in the high-dimensional space to realize the non-linear association between the odor features and the growth stages; Use the Bayesian optimization method to automatically optimize the key parameters of the support vector machine model.
6. The method for identifying the type of young seedlings according to claim 1, characterized in that The step S5 further includes: The second dynamic prediction model includes an autoregressive integrated moving average model or a long short-term memory neural network to predict the growth stage corresponding to the future time point; Preprocess the collected odor feature data, remove outliers and noise data, extract the key parameters of the odor features, and obtain a normalized odor feature data set; Clean and standardize the environmental parameter data, remove invalid data, and normalize the data to make the environmental parameter data of different dimensions comparable, forming a standardized environmental parameter data set; Align the preprocessed odor feature data and the standardized environmental parameter data according to the time stamp to construct a multi-dimensional time series data set of odor-environment-growth stage, which is used as the input of the trend analysis and prediction model; Use the autoregressive integrated moving average model to perform trend analysis on the multi-dimensional time series data, and establish a second dynamic prediction model between the odor characteristics, environmental parameters and growth stage through the autocorrelation and trend changes of historical data.
7. The method for identifying the type of young seedlings according to claim 1, wherein The step S6 further includes: According to the odor sample data of each plant species at different growth stages, extract the odor feature vectors of each sample, including the types and contents of key compounds, and associate and store the odor feature vectors with the corresponding plant species, growth stage and collection time stamp to form a structured odor feature data set; Use data preprocessing technology to clean, denoise and standardize the odor feature data, eliminate the influence of outliers and missing values, and align and synchronize the odor feature data with the environmental parameter data according to the time stamp to form a complete time series data set, and each time point contains multi-dimensional information of odor characteristics, environmental parameters and plant species; Use the dynamic time warping algorithm DTW to calculate the distance matrix between the odor feature sequences of different plant species, and then use the distance matrix as the input, apply the K-means clustering algorithm to perform clustering analysis on the odor feature time series, evaluate the clustering quality of different clustering numbers K through the silhouette coefficient index, select the optimal clustering result, and obtain the standard time series template representing the growth patterns of different plant species; For the standard time series template in each cluster, use the method of multiple linear regression, with environmental parameters as independent variables and odor characteristics as dependent variables, and establish the following regression equation: y = β0 + β1X1 + β2X2 + … + β P X P Among them, y represents the odor characteristics and is the dependent variable; β i (i = 0, 1…P) are the regression coefficients, which are used to reflect the influence degree and direction of each environmental parameter on the odor characteristics; X i (i = 1, 2…P) represents the value of the i-th environmental parameter and is the independent variable; Solve the regression coefficients by the least squares method, and perform significance tests and model diagnostics to ensure the effectiveness and interpretability of the regression model, explore the growth laws and trend changes of different plant species under different environmental conditions, construct a plant species recognition model, with odor characteristics and environmental parameters as inputs and plant species as outputs. First, perform feature selection and dimensionality reduction on the two types of features respectively, extract the key feature subsets that contribute the most to species discrimination, and then use machine learning algorithms, and realize the non-linear transformation and fusion of features through kernel function mapping or decision tree combination. For multi-classification problems, adopt the "one-vs-rest" or "one-vs-one" strategy to transform multi-classification into multiple binary classification problems for solution, and then obtain the final species label through voting or probability combination. Finally, optimize the hyperparameters of the model through grid search and cross-validation, and train to obtain the plant species discrimination model with the best performance; When a newly collected odor sample data set is input, first extract its odor feature vector and environmental parameters, and standardize them in the same preprocessing manner as the training data; Input the preprocessed feature vector into the trained plant species recognition model, calculate the similarity scores between the sample and each plant species through the decision function of the model, and select the plant species with the highest score as the recognition result.
8. The method for identifying the type of young crops according to claim 7, wherein The step S6 further includes: Combined with the standard growth trends of various plant species corresponding to the new environmental parameter data in the new sample dataset, the nearest neighbor algorithm is used to match the corresponding plant species.
9. The method for identifying the type of young seedlings according to claim 1, wherein The step S7 further includes: Obtain the morphological feature data of plants through computer vision technology, and use image segmentation, object detection, and feature extraction algorithms to convert plant images into structured morphological feature vectors; Utilize the collected odor data released by plants, extract the types and concentration characteristics of odor components, and form odor feature vectors; Input the odor feature sequence and the morphological feature sequence into two long short-term memory networks (LSTMs) respectively. Through the transmission and update of hidden states, the alignment of the feature sequences in the time dimension is achieved; In the spatial dimension, through interpolation and mapping methods, the feature data collected by different sensors are unified into the same coordinate system to construct a multi-modal plant feature dataset; Adopt methods such as Relief, Lasso, principal component analysis, and linear discriminant analysis to evaluate the importance and correlation of odor features and morphological features from multiple perspectives, select the feature subset with the greatest contribution to plant species discrimination, and achieve feature dimensionality reduction; Adopt a multi-kernel learning strategy. By defining different kernel functions, process odor features and morphological features respectively, and make selections according to the distribution characteristics of the features; Perform weighted combination of the outputs of each kernel function to obtain a unified sample similarity metric; The weights of the kernel functions are optimized through the least squares loss with L2 regularization, the Hinge loss multi-kernel learning objective function, and solved using quadratic programming or gradient descent algorithms. By alternately iterating the parameters of the kernel functions and the combined weights until convergence, the deep fusion of odor features and morphological features is achieved.
10. A green seedling variety identification system, which adopts the green seedling variety identification method described in any one of claims 1-9, is characterized in that, It includes: A collection module, where the collection module includes an odor sensor and an environmental monitoring device, and is used to collect plant odor samples and environmental parameter data; A data analysis module, where the data analysis module is used to collect plant odor samples and environmental parameter data; A feature analysis module, where the feature analysis module is used to identify key compounds and classify them to form odor features; A model processing module, where the model processing module is used to establish and optimize a support vector machine model for growth stage prediction; A dynamic analysis module, where the dynamic analysis module is used to predict the growth stage and analyze the plant growth trend; A fusion module, where the fusion module is used to perform plant species identification and adjust the output in combination with visual data.
Citation Information
Patent Citations
Bionic olfaction smell identification method
CN109932515A
Method for detecting plant disease and insect pest infection by using electronic nose
CN117129632A
Cited By
Arbor load estimation method and system based on different growth coefficients
CN120725262A