Early identification method of male and female Litsea cubeba plants based on Raman spectroscopy
The sex identification of Litsea cubeba leaves is carried out by combining Raman spectroscopy with surface treatment and classification models, which solves the time-consuming and costly problems of existing technologies, achieves rapid and accurate identification of male and female plants, and provides a simple solution for large-scale planting and breeding.
Patent Information
- Application Number
- CN202410114531.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-01-26
AI Technical Summary
Existing technologies for sex identification of male and female Litsea cubeba plants are time-consuming and costly, making it difficult to achieve rapid and accurate on-site testing.
Raman spectroscopy was used in combination with various surface treatments and classification models to identify Litsea cubeba leaves, including cleaning and drying, surface extrusion, and silver nanocolloid enhancement treatment. Raman spectroscopy was then collected and data processed, and a classification model was constructed for gender identification.
The rapid and low-cost sex identification of male and female Litsea cubeba plants was achieved, providing a simple method for large-scale plant cultivation and breeding and improving the identification accuracy.
Smart Images

Figure CN117907306B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the fields of agricultural biology and machine learning technology, and particularly relates to an early identification method for male and female Litsea cubeba plants based on Raman spectroscopy. Background Art
[0002] Litsea cubeba, also known as Litsea cubeba and Litsea cubeba, is a deciduous shrub or small tree of the genus Litsea in the Lauraceae family. It is an important economic oil crop in my country. Its unique aroma has led to its widespread use in the cosmetics industry worldwide. Furthermore, Litsea cubeba essential oil, extracted from its fruit, is rich in monoterpenes, sesquiterpenes, and alkaloids. A growing body of research indicates that it possesses antibacterial, anti-inflammatory, antioxidant, anticancer, antiviral, sedative, and therapeutic properties for rheumatoid arthritis, finding applications in a variety of fields, including food, medicine, and agriculture. Consequently, demand for Litsea cubeba cultivation is growing.
[0003] As a dioecious plant, the distribution ratio of male and female Litsea cubeba plants significantly impacts its growth. Male plants account for approximately 40% of males in a plantation of male and female seedlings, but in practice, a 10:1 ratio is sufficient within a plantation, with only a small number of males retained for pollination. However, the sex of Litsea cubeba plants can generally only be determined by flower morphology after flowering in the third year, and mature male plants must be removed. This undoubtedly increases labor costs. Therefore, early identification of Litsea cubeba sex is urgently needed in plantation management.
[0004] Currently, sex determination is based on the morphological characteristics of flowers of male and female plants during the reproductive phase, which can be time-delayed. Currently developed methods for sex differentiation in Litsea cubeba focus primarily on breeding. These include investigating sex-specific patterns in hormone levels (such as indoleacetic acid, jasmonic acid, salicylic acid, gibberellins, and trans-zeatin riboside) during flower degeneration, as well as transcriptome sequencing and transcription factor expression studies in male and female flower buds. A Chinese invention patent with the publication number "CN106987652B" discloses a SNP marker for identifying the sex of Litsea cubeba and a screening method for the SNP marker. The screening method comprises: extracting genomic DNA from the Litsea cubeba to be tested, constructing a RAD library, and sequencing to obtain raw data; optimizing and screening the raw data to obtain RAD sequencing data; counting the RAD sequencing data to obtain the number and depth of each individual RAD-tag; performing a primary classification based on the depth of the RAD-tag, and performing a secondary classification based on the number of differential bases after the primary classification, genotyping the SNP marker, and determining the SNP site according to the maximum likelihood method. However, this method is time-consuming and requires a specialized laboratory and dedicated personnel for testing. Therefore, a rapid, accurate, on-site method for identifying the sex of male and female plants is needed for agricultural applications.
[0005] In recent years, spectroscopic methods such as near-infrared spectroscopy and Raman spectroscopy have begun to be used to identify the sex of male and female plants. However, near-infrared spectroscopy is limited by water absorption, while Raman spectroscopy is insensitive to the aqueous phase and is therefore suitable for the detection of aqueous samples. Raman spectroscopy (RS) and surface-enhanced Raman spectroscopy (SERS) are non-invasive and non-destructive spectroscopic methods with the advantages of high sensitivity, high accuracy, and short testing time, which can achieve rapid detection of trace components on site. Some scholars have used portable Raman equipment to specifically identify the sex of dioecious kiwifruit (H. Jiang, et al., 2023), but the substrate preparation process used is complex and the price of screening primer DNA is relatively expensive, making it unsuitable for large-scale production. Summary of the Invention
[0006] The present invention provides a method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy, aiming to solve the problems of time-consuming and high cost in the prior art for sex identification of male and female Litsea cubeba plants.
[0007] To solve the above technical problems, the identification method proposed by the present invention comprises the following steps:
[0008] S1: Collect leaves from male and female Litsea cubeba plants;
[0009] S2: pre-treating the male and female leaves to obtain pre-treated leaves, wherein the pre-treatment is one of the following four methods: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nano-colloid after cleaning and drying, and surface enhancement using silver nano-colloid after cleaning, drying and surface extrusion;
[0010] S3: collecting the Raman spectrum of the pre-processed leaf, and performing shift shearing, smoothing filtering and background noise reduction processing on the Raman spectrum;
[0011] S4: Perform principal component analysis and linear discriminant analysis on the Raman spectrum data processed in step S3, reduce the dimensionality and classify the input data, and divide it into a training set and a test set according to a set ratio;
[0012] S5: Build a classification model, train the classification model using the training set, test the training results using the test set, and select the best model based on accuracy, precision, recall, and F1 score;
[0013] S6: Collect leaves of Litsea cubeba to be tested, and after processing in steps S2-S4, input them into the optimal model for classification, and the optimal model outputs the classification results, which include female plants and male plants.
[0014] Preferably, the preparation method of the silver nanocolloid is:
[0015] Weigh a certain amount of AgNO3 solid into a container and prepare a concentration of 1×10 -3 mol﹒ L -1 AgNO3 solution;
[0016] Stirring the AgNO3 solution at a set speed until it is heated to boiling;
[0017] After the AgNO3 solution is boiled, a certain amount of 1% sodium citrate solution is added and the solution is boiled for 30 minutes.
[0018] The solution was cooled at room temperature, and the supernatant was discarded after centrifugation to obtain a silver sol enhancer.
[0019] Preferably, the surface pressing is specifically: using a tool with a round blunt head to press the surface of the collected Litsea cubeba leaf to be tested until the surface of the leaf to be tested is damaged;
[0020] The surface enhancement using silver nano-colloid is specifically as follows: a silver sol enhancer is dripped onto the surface of the Litsea cubeba leaf where the test is to be performed, and the silver sol is naturally air-dried.
[0021] Preferably, in step S3, a portable Raman spectrometer is used to collect Raman spectra with an excitation wavelength of 785 nm.
[0022] Preferably, when collecting Raman spectra in step S3, the laser power is set to 15 mW and the collection time is set to 3 s.
[0023] Preferably, the classification model is one of decision tree, random forest, extreme gradient boosting, logistic regression, support vector machine, K nearest neighbor, and naive Bayes.
[0024] Preferably, the number of trees in the extreme gradient boosting is set to 100, 100, 100, and 100, respectively, and the maximum depth of the tree is set to 4, 0, 0, and 0, respectively; corresponding to four leaf pretreatment methods: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nanocolloid after cleaning and drying, and surface enhancement using silver nanocolloid after cleaning and drying and surface extrusion.
[0025] Preferably, the regularization parameters of the logistic regression are set to 0.1, 0.001, 1, and 0.001, respectively, and the regularization type parameters are set to 12, 12, 12, and 12, respectively; corresponding to four leaf pretreatment methods: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nanocolloids after cleaning and drying, and surface enhancement using silver nanocolloids after cleaning and drying and surface extrusion.
[0026] Preferably, the regularization parameters of the support vector machine are set to 0.1, 10, 0.1, and 0.1, respectively; the kernel function types are set to linear, rbf, linear, and linear, respectively; and the coefficients of the kernel function are set to scale, auto, scale, and scale; corresponding to four leaf pretreatment methods: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nanocolloids after cleaning and drying, and surface enhancement using silver nanocolloids after cleaning and drying and surface extrusion.
[0027] Compared with the prior art, the present invention has the following technical effects:
[0028] 1. The identification method proposed in the present invention combines multiple surface treatments and multiple classification models to identify the sex characteristics of Litsea cubeba. This method can eliminate complex pretreatment and achieve rapid and low-cost on-site detection and identification, providing a quick and simple method for adjusting the distribution of Litsea cubeba plants in nurseries. It can also be promoted and applied as a tool for early sex prediction of most plants and can be applied to large-scale plant planting and breeding programs.
[0029] 2. The identification method proposed in this invention uses sodium citrate as a reducing agent and stabilizer to prepare silver nanocolloids. When sodium citrate is used as a reducing agent, the citrate radicals strongly coordinate with the surface of the silver nanoparticles. Under appropriate pH conditions, the citrate radicals bound to the silver nanoparticles can ionize, imparting a negative charge to the surface. This can, to a certain extent, prevent the silver nanoparticles from agglomerating and ensure their good dispersibility in the aqueous phase. This enhances the Raman enhancement effect of the silver nanoparticles and improves the accuracy of surface-enhanced sample identification.
[0030] 3. The proposed identification method uses a laser power of 15 mW and an acquisition time of 3 seconds to obtain Raman spectra, ensuring that the leaves are in good condition and have a good signal-to-noise ratio. As the laser power increases, the signal intensity of the Raman spectra gradually increases, but the change is not significant. Although Raman spectroscopy performs best at 25 mW, leaves are more susceptible to dehydration or burn-through due to higher laser power, resulting in poor leaf condition. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flow chart of the identification method of the present invention;
[0032] Figure 2 is a 3D waterfall diagram of the average Raman spectra of male and female Litsea cubeba leaf samples according to an embodiment of the present invention;
[0033] Figure 3 are the average Raman spectra and differential spectra of only clean and dry leaf samples according to an embodiment of the present invention;
[0034] Figure 4 The average Raman spectrum and differential spectrum of the leaf sample after surface extrusion after cleaning and drying according to the embodiment of the present invention are shown;
[0035] Figure 5 The average Raman spectrum and differential spectrum of the leaf sample after cleaning and drying and surface enhancement with silver nanocolloids according to the embodiment of the present invention are shown;
[0036] Figure 6 These are the average Raman spectra and differential spectra of the leaf sample that was cleaned, dried, and surface-extruded and then surface-enhanced with silver nanocolloids according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in combination with specific embodiments of the present application and with reference to the accompanying drawings.
[0038] Example 1
[0039] like Figure 1 As shown, the early identification method of male and female Litsea cubeba plants based on Raman spectroscopy includes the following steps:
[0040] S1: Collect leaves from male and female Litsea cubeba plants.
[0041] In this example, for the flowering Litsea cubeba plants, the sex of the plants was first identified on-site by flower morphological characteristics and marked with a label. Three months after the new leaves grew, five male and five female plants were selected, and fresh, healthy, and complete leaves were collected from all parts of the plant. They were placed in plastic ziplock bags and numbered. The male and female plants of Litsea cubeba are defined by the flowering of the plants. The sex type of Litsea cubeba was identified according to the morphological description reported by Zilong Xu et al. (Xu, et al., 2020). These plants were identified as female and male.
[0042] S2: Pre-treating the leaves of the male and female plants to obtain pre-treated leaves, wherein the pre-treatment is one of the following four methods: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nano-colloid after cleaning and drying, and surface enhancement using silver nano-colloid after cleaning, drying and surface extrusion.
[0043] As for the pretreatment method, the surface extrusion is specifically: using a tool with a blunt head to press the surface of the collected Litsea cubeba leaves to be tested until the surface of the leaves to be tested is damaged; the surface enhancement using silver nano-colloid is specifically: adding a silver sol enhancer to the surface of the Litsea cubeba leaves to be tested, and naturally air-drying the silver sol.
[0044] Specifically, the preparation method of the silver nanocolloid is:
[0045] Weigh a certain amount of AgNO3 solid into a container and prepare a concentration of 1×10 -3 mol﹒ L -1 In this embodiment, weigh 0.0255g of AgNO3 solid in a 250mL conical flask to prepare 150mL of AgNO3 solution, the concentration of which is 1×10 -3 mol﹒ L -1 .
[0046] The AgNO3 solution is stirred at a set speed until it is heated to boil; specifically, a magnet is added and placed on a magnetic stirrer and stirred at a speed of 500 r / min until it is heated to boil.
[0047] After the boiling AgNO3 solution, a certain amount of 1% sodium citrate solution was added and the solution was boiled for 30 minutes. For this embodiment, 0.0255 g of AgNO3 solid was used, 4 mL of 1% sodium citrate solution was added and the solution was boiled for 30 minutes. During this process, the solution changed from transparent to yellow-brown and finally to gray-green.
[0048] The solution was cooled to room temperature and stored in a refrigerator for later use. The stored solution was centrifuged and the supernatant was discarded to obtain a silver sol enhancer.
[0049] When sodium citrate is used as a reducing agent, the citrate radicals strongly coordinate with the surface of the silver nanoparticles. Under appropriate pH conditions, the citrate radicals bonded to the surface of the silver nanoparticles can be ionized, giving the surface a negative charge. This can, to a certain extent, prevent the silver nanoparticles from agglomerating and ensure their good dispersion in the aqueous phase. This improves the Raman enhancement effect of the silver nanoparticles and increases the accuracy of surface-enhanced sample identification.
[0050] In step S2, this embodiment processed 1040 samples, including 520 samples of female Litsea cubeba leafs and 520 samples of male Litsea cubeba leafs. Among them, the female Litsea cubeba leaf samples were divided into 4 types, 130 each: 130 leaf samples that were only cleaned and dried, 130 leaf samples that were surface extruded after cleaning and drying, 130 leaf samples that were surface enhanced with silver nano-colloids after cleaning and drying, and 130 leaf samples that were surface enhanced with silver nano-colloids after cleaning, drying, and surface extrusion. Similarly, the male Litsea cubeba leaf samples were also divided into 4 types, 130 each: 130 leaf samples that were only cleaned and dried, 130 leaf samples that were surface extruded after cleaning and drying, 130 leaf samples that were surface enhanced with silver nano-colloids after cleaning, drying, and surface extrusion. The specific distribution of the samples is shown in Table 1 below:
[0051]
[0052]
[0053] Table 1 Sample quantity distribution
[0054] S3: collecting the Raman spectrum of the pre-processed leaf, and performing shift shearing, smoothing filtering and background noise reduction processing on the Raman spectrum.
[0055] Before collecting Raman spectra, use double-sided tape to fix the leaves to the bottom plate of the Raman spectrometer bracket to prevent the leaves from bending and affecting the distance between the laser and the leaves, thereby affecting the measurement results. Spectra are collected on leaves close to the main veins. For on-site testing of sample sex at the Litsea cubeba planting site, a portable Raman spectrometer can be used for Raman spectrum collection, with an excitation wavelength of 785nm, a fixed number of scans of 5, and a scanning range of 400-1800cm. -1 .
[0056] When the Raman spectrum is collected in step S3, the laser power is set to 15mW and the collection time is set to 3s. The optional laser power of this embodiment is 5mW-25mW. Under the same scanning time, as the laser power increases, the signal intensity of the Raman spectrum also gradually increases, but the amplitude of the change is not large. Although the Raman spectrum performs best at 25mW, since the blades are easily dehydrated or burned by higher-power laser irradiation, a laser power of 15mW is preferred to obtain a good state of the blades and a good signal-to-noise ratio of the spectrum. Similarly, the optional collection time of this embodiment is 1s-5s. Under the same laser power, the clarity of the signals collected at different collection times is different. Therefore, a 3s collection time with clear Raman signals and a shorter collection time is preferred. In short, the optimized Raman spectrum is set to a laser power of 15mW and a collection time of 3s.
[0057] The collected Raman spectra were pre-processed using the portable Raman spectrometer's built-in data processing model, including displacement shearing, smoothing and filtering, and background noise reduction. Specifically, displacement shearing is performed because the Raman signal at 785nm is affected by the filter effect, resulting in a significant background signal, and the Raman signal is prone to exceeding the measurement range. Subsequently, the Whittaker smoothing algorithm and sparse matrix technology were used to achieve optimal spectral smoothing results. There are two objectives that need to be balanced in the smoothing process, including the degree of distortion and roughness of the fitted data. The distortion can be effectively expressed as the sum of the squares of the difference between the original data and the fitted data:
[0058]
[0059] Where F is the distortion, m is the number of samples, x is the original data, and z is the smoothed data.
[0060] The roughness can be expressed as the sum of squared differences of the smoothed data z:
[0061]
[0062] Where R is the roughness.
[0063] Finally, the adaptive iterative reweighted penalized least squares algorithm airPLS is used to perform background subtraction to achieve denoising. Each step of the adaptive reweighting process can be regarded as solving a weighted penalized least squares problem:
[0064]
[0065] Where w is the weight vector, which can be obtained adaptively during the iteration process. Before the algorithm starts, the initial weight vector is given as w0=1. After initialization, w can be obtained in each step according to the following formula:
[0066]
[0067] Where t represents the number of iterations, and vector d t It is the difference between vector x and the last fitted background z during the t iterations t-1 The iterative process has a termination criterion. In the airPLS algorithm, the adaptive reweighted iterative process will terminate if the maximum number of iterations is reached or the termination criterion is reached. The termination criterion is defined as follows:
[0068] |d t |<0.001×|x|
[0069] The processed data sets are divided into four different data sets: a Raman data set of only cleaning and drying, a surface-enhanced Raman data set using silver nanocolloids after cleaning and drying, a Raman data set of surface extrusion after cleaning and drying, and a surface-enhanced Raman data set using silver nanocolloids after cleaning, drying and surface extrusion.
[0070] S4: Perform principal component analysis and linear discriminant analysis on the Raman spectrum data processed in step S3, reduce the dimension and classify the input data, and divide it into a training set and a test set according to a set ratio.
[0071] After data preprocessing, principal component analysis and linear discriminant analysis (PCA-LDA) are combined to reduce the data's dimensionality and classify it. PCA converts the original high-dimensional data into a low-dimensional feature space through linear transformation while retaining the maximum variance. The key point is that it can extract features, which helps reduce dimensionality and remove redundant information. LDA is a supervised dimensionality reduction method that projects data into a low-dimensional feature space by maximizing inter-class distances and minimizing intra-class distances. LDA can effectively improve the performance of pattern recognition and classification tasks and helps distinguish feature differences between different categories.
[0072] The 1040 spectral data collected from the four datasets were randomly divided into two parts, 80% for training and 20% for testing.
[0073] S5: Build a classification model, train the classification model using the training set, test the training results using the test set, and select the best model based on accuracy, precision, recall rate, and F1 score.
[0074] The classification model used in this embodiment is the DT (Decision Tree) model. The decision tree is a classification algorithm based on a tree structure that divides data into different categories through a series of decision rules. It makes judgments and predictions based on the conditions of different feature values by gradually segmenting the features. The decision tree can perform well-explained classification, which is convenient for understanding and explaining the decision-making process of the model. In order to avoid overfitting, the hyperparameters of the DT classifier need to be optimized. This embodiment searches and optimizes the maximum depth (max_depth) of the decision tree, the minimum number of samples required for node splitting (min_samples_leaf), and the minimum number of samples required for leaf nodes (min_samples_split).
[0075] The specific optimization parameters are as follows: for samples that are only cleaned and dried for pretreatment, the maximum depth is set to 0, the minimum number of samples required for node splitting is set to 10, and the minimum number of samples required for leaf nodes is set to 4; for samples that are surface extruded after cleaning and drying, the maximum depth is set to 0, the minimum number of samples required for node splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 1; for samples that are surface enhanced with silver nanocolloids after cleaning and drying, the maximum depth is set to 0, the minimum number of samples required for node splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 4; for samples that are surface enhanced with silver nanocolloids after cleaning, drying and surface extrusion, the maximum depth is set to 0, the minimum number of samples required for node splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 4.
[0076] This embodiment uses four indicators, namely, precision, accuracy, recall, and F1-index, to evaluate the effectiveness of the model. Accuracy is the proportion of all correctly predicted samples to the total number of samples, representing the overall accuracy of the prediction; precision is also called precision, that is, the proportion of correctly predicted positive samples to all predicted positive samples, representing the accuracy of the prediction in the positive sample results; recall is the proportion of correctly predicted positive samples to all actual positive samples. Although a high recall rate means that there may be more false detections, it will try its best to find every object that should be found; F1 score (F1-score) weighs precision (Precision) and recall (Recall). Generally speaking, accuracy and recall are negatively correlated. The introduction of F1-Score as a comprehensive indicator is to balance the effects of accuracy and recall, and to evaluate a classifier more comprehensively. F1 is the harmonic average of precision and recall. The larger the F1-score, the higher the quality of the model. The quality of the classification model needs to be evaluated based on the above indicators. The ability of the model is measured specifically by the following formula:
[0077]
[0078]
[0079]
[0080]
[0081] Where TP represents the number of true positives (TP), TN represents true negatives (TN), FP represents the number of false positives (FP), and FN represents false negatives (FN). An F1 score above 0.75 is considered acceptable, and a score above 0.8 is satisfactory.
[0082] A confusion matrix is used to record the true classification of a dataset and the classification model's predicted classification. The numbers along the main diagonal represent the correct decisions. A receiver operating characteristic (ROC) curve is used to evaluate the classifier's performance at different classification thresholds. Area under the curve (AUC) generally ranges from 0.5 to 1; the closer the AUC is to 1, the better the model's classification performance.
[0083] S6: Collect leaves of Litsea cubeba to be tested, and after processing in steps S2-S4, input them into the optimal model for classification, and the optimal model outputs the classification results, which include female plants and male plants.
[0084] In step S6, if it is necessary to detect the sex of the sample at the Litsea cubeba planting site, a portable Raman spectrometer can be used to collect Raman spectra.
[0085] Example 2
[0086] Steps S1-S4 and S6 of this embodiment are the same as those of Example 1, with the difference being step S5. The classification model adopted in step S5 of this embodiment is the RF (RandomForest) model. The RF classifier is a very representative Bagging ensemble algorithm, which is based on decision trees, introduces random feature selection, and obtains reliable classification results from predictions in a set of decision trees. Random forests will not overfit, but in the current algorithm, their accuracy is not outstanding, and hyperparameters need to be optimized. This embodiment optimizes the parameters in the random forest model: the number of trees (n_estimators), the maximum depth of the tree (max_depth), the minimum number of samples required for node splitting (min_samples_split), and the minimum number of samples required for leaf nodes (min_samples_leaf).
[0087] The specific optimization parameters are as follows: for samples that are only cleaned and dried for pretreatment, the number of trees is set to 100, the maximum depth of the tree is set to 0, the minimum number of samples required for point splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 2; for samples that are surface extruded after cleaning and drying, the number of trees is set to 100, the maximum depth of the tree is set to 0, the minimum number of samples required for point splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 1; for samples that are surface enhanced with silver nanocolloids after cleaning and drying, the number of trees is set to 100, the maximum depth of the tree is set to 0, the minimum number of samples required for point splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 2; for samples that are surface enhanced with silver nanocolloids after cleaning, drying and surface extrusion, the number of trees is set to 300, the maximum depth of the tree is set to 20, the minimum number of samples required for point splitting is set to 5, and the minimum number of samples required for leaf nodes is set to 2.
[0088] The model evaluation method of this embodiment is the same as that of the first embodiment.
[0089] Example 3
[0090] Steps S1-S4 and S6 of this embodiment are the same as those of the first embodiment, except for step S5. The classification model used in step S5 of this embodiment is the XGBoost (Extreme Gradient Boost) model. XGBoost is an integrated machine learning algorithm based on the gradient boosting decision tree (GBDT). XGBoost performs many optimizations on GBDT, for example, performing a second-order Taylor expansion on the loss function to improve the accuracy of the calculation, and using regularization terms to simplify the model to avoid overfitting. The final objective function of XGBoost can be expressed as follows:
[0091]
[0092] in, It is an expression in linear space, w is the weight vector of the leaf node, T means that the tree has T leaf nodes, G j and H j are all constants, G j Represents the cumulative sum of the first-order partial derivatives of the samples contained in leaf node j, H j Represents the cumulative sum of the second-order partial derivatives of the samples contained in leaf node j.
[0093] This embodiment performs hyperparameter optimization on the number of trees in the XGBoost model (n_estimators), the maximum learning depth of the tree (max_depth), the learning rate (learning_rate), the subsampling ratio (subsample), the minimum number of samples required for leaf nodes (min_samples_leaf), and the minimum number of samples required for node splitting (min_samples_split).
[0094] The specific optimization parameters are as follows: for samples that are only cleaned and dried for pretreatment, the number of trees is set to 100, the maximum learning depth of the tree is set to 4, the learning rate is set to 0.1, and the subsampling ratio is set to 0.8; for samples that are surface extruded after cleaning and drying, the number of trees is set to 100, the maximum depth of the tree is set to 0, the minimum number of samples required for leaf nodes is set to 1, and the minimum number of samples required for node splitting is set to 10; for samples that are surface enhanced with silver nanocolloids after cleaning and drying, the number of trees is set to 100, the maximum depth of the tree is set to 0, the minimum number of samples required for leaf nodes is set to 2, and the minimum number of samples required for node splitting is set to 5; for samples that are surface enhanced with silver nanocolloids after cleaning, drying and surface extrusion, the number of trees is set to 100, the maximum depth of the tree is set to 0, the minimum number of samples required for point splitting is set to 2, and the minimum number of samples required for leaf nodes is set to 2.
[0095] The model evaluation method of this embodiment is the same as that of the first embodiment.
[0096] Example 4
[0097] Steps S1-S4 and S6 of this embodiment are the same as those of Example 1, with the difference being step S5. The classification model adopted in step S5 of this embodiment is the LR (Logistic Regression) model. Logistic regression adds a Sigmoid function (non-linear) mapping on the basis of linear regression, making logistic regression an excellent classification algorithm. Logistic regression reduces the weight of points that are far from the classification plane through nonlinear mapping, and relatively increases the weight of data points that are most relevant to the classification. For logistic regression, since it may overfit, this embodiment adds a regularization term to it. Regularization is a strategy to minimize structural risk. This embodiment optimizes two hyperparameters: the regularization parameter (C) and the regularization type (penalty).
[0098] The specific optimization parameters are as follows: for samples that are only cleaned and dried for pretreatment, the regularization parameter is set to 0.1 and the regularization type is set to 12; for samples that are surface extruded after cleaning and drying, the regularization parameter is set to 0.001 and the regularization type is set to 12; for samples that are surface enhanced with silver nanocolloids after cleaning and drying, the regularization parameter is set to 1 and the regularization type is set to 12; for samples that are surface enhanced with silver nanocolloids after cleaning, drying and surface extrusion, the regularization parameter is set to 0.001 and the regularization type is set to 12.
[0099] The model evaluation method of this embodiment is the same as that of the first embodiment.
[0100] Example 5
[0101] Steps S1-S4 and S6 of this embodiment are the same as those of the first embodiment, with the difference being step S5. The classification model adopted in step S5 of this embodiment is the SVM (Support Vector Machine) model. The support vector machine is a binary classification model, and its basic model is a linear classifier with the largest interval defined in the feature space. The learning algorithm of SVM is an optimization algorithm for solving convex quadratic programming. SVM can successfully solve the classification and local minimum problems of high-dimensional features and has better generalization ability. Even if the feature dimension is larger than the sample data, it can still perform well.
[0102] Through the kernel function, the support vector machine can map the feature vector to a higher-dimensional space, making the originally linearly inseparable data become linearly separable in the mapped space. Therefore, this embodiment adjusts the type of kernel function (kernel) and the coefficient of the kernel function (gamma). In addition, the regularization parameter (C) is also adjusted, which adds a penalty for each misclassified data point and needs to be optimized. The larger the value of C, the smaller the margin will be selected if the hyperplane can better classify all training points.
[0103] The specific optimization parameters are as follows: for samples that are only cleaned and dried for pretreatment, the type of kernel function is set to linear, the coefficient of the kernel function is set to scale, and the regularization parameter is set to 0.1; for samples that are surface extruded after cleaning and drying, the type of kernel function is set to rbf, the coefficient of the kernel function is set to auto, and the regularization parameter is set to 10; for samples that are surface enhanced with silver nanocolloids after cleaning and drying, the type of kernel function is set to linear, the coefficient of the kernel function is set to scale, and the regularization parameter is set to 0.1; for samples that are surface enhanced with silver nanocolloids after cleaning, drying and surface extrusion, the type of kernel function is set to linear, the coefficient of the kernel function is set to scale, and the regularization parameter is set to 0.1.
[0104] The model evaluation method of this embodiment is the same as that of the first embodiment.
[0105] Example 6
[0106] Steps S1-S4 and S6 of this embodiment are the same as those of the first embodiment, with the difference being step S5. The classification model adopted in step S5 of this embodiment is the KNN (K-Nearest Neighbors) model. In K-Nearest Neighbors, it is assumed that all samples are points in n-dimensional space, and their nearest neighbor points are defined according to an appropriate distance metric. KNN has several common distance measurement methods, such as Euclidean distance, which is the straight-line distance between two points in Euclidean space. By calculating the Euclidean distance, the nearest neighbors of a given sample can be identified, and predictions can be made based on the majority class (for classification) or average value (for regression) of the neighbors. Its formula in n-dimensional space is as follows:
[0107]
[0108] The values x1, x2,…, xn on the coordinate axis are the n features of the sample data.
[0109] There is also Manhattan distance, also known as L1 distance or taxi distance, which is the sum of the absolute differences between the coordinates of two points. It represents the shortest path between points when movement is restricted to a grid-like structure, similar to a taxi driving on city streets. Compared to Euclidean distance, Manhattan distance is less sensitive to outliers because it does not square the differences. This can make it more suitable for certain data sets or problems where the presence of outliers may have a significant impact on the performance of the model. Its formula is as follows:
[0110]
[0111] This embodiment selects different numbers of neighbors (n_neighbors), different weighting methods (weights), and different distance metrics (p) in the KNN algorithm. The distance metrics include Manhattan distance and Euclidean distance.
[0112] The specific optimization parameters are as follows: for samples that were only cleaned and dried for pretreatment, the number of neighbors was set to 5, the weighting method was set to uniform, and the distance metric was set to Manhattan distance; for samples that were cleaned and dried and then surface extruded, the number of neighbors was set to 7, the weighting method was set to distance, and the distance metric was set to Manhattan distance; for samples that were surface-enhanced with silver nanocolloids after cleaning and drying, the number of neighbors was set to 5, the weighting method was set to uniform, and the distance metric was set to Euclidean distance; for samples that were surface-enhanced with silver nanocolloids after cleaning, drying, and surface extrusion, the number of neighbors was set to 5, the weighting method was set to uniform, and the distance metric was set to Manhattan distance.
[0113] The model evaluation method of this embodiment is the same as that of the first embodiment.
[0114] Example 7
[0115] Steps S1-S4 and S6 of this embodiment are the same as those of the first embodiment, except for step S5. The classification model used in step S5 of this embodiment is the NB (Naive Bayes) model. Naive Bayes is the simplest and most common classification method in Bayesian classification. Using the Bayesian rule and the strong assumption that attributes are conditionally independent, the core is the following Bayesian formula:
[0116]
[0117] For the Naive Bayes algorithm, this embodiment adjusts the smoothing parameter (var_smoothing).
[0118] The specific optimization parameters are: for samples with only clean and dry pretreatment, the smoothing parameter is set to 10-9 For the samples that were cleaned and dried before surface extrusion, the smoothing parameter was set to 10. -9 For the samples that were surface-enhanced with silver nanocolloids after cleaning and drying, the smoothing parameter was set to 10 -9 For the samples that were cleaned, dried, and surface-extruded and then surface-enhanced with silver nanocolloids, the smoothing parameter was set to 10. -9 .
[0119] The model evaluation method of this embodiment is the same as that of the first embodiment.
[0120] The performance evaluation of the classification models designed in the above seven embodiments is shown in Table 2 below:
[0121]
[0122]
[0123] Table 2. Classification model performance for identifying male and female leaves of Litsea cubeba
[0124] In Table 2, "Surface" refers to surfaces that were simply cleaned and dried, "Extrusion" refers to surfaces that were cleaned and dried and then extruded, "Surface Enhancement" refers to surfaces that were cleaned and dried and then enhanced with silver nanocolloids, and "Extrusion Enhancement" refers to surfaces that were cleaned, dried, and extruded and then enhanced with silver nanocolloids. As shown in Table 2, the combination of surface enhancement with the LR, SVM, and XGBoost models achieved the highest accuracy in distinguishing male and female leaves, reaching 85%.
[0125] The Raman spectra of the samples processed in step S2 were collected, and the Raman spectra of leaves of different sexes after data preprocessing and surface enhanced Raman spectra were averaged and calculated. The data of male and female leaves of each method were subtracted to form a differential spectrum. It can be seen in the differential spectrum that since the two spectral results of male and female leaves are similar, the spectral differences between different sexes cannot be directly identified. In addition to cellulose, lignin and pectin, proteins, chlorophyll and carotenoids in leaves can be listed as additional metabolite groups. Part of the Raman signal on the leaf surface may come from carotenoids, such as Figure 2 As shown, the blade surface is located at 1001cm -1 、1156cm -1 、1184cm -1 、1528cm -1 The Raman signals are attributed to the in-plane vibration of CH3, CC stretching vibration, CC chain stretching vibration and C=C stretching vibration. Slight shift (≤5cm -1 ) Several peaks of β-carotene and lutein were observed. The peak of the leaf is located at 1528cm-1 The C=C stretching vibration at 1156cm is stronger than that at 1156cm -1 The difference in intensity may be because lutein is an oxygen-containing carotene, and the oxidized hydroxyl group may affect the vibration of the C-C or C=C ring and adjacent positions. Therefore, it is found that the spectral information of the leaf at these positions is more biased towards lutein. Similar characteristic peaks are also observed in the Raman spectrum of the leaf squeeze, such as Figure 4 . Figure 3 and Figure 4 1605cm -1 and 1603cm -1 The signal may be derived from the aromatic stretching of lignin, 1109cm -1 The bands may correspond to the stretching of CC and CO bonds in cellulose. Raman intensity 570-725 cm -1 It may be related to the deformation vibration of C—C—O in carbohydrates. It can be observed that the Raman spectrum before and after extrusion does not change much, and the basic Raman peak position shifts slightly, only at 1528cm -1 However, the surface enhanced Raman spectrum changed significantly after adding nanosilver, and there was also a clear difference between the spectrum before and after squeezing.
[0126] Since the Raman spectroscopy (RS) signal is very weak, in order to confirm that the obtained spectral information is not the characteristics of AgNPs (silver nanoparticles), the prepared sodium citrate-coated AgNPs were characterized. It was found that under a laser power of 15 mW and an acquisition time of 3 s, the relative intensity of AgNPs was only 100-200, and its characteristic peaks were mainly located at 835, 945, 1037, and 1406 cm -1 There was no effect on the spectra of the four treatments of Litsea cubeba leaves. Figure 5 This is the Raman spectrum after surface enhancement (using silver nanocolloids for surface enhancement), and it is found that Figure 6 The data gap between extrusion enhancement (surface enhancement with silver nanocolloids after surface extrusion) is quite large. One reason is that the difference in overall Raman intensity may be attributed to the fact that surface enhancement treatment is simply to drip AgNPs on the leaf surface, while extrusion enhancement is to mix the leaf juice with AgNPs on the leaf surface. The contact area of the latter is wider, and the interaction between the active ingredients and AgNPs is stronger, which makes the Raman intensity of extrusion enhancement significantly stronger than the other three pretreatment methods, while surface enhancement is only slightly stronger than the two unenhanced methods. The second reason is that the enhanced Raman characteristic peaks are different. From Figure 5 It can be seen that the addition of AgNPs mainly enhanced the peaks at 940, 1046, 1291, and 1386 cm -1 The characteristic peak of Figure 6It is shown that AgNPs mainly enhance the 1387, 1537, and 1590 cm -1 Therefore, surface enhancement mainly enhances the peaks of carbohydrates and aliphatic groups, while extrusion enhancement mainly shows the vibrations of active ingredients contained in leaves, such as ethyl linolenate and phytol.
[0127] In order to simply count the potential spectral differences between female and male leaves, Figure 3-Figure 6 Further analysis shows that the difference in the spectra of male and female leaves collected on the surface is scattered on both sides of the dotted line, indicating that the difference in the average spectrum of male and female leaves is that the intensity of some Raman peak positions is stronger in male leaves than in female leaves, while the opposite is true in some positions. -1 、1156cm -1 、1184cm -1 、1528cm -1 ) are almost always stronger in female leaves than in male leaves, likely because female Litsea cubeba leaves have more lutein than males. Compared to the surface treatment, the enhanced treatment showed stronger male leaves than female leaves in almost every stripe position.
[0128] exist Figure 2-6 In the formula, SML, SFL, EML, EFL, SEML, SEFL, EEML, and EEFL mean the following:
[0129] SML, surface of male leaves, only clean and dry male leaves;
[0130] SFL, surface of female leaves, only clean and dry female leaves;
[0131] EML, Extrusion of male leaves, male leaves that are surface squeezed after cleaning and drying;
[0132] EFL, Extrusion of female leaves, female leaves that are surface squeezed after cleaning and drying;
[0133] SEML, surface enhancement of male leaves, male leaves that were cleaned and dried and then surface-enhanced with silver nanocolloids;
[0134] SEFL, surface enhancement of female leaves, female leaves that were cleaned and dried and then surface-enhanced with silver nanocolloids;
[0135] EEML, Extrusion enhancement of male leaves, male leaves that were cleaned, dried, and surface-extruded and then surface-enhanced with silver nanocolloids;
[0136] EEFL, Extrusion enhancement of female leaves, female leaves that are cleaned, dried, and surface-extruded and then surface-enhanced with silver nanocolloids.
[0137] The working principle of the present invention is:
[0138] There are differences in the content of substances such as lignocellulose, non-structural sugars (such as glucose, mannitol, etc.), and carotenoids on the surface of plant leaves. Lignocellulose is composed of carbohydrates (such as cellulose, hemicellulose) and aromatic polymers (lignin), while there are certain differences in chemical composition inside the leaves. Therefore, these differences can be reflected by Raman spectroscopy. The present invention not only explores the Raman spectrum and surface-enhanced Raman spectroscopy of leaves, but also compares the results of measurements on the leaf surface and squeezed juice. Based on the obtained spectral features, a PLA-LDA combined with DT, LR, KNN and other algorithm models are constructed to distinguish between female and male leaves.
[0139] Compound identification of ethanol extracts from male and female leaves of Litsea cubeba revealed 23 compounds from male leaves and 20 compounds from female leaves, listed in Table 3 below. Both male and female leaves showed the highest levels of phytols, fatty acids, and higher fatty acid esters. The highest concentration in both male and female leaves was ethyl linolenate, at 15.79% in male leaves and 12% in female leaves. This was followed by phytol, at 12.2% in male leaves and 11.62% in female leaves. Other compounds with significant concentrations included linolenic acid, ethyl linoleate, hexadecanoic acid (ethyl ester), and n-hexadecanoic acid. In addition, 0.55% linalool, 0.25% 1,4-pentadiene, 1.82% gamma-sitosterol, 0.46% nonacos-1-ene, and 0.79% 1-tetracosene are unique components in male leaves; 0.37% β-elemene is a unique component in female leaves.
[0140]
[0141]
[0142] Table 3 GC-MS analysis results of male and female leaves of Litsea cubeba
[0143] The present invention also relates to the characterization of silver nanoparticles. Silver nanocolloids were prepared using sodium citrate as a reducing agent and stabilizer. The maximum ultraviolet absorption peak of the silver nanoparticles was located at 412 nm, indicating successful synthesis of the silver nanoparticles. The particle size of the silver nanoparticles was approximately 100 nm, and the synthesized silver nanoparticles were negatively charged. This is because when sodium citrate was used as a reducing agent, the citrate radicals formed a strong coordination bond with the surface of AgNPs. The citrate radicals bound to the surface of the AgNPs can ionize under appropriate pH conditions, resulting in a negative surface charge. This can, to a certain extent, prevent the aggregation of the AgNPs and ensure their good dispersibility in the aqueous phase.
[0144] The Raman enhancement effect of silver nanoparticles was tested using R6G as a probe molecule. Silver nanoparticles were able to detect 10 - 9 The SERS (surface enhanced Raman scattering) signal of R6G has a good enhancement effect; the R6G concentration is close to that at 1512 cm -1 The characteristic peak intensity at the position shows a good linear relationship, R 2 =0.9921.
[0145] The above description is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this field, several variations and improvements can be made without departing from the creative concept of the present invention, which all fall within the scope of protection of the present invention.
Claims
1. A method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy, characterized in that: The following steps are involved: S1: Collect leaves from male and female Litsea cubeba plants; S2: The female and male plant leaves were divided into four groups, and the leaves were treated using the following four methods to obtain pre-treated leaves: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nanocolloid after cleaning and drying, and surface enhancement using silver nanocolloid after cleaning and drying and surface extrusion; S3: collecting the Raman spectrum of the pre-processed leaf, and performing shift shearing, smoothing filtering and background noise reduction processing on the Raman spectrum; S4: Perform principal component analysis and linear discriminant analysis on the Raman spectrum data processed in step S3, reduce the dimensionality and classify the input data, and divide it into a training set and a test set according to a set ratio; S5: Build a classification model, train the classification model using the training set, test the training results using the test set, and select the best model based on accuracy, precision, recall, and F1 score; S6: Collect leaves of Litsea cubeba to be tested, and after processing in steps S2-S4, input them into the optimal model for classification, and the optimal model outputs the classification results, which include female plants and male plants.
2. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 1, wherein The preparation method of the silver nano colloid is as follows: Weigh a certain amount of AgNO3 solid into a container and prepare a concentration of 1×10 -3 mol﹒ L -1 AgNO3 solution; Stirring the AgNO3 solution at a set speed until it is heated to boiling; After the AgNO3 solution is boiled, a certain amount of 1% sodium citrate solution is added and the solution is boiled for 30 minutes. The solution was cooled at room temperature, and the supernatant was discarded after centrifugation to obtain a silver sol enhancer.
3. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 1, wherein The surface extrusion is specifically as follows: using a tool with a round blunt head to press the surface of the collected Litsea cubeba leaves to be tested until the surface of the leaves to be tested is damaged; The surface enhancement using silver nano-colloid is specifically as follows: a silver sol enhancer is dripped onto the surface of the Litsea cubeba leaf where the test is to be performed, and the silver sol is naturally air-dried.
4. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 1, wherein In step S3, a portable Raman spectrometer is used to collect Raman spectra with an excitation wavelength of 785 nm.
5. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 1 or 4, wherein When the Raman spectrum is collected in step S3, the laser power is set to 15 mW and the collection time is set to 3 s.
6. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 1, wherein The classification model is one of decision tree, random forest, extreme gradient boosting, logistic regression, support vector machine, K nearest neighbor, and naive Bayes.
7. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 6, wherein: The number of trees in the extreme gradient boosting is set to 100, 100, 100, and 100, respectively, and the maximum depth of the tree is set to 4, 0, 0, and 0, respectively; corresponding to four leaf pretreatment methods: only cleaning and drying, cleaning and drying followed by surface extrusion, cleaning and drying followed by surface enhancement using silver nanocolloids, and cleaning and drying and surface extrusion followed by surface enhancement using silver nanocolloids.
8. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 6, wherein: The regularization parameters of the logistic regression were set to 0.1, 0.001, 1, and 0.001, respectively, and the regularization type parameters were set to 12, 12, 12, and 12, respectively; corresponding to four leaf pretreatment methods: only cleaning and drying, surface extrusion after cleaning and drying, surface enhancement using silver nanocolloids after cleaning and drying, and surface enhancement using silver nanocolloids after cleaning and drying and surface extrusion.
9. The method for early identification of male and female Litsea cubeba plants based on Raman spectroscopy according to claim 6, wherein: The regularization parameters of the support vector machine are set to 0.1, 10, 0.1, and 0.1, respectively; the kernel function types are set to linear, rbf, linear, and linear, respectively; and the coefficients of the kernel function are set to scale, auto, scale, and scale; corresponding to four leaf pretreatment methods: only cleaning and drying; cleaning and drying followed by surface extrusion; cleaning and drying followed by surface enhancement using silver nanocolloids; and cleaning and drying followed by surface extrusion followed by surface enhancement using silver nanocolloids.
Citation Information
Patent Citations
SNP markers for sex identification in Litsea cubeba and methods for screening these SNP markers
CN106987652B
Automated noninvasive determining the sex of an embryo of and the fertility of a bird's egg
CN111344587A
Method for non-destructive sex identification in middle stage of chicken embryo by utilizing shin color inheritance and transmission spectrum
CN115521984A