Method, device and equipment for ploidy determination of dried chrysanthemum indicum flowers
The ploidy determination model of wild chrysanthemum flower established through hyperspectral technology and machine learning algorithms solves the problems of inaccurate measurement results and complicated operations in the existing technology, realizes fast and accurate ploidy recognition, and simplifies the operation process.
Patent Information
- Application Number
- CN202510634961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art has problems such as inaccurate results and complicated operations in the determination of ploidy of wild chrysanthemums, especially the morphological identification method is susceptible to environmental factors, and the molecular marker method is costly and complicated.
Hyperspectral technology combined with machine learning algorithms is adopted to collect hyperspectral data of wild chrysanthemums through hyperspectrometers, and a ploidy measurement model is established using support vector machines, generalized linear models, linear discriminant analysis models and quadratic discriminant analysis models to preprocess the data and filter the characteristic bands to realize automated analysis.
It improves the accuracy of the determination of ploidy of wild chrysanthemum flower and simplifies the operation steps, and can quickly and accurately identify diploid and tetraploid plants, reducing the subjectivity and technical errors of artificial interpretation.
Smart Images

Figure CN120558869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ploidy determination of dried chrysanthemum wild flowers, and in particular to a method, device and equipment for determining the ploidy of dried chrysanthemum wild flowers. Background Art
[0002] Chrysanthemum indicum is a cross-pollinated plant that often hybridizes with closely related species under natural conditions, primarily existing in diploid and tetraploid forms. Currently, commercially available Chrysanthemum indicum medicinal materials exhibit mixed ploidy, with both diploid and tetraploid plants coexisting. It is noteworthy that as a medicinal plant whose flowers are used as medicine, the polyploidy of Chrysanthemum indicum significantly affects the content of its active ingredients. Among them, mongoside, a key indicator of Chrysanthemum indicum quality, is typically higher in diploid plants, while its content in tetraploid plants may be reduced due to gene loss or frameshifts. Dried flowers, a common form of Chrysanthemum indicum medicinal materials, are easy to store, transport, and use. Therefore, rapid identification of Chrysanthemum indicum medicinal materials and effective quality control are of great significance.
[0003] Existing technologies usually use morphological identification and molecular marker methods to identify the ploidy of wild chrysanthemum. However, the morphological identification method is easily affected by environmental factors. Plants under different environmental conditions exhibit different morphological characteristics, which increases the difficulty and uncertainty of identification. In addition, the morphological characteristics must be established at a specific growth stage, which limits the identification period. The molecular marker method is costly, has complex operating steps, and there are technical errors in the experimental process, which affects the accuracy of the identification results. Summary of the Invention
[0004] In view of this, the present invention provides a method for determining the ploidy of dried chrysanthemum wild flowers to solve the problem that the prior art has inaccurate results and complicated operations for determining the ploidy of dried chrysanthemum wild flowers.
[0005] In a first aspect, the present invention provides a method for determining the ploidy of dried wild chrysanthemum flowers, which obtains target hyperspectral data collected by a hyperspectrometer for the dried wild chrysanthemum flowers to be identified; the target hyperspectral data is input into a pre-trained wild chrysanthemum flower ploidy determination model, and a prediction result output by the wild chrysanthemum flower ploidy determination model is obtained; the wild chrysanthemum flower ploidy determination model is a machine learning model; and the ploidy determination result of the dried wild chrysanthemum flowers to be identified is obtained based on the prediction result.
[0006] Hyperspectral technology is an advanced remote sensing technique that combines imaging and spectral analysis. It can obtain spectral information of target objects across hundreds of continuous, narrow bands. Compared to conventional multispectral imaging (which uses only a few discrete bands), hyperspectral data provides richer spectral detail, revealing subtle characteristics of substances and is widely used in scientific research and engineering. Hyperspectral imaging uses hundreds of bands, typically 200-300, far exceeding the 3-10 discrete bands of multispectral imaging. Band widths are less than 10 nanometers, forming a continuous spectral curve that can capture subtle characteristics of substances. Each pixel contains not only spatial information but also a complete spectral signature. The spectral signatures of different substances are unique. Leveraging these properties, hyperspectral technology can be used for precise classification and identification. In use, hyperspectral technology typically collects data through a hyperspectral sensor, preprocesses the data, extracts characteristic bands, and then uses machine learning or spectral matching algorithms to classify and identify substances. Currently, hyperspectral technology is mainly used for agricultural monitoring, environmental monitoring, geological exploration, pathological tissue analysis, and drug component detection. The principle of imaging using hyperspectral technology is primarily based on high-resolution spectral coverage, revealing the unique spectral fingerprint of a substance. Specifically, traditional imaging relies on recording three broad bands of red, green, and blue, or expanding them into 5-10 discrete bands, but is unable to capture continuous spectral details. Hyperspectral imaging, on the other hand, uses spectrometers to decompose incident light into hundreds of continuous, narrow bands. Each pixel not only records brightness but also contains a complete spectral curve, reflecting the reflection / absorption characteristics of the target at that location for different wavelengths of light, thereby enabling identification or classification of the target. Currently, hyperspectral technology expands the adaptability of target data collection by selecting multiple data collection methods, while also using various correction methods to remove erroneous data to obtain accurate data. Spectral feature extraction is performed on the collected data to reduce redundant data, and corresponding results are matched through training models. The present invention utilizes hyperspectral data that combines spectral information with image data to simultaneously obtain spectral information for each pixel in the image, enabling detailed analysis of heterogeneous solid samples. Furthermore, to address the influence of peaks, noise, and external interference in spectral data, the present invention incorporates machine learning algorithms into the automated analysis of spectral data. Machine learning algorithms, as a fundamental algorithm in the field of artificial intelligence, can simulate human decision-making capabilities. Analysis using machine learning algorithms can eliminate the subjectivity of manual interpretation and generate repeatable results that can be evaluated using performance indicators. In addition, machine learning technology is good at identifying subtle changes in spectra. After proper training, it can generate results quickly and show high reliability in continuous monitoring.
[0007] In an optional embodiment, the ploidy determination model of dried chrysanthemum wild flowers includes at least one of a support vector machine model, a generalized linear model, a linear discriminant analysis model, and a quadratic discriminant analysis model. The present invention utilizes a support vector machine model, a generalized linear model, a linear discriminant analysis model, and a quadratic discriminant analysis model to establish a correspondence between hyperspectral data and the ploidy label results of dried chrysanthemum wild flowers. Among them, the support vector machine model has strong high-order data processing capabilities and nonlinear classification capabilities, can maximize the classification of hundreds of hyperspectral data, and can solve the problem of classification boundaries. The generalized linear model processes data flexibly and has high computational efficiency, and can quickly process hyperspectral data. The linear discriminant analysis model has good classification capabilities and can effectively distinguish hyperspectral data from ploidy label data. The quadratic discriminant analysis has a strong ability to model boundaries and clearly distinguish different types of data.
[0008] In an optional embodiment, the wavelength of the preset wavelength band light is 400-1000 nm.
[0009] In an optional embodiment, before inputting the target hyperspectral data into a pre-trained wild chrysanthemum dried flower ploidy determination model, the method further includes: preprocessing the target hyperspectral data, and the preprocessing method includes at least one of an SG smoothing algorithm, a standard normal variable transformation algorithm, or a first-order derivative algorithm.
[0010] The present invention uses at least one of the SG smoothing algorithm, the standard normal variable transformation algorithm or the first-order derivative algorithm to pre-process the hyperspectral data to improve the accuracy of the measurement results. Among them, the SG smoothing algorithm smoothes the spectral curve by local polynomial fitting, effectively suppresses random noise, and retains the key features of the spectrum such as peak shape and peak width, and can select the window size and polynomial order according to the characteristics of the data, based on the convolution operation, to process large-scale hyperspectral data; the standard normal variable transformation algorithm eliminates the scattering interference caused by the surface roughness of the sample, particle size or optical path difference by performing individual standardization on each sample spectrum, and can map the spectra of different samples to the same scale, which is more conducive to the comparison of spectral features across samples, that is, it can specifically identify dried wild chrysanthemum flowers of different germplasm sources; the first-order derivative algorithm effectively removes the baseline drift in the spectrum by calculating the difference between adjacent wavelengths, enhances the resolution of spectral details, and improves the accuracy of the measurement results.
[0011] In an optional embodiment, the target hyperspectral data is input into a pre-trained wild chrysanthemum dried flower ploidy determination model, comprising: using a characteristic band screening method to screen the target hyperspectral data to obtain hyperspectral data of a target spectral band; inputting the hyperspectral data of the target spectral band into a pre-trained wild chrysanthemum dried flower ploidy determination model; wherein the characteristic band screening method comprises: at least one of an uninformative variable elimination algorithm, a competitive adaptive reweighted sampling algorithm, and a continuous projection algorithm.
[0012] Since hyperspectral data usually includes hundreds of bands of data, high-dimensional data will lead to data sparsity, making it difficult for the model to effectively learn features. At the same time, it will also lead to a surge in computational complexity, exponential growth in algorithm training time and storage requirements, and highly similar data information will also cause information duplication. The present invention uses a method of feature band screening to screen the target hyperspectral data, remove redundant data, and screen key bands with rich information and low redundancy, which can not only optimize model performance but also improve analysis efficiency. Among them, the non-information variable elimination method can analyze the stability and uncertainty of the regression coefficient, use the random disturbance introduced by the variable as a reference threshold, and eliminate bands that do not contribute to the model or whose contribution is unstable; competitive adaptive reweighted sampling gradually eliminates low-weight bands through Monte Carlo sampling and exponential decay function, and adaptively retains bands with large information content and high correlation with the target variable; the continuous projection algorithm selects bands with strong orthogonality to each other through vector projection, ensuring the information complementarity of the selected bands to reduce redundancy.
[0013] In an optional embodiment, before inputting the target hyperspectral data into the pre-trained wild chrysanthemum dried flower ploidy determination model, it also includes: obtaining the wild chrysanthemum dried flower ploidy determination model to be trained; obtaining sample hyperspectral data collected by a hyperspectrometer for sample wild chrysanthemum dried flowers and the corresponding wild chrysanthemum ploidy labels, and preprocessing the sample hyperspectral data; the preprocessing includes: at least one of the SG smoothing algorithm, the standard normal variable transformation algorithm or the first-order derivative algorithm; and using the preprocessed sample hyperspectral data and the corresponding wild chrysanthemum ploidy labels to train, test and verify the wild chrysanthemum dried flower ploidy determination model to be trained.
[0014] The present invention obtains a ploidy determination model for dried chrysanthemum flowers to be trained, obtains pre-processed hyperspectral data and ploidy labels for dried chrysanthemum flowers, and uses the pre-processed hyperspectral data and ploidy labels to train, test, and verify the ploidy determination model for dried chrysanthemum flowers to be trained, thereby improving the accuracy of the model measurement results. Among them, the method for training the measurement model can be divided into a training set, a validation set, and a test set, and the data is trained by model architecture, defining optimization goals, selecting algorithms, etc.; the method for verifying the measurement model can adopt a cross-validation method, an early stopping method, or a hyperparameter tuning method, wherein the cross-validation method can select a K-fold cross-validation, an outflow validation, or a stratified sampling method to verify the model; the method for testing the measurement model can test the model through a single evaluation, setting a reasonable evaluation index, and robustness analysis; it should be noted that the present invention isolates the data of the training set, the validation set, and the test set from each other, which can effectively improve the accuracy of the constructed model detection.
[0015] In an optional embodiment, the preprocessed sample hyperspectral data and the corresponding wild chrysanthemum ploidy label are used to train, test and verify the wild chrysanthemum dried flower ploidy determination model to be trained, including: characteristic band screening of the preprocessed sample hyperspectral data to obtain sample hyperspectral data of the target spectral band, the characteristic band screening method includes at least one of an uninformative variable elimination algorithm, a competitive adaptive reweighted sampling and a continuous projection algorithm; the sample hyperspectral data of the target spectral band and the corresponding wild chrysanthemum ploidy label are used to train, test and verify the wild chrysanthemum dried flower ploidy determination model to be trained.
[0016] The present invention performs characteristic band screening on pre-processed hyperspectral data to obtain target data for model training, testing and verification, which can not only improve the accuracy of model measurement, but also improve the efficiency of model optimization.
[0017] In an optional embodiment, the preprocessed sample hyperspectral data and the corresponding wild chrysanthemum ploidy label are used to train, test and verify the wild chrysanthemum dried flower ploidy determination model to be trained, including: dividing the preprocessed sample hyperspectral data and the corresponding wild chrysanthemum ploidy label into a training set, a test set and a validation set; for the training set, using the K-fold cross-validation method, the wild chrysanthemum dried flower ploidy determination model to be trained is trained to obtain the trained wild chrysanthemum dried flower ploidy determination model; for the trained wild chrysanthemum dried flower ploidy determination model, the test set is used for testing and the validation set is used for verification.
[0018] The present invention adopts the K-fold cross-validation method to train the ploidy determination model of dried chrysanthemum wild flowers to be trained. K-fold cross-validation is a model verification method for efficient data utilization. By dividing the data set multiple times, the value of the data is fully explored and the evaluation bias caused by the randomness of a single data division is reduced. Specifically, the data set is divided into K mutually non-overlapping subsets, and K rounds of iterations are performed. The K verification results are averaged as the final performance index of the model, wherein K can be selected as 5 or 10. The present invention adopts K-fold cross-validation to make full use of data, reduce evaluation variance, reduce the random influence of a single division, and objectively reflect the accuracy of the model.
[0019] In an optional embodiment, the hyperspectral instrument has a movement speed of 1.6 mm / s to 1.8 mm / s and an exposure time of 5 ms to 7 ms when collecting hyperspectral data. The present invention limits the movement speed and exposure time when collecting hyperspectral data to more accurately obtain hyperspectral data and reduce errors.
[0020] In a second aspect, the present invention provides a device for determining the ploidy of dried wild chrysanthemum flowers, the device comprising: a data acquisition module for collecting target hyperspectral data of dried wild chrysanthemum flowers to be identified, and preprocessing the target hyperspectral data to obtain preprocessed sample hyperspectral data; a data processing module for using a target machine learning model to predict the preprocessed sample hyperspectral data and output a prediction result; and a ploidy identification module for obtaining a corresponding ploidy identification result based on the prediction result.
[0021] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the above-mentioned method for determining the ploidy of dried chrysanthemum indicum by executing the computer instructions.
[0022] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for determining the ploidy of dried chrysanthemum indicum according to the first aspect or any corresponding embodiment thereof.
[0023] In a fifth aspect, the present invention provides a computer program product comprising computer instructions for causing a computer to execute the method for determining the ploidy of dried chrysanthemum indicum according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 1 is a flow chart of a method for determining the ploidy of dried chrysanthemum indicum flowers according to an embodiment of the present invention;
[0026] Figure 2 1 is a schematic flow chart of a second method for determining the ploidy of dried chrysanthemum indicum flowers according to an embodiment of the present invention;
[0027] Figure 3 1 is a schematic flow chart of a third method for determining the ploidy of dried chrysanthemum indicum flowers according to an embodiment of the present invention;
[0028] Figure 4 4 is a schematic flow chart of a method for determining the ploidy of dried chrysanthemum indicum flowers according to an embodiment of the present invention;
[0029] Figure 5 1 is a schematic flow chart of a fifth method for determining the ploidy of dried chrysanthemum indicum flowers according to an embodiment of the present invention;
[0030] Figure 6 is a graph showing the change in the original average hyperspectral reflectance of the spectrum of the sample in different bands according to an embodiment of the present invention;
[0031] Figure 7 1 is a reflectivity curve diagram after preprocessing by the SG smoothing algorithm in an embodiment of the present invention;
[0032] Figure 8 1 is a reflectivity curve diagram after preprocessing by the SG smoothing algorithm-first-order derivative algorithm in an embodiment of the present invention;
[0033] Figure 9 3. This is a reflectivity curve diagram after preprocessing by the SG smoothing algorithm-standard normal variable transformation algorithm in an embodiment of the present invention;
[0034] Figure 10 This is a feature extraction diagram obtained by filtering data after preprocessing by the SG smoothing algorithm and then using the non-information variable elimination method in an embodiment of the present invention;
[0035] Figure 11 This is a feature extraction diagram of data pre-processed by the SG smoothing algorithm in an embodiment of the present invention and screened by competitive adaptive reweighted sampling;
[0036] Figure 12 This is a feature extraction graph obtained by filtering the data after preprocessing by the SG smoothing algorithm and then filtering by the continuous projection algorithm in an embodiment of the present invention;
[0037] Figure 13 It is a feature extraction diagram of data pre-processed by SG smoothing algorithm-standard normal variable transformation and then screened by competitive adaptive reweighted sampling in an embodiment of the present invention;
[0038] Figure 14 This is a feature extraction diagram of data pre-processed by the SG smoothing algorithm-standard normal variable transformation and screened by the continuous projection algorithm in the embodiment of the present invention;
[0039] Figure 15 This is a feature extraction diagram of data after preprocessing by the SG smoothing algorithm-standard normal variable transformation and screening by the non-information variable elimination method in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following examples are provided for a better understanding of the present invention and are not intended to limit the best mode of implementation. They do not limit the content and scope of protection of the present invention. Any product identical or similar to the present invention obtained by anyone under the guidance of the present invention or by combining the features of the present invention with other prior arts shall fall within the scope of protection of the present invention.
[0041] If no specific experimental steps or conditions are specified in the examples, the conventional experimental steps or conditions described in the literature in this field can be used. If the manufacturer of the reagents or instruments is not specified, they are all commercially available conventional reagents.
[0042] An embodiment of the present invention provides a method for determining the ploidy of dried chrysanthemum wild flowers. A determination model is constructed by utilizing hyperspectral data and ploidy label data of dried chrysanthemum wild flowers. The ploidy of the dried chrysanthemum wild flowers to be identified is determined by the model, thereby achieving the effect of improving the accuracy of the ploidy determination results of the dried chrysanthemum wild flowers and simplifying the determination steps.
[0043] According to an embodiment of the present invention, an embodiment of a method for determining the ploidy of dried wild chrysanthemum flowers is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of executable computer instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] In this embodiment, a method for determining the ploidy of dried chrysanthemum wild flowers is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 1 Flowchart of a method for determining ploidy of dried chrysanthemum indicum flowers according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0045] Step S101, obtaining target hyperspectral data collected by a hyperspectrometer for the dried chrysanthemum to be identified;
[0046] Step S102: inputting the target hyperspectral data into a pre-trained ploidy determination model for dried chrysanthemums of wild flowers, and obtaining a prediction result output by the ploidy determination model for dried chrysanthemums of wild flowers; the ploidy determination model for dried chrysanthemums of wild flowers is a machine learning model;
[0047] Step S103, obtaining the ploidy determination result of the dried chrysanthemum wild chrysanthemum flower to be identified according to the prediction result.
[0048] The method for determining the ploidy of dried chrysanthemum wild flowers provided by the present invention utilizes hyperspectral data that combines spectral information with image data, thereby obtaining the spectral information of each pixel in the image while acquiring the image, and performing a detailed analysis of the inhomogeneous solid sample. At the same time, due to the influence of peaks, noise and external interference in the spectral data, the present invention introduces a machine learning algorithm into the automated analysis of spectral data. As a basic algorithm in the field of artificial intelligence, the machine learning algorithm can simulate human decision-making ability. Through the analysis of the machine learning algorithm, the subjectivity of manual interpretation can be eliminated, and repeatable results that can be evaluated by performance indicators can be generated. In addition, machine learning technology is good at identifying subtle changes in the spectrum, and can quickly generate results after appropriate training, and show high reliability in continuous monitoring. Wherein, the ploidy determination model of dried chrysanthemum wild flowers includes at least one of a support vector machine model, a generalized linear model, a linear discriminant analysis model and a quadratic discriminant analysis model. The present invention uses a support vector machine model, a generalized linear model, a linear discriminant analysis model and a quadratic discriminant analysis model to establish a correspondence between hyperspectral data and the ploidy label results of dried wild chrysanthemum flowers. Among them, the support vector machine model has strong high-order data processing capabilities and nonlinear classification capabilities. It can maximize the classification of hundreds of hyperspectral data and solve the problem of classification boundaries. The generalized linear model processes data flexibly and has high computational efficiency, and can quickly process hyperspectral data. The linear discriminant analysis model has good classification capabilities and can effectively distinguish hyperspectral data from ploidy label data. The quadratic discriminant analysis has a strong boundary modeling capability and clearly distinguishes different types of data.
[0049] In this embodiment, a method for determining the ploidy of dried chrysanthemum indicum flowers is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 2 Flowchart of the method for determining the ploidy of dried chrysanthemum wild flowers according to an embodiment of the present invention, Figure 2 As shown, the process includes the following steps:
[0050] Step S201, see Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0051] Step S202: preprocessing the target hyperspectral data.
[0052] Specifically, the preprocessing method includes at least one of the SG smoothing algorithm, the standard normal variable transformation algorithm, or the first-order derivative algorithm. The embodiment of the present invention uses at least one of the SG smoothing algorithm, the standard normal variable transformation algorithm, or the first-order derivative algorithm to preprocess the hyperspectral data to improve the accuracy of the measurement results. Among them, the SG smoothing algorithm smoothes the spectral curve through local polynomial fitting, effectively suppressing random noise while retaining key features such as the peak shape and peak width of the spectrum, and can select the window size and polynomial order according to the characteristics of the data, based on convolution operations, to process large-scale hyperspectral data; the standard normal variable transformation algorithm eliminates the surface roughness of the sample by individually standardizing each sample spectrum. The scattering interference caused by particle size or optical path difference can be eliminated, and the spectra of different samples can be mapped to the same scale, which is more conducive to the comparison of spectral features across samples, that is, it can specifically identify dried wild chrysanthemum flowers from different germplasm sources; the first-order derivative algorithm effectively removes the baseline drift in the spectrum by calculating the difference between adjacent wavelengths, enhances the resolution of spectral details, and improves the accuracy of the measurement results. Furthermore, the present invention adopts two or three methods of SG smoothing algorithm, standard normal variable transformation algorithm or first-order derivative algorithm to comprehensively process the collected hyperspectral data, and the obtained data is more accurate.
[0053] More specifically, the collected reflectivity data is preprocessed using the SG smoothing algorithm, which is shown in formula (I):
[0054]
[0055] Among them, y′ i is the processed data point, y i+j is the original data point, which refers to the reflectivity collected at different wavelengths, c j is a coefficient obtained by polynomial fitting, k is half of the window size, and preferably, in this embodiment, k is 15.
[0056] The data obtained after SG smoothing algorithm in 1.4.1 is preprocessed using the standard normal variable transformation algorithm (SNV). The SNV algorithm is shown in formula (II):
[0057]
[0058] Among them, z ij is the processed data point, x ij is the value of the jth variable of the i-th sample, that is, the spectral reflectance of the i-th wild chrysanthemum sample under the j-th band, is the mean value of all variables in the i-th sample, s i is the standard deviation of all variables for the jth sample.
[0059] The first-order derivative algorithm (D1) is used to preprocess the data obtained after the SG smoothing algorithm in 1.4.1. The D1 algorithm is shown in formula (III):
[0060]
[0061] in, is the processed data point, R(λ) is the spectral reflectance, that is, the spectral reflectance of each wild chrysanthemum sample in the λ band, λ is the wavelength, △λ is the wavelength increment, and in this embodiment, △λ is 1.
[0062] Formula (III) represents the first derivative of the spectral reflectance R(λ) with respect to the wavelength λ at the wavelength λ.
[0063] For details of steps S203 and S204, please refer to Figure 1 Step S102 and step S103 of the illustrated embodiment will not be described in detail here.
[0064] In this embodiment, a method for determining the ploidy of dried chrysanthemum indicum flowers is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 3 Flowchart of the method for determining the ploidy of dried chrysanthemum wild flowers according to an embodiment of the present invention, Figure 3 As shown, the process includes the following steps:
[0065] Step S301, see Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0066] Step S302 : performing characteristic band screening on the sample hyperspectral data to obtain sample hyperspectral data of a target spectral band.
[0067] Specifically, the characteristic band screening method includes at least one of an uninformative variable elimination algorithm, a competitive adaptive reweighted sampling algorithm, and a continuous projection algorithm. Since hyperspectral data typically includes hundreds of bands, the high dimensionality of the data leads to data sparsity, making it difficult for the model to effectively learn features. This also leads to a surge in computational complexity, exponentially increasing algorithm training time and storage requirements. Highly similar data can also cause information duplication. The present invention utilizes a characteristic band screening method to screen the target hyperspectral data, remove redundant data, and select key bands with rich information and low redundancy. This not only optimizes model performance but also improves analysis efficiency. The uninformative variable elimination method analyzes the stability and uncertainty of the regression coefficients and uses the random perturbations introduced by the variables as a reference threshold to eliminate bands that do not contribute to the model or whose contributions are unstable. The competitive adaptive reweighted sampling algorithm uses Monte Carlo sampling and an exponential decay function to gradually eliminate low-weight bands while adaptively retaining bands with high information content and high correlation with the target variable. The continuous projection algorithm uses vector projection to select bands with strong orthogonality to each other, ensuring information complementarity in the selected bands to reduce redundancy.
[0068] For details of steps S303 and S304, please refer to Figure 1 Step S102 and step S103 of the illustrated embodiment will not be described in detail here.
[0069] In this embodiment, a method for determining the ploidy of dried chrysanthemum indicum flowers is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 4 Flowchart of the method for determining the ploidy of dried chrysanthemum wild flowers according to an embodiment of the present invention, Figure 4 As shown, the process includes the following steps:
[0070] Step S401, see Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0071] Step S402: pre-processing the target hyperspectral data.
[0072] Step S403 , using the pre-processed sample hyperspectral data and the corresponding Chrysanthemum indica ploidy label to train, test and verify the Chrysanthemum indica dried flower ploidy determination model to be trained.
[0073] The present invention utilizes the pre-processed sample hyperspectral data and the corresponding wild chrysanthemum ploidy label to train, test and verify the wild chrysanthemum dried flower ploidy determination model to be trained to improve the accuracy of the model determination results. The method for training the determination model can be to divide the collected data into a training set, a validation set and a test set, and train the data by means of model architecture, definition of optimization goals, selection of algorithms, etc.; the method for verifying the determination model can adopt a cross-validation method, an early stopping method or a hyperparameter tuning method, wherein the cross-validation method can select a K-fold cross-validation, an outflow validation or a stratified sampling method to verify the model; the method for testing the determination model can test the model through a single evaluation, setting a reasonable evaluation index and a robustness analysis; it should be noted that the present invention isolates the data of the training set, the validation set and the test set from each other, which can effectively improve the accuracy of the constructed model detection.
[0074] For details of steps S404 and S405, please refer to Figure 1 Step S102 and step S103 of the illustrated embodiment will not be described in detail here.
[0075] In this embodiment, a method for determining the ploidy of dried chrysanthemum indicum flowers is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 5 Flowchart of the method for determining the ploidy of dried chrysanthemum wild flowers according to an embodiment of the present invention, Figure 5 As shown, the process includes the following steps:
[0076] Step S501, see Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0077] Step S502, preprocessing the target hyperspectral data;
[0078] Step S503, performing characteristic band screening on the pre-processed sample hyperspectral data to obtain sample hyperspectral data of a target spectral band;
[0079] Step S504 , using the sample hyperspectral data of the target spectral band and the corresponding Chrysanthemum indica ploidy label to train, test and verify the Chrysanthemum indica dried flower ploidy determination model to be trained.
[0080] The present invention performs characteristic band screening on pre-processed hyperspectral data to obtain target data for model training, testing and verification, which can not only improve the accuracy of model measurement, but also improve the efficiency of model optimization.
[0081] For details of steps S505 and S506, see Figure 1 Step S102 and step S103 of the illustrated embodiment will not be described in detail here.
[0082] More specifically, when the hyperspectral data is collected using a hyperspectral instrument in the embodiment of the present invention, the moving speed of the hyperspectral instrument is 1.6 mm / s to 1.8 mm / s, and the exposure time is 5 ms to 7 ms.
[0083] As a specific application example of the embodiment of the present invention, the embodiment of the present invention provides a method for determining the ploidy of dried chrysanthemum indicum flowers, and the specific steps and parameters are as follows:
[0084] (1) Ten germplasm wild chrysanthemum plants collected from different regions and planted on the substrate were selected as samples.
[0085] (2) The chromosome counting method was used to identify the ploidy of various accessions of Chrysanthemum indicum. It was determined that 5 accessions were tetraploid (2n=36) and 5 accessions were diploid (2n=18).
[0086] (3) Collect 50 wild chrysanthemums of various qualities and dry them until the moisture content meets the requirements of the "Chinese Pharmacopoeia" (2020 edition). Collect the hyperspectral reflectance data of 400-1000nm of the dried flowers, and the spectral data is the average spectral reflectance of each band in the collected dried flower area. Set the platform movement speed of the spectrometer to 1.71mm / s and the exposure time to 6ms. Before collecting the spectrum, preheat the equipment for 30 minutes and collect black and white board images. When collecting spectral information, randomly select 50 complete dried flowers of various qualities and place them on the displacement stage for scanning. The changes in the hyperspectral original average hyperspectral reflectance of wild chrysanthemums of various qualities are as follows: Figure 6 shown.
[0087] (4) The original hyperspectral data in step 3 are preprocessed using a preprocessing method; SG smoothing can eliminate noise and fluctuations while retaining the data trend, thereby improving the spectral signal-to-noise ratio; D1 can enhance the peak and valley information in the spectrum; SNV can eliminate the scattering effect and baseline drift in the spectral data, making the features in the spectrum more obvious. The results are shown in Figure 7-Figure 9 .
[0088] (5) Among the 10 wild chrysanthemum accessions, 4 diploid accessions and 4 tetraploid accessions were selected, and the pre-processed hyperspectral full spectrum data were divided into training set and test set according to 7:3; the remaining 1 diploid accession and 1 tetraploid accession spectral data were used as independent validation sets. Among them, the training set was used to establish the model structure and parameters, and the prediction set and independent validation set were used to evaluate the model robustness and prediction ability. The SVM, GLM, and LDA models were established using the 10-fold cross-validation method, which is the rapid and non-destructive discrimination model for wild chrysanthemum ploidy.
[0089] The best SVM, GLM, and LDA models constructed based on the full hyperspectral spectrum with different preprocessing methods are shown in Table 1. The performance of the prediction models was evaluated using the accuracy and F1 score of the prediction set and the independent validation set. The closer the accuracy is to 100% and the closer the F1 score is to 1, the better the model performance.
[0090] Table 1 Complete band modeling results under three preprocessing conditions
[0091]
[0092] Table 1 shows that the LDA model performs well under all three preprocessing methods, achieving high accuracy and F1 scores for the prediction set. Among the nine models constructed, the one using SG-SNV as the preprocessing method achieves the best classification performance, achieving 93.33% accuracy and 0.94 F1 scores for the test set, respectively.
[0093] As another specific application embodiment of the present invention, the present invention provides a method for determining the ploidy of dried chrysanthemum indicum flowers, and the specific steps and parameters are as follows:
[0094] (1) Ten germplasm wild chrysanthemum plants collected from different regions and planted on the substrate were selected as samples.
[0095] (2) The chromosome counting method was used to identify the ploidy of various accessions of Chrysanthemum indicum. It was determined that 5 accessions were tetraploid (2n=36) and 5 accessions were diploid (2n=18).
[0096] (3) Fifty flowers of various qualities of wild chrysanthemum were collected and dried until the moisture content met the requirements of the Chinese Pharmacopoeia (2020 edition). Hyperspectral reflectance data of the dried flowers at 400-1000 nm were collected.
[0097] (4) Preprocessing the original hyperspectral data in step 3 using a preprocessing method; wherein the preprocessing method is the SG smoothing method.
[0098] (5) Three feature extraction methods were used to process the full hyperspectral spectrum pre-processed by SG in step (4) of Example 1.
[0099] (6) Among the 10 wild chrysanthemum germplasms, 4 diploid germplasms and 4 tetraploid germplasms were selected, and the pre-processed hyperspectral full spectrum data were divided into training set and test set according to 7:3; the remaining 1 diploid germplasm and 1 tetraploid germplasm spectral data were used as independent validation sets. Among them, the training set was used to establish the model structure and parameters, and the prediction set and independent validation set were used to evaluate the model robustness and prediction ability. The SVM, GLM, and LDA models were established using the 10-fold cross-validation method, which is the rapid and non-destructive discrimination model for wild chrysanthemum ploidy.
[0100] The screening results of UVE, CARS, and SPA are shown in Figure 10-12 The hyperspectral characteristic band model established based on this is shown in Table 2.
[0101] Table 2 Characteristic band modeling results after SG processing
[0102]
[0103]
[0104] Table 2 shows that when the SG preprocessing method is used, the CARS and SPA algorithms outperform the UVE algorithm in predictive performance. The SPA algorithm screened 16 ploidy-related characteristic variables from 765 variables, accounting for only 2.09% of the total frequency. Among the models shown in the table, the SG-SPA-QDA model used 16 bands and achieved accuracies of 96.67% and 94.00% on the prediction set and independent validation set, respectively.
[0105] As another specific application embodiment of the present invention, the present invention provides a method for determining the ploidy of dried chrysanthemum indicum flowers, and the specific steps and parameters are as follows:
[0106] (1) Ten germplasm wild chrysanthemum plants collected from different regions and planted on the substrate were selected as samples.
[0107] (2) The chromosome counting method was used to identify the ploidy of various accessions of Chrysanthemum indicum. It was determined that 5 accessions were tetraploid (2n=36) and 5 accessions were diploid (2n=18).
[0108] (3) Fifty wild chrysanthemums of various qualities were collected and dried until the moisture content met the requirements of the Chinese Pharmacopoeia (2020 edition). Hyperspectral reflectance data of the dried flowers at 400-1000 nm were collected.
[0109] (4) Preprocessing the original hyperspectral data in step 3 using a preprocessing method; wherein the preprocessing method is the SG smoothing method and the SNV method.
[0110] (5) Three feature extraction methods were used to process the full hyperspectral spectrum preprocessed by SG-SNV.
[0111] (6) Among the 10 wild chrysanthemum germplasms, 4 diploid germplasms and 4 tetraploid germplasms were selected, and the pre-processed hyperspectral full spectrum data were divided into training set and test set according to 7:3; the remaining 1 diploid germplasm and 1 tetraploid germplasm spectral data were used as independent validation sets. Among them, the training set was used to establish the model structure and parameters, and the prediction set and independent validation set were used to evaluate the model robustness and prediction ability. The SVM, GLM, and LDA models were established using the 10-fold cross-validation method, which is the rapid and non-destructive discrimination model for wild chrysanthemum ploidy.
[0112] The screening results of UVE, CARS, and SPA are shown in Figure 13-15 The hyperspectral characteristic band model established based on this is shown in Table 3.
[0113] Table 3 Characteristic band modeling results after SG-SNV processing
[0114]
[0115] Table 3 shows that when the SG-SNV preprocessing method was used, the SPA algorithm outperformed the UVE and CARS algorithms in predictive performance. The SPA algorithm screened 10 ploidy-related characteristic variables from 765 variables, accounting for only 1.31% of the total frequency. Compared with the SG-SPA-QDA model constructed in Example 2, the SG-SNV-SPA-GLM model constructed in this example used fewer bands and achieved superior accuracy and F1 scores on the independent validation set, demonstrating that this model is capable of discriminating ploidy across accessions of Chrysanthemum indicum. The results showed that the SG-SNV-SPA-GLM model was the optimal model for discriminating diploids and tetraploids in Chrysanthemum indicum, achieving accuracy of 93.33% and 97.00% on the test set and independent validation set, respectively, with F1 scores of 0.93 and 0.97, respectively.
[0116] In this embodiment, a device for determining the ploidy of dried chrysanthemum wild flowers is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be repeated here. As used below, the term "module" can implement a combination of software and / or hardware that has a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceived. Specifically, this embodiment provides a device for determining the ploidy of dried chrysanthemum wild flowers, comprising a data acquisition module for collecting target hyperspectral data of dried chrysanthemum wild flowers to be identified, and preprocessing the target hyperspectral data to obtain preprocessed sample hyperspectral data; a data processing module for using a target machine learning model to predict the preprocessed sample hyperspectral data and output a prediction result; and a ploidy identification module for obtaining a corresponding ploidy identification result based on the prediction result.
[0117] The ploidy determination device for dried chrysanthemum wild flowers in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0118] An embodiment of the present invention also provides a computer device, which includes: one or more processors, a memory, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system).
[0119] The processor may be a central processing unit, a network processor, or a combination thereof. The processor may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0120] The memory stores instructions that can be executed by at least one processor, so that the at least one processor executes the method shown in the above embodiment.
[0121] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory may also include a combination of the above types of memory.
[0123] The computer device also includes an input device and an output device. The processor, memory, input device 30 and output device can be connected via a bus or other means.
[0124] The input device can receive input digital or character information and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device can include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0125] The computer device further includes a communication interface for the computer device to communicate with other devices or a communication network.
[0126] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0127] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0128] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for determining the ploidy of dried chrysanthemum flowers, characterized in that: The following steps are included: Obtain target hyperspectral data collected by the hyperspectrometer for the dried wild chrysanthemum flowers to be identified; Inputting the target hyperspectral data into a pre-trained ploidy determination model for dried chrysanthemum flowers, and obtaining a prediction result output by the ploidy determination model for dried chrysanthemum flowers, wherein the ploidy determination model for dried chrysanthemum flowers is a machine learning model; The ploidy determination result of the dried chrysanthemum indicum to be identified is obtained according to the prediction result.
2. The method for determining the ploidy of dried chrysanthemum indicum flowers according to claim 1, wherein The ploidy determination model for dried chrysanthemum wild flowers includes at least one of a support vector machine model, a generalized linear model, a linear discriminant analysis model, and a quadratic discriminant analysis model.
3. The method for determining the ploidy of dried chrysanthemum indicum flowers according to claim 1, wherein Before inputting the target hyperspectral data into the pre-trained ploidy determination model for dried chrysanthemum indicum, the method further comprises: The target hyperspectral data is preprocessed, and the preprocessing method includes at least one of an SG smoothing algorithm, a standard normal variable transformation algorithm, or a first-order derivative algorithm.
4. The method for determining the ploidy of dried chrysanthemum indicum flowers according to claim 1 or 3, wherein: The target hyperspectral data is input into a pre-trained ploidy determination model for dried chrysanthemum indicum, comprising: Using a characteristic band screening method to screen the target hyperspectral data to obtain hyperspectral data of a target spectral band; Inputting the hyperspectral data of the target spectral band into a pre-trained ploidy determination model for dried chrysanthemum flowers; The characteristic band screening method includes at least one of an uninformative variable elimination algorithm, a competitive adaptive reweighted sampling algorithm and a continuous projection algorithm.
5. The method for determining the ploidy of dried chrysanthemum indicum flowers according to claim 1, wherein Before inputting the target hyperspectral data into the pre-trained ploidy determination model for dried chrysanthemum indicum, the method further comprises: Obtaining a ploidy determination model for dried chrysanthemum indicum flowers to be trained; Acquire sample hyperspectral data collected by a hyperspectrometer for sample dried chrysanthemum wild flowers and corresponding ploidy labels of the dried chrysanthemum wild flowers, and preprocess the sample hyperspectral data, wherein the preprocessing includes at least one of an SG smoothing algorithm, a standard normal variable transformation algorithm, or a first-order derivative algorithm; The pre-processed sample hyperspectral data and the corresponding Chrysanthemum indica ploidy labels are used to train, test and verify the Chrysanthemum indica dried flower ploidy determination model to be trained.
6. The method for determining the ploidy of dried chrysanthemum indicum flowers according to claim 5, wherein: The method of training, testing, and verifying the ploidy determination model for dried chrysanthemum indica flowers to be trained by using the preprocessed sample hyperspectral data and the corresponding chrysanthemum indica ploidy label comprises: Performing characteristic band screening on the preprocessed sample hyperspectral data to obtain sample hyperspectral data of a target spectral band, wherein the characteristic band screening method includes at least one of an uninformative variable elimination algorithm, a competitive adaptive reweighted sampling algorithm, and a continuous projection algorithm; The ploidy determination model for dried chrysanthemum indica flowers to be trained is trained, tested and verified using the sample hyperspectral data of the target spectral band and the corresponding chrysanthemum indica ploidy label.
7. The method for determining the ploidy of dried chrysanthemum indicum flowers according to claim 5 or 6, wherein: The method of training, testing, and verifying the ploidy determination model for dried chrysanthemum indica flowers to be trained by using the preprocessed sample hyperspectral data and the corresponding chrysanthemum indica ploidy label comprises: Dividing the preprocessed sample hyperspectral data and the corresponding Chrysanthemum indicum ploidy labels into a training set, a test set, and a validation set; For the training set, the ploidy determination model of the dried chrysanthemum wild flower to be trained is trained using a K-fold cross-validation method to obtain the trained ploidy determination model of the dried chrysanthemum wild flower; The trained model for determining the ploidy of dried chrysanthemum indicum flowers is tested using the test set and verified using the validation set.
8. The method according to claim 1 or 5, characterized in that The moving speed of the hyperspectrometer when collecting hyperspectral data is 1.6 mm / s to 1.8 mm / s, and the exposure time is 5 ms to 7 ms.
9. A device for determining the ploidy of dried chrysanthemum wild flowers, characterized in that: The device comprises: A data acquisition module is used to collect target hyperspectral data of dried wild chrysanthemum flowers to be identified, and to obtain preprocessed sample hyperspectral data by preprocessing the target hyperspectral data; The data processing module is used to use the target machine learning model to predict the preprocessed sample hyperspectral data and output the prediction results; The ploidy identification module is used to obtain the corresponding ploidy identification result based on the prediction result.
10. A computer device, characterized in that: include A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for determining the ploidy of dried chrysanthemum indicum according to any one of claims 1 to 8 by executing the computer instructions.