Method, device and equipment for detecting content of ferulic acid in ligusticum wallichii and storage medium

Through the decision tree and random forest model screening key wavelengths, combined with the Transformer model to capture long-term dependence, the problem of variability in the ferulic acid content detection results of Chuanxiong extract was solved, and efficient and stable detection effect was achieved.

CN120408392APending Publication Date: 2025-08-01LONGSHUNRONG PHARMA FACTORY TIANJIN ZHONGXIN PHARMA GRP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503391.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to stably detect the ferulic acid content in Chuanxiong extract over a long span, resulting in high variability in the detection results and unable to meet the quality control requirements.

Method used

The decision tree model is used to screen key wavelengths, the random forest model is used to enhance features, and the long-term dependencies are captured by the Transformer model, and nonlinear mapping is performed through the multi-head attention mechanism to generate ferulic acid content prediction.

Benefits of technology

It improves the accuracy and stability of ferulic acid content detection, adapts to the variability between Chuanxiong extract, and improves the reliability and consistency of the test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408392A_ABST
    Figure CN120408392A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method, device and equipment for detecting the content of ferulic acid in ligusticum wallichii and a storage medium, and the method comprises the following steps: screening a key Top-N wavelength of near-infrared hyperspectral data by using a decision tree, and constructing the decision tree again to predict the content of ferulic acid and generate a label; forming a key wavelength sample set with near-infrared hyperspectral data, performing feature enhancement by using a random forest, and fusing with time sequence features to form joint features; and the ferulic acid content is predicted by utilizing Transform. The decision tree is suitable for extracting key wavelengths from high-dimensional spectral data, the random forest enhances features through Bootstrap sampling and random features, Transform captures a long-distance dependency relationship of time sequence data by using multi-head self-attention, multi-component coexistence and large variability of a ligusticum wallichii extracting solution are better adapted through cross-modal complementation, the detection precision of the content of ferulic acid is improved, and the detection accuracy of the content of ferulic acid in the ligusticum wallichii extracting solution is improved. And the interpretability, the stability and the consistency of the data are kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of detection of traditional Chinese medicine components, and particularly to a method, device, equipment and storage medium for detecting the content of ferulic acid in Ligusticum chuanxiong Hort. Background Art

[0002] With the development of the modern traditional Chinese medicine industry and the expansion of the international market, the quality control requirements for Ligusticum chuanxiong Hort. extracts and their preparations are increasing day by day. The components of Ligusticum chuanxiong Hort. extracts are complex, containing a variety of active substances and interfering substances with large differences in structure and properties. Ferulic acid is one of the main active components of Ligusticum chuanxiong Hort., which has effects such as anti-radiation, antioxidant, antibacterial and antiviral. However, the content stability of ferulic acid in Ligusticum chuanxiong Hort. is often affected by factors such as diverse origin and different processing techniques. Traditional methods for detecting ferulic acid content are time-consuming and destructive. Modern technologies such as near-infrared spectroscopy and machine learning for rapid and non-destructive detection are the current trend.

[0003] The composition of Ligusticum chuanxiong Hort. extract can be detected by near-infrared spectroscopy technology to further detect the content of ferulic acid. Using traditional multivariate calibration models such as partial least squares (PLS) and principal component regression (PCR) can reduce the influence of spectral overlap among various components in Ligusticum chuanxiong Hort. extract. However, due to the large number of component types in Ligusticum chuanxiong Hort. extract, it is difficult for traditional models to capture the internal relationship of spectral data well when dealing with complex spectral data. Especially when the variability between different batches of Ligusticum chuanxiong Hort. extracts during the growth cycle is large, it is difficult to capture the dependence of the change in ferulic acid content over a long time span, resulting in difficulty in ensuring the stability and reliability of the detection results and not meeting the quality control requirements for Ligusticum chuanxiong Hort. extracts and their preparations. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device, equipment and storage medium for detecting the content of ferulic acid in Ligusticum chuanxiong Hort. to solve the technical problem that the variability of detection results is large due to the long time span between Ligusticum chuanxiong Hort. extracts.

[0005] In a first aspect, embodiments of the present invention provide a method for detecting the content of ferulic acid in Ligusticum chuanxiong Hort., including:

[0006] Using the trained first decision tree model to screen key wavelengths from the obtained near-infrared hyperspectral data to obtain key Top-N wavelengths;

[0007] Using the trained second decision tree model to predict the ferulic acid content according to the key Top-N wavelengths to obtain a ferulic acid content label;

[0008] Matching the ferulic acid content label with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and using the trained random forest model to perform feature enhancement on the key wavelength sample set to obtain an enhanced wavelength feature matrix;

[0009] Based on the temporal features of the enhanced wavelength feature matrix and the near-infrared hyperspectral data, a joint feature is formed. The trained Transformer model captures the global dependencies of the enhanced wavelength feature matrix over a long time series based on the multi-head attention mechanism, and performs a non-linear mapping through a fully connected layer to generate the predicted ferulic acid content.

[0010] Further, the step of forming a joint feature based on the temporal features of the enhanced wavelength feature matrix and the near-infrared hyperspectral data, using the trained Transformer model to capture the global dependencies of the enhanced wavelength feature matrix over a long time series based on the multi-head attention mechanism, and performing a non-linear mapping through a fully connected layer to generate the predicted ferulic acid content includes:

[0011] According to the correspondence between the enhanced wavelength feature matrix and the near-infrared hyperspectral data, the temporal features of the near-infrared hyperspectral data are fused with the corresponding enhanced wavelength feature matrix to generate a joint feature matrix;

[0012] The trained Transformer model uses the multi-head attention mechanism. Each attention head generates queries, keys, and values respectively according to the joint feature matrix, calculates the attention scores of each attention head, and concatenates the outputs of all attention heads to generate a multi-head attention feature matrix;

[0013] The fully connected layer is used to perform a linear transformation on the multi-head attention feature matrix, and then a non-linear transformation is performed through an activation function to generate a ferulic acid content feature vector. The predicted ferulic acid content is determined according to the ferulic acid content feature vector.

[0014] Further, the step of matching the ferulic acid content labels with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and using the trained random forest model to perform feature enhancement on the key wavelength sample set to obtain an enhanced wavelength feature matrix includes:

[0015] Perform data preprocessing on the key wavelength sample set to generate a preprocessed sample set;

[0016] When the trained random forest model splits nodes in each decision tree according to the preprocessed sample set, it randomly selects some candidate features to form a feature subset, and then each decision tree continues to split nodes according to the feature subset to generate random features of each decision tree;

[0017] Integrate the random features output by all decision trees to generate an enhanced wavelength feature matrix.

[0018] Further, the step of using the trained second decision tree model to predict the ferulic acid content according to the key Top-N wavelengths to obtain the ferulic acid content labels includes:

[0019] Using the trained second decision tree model, node splitting is performed according to the key Top-N wavelengths, and when splitting the nodes, splitting is performed according to the feature with the largest information gain value to generate child nodes;

[0020] Generate the ferulic acid content label according to the mean value of all leaf nodes of the second decision tree model.

[0021] Further, the using the trained first decision tree model to screen the key wavelengths from the acquired near-infrared hyperspectral data to obtain the key Top-N wavelengths includes:

[0022] Performing near-infrared hyperspectral acquisition on the ligusticum wallichii extract to obtain the original near-infrared hyperspectral data and form an original sample set;

[0023] Using the trained first decision tree model to calculate the information gain of the original sample set, and sorting according to the information gain, and selecting the top N wavelengths with the largest information gain values to generate the key Top-N wavelengths.

[0024] Further, the method further includes: using a sliding window to extract time series features from the near-infrared hyperspectral data to obtain the time series features of the near-infrared hyperspectral data.

[0025] Further, the generating a joint feature matrix by fusing the time series features of the near-infrared hyperspectral data with the corresponding enhanced wavelength feature matrix according to the correspondence between the enhanced wavelength feature matrix and the near-infrared hyperspectral data includes:

[0026] Performing feature splicing on the time series features of the near-infrared hyperspectral data and the enhanced wavelength feature matrix along the channel dimension to generate a joint feature matrix.

[0027] In a second aspect, an embodiment of the present invention provides a device for detecting the ferulic acid content in ligusticum wallichii, including:

[0028] A key wavelength screening module, configured to screen the key Top-N wavelengths in the near-infrared hyperspectrum using a decision tree model;

[0029] A ferulic acid content label generation module, configured to predict the ferulic acid content using a decision tree model according to the key Top-N wavelengths and generate a label;

[0030] A feature enhancement module, configured to enhance the features of the key wavelength samples using a random forest model;

[0031] A ferulic acid content prediction module, configured to predict the ferulic acid content according to the joint features fused with the time series features.

[0032] In a third aspect, an embodiment of the present invention provides a device, including:

[0033] One or more processors;

[0034] A storage device for storing one or more programs,

[0035] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for detecting the ferulic acid content in the above-mentioned Chuanxiong.

[0036] In a fourth aspect, an embodiment of the present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for detecting the ferulic acid content in the above-mentioned Chuanxiong when executed by a computer processor.

[0037] A method, device, equipment and storage medium for detecting the ferulic acid content in Chuanxiong provided by an embodiment of the present invention. The method screens key wavelength features through a decision tree algorithm, enhances features through a random forest, and then captures long-term dependencies through a Transformer. The decision tree algorithm is suitable for capturing the interactions between non-linear features in spectral data and extracting key wavelengths from high-dimensional spectral data, which is particularly effective for the extracted solution of Chuanxiong with coexisting multiple components; the random forest obtains enhanced features through its own Bootstrap sampling and random feature selection, improves the noise robustness of the model, and reduces the risk of overfitting; the Transformer uses a multi-head self-attention mechanism to capture the long-range dependencies of time-series data. The features output by the random forest are fused with the dynamic characteristics of time-series features, and through cross-modal complementarity, data with both high-dimensional non-linear features and time-series dependencies is processed efficiently, making up for the deficiencies of traditional models in modeling time-series patterns. Compared with the method of a single model, the data utilization efficiency, interpretability and the ability to mine complex time-series patterns are improved, and it has better adaptability in generalization performance, noise resistance and complex scenarios. Description of the Drawings

[0038] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0039] [[ID=SS]] Figure 1 Is a flowchart of a method for detecting the ferulic acid content in Chuanxiong according to Embodiment 1 of the present invention;

[0040] Figure 2 Is a flowchart of a method for detecting the ferulic acid content in Chuanxiong according to Embodiment 2 of the present invention;

[0041] Figure 3 Is a flowchart of a method for detecting the ferulic acid content in Chuanxiong according to Embodiment 3 of the present invention;

[0042] Figure 4 It is a flowchart of a method for detecting the content of ferulic acid in Ligusticum chuanxiong described in Embodiment 4 of the present invention;

[0043] Figure 5 It is a schematic structural diagram of a device for detecting the content of ferulic acid in Ligusticum chuanxiong described in Embodiment 5 of the present invention;

[0044] Figure 6 It is a structural diagram of the device described in Embodiment 6 of the present invention. Detailed implementation manners

[0045] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0046] Embodiment 1

[0047] Figure 1 It is a flowchart of a method for detecting the content of ferulic acid in Ligusticum chuanxiong described in Embodiment 1 of the present invention. In this embodiment, the content of ferulic acid is predicted by using a Transformer model to model the joint features fused with temporal features. The specific steps are as follows:

[0048] S101, Use the trained first decision tree model to screen the key wavelengths from the acquired near-infrared hyperspectral data to obtain the key Top-N wavelengths.

[0049] After collecting the near-infrared hyperspectral data of the Ligusticum chuanxiong extract, use the first decision tree model to calculate the information gain of the near-infrared hyperspectral data. According to the sorting of the information gain, select the top N wavelengths with the greatest contribution to the detection of the ferulic acid content to form the key Top-N wavelengths.

[0050] S102, Use the trained second decision tree model to predict the ferulic acid content according to the key Top-N wavelengths to obtain the ferulic acid content label.

[0051] Use the screened key Top-N wavelengths to construct a new decision tree again to form the second decision tree model. The second decision tree model selects the optimal feature with the greatest information gain among the Top-N wavelengths when splitting nodes. Finally, when the stopping condition of the second decision tree model is reached, use the mean value of all leaf nodes as the output of the second decision tree model to obtain the predicted value of the ferulic acid content and form the ferulic acid content label for matching with the acquired near-infrared hyperspectral data.

[0052] S103. Match the ferulic acid content label with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and use the trained random forest model to enhance the features of the key wavelength sample set to obtain an enhanced wavelength feature matrix.

[0053] Use the ferulic acid label to match with the near-infrared hyperspectral data to form a new key wavelength sample set. The key wavelength sample set selects the wavelength features with higher contribution to the detection of ferulic acid concentration, and at the same time removes the wavelength features with smaller contribution and irrelevant ones, reducing the computational amount and model complexity, while reducing the influence of irrelevant factors and improving the model performance. Use the random forest model to enhance the features of the key wavelength sample set. Due to the characteristics of the random forest itself, when splitting the nodes of each decision tree in the random forest, it will randomly select some features to form a feature subset, and then continue to split according to the optimal feature in the feature subset. Integrate the outputs of all decision trees to form the final output of the random forest, that is, the enhanced wavelength feature matrix. Through the randomness of feature selection, the robustness of the model is improved, which can make up for the anti-overfitting of the decision tree algorithm, has good high-dimensional data processing ability, and improves the expression ability of the wavelengths with greater contribution to the detection of ferulic acid.

[0054] S104. Form joint features according to the temporal features of the enhanced wavelength feature matrix and the near-infrared hyperspectral data, and use the trained Transformer model to capture the global dependence of the enhanced wavelength feature matrix in the long time series based on the multi-head attention mechanism, and perform non-linear mapping through the fully connected layer to generate the predicted ferulic acid content.

[0055] When collecting near-infrared hyperspectral data, temporal features will be formed according to the collection time. The temporal features can reflect the differences between different batches of Ligusticum chuanxiong extract and the variability of the same batch of Ligusticum chuanxiong extract over time. There will be differences in the near-infrared hyperspectral data collected at different batches and at different times of the same batch. Match the temporal features with the enhanced wavelength feature matrix of the corresponding samples to form joint features that integrate the temporal features and the enhanced wavelength features. Use the characteristic that the Transformer model can capture the global dependence between features in the long time series to improve the accuracy of ferulic acid content prediction. The multi-head self-attention mechanism of the Transformer model maps the input into queries (Q), keys (K), and values (V) through each self-attention head, calculates the scores and weights of each attention head respectively, and then concatenates the outputs of all attention heads to form the output of the multi-head attention, which can assign weights to different features. Then, in the fully connected layer, first perform a linear mapping on the output of the multi-head attention, and then perform a non-linear mapping through the activation function to map the high-dimensional data, and finally output a vector representing the predicted ferulic acid content, thereby determining the ferulic acid content corresponding to the collected sample.

[0056] In this embodiment, the decision tree algorithm is used to screen the key wavelength features, the random forest is used for feature enhancement, and then the Transformer is used to capture the dependencies in the long time series. The decision tree algorithm is suitable for capturing the interactions between non-linear features in spectral data and extracting key wavelengths from high-dimensional spectral data, which is particularly effective for the Chuanxiong extract with coexisting multiple components; the random forest obtains enhanced features through its own Bootstrap sampling and random feature selection, improves the noise robustness of the model, and reduces the risk of overfitting; the Transformer uses the multi-head self-attention mechanism to capture the long-distance dependencies of time series data. The features output by the random forest are fused with the dynamic characteristics of the time series features, and through cross-modal complementarity, data with both high-dimensional non-linear features and time series dependencies are processed efficiently, making up for the deficiencies of traditional models in modeling time series patterns. Compared with the method of a single model, the data utilization efficiency, interpretability, and the ability to mine complex time series patterns are improved, and it has better adaptability in generalization performance, noise resistance, and complex scenarios.

[0057] Embodiment 2[[ID=E5]]

[0058] Figure 2 FIG. is a flowchart of a method for detecting the content of ferulic acid in Chuanxiong according to Embodiment 2 of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, a joint feature will be formed according to the enhanced wavelength feature matrix and the time series features of the near-infrared hyperspectral data. The trained Transformer model uses the multi-head attention mechanism to capture the global dependencies of the enhanced wavelength feature matrix in the long time series, and performs non-linear mapping through a fully connected layer to generate the predicted ferulic acid content. The specific optimization is as follows:

[0059] According to the correspondence between the enhanced wavelength feature matrix and the near-infrared hyperspectral data, the time series features of the near-infrared hyperspectral data are fused with the corresponding enhanced wavelength feature matrix to generate a joint feature matrix;

[0060] The trained Transformer model uses the multi-head attention mechanism. Each attention head generates queries, keys, and values respectively according to the joint feature matrix, calculates the attention scores of each attention head, and splices the outputs of all attention heads to generate a multi-head attention feature matrix;

[0061] The fully connected layer is used to perform a linear transformation on the multi-head attention feature matrix, and then a non-linear transformation is performed through an activation function to generate a ferulic acid content feature vector, and the predicted ferulic acid content is determined according to the ferulic acid content feature vector.

[0062] Correspondingly, the method for detecting the content of ferulic acid in Chuanxiong provided in this embodiment specifically includes:

[0063] S201, Use the trained first decision tree model to screen the obtained near-infrared hyperspectral data to obtain the key Top-N wavelengths.

[0064] S202, Use the trained second decision tree model to predict the ferulic acid content based on the key Top-N wavelengths to obtain the ferulic acid content labels.

[0065] S203, Match the ferulic acid content labels with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and use the trained random forest model to enhance the features of the key wavelength sample set to obtain an enhanced wavelength feature matrix.

[0066] S204, According to the corresponding relationship between the enhanced wavelength feature matrix and the near-infrared hyperspectral data, fuse the temporal features of the near-infrared hyperspectral data with the corresponding enhanced wavelength feature matrix to generate a joint feature matrix.

[0067] In order to utilize the characteristic that the Transformer model can flexibly process temporal data, capture the correlation between the time characteristics and spectral wavelength characteristics of the near-infrared hyperspectral data collected from the Ligusticum chuanxiong extract, and improve the adaptability to the time difference in the collection of near-infrared hyperspectral data, the differences between batches of Ligusticum chuanxiong extract, and the variability of the same batch of Ligusticum chuanxiong extract over time, the enhanced wavelength feature matrix after random forest enhancement needs to be feature-fused with the temporal features. The specific method can be to perform feature splicing on the enhanced wavelength feature matrix and the temporal features. The generated joint features include the wavelength features of the spectral data and the relevant time features, which can better capture the global dependence between wavelengths in long time series through the Transformer model, improve the prediction accuracy of dynamic detection data, and achieve temporal adaptability.

[0068] S205, The trained Transformer model, through the multi-head attention mechanism, each attention head generates queries, keys, and values respectively according to the joint feature matrix, calculates the attention scores of each attention head, and splices the outputs of all attention heads to generate a multi-head attention feature matrix.

[0069] The multi-head attention mechanism is to pass the input data through multiple self-attention heads. Each attention head only focuses on a part of the input data. By mapping the input of each attention head to queries Q, keys K, and values V, calculating the attention scores and weight matrices of each attention head, and then weighting different features of the output data with different weights. Exemplarily, the matrix representations of queries Q, keys K, and values V are:

[0070] Q = XW Q

[0071] K = XW K

[0072] V = XW V

[0073]

[0074] Wherein, X represents the combined feature matrix in the multi-head attention mechanism, and W Q , W K , W V respectively represent the weight matrices of the query Q, key K, and value V mapped by X, R represents the set of real numbers (the element values are real numbers), and d k represents the dimension of the key vector.

[0075] The weight calculation formula for the attention head is:

[0076]

[0077] Wherein, K ′ represents the transpose matrix of the key matrix K, d k = 64 is the dimension of the key vector. To avoid gradient saturation, 8 groups of attention heads are calculated in parallel, and then the outputs of all attention heads are concatenated to obtain the output of the multi-head attention mechanism, that is, the multi-head attention feature matrix. When facing an environment with limited computing resources, channel pruning (removing channels with an absolute weight value < 0.01) and 8-bit integer quantization can also be used to reduce the data volume and the amount of calculation, and the volume of the model can be compressed to 25% of the original size.

[0078] S206. Use the fully connected layer to perform a linear transformation on the multi-head attention feature matrix, and then perform a non-linear transformation through the activation function to generate the ferulic acid content feature vector, and determine the predicted ferulic acid content according to the ferulic acid content feature vector.

[0079] After the multi-head attention weights the features, the output multi-head attention feature matrix is input into the fully connected layer. In the fully connected layer, first perform a linear mapping, and then perform a non-linear mapping through the activation function ReLU, which is expressed as:

[0080] h = ReLU(W fc1 ·F fused + b fc1 )

[0081] Wherein, W fc1 is the multi-head attention weight, F fused is the combined feature, b fc1 is the bias term, and is, and d h is the hidden layer dimension, T represents the sequence length, that is, the number of spectral samples collected, and d t represents the time step, indicating the time interval between adjacent collection points.

[0082] The structure of multiple fully connected layers can be constructed according to the dimension of the hidden layer for regression prediction. Finally, the output of the intermediate hidden layer is mapped to the target space to generate the predicted ferulic acid content feature vector. The predicted ferulic acid content, which is a scalar, can be obtained through the ferulic acid content feature vector and the sampling time of the corresponding sample. The formula is as follows:

[0083]

[0084] Among them, W fc2 and b fc2 are the weight matrix and bias term of the fully connected layer. Through backpropagation optimization, the output dimension is fixed at 1, corresponding to 1 target variable, i.e., the ferulic acid content. R represents the set of real numbers, and d h represents the output dimension of the intermediate hidden layer. During the training process of the model, the model can also be optimized through a loss function. For example, the mean squared error (MSE) loss function is expressed as:

[0085]

[0086] Among them, m represents the number of training samples, represents the predicted value of the i-th sample, y i represents the true value of the i-th sample. By minimizing the loss function L, the parameters W fc2 and b fc2 of the fully connected layer and the weights of the previous network are updated through backpropagation.

[0087] In this embodiment, the enhanced features output by the random forest model are fused with the temporal features in the near-infrared hyperspectral data to obtain a joint feature matrix. The characteristics of the Transformer model are used to capture the variation law of features in the time dimension. The Transformer captures the long-range dependence relationship of temporal data through the multi-head self-attention mechanism, deeply models the long-range dependence relationship of temporal data such as the growth cycle and environmental changes of Ligusticum chuanxiong Hort., captures the dynamic interaction mode between spectral features and ferulic acid content at different time points, can improve the ability to mine the temporal relationship of complex data, break through the local modeling limitation of traditional models for temporal dynamics, enable the model to adapt to the large variability between Ligusticum chuanxiong Hort. extracts, and improve the detection accuracy when detecting the ferulic acid content, while maintaining the interpretability, stability, and consistency of the data.

[0088] An optional implementation manner of this embodiment is to use a sliding window to extract the time series features of the near-infrared hyperspectral data to obtain the temporal features of the near-infrared hyperspectral data.

[0089] During the acquisition of near-infrared hyperspectral data, temporal data will be generated according to the acquisition time, sampling batch, etc. The time series features can be extracted through a sliding window: Among them, dt is the time step.

[0090] Specifically, according to the correspondence between the enhanced wavelength feature matrix and the near-infrared hyperspectral data, the temporal features of the near-infrared hyperspectral data are fused with the corresponding enhanced wavelength feature matrix to generate a joint feature matrix, including:

[0091] The temporal features of the near-infrared hyperspectral data are concatenated with the enhanced wavelength feature matrix along the channel dimension to generate a joint feature matrix.

[0092] After the key wavelength sample set is feature-enhanced by the random forest model, the enhanced wavelength feature matrix F is obtained RF . The sliding window is used to extract the time series features of the near-infrared hyperspectral data to obtain the temporal features F time .

[0093] The enhanced wavelength feature matrix output by the random forest and the temporal features are concatenated along the channel dimension to generate a joint feature, denoted as:

[0094] F fused = Concat(F RF , F time )

[0095] where F fused represents the joint feature, and R represents the set of real numbers, T represents the channel dimension size of the enhanced wavelength feature matrix F RF , d t represents the channel dimension size of the temporal feature F time , F RF represents the enhanced wavelength feature matrix, F time represents the temporal feature, and Concat represents feature concatenation.

[0096] The near-infrared hyperspectral data of the Ligusticum chuanxiong extract contains rich spectral information, and the temporal features can reflect the differences between different batches and the variability over time within the same batch. In the dimensional representation of the data, the channel dimension is an important dimension direction for feature combination, and it is a reasonable choice to fuse the spectral features and the temporal features. By concatenating the two, the spectral features and the time features can be combined, enabling the model to simultaneously learn the features related to the ferulic acid content in the spectral data and the variation law of these features over time, providing more comprehensive information for subsequent prediction. Through this cross-modal complementary method, data with both high-dimensional non-linear features (spectral features) and temporal dependencies (time features) can be efficiently processed. Compared with single-feature input, the joint feature can enable the model to better capture the complex relationships in the data and improve the accuracy of predicting the ferulic acid content.

[0097] Example Three

[0098] Figure 3 This is a flowchart of a method for detecting the content of ferulic acid in Ligusticum chuanxiong described in Embodiment 3 of the present invention. This embodiment is optimized based on the above-mentioned embodiment. In this embodiment, the ferulic acid content label is matched with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and the trained random forest model is used to enhance the features of the key wavelength sample set to obtain an enhanced wavelength feature matrix. The specific optimization is as follows:

[0099] Perform data preprocessing on the key wavelength sample set to generate a preprocessed sample set;

[0100] When the trained random forest model splits the nodes of each decision tree according to the preprocessed sample set, randomly select some candidate features to form a feature subset, and then each decision tree continues to split the nodes according to the feature subset to generate the random features of each decision tree;

[0101] Integrate the random features output by all decision trees to generate an enhanced wavelength feature matrix.

[0102] Correspondingly, the method for detecting the content of ferulic acid in Ligusticum chuanxiong provided in this embodiment specifically includes:

[0103] S301, Use the trained first decision tree model to screen the key wavelengths from the obtained near-infrared hyperspectral data to obtain the key Top-N wavelengths.

[0104] S302, Use the trained second decision tree model to predict the ferulic acid content according to the key Top-N wavelengths to obtain the ferulic acid content label.

[0105] S303, Perform data preprocessing on the key wavelength sample set to generate a preprocessed sample set.

[0106] The input key wavelength sample set is expressed as:

[0107] D = {(x1, y1), (x2, y2), …, (x m , y m )}

[0108] Among them, x m is the near-infrared hyperspectral feature, reflecting the wavelengths with higher contribution, and y m is the ferulic acid content label.

[0109] Through normalization processing, the feature values are mapped to a preprocessed sample set with a mean of 0 and a variance of 1 distribution, which is convenient for enhancing the convergence stability of the model during subsequent calculations.

[0110] In S304, when the trained random forest model performs node splitting in each decision tree according to the preprocessed sample set, it randomly selects some candidate features to form a feature subset. Then, each decision tree continues to perform node splitting according to the feature subset to generate the random features of each decision tree.

[0111] The decision trees in the random forest model are represented as:

[0112] {G1(x), G2(x), …, G T (x)}

[0113] For each decision tree G T (x), when performing node splitting, n sub candidate features are randomly selected from all features, and n sub < n, to form a new feature subset A sub , which is represented as:

[0114] A sub =RandomSelect(A, n sub )

[0115] Where

[0116] The characteristic of the random forest is node classification optimization. Calculate the information gain on the feature subset A sub , select the optimal feature for node splitting, and form a new decision tree. The calculation formula of information gain is:

[0117]

[0118] Where a is the candidate feature and D v is the subset after node splitting.

[0119] In S305, integrate the random features output by all decision trees to generate an enhanced wavelength feature matrix.

[0120] The final output of the random forest is obtained by integrating the outputs of each decision tree. Integrate the feature vectors output by all decision trees through average or majority voting to form an enhanced wavelength feature matrix F RF ∈R T to enhance the expression ability of the wavelengths with greater contribution in the detection process of ferulic acid content.

[0121] In S306, form joint features according to the enhanced wavelength feature matrix and the temporal features of the near-infrared hyperspectral data. Use the trained Transformer model based on the multi-head attention mechanism to capture the global dependence of the enhanced wavelength feature matrix in the long time series, and perform non-linear mapping through the fully connected layer to generate the predicted ferulic acid content.

[0122] In this embodiment, the random forest is used to enhance the features of the wavelengths with greater contribution. By leveraging the randomness of feature subset selection and the ensemble strategy of the random forest, the robustness of small samples is improved. It can accurately screen out the combination of non-linear key features from the high-dimensional spectral data of the Chuanxiong extract, suppress noise interference, reduce the interference of redundant features on the model, and reduce the data dimension, enabling the model to be deployed in resource-constrained environments such as embedded devices, and improving the training efficiency and generalization ability of the model. Instead of directly outputting the prediction result, it forms a complementary relationship with the decision tree algorithm, compensates for the relatively simple structure of the decision tree algorithm, reduces the risk of overfitting, and enhances the expression ability of the wavelengths with greater contribution in the process of ferulic acid content detection, thereby improving the accuracy of ferulic acid content detection.

[0123] Embodiment Four

[0124] Figure 4 FIG. is a flowchart of a method for detecting the ferulic acid content in Chuanxiong according to Embodiment Four of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, the trained second decision tree model will be used to predict the ferulic acid content according to the key Top-N wavelengths to obtain the ferulic acid content label. The specific optimization is as follows:

[0125] Using the trained second decision tree model, perform node splitting according to the key Top-N wavelengths, and when splitting the nodes, split according to the feature with the largest information gain value to generate child nodes;

[0126] Generate the ferulic acid content label according to the mean value of all leaf nodes of the second decision tree model.

[0127] Correspondingly, the method for detecting the ferulic acid content in Chuanxiong provided in this embodiment specifically includes:

[0128] S401, using the trained first decision tree model to screen the key wavelengths from the acquired near-infrared hyperspectral data to obtain the key Top-N wavelengths.

[0129] S402, using the trained second decision tree model, perform node splitting according to the key Top-N wavelengths, and when splitting the nodes, split according to the feature with the largest information gain value to generate child nodes.

[0130] Using the optimal feature with the largest information gain among the key Top-N wavelengths as the root node, divide the sample set D into v subsets {D 1 , D 2 , ……, D v}, where each subset corresponds to a specific value branch of the feature a. The branch weight is determined by the proportion of subset samples:

[0131]

[0132] When splitting nodes in the second decision tree model, at each child node, the new feature with the largest information gain in the current subset is selected for classification to generate the next-level child nodes until a preset condition is reached, such as the node purity meeting the threshold or the decision tree reaching the preset depth.

[0133] S403. Generate the ferulic acid content label according to the mean values of all leaf nodes of the second decision tree model.

[0134] In the second decision tree model, the mean value of the ferulic acid concentration of all samples in each leaf node is stored. The mean values of all leaf nodes are used as the final output of the second decision tree model, that is, the content of ferulic acid in the current sample (unit: mg / g). Taking the ferulic acid content as a label, a new sample set can be formed with the original spectral samples for subsequent processing. The decision tree model can also be optimized by pruning strategies, such as pre-pruning or post-pruning, to remove redundant branches of the decision tree, reduce the data volume, and lower the model complexity.

[0135] S404. Match the ferulic acid content label with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and use the trained random forest model to enhance the features of the key wavelength sample set to obtain an enhanced wavelength feature matrix.

[0136] S405. Form joint features according to the enhanced wavelength feature matrix and the temporal features of the near-infrared hyperspectral data. Use the trained Transformer model to capture the global dependence of the enhanced wavelength feature matrix in the long time series based on the multi-head attention mechanism, and perform non-linear mapping through the fully connected layer to generate the predicted ferulic acid content.

[0137] In this embodiment, the first decision tree model is used to screen key features. Using the screened key features, the second decision tree model is used again to select the optimal features to generate a tree structure for the preliminary prediction of the ferulic acid content, form the ferulic acid content label, and mark the original sample set for subsequent calculations. The decision tree algorithm can evaluate the importance of different features according to the information gain, select important features for node splitting, automatically select the most important features, which helps to reduce the data dimension and improve the prediction accuracy and efficiency of the model.

[0138] Optionally, the step of using the trained first decision tree model to screen key wavelengths from the acquired near-infrared hyperspectral data to obtain the key Top-N wavelengths includes:

[0139] Collect near-infrared hyperspectral data of the chuanxiong extract to obtain the original near-infrared hyperspectral data and form an original sample set.

[0140] Sampling of the Ligusticum chuanxiong extract was carried out using a near-infrared hyperspectral imager. Exemplarily, the spectral range of the near-infrared hyperspectrum is 700 - 2500 nm, and the resolution is better than 10 nm. When collecting, a sample stage with adjustable height can be designed and equipped with a fixed fixture to ensure that the sample is evenly illuminated and does not block the spectral collection area. The original sample set D is constituted by using the obtained spectral reflectance information. The sample set D contains wavelength information and time information in the full band, and the time information is the collection time of the spectrum.

[0141] The information gain of the original sample set was calculated using the trained first decision tree model, and sorted according to the information gain. The top N wavelengths with the largest information gain values were selected to generate the key Top-N wavelengths.

[0142] With its own non-linear modeling and feature screening capabilities, the first decision tree model can extract key feature combinations from high-dimensional spectral data. Using the decision tree algorithm to calculate the information gain of the full-band spectrum of near-infrared hyperspectral data, it can evaluate the contribution degree of each wavelength in the spectral data to the detection of ferulic acid according to the change in the dispersion degree of the ferulic acid content distribution represented by the wavelength, and screen out the key Top-N wavelengths according to the contribution degree ranking. Screening the spectral data and selecting more important spectral features for subsequent calculations reduces the data volume and computational complexity, reduces the influence of irrelevant factors, and can improve the accuracy of ferulic acid content detection. Exemplarily, for each discrete wavelength feature a (regarded as a candidate feature), its information gain for the sample set D is calculated, and the formula is as follows:

[0143] Gain(D,a)=Ent(D)-Ent(D∣a)

[0144] Among them, the information entropy is expressed as:

[0145]

[0146] Among them, p k is the proportion of the k-th category in the sample.

[0147] The conditional entropy is expressed as:

[0148]

[0149] Among them, D v is the subset when the feature a takes the value of a v .

[0150] Sort according to the information gain of each wavelength, and select the top N wavelengths with the highest contribution degree to the detection of ferulic acid, which are determined as the key Top-N wavelengths.

[0151] It should be noted that after the spectral data is collected, the data still needs to be preprocessed to achieve data standardization so that it can be applied to subsequent algorithm processing. At the same time, during the training process, the data set can be divided into a training set and a test set by using random sampling to ensure the representativeness of the training set and the test set. The decision tree regression algorithm can also be used to select the optimal parameters through cross-validation, such as determining the depth of the decision tree, the minimum number of samples in the leaf nodes, etc., to improve the performance of the decision tree model.

[0152] Embodiment 5

[0153] Figure 5 FIG. is a schematic structural diagram of a device for detecting the ferulic acid content in Ligusticum chuanxiong according to Embodiment 5 of the present invention. In this embodiment, the device for detecting the ferulic acid content in Ligusticum chuanxiong includes:

[0154] A key wavelength screening module 810, configured to screen the key Top-N wavelengths in the near-infrared hyperspectrum by using a first decision tree model;

[0155] A ferulic acid content label generation module 820, configured to predict the ferulic acid content by using a second decision tree model according to the key Top-N wavelengths and generate labels;

[0156] A feature enhancement module 830, configured to enhance the features of the key wavelength samples by using a random forest model;

[0157] A ferulic acid content prediction module 840, configured to predict the ferulic acid content according to the combined features fused with time series features.

[0158] In this embodiment, the key Top-N wavelengths in the near-infrared hyperspectrum are screened by the key wavelength screening module. The ferulic acid content prediction label is generated by the ferulic acid content label generation module based on the Top-N wavelengths. The key wavelength samples are feature-enhanced by the feature enhancement module. The ferulic acid content is predicted by the ferulic acid content prediction module by fusing the temporal features and based on the fused joint features. The decision tree algorithm is used to screen the key wavelength features, the random forest is used for feature enhancement, and then the Transformer is used to capture the dependencies in the long time series. The decision tree algorithm is suitable for capturing the interactions between non-linear features in spectral data and extracting key wavelengths from high-dimensional spectral data, which is particularly effective for the Ligusticum chuanxiong extract with co-existing multiple components. The random forest obtains enhanced features through its own Bootstrap sampling and random feature selection, improving the noise robustness of the model and reducing the risk of overfitting. The Transformer uses the multi-head self-attention mechanism to capture the long-range dependencies in the time series data. The features output by the random forest are fused with the dynamic characteristics of the temporal features, and through cross-modal complementarity, the data with both high-dimensional non-linear features and temporal dependencies is efficiently processed, making up for the deficiencies of traditional models in modeling temporal patterns. Compared with the single-model method, the data utilization efficiency, interpretability, and the ability to mine complex temporal patterns are improved, and it has better adaptability in generalization performance, noise resistance, and complex scenarios.

[0159] The ferulic acid content detection device in Ligusticum chuanxiong provided by the embodiment of the present invention can execute the ferulic acid content detection method in Ligusticum chuanxiong provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0160] Embodiment Six

[0161] Figure 6 It is a structural diagram of a device according to Embodiment Six of the present invention. Figure 6 It shows a block diagram of an exemplary device 12 suitable for implementing the embodiments of the present invention. Figure 6 The shown device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0162] As Figure 6 shown, the device 12 is presented in the form of a general-purpose computing device. The components of the device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0163] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor bus, or a local bus using any of a variety of bus architectures. By way of example, and not limitation, these architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MAC) buses, Enhanced ISA buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.

[0164] Device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by device 12, including both volatile and nonvolatile media, removable and non-removable media.

[0165] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Device 12 can further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing non-removable, nonvolatile magnetic media ( Figure 6 not shown and typically called a "hard disk drive"). Although Figure 6 not shown in the figures, a disk drive for reading and writing a removable, nonvolatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable, nonvolatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) can be provided. In these instances, each drive can be connected to bus 18 by one or more data media interfaces. Memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the present invention.

[0166] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a networking environment. The program modules 42 typically carry out the functions and / or methods of the embodiments described herein.

[0167] Device 12 may also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and may also communicate with one or more devices that enable a user to interact with the device 12 / server / computer, and / or communicate with any device that enables the device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 22. Also, device 12 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, network adapter 20 communicates with other modules of device 12 through bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0168] The processing unit 16 performs various functional applications and data processing by running programs stored in the system memory 28, for example, implementing the method for detecting the content of ferulic acid in Ligusticum chuanxiong provided by the embodiments of the present invention.

[0169] Embodiment Seven

[0170] Embodiment Seven of the present invention also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for detecting the content of ferulic acid in Ligusticum chuanxiong provided by the above embodiments when executed by a computer processor.

[0171] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0172] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0173] The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including - but not limited to - wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0174] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, it can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0175] Note that the above is only the preferred embodiment of the present invention and the applied technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for detecting the ferulic acid content in Ligusticum chuanxiong, characterized in that, Including: Using the trained first decision tree model to screen the key wavelengths from the acquired near-infrared hyperspectral data, obtaining the key Top-N wavelengths; Using the trained second decision tree model to predict the ferulic acid content based on the key Top-N wavelengths, obtaining the ferulic acid content label; Matching the ferulic acid content label with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and using the trained random forest model to enhance the features of the key wavelength sample set, obtaining the enhanced wavelength feature matrix; Forming joint features based on the enhanced wavelength feature matrix and the temporal features of the near-infrared hyperspectral data, using the trained Transformer model to capture the global dependence of the enhanced wavelength feature matrix in the long time series based on the multi-head attention mechanism, and performing non-linear mapping through the fully connected layer to generate the predicted ferulic acid content.

2. The method according to claim 1, characterized in that, The forming joint features based on the enhanced wavelength feature matrix and the temporal features of the near-infrared hyperspectral data, using the trained Transformer model to capture the global dependence of the enhanced wavelength feature matrix in the long time series based on the multi-head attention mechanism, and performing non-linear mapping through the fully connected layer to generate the predicted ferulic acid content includes: According to the corresponding relationship between the enhanced wavelength feature matrix and the near-infrared hyperspectral data, fusing the temporal features of the near-infrared hyperspectral data with the corresponding enhanced wavelength feature matrix to generate a joint feature matrix; The trained Transformer model, through the multi-head attention mechanism, each attention head generates queries, keys, and values respectively according to the joint feature matrix, calculates the attention scores of each attention head, and splices the outputs of all attention heads to generate a multi-head attention feature matrix; Using the fully connected layer to perform a linear transformation on the multi-head attention feature matrix, and then performing a non-linear transformation through the activation function to generate a ferulic acid content feature vector, and determining the predicted ferulic acid content according to the ferulic acid content feature vector.

3. The method according to claim 1, characterized in that, The matching the ferulic acid content label with the corresponding near-infrared hyperspectral data to form a key wavelength sample set, and using the trained random forest model to enhance the features of the key wavelength sample set to obtain The enhanced wavelength feature matrix includes: Performing data preprocessing on the key wavelength sample set to generate a preprocessed sample set; When the trained random forest model performs node splitting for each decision tree according to the preprocessed sample set, randomly selecting some candidate features to form a feature subset, and then each decision tree continues to perform node splitting according to the feature subset to generate the random features of each decision tree; Integrating the random features output by all decision trees to generate an enhanced wavelength feature matrix.

4. The method according to claim 1, characterized in that, The using the trained second decision tree model to predict the ferulic acid content based on the key Top-N wavelengths, obtaining the ferulic acid content label includes: Using the trained second decision tree model to perform node splitting according to the key Top-N wavelengths, and splitting to generate child nodes according to the feature with the largest information gain value during node splitting; Generating the ferulic acid content label according to the mean value of all leaf nodes of the second decision tree model.

5. The method according to claim 1, wherein Performing key wavelength screening on the acquired near-infrared hyperspectral data by using the trained first decision tree model to obtain the key Top-N wavelengths, including: Performing near-infrared hyperspectral collection on the Chuanxiong extract to obtain the original near-infrared hyperspectral data and forming an original sample set; Calculating the information gain of the original sample set by using the trained first decision tree model, sorting according to the information gain, and selecting the top N wavelengths with the largest information gain values to generate the key Top-N wavelengths.

6. The method according to claim 2, characterized in that, The method further includes: extracting time series features of the near-infrared hyperspectral data by using a sliding window to obtain the time series features of the near-infrared hyperspectral data.

7. The method according to claim 6, characterized in that, According to the corresponding relationship between the enhanced wavelength feature matrix and the near-infrared hyperspectral data, performing feature fusion on the time series features of the near-infrared hyperspectral data and the corresponding enhanced wavelength feature matrix to generate a joint feature matrix, including: Performing feature splicing on the time series features of the near-infrared hyperspectral data and the enhanced wavelength feature matrix along the channel dimension to generate a joint feature matrix.

8. A detection device for the content of ferulic acid in Ligusticum chuanxiong, characterized in that, Including: A key wavelength screening module, configured to screen the key Top-N wavelengths in the near-infrared hyperspectrum by using a decision tree model; A ferulic acid content label generation module, configured to predict the ferulic acid content by using a decision tree model according to the key Top-N wavelengths and generate a label; A feature enhancement module, configured to perform feature enhancement on the key wavelength samples by using a random forest model; A ferulic acid content prediction module, configured to predict the ferulic acid content according to the joint features fused with the time series features.

9. A device, characterized in that, The device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for detecting the ferulic acid content in Chuanxiong as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the method for detecting the ferulic acid content in Chuanxiong as described in any one of claims 1-7 when executed by a computer processor.