Typha angustifolia processing fire judgment method based on multi-source data fusion and lightweight attention mechanism

By integrating near-infrared spectroscopy, electronic nose, and electronic eye data through the SAINT model of multi-source data fusion, the problem of controlling the processing temperature of Typha pollen was solved, achieving efficient and objective temperature judgment, improving accuracy and robustness, and providing technical support for the quality control of charred Chinese medicine.

CN120974272APending Publication Date: 2025-11-18JINAN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511090338.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies for processing cattail pollen have problems such as difficulty in accurately controlling the heat, high subjectivity, high testing costs, low efficiency, and inability to achieve real-time monitoring. In particular, the lack of collaborative analysis of multimodal data results in a single dimension of quality control, making it difficult to comprehensively evaluate the quality of the processing technology.

Method used

The SAINT model, based on multi-source data fusion and a lightweight attention mechanism, integrates near-infrared spectroscopy, electronic nose, and electronic eye data. Through self-attention and inter-sample attention mechanisms, it captures the complex relationships between features and between samples, enabling non-destructive, real-time online monitoring of the processing time of Typha pollen.

Benefits of technology

It enables precise and objective judgment of the processing temperature of cattail pollen, improves the accuracy and robustness of the judgment, and achieves an average classification accuracy of 99.60%. It overcomes the subjective limitations of manual judgment and provides an innovative solution for the standardization and digital quality control of charred Chinese medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974272A_ABST
    Figure CN120974272A_ABST
Patent Text Reader

Abstract

The invention relates to a method for judging the degree of heating of pollen typhae processing, in particular to a method for judging the degree of heating of pollen typhae processing based on multi-source data fusion and a lightweight attention mechanism. In view of the remarkable advantages of automatically extracting deep features and establishing nonlinear association among multi-source information through a cross-modal interaction mechanism in the process of processing high-dimensional heterogeneous data fusion by an emerging deep learning method, the invention provides a pollen typhae processing fire judgment method based on multi-source data fusion and a lightweight attention mechanism. According to the method, near infrared spectrum (chemical components), electronic nose (smell) and electronic eye (color) data are fused for the first time and applied to cattail pollen processing fire judgment; meanwhile, the deep learning algorithm SAINT is subjected to lightweight improvement for the first time, and the purpose is to construct a pollution-free, lossless and rapid pollen typhae charcoal processing quality multi-dimensional intelligent comprehensive judgment model capable of achieving real-time online monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for judging the degree of processing of Typhae Pollen, in particular to a method for judging the degree of processing of Typhae Pollen based on multi-source data fusion and light attention mechanism BACKGROUND

[0002] Typhae Pollen Carbonisatum is a commonly used fried carbon traditional Chinese medicine. Its processing quality is directly related to the clinical efficacy and safety of drug use. The 2025 edition of Chinese Pharmacopoeia stipulates that Typhae Pollen Carbonisatum is a brownish black or blackish brown, fried to brownish black, with brownish black or blackish brown, smoky aroma and slightly bitter and astringent taste. Typhae Pollen Carbonisatum has the effect of stopping bleeding and is used for various blood conditions. Traditional Chinese medicine theory emphasizes that fried carbon should retain its nature, but Typhae Pollen belongs to pollen medicine with loose texture, and its fried carbon degree is difficult to control accurately, and it is easy to appear "not enough" or "too much" phenomenon. Current production mainly relies on the observation of color change to judge the end point of processing, which has the problems of strong subjectivity, poor stability and low efficiency. Although the current quality standard covers color, moisture, ash content, extract and other indicators, and the literature also reports the use of HPLC fingerprint for quality evaluation, these methods all need to rely on instruments and chemical reagents for sample preparation and pretreatment, which is high in cost and time-consuming, and all are quality detection of terminal processing products, which cannot realize real-time monitoring of the processing process. The existing technology such as patent CN111487202A proposes to determine the color parameters (L * ,a * ,b * value) based on a spectrophotometer, and uses color difference value or discriminant function to determine the degree of processing, but its discrimination accuracy and detection efficiency still need to be improved. Another type of technology such as patents CN112414967B and CN113030008B uses near infrared spectroscopy for real-time detection or online quality control, but this kind of method fails to integrate the key sensory attribute data such as appearance color and odor change in the processing process, resulting in the fragmentation of the correlation between spectrum (chemical component representation) and color, odor (sensory attribute), and lacking the ability of collaborative analysis of multi-modal data, which makes the quality control dimension single and difficult to comprehensively evaluate the processing technology. SUMMARY

[0003] OBJECTIVE

[0004] In view of the significant advantages of emerging deep learning methods in handling high-dimensional heterogeneous data fusion, such as automatic extraction of deep features and establishment of nonlinear correlation between multi-source information through cross-modal interaction mechanism, the present application proposes a method for judging the degree of processing of Typhae Pollen based on multi-source data fusion and light attention mechanism. This method first integrates near infrared spectroscopy (chemical components), electronic nose (odor) and electronic eye (color) data, and applies it to the judgment of the degree of processing of Typhae Pollen; at the same time, the deep learning algorithm SAINT is improved for the first time, aiming to build a multi-dimensional intelligent comprehensive judgment model for the processing quality of Typhae Pollen Carbonisatum, which is pollution-free, non-destructive, fast and can realize real-time online monitoring.

[0005] Technical solution

[0006] A processing time discrimination method of pollen based on multi-source data fusion and light attention mechanism, characterized by the following steps:

[0007] (1) Collect and preprocess olfactory information: detect the pollen processing sample by electronic nose; and preprocess the sensor data obtained by electronic nose, that is, use principal component analysis for feature extraction, and optimize the number of principal components;

[0008] (2) Visual information collection: detect the pollen processing sample by electronic eye detector, record the colorimetric values L * , a * , b * ; according to Eab * =(L *2 +a *2 +b *2 ) 1 / 2 (1) Calculate the total colorimetric value Eab * ;

[0009] (3) Near-infrared information collection:

[0010] Use XDS near-infrared rapid component analyzer to collect spectrum data of pollen processing sample in diffuse reflection mode; preprocess high-dimensional near-infrared spectrum data, and use principal component analysis for dimension reduction processing to extract principal components;

[0011] (4) Data fusion: fuse the electronic eye data, preprocessed near-infrared spectrum and preprocessed electronic nose feature data at the feature level; Z-score standardization is performed on the fused data set,

[0012] The formula of Z-score standardization is as follows:

[0013]

[0014] Where μ is the mean of the feature, σ is the standard deviation of the feature, X is the original feature value, and X scaled is the standardized feature value. (5) Use SAINT model to realize the processing time discrimination of pollen:

[0015] By fusing near-infrared spectrum, electronic nose and electronic eye data, use self-attention mechanism and sample attention mechanism to capture the complex relationship between features and samples, as follows:

[0016] A. Input Embedding Layer: The embedding layer maps the fused multi-source data dimensionally to a high-dimensional feature space. Since our dataset does not contain categorical features, all features are numerical continuous features. Therefore, each feature dimension is directly processed through an independent linear layer to generate an embedding vector. Through independent linear transformations, the embedding layer provides a unified representation for data from different sources. The linear transformation formula is as follows:

[0017] T(x)=Ax+b (3)

[0018] Where T is the linear transformation, x is the input vector, A is the linear transformation matrix, and b is the offset vector;

[0019] B. Self-Attention Layer: By applying a self-attention mechanism along the feature dimension, it captures the correlation and complementary information between features in the sample, adaptively focusing on the key features for judging the heat level. For a given sample, the self-attention mechanism calculates the attention weights based on the query, key, and value vectors obtained from the feature embedding, as shown in the following formula:

[0020]

[0021] Where Q, K, and V are the query, key, and value matrices, respectively, and d k The key vector dimension, scaling factor To prevent the value from being too large;

[0022] C. Inter-sample attention layer: The calculation of inter-sample attention is also based on formula (4), but it is calculated for different samples (rows of the data matrix) in a given batch, rather than just the features of a single sample.

[0023] D. Feedforward Neural Network (FFN): A two-layer feedforward network is applied after the attention layer, and the ReLU activation function formula (5) is used in the middle for nonlinear transformation. The FFN further processes the output of the attention layer to capture the complex feature interactions that are not fully modeled by the attention mechanism.

[0024] The formula for the ReLU activation function is as follows:

[0025] g(x) = max(0,x) (5)

[0026] Here, x represents the output of a neuron in the previous layer.

[0027] E. Processing Temperature Determination: Finally, the processing temperature determination results of the Typhae Pollen samples are output using a softmax-based classification head. The softmax formula is as follows:

[0028]

[0029] Where, y(Z) iZ represents the output result, Z represents the input vector, and k represents the total number of categories.

[0030] Beneficial effects

[0031] (1) This invention empirically demonstrates the significant applicability of the SAINT model in the task of determining the processing temperature of Typha pollen through multi-source data fusion. To date, this is the first invention to collaboratively integrate near-infrared spectroscopy (chemical composition), electronic nose (odor), and electronic eye (color) data into the determination of the processing temperature of Typha pollen. Compared with existing literature methods, the improved deep learning algorithm—SNAIT—in this invention exhibits higher accuracy and stronger robustness. This breakthrough provides a theoretical basis and data support for solving the core challenges of strong reliance on subjective experience and fragmented multimodal data in the traditional quality control of Typha pollen charcoal, and provides an innovative solution for promoting the transformation of the quality control paradigm of "standardization-digitalization" synergistic development of traditional Chinese medicine charcoal.

[0032] (2) The SAINT fusion discriminant model achieved an average classification accuracy of 99.60% on the test set, and its discrimination results were basically consistent with the traditional manual experience discrimination results based on pharmacopoeia descriptions.

[0033] (3) It is worth noting that during the model discrimination process, one batch of samples that was judged as "moderately charred (S-CTP)" based on human experience was accurately classified as "over-charred (H-CTP)" by the model. In-depth analysis shows that this difference is more likely due to the inherent subjectivity and visual evaluation errors in human experience discrimination, rather than model misjudgment. This phenomenon further corroborates that the constructed SAINT multi-source information fusion discrimination model has higher objectivity, consistency and reliability, and can effectively overcome the subjective limitations of human discrimination, providing strong technical support for the accurate and objective judgment of the degree of processing of cattail pollen charcoal. Attached Figure Description

[0034] Figure 1 Odor radar map (a) and PCA loading map (b) of the maximum response values ​​of 10 sensors for charred cattail pollen at different processing levels (PC1 to PC5: the first five principal components);

[0035] Figure 2 Feature-level data fusion strategies;

[0036] Figure 3 Visualization results of confusion matrices for SAINT(a), MLP(b), and SVM(c) under random seed 2025. Detailed Implementation

[0037] Example 1

[0038] 1. Preparation of processed products made from cattail pollen charcoal:

[0039] Raw cattail pollen was purchased from Jiangsu and Inner Mongolia, with batch numbers 220701 and 181201 respectively. It was identified by the School of Pharmacy, Jinan University, as dried pollen of *Typha angustifolia* L. Following the preparation method of cattail pollen charcoal in the 2020 edition of the *Chinese Pharmacopoeia*, appropriate amounts of raw cattail pollen were weighed, and 125 batches of cattail pollen charcoal samples with different degrees of processing were prepared under the guidance of senior processing experts. Among them, 33 batches of samples with "insufficient" charcoal processing, with a yellow to yellowish-brown surface and a charcoal processing temperature of 100-160℃, were marked as L-CTP; 45 batches of samples with "moderate" charcoal processing, with a brownish-black or dark brown surface and a charcoal processing temperature of 190-220℃, were marked as S-CTP; and 47 batches of samples with "excessive" charcoal processing, with a black surface and a charcoal processing temperature of 240-290℃, were marked as H-CTP.

[0040] 2. Olfactory Information Collection

[0041] 2.1 Detection Conditions: The carrier gas for the electronic nose was natural air treated by an activated carbon cleaning device. The detection room temperature was adjusted to 25℃, and the device was preheated for 30 minutes to generate a stable headspace gas before measurement began. The sensor flow rate was 600 mL / min, the carrier gas flow rate was 600 mL / min, the cleaning time was 60 s, and the sampling time was 80 s. The PEN3 electronic nose mainly consists of 10 metal oxide sensor arrays with different performance characteristics, which can effectively analyze different volatile substances. Details of each sensor array and its performance are shown in Table 1.

[0042] Table 1. Introduction to the PEN3 Electronic Nose Sensor

[0043]

[0044] 2.2 Methodological Examination

[0045] A random batch of samples was placed in a container, sealed with double-layer plastic wrap, and left for at least 30 minutes. Electronic nose detection was performed, repeated six times in parallel, recording the values ​​from 10 sensors. The RSD% value was calculated for each sample to assess the instrument's precision. Six samples from the same batch were also placed in a container, sealed with double-layer plastic wrap, and left for at least 30 minutes. Electronic nose detection was performed, and the values ​​from 10 sensors were recorded. The RSD% value was calculated for each sample to assess the method's repeatability. Finally, the same batch of samples was placed in a container, sealed with double-layer plastic wrap, and left for at least 30 minutes. Electronic nose detection parameters were recorded for 0, 2, 4, 6, 8, and 10 hours, and the RSD% value was calculated for each sample to assess the sample's stability.

[0046] Table 2. Methodological Results (RSD%) of 10 Sensors for the Electronic Nose

[0047]

[0048] The results showed that the RSD of the stable response values ​​of each sensor was less than 5%, indicating that the instrument had good precision, the method had good repeatability, and the sample had good stability within 10 hours.

[0049] 2.3 Sample Determination:

[0050] Take 1g of each powder and place it in a 50mL sample cup, then seal it with double layers of plastic wrap. Following the detection conditions in "2.1", collect samples three times for each sample, and analyze the stable signal obtained from the sensor after 73s.

[0051] 3. Visual information acquisition

[0052] 3.1 Measurement conditions:

[0053] The measurement was performed using a D65 light source with a viewing angle of 10° and a measuring aperture of Φ30mm / Φ25.4mm (reflection). The measurement was conducted in SCE mode and after black and white correction.

[0054] 3.2 Methodological Examination

[0055] Randomly select an appropriate amount of samples from the same batch, measure the color at the same time point, repeat the measurement 6 times, and calculate L. * a * b * The RSD% value is used to assess the instrument's precision. Six samples from the same batch are placed in the same container, and their colors are measured. The L value is then calculated. * a * b * The RSD% value was used to assess the repeatability of the method. An appropriate amount of the same batch of samples was placed in the same container, and the color values ​​were measured at 0h, 1h, 2h, 3h, 4h, 5h, and 6h. The measurements were repeated 6 times, and L was calculated. * a * b * The RSD% value is used to assess the stability of the sample.

[0056] Table 3. Methodological Results of the Electronic Eye Detector (RSD%)

[0057]

[0058] The results showed that L * a * b * The RSD < 0.40% indicates that the instrument has good precision, the method has good repeatability, and the sample has good stability within 12 hours.

[0059] 3.3 Sample Determination

[0060] Take an appropriate amount of sample powder, spread it evenly in a dish, and measure each sample three times. Record the colorimetric value: L * a * b * And according to Eab * =(L *2 +a *2 +b *2 ) 1 / 2 (1) Calculate the total chromaticity value Eab * .

[0061] 4. Near-infrared information acquisition

[0062] Sample spectral data were acquired using an XDS near-infrared rapid component analyzer in diffuse reflectance mode. The spectral data acquisition range was 400 to 2498 nm, with a wavelength interval of 2 nm. An appropriate amount of sample was placed in the sample loop, evenly distributed, with the instrument's built-in background used as a reference. During acquisition, the temperature was controlled at 25.0 ± 1.0 °C, and the humidity at 45.0 ± 1.0%. Each sample underwent three parallel scans, and the average value was used for analysis.

[0063] 5. Data fusion and preprocessing

[0064] 5.1 Electronic nose data preprocessing

[0065] The original 10 sensor data acquired by the electronic nose were filtered to identify sensor response data that significantly contributed to the determination of the degree of processing. This was done using radar charts (…). Figure 1 a) shows the maximum response values ​​of 10 sensors. The maximum value of the sensor response signal is closely related to the compound content. It can be seen that the maximum response values ​​of the sensors are different, indicating that the contents of volatile compounds are different. Based on this, PCA was used for feature extraction, and the number of principal components (PCs) was optimized. Figure 1 b shows the loading plots of the first two PCs in the electronic nose data, where PC1 and PC2 explain 87.33% of the data. Using a factor loading value greater than 0.5 as a threshold, sensors that contribute significantly to PC1 and PC2 were selected as W5S, W6S, W1S, W1W, W2S, W2W, and W3S, indicating that nitrogen oxides, methane, alcohols, aldehydes, ketones, aromatics, and organosulfur compounds are likely the main compound types.

[0066] 5.2 Near-infrared data preprocessing

[0067] For high-dimensional near-infrared spectroscopy (NIRS) data, directly inputting the original dimensions can easily lead to a decline in model performance and a weakening of generalization ability. Therefore, principal component analysis (PCA) is used for dimensionality reduction to extract key principal components (PCs). PCA can effectively retain the main information entropy while ensuring rapid dimensionality reduction, which is more in line with the overall predictive performance requirements of the model. Under the guidance of the data fusion strategy, the impact of the number of principal components (PCs) on the performance of the SAINT model during PCA dimensionality reduction of NIRS data was analyzed and compared. Table 4 shows that under the condition of 4 independent random seeds, when the number of PCs is set to 20, SAINT achieved the highest average accuracy (99.60%) in stratified 5-fold cross-validation. By calculating the mean and standard deviation of the accuracy across seeds for each number of PCs (Table 5), the impact mechanism of dimensionality reduction on model performance can be evaluated intuitively and systematically. The higher the mean and the smaller the standard deviation, the better the overall performance and robustness of the SAINT model. At PCs=20, the SAINT model reaches its peak performance with an average accuracy of 99.60% and a standard deviation of 0.46%, confirming that this dimension can fully and effectively retain the key information for fire determination in near-infrared spectral data while maintaining model stability. When the dimension increases to 25, the accuracy decreases and the standard deviation increases, indicating a decline in robustness. This suggests that excessive dimensionality introduces noise or redundant features, leading to uncontrolled model complexity. The accuracy of low-dimensional scenarios (PCs=5, 10, 15) all significantly deteriorates, reflecting that the effective information contained in the data is not fully utilized and information extraction is insufficient. The smaller standard deviation at PCs=5 reveals the cross-modal compensation effect of multi-source data, but the lack of core information still limits further improvement in the accuracy of the SAINT model. This invention empirically demonstrates that the dimensionality of near-infrared spectral PCA dimensionality reduction has a significant regulatory effect on the performance and robustness of the SAINT model. Setting the number of components to 20 maximizes the retention of effective discrimination information while maintaining model stability, achieving a balance between minimizing noise interference.

[0068] Table 4. Accuracy distribution of the model based on hierarchical 5-fold cross-validation with different numbers of principal components (PCs) (4 random seeds)

[0069]

[0070] Table 5. Distribution of mean and standard deviation of model accuracy for different numbers of principal components (PCs)

[0071]

[0072] 5.3 Data Fusion

[0073] The electronic eye data (which retains all original features due to its low inherent dimensionality and excellent signal-to-noise ratio), preprocessed near-infrared spectra, and electronic nose feature data are fused at the feature level using multi-source heterogeneous data. See the flowchart for the feature-level fusion strategy.Figure 2 。

[0074] Specifically, the above multi-source feature data of each batch of samples are subjected to feature-level fusion and integrated into a single feature vector.

[0075] Finally, the fused dataset is subjected to Z-score normalization to eliminate the influence of different feature dimensions on scale-sensitive algorithms and improve the overall performance and training convergence efficiency of the model. The formula for Z-score normalization is as follows:

[0076]

[0077] where μ is the mean of the feature, σ is the standard deviation of the feature, X is the original feature value, and X scaled is the normalized feature value.

[0078] 6. Modeling method

[0079] Self-Attention and Intersample Attention Transformer (SAINT) is a deep learning model based on a lightweight attention mechanism. Different from traditional transformer models mainly developed for sequence data, SAINT integrates feature-level and sample-level interactions and can well adapt to classification tasks based on multi-source data. In this invention, the SAINT model is used to implement the discrimination of the processing火候 of Pollen Typhae. By fusing near-infrared spectroscopy, electronic nose, and electronic eye data, the self-attention mechanism (Self-Attention) and inter-sample attention mechanism (Intersample Attention) are used to capture the complex relationships between features and samples.

[0080] The SAINT architecture consists of the following key parts:

[0081] (1) Input embedding layer:

[0082] The embedding layer maps the fused multi-source data dimension by dimension to a high-dimensional feature space. Since the dataset in this invention does not contain categorical features and all features are numerical continuous features, each feature dimension is directly processed through an independent linear layer to generate an embedding vector. Through independent linear transformation, the embedding layer provides a unified representation for data from different sources and enhances the expression ability of the original features. The linear transformation formula is as follows:

[0083] T(x) = Ax + b (3)

[0084] where T is the linear transformation, x is the input vector, A is the matrix of the linear transformation, and b is the offset vector.

[0085] (2) Self-attention layer:

[0086] By applying a self-attention mechanism along the feature dimension, the correlation and complementary information between features in the sample can be captured, adaptively focusing on key features for heat determination. For a given sample, the self-attention mechanism calculates attention weights based on the query, key, and value vectors obtained from the feature embedding, as shown in the following formula:

[0087]

[0088] Where Q, K, and V are the query, key, and value matrices, respectively, and d k The key vector dimension, scaling factor To prevent the value from being too large.

[0089] (3) Inter-sample attention layer:

[0090] The SAINT architecture is unique in that it introduces an inter-sample attention mechanism (a type of row attention) to capture similarities (same category) and differences (different categories) between samples. Here, attention is calculated for different samples (rows of the data matrix) in a given batch, not just the features of a single sample. The calculation of inter-sample attention is also based on formula (4), but it is applied to samples (rows) rather than features. In our task, different batches of *Pyrola rotundifolia* samples may have similar or significantly different spectral or odor patterns. By mining the relationships between samples, inter-sample attention can improve the model's predictive ability for unseen samples, thereby improving classification accuracy.

[0091] (4) Feedforward Neural Network (FFN):

[0092] A two-layer feedforward network is applied after the attention layer, with the ReLU activation function used in between; this approach is the same as the traditional Transformer architecture; by introducing formula (5) for nonlinear transformation, FFN further processes the output of the attention layer to capture complex feature interactions that are not fully modeled by the attention mechanism;

[0093] The ReLU activation function is as follows:

[0094] g(x) = max(0,x) (5)

[0095] Here, x represents the output of a neuron in the previous layer.

[0096] (5) Processing temperature determination: Finally, the processing temperature determination results of the Typha pollen samples are output through a softmax-based classification head. The softmax formula is as follows:

[0097]

[0098] Wherein, γ(Z) iZ represents the output result, Z represents the input vector, and k represents the total number of categories.

[0099] 7. Modeling Results (Comparison of different machine learning and deep learning methods)

[0100] To systematically evaluate the comprehensive performance and robustness of the SAINT model in the task of determining the processing time of *Coptis chinensis*, this invention systematically compares it with traditional machine learning methods (Support Vector Machine, SVM) and deep learning methods (Multilayer Perceptron, MLP) by fixing multi-source data input, unifying the preprocessing process, and using a hierarchical 5-fold cross-validation framework. Except for differences in model architecture, all other experimental variables are kept strictly consistent. The modeling results of data fusion (Tables 6 and 7) show that the SAINT, MLP, and SVM models perform well, with average accuracies of 99.60%, 98.60%, and 98.40%, respectively. Compared with the literature... [1] The model for determining the processing time of Typha pollen based solely on near-infrared spectral data (with a sample size of n=166 and a training set ratio of 80%) achieved an accuracy of 95.4%. This invention significantly improves performance by employing multi-source fusion data modeling. Even when using traditional models (MLP, SVM) with the same validation strategy (stratified 5-fold cross-validation with 4 random seeds and an equivalent training set ratio of 80%), the average accuracy reaches 98.60% and 98.40% respectively, exceeding the literature benchmark by 3.20-3.00 percentage points, thus demonstrating the effectiveness of multi-source data fusion.

[0101] The SAINT model further demonstrates superior performance. Under the same validation framework, its accuracies using randomized splitting and cross-validation methods are 99.2%, 100%, 99.2%, and 100%, respectively, with an average of 99.60% ± 0.46% and the lowest standard deviation among the three models. (Compared to patents) [2,3] The near-infrared single-mode method improves the average accuracy by ≥1.7 percentage points, compared to the patented method. [4] The color feature method achieves an average accuracy improvement of ≥9.7 percentage points and, for the first time, realizes collaborative modeling of chemical composition (near-infrared), odor (electronic nose), and color (electronic eye). Taking random seed 2025 as an example, the visualization results of the confusion matrix for each test fold are as follows: Figure 3 As shown. SAINT’s superiority is attributed to two core mechanisms: (1) The self-attention module effectively captures the nonlinear interaction between near-infrared spectroscopy, electronic nose and electronic eye data through dynamic weight allocation, which has a significant modeling advantage over the shallow nonlinear modeling of MLP and the sigmoid kernel function of SVM; (2) The inter-sample attention mechanism deeply mines the potential distribution patterns between the three heat levels, which significantly enhances the model’s generalization ability in small sample scenarios.

[0102] Table 6. Accuracy distribution of different models in hierarchical 5-fold cross-validation (4 random seeds)

[0103]

[0104] Table 7 Comparison of mean and standard deviation of accuracy for different models

[0105]

[0106] Example 2

[0107] Based on the morphological standards of charred cattail pollen slices in the Chinese Pharmacopoeia (2025 edition), the processing level of 125 batches of samples with different processing degrees was determined. According to the pharmacopoeia's morphological description (mainly based on the color change of the slices' surface), the samples were divided into three processing levels:

[0108] (1) L-CTP (Low-Carbonized Powder): 33 batches in total. The surface is yellow to yellowish-brown, and the frying temperature is 100-160℃.

[0109] (2) Moderately charred (S-CTP): 45 batches in total. The surface is brown or dark brown, and the charring temperature is 190-220℃.

[0110] (3) Over-charring (H-CTP): A total of 47 batches. The surface is black, and the charring temperature is between 240-290℃.

[0111] To establish an objective and quantitative model for judging the temperature of samples, this study integrated multi-source sensing technologies, including electronic nose, electronic eye, and near-infrared spectroscopy, to acquire raw data from all 125 batches of samples. Through dimensionality reduction and feature extraction, a feature-level data fusion strategy was adopted to effectively integrate key feature information from different sensing modalities, constructing a comprehensive feature set characterizing the temperature of the samples.

[0112] Based on the fused feature set, the SAINT neural network model under the deep learning framework is introduced to perform heat level discrimination modeling.

[0113] Model validation and results analysis:

[0114] (1) The SAINT fusion discriminant model achieved an average classification accuracy of 99.60% on the test set, and its discrimination results were basically consistent with the traditional manual experience discrimination results based on pharmacopoeia characteristics description.

[0115] (2) It is worth noting that during the model discrimination process, one batch of samples that was judged as "moderately charred (S-CTP)" based on human experience was accurately classified as "over-charred (H-CTP)" by the model. In-depth analysis shows that this difference is more likely due to the inherent subjectivity and visual evaluation errors in human experience discrimination, rather than model misjudgment. This phenomenon further corroborates that the constructed SAINT multi-source information fusion discrimination model has higher objectivity, consistency and reliability, and can effectively overcome the subjective limitations of human discrimination, providing strong technical support for the accurate and objective judgment of the degree of processing of cattail pollen charcoal.

[0116] References

[0117] [1] Chen Chengwu, Wang Tianshu, Hu Kongfa, et al. Near-infrared discrimination method for processed cattail pollen based on convolutional neural network and voting mechanism [J]. Spectroscopy and Spectral Analysis, 2022, 42(11):3361-3367.

[0118] [2] A rapid real-time detection method for near-infrared quality control of processed cattail pollen charcoal [P]. Chinese Patent: CN112414967B. 2023.08.18.

[0119] [3] A near-infrared online quality detection method for cattail pollen charcoal products [P]. Chinese Patent: CN113030008B.2022.08.26.

[0120] [4] A method for online control of the processing of cattail pollen [P]. Chinese Patent: CN111487202A. 2020.08.28.

Claims

1. A method for determining the processing time of Typha pollen based on multi-source data fusion and a lightweight attention mechanism, characterized in that, To achieve this, follow these steps: (1) Olfactory information acquisition and preprocessing: The processed sample of cattail pollen was detected by an electronic nose, and the sensor data acquired by the electronic nose was preprocessed, that is, feature extraction was performed by principal component analysis, and the number of principal components was optimized. (2) Visual information acquisition: The processed Typha pollen sample was detected by an electronic eye detector, and the colorimetric value L was recorded. * a * b * According to Eab * =(L *2 +a *2 +b *2 ) 1 / 2 (1) Calculate the total chromaticity value Eab * ; (3) Near-infrared information acquisition: The XDS near-infrared rapid component analyzer was used to acquire spectral data of processed Typha pollen samples in diffuse reflectance mode; the high-dimensional near-infrared spectral data were preprocessed and principal component analysis was used to perform dimensionality reduction to extract principal components. (4) Data fusion: The electronic eye data, preprocessed near-infrared spectrum, and preprocessed electronic nose feature data are fused at the feature level, and the fused dataset is standardized by Z-score. The formula for Z-score standardization is as follows: Where μ is the mean of the feature, σ is the standard deviation of the feature, and X is the original feature value. scaled These are the standardized eigenvalues; (5) Using the SAINT model to determine the processing time of Typha pollen: By fusing near-infrared spectroscopy, electronic nose, and electronic eye data, complex relationships between features and between samples are captured using self-attention and inter-sample attention mechanisms, as detailed below: A. Input Embedding Layer: The embedding layer maps the multi-source data fused in step (4) dimension-by-dimensional to a high-dimensional feature space. Since our dataset does not contain categorical features, all features are numerical continuous features. Therefore, each feature dimension is directly processed by an independent linear layer to generate an embedding vector. Through independent linear transformation, the embedding layer provides a unified representation for data from different sources. The linear transformation formula is as follows: T(x)=Ax+b (3) Where T is the linear transformation, x is the input vector, A is the linear transformation matrix, and b is the offset vector; B. Self-Attention Layer: By applying a self-attention mechanism along the feature dimension, it captures the correlation and complementary information between features in the sample, adaptively focusing on the key features for judging the heat level. For a given sample, the self-attention mechanism calculates the attention weights based on the query, key, and value vectors obtained from the feature embedding, as shown in the following formula: Where Q, K, and V are the query, key, and value matrices, respectively, and d k The key vector dimension, scaling factor To prevent the value from being too large; C. Inter-sample attention layer: The calculation of inter-sample attention is also based on formula (4), but it is calculated for different samples (rows of the data matrix) in a given batch, rather than just the features of a single sample. D. Feedforward Neural Network (FFN): A two-layer feedforward network is applied after the attention layer, and the ReLU activation function formula (5) is used in the middle for nonlinear transformation. The FFN further processes the output of the attention layer to capture the complex feature interactions that are not fully modeled by the attention mechanism. The formula for the ReLU activation function is as follows: g(x) = max(0,x) (5) Where x represents the output of a neuron in the previous layer. E. Processing Temperature Judgment: Finally, the processing temperature judgment results of the Typhae Pollen samples are output through a softmax-based classification head; the softmax formula is as follows: Where, y(Z) i Z represents the output result, Z represents the input vector, and k represents the total number of categories.

Citation Information

Patent Citations

  • Method for controlling processing process of Pollen Typhae on line

    CN111487202A

  • A rapid real-time near-infrared quality control method for detecting processed cattail pollen charcoal

    CN112414967B

  • A near-infrared online quality detection method for cattail pollen charcoal products

    CN113030008B