American ginseng quality high-throughput detection method and system based on near infrared spectrum and multi-task deep learning

By optimizing feature extraction and cross-task interaction through a detection method based on near-infrared spectroscopy and multi-task deep learning, the problem of high-throughput, accurate, and non-destructive detection of multiple quality indicators of American ginseng was solved, achieving high efficiency, accuracy, and transparency in the quality detection of American ginseng.

CN121601085APending Publication Date: 2026-03-03HENAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511442321.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-throughput, accurate, and non-destructive testing of multiple quality indicators in American ginseng. Traditional methods are time-consuming, labor-intensive, and highly destructive to samples. Existing multi-task learning models cannot effectively capture long-distance band correlation information in the near-infrared spectrum and lack efficient cross-task feature interaction mechanisms, resulting in insufficient detection accuracy.

Method used

A detection method based on near-infrared spectroscopy and multi-task deep learning is adopted. By sharing a feature extraction module, a task-specific feature module, a feature interaction module, and a parallel output module, and combining a selective kernel network and a channel cross-attention module, feature extraction and cross-task interaction are optimized to achieve efficient detection of PDG and PTG content in American ginseng.

Benefits of technology

It enables high-throughput, accurate, and non-destructive testing of multiple quality indicators of American ginseng, improving testing accuracy and efficiency. The model has high transparency and is suitable for industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601085A_ABST
    Figure CN121601085A_ABST
Patent Text Reader

Abstract

The invention discloses a near infrared spectrum and multi-task deep learning-based American ginseng quality high-throughput detection method and system, and relates to the technical field of food quality detection and intelligent analysis, and the detection method comprises the following steps: collecting near infrared spectrum data of American ginseng; inputting the near infrared spectrum data into a pre-trained detection model, and analyzing and outputting detection results of the content of the PDG and the PTG of the American ginseng by utilizing the detection model; wherein the detection model comprises a shared feature extraction module, a task specific feature module, a feature interaction module and a parallel output module. The method is used for realizing high-throughput detection of multiple quality indexes of the American ginseng, and the quality detection precision and efficiency of the American ginseng are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of food quality testing and intelligent analysis technology, specifically a high-throughput detection method and system for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning. Background Technology

[0002] American ginseng, a precious traditional Chinese medicine with both medicinal and health benefits, directly determines its pharmacological efficacy, market value, and safety for consumption. Precise detection of key quality indicators such as saponin content (e.g., protopanaxadiol-type PDG and protopanatriol-type PTG), polysaccharides, and amino acids is crucial for quality control, protecting consumer health, and regulating the industry. With the large-scale development of the American ginseng industry chain, traditional quality evaluation methods relying on single-indicator testing and involving time and effort are no longer sufficient to meet the industry's needs for simultaneous analysis of multiple indicators, high-throughput rapid screening, and non-destructive green testing. Therefore, breakthroughs in efficient and accurate multi-dimensional quality assessment technologies are urgently needed.

[0003] In existing technologies, methods for quality detection of American ginseng are mainly divided into two categories: one is chemical analysis methods based on high performance liquid chromatography (HPLC), gas chromatography-mass spectrometry (GC-MS), etc. Although these methods can achieve accurate quantification of target components, they require complex sample pretreatment (such as crushing, extraction, centrifugation, filtration, etc.), have long operation cycles (several hours for a single detection), high consumable costs, and are destructive to the samples, making them unsuitable for rapid detection of large batches of samples and difficult to apply to real-time quality monitoring in the production and distribution process; the other is rapid detection technology based on near-infrared spectroscopy (NIR). Near-infrared spectroscopy can indirectly analyze chemical components and physical properties by capturing the characteristic absorption information of sample molecular vibrations. It has the advantages of being non-destructive, rapid, low-cost, and capable of in-situ detection, and has become one of the mainstream technologies for rapid quality analysis of Chinese medicinal materials.

[0004] To achieve simultaneous prediction of multiple quality indicators of American ginseng, multi-task learning (MTL) has been increasingly applied to near-infrared spectroscopy analysis. MTL, by jointly learning multiple related tasks within a shared framework, can fully utilize the potential correlations between indicators, improve the model's ability to model complex systems, and reduce the information fragmentation problem of single-task modeling. In the technical approach combining MTL and near-infrared spectroscopy, the feature extraction architecture and cross-task interaction mechanism are the core determinants of model performance.

[0005] Early multi-task models often used one-dimensional convolutional neural networks (1D-CNN) as the backbone for feature extraction. While 1D-CNN can effectively capture local feature patterns in the spectrum and suppress noise interference, it is limited by the local receptive field of the convolution kernel and cannot effectively model long-distance band correlation information in near-infrared spectra (such as the feature responses of different saponin components in similar or overlapping bands). This leads to insufficient feature extraction and inadequate prediction accuracy when processing American ginseng samples with complex composition and high spectral overlap. In recent years, the Transformer architecture has been introduced into spectral analysis due to its advantage of modeling global dependencies through its self-attention mechanism. However, this architecture suffers from high computational complexity and stringent hardware resource requirements. It also performs poorly in capturing subtle local spectral features (such as weak absorption peaks of low-content saponins), making it difficult to meet the high-throughput analysis needs of on-site detection and resource-constrained scenarios. Furthermore, its application in near-infrared spectroscopy, a specific one-dimensional sequence data, has not yet developed into a mature solution, especially lacking customized optimizations for the spectral characteristics of American ginseng.

[0006] Besides feature extraction, feature fusion and cross-task interaction are also key to the success of the MTL framework. By jointly learning shared representations and leveraging task relevance, the MTL framework demonstrates superior performance in the multi-quality detection task of American ginseng. However, most existing MTL frameworks still rely on traditional designs, typically including a shared encoder and independent task heads. This design lacks a modeling mechanism for the complex interactions between private and shared features, which may lead to suboptimal feature learning, inter-task interference, and degraded prediction performance.

[0007] In the prior art, Chinese invention patent CN119622520A discloses a method for quality assessment of American ginseng based on multi-task deep learning, a computer-readable storage medium, and a computer program product. Chinese invention patent CN120558902A discloses a method and system for quality detection of American ginseng based on S-transform and multi-task deep learning. Both disclose methods for quality assessment of American ginseng based on multi-task deep learning, which improves detection efficiency to some extent. However, for the detection requirement of simultaneously predicting the content of multiple components (such as diol and triol ginsenosides) in American ginseng, both use conventional convolutional neural networks (CNNs) as the feature extraction backbone. This fails to effectively capture the long-distance band correlation information of different saponin components in the near-infrared spectrum, lacks an efficient cross-task feature interaction mechanism, and struggles to balance the synergistic correlation and feature independence between indicators, resulting in insufficient model generalization ability.

[0008] Therefore, there is an urgent need to develop a new method that integrates near-infrared spectroscopy technology with an advanced multi-task deep learning architecture. By optimizing feature extraction capabilities, constructing an efficient cross-task interaction mechanism, and enhancing model interpretability, this method can achieve high-throughput, accurate, and non-destructive prediction of multiple quality indicators of American ginseng, providing technical support for quality control, quality traceability, and scientific research in the American ginseng industry. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides a high-throughput detection method and system for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning, thereby achieving high-throughput detection of multiple quality indicators of American ginseng and improving the accuracy and efficiency of American ginseng quality detection.

[0010] To achieve the above objectives, the specific solution adopted by this invention is as follows: a high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning, comprising:

[0011] Near-infrared spectral data of American ginseng were collected;

[0012] Near-infrared spectral data is input into a pre-trained detection model, which is then used to analyze and output the detection results of PDG (Protopanaxadiol-type Ginsenosides) and PTG (Protopanaxatriol-type Ginsenosides) content in American ginseng.

[0013] The detection model includes a shared feature extraction module, a task-specific feature module, a feature interaction module, and a parallel output module; the processing methods of the detection model include:

[0014] Near-infrared spectral data is input into the shared feature extraction module to obtain shared feature X3;

[0015] Input the shared feature X3 into the task-specific feature module to obtain the private features of the two tasks;

[0016] The private features of the two tasks are input into the feature interaction module for feature optimization, resulting in two refined private features;

[0017] Two refined private features are fed into the output module for parallel regression to obtain the detection results of PDG and PTG content in American ginseng.

[0018] As a further optimization of the above technical solution, the processing method of the shared feature extraction module includes: linearly projecting the near-infrared spectral data to a fixed hidden dimension and normalizing it, and then processing the normalized near-infrared spectral data through two cascaded one-dimensional selective scanning blocks to output the shared feature X3.

[0019] As a further optimization of the above technical solution, the processing method of the task-specific feature module includes: copying the shared feature X3 and inputting it into two parallel 1D convolutional branches respectively. Each 1D convolutional branch contains a 1D convolutional layer, a batch normalization layer and a ReLU activation function. The private features of the two tasks are extracted and output through the two branches respectively.

[0020] As a further optimization of the above technical solution, the private features of the two tasks are the private features of the PDG detection task and the private features of the PTG detection task.

[0021] As a further optimization of the above technical solution, the processing method of the feature interaction module includes: optimizing private features through a selective kernel network and a channel cross-attention module; wherein, the selective kernel network performs multi-scale feature fusion using statistical cues of private features, the channel cross-attention module generates a channel attention map from the aggregated features of private features, refines the private features of the current task, and finally outputs the refined private features of the two tasks.

[0022] As a further optimization of the above technical solution, the processing method of the parallel output module includes: firstly, flattening the refined private features of the two tasks into one-dimensional vectors respectively, and then inputting each one-dimensional vector into an independent fully connected sub-network. Each fully connected sub-network contains two linear layers with ReLU activation functions. Through parallel regression calculation, the detection results of PDG content and PTG content of American ginseng are output.

[0023] As a further optimization of the above technical solution, the near-infrared spectral data is preprocessed and then input into the shared feature extraction module. The preprocessing process includes: sequentially processing the near-infrared spectral data with Savitzky-Golay smoothing and standard normal variable transformation. The Savitzky-Golay smoothing reduces high-frequency noise by fitting the spectral shift window with a low-order polynomial, and the standard normal variable transformation corrects baseline shift and scattering effects by centering a single spectrum to zero mean and scaling it according to the standard deviation.

[0024] As a further optimization of the above technical solution, the specific process for collecting near-infrared spectral data of American ginseng includes:

[0025] The American ginseng to be tested was sliced ​​and pre-treated, including refrigeration at 2°C and standing at 20°C.

[0026] Near-infrared spectral data of pre-treated American ginseng slices to be tested were collected in the range of 950 to 1650 nm using a near-infrared spectroscopy analyzer.

[0027] As a further optimization of the above technical solution, the pre-training method for the detection model is as follows:

[0028] Near-infrared spectral data of American ginseng samples were collected and the PDG and PTG contents of the American ginseng samples were determined to construct a training dataset;

[0029] The training dataset is divided into a training set, a validation set, and a test set.

[0030] An unsupervised spectral data augmentation strategy was adopted for the training set to increase the sample size of the training set tenfold.

[0031] Using near-infrared spectral data from the training set as input, PDG and PTG content as targets, and a weighted sum of task-specific MSE losses as the overall loss function, the detection model is iteratively trained. The formula for the overall loss function is: Where t=1 corresponds to the PDG prediction task, and t=2 corresponds to the PTG prediction task;

[0032] The trained detection model was tested and evaluated using a test set.

[0033] The final detection model will be determined based on the evaluation results.

[0034] A high-throughput detection system for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning includes a data acquisition module and a detection module. The data acquisition module is used to acquire near-infrared spectral data of American ginseng, and the detection module includes a pre-trained detection model. The detection module is used to input near-infrared spectral data and analyze and output the detection results of PDG and PTG content of American ginseng.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] 1. The detection model MMINet of this invention is not a general single-task model, but is specifically designed for predicting multiple quality indicators of American ginseng. By introducing a task-specific module on the basis of a shared feature extraction module, the model can simultaneously learn multiple quality indicators such as chemical components and nutritional components, achieving true "multi-indicator joint prediction" and avoiding the low efficiency and information fragmentation problems caused by task-based modeling.

[0037] 2. In the feature extraction stage, this invention introduces a two-level Mamba backbone, which can effectively capture the global contextual dependencies and subtle local changes in the near-infrared spectrum. This design is better at handling spectral overlap problems than traditional CNNs, while its computational complexity is much lower than that of Transformers, making it particularly suitable for the refined analysis of complex ginseng samples.

[0038] 3. The feature interaction module of this invention achieves multi-scale information sharing and refinement of differentiated features among tasks through Selective Kernel Network (SKN) and Channel Cross-Attention (CCAM). This mechanism, while ensuring the independence of each task, enhances the model's ability to jointly model correlation indicators, thus outperforming traditional independent task learning frameworks in both prediction accuracy and robustness.

[0039] 4. The SHAP method was used to visualize the prediction results, clearly demonstrating the key spectral bands that the model focuses on in predicting the quality of American ginseng, and establishing their correspondence with known absorption peaks in NIR. This not only improves the transparency and interpretability of the model but also provides a reliable basis for scientific research and application promotion in the quality analysis of American ginseng.

[0040] 5. Benefiting from the rapid and non-destructive nature of near-infrared spectroscopy and the efficient architecture of multi-task deep learning, this invention can achieve multi-index detection of a large number of American ginseng samples in a short time. Furthermore, this method has strong versatility and can be extended to the quality detection and component analysis of other Chinese medicinal materials and edible agricultural products, possessing industrialization and large-scale application value. Attached Figure Description

[0041] Figure 1 Image of a ginseng sample;

[0042] Figure 2 Image of a near-infrared spectrometer (DA7250);

[0043] Figure 3 The original NIR spectrum of American ginseng used to train the detection model;

[0044] Figure 4 A schematic diagram of the structure of the MMINet network for the detection model;

[0045] Figure 5 This is a schematic diagram of the SS1D block structure;

[0046] Figure 6 A structural diagram of a task-specific feature module;

[0047] Figure 7 This is a structural diagram of the feature interaction module;

[0048] Figure 8 This is a schematic diagram of the parallel output module.

[0049] Figure 9 This is a schematic diagram of the algorithm flow for the detection model;

[0050] Figure 10 A visualization of SHAP. Detailed Implementation

[0051] The technical solution of the present invention will be further described in detail below with reference to specific embodiments. Parts not described or disclosed in detail in the following embodiments of the present invention should be understood as prior art known or should be known by those skilled in the art.

[0052] A high-throughput method for detecting the quality of American ginseng based on near-infrared spectroscopy and multi-task deep learning includes:

[0053] Near-infrared spectral data of American ginseng were collected;

[0054] Near-infrared spectral data is input into a pre-trained detection model, which is then used to analyze and output the detection results of PDG and PTG content in American ginseng.

[0055] The detection model includes a shared feature extraction module, a task-specific feature module, a feature interaction module, and a parallel output module; the processing methods of the detection model include:

[0056] Near-infrared spectral data is input into the shared feature extraction module to obtain shared feature X3;

[0057] Input the shared feature X3 into the task-specific feature module to obtain the private features of the two tasks;

[0058] The private features of the two tasks are input into the feature interaction module for feature optimization, resulting in two refined private features;

[0059] Two refined private features are fed into the output module for parallel regression to obtain the detection results of PDG and PTG content in American ginseng.

[0060] Specifically, the detection method includes the following steps:

[0061] S1. Collect near-infrared spectral data of American ginseng;

[0062] The American ginseng to be tested was sliced ​​and pre-treated. Specifically, the ginseng was manually sliced ​​into thin sections, approximately 1 mm thick, with a diameter ranging from 8.2 mm to 15.6 mm. Pre-treatment included refrigeration at 2°C and resting at 20°C, meaning all ginseng slices were kept healthy and intact at 2°C. Before measuring near-infrared spectral data, the slices were removed and kept at 20°C for 8 hours.

[0063] Near-infrared spectral data of pre-treated ginseng slices were acquired in the range of 950 to 1650 nm using a DA7250 NIR analyzer (Perten, Sweden). The resolution was 0.5 nm. The device has a built-in tungsten halide light source. After a 30-minute preheating period, spectral data were acquired, and a sliding window technique was used to reduce the original spectrum from 1401 dimensions to 280 dimensions to reduce redundant information.

[0064] S2. Preprocessing of near-infrared spectral data;

[0065] Two commonly used preprocessing techniques were applied: Savitzky-Golay (SG) smoothing and Standard Normal Variable (SNV) transformation. SG smoothing reduces high-frequency noise by fitting a low-order polynomial to a moving window on the spectrum while preserving the original spectral characteristics. After smoothing, SNV corrects for baseline shift and scattering effects by centering each individual spectrum to zero mean and scaling it by its standard deviation. This combination effectively minimizes variations caused by sample surface properties and improves spectral consistency, facilitating the extraction of meaningful information for subsequent detection.

[0066] S3. The PDG and PTG content of American ginseng was detected using a detection model;

[0067] like Figure 4 As shown, the detection model is MMINet, which consists of four parts: a shared feature extraction module, a task-specific feature module, a feature interaction (FI) module, and a parallel output module.

[0068] S301. Linearly project the near-infrared spectral data to a fixed hidden dimension and normalize it. Then, process the normalized near-infrared spectral data through two cascaded one-dimensional selective scanning blocks and output the shared feature X3.

[0069] The shared feature extraction module introduces a two-level Mamba backbone, effectively capturing global contextual dependencies and subtle local variations in near-infrared spectra. Its aim is to extract shared, information-rich, and global spectral representations from the original one-dimensional spectra, serving as the foundation for subsequent task-specific branches within MMINet. Formally, given an input spectral vector X, the data is first linearly projected onto a fixed hidden dimension and normalized. Subsequently, as... Figure 5 As shown, X1 is a stack of two one-dimensional selective scan (SS1D) blocks used to capture long-range dependencies and contextual features on the spectral sequence. The final shared feature representation is X3. Each SS1D block contains a residual connection to preserve information and enhance gradient flow.

[0070] S302, such as Figure 6 As shown, the task-specific feature module includes two parallel 1D convolution (Conv) branches, each of which contains a 1D convolutional layer, a batch normalization (BatchNorm) layer, and a ReLU activation function.

[0071] The shared feature X3 is copied and fed into two parallel 1D convolutional (Conv) branches. These branches then extract and output the private features for each of the two tasks. These private features are for the PDG detection task and the PTG detection task, respectively.

[0072] S303, such as Figure 7 As shown, the feature interaction module consists of two core components: a Selective Kernel Network (SKN) and a Channel Cross-Attention Module (CCAM). The former performs task-guided multi-scale feature fusion using statistical cues from other tasks, while the latter further refines task-specific representations by applying channel attention maps derived from aggregated features from other tasks. This allows the model to uncover effective collaborative relationships between tasks while maintaining the independence of feature representations for each task.

[0073] In the feature interaction module, the private features are optimized through a selective kernel network and a channel cross-attention module. The selective kernel network performs multi-scale feature fusion using statistical cues of the private features, and the channel cross-attention module generates a channel attention map from the aggregated features of the private features, which refines the private features of the current task and finally outputs the refined private features of the two tasks (PDG detection task and PTG detection task).

[0074] S304, such as Figure 8 As shown, in the parallel output module, the refined private features of the two tasks are first flattened into one-dimensional vectors, and then each one-dimensional vector is input into an independent fully connected sub-network. Each fully connected sub-network contains two linear layers with ReLU activation functions. Through parallel regression calculation, the detection results of PDG and PTG content of American ginseng are output.

[0075] This invention achieves effective information exchange between two regression tasks by using a single model to perform two tasks simultaneously, avoiding feature learning conflicts. It maintains task-specific feature representations, reduces negative transfer, and improves prediction accuracy. Furthermore, it enhances the model's generalization ability and stability, adapting to complex near-infrared spectral scenarios.

[0076] This invention also discloses a high-throughput detection system for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning, including a data acquisition module and a detection module. The data acquisition module is used to acquire near-infrared spectral data of American ginseng, and the detection module includes a pre-trained detection model. The detection module is used to input near-infrared spectral data and analyze and output the detection results of PDG and PTG content of American ginseng.

[0077] The pre-training process of the detection model is designed as follows:

[0078] Step 1: Sample Preparation

[0079] like Figure 1 As shown, the American ginseng samples were identified by experts from Kaifeng Municipal Hospital of Traditional Chinese Medicine and provided by Henan Zhangzhongjing Pharmacy Co., Ltd. (Zhengzhou, China). The ginseng originated from four different growing regions: Weihai City, Baishan City, Montreal, Canada, and Wisconsin, USA. To ensure experimental uniformity, the entire ginseng root was manually cut into thin sections approximately 1 mm thick, with diameters ranging from 8.2 mm to 15.6 mm. For each origin, 35 to 40 samples were prepared, resulting in a total of 150 samples. All samples were healthy, intact, and stored at 2°C. Before measuring near-infrared spectral data and quality indicators, all samples were removed and kept at 20°C for 8 hours to ensure data consistency during near-infrared spectroscopy analysis.

[0080] Step 2: Near-infrared spectroscopy (NIR) data acquisition

[0081] like Figure 3 The image shown is the original NIR spectrum of American ginseng. The near-infrared spectral data of the American ginseng sample were obtained through methods such as... Figure 2 The DA7250 NIR analyzer shown (Perten, Sweden) acquires data in the range of 950 to 1650 nm with a resolution of 0.5 nm. The device incorporates a tungsten halide light source and acquires spectral data after a 30-minute warm-up period. A sliding window technique is used to reduce the original spectrum from 1401 dimensions to 280 dimensions to minimize redundant information.

[0082] Step 3: Sample quality index testing

[0083] After acquiring near-infrared spectral data, targeted chemical analysis was conducted to determine the key internal components of the American ginseng samples.

[0084] For American ginseng, the focus is on quantifying its main bioactive components, especially ginsenosides, whose pharmacological efficacy is widely recognized. Based on their structure, they can be divided into two main categories: protopanaxadiol type (PDG: Rb1, Rc, Rb2, and Rd) and protopanaxtriol type (PTG: Rg1 and Re).

[0085] The sample pretreatment steps are as follows: First, the American ginseng slices were dried and pulverized, and then passed through an 80-mesh sieve to ensure particle uniformity. 100 mg of American ginseng powder was accurately weighed and added to 3 mL of 20% (v / v) ethanol solution. Ultrasonic extraction was performed at 50℃ (40 kHz, 100 W) for 60 minutes. The extract was centrifuged at 12,000 rpm for 15 minutes (25℃) to remove insoluble particles. The supernatant was then filtered through a 0.45 μm polytetrafluoroethylene (PTFE) membrane for subsequent chromatographic analysis.

[0086] Quantification of ginsenosides was performed using high-performance liquid chromatography (HPLC) on a Shimadzu LC-2030C3D Plus system. Separation was achieved using an Agilent ZORBAX XDB-C18 column (4.6 mm × 150 mm, 5 μm particle size). The mobile phase was a binary gradient elution system consisting of acetonitrile (A) and ultrapure water (B), with the program set as follows: 0–20 min, 20% A; 20–35 min, linear gradient to 35% A; 35–42 min, hold at 35% A. The column temperature was controlled at 30 °C, the flow rate at 1.0 mL / min, and the injection volume at 10 μL. The detection wavelength was set to 203 nm.

[0087] Table 1. Statistics on the content of major internal components (g / 1000g) of American ginseng samples.

[0088] Step 4: Spectral Preprocessing

[0089] This invention employs two commonly used preprocessing techniques: Savitzky-Golay (SG) smoothing and Standard Normal Variable (SNV) transformation. SG smoothing reduces high-frequency noise by fitting a low-order polynomial to a moving window on the spectrum while preserving the original spectral characteristics. After smoothing, SNV corrects for baseline shifts and scattering effects by centering each individual spectrum to zero mean and scaling it by its standard deviation. This combination effectively minimizes variations caused by sample surface properties and improves spectral consistency, facilitating the extraction of meaningful information for subsequent modeling.

[0090] Step 5: Construct the detection model

[0091] The detection model comprises four parts: a shared feature extraction module, a task-specific feature module, a feature interaction (FI) module, and a parallel output module, such as... Figure 4 As shown. The structure and operation of each part are detailed in section S3 of the detection method above, and will not be repeated here.

[0092] It should be noted that, to adapt to MTL, the overall loss is expressed as a weighted sum of the task-specific MSE losses, enabling the model to balance the learning of multiple quality metrics during training. The formula is as follows:

[0093] Where t=1 corresponds to the PDG prediction task, and t=2 corresponds to the PTG prediction task.

[0094] Step Six: Experiment Setup

[0095] To address the limited size of the original American ginseng dataset, an unsupervised spectral data augmentation strategy was employed to simulate spectral variability under different conditions. Stratified sampling based on sample origin was used to divide the original American ginseng dataset into training, validation, and test sets in a 5:3:2 ratio. To prevent data leakage and ensure fair evaluation, augmentation was applied exclusively to the training set, increasing its size tenfold. Each comparison method was trained and validated separately to determine its optimal configuration, and then evaluated on the test set. This process was repeated in ten randomized splits, with the model trained and evaluated independently in each run. The final result is reported as the average performance across all ten repetitions.

[0096] Step 7: Model Evaluation

[0097] To evaluate the model's performance in predicting the content of quality indicators in American ginseng, the following metrics were used:

[0098] Using the coefficient of determination (R²) 2 The root mean square error (RMSE) and residual prediction bias (RPD) measure the goodness of fit between the predicted and actual values, error, and prediction stability, respectively.

[0099] Step 8: Visualization

[0100] The SHapley Additive exPlanations (SHAP) are used to explain the characteristic contributions of the proposed MMINet model in predicting the quality indicators of American ginseng. For example... Figure 10 As shown, the generated SHAP summary plots display the top 20 most influential spectral wavelengths for each quality metric in the dataset. In each SHAP summary plot, the y-axis represents the spectral feature (wavelength), and the x-axis represents its SHAP value. Each point corresponds to a test sample, and the color reflects the original feature value. SHAP values ​​close to zero indicate minimal impact on prediction, while larger positive or negative values ​​indicate stronger impact. Positive SHAP values ​​imply a direct correlation with the target variable, while negative values ​​indicate an inverse relationship. A wider distribution of SHAP values ​​at a given wavelength reflects greater variability in its contribution across samples, highlighting its relevance to model predictions.

[0101] The algorithm flow of the model is as follows Figure 9 As shown, since the model uses full-band American ginseng spectral data, the relationship between the bands of interest for the model's prediction of American ginseng quality indicators and the characteristic absorption peaks of the American ginseng NIR spectrum is obtained. This verifies the effectiveness and interpretability of the model.

[0102] The experimental results are shown in Table 2.

[0103] Table 2

[0104] This invention achieves the following objectives by introducing a two-level Mamba backbone network, task-specific features and feature interaction modules, and combining the SHAP algorithm to enhance interpretability:

[0105] At the same time, it can efficiently and accurately predict multiple quality indicators of American ginseng, improving the comprehensiveness and practicality of the test.

[0106] Capture long-range dependencies and global features in near-infrared spectra with low computational overhead, enhancing the modeling capabilities for complex samples;

[0107] Enhance the interaction and decoupling between shared and private features, alleviate interference between tasks, and improve multi-task learning performance;

[0108] By generating feature contribution maps to reveal key spectral bands, the transparency and interpretability of the model in high-throughput detection of American ginseng are improved;

[0109] To develop a rapid, non-destructive, and scalable new method for multi-quality testing of American ginseng, providing reliable technical support for food safety, nutritional evaluation, and industrial applications.

[0110] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A high-throughput method for detecting the quality of American ginseng based on near-infrared spectroscopy and multi-task deep learning, characterized in that, include: Near-infrared spectral data of American ginseng were collected; Near-infrared spectral data is input into a pre-trained detection model. The model analyzes and outputs the detection results of PDG (Protopanaxadiol-type Ginsenosides) and PTG (Protopanaxatriol-type Ginsenosides) content in American ginseng. The detection model includes a shared feature extraction module, a task-specific feature module, a feature interaction module, and a parallel output module. The processing methods of the detection model include: Near-infrared spectral data is input into the shared feature extraction module to obtain shared feature X3; Input the shared feature X3 into the task-specific feature module to obtain the private features of the two tasks; The private features of the two tasks are input into the feature interaction module for feature optimization, resulting in two refined private features; Two refined private features are fed into the output module for parallel regression to obtain the detection results of PDG and PTG content in American ginseng.

2. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The processing method of the shared feature extraction module includes: linearly projecting the near-infrared spectral data to a fixed hidden dimension and normalizing it, and then processing the normalized near-infrared spectral data through two cascaded one-dimensional selective scanning blocks to output the shared feature X3.

3. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The processing method for task-specific feature modules includes: copying the shared feature X3 and inputting it into two parallel 1D convolutional branches. Each 1D convolutional branch contains a 1D convolutional layer, a batch normalization layer, and a ReLU activation function. The private features of the two tasks are extracted and output through the two branches respectively.

4. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The private features of the two tasks are the private features of the PDG detection task and the private features of the PTG detection task.

5. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The processing method of the feature interaction module includes: optimizing private features through a selective kernel network and a channel cross-attention module; wherein, the selective kernel network performs multi-scale feature fusion using statistical cues of private features, the channel cross-attention module generates a channel attention map from the aggregated features of private features, refines the private features of the current task, and finally outputs the refined private features of the two tasks.

6. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The parallel output module's processing method includes: first, flattening the refined private features of the two tasks into one-dimensional vectors, and then inputting each one-dimensional vector into an independent fully connected sub-network. Each fully connected sub-network contains two linear layers with ReLU activation functions. Through parallel regression calculation, the detection results of PDG and PTG content in American ginseng are output.

7. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, After preprocessing, the near-infrared spectral data is input into the shared feature extraction module. The preprocessing process includes sequentially processing the near-infrared spectral data using Savitzky-Golay smoothing and standard normal variable transformation. The Savitzky-Golay smoothing reduces high-frequency noise by fitting the spectral shift window with a low-order polynomial, and the standard normal variable transformation corrects baseline shift and scattering effects by centering a single spectrum to zero mean and scaling it according to the standard deviation.

8. The high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The specific process of collecting near-infrared spectral data of American ginseng includes: The American ginseng to be tested was sliced ​​and pre-treated, including refrigeration at 2°C and static treatment at 20°C. Near-infrared spectral data of pre-treated American ginseng slices to be tested were collected in the range of 950 to 1650 nm using a near-infrared spectroscopy analyzer.

9. A high-throughput detection method for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning according to claim 1, characterized in that, The pre-training method for the detection model is as follows: Near-infrared spectral data of American ginseng samples were collected and the PDG and PTG contents of the American ginseng samples were determined to construct a training dataset; The training dataset is divided into a training set, a validation set, and a test set. An unsupervised spectral data augmentation strategy was adopted for the training set to increase the sample size of the training set tenfold. Using near-infrared spectral data from the training set as input, PDG and PTG content as targets, and a weighted sum of task-specific MSE losses as the overall loss function, the detection model is iteratively trained. The formula for the overall loss function is: Where t=1 corresponds to the PDG prediction task, and t=2 corresponds to the PTG prediction task; The trained detection model was tested and evaluated using a test set. The final detection model will be determined based on the evaluation results.

10. A high-throughput detection system for American ginseng quality based on near-infrared spectroscopy and multi-task deep learning, characterized in that, It includes a data acquisition module and a detection module. The data acquisition module is used to collect near-infrared spectral data of American ginseng, and the detection module includes a pre-trained detection model. The detection module is used to input near-infrared spectral data and analyze and output the detection results of PDG and PTG content of American ginseng.

Citation Information

Patent Citations

  • American ginseng quality evaluation method based on multi-task deep learning, computer readable storage medium and computer program product

    CN119622520A

  • American ginseng quality detection method and system based on S transformation and multi-task deep learning

    CN120558902A