Traditional Chinese medicinal material heavy metal content detection system based on multispectral imaging and AI fusion
The detection system, which integrates multispectral imaging and AI, solves the problems of complex procedures, expensive equipment, and sample destructiveness in the detection of heavy metal content in Chinese medicinal materials. It enables rapid and accurate detection of heavy metal content in Chinese medicinal materials and is suitable for grassroots institutions and Chinese medicinal material planting enterprises.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for detecting heavy metal content in Chinese medicinal materials are complex, time-consuming, labor-intensive, require expensive equipment, have low detection efficiency, and are highly destructive to samples, failing to meet the needs of the rapidly developing Chinese medicinal materials industry.
A detection system based on multispectral imaging and AI fusion is adopted, including a data acquisition and preprocessing module, a spectral combination and subset extraction module, a constraint factor construction module, and a model training module. Data is acquired through multispectral imaging, inconsistency constraint factors are constructed, and an AI regression model is trained to achieve rapid and accurate detection of heavy metal content.
It enables rapid batch testing, reduces testing costs, preserves sample integrity, and improves testing accuracy and adaptability, making it suitable for grassroots institutions and Chinese medicinal herb cultivation enterprises.
Smart Images

Figure CN121783880A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quality testing of Chinese medicinal materials, and more specifically, to a system for detecting the heavy metal content of Chinese medicinal materials based on the fusion of multispectral imaging and AI. Background Technology
[0002] In the planting, processing, and storage of Chinese medicinal herbs, heavy metal contamination is common due to soil pollution, the improper use of pesticides and fertilizers, and cross-contamination during processing. Heavy metals such as lead, mercury, cadmium, and arsenic can accumulate in the human body to a certain level, causing serious harm to human health, such as damaging the nervous system, liver, kidneys, and other vital organs, and affecting normal physiological functions. Therefore, accurate detection of heavy metal content in Chinese medicinal herbs is crucial for ensuring their quality and safe use.
[0003] Traditional methods for detecting heavy metal content in Chinese medicinal herbs mainly include atomic absorption spectrometry and inductively coupled plasma mass spectrometry. While these methods offer high accuracy, they also have some significant limitations:
[0004] The detection process is complex, usually requiring complicated pretreatment of Chinese medicinal materials samples, such as digestion and extraction, to convert heavy metals into detectable forms. These pretreatment steps are not only time-consuming and labor-intensive, but also prone to introducing errors. The equipment is expensive; traditional detection equipment is costly, with high purchase and maintenance costs, limiting its widespread application in some grassroots testing institutions and Chinese medicinal material planting enterprises. The detection efficiency is low; each test can only be performed on a single sample, making it difficult to achieve rapid, large-scale testing, which cannot meet the needs of the rapidly developing Chinese medicinal material industry. The test is also highly destructive to the samples; the samples are usually destroyed during the testing process, making it impossible to retain complete samples for subsequent research or other uses.
[0005] With the continuous development of multispectral imaging and artificial intelligence technologies, new ideas and methods have been provided for the detection of heavy metal content in traditional Chinese medicine (TCM) herbs. Multispectral imaging technology can acquire spectral information of samples at different wavelengths, reflecting the physical and chemical properties of the samples; artificial intelligence algorithms can perform in-depth mining and analysis of large amounts of spectral data, establish a mapping relationship between spectral features and heavy metal content, and achieve rapid and accurate detection. Therefore, developing a TCM herbal heavy metal content detection system based on the fusion of multispectral imaging and AI has significant practical implications. Summary of the Invention
[0006] The purpose of this invention is to provide a heavy metal content detection system for traditional Chinese medicine (TCM) based on multispectral imaging and AI fusion. This system addresses the shortcomings of existing methods for heavy metal content detection in TCM, which are characterized by complex processes requiring extensive pretreatment of samples, such as digestion and extraction, to convert heavy metals into detectable forms. These pretreatment steps are not only time-consuming and labor-intensive but also prone to introducing errors. Furthermore, the equipment is expensive, with traditional detection equipment incurring high purchase and maintenance costs, limiting its widespread application in some grassroots testing institutions and TCM cultivation enterprises. Additionally, the detection efficiency is low, as each test can only target a single sample, making rapid, large-scale testing difficult and failing to meet the needs of the rapidly developing TCM industry. Finally, the system is highly destructive to samples, often damaging them during the testing process and preventing the preservation of intact samples for subsequent research or other uses, thus failing to meet the required standards.
[0007] This invention achieves the above objective through the following technical solution: a system for detecting heavy metal content in traditional Chinese medicinal materials based on multispectral imaging and AI fusion, the system comprising:
[0008] The system includes a data acquisition and preprocessing module, a spectral combination and subset extraction module, a constraint factor construction module, a model training module, and a detection output module.
[0009] The data acquisition and preprocessing module is used to acquire the original multispectral imaging data of the Chinese medicinal materials to be detected and to perform preprocessing.
[0010] The spectral combination and subset extraction module is used to construct multiple spectral combinations, extract preprocessed multispectral imaging data in groups, and obtain imaging subsets corresponding to different spectral combinations.
[0011] The constraint factor construction module is used to calculate the consistency evaluation index between each spectral combination imaging subset and to construct the inconsistency constraint factor based on the nonlinear mismatch characteristics of the spectral response caused by heavy metals.
[0012] The model training module is used to input the inconsistency constraint factor into the AI regression model, and learn the mapping relationship between the constraint factor and the heavy metal content through model training to establish a quantitative discrimination model.
[0013] The detection output module is used to analyze the inconsistency constraint factors of unknown Chinese medicinal material samples using a trained quantitative discrimination model, and output the detection results of heavy metal content of the Chinese medicinal materials.
[0014] Furthermore, when the data acquisition and preprocessing module acquires the raw multispectral imaging data, it includes the following steps:
[0015] Data on Chinese medicinal materials samples that maintain their complete physical morphology and are free from damage, mold, and obvious impurities were collected using multispectral imaging equipment.
[0016] During the data collection process, the ambient light intensity is controlled within a preset range;
[0017] The vertical distance between the fixed sample and the imaging lens;
[0018] The multispectral imaging data includes spectral images of different wavelength channels covering the visible to near-infrared bands, and the image size of each wavelength channel is not less than a preset pixel standard.
[0019] Furthermore, the preprocessing process of the data acquisition and preprocessing module includes:
[0020] The image denoising, geometric correction, and radiometric correction steps are performed sequentially.
[0021] The image denoising uses an adaptive filtering algorithm to dynamically adjust the window size to remove environmental noise and device electronic noise.
[0022] The geometric correction establishes a coordinate mapping relationship through a calibration plate, corrects imaging distortion errors, and ensures that the spatial alignment accuracy of images from different wavelength channels meets the preset requirements.
[0023] The radiation correction uses a normalization formula to eliminate the influence of uneven illumination and differences in device response, and normalizes the corrected grayscale values to a preset range.
[0024] Furthermore, the spectral combination and subset extraction module constructs the spectral combination by including the following steps:
[0025] Based on preset spectral combination rules, multiple spectral combinations covering different band types are constructed. Different channel combinations are selected from all wavelength channels to form non-repeating spectral combinations. Each spectral combination contains at least 2 wavelength channels, and the intersection of channels between any two combinations does not exceed a preset proportion of the total number of channels.
[0026] During the extraction of the imaging subset, for each spectral combination, the corresponding imaging data is extracted from the preprocessed complete dataset according to the channel index, while strictly preserving the spatial location information and grayscale features of each wavelength channel.
[0027] Furthermore, when the constraint factor construction module calculates the consistency evaluation index, it includes the following steps:
[0028] Two complementary indices, structural similarity index and spectral angle matching degree, are used. The structural similarity index focuses on the brightness, contrast and structural consistency of the image, reflecting the degree of imaging matching in the spatial domain.
[0029] The spectral angle matching degree focuses on the pixel-level consistency of the spectral vector direction, reflecting the degree of response matching in the spectral domain;
[0030] The inconsistency constraint factor is obtained by weighted fusion of the deviation values of the two consistency evaluation indicators. The weighting coefficient is determined by cross-validation combined with grid search to ensure the balanced contribution of consistency information in the spatial domain and the spectral domain.
[0031] Furthermore, the constraint factor construction module also performs an ordered integration of the inconsistency constraint factors of all spectral combination pairs;
[0032] The system uses a pre-defined outlier removal rule to remove factor values that deviate from the normal range, forming a feature vector for input to the AI model.
[0033] The feature vector comprehensively characterizes the imaging inconsistencies of Chinese medicinal materials under different spectral combinations.
[0034] Furthermore, the AI regression model of the model training module adopts a three-stage architecture of feature extraction, nonlinear mapping, and output regression, including:
[0035] Input layer, feature extraction layer, intermediate mapping layer, and output layer;
[0036] The input layer receives the feature vector of the inconsistency constraint factor, and the input data type is a floating-point number with a preset precision;
[0037] The feature extraction layer employs either a convolutional neural network or a Transformer architecture.
[0038] The intermediate mapping layer contains multiple fully connected layers, each configured with a corresponding activation function, normalization layer and dropout layer;
[0039] The output layer consists of a single neuron and uses a linear activation function.
[0040] Furthermore, the model training module also includes:
[0041] The training dataset construction and preprocessing unit is used to construct a training dataset containing various types of Chinese medicinal herbs and various common heavy metal elements.
[0042] The true heavy metal content of the sample is obtained through high-precision detection methods and used as a label.
[0043] The training set, validation set, and test set are divided according to a preset ratio. The input features are standardized and feature perturbation enhancement is used to improve the model's generalization ability.
[0044] Furthermore, the training process of the model training module adopts a preset optimizer, learning rate scheduling strategy and loss function, sets a maximum number of training rounds and adopts an early stopping strategy to monitor the model convergence status, and performs model training according to a preset batch size.
[0045] During training, the model parameters are initialized, data is read in batches, the predicted values are obtained through forward propagation, the loss value is calculated, and the model is optimized by backpropagation and parameter updates.
[0046] After each preset number of training rounds, the model performance is evaluated on the validation set. The model with the best performance on the validation set is selected as the final quantitative discrimination model and saved in a preset format.
[0047] Furthermore, when the detection output module processes unknown Chinese medicinal material samples, it first preprocesses the samples, then collects multispectral imaging data according to the same equipment parameters and combination rules as the training phase, performs preprocessing, constructs spectral combinations, extracts imaging subsets, and calculates the inconsistency constraint factor feature vector.
[0048] The feature vector is processed using standardized parameters saved during the training phase and then input into a quantitative discrimination model. The predicted heavy metal content is obtained through forward inference of the model.
[0049] The detection output module also performs a confidence assessment on the prediction results. It sets a preset threshold according to the detection scenario. When the prediction uncertainty does not exceed the preset threshold, it outputs the prediction value and the corresponding confidence interval; otherwise, it outputs a detection anomaly prompt and related suggestions.
[0050] The beneficial effects of this invention are as follows:
[0051] 1. This system constructs multiple spectral combinations and extracts imaging subsets, making full use of multispectral information. It calculates inconsistency constraint factors using structural similarity index and spectral angle matching degree, taking into account both spatial and spectral domains. The three-stage architecture of the AI regression model can automatically learn complex nonlinear relationships, with multi-stage collaboration, accurately capturing spectral changes caused by heavy metals, greatly improving detection accuracy, and providing a reliable basis for quality control of Chinese medicinal materials.
[0052] 2. Multispectral imaging equipment can acquire multispectral images of samples at once, skipping complex preprocessing and individual detection processes, greatly shortening detection time, and enabling rapid batch detection. The trained model can quickly analyze unknown samples and output results in the detection output module, which meets the urgent need of the Chinese medicinal materials industry for efficient detection and powerfully promotes the rapid development of the industry.
[0053] 3. The system uses multispectral imaging technology, which does not damage the Chinese medicinal material samples during the entire detection process and preserves the samples completely. This not only avoids sample waste, but also provides convenience for subsequent research, preservation and reuse. It is especially suitable for the detection of precious or rare Chinese medicinal materials, which helps to explore the value of Chinese medicinal materials and ensure the sustainable use of resources.
[0054] 4. Compared with expensive detection equipment such as traditional atomic absorption spectrometry and inductively coupled plasma mass spectrometry, multispectral imaging equipment is more affordable. Moreover, the system integrates artificial intelligence algorithms, reducing reliance on high-cost instruments and significantly lowering the overall detection cost. This makes it affordable for grassroots testing institutions and Chinese medicinal herb planting enterprises, which is conducive to the widespread promotion and application of the technology.
[0055] 5. The system's training dataset covers a variety of Chinese medicinal herbs and common heavy metal elements. After training, it can adapt to the detection of different types and elements, and has strong versatility. The detection output module evaluates the confidence of the prediction results, sets thresholds, and provides anomaly prompts and suggestions. It can maintain reliability and stability in complex scenarios, providing comprehensive protection for the quality detection of Chinese medicinal herbs. Attached Figure Description
[0056] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0057] Figure 1 This is a flowchart of the data acquisition and preprocessing module of the present invention;
[0058] Figure 2 This is a flowchart of the spectral combination and subset extraction module of the present invention;
[0059] Figure 3 This is a flowchart of the constraint factor construction module of the present invention;
[0060] Figure 4 This is a flowchart of the model training and detection output module of the present invention. Detailed Implementation
[0061] The present application will now be described in further detail with reference to the accompanying drawings. It should be noted that the following specific embodiments are only used to further illustrate the present application and should not be construed as limiting the scope of protection of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application based on the above application content.
[0062] Example 1:
[0063] Please see Figure 1-4 This invention provides a technical solution: a system for detecting heavy metal content in traditional Chinese medicinal materials based on multispectral imaging and AI fusion, the system comprising:
[0064] The module includes: data acquisition and preprocessing module, spectral combination and subset extraction module, constraint factor construction module, model training module, and detection output module.
[0065] The data acquisition and preprocessing module is used to acquire and preprocess the raw multispectral imaging data of the Chinese medicinal materials to be tested.
[0066] Among them, the raw multispectral imaging data is the initial data obtained by imaging Chinese medicinal materials with multispectral imaging equipment. This data contains image information of Chinese medicinal materials in different spectral bands, reflecting the absorption and reflection characteristics of Chinese medicinal materials to different spectra. Preprocessing is a series of processing operations performed on the acquired raw multispectral imaging data. The purpose is to remove noise, correct distortion, and enhance image quality, so as to analyze and process the data more accurately in the future. Common preprocessing methods include image denoising, geometric correction, and radiometric correction.
[0067] The spectral combination and subset extraction module is used to construct multiple spectral combinations, extract preprocessed multispectral imaging data in groups, and obtain imaging subsets corresponding to different spectral combinations.
[0068] Among them, spectral combination is a combination of multiple bands selected from the many spectral bands covered by multispectral imaging according to certain rules or strategies. Different spectral combinations may highlight the characteristic information of Chinese medicinal materials in different aspects. Group extraction is to divide the preprocessed multispectral imaging data according to the constructed multiple spectral combinations, extract the data corresponding to each spectral combination, and form different imaging subsets. Each imaging subset contains only image information under a specific spectral combination.
[0069] The constraint factor construction module is used to calculate the consistency evaluation index between each spectral combination imaging subset. Based on the nonlinear mismatch characteristics of the spectral response caused by heavy metals, the inconsistency constraint factor is constructed.
[0070] Among them, consistency evaluation indicators are quantitative indicators used to measure the similarity or consistency between different spectral combination imaging subsets. By calculating these indicators, the differences between imaging subsets under different spectral combinations can be understood. Common consistency evaluation indicators may include correlation coefficient, structural similarity index (SSIM), etc. The nonlinear mismatch feature of spectral response is due to the presence of heavy metals in Chinese medicinal materials, which causes nonlinear changes in the spectral response of Chinese medicinal materials. That is, the response relationship under different spectral bands no longer follows a simple linear law. This nonlinear mismatch feature reflects the influence of heavy metals on the spectral characteristics of Chinese medicinal materials. The inconsistency constraint factor is a factor constructed based on the nonlinear mismatch feature of spectral response caused by heavy metals and the consistency evaluation indicators between different spectral combination imaging subsets. This factor can reflect the degree of inconsistency between different spectral combination imaging subsets caused by the presence of heavy metals and is used to establish the correlation with heavy metal content in subsequent model training.
[0071] The model training module is used to input inconsistency constraint factors into the AI regression model, and learn the mapping relationship between constraint factors and heavy metal content through model training to establish a quantitative discrimination model.
[0072] AI regression models are regression analysis models built using artificial intelligence technology. They aim to learn the mapping relationship between input variables (here, inconsistency constraint factors) and output variables (heavy metal content). Common AI regression models include Support Vector Regression (SVR) and Neural Network Regression (such as Multilayer Perceptron (MLP) and Convolutional Neural Networks (CNN) used for regression tasks). Model training involves inputting the inconsistency constraint factors as input data and the known heavy metal content of the medicinal herbs as output data into the AI regression model. By continuously adjusting the model's parameters, the model can learn the mapping relationship between input and output as accurately as possible, thereby establishing a quantitative discriminant model. This process typically requires iterative optimization using a large amount of sample data. The quantitative discriminant model, obtained after model training, is a model that can quantitatively output the heavy metal content of medicinal herbs based on the input inconsistency constraint factors. This model can be used to detect the heavy metal content of new samples.
[0073] The detection output module is used to analyze the inconsistency constraint factors of unknown Chinese medicinal material samples using the trained quantitative discrimination model, and output the detection results of heavy metal content of Chinese medicinal materials.
[0074] Among them, unknown Chinese medicinal material samples are those whose heavy metal content is unknown and are the objects to be tested; the analysis involves inputting the inconsistency constraint factor calculated from the unknown Chinese medicinal material sample into the trained quantitative discriminant model, and using the mapping relationship already learned by the model to infer and analyze the heavy metal content of the sample; the heavy metal content detection result is the specific value or range of heavy metal content in the unknown Chinese medicinal material sample obtained after analysis by the quantitative discriminant model, which is the final output of the entire detection system and is used to evaluate whether the Chinese medicinal material meets the heavy metal content standards and other related indicators.
[0075] It should be noted that during use, the data acquisition and preprocessing module acquires and processes the raw data, removing noise and other interference to provide a high-quality foundation for subsequent analysis. The spectral combination and subset extraction module constructs multiple combinations and extracts them in groups, which can explore the characteristics of Chinese medicinal materials from different spectral perspectives and comprehensively reflect their properties. The constraint factor construction module calculates the consistency index and constructs the inconsistency constraint factor to accurately capture the nonlinear changes in spectral response caused by heavy metals and highlight the impact of heavy metals. The model training module uses an AI regression model to learn the mapping relationship between constraint factors and heavy metal content, establishes a quantitative discrimination model, and improves the accuracy and objectivity of detection. The detection output module uses the trained model to analyze unknown samples and quickly output the heavy metal content results, achieving efficient detection, which helps to ensure the quality and safety of Chinese medicinal materials and promote the healthy development of the Chinese medicine industry.
[0076] In one embodiment, acquiring and preprocessing the raw multispectral imaging data of the Chinese medicinal material to be detected includes:
[0077] The raw imaging data of the Chinese medicinal materials to be tested were acquired using multispectral imaging equipment. The medicinal material samples must maintain their complete physical morphology, free from damage, mold, and obvious impurities. During the acquisition process, the ambient light intensity was controlled to remain stable. The vertical distance between the sample and the imaging lens is fixed at 1. To ensure image clarity and consistency; multispectral imaging data contains spectral images of N different wavelength channels, denoted as the dataset:
[0078]
[0079] in, This represents the two-dimensional image data of the k-th wavelength channel, with an image size of at least 512×512 pixels. , The value range is 3-20, and the wavelength channel covers the visible to near-infrared band. It includes characteristic sensitive wavelengths, such as 650nm, 940nm, and 1450nm, where heavy metals have significant responses.
[0080] The raw imaging data is preprocessed by sequentially performing image denoising, geometric correction, and radiometric correction steps. Image denoising employs an adaptive median filtering algorithm, with the window size dynamically adjusted based on noise intensity. to This effectively removes environmental noise and equipment electronic noise; geometric correction establishes a mapping relationship between pixel coordinates and actual spatial coordinates through a calibration board, correcting distortion errors during the imaging process and ensuring spatial alignment accuracy of images from different wavelength channels is better than one pixel; radiometric correction uses the following formula for normalization to eliminate the effects of uneven illumination and differences in equipment response:
[0081]
[0082] In the formula, For the corrected first Each wavelength channel at the pixel The grayscale value at that location is normalized to a range of 1000. , and The first The minimum and maximum gray values of all pixels in the image of each wavelength channel are calculated by traversing all pixels in the image.
[0083] This design allows for multispectral imaging covering multiple bands, including those with significant heavy metal responses, enabling comprehensive capture of heavy metal information. Preprocessing steps are performed sequentially, with adaptive median filtering providing flexible and effective noise reduction, geometric correction correcting distortion to ensure spatial alignment, and radiometric correction eliminating the effects of lighting and equipment differences. High-quality raw data and precise preprocessing lay the foundation for subsequent analysis, removing various interference factors and making the data more accurately reflect the characteristics of Chinese medicinal materials, improving detection accuracy, avoiding misjudgments due to data issues, and ensuring the reliability of the entire detection system from the source.
[0084] In one embodiment, multiple spectral combinations are constructed, and the preprocessed multispectral imaging data is grouped and extracted to obtain imaging subsets corresponding to different spectral combinations, including:
[0085] Multiple spectral combinations are constructed based on preset spectral combination rules. These rules must cover combinations of different wavelength types, such as visible-visible, visible-near-infrared, and near-infrared-near-infrared, to ensure comprehensive capture of spectral response differences caused by heavy metals. Different channel combinations are selected from N wavelength channels to form M unique spectral combinations, denoted as .
[0086]
[0087] Each spectral combination It contains at least two wavelength channels, and the intersection of any two combinations of channels does not exceed a certain percentage of the total number of channels. , , The range of values is to (Number of combinations), ensuring combination diversity;
[0088] For each spectral combination Based on the channel index from the preprocessed complete dataset:
[0089]
[0090] Extract the corresponding imaging data to form an imaging subset:
[0091]
[0092] The imaging subset strictly preserves the spatial location information of each wavelength channel, i.e., the pixel. The one-to-one correspondence and grayscale value features are not subjected to additional spatial downsampling or feature compression processing, ensuring the accuracy of subsequent consistency evaluation.
[0093] This design pre-determines multiple band combinations to construct spectral combinations, ensuring comprehensive capture of spectral response differences caused by heavy metals. It selects non-repeating combinations from multiple channels to guarantee combination diversity. When extracting imaging subsets, it retains spatial location and grayscale value characteristics without additional processing. Diverse spectral combinations can analyze the impact of heavy metals from different angles, comprehensively mine data information, retain original features to ensure accurate subsequent consistency evaluation, avoid the loss of key information due to downsampling or compression, and make the detection results more accurately reflect the heavy metal status of Chinese medicinal materials.
[0094] In one embodiment, a consistency evaluation index is calculated among the various spectral combination imaging subsets. Based on the nonlinear mismatch characteristics of the spectral response caused by heavy metals, an inconsistency constraint factor is constructed, including:
[0095] For any two spectral combinations and Corresponding imaging subset and The consistency between the two is calculated using two complementary indices: the structural similarity index (SSIM) and the spectral angle matching degree (SAM). The structural similarity index focuses on the brightness, contrast and structural consistency of the image, reflecting the degree of imaging matching in the spatial domain; the spectral angle matching degree focuses on the pixel-level spectral vector direction consistency, reflecting the degree of response matching in the spectral domain. The combination of the two can comprehensively characterize the imaging consistency between spectral combinations.
[0096] The formula for calculating the structural similarity index is:
[0097] in, , Imaging subsets , The mean of all pixel grayscale values in the image. , These are the variances of the grayscale values of the two entities, respectively. Let covariance be the gray values of the two subsets. , It is a constant. This is the maximum gray value after normalization, used to avoid the denominator being zero and to ensure calculation stability;
[0098] The formula for calculating spectral angle matching is:
[0099] In the formula, , These are the wavelength channel indices commonly contained in the two spectral combinations, for each pixel. Calculate the local spectral angles, then average the spectral angles of all pixels in the entire image to obtain the overall spectral angle matching degree between the two imaging subsets. The value range is [value range missing]. The smaller the angle, the more consistent the spectral response;
[0100] Based on the nonlinear mismatch characteristics of spectral response caused by heavy metals—that is, the accumulation of heavy metals in Chinese medicinal materials disrupts their normal spectral absorption and reflection patterns, leading to irregular deviations in imaging responses under different spectral combinations—the deviation values of consistency evaluation indicators are modeled as inconsistency constraint factors. This factor quantifies the degree of imaging inconsistency between spectral combinations caused by heavy metals, and is calculated using the following formula:
[0101] In the formula, , For the weighting coefficients, satisfying The determination rule is as follows: a combination of 5-fold cross-validation and grid search is used, with the parameter range of the grid search being... ∈[0.3,0.7]、 ∈[0.3,0.7], with a step size of 0.05, select the weight combination that maximizes the coefficient of determination (R²) and minimizes the mean squared error (MSE) of heavy metal content prediction on the validation set. The optimal value range is typically [0.3,0.7]. ∈[0.4-0.6], ∈[0.4-0.6], ensuring a balanced contribution of consistent information between the spatial and spectral domains;
[0102] For all spectral combinations, The inconsistent constraint factors are systematically integrated, outliers are eliminated, and through... After removing factor values that deviate from the mean by more than three standard deviations, the criterion forms a dimension of Feature vectors:
[0103] This feature vector comprehensively characterizes the imaging inconsistencies of Chinese medicinal materials under different spectral combinations, serving as the core input feature of the AI model.
[0104] This design employs two indices: structural similarity index and spectral angle matching degree. It comprehensively characterizes the imaging consistency between spectral combinations from both the spatial and spectral domains. Based on the nonlinear mismatch characteristics of heavy metal spectral response, deviation values are modeled as inconsistency constraint factors. After integrating and eliminating outliers, feature vectors are formed. These two complementary indices comprehensively and accurately measure consistency, while the inconsistency constraint factor quantifies the degree of imaging inconsistency caused by heavy metals. The feature vectors comprehensively characterize the features, providing core input for the AI model. This enables the model to better learn the relationship between heavy metal content and features, thereby improving detection accuracy.
[0105] In one embodiment, an inconsistency constraint factor is input into an AI regression model. The model is trained to learn the mapping relationship between the constraint factor and heavy metal content, establishing a quantitative discrimination model, including:
[0106] (1) Model architecture determination
[0107] We select a deep learning regression model as the base model and adopt a three-stage architecture of feature extraction, nonlinear mapping, and output regression.
[0108] Input layer: Receives feature vectors of inconsistency constraint factors. The dimension is the same as the feature vector dimension. Dimension, the input data type is a 32-bit floating-point number;
[0109] Feature extraction layer: Optional Convolutional Neural Network (CNN) or Transformer architecture, choose one of the two configurations:
[0110] CNN configuration: 3 one-dimensional convolutional layers. The first layer has 64 kernels, a size of 3×1, a stride of 1, and a padding pattern of "same". The second layer has 128 kernels, a size of 5×1, a stride of 1, and a padding pattern of "same". The third layer has 64 kernels, a size of 3×1, a stride of 1, and a padding pattern of "same". Each convolutional layer is followed by a batch normalization (BatchNorm1d) layer and a ReLU activation function, with the dropout probability set to 0.2.
[0111] Transformer configuration: 2-layer encoder, each encoder contains an 8-head self-attention mechanism, head dimension = feature vector dimension / 8 and a feedforward neural network, hidden layer dimension = 256, the self-attention mechanism uses a lower triangular mask to prevent future information leakage, the feedforward neural network activation function is GELU, the post-encoder connection layer is normalized (LayerNorm), and the dropout probability is set to 0.2;
[0112] Intermediate mapping layer: 2 fully connected layers. The first layer has 128 neurons, the activation function is LeakyReLU, the negative slope is 0.01, and it connects the batch normalization layer and the dropout layer with a probability of 0.2. The second layer has 64 neurons, the activation function is LeakyReLU, the negative slope is 0.01, there is no batch normalization layer, and the dropout layer is retained with a probability of 0.2.
[0113] Output layer: a single neuron, using a linear activation function, with the output range mapped to [0,5] mg / kg, covering the common heavy metal limits in Chinese medicinal materials.
[0114] (2) Training dataset construction and preprocessing
[0115] Dataset size: Includes A sample of Chinese medicinal herbs with known heavy metal content. It covers at least three different types of Chinese medicinal materials, such as ginseng, wolfberry, and astragalus, and three common heavy metal elements, such as lead, cadmium, and mercury. The number of samples for each combination of Chinese medicinal materials and heavy metals is no less than 50.
[0116] Tag Acquisition: Actual Heavy Metal Content Values Measured by inductively coupled plasma mass spectrometry (ICP-MS), the measurement accuracy is better than Label data is rounded to two decimal places;
[0117] Data partitioning: The training set, validation set, and test set are randomly partitioned in a ratio of 7:2:1. Before partitioning, stratified sampling is used to ensure that the types of Chinese medicinal materials and the distribution of heavy metal elements in each subset are consistent with the original dataset.
[0118] Input feature preprocessing: processing the feature vector of inconsistency constraint factors Standardization is performed, and the formula is: ,in and These are the mean and standard deviation of the feature vectors in the training set, calculated only based on the training set to avoid data leakage;
[0119] Data augmentation: Feature perturbation augmentation is applied to the training set, and random additions are made to the data. Gaussian noise is used, and 5% of the feature dimensions are randomly selected for 0-value masking to enhance the model's generalization ability.
[0120] (3) Training parameter configuration
[0121] Optimizer: The AdamW optimizer is used, with parameters set as follows: initial learning Weight decay coefficient ;
[0122] Learning rate scheduling strategy: Cosine Annealing (LR), T_max = 50 rounds, meaning one cosine cycle is completed in 50 rounds, eta_min = 1e-6, the minimum learning rate, and the learning rate update formula for each round is as follows: ;
[0123] Loss function: Weighted mean squared error (WMSE) loss is used. For high-content samples, Assign a weight of 2.0 to the low-content sample. Assign a weight of 1.0, and the calculation formula is as follows:
[0124] In the formula, , or , , The number of samples in the training set. These are the model's predicted values;
[0125] Training rounds: Maximum training rounds = 200 rounds, adopting an early stopping strategy, the monitoring metric is the validation set WMSE loss, the early stopping threshold is determined by the following rule: preset minimum improvement threshold = 1e-4, when the decrease in validation set loss is less than this threshold for 10 consecutive rounds, the model is judged to have converged and training is stopped, and the current optimal model weights are saved.
[0126] Batch Size: Set to 16, 32, or 64 depending on hardware resources. 32 is preferred to balance training efficiency and stability. If there is insufficient video memory, reduce it to 16.
[0127] (4) Model training steps
[0128] Initialize model parameters: Use He normal distribution to initialize the weights of convolutional and fully connected layers, and initialize the bias term to 0; use Xavier uniform distribution to initialize the weights of the Transformer self-attention layer.
[0129] Data loading: Batch loading is used. Shuffling is enabled for the training set (shuffle=True), while shuffling is disabled for the validation and test sets (shuffle=False). Multi-threading is used during loading, with num_workers=4 to accelerate data reading.
[0130] Forward propagation: Batch data from the training set is input into the model, passing sequentially through the input layer, feature extraction layer, intermediate mapping layer, and output layer to obtain the predicted value. ;
[0131] Loss calculation: Calculate the loss between the predicted value and the actual value based on the WMSE loss function;
[0132] Backpropagation: The gradient of the loss function with respect to each model parameter is calculated using an automatic differentiation mechanism, and the gradient clipping threshold is set to 1.0 to prevent gradient explosion;
[0133] Parameter update: Update all trainable parameters of the model based on gradients using the AdamW optimizer;
[0134] Validation and logging: After each training round, calculate WMSE loss, coefficient of determination (R²), and mean absolute error (MAE) on the validation set, and record information such as model parameters, learning rate, and training time;
[0135] Model selection: After training, the model with the largest R² on the validation set and the smallest WMSE loss is selected as the final quantitative discrimination model. The model's performance metrics on the test set must meet the following requirements: ;
[0136] Model saving: Save the optimal model as Save in a format that includes model architecture configuration, layer weight parameters, and input feature standardization parameters. , Information such as training parameters and logs is available to facilitate subsequent inference calls.
[0137] This design defines the architecture of the deep learning regression model, clarifies the configuration of each layer, constructs a large-scale training dataset, rigorously divides and preprocesses the data, employs data augmentation to improve generalization ability, rationally configures training parameters, specifies training steps in detail, rigorously selects and preserves the optimal model, and the reasonable architecture and parameter configuration enable the model to efficiently learn data features. The large-scale and diverse datasets enhance the model's generalization ability, and the strict training process and selection criteria ensure excellent model performance, achieving high accuracy requirements on the test set, thus providing a reliable model for accurately detecting the heavy metal content of Chinese medicinal materials.
[0138] In one embodiment, a trained quantitative discriminant model is used to analyze the inconsistency constraint factors of unknown Chinese medicinal material samples, and the heavy metal content detection results of the Chinese medicinal materials are output, including:
[0139] For samples of unknown Chinese medicinal materials to be tested, sample preprocessing is first performed to remove surface impurities and ensure flat placement. Then, multispectral imaging data is acquired. After denoising, geometric correction, and radiometric correction, spectral combinations are constructed and imaging subsets are extracted. Finally, the inconsistency constraint factor feature vector is calculated. The equipment parameters and combination rules used in the processing are kept consistent with those used in the training phase to avoid systematic errors.
[0140] Will Using standardized parameters saved during the training phase, , Standardization process is performed to obtain ;
[0141] Will Input the trained quantitative discrimination model and obtain the predicted heavy metal content of the sample through forward inference. During the prediction process, the variance of the output features of each layer of the model is recorded for subsequent confidence assessment.
[0142] The confidence level of the prediction results is assessed using the prediction standard deviation. Indicates forecast uncertainty. The Monte Carlo dropout method was used to calculate the predictions. All dropout layers were enabled during the inference phase, and the inference was repeated 10 times. The standard deviation of the 10 predictions was then taken. The preset threshold was determined dynamically based on the accuracy requirements of the detection scenario and the limits for heavy metal elements.
[0143] High-precision detection scenarios, such as laboratory quantitative analysis: threshold The determination is based on the measurement uncertainty of the ICP-MS method. This is 1.5 times the standard error, ensuring that the prediction uncertainty does not exceed the measurement error of the reference method.
[0144] In routine screening scenarios, such as rapid testing in production sites: threshold The determination is based on the limit standards for corresponding heavy metal elements in traditional Chinese medicine, such as the limit for lead. One-third of the samples must be rejected to ensure that the false negative rate for non-compliant samples does not exceed 5%.
[0145] Custom scenarios: Allow users to set a threshold based on the 95th percentile of the prediction error, according to specific needs, such as specific Chinese medicinal materials or target heavy metal elements, based on the prediction error distribution of the training set samples.
[0146] when When the preset threshold is met, the predicted value is output. Retain two decimal places and the corresponding 95% confidence interval. The result will be used as the final test result; otherwise, an abnormal test result will be output, including a prediction uncertainty exceeding the threshold, and suggestions to recollect the sample, check whether the sample is placed flat, whether there are impurities, or check the equipment status, calibrate the imaging lens, and check the stability of the light source.
[0147] This design involves processing the samples to be tested according to the training phase process to obtain and standardize the inconsistency constraint factors, inputting them into the model to obtain predictions, using the Monte Carlo dropout method to evaluate the confidence of the prediction results, setting thresholds according to different scenarios, and outputting results or indicating anomalies accordingly. The unified processing flow avoids systematic errors, ensures detection consistency, and the confidence assessment and dynamic threshold setting make the detection results more reliable. It can be flexibly adjusted according to the needs of different scenarios, improving the practicality and accuracy of detection, and providing an effective basis for the quality control of Chinese medicinal materials.
[0148] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0149] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A system for detecting heavy metal content in traditional Chinese medicinal materials based on multispectral imaging and AI fusion, characterized in that, The system includes: The system includes a data acquisition and preprocessing module, a spectral combination and subset extraction module, a constraint factor construction module, a model training module, and a detection output module. The data acquisition and preprocessing module is used to acquire the original multispectral imaging data of the Chinese medicinal materials to be detected and to perform preprocessing. The spectral combination and subset extraction module is used to construct multiple spectral combinations, extract preprocessed multispectral imaging data in groups, and obtain imaging subsets corresponding to different spectral combinations. The constraint factor construction module is used to calculate the consistency evaluation index between each spectral combination imaging subset and to construct the inconsistency constraint factor based on the nonlinear mismatch characteristics of the spectral response caused by heavy metals. The model training module is used to input the inconsistency constraint factor into the AI regression model, and learn the mapping relationship between the constraint factor and the heavy metal content through model training to establish a quantitative discrimination model. The detection output module is used to analyze the inconsistency constraint factors of unknown Chinese medicinal material samples using a trained quantitative discrimination model, and output the detection results of heavy metal content of the Chinese medicinal materials.
2. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 1, characterized in that, When the data acquisition and preprocessing module acquires raw multispectral imaging data, it includes the following steps: Data on Chinese medicinal materials samples that maintain their complete physical morphology and are free from damage, mold, and obvious impurities were collected using multispectral imaging equipment. During the data collection process, the ambient light intensity is controlled within a preset range; The vertical distance between the fixed sample and the imaging lens; The multispectral imaging data includes spectral images of different wavelength channels covering the visible to near-infrared bands, and the image size of each wavelength channel is not less than a preset pixel standard.
3. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 2, characterized in that, The preprocessing process of the data acquisition and preprocessing module includes: The image denoising, geometric correction, and radiometric correction steps are performed sequentially. The image denoising uses an adaptive filtering algorithm to dynamically adjust the window size to remove environmental noise and device electronic noise. The geometric correction establishes a coordinate mapping relationship through a calibration plate, corrects imaging distortion errors, and ensures that the spatial alignment accuracy of images from different wavelength channels meets the preset requirements. The radiation correction uses a normalization formula to eliminate the influence of uneven illumination and differences in device response, and normalizes the corrected grayscale values to a preset range.
4. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 1, characterized in that, The spectral combination and subset extraction module constructs the spectral combination by including the following steps: Based on preset spectral combination rules, multiple spectral combinations covering different band types are constructed. Different channel combinations are selected from all wavelength channels to form non-repeating spectral combinations. Each spectral combination contains at least 2 wavelength channels, and the intersection of channels between any two combinations does not exceed a preset proportion of the total number of channels. During the extraction of the imaging subset, for each spectral combination, the corresponding imaging data is extracted from the preprocessed complete dataset according to the channel index, while strictly preserving the spatial location information and grayscale features of each wavelength channel.
5. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 1, characterized in that, When the constraint factor construction module calculates the consistency evaluation index, it includes the following steps: Two complementary indices, structural similarity index and spectral angle matching degree, are used. The structural similarity index focuses on the brightness, contrast and structural consistency of the image, reflecting the degree of imaging matching in the spatial domain. The spectral angle matching degree focuses on the pixel-level consistency of the spectral vector direction, reflecting the degree of response matching in the spectral domain; The inconsistency constraint factor is obtained by weighted fusion of the deviation values of the two consistency evaluation indicators. The weighting coefficient is determined by cross-validation combined with grid search to ensure the balanced contribution of consistency information in the spatial domain and the spectral domain.
6. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 5, characterized in that: The constraint factor construction module also performs an ordered integration of the inconsistency constraint factors for all spectral combination pairs. The system uses a pre-defined outlier removal rule to remove factor values that deviate from the normal range, forming a feature vector for input to the AI model. The feature vector comprehensively characterizes the imaging inconsistencies of Chinese medicinal materials under different spectral combinations.
7. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 1, characterized in that, The AI regression model in the model training module adopts a three-stage architecture of feature extraction, nonlinear mapping, and output regression, including: Input layer, feature extraction layer, intermediate mapping layer, and output layer; The input layer receives the feature vector of the inconsistency constraint factor, and the input data type is a floating-point number with a preset precision; The feature extraction layer employs either a convolutional neural network or a Transformer architecture. The intermediate mapping layer contains multiple fully connected layers, each configured with a corresponding activation function, normalization layer and dropout layer; The output layer consists of a single neuron and uses a linear activation function.
8. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 7, characterized in that, The model training module also includes: The training dataset construction and preprocessing unit is used to construct a training dataset containing multiple types of Chinese medicinal herbs and multiple common heavy metal elements. The true heavy metal content of a sample is obtained through high-precision detection methods and used as a label. The training set, validation set, and test set are divided according to a preset ratio. The input features are standardized and feature perturbation enhancement is used to improve the model's generalization ability.
9. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 8, characterized in that: The training process of the model training module adopts a preset optimizer, learning rate scheduling strategy and loss function, sets a maximum number of training rounds and adopts an early stopping strategy to monitor the model convergence status, and performs model training according to a preset batch size. During training, the model parameters are initialized, data is read in batches, the predicted values are obtained through forward propagation, the loss value is calculated, and the model is optimized by backpropagation and parameter updates. After each preset number of training rounds, the model performance is evaluated on the validation set. The model with the best performance on the validation set is selected as the final quantitative discrimination model and saved in a preset format.
10. The heavy metal content detection system for traditional Chinese medicinal materials according to claim 1, characterized in that: When the detection output module processes unknown Chinese medicinal material samples, it first preprocesses the samples, then collects multispectral imaging data according to the same equipment parameters and combination rules as the training phase, performs preprocessing, constructs spectral combinations, extracts imaging subsets, and calculates the feature vector of inconsistency constraint factors. The feature vector is processed using standardized parameters saved during the training phase and then input into a quantitative discrimination model. The predicted heavy metal content is obtained through forward inference of the model. The detection output module also performs a confidence assessment on the prediction results. It sets a preset threshold according to the detection scenario. When the prediction uncertainty does not exceed the preset threshold, it outputs the prediction value and the corresponding confidence interval; otherwise, it outputs a detection anomaly prompt and related suggestions.