Nondestructive testing method and system for dyeing pigment of traditional Chinese medicine salvia miltiorrhiza based on hyperspectral imaging
By using hyperspectral imaging technology to screen core feature bands and construct a pixel-level classification model, the problems of low non-destructive testing accuracy and difficulty in identifying local staining in the staining detection of the traditional Chinese medicine Danshen were solved, thus realizing efficient and reliable non-destructive testing of Danshen decoction pieces.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-17
AI Technical Summary
Existing staining detection methods for the traditional Chinese medicine Danshen have problems such as low non-destructive testing accuracy, difficulty in identifying local staining, poor data reliability, and low detection efficiency. In particular, in the application of hyperspectral imaging technology, the curse of dimensionality and difficulty in identifying weak staining features are prone to occur.
Using hyperspectral imaging technology, a competitive adaptive reweighting algorithm is used to select core feature bands. Combined with principal component analysis and a pixel-level classification model, pixel-by-pixel identification and staining contamination ratio calculation are achieved. Real-time quality control is performed using reference samples from the skin and cross-sections, and a pixel-level classification model is constructed for judgment.
It enables non-destructive and accurate identification of locally stained areas in traditional Chinese medicine Danshen slices, reduces the probability of boundary misjudgment, improves detection efficiency and data reliability, and meets the high-throughput quality control needs of the modern Chinese medicine industry.
Smart Images

Figure CN121877765A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of non-destructive quality testing technology of traditional Chinese medicine, and in particular relates to a non-destructive testing method and system for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Danshen, a traditional Chinese medicine, is an important herb widely used in the clinical treatment of cardiovascular diseases and other conditions. Its quality directly affects clinical efficacy and medication safety. However, in the distribution of Danshen, some unscrupulous merchants, driven by profit, often use illegal dyeing methods to conceal defects such as mold, insect infestation, excessive sulfur fumigation, or poor quality. These dyed Danshen slices not only significantly reduce their efficacy but may also pose potential health risks and threaten public medication safety.
[0004] Currently, the main methods for detecting staining in the traditional Chinese medicine Danshen (Salvia miltiorrhiza) include traditional sensory identification, chemical detection, and conventional spectroscopic detection. Traditional sensory identification relies on the visual and olfactory experience of the testing personnel, making it highly subjective and inadequate for identifying slight or localized staining. It is also easily affected by factors such as the testing personnel's skill level and work status. While chemical detection can achieve qualitative and quantitative analysis of specific pigment components, it requires pretreatment such as sample crushing and extraction, resulting in sample damage, long testing cycles, complex operations, and high reagent consumption. Conventional spectroscopic detection, although offering the advantage of non-destructive testing, often involves spectral acquisition and analysis of the entire medicinal slice, making it difficult to accurately locate locally stained areas, leading to insufficient reliability and affecting the accuracy of the test results.
[0005] Currently, some research is applying hyperspectral imaging technology to the precise detection of illegally stained Chinese medicinal herbs. Hyperspectral imaging is an emerging non-destructive testing technology that combines imaging technology and spectroscopy, capable of simultaneously acquiring spatial and spectral information of the object being tested. However, its application faces the following problems: First, hyperspectral data contains hundreds of bands, resulting in extremely high information dimensionality, while the effective sample size for a specific staining problem is usually limited, easily leading to the "curse of dimensionality" and overfitting problems in subsequent modeling. Second, during the identification process, the mixing of spectra at the boundary between stained areas and normal tissue makes it difficult to identify weak staining features. Summary of the Invention
[0006] To overcome the shortcomings of the existing technologies, this invention provides a non-destructive detection method and system for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging. It aims to solve the problems of low non-destructive detection accuracy, difficulty in identifying local staining, poor data reliability, and low detection efficiency in the existing Danshen staining detection, and to meet the urgent needs of the modern Chinese medicine industry for high-throughput quality control.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a non-destructive detection method for staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging; A non-destructive detection method for staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging includes: Hyperspectral image data of the Chinese herbal medicine pieces to be tested and the reference sample were acquired using a hyperspectral imaging system. The collected hyperspectral data were preprocessed to extract the average spectrum of each medicinal slice to form a medicinal slice-level spectral dataset. A competitive adaptive reweighting algorithm was used to select the core feature bands from the entire spectrum to distinguish between stained and normal categories. The medicinal slice-level spectral dataset was analyzed by principal component analysis to output the separation effect between stained and normal Salvia miltiorrhiza slices. Based on the core feature bands, pixel-level data of the hyperspectral image is extracted; pixel-level annotation is performed using the rule of inward shrinking and coloring the visible boundary to construct a training set and train a pixel-level classification model. The image data of the medicinal slices to be tested under the core feature band is input into the pixel-level classification model for pixel-by-pixel identification to obtain the classification results of stained pixels and normal pixels; the proportion of staining contamination is calculated based on the classification results, and a judgment conclusion is output according to a preset threshold.
[0008] As a further technical solution, a hyperspectral imaging system is used to acquire hyperspectral image data of the Chinese herbal medicine pieces to be tested and the reference sample, including: The tested Danshen slices and the reference sample were laid flat on the stage in a dark room, and hyperspectral image data in a preset band range were acquired using a hyperspectral imaging system. The reference samples include the bark and cross-sectional samples of the Chinese herbal medicine slices, used to simultaneously monitor the spectral stability during the collection process.
[0009] As a further technical solution, the collected hyperspectral data is preprocessed to extract the average spectrum of each medicinal slice to construct a medicinal slice-level spectral dataset, including: After performing black-and-white correction on the collected hyperspectral data, the corrected data is segmented by thresholding. The average spectral data of each medicinal slice is extracted by removing the background, thus forming a medicinal slice-level spectral dataset.
[0010] As a further technical solution, a competitive adaptive reweighting algorithm is used to select core feature bands from the entire band to distinguish between stained and normal categories, including: The number of Monte Carlo sampling times was set, and K-fold cross-validation was adopted as the cross-validation strategy. The root mean square error was used as the evaluation index for variable selection. The maximum number of latent variables in the partial least squares model was set. The input herbal medicine slice-level spectral dataset is preprocessed by centering to eliminate the impact of baseline drift on feature selection; Based on the preprocessed herbal medicine slice spectral dataset, a preset number of Monte Carlo sampling iterations are performed to generate a frequency matrix, and the selected frequency of each band variable is calculated based on the frequency matrix. By setting a frequency threshold, the core characteristic bands used to distinguish between stained and normal Danshen are obtained.
[0011] As a further technical solution, the analysis of the spectral dataset of processed medicinal herbs using principal component analysis to output the separation effect between stained and normal Salvia miltiorrhiza slices includes: The selected characteristic band spectral dataset of medicinal slices was used as input and automatically standardized preprocessed. A PCA model is built based on the preprocessed dataset. The spectral data in the original dimensional feature space is projected onto the low-dimensional basis vectors through linear transformation, and the principal components are set and preserved. The preprocessed dataset is then projected onto the principal component space to generate a score matrix. Principal component score maps are plotted based on the score matrix, and the coordinate positions of stained samples and normal samples in the principal component space are marked and the separation effect is output.
[0012] As a further technical solution, pixel-level data of the hyperspectral image is extracted based on the core feature bands; pixel-level annotation is performed using a rule of inward shrinking and coloring visible boundaries to construct a training set and train a pixel-level classification model, including: Based on the acquired core feature bands, pixel-level data of the hyperspectral image is extracted; For each image of Salvia miltiorrhiza slices, it was divided into a stained core area, a normal Salvia miltiorrhiza area, and a background area by labeling; Identify the visible edges of stained patches in the image, use these edges as a reference, shrink the pixel distance along the inside of the image, and use the shrunken boundary as the outline to delineate the region. Traverse all pixels within each region, obtain the reflectance value of each pixel under the core feature band, form the spectral feature vector of the pixel, integrate the spectral feature vectors to construct a pixel-level training dataset. The model was trained based on the constructed pixel-level training dataset to obtain a classification model for pixel-level identification of staining in Danshen slices.
[0013] As a further technical solution, the image data of the medicinal slices to be tested under the core feature band is input into the pixel-level classification model for pixel-by-pixel identification to obtain the classification results of stained pixels and normal pixels; based on the classification results, the proportion of staining contamination is calculated, and a judgment conclusion is output according to a preset threshold, including: Iterate through all pixels in the feature image of the sample to be tested and extract the spectral feature vector of each pixel one by one; The extracted spectral feature vectors are input into the trained pixel-level classification model. The pixel-level classification model judges the category of the input pixel spectral feature vectors based on the differences in spectral features of the stained core region, normal danshen region and background region learned during the training phase, and obtains a pixel-level classification result matrix covering the entire spatial range of the danshen slice image to be tested. Traverse all elements in the result matrix, count the total number of stained pixels and normal Danshen pixels, and output the judgment result by calculating the percentage of staining contamination.
[0014] The second aspect of the present invention provides a non-destructive detection system for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging.
[0015] A non-destructive detection system for staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging includes: The image data acquisition module is configured to acquire hyperspectral image data of the Chinese herbal medicine pieces to be tested and the reference sample using a hyperspectral imaging system; The feature band screening module is configured to: preprocess the collected hyperspectral data, extract the average spectrum of each medicinal slice to form a medicinal slice-level spectral dataset; use a competitive adaptive reweighting algorithm to screen out the core feature bands from the entire spectrum to distinguish between stained and normal categories, and analyze the medicinal slice-level spectral dataset using principal component analysis to output the separation effect between stained and normal Danshen medicinal slices; The classification model construction module is configured to: extract pixel-level data from the hyperspectral image based on the core feature bands; perform pixel-level annotation using the rule of inward shrinking and coloring visible boundaries to construct a training set and train a pixel-level classification model; The identification and judgment module is configured to: input the image data of the medicinal slices to be tested under the core feature band into the pixel-level classification model, perform pixel-by-pixel identification, and obtain the classification results of stained pixels and normal pixels; calculate the proportion of staining contamination based on the classification results, and output the judgment conclusion according to the preset threshold.
[0016] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the method for non-destructive detection of pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in the first aspect of the present invention.
[0017] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the method for non-destructive detection of pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in the first aspect of the present invention.
[0018] The above one or more technical solutions have the following beneficial effects: This invention utilizes hyperspectral imaging technology for detection, eliminating the need for destructive processing such as crushing or extraction of the Salvia miltiorrhiza slices throughout the entire process, thus preserving the sample's morphology and quality intact. Through pixel-level analysis strategies combined with refined annotation rules, it effectively avoids spectral mixing interference at the boundary between stained and normal areas, accurately identifying slightly stained localized areas on the surface of the slices. This solves the industry pain point that traditional sensory identification and conventional spectral detection struggle to detect localized staining.
[0019] During the data acquisition phase, reference samples including the skin and cross-sections were introduced. The acquisition process was monitored in real time by setting segmented RSD thresholds to eliminate data deviations caused by instrument fluctuations and environmental interference at the source. Through CARS screening and PCA validation, core feature bands were selected and their spectral separability was verified, further ensuring data validity. The classification model trained using the selected core feature bands maintained stable performance in the detection of Salvia miltiorrhiza slices from different batches and origins, reducing the probability of boundary misclassification.
[0020] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0022] Figure 1 This is a flowchart of the method in the first embodiment.
[0023] Figure 2 This is a schematic diagram of the hyperspectral imaging system in the first embodiment.
[0024] Figure 3 This provides a reusable workflow for visual SVM classification built in the first embodiment when there are a large number of images.
[0025] Figure 4 This is a schematic diagram comparing the hyperspectral data of stained tanshinone in the first embodiment with the staining distribution thermogram of stained tanshinone.
[0026] Figure 5 This is a system structure diagram of the second embodiment. Detailed Implementation
[0027] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0028] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0029] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0030] Example 1 like Figure 1 As shown, this embodiment discloses a non-destructive detection method for staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging. When collecting hyperspectral data of Danshen slices, reference samples including the peel and cross-section are introduced. Real-time quality control is achieved by continuously acquiring and calculating the relative standard deviation (RSD) of the spectrum. Subsequently, a two-stage strategy of CARS screening and PCA verification is used to obtain core feature bands. A training set is then constructed based on refined annotation rules to train a classification model for pixel-by-pixel recognition. Finally, by calculating the staining contamination ratio and comparing it with a preset threshold, a binary judgment conclusion and a visualized heatmap are output. This method effectively solves the industry problem of difficult detection of local staining in Danshen. The specific operating steps are as follows: Step S1: Use a hyperspectral imaging system to acquire hyperspectral image data of the Chinese herbal medicine slices to be tested and the reference sample.
[0031] like Figure 2 As shown, the tested Danshen (Salvia miltiorrhiza) slices and reference samples were laid flat on the stage in a darkroom environment. A hyperspectral imaging system was used to acquire hyperspectral image data in the 400-1000 nm wavelength range. The reference samples included the peel and cross-sectional samples of the herbal slices. Multiple consecutive acquisitions of the reference samples were performed, and the relative standard deviation (RSD) between the average spectral curves of each acquisition was calculated to monitor and ensure the reliability of the data for the same batch in real time. When the RSD value was lower than a preset threshold, the spectral data acquired in that batch was considered stable and reliable; otherwise, the equipment stability needed to be checked and the data reacquired.
[0032] Step S2: Preprocess the collected hyperspectral data, extract the average spectrum of each slice to form a slice-level spectral dataset; use a competitive adaptive reweighting algorithm to select the core feature bands from the entire band to distinguish between stained and normal categories, and analyze the slice-level spectral dataset using principal component analysis to output the separation effect between stained and normal Salvia miltiorrhiza slices. Step S21: Perform black-and-white correction on the acquired hyperspectral data, perform threshold segmentation and background subtraction on the corrected hyperspectral data, and extract the average spectral data of each medicinal slice to form a medicinal slice-level spectral dataset. The medicinal slice-level spectral dataset consists of a... It consists of a two-dimensional matrix, where The number of medicinal slices samples. This is spectral data. Each row in the matrix represents the average spectral vector of a medicinal slice, and each column represents the reflectance value of a specific wavelength band.
[0033] Step S22: A competitive adaptive reweighting algorithm is used to select the core feature bands from the entire band to distinguish between stained and normal categories.
[0034] Among them, the Competitive Adaptive Reweighting Algorithm (CARS) takes the competition of partial least squares (PLS) regression coefficient weights as its core, and selects the feature bands that are most discriminative in distinguishing between stained or normal categories through an adaptive sampling-exponential decay elimination-cross-validation process.
[0035] The number of Monte Carlo sampling iterations is set to 30 in this embodiment. This parameter determines the robustness of the algorithm. Multiple sampling and modeling reduce the impact of randomness and ensure that the selected feature bands are stable and reliable. The cross-validation method is set to 10-fold cross-validation, with the root mean square error (RMSE) used as the model performance evaluation metric. The maximum number of latent variables in the partial least squares (PLS) model within the CARS algorithm is configured to 30 to fully capture the linear relationship between spectral data and category labels (i.e., "stained" and "normal" labels).
[0036] The input herbal medicine slice-grade spectral dataset undergoes a centering preprocessing step, which involves calculating the mean of each band variable and subtracting the corresponding mean from the reflectance values of all samples in that band. The formula is as follows: ( (The mean value for this band) is used to eliminate the impact of baseline drift on feature selection.
[0037] Based on the initialized parameters and the preprocessed herbal medicine slice-level spectral dataset, Monte Carlo sampling is performed a preset number of times. Each time, a subset of samples is randomly selected from the preprocessed spectral data training set to construct a training subset, simulating data variability for PLS model training. Based on the PLS model constructed from each sampling, the absolute value of the regression coefficient of each band variable is calculated. According to the elimination ratio set by the exponential decay function (EDF), band variables with smaller absolute values of regression coefficients are forcibly removed to achieve preliminary variable screening, as shown in the following formula:
[0038] Where k is the current iteration number, K is the total number of iterations, and c is the decay coefficient (used to control the elimination rate). Each round retains the initial int(p×r) k ) high-weight bands.
[0039] For the band variables retained after the initial screening, the band weights are calculated based on the absolute values of their regression coefficients. , (in (where is the regression coefficient of the i-th band), and the weights are positively correlated with the absolute values of the regression coefficients. A weighted sampling method is used to determine the variable subset for the next iteration, increasing the probability that more important band variables will be retained in subsequent iterations. After each iteration, 10-fold cross-validation is used to calculate the RMSE value of the PLS model built with the current variable subset, and the band variables retained during the iteration are recorded. The calculation formula is:
[0040] in, For real labels, is the predicted value, and n is the number of samples in the validation set.
[0041] After completing a preset number of Monte Carlo sampling iterations, a frequency matrix F is generated, which records the number of times each band variable is retained in all iterations; the selected frequency E of each band variable is calculated, which is the ratio of the number of times it is retained to the total number of samplings. (in This indicates whether the i-th band is selected in the k-th sampling (1 for selected, 0 for unselected). A frequency threshold is set, which is 0.3 in this embodiment. This means that a band variable must be considered important in more than 9 (30×30%) samplings to be finally selected, thus ensuring the high reliability and robustness of the selected features. Finally, all band variables with a selection frequency greater than 0.3 are determined as the final subset of feature bands, filtering out the core feature bands.
[0042] Step S23: Analyze the spectral dataset of processed medicinal slices using principal component analysis and output the separation effect between stained and normal Danshen slices.
[0043] Filtered results The characteristic band spectral dataset of medicinal slices was used as input for PCA analysis, whereby... This represents the total number of Danshen (Salvia miltiorrhiza) slices samples (including stained and normal samples). To determine the number of core feature bands selected, the SPXY partitioning method was used to split the data into a training set and a validation set in a 7:3 ratio, so as to systematically train and verify the effectiveness and generalization ability of the core feature bands.
[0044] PCA modeling based on the training set involves automatically standardizing the input dataset using the following formula:
[0045] in, The original reflectance value under a certain wavelength band. This represents the mean reflectance of all samples within this wavelength band. This represents the standard deviation of the reflectance of all samples in this band. The standardized reflectance value is processed to make the mean of each characteristic band variable 0 and the standard deviation 1, thus eliminating the dimensional differences between different bands.
[0046] For the standardized data matrix (denoted as Principal component analysis (PCA) is performed. PCA projects the original high-dimensional data onto a new set of orthogonal low-dimensional basis vectors (called principal components, PCs) through a linear transformation, while retaining the most important variation information in the data.
[0047] Calculate the covariance matrix using the processed data matrix. Since the mean of each variable is 0 and the standard deviation is 1 after standardization, the covariance matrix and the correlation coefficient matrix are completely equivalent. The elements in the matrix reflect both the covariance between variables and directly represent the correlation coefficient between variables. Eigenvalue decomposition of the covariance matrix yields eigenvalues λ1≥λ2≥…≥λ n and the corresponding eigenvectors p1, p2, ..., p n The eigenvectors represent the principal component directions, and the eigenvalues characterize the variance contribution of the corresponding principal components. The projection of a sample onto the principal component space is achieved through a linear transformation, where the original eigenvector x of the i-th sample is... i The score (coordinates) on the k-th principal component is .
[0048] We define and retain K=20 principal components, which are the K directions with the largest variance contributions in the dataset. The contribution rate is calculated by recognizing the variance contribution percentage of each individual principal component. The first K principal components collectively contribute over 99.5% of the variation information in the original data, sufficient to characterize the core features of the data. The preprocessed data... Projecting the dimension dataset onto a new space composed of K principal components generates A score matrix of dimension , where each row of the matrix represents the coordinates of a sample in the principal component space.
[0049] To objectively evaluate the robustness and generalization ability of the constructed PCA model and avoid overfitting, a 10-fold cross-validation method was used to validate the model. The entire dataset was randomly divided into 10 mutually exclusive subsets (folds). One subset was used as the test set, and the remaining nine subsets were used as the training set to fit the PCA model. The projection error of the test set data was calculated, and the root mean square error of cross-validation (RMSEcv) was used to evaluate the model performance, as shown below:
[0050] Where 0.7m is the total number of samples in the dataset (2052 in this case). Let x be the test set for the i-th round of cross-validation, and let x be the original feature value (reflectance value after standardization) of the samples in the test set. The original values reconstructed from the training set model for the test samples (the original spatial approximations obtained by back-deriving the principal component scores) are repeated 10 times to ensure that each subset is used as a test set once. The calculated values are... = 0.00125945, an extremely low value, indicating that the model has high prediction accuracy, strong robustness, and no overfitting phenomenon.
[0051] Finally, principal component score maps are plotted based on the score matrix, and the coordinate positions of stained samples and normal samples in the principal component space are marked with different labels (such as different colors and shapes). The separation effect is output by observing the spatial distribution of the two types of samples in the score map. If the stained samples and normal samples form independent clusters in the principal component space, the two types of samples have no overlapping areas, and the correct recognition rate calculated by the confusion matrix is 100%, then the stained and normal Danshen slices are completely separated, that is, the feature band dataset has the ability to effectively distinguish between the two types of samples; otherwise, the separation effect is deemed unsatisfactory, and feature band screening needs to be repeated.
[0052] Step S3: Based on the core feature bands, extract pixel-level data from the hyperspectral image; use the rule of inward shrinking and coloring the visible boundary to perform pixel-level annotation in order to construct a training set and train a pixel-level classification model.
[0053] Pixel-level data from hyperspectral images are extracted based on the acquired core feature bands.
[0054] For each image of Salvia miltiorrhiza slices, it is divided into the following areas by labeling: the stained core area, which refers to the core part of the Salvia miltiorrhiza slice that has been illegally stained; the normal Salvia miltiorrhiza area, which refers to the normal tissue part of the Salvia miltiorrhiza slice that has not been stained; and the background area, which refers to the blank area in the image other than the Salvia miltiorrhiza slices.
[0055] Identify the edges of visible stained patches in the image. Using these edges as a reference, precisely shrink the area by 3-5 pixels along the image's interior direction. Use this shrunken boundary as the outline to delineate the region; this area is the core stained region. This rule avoids spectral mixing pixels at the boundary between stained and normal areas, ensuring the purity of the spectral characteristics of the core stained region.
[0056] Pixel-level spectral data extraction was performed on the three labeled regions: all pixels within each region were traversed, and the reflectance values of each pixel under K core feature bands were obtained to form the spectral feature vector of that pixel. A corresponding category label (label for stained core region, label for normal Salvia miltiorrhiza region, and label for background region) was added to each spectral feature vector. All labeled spectral feature vectors were integrated to construct a pixel-level training dataset. This dataset contains a feature matrix and label vectors. The feature matrix has a dimension of P×K, where P is the total number of extracted pixels, and the label vector has a dimension of P×1.
[0057] Select a machine learning classification algorithm (such as support vector machine, random forest, etc.) as the basic model framework, and initialize the model hyperparameters (such as the kernel function type of support vector machine, the number of decision trees of random forest, etc.) according to the size and feature dimension of the training dataset.
[0058] In this embodiment, for a single image, a support vector machine can be selected as the basic model framework. First, redundant bands are removed from the 301 bands of the hyperspectral data, retaining only 27 core feature bands. These 27 bands are then extracted as a new image. An entire image is selected as the training set. Three regions of interest (ROIs)—background, *Salvia miltiorrhiza*, and stained areas—are manually labeled. The separation degree of these three ROIs is calculated using the Jeffries-Matusita (JM) distance and Transformed Divergence (TD) formulas, respectively:
[0059]
[0060] If the separation is greater than 1.9, proceed to the next step of model parameter optimization. The optimal hyperparameters for the support vector machine are determined through grid search, with the kernel function being a radial basis function (RBF). The formula is as follows:
[0061] Where x and x' are the feature vectors of the two samples. , is the square of their Euclidean distance. The kernel parameter is set to 0.006. The penalty parameter (Penalty Para 0.7meter, C parameter) is used to balance classification error and model complexity, and is set to 1800.
[0062] For scenarios with a large number of images, a reusable workflow for visual SVM classification can be built using the ENVI Modeler+ machine learning plugin. The reusable workflow includes, for example: Figure 3 As shown, firstly, redundant bands in the 301 bands of the hyperspectral data are deleted, leaving only 27 core feature bands. These 27 bands are then extracted as a new image. In this way, 3 stained sample images and 1 normal sample image are prepared, along with the corresponding regions of interest (with the category regions labeled). For the stained sample, the regions of interest are manually labeled as background and stained. For the normal sample, the regions of interest are manually labeled as background and Salvia miltiorrhiza.
[0063] The constructed pixel-level training dataset is divided into training and validation subsets according to a preset ratio. The feature matrix and label vector of the training subset are input into the initialized model. Based on the spectral information of K core feature bands, the model learns the spectral feature differences of different categories of regions (stained core region, normal region, and background region). The model parameters are continuously adjusted through iterative calculations to gradually reduce the error between the model's classification prediction results for the training subset and the actual labels. During training, the model performance is evaluated in real time using the validation subset, with classification accuracy as the core evaluation metric. When the classification accuracy of the validation subset stabilizes and reaches a preset threshold, model training is stopped. Finally, a classification model that can be used for pixel-level recognition of staining in Danshen slices is obtained.
[0064] Step S4: Input the image data of the medicinal slices to be tested under the core feature band into the pixel-level classification model, perform pixel-by-pixel recognition, and obtain the classification results of stained pixels and normal pixels; calculate the proportion of staining contamination based on the classification results, and output the judgment conclusion according to the preset threshold.
[0065] The process iterates through all pixels in the feature image of the test sample, extracting the spectral feature vector of each pixel and inputting it into a pre-trained pixel-level classification model. This pixel-level classification model, based on the differences in spectral features between the stained core region, the normal *Salvia miltiorrhiza* region, and the background region learned during training, classifies the input pixel spectral feature vectors and outputs the corresponding classification label. The labels include three categories: "stained pixel," "normal *Salvia miltiorrhiza* pixel," and "background pixel." After the iteration is complete, a pixel-level classification result matrix covering the entire spatial range of the *Salvia miltiorrhiza* slice image is obtained. The matrix dimension is consistent with the spatial dimension of the cube of the feature image of the test sample, and each element in the matrix is the classification result label of the corresponding pixel.
[0066] The algorithm iterates through all elements in the result matrix, counting the total number of pixels labeled as stained pixels and the total number of pixels labeled as normal Danshen pixels, ignoring the counts of pixels labeled as background. The determination result is output by calculating the staining contamination ratio. Simultaneously, a visual staining distribution heatmap is generated based on the pixel-level classification result matrix, visually displaying the spatial distribution and clustering degree of stained and normal Danshen pixels. Finally, the binary determination conclusion, the specific numerical value of the staining contamination ratio, and the visual staining distribution heatmap are output simultaneously.
[0067] Furthermore, to verify the effectiveness of the present invention, the method of the present invention was used to perform verification on relevant datasets, and the verification results are as follows: Figure 4 As shown. Among them, Figure 4 Figure a shows the obtained hyperspectral data of stained Salvia miltiorrhiza. Figure 4 Figure b in the figure is a heat map of the staining distribution of stained Salvia miltiorrhiza. Combining the two figures, it can be seen that the method used in this invention can efficiently and accurately identify the location of stained Salvia miltiorrhiza.
[0068] Example 2 This embodiment discloses a non-destructive detection system for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging; like Figure 5 As shown, the non-destructive detection system for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging includes: The image data acquisition module is configured to acquire hyperspectral image data of the Chinese herbal medicine pieces to be tested and the reference sample using a hyperspectral imaging system; The feature band screening module is configured to: preprocess the collected hyperspectral data, extract the average spectrum of each medicinal slice to form a medicinal slice-level spectral dataset; use a competitive adaptive reweighting algorithm to screen out the core feature bands from the entire spectrum to distinguish between stained and normal categories, and analyze the medicinal slice-level spectral dataset using principal component analysis to output the separation effect between stained and normal Danshen medicinal slices; The classification model construction module is configured to: extract pixel-level data from the hyperspectral image based on the core feature bands; perform pixel-level annotation using the rule of inward shrinking and coloring visible boundaries to construct a training set and train a pixel-level classification model; The identification and judgment module is configured to: input the image data of the medicinal slices to be tested under the core feature band into the pixel-level classification model, perform pixel-by-pixel identification, and obtain the classification results of stained pixels and normal pixels; calculate the proportion of staining contamination based on the classification results, and output the judgment conclusion according to the preset threshold. Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0069] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for non-destructive detection of pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in Example 1.
[0070] Example 4 The purpose of this embodiment is to provide an electronic device.
[0071] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the non-destructive detection method for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging as described in Example 1.
[0072] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0073] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0074] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A non-destructive detection method for the colorant of Chinese herbal medicine Danshen based on hyperspectral imaging, characterized in that, include: Hyperspectral image data of the Chinese herbal medicine pieces to be tested and the reference sample were acquired using a hyperspectral imaging system. The collected hyperspectral data were preprocessed to extract the average spectrum of each medicinal slice to form a medicinal slice-level spectral dataset. A competitive adaptive reweighting algorithm was used to select the core feature bands from the entire spectrum to distinguish between stained and normal categories. The medicinal slice-level spectral dataset was analyzed by principal component analysis to output the separation effect between stained and normal Salvia miltiorrhiza slices. Based on the core feature bands, pixel-level data of the hyperspectral image is extracted; pixel-level annotation is performed using the rule of inward shrinking and coloring the visible boundary to construct a training set and train a pixel-level classification model. The image data of the medicinal slices to be tested under the core feature band is input into the pixel-level classification model for pixel-by-pixel identification to obtain the classification results of stained pixels and normal pixels; the proportion of staining contamination is calculated based on the classification results, and a judgment conclusion is output according to a preset threshold.
2. The hyperspectral imaging-based nondestructive detection method for the colorant of Chinese medicine Danshen as claimed in claim 1, characterized in that, Hyperspectral image data of the Chinese herbal medicine slices to be tested and the reference sample were acquired using a hyperspectral imaging system, including: The tested Danshen slices and the reference sample were laid flat on the stage in a dark room, and hyperspectral image data in a preset band range were acquired using a hyperspectral imaging system. The reference samples include the bark and cross-sectional samples of the Chinese herbal medicine slices, used to simultaneously monitor the spectral stability during the collection process.
3. The method for non-destructive detection of staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in claim 1, characterized in that, The collected hyperspectral data were preprocessed to extract the average spectrum of each medicinal slice to construct a medicinal slice-level spectral dataset, including: After performing black-and-white correction on the collected hyperspectral data, the corrected data is segmented by thresholding. The average spectral data of each medicinal slice is extracted by removing the background, thus forming a medicinal slice-level spectral dataset.
4. The method for non-destructive detection of staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in claim 1, characterized in that, A competitive adaptive reweighting algorithm was used to select core feature bands from the entire band to distinguish between stained and normal categories, including: The number of Monte Carlo sampling times was set, and K-fold cross-validation was adopted as the cross-validation strategy. The root mean square error was used as the evaluation index for variable selection. The maximum number of latent variables in the partial least squares model was set. The input herbal medicine slice-level spectral dataset is preprocessed by centering to eliminate the impact of baseline drift on feature selection; Based on the preprocessed herbal medicine slice spectral dataset, a preset number of Monte Carlo sampling iterations are performed to generate a frequency matrix, and the selected frequency of each band variable is calculated based on the frequency matrix. By setting a frequency threshold, the core characteristic bands used to distinguish between stained and normal Danshen are obtained.
5. The method for non-destructive detection of staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in claim 1, characterized in that, The analysis of the spectral dataset of processed medicinal herbs using principal component analysis, and the output of the separation effect between stained and normal Salvia miltiorrhiza slices, includes: The selected characteristic band spectral dataset of medicinal slices was used as input and automatically standardized preprocessed. A PCA model is built based on the preprocessed dataset. The spectral data in the original dimensional feature space is projected onto the low-dimensional basis vectors through linear transformation, and the principal components are set and preserved. The preprocessed dataset is then projected onto the principal component space to generate a score matrix. Principal component score maps are plotted based on the score matrix, and the coordinate positions of stained samples and normal samples in the principal component space are marked and the separation effect is output.
6. The method for non-destructive detection of staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in claim 1, characterized in that, Based on the core feature bands, pixel-level data of the hyperspectral image is extracted; Pixel-level annotation is performed using a rule of inward shrinking the visible boundary color to construct a training set, and a pixel-level classification model is trained, including: Based on the acquired core feature bands, pixel-level data of the hyperspectral image is extracted; For each image of Salvia miltiorrhiza slices, it was divided into a stained core area, a normal Salvia miltiorrhiza area, and a background area by labeling; Identify the visible edges of stained patches in the image, use these edges as a reference, shrink the pixel distance along the inside of the image, and use the shrunken boundary as the outline to delineate the region. Traverse all pixels within each region, obtain the reflectance value of each pixel under the core feature band, form the spectral feature vector of the pixel, integrate the spectral feature vectors to construct a pixel-level training dataset. The model was trained based on the constructed pixel-level training dataset to obtain a classification model for pixel-level identification of staining in Danshen slices.
7. The method for non-destructive detection of staining pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in claim 1, characterized in that, The image data of the medicinal slices to be tested under the core feature band is input into the pixel-level classification model for pixel-by-pixel identification to obtain the classification results of stained pixels and normal pixels. The percentage of staining contamination is calculated based on the classification results, and a judgment conclusion is output according to a preset threshold, including: Iterate through all pixels in the feature image of the sample to be tested and extract the spectral feature vector of each pixel one by one; The extracted spectral feature vectors are input into the trained pixel-level classification model. The pixel-level classification model judges the category of the input pixel spectral feature vectors based on the differences in spectral features of the stained core region, normal danshen region and background region learned during the training phase, and obtains a pixel-level classification result matrix covering the entire spatial range of the danshen slice image to be tested. Traverse all elements in the result matrix, count the total number of stained pixels and normal Danshen pixels, and output the judgment result by calculating the percentage of staining contamination.
8. A non-destructive detection system for staining pigments of the traditional Chinese medicine Danshen based on hyperspectral imaging, characterized in that: include: The image data acquisition module is configured to acquire hyperspectral image data of the Chinese herbal medicine pieces to be tested and the reference sample using a hyperspectral imaging system; The feature band screening module is configured to: preprocess the collected hyperspectral data, extract the average spectrum of each slice to form a slice-level spectral dataset; use a competitive adaptive reweighting algorithm to screen out the core feature bands from the entire spectrum to distinguish between stained and normal categories, and analyze the slice-level spectral dataset using principal component analysis to output the separation effect between stained and normal Salvia miltiorrhiza slices; The classification model construction module is configured to: extract pixel-level data from the hyperspectral image based on the core feature bands; perform pixel-level annotation using the rule of inward shrinking and coloring visible boundaries to construct a training set and train a pixel-level classification model; The identification and judgment module is configured to: input the image data of the medicinal slices to be tested under the core feature band into the pixel-level classification model, perform pixel-by-pixel identification, and obtain the classification results of stained pixels and normal pixels; calculate the proportion of staining contamination based on the classification results, and output the judgment conclusion according to the preset threshold.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the method for non-destructive detection of pigments in the traditional Chinese medicine Danshen based on hyperspectral imaging as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for non-destructive detection of pigments in traditional Chinese medicine Danshen based on hyperspectral imaging as described in any one of claims 1-7.