Traditional Chinese medicine formula granule quality uniformity evaluation system based on machine learning
By constructing a spectral benchmark library and calculating a comprehensive anomaly index, the problem of identifying unknown anomalies in the quality control of traditional Chinese medicine formula granules was solved, enabling comprehensive assessment and early warning of the quality of traditional Chinese medicine formula granules, and improving the coverage and reliability of quality control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEZE INST OF FOOD & DRUG INSPECTION & TESTING
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, quantitative models based on near-infrared spectroscopy are difficult to identify unknown anomalies that emerge during the production process, which leads to potential risks in the quality control of traditional Chinese medicine formula granules. Traditional release models cannot effectively detect the confusion or anomalies of non-target medicinal materials.
A machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules was constructed, including a benchmark maintenance module, an anomaly detection module, a decision control module, a health early warning module, and an iterative update module. By constructing a spectral benchmark library, a comprehensive anomaly index and a drift index were calculated to achieve the detection of unknown anomalies and a comprehensive assessment of quality risks.
It enables a comprehensive assessment of the quality of traditional Chinese medicine formula granules, breaking through the limitations of traditional models, significantly improving the coverage and reliability of quality control, effectively detecting unknown anomalies, and ensuring the long-term applicability of quality evaluation through automated review and early warning mechanisms.
Smart Images

Figure CN121899071A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traditional Chinese medicine quality testing technology, and more specifically, to a machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules. Background Technology
[0002] Traditional Chinese medicine granules are a modern dosage form of traditional Chinese medicine made from single-herb decoction pieces through processes such as extraction, concentration, and drying. In its production chain, precise control of batch-to-batch quality uniformity is the cornerstone of ensuring its clinical efficacy and safety, and also the core bottleneck for promoting the industry towards standardization and intelligence.
[0003] In existing technologies, quantitative models based on near-infrared spectroscopy are commonly used to achieve rapid quality evaluation of traditional Chinese medicine (TCM) formula granules. This involves collecting qualified samples of TCM formula granules and measuring the content of preset key indicator components to train the release model. The release model judges whether the content of indicator components in the sample meets the standards. However, this model can only perceive the limited targets preset during training and is unable to detect unknown anomalies that emerge during production and were never seen during training. For example, if the release model uses the content of indicator components A and B as training targets, and a non-target herb C is accidentally mixed in with the product, as long as this mixing does not significantly affect the predicted content of the monitored indicator components A and B, even if the introduction of herb C has left clear abnormal characteristics on the spectrum, the release model will still misjudge such substantially abnormal samples as qualified and release them due to its fixed judgment logic. This poses a serious hidden danger to the quality and safety of TCM formula granule products.
[0004] In view of this, the present invention proposes a machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules to solve the above problems. Summary of the Invention
[0005] To overcome the above-mentioned defects of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules, comprising: a benchmark maintenance module, used to construct a spectral benchmark library containing characteristic data of known qualified samples, analyze the spectral benchmark library, and determine the distribution boundary of known qualified samples in the spectral feature space; The anomaly detection module is used to obtain the spectral feature vector and spectral reconstruction residual vector of the sample to be evaluated. Based on the spectral feature vector and spectral reconstruction residual vector of the sample to be evaluated, the relative deviation distance from the distribution boundary is analyzed to determine the comprehensive anomaly index of the sample to be evaluated. The decision control module is used to perform hierarchical decisions based on the comprehensive anomaly index, respond to the triggered secondary alarm, calculate the spectral characteristics of the sample to be evaluated and the information entropy of the spectral benchmark library, generate a recommended review list, and perform manual review based on the recommended review list; The health early warning module is used to analyze the current data distribution of all samples to be evaluated within the evaluation period and the historical data distribution of the spectral benchmark library, determine the comprehensive drift index, and further determine whether to generate a concept drift early warning signal based on the comprehensive drift index. The iterative update module updates the spectral reference library or distribution boundaries based on new qualified samples in response to concept drift warning signals or confirmed manual review conclusions.
[0006] Furthermore, a spectral benchmark library containing known qualified sample characteristic data is constructed, including: Near-infrared spectral data of known qualified samples from traditional Chinese medicine formula granules were obtained; principal component analysis was performed on the preprocessed near-infrared spectral data to obtain the principal component variance contribution rate, principal component loading matrix, principal component eigenvalues and principal component score matrix of each known qualified sample; Based on the top principal components whose cumulative variance contribution rate exceeds the set contribution threshold, extract the principal component scores corresponding to each known qualified sample from the principal component score matrix to determine the spectral feature vector; Based on the product of the principal component score matrix and the principal component loading matrix, the spectral data is reconstructed, the difference between the reconstructed spectral data and the preprocessed near-infrared spectral data is calculated, and the spectral reconstruction residual vector of each known qualified sample is generated. The spectral feature vectors, spectral reconstruction residual vectors, and principal component eigenvalues of each known qualified sample are summarized by timestamp to form a spectral benchmark library.
[0007] Further, the distribution boundaries of known qualified samples in the spectral feature space are determined, including: Based on the spectral feature vectors and principal component eigenvalues in the spectral benchmark library, the principal space variation statistic for each known qualified sample is calculated; based on the distribution of the principal space variation statistics for all known qualified samples, the principal space boundary threshold is determined. The residual space variation statistic of each known qualified sample is calculated based on the residual vector of spectral reconstruction. Based on the distribution of the residual space variation statistic of all known qualified samples, the residual space boundary threshold is determined. Based on the region jointly defined by the principal space boundary threshold and the residual space boundary threshold, the distribution boundary of the known qualified samples in the spectral feature space is determined.
[0008] Furthermore, obtain the comprehensive anomaly index of the sample to be evaluated, including: Principal component analysis was performed on the near-infrared spectral data of the samples to be evaluated to obtain the spectral feature vector and spectral reconstruction residual vector of the samples to be evaluated. Based on the principal component eigenvalues in the spectral reference library, the spectral feature vector of the sample to be evaluated is analyzed to obtain the principal space variation statistic of the sample to be evaluated; the ratio of the principal space variation statistic of the sample to be evaluated to the principal space boundary threshold is calculated to determine the principal space deviation. Calculate the sum of squares of the residual vectors of the spectral reconstruction of the sample to be evaluated, determine the residual space variability statistic, calculate the ratio of the residual space variability statistic of the sample to be evaluated to the residual space boundary threshold, and determine the residual space deviation. The deviation of the principal space and the deviation of the residual space are added together to determine the comprehensive anomaly index of the sample to be evaluated.
[0009] Furthermore, hierarchical decision-making is performed based on the comprehensive anomaly index, including: The comprehensive anomaly index of the sample to be evaluated is compared with the preset first threshold and second threshold. If the comprehensive anomaly index is less than or equal to the first threshold, the sample to be evaluated is determined to be of uniform quality and is released. If the comprehensive anomaly index is greater than the second threshold, a first-level alarm is triggered and the sample to be evaluated is rejected. If the comprehensive anomaly index is greater than the first threshold and less than or equal to the second threshold, a second-level alarm is triggered and the sample to be evaluated is marked as a key review sample.
[0010] Furthermore, a recommended review checklist is generated, including: From the preprocessed near-infrared spectral data of the key review samples, the spectral absorbance values of the key characteristic bands related to the main active ingredients were selected; Obtain the spectral absorbance values of all known qualified samples in the spectral reference library in the same key characteristic band; Calculate the probability distribution difference of spectral absorbance values in key characteristic bands between each key review sample and all known qualified samples in the spectral benchmark library, and quantify the probability distribution difference to determine the information entropy. The samples to be reviewed are sorted from highest to lowest based on their information entropy values, and a recommended review list is generated.
[0011] Furthermore, manual review is conducted based on the recommended review checklist, including: The recommended review list is assigned to the quality inspectors' work terminals. The quality inspectors manually review each key review sample according to the recommended review list and determine the review conclusion for each sample to be evaluated. The review conclusion includes qualified or unqualified.
[0012] Furthermore, the overall drift index is determined, including: Collect the spectral feature vectors of all samples to be evaluated within the evaluation period to form the current data distribution set; randomly select the same number of spectral feature vectors as the current data distribution set from the spectral benchmark library to form the historical data distribution set; After merging the current data distribution set with the historical data distribution set, principal component analysis is performed, and a two-dimensional projection plane is constructed based on the first two principal components obtained. Project the current data distribution set and the historical data distribution set onto a two-dimensional projection plane, and calculate the corresponding two-dimensional probability density distributions based on the projection coordinates; calculate the difference measure between the two two-dimensional probability densities to obtain the comprehensive drift index.
[0013] Furthermore, based on the comprehensive drift index, it is further determined whether to generate a concept drift warning signal, including: The comprehensive drift index is compared with the preset warning threshold. If the comprehensive drift index is greater than the warning threshold, it is determined that the distribution boundary has undergone concept drift. In response to the determination of concept drift, the current timestamp, comprehensive drift index and identification information of the samples to be evaluated in the current evaluation period are extracted and combined to generate a concept drift warning signal.
[0014] Furthermore, the spectral reference library or distribution boundaries are updated based on new qualified samples, including: The samples to be evaluated that were deemed qualified by manual review during the evaluation period were identified as new qualified samples. The spectral feature vector, spectral reconstruction residual vector, and principal component eigenvalues of the new qualified samples were extracted and expanded into the spectral benchmark library after being indexed by timestamp. In response to the concept drift warning signal, principal component analysis was re-executed based on the expanded spectral benchmark library to update the distribution boundary.
[0015] The technical effects and advantages of the machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules of this invention are as follows: 1. This invention establishes a quantifiable digital standard material for traditional Chinese medicine formula granules by constructing a spectral benchmark library containing characteristic data of known qualified samples and determining the distribution boundary; by calculating the deviation of the sample to be evaluated in the principal space and residual space and synthesizing them into a comprehensive anomaly index, a comprehensive assessment of quality risks is achieved; by performing hierarchical decision-making through the comprehensive anomaly index, both efficient release of qualified batches and precise identification of potentially risky batches are ensured, breaking through the limitation of traditional release models that can only identify preset target components, realizing effective detection of unknown anomalies, and significantly improving the coverage and reliability of quality control.
[0016] 2. This invention generates a recommended review list by calculating the spectral characteristics of the samples to be evaluated and the information entropy of the spectral reference library. It automatically prioritizes limited manual review resources to samples with the greatest difference from historical pass status and the highest uncertainty, significantly improving the relevance of the review and the efficiency of problem detection. By analyzing the current and historical data distributions of all samples to be evaluated within the evaluation period to determine the comprehensive drift index and issue early warnings, it can proactively detect background drift in the production system caused by changes in raw materials, processes, and environment. This realizes the transformation from "post-event inspection" to "pre-event early warning" and ensures the long-term applicability of the quality evaluation system. Attached Figure Description
[0017] Figure 1 This is a system principle diagram of the machine learning-based traditional Chinese medicine formula granule quality uniformity evaluation system of the present invention; Figure 2 A schematic diagram illustrating the process of determining the distribution boundary in this invention; Figure 3 This is a schematic diagram of the process for obtaining the comprehensive anomaly index according to the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example
[0019] Please see Figures 1-3 As shown in this embodiment, the main design contents of the machine learning-based traditional Chinese medicine formula granule quality uniformity evaluation system are as follows: Existing release models determine compliance by predicting whether the content of indicator components in the test sample meets the standards, for example, by using the content of indicator components A and B as training targets. This approach has a fundamental limitation: when anomalies not covered in the training set appear in the Chinese herbal medicine granules, such as "cleaning agent residue C" or "accidental confusion of non-target medicinal materials D" introduced during the production process, the model can only respond to the defect patterns of "indicator components A and B" because these anomaly patterns are irrelevant to the model's existing component content prediction task. It exhibits a systematic "blind spot" for anomalies outside the training targets.
[0020] Based on this, a machine learning-based system for evaluating the quality uniformity of traditional Chinese medicine formula granules is designed, including: The benchmark maintenance module is used to build a spectral benchmark library containing characteristic data of known qualified samples, analyze the spectral benchmark library, and determine the distribution boundaries of known qualified samples in the spectral feature space.
[0021] It should be explained that the known qualified samples are sufficiently representative and cover all qualified situations involving raw materials, processes, and equipment within the normal fluctuation range. For example, the known qualified samples should cover products collected from at least 24 consecutive production batches.
[0022] Construct a spectral benchmark library containing characteristic data of known qualified samples, including: Obtain near-infrared spectral data of known qualified samples from traditional Chinese medicine (TCM) formula granule samples. From routinely produced TCM formula granule products that have passed all legal inspections and been deemed qualified, select several samples (the samples must cover multiple independent production batches). Collect near-infrared spectral data of the known qualified samples using a near-infrared spectrometer. Spectral acquisition parameters can be set as follows: spectral range: 800~2500nm; resolution: 4nm; number of scans: 64. At least three acquisition points should be collected for each known qualified sample, and the average spectrum should be taken as the near-infrared spectral data of that known qualified sample.
[0023] It should be explained that each near-infrared spectral data needs to be preprocessed after acquisition: each near-infrared spectral data is subjected to standard normal variable transformation to eliminate the spectral baseline drift and amplification effect caused by known qualified sample particle size, distribution density and surface scattering; the spectral data after standard normal variable transformation is processed with the first derivative to eliminate the baseline translation and rotation shift of the near-infrared spectrum, while enhancing the absorption peak characteristics of the spectrum and improving the resolution.
[0024] Principal component analysis (PCA) was performed on the preprocessed near-infrared spectral data to obtain the principal component variance contribution rate, principal component loading matrix, principal component eigenvalues, and principal component score matrix for each known qualified sample. The preprocessed near-infrared spectral data of all known qualified samples were organized into a spectral data matrix, where each row represents a known qualified sample and each column represents the wavelength variable of the near-infrared spectrum. Centering was performed on each column of wavelength variables in the spectral data matrix (subtracting the arithmetic mean of the data in that column) to obtain a centered spectral data matrix. Singular value decomposition (SVD) was performed on the centered spectral data matrix to form three matrices: a left singular vector matrix, a singular value diagonal matrix, and the transpose of the right singular vector matrix.
[0025] The right singular vector matrix is defined as the principal component loading matrix representing the projection direction of the feature space. The diagonal elements of the singular value diagonal matrix are extracted, and the squares of these elements are calculated and divided by the number of known qualified samples minus one. The result is defined as the principal component eigenvalue, which measures the degree of variability in each principal component direction. The ratio of each principal component eigenvalue to the sum of all principal component eigenvalues is calculated, and this ratio is defined as the principal component variance contribution rate, which measures the importance of each principal component. The centered spectral data matrix and the principal component loading matrix are multiplied together, and the resulting projection matrix is defined as the principal component score matrix representing the coordinate positions of each known qualified sample in the feature space.
[0026] Based on the principal components whose cumulative variance contribution rate exceeds a set contribution threshold, the principal component scores corresponding to each known qualified sample are extracted from the principal component score matrix to determine the spectral feature vector. The spectral feature vector is used to characterize the core quality attributes of traditional Chinese medicine formula granules. The variance contribution rates of each principal component are arranged in descending order of their corresponding principal component feature values. Starting from the first item, the variance contribution rates of the arranged principal components are summed sequentially to obtain the corresponding cumulative variance contribution rate. This process continues until the cumulative variance contribution rate first reaches the set contribution threshold. The first k principal components required to achieve this cumulative variance contribution rate are recorded.
[0027] From the principal component score matrix, sequentially read the scores of the first k principal components corresponding to each known qualified sample. Combine the scores of the first k principal components corresponding to each known qualified sample into a k-dimensional vector in order; this k-dimensional vector is the spectral feature vector of the known qualified sample.
[0028] It's important to explain that the core purpose of setting the contribution threshold is to achieve an optimal balance between data dimensionality reduction and information preservation. Specifically, this means retaining as many of the main spectral features (effective variations) characterizing the chemical composition of known qualified samples as possible, while eliminating redundant information composed of high-frequency random noise and non-specific background interference. The specific value of the contribution threshold is determined based on a trade-off between long-term practical consensus in the field of spectral analysis and specific application needs. Commonly used contribution thresholds are set between 95% and 99%. For example, setting the contribution threshold to 95% means that the top k principal components selected collectively carry 95% of the total variation information in the original near-infrared spectral data.
[0029] The spectral data is reconstructed based on the product of the principal component score matrix and the principal component loading matrix. The difference between the reconstructed spectral data and the preprocessed near-infrared spectral data is calculated to generate the spectral reconstruction residual vector for each known qualified sample. The spectral reconstruction residual vector is used to characterize interference or anomalous information (such as trace impurities or unknown contamination) that is not explained by the core features, and is used to assist in the identification of hidden quality problems. Based on the first k principal components, the corresponding first k column vectors are extracted from the principal component loading matrix to form a reduced-dimensional principal component loading matrix. The principal component score matrix and the transpose of the reduced-dimensional principal component loading matrix are multiplied to obtain the reconstructed spectral data matrix. The difference between the preprocessed near-infrared spectral data matrix and the reconstructed spectral data matrix is calculated to obtain the residual data matrix. From the residual data matrix, the residual data at all wavelength variable points corresponding to each known qualified sample are extracted row by row, and each residual data is combined into a vector according to the original wavelength order; thus forming the spectral reconstruction residual vector for the corresponding known qualified sample.
[0030] The spectral feature vectors, spectral reconstruction residual vectors, and principal component eigenvalues of each known qualified sample are summarized by timestamp to form a spectral benchmark library.
[0031] Determine the distribution boundary of known qualified samples in the spectral feature space, including: Based on the spectral eigenvectors and principal component eigenvalues in the spectral reference library, the principal space variability statistic is calculated for each known qualified sample. The principal space variability statistic quantifies the deviation of a known qualified sample from the center of the known qualified sample matrix in the principal space; it measures how far a known qualified sample is from the center within the internal space of the "normal variability model." For each known qualified sample, the square of each element in its spectral eigenvector (i.e., the principal component score) is calculated. Each squared principal component score is then divided by its corresponding principal component eigenvalue. All the divided results are summed to obtain the principal space variability statistic.
[0032] Based on the distribution of the principal space variation statistics of all known qualified samples, the principal space boundary threshold is determined. The principal space boundary threshold characterizes the boundary of the allowable normal variation range of known qualified samples in the principal space. All principal space variation statistics are sorted, and their specific statistical percentiles are calculated. The calculated percentile values are used to determine the principal space boundary threshold.
[0033] For example, setting the principal space boundary threshold to the 99th percentile means that among all known qualified samples in the spectral reference library, 99% of the samples have a principal space variation statistic value lower than this principal space boundary threshold.
[0034] It should be explained that if the principal space variation statistic of a sample is lower than or equal to the principal space boundary threshold, it is considered "normal" in the principal space dimension; if the principal space variation statistic of a sample is higher than the principal space boundary threshold, it indicates that it has exceeded the acceptable range in terms of the main variation trend.
[0035] The residual space variability statistic for each known qualified sample is calculated based on the spectral reconstruction residual vector. The residual space boundary threshold is then determined based on the distribution of the residual space variability statistics for all known qualified samples. The residual space variability statistic quantifies the degree of variation of a known qualified sample within the residual space; the residual space boundary threshold defines the boundaries of acceptable normal noise or unexplained information for qualified samples within the residual space. For each known qualified sample, the squares of all elements in its spectral reconstruction residual vector are calculated, and the summation of these squares yields the residual space variability statistic for that known qualified sample. The residual space variability statistics for all known qualified samples are sorted in ascending order, and the statistical percentiles at the same confidence level as the residual space boundary threshold are calculated to determine the residual space boundary threshold.
[0036] It should be explained that if the residual space variation statistic of a sample is lower than or equal to the residual space boundary threshold, it is considered "normal" in the residual space dimension; if the residual space variation statistic of a sample is higher than the residual space boundary threshold, it indicates that it contains abnormal information that exceeds the normal level and cannot be explained, and may be an "unknown abnormal" sample.
[0037] Based on the region jointly defined by the principal space boundary threshold and the residual space boundary threshold, the distribution boundary of known qualified samples in the spectral feature space is determined. The principal space boundary threshold is used as an independent criterion: the principal space variation statistic of any sample must not exceed the principal space boundary threshold. The residual space boundary threshold is used as another independent criterion: the residual space variation statistic of any sample must not exceed the residual space boundary threshold. The condition that simultaneously does not exceed both the principal space boundary threshold and the residual space boundary threshold is defined as the distribution boundary of known qualified samples in the spectral feature space.
[0038] The anomaly detection module is used to obtain the spectral feature vector and spectral reconstruction residual vector of the sample to be evaluated. Based on the spectral feature vector and spectral reconstruction residual vector of the sample to be evaluated, the relative deviation distance from the distribution boundary is analyzed to determine the comprehensive anomaly index of the sample to be evaluated.
[0039] Obtain the comprehensive anomaly index of the sample to be evaluated, including: Principal component analysis was performed on the near-infrared spectral data of the samples to be evaluated to obtain the spectral feature vector and spectral reconstruction residual vector of the samples to be evaluated. Based on the principal component eigenvalues in the spectral benchmark library, the spectral feature vector of the samples to be evaluated was analyzed to obtain the principal space variation statistics of the samples to be evaluated.
[0040] It should be explained that the near-infrared spectral data acquisition of the sample to be evaluated was conducted under the same spectral acquisition parameters and conditions as those of the known qualified samples.
[0041] Calculate the ratio of the principal space variance statistic of the sample to be evaluated to the principal space boundary threshold to determine the principal space deviation. Principal space deviation = Principal space variance statistic of the sample to be evaluated ÷ Principal space boundary threshold; the principal space deviation is used to quantify the degree of variation of the sample to be evaluated within the "principal space", which is the proportion of deviation relative to the normal range of variation (principal space boundary threshold) allowed by known qualified samples.
[0042] It should be explained that when the principal space deviation is ≤1, it indicates that the principal space variation statistic of the sample to be evaluated is within the acceptable range, and the degree of variation of its core chemical characteristics is judged to be normal; when the principal space deviation is >1, it indicates that the principal space variation statistic of the sample to be evaluated has exceeded the acceptable range; and the larger the deviation value, the more serious the difference from the acceptable standard in terms of main chemical components and macroscopic characteristics.
[0043] Calculate the sum of squares of the residual vectors of the spectral reconstruction of the sample to be evaluated, determine the residual space variability statistic, calculate the ratio of the residual space variability statistic of the sample to be evaluated to the residual space boundary threshold, and determine the residual space deviation; residual space deviation = residual space variability statistic of the sample to be evaluated ÷ residual space boundary threshold; residual space deviation is used to quantify the amount of noise or anomalous information in the "residual space" of the sample to be evaluated, and the proportion of its deviation from the normal level (residual space boundary threshold) allowed by known qualified samples.
[0044] It should be explained that when the residual space deviation is ≤1, it indicates that the residual statistics of the sample to be evaluated are within the normal noise range and its spectral details are consistent with the known qualified sample library; when the residual space deviation is >1, it indicates that the sample to be evaluated contains abnormal information that exceeds the normal level and cannot be explained by the principal component model, and the larger the residual space deviation value, the higher the possibility of unknown contamination, confusion or special anomalies.
[0045] The overall anomaly index of the sample under evaluation is determined by adding the deviation of the principal space to the deviation of the residual space. The overall anomaly index measures the degree of deviation of a sample under evaluation from the acceptance standard.
[0046] The decision control module is used to perform hierarchical decisions based on the comprehensive anomaly index. In response to the triggered secondary alarm, it calculates the spectral characteristics of the sample to be evaluated and the information entropy of the spectral benchmark library, generates a recommended review list, and performs manual review based on the recommended review list.
[0047] Hierarchical decision-making based on comprehensive anomaly index includes: The comprehensive anomaly index of the sample to be evaluated is compared with a preset first threshold and a second threshold. If the comprehensive anomaly index is less than or equal to the first threshold, the sample to be evaluated is deemed to be of uniform quality and is released. The first threshold represents the maximum risk limit for automatic release; the second threshold represents the minimum risk limit for automatic rejection. Releasing the sample indicates that the spectral characteristics of the sample to be evaluated are highly consistent with the qualified sample library, and any minor fluctuations are entirely within the normal statistical variation range inherent in the production process.
[0048] If the overall anomaly index exceeds the second threshold, a Level 1 alarm is triggered, and the sample to be evaluated is rejected and released. The triggered Level 1 alarm automatically locks the identification information of the sample to be evaluated and generates a rejection instruction. The rejection instruction is sent to the release control terminal to execute the rejection and release operation, physically isolating the sample or marking it as a non-conforming product.
[0049] If the comprehensive anomaly index is greater than the first threshold and less than or equal to the second threshold, a level two alarm is triggered, and the sample to be evaluated is marked as a key sample for review.
[0050] It should be explained that the first threshold is set based on the distribution of the comprehensive anomaly index of historical qualified samples. For example, the 99th percentile of this distribution is usually taken directly. This ensures that the vast majority of qualified samples can be automatically released, thereby ensuring detection efficiency. The setting of the second threshold is an optimization process that comprehensively considers the control of the risk of false negatives and the frequency of false positives. The method is as follows: based on historical data, a critical value is found that can ensure that all known problematic batches are effectively identified (i.e., the comprehensive anomaly index is higher than this value). The second threshold is set at the value point that can highly intercept known problematic samples while having the lowest false positive rate.
[0051] Generate a recommended review checklist, including: From the preprocessed near-infrared spectral data of the key verification samples, the spectral absorbance values of the key characteristic bands related to the main active ingredients were selected. Key characteristic bands refer to the specific bands in the near-infrared spectral data that correspond to the characteristic absorption peaks that are significantly distinguishable and representative of the specific chemical functional groups of the main active ingredients (such as saponins, flavonoids, alkaloids, etc.) of the target traditional Chinese medicine formula granules; the spectral absorbance value represents the spectral intensity data of the specific band.
[0052] Obtain the spectral absorbance values of all known qualified samples in the spectral reference library at the same key feature bands. Read the preprocessed near-infrared spectral data of all known qualified samples, determine the set of wavelength points of the key feature bands that are the same as those of the key review samples, and extract the spectral absorbance values corresponding to the positions of the set of wavelength points of the key feature bands.
[0053] Calculate the difference in the probability distribution of the spectral absorbance values of each key review sample and all known qualified samples in the spectral reference library at the key feature bands, and quantify the difference in the probability distribution to determine the information entropy. The information entropy value is used to measure the "peculiarity degree" or "uncertainty degree" of a key review sample; normalize the spectral absorbance values of the key feature bands of the key review samples so that the sum of all absorbance values within the key feature bands is equal to 1, forming a spectral probability distribution to be evaluated that can represent the spectral energy distribution characteristics of the key review sample. Calculate the arithmetic mean of the spectral absorbance values of all known qualified samples at the same key feature bands to determine an average spectral curve representing the average level of the known qualified population, and perform the same normalization process on this average spectral curve to construct an average spectral probability distribution of the known qualified samples. Use the Jensen-Shannon divergence algorithm to calculate the relative entropy variant value between the spectral probability distribution to be evaluated and the average spectral probability distribution of the known qualified samples to determine the information entropy.
[0054] Sort the key review samples in descending order according to the information entropy values to generate a recommended review list.
[0055] It should be noted that the information entropy value is positively correlated with the degree of distribution difference between the key review samples and the spectral reference library. The higher the information entropy, the greater the difference in the spectral data of the corresponding key review samples at the key feature bands from the spectral data distribution of the spectral reference library, showing a high degree of "non - typicality", and the greatest uncertainty in its quality status, so it is given the highest review priority; the lower the information entropy, it indicates that the spectral data of the sample is more similar to the distribution of the spectral reference library, its "non - typicality" is weaker, the uncertainty of the quality status is relatively small, and the review priority is correspondingly lower.
[0056] Conduct manual review based on the recommended review list, including: Dispatch the recommended review list to the work terminals of quality inspectors. The quality inspectors conduct manual review on each key review sample according to the recommended review list to determine the review conclusions of each sample to be evaluated. The review conclusions include qualified or unqualified. The recommended review list includes the identification information, comprehensive anomaly index, and information entropy value of each key review sample; based on the identification information of each key review sample in the recommended review list, the quality inspector retrieves the physical samples corresponding to the key review samples and conducts laboratory tests according to legal standards (for example, using high - performance liquid chromatography to determine the content of specific active ingredients in the sample).
[0057] The laboratory test results are compared with the acceptable range specified in the statutory standards. If the laboratory test results fall within the acceptable range, the review conclusion of the key review sample is determined to be acceptable; if the laboratory test results exceed the acceptable range, the review conclusion of the key review sample is determined to be unacceptable.
[0058] The health early warning module is used to analyze the current data distribution of all samples to be evaluated within the evaluation period and the historical data distribution of the spectral benchmark library, determine the comprehensive drift index, and further determine whether to generate a concept drift early warning signal based on the comprehensive drift index.
[0059] Determine the overall drift index, including: Collect the spectral feature vectors of all samples to be evaluated within the evaluation period to form the current data distribution set. The current data distribution set is used to represent the overall quality data distribution of traditional Chinese medicine formula granule products under the current production status.
[0060] It should be explained that the duration of the evaluation period is determined by a fixed number of production batches to ensure that the current data distribution set always contains a sufficient number of samples, so as to stably represent the current state of the production process and avoid fluctuations or misjudgments in the calculation of the comprehensive drift index due to insufficient data. For example, the evaluation period can be set to the duration of 30 consecutive production batches.
[0061] Spectral feature vectors with the same number of samples as the current data distribution set are randomly selected from the spectral benchmark library to form a historical data distribution set; the historical data distribution set is used to represent the overall quality data distribution of qualified traditional Chinese medicine formula granule products under the best historical production conditions.
[0062] It should be explained that random selection is used to ensure that every known qualified sample in the spectral benchmark library has an equal chance of being selected, thus unbiasedly representing the "historical qualification status" described by the entire spectral benchmark library. Ensuring that the sample sizes of the current data distribution set and the historical data distribution set are consistent guarantees that the two sets have equal weights in the calculation, improving the accuracy and comparability of the difference measurement results.
[0063] Principal component analysis (PCA) is performed after merging the current and historical data distribution sets. A two-dimensional projection plane is constructed based on the first two principal components. All spectral eigenvectors from the current and historical data distribution sets are merged to form a joint data matrix. This joint data matrix is then mean-centered to ensure that the mean of each dimension is 0 and the variance is 1. The covariance matrix of the standardized joint data matrix is calculated, and eigenvalues and eigenvectors are obtained through eigenvalue decomposition. The eigenvalues are sorted in descending order, and the eigenvectors corresponding to the first two eigenvalues are selected as principal components (the first and second principal components). The eigenvectors corresponding to the first principal component are used as the X-axis projection direction, and the eigenvectors corresponding to the second principal component are used as the Y-axis projection direction, forming a two-dimensional projection plane.
[0064] Project the current data distribution set and the historical data distribution set onto a two-dimensional projection plane. Perform dot product operations on each spectral feature vector in the current data distribution set with the determined X-axis projection direction and Y-axis projection direction respectively to calculate all two-dimensional projection coordinate points of the current data distribution set on the two-dimensional projection plane; use the same dot product operation to calculate all two-dimensional projection coordinate points of the historical data distribution set on the two-dimensional projection plane.
[0065] The corresponding two-dimensional probability density distributions are calculated based on the projected coordinates. A measure of the difference between the two two-dimensional probability densities is calculated to obtain the comprehensive drift index. Two-dimensional kernel density estimation is performed on all two-dimensional projected coordinate points of the current data distribution set to construct the current two-dimensional probability density distribution describing the density of the current data distribution on the plane. Two-dimensional kernel density estimation is also performed on all two-dimensional projected coordinate points of the historical data distribution set to construct the historical two-dimensional probability density distribution. The Jensen-Shannon divergence between the current and historical two-dimensional probability density distributions is calculated, and the calculated Jensen-Shannon divergence value is determined as the comprehensive drift index.
[0066] Further judgment based on the comprehensive drift index is needed to determine whether a concept drift warning signal should be generated, including: The overall drift index is compared with a preset warning threshold. If the overall drift index is greater than the warning threshold, it is determined that the chemical background of the production system has shifted significantly, meaning that the distribution boundary of known qualified samples in the spectral feature space has undergone conceptual drift. If the overall drift index is less than or equal to the preset warning threshold, it is determined that the current data distribution is consistent with the historical data distribution, and no adjustment to the distribution boundary is required.
[0067] It needs to be explained that the method for setting the warning threshold is as follows: Select a period with known stable production and qualified quality as the baseline period; randomly divide the spectral feature vectors of all known qualified samples within the baseline period into multiple data pairs, each data pair containing a "simulated current set" and a "simulated historical set", and calculate the comprehensive drift index for each data pair; thereby obtaining a set of "background distributions" representing the normal fluctuations of the comprehensive drift index under stable production conditions; take a specific high quantile of this distribution (such as the 95th or 99th quantile) and set it as the warning threshold.
[0068] In response to the determination of concept drift, the current timestamp, comprehensive drift index and identification information of the samples to be evaluated in the current evaluation period are extracted and combined to generate a concept drift early warning signal.
[0069] The iterative update module updates the spectral reference library or distribution boundaries based on new qualified samples in response to concept drift warning signals or confirmed manual review conclusions.
[0070] Update the spectral reference library or distribution boundaries based on new qualified samples, including: Samples deemed qualified by manual review during the evaluation period are designated as new qualified samples. The spectral feature vectors, spectral reconstruction residual vectors, and principal component eigenvalues of these new qualified samples are extracted and expanded into the spectral benchmark library after being indexed by timestamps. In response to the concept drift warning signal, principal component analysis is re-executed based on the expanded spectral benchmark library to update the distribution boundaries.
[0071] In this embodiment, the present invention establishes a quantifiable digital standard material for traditional Chinese medicine formula granules by constructing a spectral benchmark library containing characteristic data of known qualified samples and determining the distribution boundary; by calculating the deviation of the sample to be evaluated in the principal space and residual space and synthesizing them into a comprehensive anomaly index, a comprehensive assessment of quality risk is achieved; by performing hierarchical decision-making through the comprehensive anomaly index, both efficient release of qualified batches and precise identification of potentially risky batches are ensured, breaking through the limitation of traditional release models that can only identify preset target components, realizing effective detection of "unknown anomalies", and significantly improving the coverage and reliability of quality control.
[0072] This invention generates a recommended review list by calculating the spectral characteristics of the samples to be evaluated and the information entropy of the spectral reference library. It automatically prioritizes limited manual review resources to samples with the greatest difference from historical pass status and the highest uncertainty, significantly improving the relevance of the review and the efficiency of problem detection. By analyzing the current and historical data distributions of all samples to be evaluated within the evaluation period to determine the comprehensive drift index and issue early warnings, it can proactively detect background drift in the production system caused by changes in raw materials, processes, and environment. This realizes the transformation from "post-event inspection" to "pre-event early warning" and ensures the long-term applicability of the quality evaluation system.
[0073] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0074] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only one method, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0075] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
[0076] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules, characterized in that, include: The benchmark maintenance module is used to build a spectral benchmark library containing characteristic data of known qualified samples, analyze the spectral benchmark library, and determine the distribution boundaries of known qualified samples in the spectral feature space. The anomaly detection module is used to obtain the spectral feature vector and spectral reconstruction residual vector of the sample to be evaluated. Based on the spectral feature vector and spectral reconstruction residual vector of the sample to be evaluated, the relative deviation distance from the distribution boundary is analyzed to determine the comprehensive anomaly index of the sample to be evaluated. The decision control module is used to perform hierarchical decisions based on the comprehensive anomaly index, respond to the triggered secondary alarm, calculate the spectral characteristics of the sample to be evaluated and the information entropy of the spectral benchmark library, generate a recommended review list, and perform manual review based on the recommended review list; The health early warning module is used to analyze the current data distribution of all samples to be evaluated within the evaluation period and the historical data distribution of the spectral benchmark library, determine the comprehensive drift index, and further determine whether to generate a concept drift early warning signal based on the comprehensive drift index. The iterative update module updates the spectral reference library or distribution boundaries based on new qualified samples in response to concept drift warning signals or confirmed manual review conclusions.
2. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 1, characterized in that, The construction of the spectral benchmark library containing known qualified sample feature data includes: Near-infrared spectral data of known qualified samples from traditional Chinese medicine formula granules were obtained; principal component analysis was performed on the preprocessed near-infrared spectral data to obtain the principal component variance contribution rate, principal component loading matrix, principal component eigenvalues and principal component score matrix of each known qualified sample; Based on the top principal components whose cumulative variance contribution rate exceeds the set contribution threshold, extract the principal component scores corresponding to each known qualified sample from the principal component score matrix to determine the spectral feature vector; Based on the product of the principal component score matrix and the principal component loading matrix, the spectral data is reconstructed, the difference between the reconstructed spectral data and the preprocessed near-infrared spectral data is calculated, and the spectral reconstruction residual vector of each known qualified sample is generated. The spectral feature vectors, spectral reconstruction residual vectors, and principal component eigenvalues of each known qualified sample are summarized by timestamp to form a spectral benchmark library.
3. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 2, characterized in that, Determining the distribution boundary of known qualified samples in the spectral feature space includes: Based on the spectral feature vectors and principal component eigenvalues in the spectral benchmark library, the principal space variation statistic for each known qualified sample is calculated; based on the distribution of the principal space variation statistics for all known qualified samples, the principal space boundary threshold is determined. The residual space variation statistic of each known qualified sample is calculated based on the residual vector of spectral reconstruction. Based on the distribution of the residual space variation statistic of all known qualified samples, the residual space boundary threshold is determined. Based on the region jointly defined by the principal space boundary threshold and the residual space boundary threshold, the distribution boundary of the known qualified samples in the spectral feature space is determined.
4. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 3, characterized in that, The process of obtaining the comprehensive anomaly index of the sample to be evaluated includes: Principal component analysis was performed on the near-infrared spectral data of the samples to be evaluated to obtain the spectral feature vector and spectral reconstruction residual vector of the samples to be evaluated. Based on the principal component eigenvalues in the spectral reference library, the spectral feature vector of the sample to be evaluated is analyzed to obtain the principal space variation statistic of the sample to be evaluated; the ratio of the principal space variation statistic of the sample to be evaluated to the principal space boundary threshold is calculated to determine the principal space deviation. Calculate the sum of squares of the residual vectors of the spectral reconstruction of the sample to be evaluated, determine the residual space variability statistic, calculate the ratio of the residual space variability statistic of the sample to be evaluated to the residual space boundary threshold, and determine the residual space deviation. The deviation of the principal space and the deviation of the residual space are added together to determine the comprehensive anomaly index of the sample to be evaluated.
5. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 4, characterized in that, The hierarchical decision-making based on the comprehensive anomaly index includes: The comprehensive anomaly index of the sample to be evaluated is compared with the preset first threshold and second threshold. If the comprehensive anomaly index is less than or equal to the first threshold, the sample to be evaluated is determined to be of uniform quality and is released. If the comprehensive anomaly index is greater than the second threshold, a first-level alarm is triggered and the sample to be evaluated is rejected. If the comprehensive anomaly index is greater than the first threshold and less than or equal to the second threshold, a second-level alarm is triggered and the sample to be evaluated is marked as a key review sample.
6. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 5, characterized in that, The generation of the recommended review list includes: From the preprocessed near-infrared spectral data of the key review samples, the spectral absorbance values of the key characteristic bands related to the main active ingredients were selected; Obtain the spectral absorbance values of all known qualified samples in the spectral reference library in the same key characteristic band; Calculate the probability distribution difference of spectral absorbance values in key characteristic bands between each key review sample and all known qualified samples in the spectral benchmark library, and quantify the probability distribution difference to determine the information entropy. The samples to be reviewed are sorted from highest to lowest based on their information entropy values, and a recommended review list is generated.
7. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 6, characterized in that, The manual review based on the recommended review checklist includes: The recommended review list is assigned to the quality inspectors' work terminals. The quality inspectors manually review each key review sample according to the recommended review list and determine the review conclusion for each sample to be evaluated. The review conclusion includes qualified or unqualified.
8. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 7, characterized in that, The determination of the comprehensive drift index includes: Collect the spectral feature vectors of all samples to be evaluated within the evaluation period to form the current data distribution set; randomly select the same number of spectral feature vectors as the current data distribution set from the spectral benchmark library to form the historical data distribution set; After merging the current data distribution set with the historical data distribution set, principal component analysis is performed, and a two-dimensional projection plane is constructed based on the first two principal components obtained. Project the current data distribution set and the historical data distribution set onto a two-dimensional projection plane, and calculate the corresponding two-dimensional probability density distributions based on the projection coordinates; calculate the difference measure between the two two-dimensional probability densities to obtain the comprehensive drift index.
9. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 8, characterized in that, The further determination based on the comprehensive drift index, including whether to generate a concept drift warning signal, includes: The comprehensive drift index is compared with the preset warning threshold. If the comprehensive drift index is greater than the warning threshold, it is determined that the distribution boundary has undergone concept drift. In response to the determination of concept drift, the current timestamp, comprehensive drift index and identification information of the samples to be evaluated in the current evaluation period are extracted and combined to generate a concept drift warning signal.
10. The machine learning-based evaluation system for the quality uniformity of traditional Chinese medicine formula granules according to claim 9, characterized in that, The updating of the spectral reference library or distribution boundary based on new qualified samples includes: The samples to be evaluated that were deemed qualified by manual review during the evaluation period were identified as new qualified samples. The spectral feature vector, spectral reconstruction residual vector, and principal component eigenvalues of the new qualified samples were extracted and expanded into the spectral benchmark library after being indexed by timestamp. In response to the concept drift warning signal, principal component analysis was re-executed based on the expanded spectral benchmark library to update the distribution boundary.