Method for identifying origin of genuine medicinal material based on endogenous nano property in combination with machine learning

By characterizing the intrinsic nanoparameters of Chinese herbal medicine extracts and using machine learning algorithms, the high cost and complexity of identifying the origin of Chinese herbal medicines are solved, and rapid and accurate identification of the origin of Chinese herbal medicines is achieved, which is applicable to a variety of Chinese herbal medicines.

CN120761342APending Publication Date: 2025-10-10HANGZHOU DIETOTHERAPY JINGYUAN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510825926.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The existing methods for identifying the origin of Chinese medicinal materials rely on high-cost and time-consuming biomarker testing, and are easily interfered with by counterfeiting, which makes identification more difficult and affects the development of the Chinese medicine industry.

Method used

By characterizing the intrinsic nanoparameters of Chinese medicine extracts such as light scattering intensity, distribution curve, photoelectric coefficient, and average particle size, and combining it with machine learning algorithms to build a classification model, accurate identification of the origin of Chinese medicinal materials can be achieved.

Benefits of technology

It realizes the rapid and accurate identification of the origin of Chinese medicinal materials, reduces the detection cost, improves the identification efficiency and accuracy, is applicable to a variety of Chinese medicinal materials, and is widely applicable and economical.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120761342A_ABST
    Figure CN120761342A_ABST
Patent Text Reader

Abstract

The invention provides a genuine medicinal material producing area identification method based on endogenous nanometer properties in combination with machine learning. The genuine medicinal material producing area identification method comprises the following steps: collecting standard substances of different producing areas and a sample to be detected, and performing pretreatment to obtain filtrate; respectively carrying out nanometer characteristic characterization on the obtained filtrate based on a dynamic light scattering measurement method to obtain nanometer characteristic parameters of standard products and samples to be detected from different producing areas; performing equation fitting on the obtained nanometer characteristic parameters of the standard products from the different producing areas by utilizing a machine learning algorithm to obtain a classification model for identifying the producing areas of the medicinal materials; using the obtained classification model to classify the sample to be detected and the standard products of different producing areas to obtain a classification result presented in the form of a confusion matrix, and judging the producing area of the sample to be detected according to the classification result. According to the method, various endogenous nano parameters of a traditional Chinese medicine extracting solution are represented, and statistical analysis and machine learning are combined, so that different producing areas of traditional Chinese medicinal materials are accurately distinguished.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of food detection, specifically relates to the field of traditional Chinese medicine origin identification, and more specifically relates to a method for authenticating genuine medicinal materials based on endogenous nano properties combined with machine learning. BACKGROUND

[0002] The quality and efficacy of traditional Chinese medicinal materials are largely influenced by their growth environment and origin. Due to differences in geographical climate, soil composition, and water conditions, the same type of traditional Chinese medicinal material from different origins will exhibit significant differences in chemical composition, physical properties, and other aspects. Therefore, the concept of "genuine medicinal materials" has emerged and become an important research and application basis in the field of traditional Chinese medicine. However, accurately identifying the origin of traditional Chinese medicinal materials still faces severe challenges. Currently, the main methods for identifying the origin of traditional Chinese medicinal materials include chemical composition analysis, DNA fingerprinting, and sensory identification. These methods rely on screening specific biomarkers and often require the use of sophisticated instruments such as high-performance liquid chromatography and high-throughput mass spectrometry, resulting in high detection costs, long time consumption, and complex operations. In addition, due to the increasing diversification of counterfeit methods, the behavior of adding additional, relatively low-value biomarkers to interfere with the detection results further increases the difficulty of identifying genuine medicinal materials, seriously affecting the healthy development of the traditional Chinese medicine industry.

[0003] The quality of traditional Chinese medicinal materials is closely related to the content of their chemical components such as proteins, polyphenols, and polysaccharides. However, when these components are used as biomarkers for quality control and origin identification, there are problems such as multiple screening objects, large screening workload, and high screening cost. SUMMARY

[0004] In order to solve the problems of multiple screening objects, large screening workload, and high screening cost in the prior art when using biomarkers to identify the origin of traditional Chinese medicinal materials, the present application provides a method for authenticating genuine medicinal materials based on endogenous nano properties combined with machine learning. This method characterizes multiple endogenous nano parameters of traditional Chinese medicine extract solutions such as light scattering intensity, distribution curve, photoelectric coefficient, average particle size, and polydispersity index, and combines statistical analysis and machine learning to accurately distinguish different origins of traditional Chinese medicinal materials.

[0005] The technical solution adopted by the present application is: a method for authenticating genuine medicinal materials based on endogenous nano properties combined with machine learning, comprising the following steps: S1. Collect standard samples from different origins and samples to be tested, heat them in water respectively, filter and cool them, then perform centrifugal treatment and membrane treatment, and take the filtrate; S2. Perform nano property characterization on the obtained filtrate based on dynamic light scattering measurement method, and obtain nano property parameters of the standard samples from different origins and the samples to be tested; S3. Use a machine learning algorithm to fit the nanoscale characteristic parameters of the obtained standards from different origins to obtain a classification model for identifying the origin of the medicinal material. S4. Use the obtained classification model to classify the test samples and standards from different origins, obtain the classification results in the form of a confusion matrix, and determine the origin of the test samples based on the classification results.

[0006] Research has found that the naturally formed nanoaggregates of chemical components in traditional Chinese medicines (such as proteins, polyphenols, and polysaccharides) during the processing of these herbs exhibit single-target, stable properties and specific nanoscale properties, such as particle size distribution, dispersibility, photoelectric coefficient, and light scattering intensity. These nanoaggregates are not only important carriers or manifestations of the active ingredients in traditional Chinese medicines, but also significantly influence their bioactivity, stability, and absorption characteristics. Therefore, in theory, these nanoscale properties can effectively reflect the intrinsic quality of traditional Chinese medicines and their origin, providing a new scientific basis for quality control and authenticity verification of these materials.

[0007] Therefore, this application, based on the typical properties of nanoaggregates in traditional Chinese medicine (TCM), incorporates artificial intelligence technology to propose an efficient, economical, and non-destructive solution for identifying the origin of TCM materials. First, nanoproperty detection eliminates the need for complex sample pretreatment, preserving the sample's original structure. This approach is particularly suitable for studying complex systems such as TCM extracts. By measuring key parameters of TCM extracts (such as light scattering intensity, particle size distribution, polydispersity index, photoelectric coefficient, etc.), the physical properties of the samples are extracted, providing a scientific basis for quality control and origin identification of TCM materials. Second, to improve the efficiency and accuracy of origin identification, this application selects multiple nanoproperty parameters and employs machine learning algorithms (such as random forests, support vector machines, gradient boosting algorithms, convolutional neural networks, and K-means clustering) for training and fitting analysis. Leveraging the advantages of machine learning in high-dimensional data processing and pattern recognition, key features can be extracted from complex parameters, enabling accurate classification and effective differentiation of TCM origins. This solution significantly reduces the reliance on complex sample pretreatment while ensuring the rapidity and non-destructive nature of the detection process.

[0008] Experimental results show that the classification accuracy of origin in this application is stable at over 95%, with good robustness and wide applicability. This technical approach not only provides a scientific, economical, and efficient solution for the quality control and origin traceability of Chinese medicinal materials, but also opens up new directions for intelligent analysis and precise identification in the field of Chinese medicine research. At the same time, this solution helps to further explore the potential relationship between the origin and intrinsic quality of Chinese medicinal materials, providing important support for the development and modernization of the Chinese medicine industry.

[0009] Preferably, the medicinal materials are selected from licorice, polygonatum, ginseng, pseudoginseng, wolfberry, yam, ganoderma lucidum, Panax notoginseng, three leaves green, astragalus, angelica, Chuanxiong, Atractylodes macrocephala, Codonopsis pilosula, Poria cocos, Dendrobium, American ginseng, Rhodiola rosea, Salvia miltiorrhiza, Cordyceps sinensis, Ophiopogon japonicus, chrysanthemum, forsythia, honeysuckle, Pinellia ternata, coix seed, lily, ginger, red date, mulberry, raspberry, Cistanche deserticola, Rehmannia root, Morinda officinalis, cassia seed, papaya, burdock, longan, black plum, mint, purple Any one of perilla, golden cherry fruit, bitter orange, sophora japonica seed, ginkgo, chamomile, hawthorn, jujube seed, cornus officinalis, saffron, angelica, psoralea corylifolia, fritillaria, Chinese clematis, oriental water plantain, ligusticum chuanxiong, rhodiola rosea, magnolia officinalis flower, motherwort, sabdariffa, kudingcha, golden buckwheat, golden tassel, green peel, magnolia officinalis flower, gastrodia elata, sterculia lychnophora, poria, patchouli, burdock root, plantain seed, and lotus seed core, more preferably any one of licorice, polygonatum, ginseng, wolfberry fruit, and yam.

[0010] Preferably, in step S1, the centrifugation conditions are centrifugation at 3000-10000 g for 10-30 min; and the membrane treatment conditions are ultrafiltration through a 10-100 kDa filter membrane at 3000-5000 g for 20-50 min. Step S1 also includes, before the heating step, pre-treating the standard and the sample to be tested, wherein the pre-treatment includes powdering, slicing, or slicing followed by powdering. The powder is sieved through a 60-80 mesh screen. The standard and the sample to be tested, obtained with or without pre-treatment, are heated in water at 80-100°C for 10-90 min.

[0011] Preferably, in step S2, the nano-characteristic parameters include one or more of light scattering intensity, distribution curve, photoelectric coefficient, average particle size, and polydispersity index, more preferably light scattering intensity, average particle size, and polydispersity index.

[0012] Preferably, in step S3, the machine learning algorithm includes one or more of random forest, support vector machine, gradient boosting algorithm, convolutional neural network, and K-means clustering, more preferably random forest algorithm.

[0013] Preferably, in step S3, the model data of the standard products from different origins are no less than 150.

[0014] Preferably, in step S3, the classification accuracy of the classification model is not less than 95%.

[0015] Beneficial effects of the present invention: Endogenous nanoparticles, as important physical and chemical property carriers in traditional Chinese medicine extracts, play a key role in the study of the quality and biological activity of traditional Chinese medicine. The particle size, dispersity and dynamic behavior of these particles not only directly affect the absorption and efficacy mechanism of traditional Chinese medicine, but also can reflect the origin characteristics and processing technology of traditional Chinese medicine. Based on these characteristics of endogenous nanoparticles, the present invention realizes the rapid and accurate identification of the origin of traditional Chinese medicine through dynamic light scattering technology, which has the significant advantages of simple operation, short detection time and high accuracy. The present invention uses light scattering intensity, distribution curve, photoelectric coefficient, average particle size and polydispersity index as key endogenous nano characteristic parameters, without the need for complex sample pretreatment, ensuring the non-destructiveness and objectivity of the method, while greatly reducing the influence of human subjective judgment, and the classification accuracy rate is more than 95%. In addition, this method is suitable for the identification of the origin of a variety of common traditional Chinese medicines, has wide applicability and economy, and can realize traceability analysis of mixed samples by establishing a database, providing a kind of efficient, economical and scientific solution for traditional Chinese medicine quality control and market supervision. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a principal component analysis (PCA) score diagram of licorice herbs from different regions in Example 1 of the present invention.

[0017] Figure 2 This is the machine learning confusion matrix (top) and area under the curve (AUC) (bottom) of licorice from different origins in Example 1 of the present invention.

[0018] Figure 3 This is the machine learning confusion matrix (top) and area under the curve (AUC) (bottom) of Polygonatum sibiricum from different origins in Example 2 of the present invention.

[0019] Figure 4 This is the machine learning confusion matrix (top) and area under the curve (AUC) (bottom) of ginseng from different origins in Example 3 of the present invention.

[0020] Figure 5 This is the machine learning confusion matrix (top) and area under the curve (AUC) (bottom) of wolfberries from different origins in Example 4 of the present invention.

[0021] Figure 6 This is the machine learning confusion matrix (top) and area under the curve (AUC) (bottom) of yam from different origins in Example 5 of the present invention. DETAILED DESCRIPTION

[0022] The following describes the embodiments of the present invention by specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the case of no conflict, the features in the following examples and embodiments can be combined with each other. In the embodiments of the present invention, unless otherwise specified, the methods used are all conventional methods, and the reagents used can be obtained from commercial sources.

[0023] Example 1: Identification method of liquorice from different origins S1. Pre-treatment of medicinal materials: Licorice standards were collected from various origins, including samples from Russia, Gansu, Ningxia, Inner Mongolia, Shanxi, and Xinjiang. Each licorice sample was washed, dried to constant weight, and then crushed into an 80-mesh sieved powder. The powder was washed in warm water and boiled with purified water at a ratio of 1:30 (w / v) for 60 minutes. The resulting herbal decoction was centrifuged at 5000 g for 30 minutes to remove solid residues. The supernatant was collected and ultrafiltered through a 100 kDa membrane at 3000 g for 30 minutes. The retained solution was washed twice with Milli-Q water and centrifuged again. The retained crystals were resuspended in an equal volume of Milli-Q water to obtain a licorice solution and stored at −20°C for further use.

[0024] S2. Nano-property characterization: The particle size, polydispersity index, and light scattering intensity of the liquorice solution were determined by dynamic light scattering measurement using a Zetasizer Nano ZS (Malvern Instruments Ltd, UK) at room temperature (25°C).

[0025] S3. Classification model construction: Random forest is used to fit the nano-property parameters of licorice standards from different origins to obtain a classification model for identifying the origin of medicinal materials; the model data of standards from different origins are no less than 150.

[0026] S4. Identify the origin: To further verify the accuracy of distinguishing licorice from different origins, this example also performed principal component analysis (PCA) on licorice from different origins. First, the data of the nano characteristic parameters were standardized. To eliminate unit differences, each column in the data matrix was standardized to zero mean and unit variance. The calculation formula is: ; Where: X is the original data, μ is the mean, and σ is the standard deviation.

[0027] A covariance matrix is ​​calculated for the standardized data to reflect the correlations between parameters. Eigenvalue decomposition is then performed on the covariance matrix to extract the eigenvalues ​​and corresponding eigenvectors. The directions of the principal components are determined by the eigenvectors, and the eigenvalues ​​reflect the contribution of each principal component to the data variance. Three-dimensional data are reduced to two dimensions using PC1 and PC2. The calculated PC1 and PC2 are used as the horizontal and vertical coordinates to plot a PCA distribution plot, with samples from different origins marked with different colors.

[0028] The results are as follows Figure 1 As shown, the contribution rate of PC1 is 66.2%, the contribution rate of PC2 is 31.3%, and the cumulative contribution rate of the two is 97.5%. In addition, licorice samples from different origins formed significant spatial distribution differences on PC1 and PC2, and licorice samples from Gansu and Xinjiang were clearly distinguished in the PC2 direction. Licorice samples from Inner Mongolia are far away from samples from other origins in both PC1 and PC2 directions. The classification results are consistent with the results in step S4. Further, the obtained classification model is used to classify standard samples from different origins, and the classification results presented in the form of a confusion matrix are obtained. The origin of the sample to be tested is determined based on the classification results. Figure 2 (Top) Shows the confusion matrix and pairwise significance analysis of the machine learning algorithm's results for distinguishing licorice from six different origins using particle size, polydispersity index, and light scattering intensity. Aside from a small number of prediction errors (<3%) between Shanxi and Ningxia licorice, the remaining results demonstrate the superiority of the machine learning algorithm and its high accuracy in distinguishing and classifying licorice. Figure 2 The AUC (area under the curve) results (below) demonstrate significant differentiation among various licorice varieties, particularly those from Russia, Gansu, and Inner Mongolia. This method leverages the efficiency and robustness of machine learning algorithms, providing a reliable means for rapid and efficient identification of licorice samples by origin.

[0029] Example 2: Identification Method of Polygonatum sibiricum from Different Origins Polygonatum sibiricum samples were collected from Jilin, Heilongjiang, Liaoning, Shandong, and South Korea. The samples were washed, dried to constant weight, and then crushed to pass through an 80-mesh sieve. The powder was washed in 40°C warm water and then soaked in pure water at a ratio of 1:50 (w / v) for 30 minutes. The mixture was then heated and boiled for 45 minutes. After cooling, the mixture was centrifuged at 8000g for 20 minutes to remove coarse impurities. The supernatant was collected and centrifuged at 3000g for 30 minutes using a 100 kDa ultrafiltration membrane. The retained solution was diluted 1:1 with Milli-Q water and then centrifuged and washed twice. The filtrate was resuspended in an equal volume of Milli-Q water to obtain a Polygonatum sibiricum solution, which was stored at -20°C until use. Subsequent procedures were similar to those in Example 1 for nano-property characterization, classification model construction, and origin identification. Figure 3(Top) shows the confusion matrix results generated by the machine learning algorithm and pairwise significance analysis of particle size, polydispersity index and light scattering intensity parameters, which clearly shows the classification effect of the model on five different origins of Polygonatum. AUC results ( Figure 3 (Right) First, it shows that all Polygonatum sibiricum from different origins were well distinguished (AUC > 0.9). Therefore, model validation showed that Polygonatum sibiricum classification accuracy was extremely high across all origins, demonstrating the effectiveness and stability of machine learning algorithms in identifying Polygonatum sibiricum origins.

[0030] Example 3: Identification Method for Ginseng from Different Origins Ginseng samples from Jilin, Heilongjiang, Liaoning, Henan, and Shandong were collected, peeled, sliced, dried at 50°C, and then pulverized through an 80-mesh sieve. The powder was washed with warm water, added to purified water at a 1:20 (w / v) ratio, and decocted in a 95°C waterbath for 70 minutes. After cooling, the decoction was centrifuged at 6000g for 25 minutes to remove the precipitate. The supernatant was passed through a 100 kDa ultrafiltration membrane at 3500g for 35 minutes. The retained solution was washed once with Milli-Q water and centrifuged again under the same conditions. The final solution was resuspended in Milli-Q water to the original volume and stored at -20°C as the ginseng extract. The samples were characterized for nanostructured properties, a classification model was constructed, and the origin was identified using the methods described in Example 1. Figure 4 (Top) Shows the confusion matrix generated by the machine learning algorithm based on particle size, polydispersity index, and light scattering intensity parameters, clearly demonstrating the model's ability to predict and classify ginseng from five different origins. Figure 4 The AUC model (shown below) also demonstrates a very high degree of discrimination between ginseng samples from different origins (>0.91). Model validation showed that the model achieved an overall accuracy of over 97% for ginseng classification, with good discrimination across samples from all origins.

[0031] Example 4: Method for identifying wolfberries from different origins Wolfberry samples were collected from Ningxia, Gansu, Qinghai, Xinjiang, and Shaanxi. The samples were washed and dried at 40°C. The dried wolfberries were directly heated at 80°C for 30 minutes in a 1:50 (w / v) ratio in pure water. After cooling, the dried wolfberries were centrifuged at 9000g for 15 minutes to remove coarse particles and pectin. The supernatant was filtered through a 100 kDa ultrafiltration membrane at 3000g for 45 minutes. The retained solution was washed twice with Milli-Q water and centrifuged again. The resulting solution was resuspended in the same volume of Milli-Q water to obtain a wolfberry solution, which was stored at -20°C until use. The samples were characterized for nanoscale properties, a classification model was constructed, and the origin was identified using the methods described in Example 1. Figure 5(Top) shows the confusion matrix generated by the machine learning algorithm based on these three parameters, clearly demonstrating the model's ability to classify goji berry samples from five different origins. Furthermore, the AUC model demonstrates a very high degree of differentiation between goji berries from different origins (>0.92). Model validation demonstrated that the model achieved classification accuracy exceeding 98%, significantly distinguishing goji berry samples from different origins.

[0032] Example 5: Identification Method for Chinese Yams from Different Origins Yam samples were collected from Henan, Shandong, Anhui, Jiangsu, and Hubei provinces. The yam was washed, sliced, dried to constant weight, and then ground through a 60-mesh sieve. The powder was decocted in pure water at a 1:100 (w / v) ratio for 10 minutes. After cooling, the decoction was centrifuged at 7000 g for 30 minutes to remove crude fiber and starch residues. The supernatant was ultrafiltered using a 10 kDa membrane at 4000 g for 30 minutes to remove large polysaccharides. The filtrate was diluted with Milli-Q water, washed twice, and centrifuged again. The final solution was resuspended in an equal volume of Milli-Q water and stored at -20°C as a yam solution. The samples were characterized for nanostructured properties, a classification model was constructed, and the origin was identified using the methods described in Example 1. Figure 6 (Top) Shows the confusion matrix results generated by the machine learning algorithm based on these three parameters, which clearly demonstrates the model's ability to classify yam samples from five different origins. Model validation shows that the model's classification accuracy is over 96%, and it performs well in samples from all origins. Figure 6 (Bottom) shows the AUC curve results, which shows that yams from different origins have very high discrimination (>0.9).

[0033] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection of the present invention.

Claims

1. A method for identifying the origin of authentic medicinal materials based on endogenous nano properties combined with machine learning, characterized in that: The following steps are involved: S1. Collect standards and test samples from different origins, heat them in water, filter and cool them, centrifuge them, filter them through a membrane, and collect the filtrate. S2. Characterize the nanostructures of the filtrates using dynamic light scattering to determine the nanostructure parameters of the standards and test samples from different origins. S3. Use a machine learning algorithm to fit the nanoscale characteristic parameters of the obtained standards from different origins to obtain a classification model for identifying the origin of the medicinal materials. S4. Use the obtained classification model to classify the test samples and standards from different origins, obtain the classification results in the form of a confusion matrix, and determine the origin of the test samples based on the classification results.

2. The method according to claim 1, wherein The medicinal materials are selected from licorice, polygonatum, ginseng, pseudoginseng, wolfberry, yam, ganoderma lucidum, Panax notoginseng, three leaves green, astragalus, angelica, Chuanxiong, Atractylodes macrocephala, Codonopsis pilosula, Poria cocos, Dendrobium, American ginseng, Rhodiola rosea, Salvia miltiorrhiza, Cordyceps sinensis, Ophiopogon japonicus, chrysanthemum, forsythia, honeysuckle, Pinellia ternata, coix seed, lily, ginger, red date, mulberry, raspberry, Cistanche deserticola, Rehmannia root, Morinda officinalis, Cassia seed, papaya, burdock, Any one of longan, black plum, mint, perilla, golden cherry fruit, fructus aurantii, sophora japonica seed, ginkgo, chamomile, hawthorn, jujube seed, cornus officinalis, saffron, angelica, psoralea corylifolia, fritillaria, chyranthes bidentata, oriental water plantain, ligusticum chuanxiong, rhodiola rosea, magnolia officinalis flower, motherwort, sabdariffa, kudingcha, golden buckwheat, golden tassel fruit, tangerine peel, magnolia officinalis flower, gastrodia elata, sterculia lychnophora, poria, patchouli, burdock root, plantain seed, lotus seed core.

3. The method according to claim 1, wherein In step S1, the centrifugation condition is 3000-10000 g for 10-30 min; the membrane treatment condition is ultrafiltration through a 10-100 kDa filter membrane at 3000-5000 g for 20-50 min.

4. The method according to claim 1, wherein In step S2, the nano-characteristic parameters include one or more of light scattering intensity, distribution curve, photoelectric coefficient, average particle size, and polydispersity index.

5. The method according to claim 1, wherein In step S3, the machine learning algorithm includes one or more of a random forest algorithm, a support vector machine, a gradient boosting algorithm, a convolutional neural network, and a K-means clustering algorithm.

6. The method according to claim 1, wherein In step S3, the model data of the standard products from different origins are no less than 150.

7. The method according to claim 1, wherein In step S3, the classification accuracy of the classification model is not less than 95%.