Traditional Chinese medicine classification method, electronic equipment and storage medium
Through electronic nasal and hyperspectral imaging technology combined with Dempster-Shafer theory and machine learning model, the time-consuming, cost and destructive problems of traditional Chinese medicine identification are solved, and the rapid and accurate classification of traditional Chinese medicine is achieved, which improves the objectivity and completeness of the identification.
Patent Information
- Application Number
- CN202510532668.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The identification methods of traditional Chinese medicine products in the prior art are time-consuming, costly and subjective, and require damage to sample integrity, making it difficult to achieve rapid, non-destructive and objective identification.
Electronic nasal and hyperspectral imaging technology are used to obtain information about traditional Chinese medicine, combined with Dempster-Shafer theory and machine learning model, and through the fusion of multiple classification models, BPA functions and combined reliability are calculated to achieve accurate classification of the properties of traditional Chinese medicine.
It improves the accuracy and reliability of Chinese medicine identification, reduces the deviation of a single model, realizes fast and non-destructive Chinese medicine classification, shortens the detection cycle, and avoids sample damage.
Smart Images

Figure CN120448965A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of traditional Chinese medicine identification, and in particular to a traditional Chinese medicine classification method, electronic equipment and storage medium. Background Art
[0002] There are a wide variety of traditional Chinese medicine products on the market. Currently, identification of these products, such as fermented Cordyceps sinensis powder, musk authenticity, origin, and age, relies primarily on manual experience or chromatographic techniques. These methods are time-consuming, costly, subjective, and require sample integrity. Therefore, a rapid, non-destructive, and relatively objective classification and identification method is needed. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a Chinese medicine classification method, electronic equipment and storage medium to improve the accuracy of Chinese medicine identification.
[0004] In a first aspect, the present application provides a method for classifying traditional Chinese medicine, comprising: obtaining an electronic nose response value and / or a hyperspectral image of a target traditional Chinese medicine; inputting the electronic nose response value and / or the hyperspectral image into at least two classification models to obtain classification results output by the at least two classification models, wherein the classification result characterizes the traditional Chinese medicine properties corresponding to the target traditional Chinese medicine, and the classification model is determined based on a sample data set corresponding to the sample traditional Chinese medicine, and the sample data set includes the electronic nose sample value and / or the hyperspectral sample image corresponding to the sample traditional Chinese medicine; based on the Dempster-Shafer theory and the classification results output by each classification model, calculating the value of the BPA function corresponding to each classification model; according to the value of the BPA function and the DS fusion rule, obtaining the joint reliability corresponding to each target traditional Chinese medicine property; based on the joint reliability, determining the final target traditional Chinese medicine property according to the reliability rule.
[0005] In the above scheme, electronic nose and hyperspectral imaging technology can obtain information about traditional Chinese medicine from different angles. Combined with the analysis of classification models, they can comprehensively reflect the multidimensional properties of traditional Chinese medicine and provide a more comprehensive basis for the quality control and quality evaluation of traditional Chinese medicine. During the detection process, multiple classification models trained and established based on a large number of sample data sets each have their own advantages in identifying and distinguishing the subtle differences between different target traditional Chinese medicines. By integrating the classification results of multiple classification models through the Dempster-Shafer theory, the possible deviations and errors of a single model can be effectively reduced, and the advantages and disadvantages of each model can be balanced, thereby improving the overall classification accuracy and reliability.
[0006] As an optional method, the value of the BPA function is calculated using the following formula: Wherein, i is the i-th classification model, I is the number of the classification models, j is the j-th target Chinese medicine property, and m is ij is the value of the BPA function, representing the reliability value of the i-th classification model for the j-th target Chinese medicine property, and the P ij is the recognition accuracy of the i-th classification model for the j-th target Chinese medicine property calculated in advance on the test set, and the R i The value of is 1 or 0, 1 represents that the classification result of the target Chinese medicine output by the i-th classification model is the j-th target Chinese medicine property, and 0 represents that the classification result of the target Chinese medicine output by the i-th classification model is not the j-th target Chinese medicine property.
[0007] In the above scheme, by using the value of the BPA function, the Dempster-Shafer theory is able to more comprehensively represent the uncertainty of information and provides an effective method to merge classification results from different classification models.
[0008] As an optional method, before obtaining the electronic nose response value and / or hyperspectral image of the target traditional Chinese medicine, the traditional Chinese medicine classification method also includes: training the machine learning model using the following steps to obtain the classification model: obtaining the sample data set corresponding to the sample traditional Chinese medicine; using the sample data set to train the machine learning model to obtain the classification model.
[0009] In this approach, a more accurate classification model can be constructed by training the machine learning model. The training process enables the model to learn the characteristic patterns of traditional Chinese medicines from a large amount of sample data, thereby enabling more accurate classification of target traditional Chinese medicines in practical applications.
[0010] As an optional manner, the obtaining of the sample data set corresponding to the sample Chinese medicine includes: if the sample data set includes the electronic nose sample value, obtaining the electronic nose sample value corresponding to the sample Chinese medicine, wherein the electronic nose sample value includes the electronic nose sample value corresponding to a plurality of gas sensors respectively; screening the effective sensors among the plurality of gas sensors, and the sample data set includes the electronic nose sample value corresponding to the effective sensors; if the sample data set includes the hyperspectral sample image, obtaining the hyperspectral sample image corresponding to the sample Chinese medicine, wherein the hyperspectral sample image includes the hyperspectral sample image corresponding to a plurality of spectral bands respectively; screening the effective sensors among the plurality of spectral bands an effective wavelength, the sample data set includes a hyperspectral sample image corresponding to the effective wavelength; if the sample data set includes the electronic nose sample value and the hyperspectral sample image, the electronic nose sample value and the hyperspectral sample image corresponding to the sample Chinese medicine are obtained, wherein the electronic nose sample value includes the electronic nose sample value corresponding to a plurality of gas sensors, and the hyperspectral sample image includes the hyperspectral sample image corresponding to a plurality of spectral bands; effective sensors among the plurality of gas sensors are screened, and effective wavelengths among the plurality of spectral bands are screened, the sample data set includes the electronic nose sample value corresponding to the effective sensor, and the hyperspectral sample image corresponding to the effective wavelength.
[0011] In the above scheme, by screening effective sensors and effective wavelengths, redundant and irrelevant data can be removed, noise and interference can be reduced, thereby improving the quality of the sample data set and enabling the model to learn the classification rules of traditional Chinese medicine more accurately.
[0012] As an optional method, the screening of effective sensors among the multiple gas sensors includes: calculating the importance of the characteristics of the sample Chinese medicine corresponding to each gas sensor; if the importance exceeds the predetermined threshold corresponding to the characteristic, determining that the gas sensor corresponding to the characteristic is the effective sensor; or, the screening of effective wavelengths among the multiple spectral bands includes: calculating the importance of the characteristics of the sample Chinese medicine corresponding to each spectral band; if the importance exceeds the predetermined threshold corresponding to the characteristic, determining that the spectral band corresponding to the characteristic includes the effective wavelength.
[0013] In the above scheme, by calculating the importance of features and screening out important features, it can ensure that the classification model focuses on the information that contributes most to the classification results, reduce the interference of noise and irrelevant features, and help improve the prediction accuracy of the model.
[0014] As an optional method, before using the sample data set to train the machine learning model to obtain the classification model, it also includes: preprocessing the electronic nose sample value according to the ratio of the maximum sample value of the electronic nose to the baseline value of the electronic nose in clean air.
[0015] In the above scheme, by performing a ratio processing on the maximum response value of the electronic nose and the baseline value in clean air, the baseline drift phenomenon that may occur in the sensor under different environmental conditions can be effectively eliminated.
[0016] As an optional method, before using the sample data set to train the machine learning model to obtain the classification model, it also includes: performing black and white correction on the hyperspectral sample image to obtain an intermediate hyperspectral image; extracting a region of interest based on the intermediate hyperspectral image; and determining a final hyperspectral sample image based on the region of interest.
[0017] In this approach, black-white correction eliminates the effects of uneven light source intensity distribution and camera dark current on hyperspectral images, improving image quality and making subsequent feature extraction and analysis more accurate. By extracting regions of interest, we can further focus on key areas of the sample, extracting more representative features.
[0018] As an optional manner, if the sample data set includes the electronic nose sample values and the hyperspectral sample images, before using the sample data set to train the machine learning model to obtain the classification model, the method further includes: normalizing the electronic nose sample values and the hyperspectral sample image data.
[0019] In the above scheme, normalization can ensure that data from different sources have similar numerical ranges, thereby preventing some data from having an uneven impact on the classification model due to a large value range.
[0020] In a second aspect, the present application provides an electronic device comprising: a processor, a memory and a bus, wherein the processor and the memory communicate with each other through the bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method steps of the first aspect.
[0021] In a third aspect, the present application provides a computer-readable storage medium, comprising: the computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method steps of the first aspect.
[0022] Other features and advantages of the present application will be described in the subsequent description, and in part will become apparent from the description, or will be understood by practicing the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A schematic diagram of a process for classifying traditional Chinese medicine provided in an embodiment of the present application;
[0025] Figure 2 A schematic diagram of the electronic device structure provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following embodiments of the technical solution of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present application and are therefore only examples and are not intended to limit the scope of protection of the present application.
[0027] It should be noted that all technical and scientific terms used herein have the same meanings as those commonly understood by technicians in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions.
[0028] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.
[0029] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0030] It can be understood that the traditional Chinese medicine classification method provided in the embodiment of the present application can be applied to a terminal device (also referred to as an electronic device) or a server; the terminal device can specifically be a smart phone, a tablet computer, a computer, a personal digital assistant (PDA), etc.; the server can specifically be an application server or a web server.
[0031] For ease of understanding, the technical solution provided in the embodiment of the present application is introduced below using the server as an example of the execution entity to introduce the application scenario of the traditional Chinese medicine classification method provided in the embodiment of the present application.
[0032] Traditional Chinese medicine products are diverse and complex in composition. Even the same type of Chinese medicine can be subdivided into different categories based on differences in properties or different processing techniques. Therefore, there are various quality differences and problems in both the production process and the Chinese medicine products entering the market. For example, precious and rare medicinal materials are prone to the flow of counterfeit and inferior products. Unscrupulous merchants pass off low-quality Chinese medicines as high-quality Chinese medicines. The endpoint cannot be accurately controlled during the product production process (for example, the fermentation degree of the same fermented Cordyceps powder is divided into insufficient fermentation, appropriate fermentation, and excessive fermentation). The above methods often rely on manual experience and have a certain degree of subjectivity. Traditional chromatographic identification methods are time-consuming, require sample destruction, and the results are not fully displayed, making the results less reliable.
[0033] Reference Figure 1 , Figure 1 A schematic flow chart of a method for classifying traditional Chinese medicine provided in an embodiment of the present application, the method comprising the following steps:
[0034] Step S10: obtaining the electronic nose response value and / or hyperspectral image of the target traditional Chinese medicine.
[0035] The target Chinese medicine is the same type of Chinese medicine that needs to be identified and classified, such as fermented Cordyceps sinensis powder or musk. The following is a possible application scenario of the Chinese medicine classification method of this application: classifying the main fermented Cordyceps sinensis products on the market. Known fermented Cordyceps sinensis products include Jinshuibao Capsules, Bailing Capsules, Ningxinbao Capsules, Zhiling Capsules, and Xinganbao Capsules. These fermented Cordyceps sinensis products are all fermented by different fungi. For example, Jinshuibao Capsules are CS-4. Therefore, the method provided in the embodiments of this application can be used to classify different fermented Cordyceps sinensis products.
[0036] An electronic nose simulates a biological olfactory system. Its basic principle is to use a gas sensor array to detect the odor or chemical substance of a target substance and convert the detected signal into an electrical output. Analysis of the electrical signal can rapidly detect and identify complex odors, identifying the overall characteristics of the odors of different substances. The electronic nose's gas sensor array contains multiple metal oxide sensors, each of which responds differently to different gases, allowing it to identify distinct odor characteristics.
[0037] Taking fermented Cordyceps sinensis powder as the target traditional Chinese medicine as an example, the electronic nose response value of fermented Cordyceps sinensis powder can be obtained by the following steps: take 5g of fermented Cordyceps sinensis powder, place it in a headspace sampling bottle, let it stand to allow the aroma to fully evaporate, and then use an electronic nose device to measure the response values of multiple metal oxide sensors, that is, the electronic nose response value.
[0038] Hyperspectral imaging technology integrates two-dimensional imaging and spectroscopy. Its hyperspectral data contains both spectral and image information, which can simultaneously characterize the external features and internal information of the object being measured. The following steps can be used to obtain a hyperspectral image of fermented Cordyceps sinensis powder:
[0039] Fermented Cordyceps sinensis powder was placed in a culture dish and spread until the bottom of the dish was no longer visible. Spectral image data of the sample was collected using a hyperspectral camera.
[0040] In order to classify the target Chinese medicine, only the electronic nose response value can be obtained, or only the hyperspectral image can be obtained, or both the electronic nose response value and the hyperspectral image can be obtained.
[0041] Step S20: Input the electronic nose response value and / or the hyperspectral image into at least two classification models to obtain classification results output by the at least two classification models, wherein the classification results characterize the properties of the target Chinese medicine corresponding to the target Chinese medicine, and the classification model is determined based on a sample data set corresponding to the sample Chinese medicine, and the sample data set includes the electronic nose sample value and / or the hyperspectral sample image corresponding to the sample Chinese medicine.
[0042] The electronic nose response values and / or hyperspectral images are fed as input into multiple established classification models with the same directional function. Each classification model outputs a corresponding classification result. Classification models can vary depending on their function. For example, a classification model for distinguishing genuine and fake musk outputs a true or false classification result; a classification model for identifying origin outputs a location classification result; and a classification model for identifying the fermented fungi in fermented Cordyceps sinensis powder products outputs the names of various fermented fungi. The function of the classification model and the input data must correspond. For example, a classification model for distinguishing genuine and fake musk requires the electronic nose response values and / or hyperspectral images of musk, not Cordyceps sinensis powder or other traditional Chinese medicines.
[0043] There are many algorithms used to build classification models. Each algorithm can be used to build a classification model. These algorithms include Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Tree (GBT), and others. Classification is primarily based on the differences in odor and chemical composition between traditional Chinese medicines. Electronic noses and hyperspectral imaging technologies detect and present these differences in data. Therefore, classification models built from electronic nose data and / or hyperspectral image data can accurately identify and classify traditional Chinese medicines.
[0044] If the electronic nose response value is obtained, two or more classification models can be used to classify the target Chinese medicine; if the hyperspectral image is obtained, two or more classification models can be used to classify the target Chinese medicine; if the electronic nose response value and hyperspectral image are obtained, two or more classification models can be used to classify the target Chinese medicine.
[0045] When training a classification model, there are three implementation methods based on the different training data: using only electronic nose sample values to train the classification model, using only hyperspectral sample images to train the classification model, and using both electronic nose sample values and hyperspectral sample images to train the classification model. It should be noted that when using a classification model, the type of input data must be consistent with the type of data used when training the model. For example, if electronic nose sample values are used to train the classification model, only electronic nose response values can be input to obtain classification results; if hyperspectral sample images are used to train the classification model, only hyperspectral images can be input to obtain classification results; and if both electronic nose sample values and hyperspectral sample images are used to train the classification model, only electronic nose response values and hyperspectral images can be input to obtain classification results.
[0046] Step S30: Based on the Dempster-Shafer theory and the classification results output by each classification model, the value of the BPA function corresponding to each classification model is calculated.
[0047] Because each classification model has its own strengths and weaknesses, the classification results obtained by multiple classification models may differ. When determining the properties of the final target traditional Chinese medicine from these multiple classification results, the classification results of each classification model can be integrated based on the Dempster-Shafer theory. The Dempster-Shafer theory (also known as evidence theory or trust function theory) is a mathematical framework for processing uncertain information, particularly suitable for situations where information is incomplete, ambiguous, or conflicting. When multiple independent classification models provide different classification results for the same target traditional Chinese medicine, Dempster's combination rule can be used to integrate these classification results to form a comprehensive trust assessment.
[0048] The reason why the Dempster-Shafer theory allows for a more extensive expression and processing of uncertainty is the Basic Probability Assignment (BPA). The Dempster-Shafer theory uses BPA to express the degree of support for a proposition or hypothesis. BPA can assign probabilities not only to a single event, but also to a group of events (i.e., a collection of propositions). The recognition accuracy of each classification model for each target Chinese herbal medicine property is obtained on the test set. Then, the multiple classification results of the output target Chinese herbal medicine are combined and fused through the full probability formula to obtain the support of each classification model for the classification result. After normalization, the value of the BPA function can be obtained, which represents the support (confidence value) of each classification model for each classification result.
[0049] The following example illustrates how to obtain the value of the BPA function. For example, i (i = 1, 2, 3) classification models classify target Chinese medicines with j (j = 1, 2, 3, 4, 5) target Chinese medicine properties. In advance, the feature vectors of the j target Chinese medicine properties are input into the i classification models on the test set. The recognition accuracy of the i classification models for the j target Chinese medicine properties is P ij Then, the electronic nose response value and / or hyperspectral image of the target Chinese medicine is input into i classification models, and the classification results output by i classification models are R i , R i =1 or R i =0, when R i =1, indicating that the classification result is j kinds of target Chinese medicine properties. i = 0, indicating that the classification result is not the j-type target Chinese medicine property, and then the support M of the i-type classification model for the j-type target Chinese medicine property is preliminarily obtained through the total probability formula ij , the specific formula is as follows:
[0050] M ij =P ij ×R i +(1-P ij )×(1-R i ).
[0051] Then, according to the characteristic that the sum of the confidences of the three classification models of BPA on the recognition frame power set is equal to 1, the above formula can be normalized to obtain the value of the BPA function.
[0052] Step S40: According to the value of the BPA function and the DS fusion rule, the joint reliability corresponding to each target Chinese medicine property is obtained.
[0053] The DS fusion rule is a method used in Dempster-Shafer evidence theory to combine BPAs from different information sources. The values of the BPA functions corresponding to each classification model are fused using the DS fusion rule to obtain the joint reliability of each target TCM property.
[0054] Step S50: Based on the joint reliability, the final target Chinese medicine property is determined according to the reliability rule.
[0055] The target Chinese herbal medicine properties are finally classified and identified according to the reliability rules. Specifically, if the joint reliability of the target Chinese herbal medicine properties finally identified is I, then I should satisfy the reliability rules. As an implementation method, the reliability rules can be:
[0056] 1): I is the maximum value of the joint reliability of all target Chinese medicine properties.
[0057] 2): The value of I must be greater than the threshold x.
[0058] 3): The difference between the joint reliability of I and the other target Chinese medicine properties is always greater than the threshold y.
[0059] If the confidence rule is not met, the final recognition result is "uncertain target Chinese medicine properties". The thresholds x and y are determined based on experiments.
[0060] In this approach, electronic noses and hyperspectral imaging technologies can acquire information about traditional Chinese medicines (TCMs) from different perspectives. The electronic nose focuses on the odor characteristics of the target TCM, while hyperspectral imaging focuses on its composition. These two technologies complement each other to form two dimensions for identifying the target TCM. Combined with classification model analysis, they can comprehensively reflect the multidimensional properties of TCMs, providing a more comprehensive basis for quality control and evaluation. During the detection process, multiple classification models trained on a large sample dataset each have their own strengths in identifying and distinguishing the subtle differences between target TCMs. Fusion of the classification results from multiple models using the Dempster-Shafer theory effectively reduces potential bias and error inherent in individual models, balances the strengths and weaknesses of each model, and thus improves overall classification accuracy and reliability. Both electronic noses and hyperspectral imaging technologies offer rapid response times, enabling rapid identification and classification of target TCMs, significantly shortening the experimental cycle required by traditional physical and chemical analysis methods. Furthermore, neither electronic nose nor hyperspectral imaging requires damage to the integrity of the target TCM, avoiding the damage typically associated with traditional methods, allowing the TCM to retain its original form and properties.
[0061] In some embodiments, the value of the BPA function is calculated using the following formula:
[0062]
[0063] Wherein, i is the i-th classification model, I is the number of the classification models, j is the j-th target Chinese medicine property, and m is ij is the value of the BPA function, representing the reliability value of the i-th classification model for the j-th target Chinese medicine property, and the P ij is the recognition accuracy of the i-th classification model for the j-th target Chinese medicine property calculated in advance on the test set, and the R i The value of is 1 or 0, 1 represents that the classification result of the target Chinese medicine output by the i-th classification model is the j-th target Chinese medicine property, and 0 represents that the classification result of the target Chinese medicine output by the i-th classification model is not the j-th target Chinese medicine property.
[0064] In the above scheme, by using the value of the BPA function, the Dempster-Shafer theory is able to more comprehensively represent the uncertainty of information and provides an effective method to merge classification results from different classification models.
[0065] In some embodiments, before step S10, the Chinese medicine classification method further includes:
[0066] Use the following steps to train the machine learning model to obtain a classification model:
[0067] Get the sample data set corresponding to the sample Chinese medicine.
[0068] The machine learning model is trained using the sample data set to obtain a classification model.
[0069] There are three ways to obtain sample datasets: obtaining electronic nose sample values, obtaining hyperspectral sample images, and obtaining both electronic nose sample values and hyperspectral sample images. Sample datasets obtained in any of these three ways can be used to build classification models.
[0070] If electronic nose sample values and hyperspectral sample images are obtained when the sample data set is obtained, both the electronic nose sample values and the hyperspectral sample images can be used together to train the classification model. This is a direct use of the two types of data. In addition, standards can be set to filter out parts of the electronic nose sample values and hyperspectral images for training the classification model.
[0071] In this approach, a more accurate classification model can be constructed by training the machine learning model. The training process enables the model to learn the characteristic patterns of traditional Chinese medicines from a large amount of sample data, thereby enabling more accurate classification of target traditional Chinese medicines in practical applications.
[0072] In some embodiments, obtaining a sample dataset corresponding to a sample Chinese medicine includes:
[0073] If the sample data set includes electronic nose sample values, obtain the electronic nose sample values corresponding to the medicine in the sample, wherein the electronic nose sample values include electronic nose sample values corresponding to multiple gas sensors respectively; select valid sensors from the multiple gas sensors, and the sample data set includes the electronic nose sample values corresponding to the valid sensors.
[0074] If the sample data set includes hyperspectral sample images, obtain the hyperspectral sample images corresponding to the sample Chinese medicine, wherein the hyperspectral sample images include hyperspectral sample images corresponding to multiple spectral bands respectively; filter the effective wavelengths in the multiple spectral bands, and the sample data set includes the hyperspectral sample images corresponding to the effective wavelengths.
[0075] If the sample data set includes electronic nose sample values and hyperspectral sample images, obtain the electronic nose sample values and hyperspectral sample images corresponding to the sample Chinese medicine, wherein the electronic nose sample values include electronic nose sample values corresponding to multiple gas sensors, and the hyperspectral sample images include hyperspectral sample images corresponding to multiple spectral bands; screen effective sensors from the multiple gas sensors, and screen effective wavelengths from the multiple spectral bands, the sample data set includes the electronic nose sample values corresponding to the effective sensors, and the hyperspectral sample images corresponding to the effective wavelengths.
[0076] In a sample dataset consisting of electronic nose sample values and / or hyperspectral sample images, not all sample data is required for subsequent model training. Sample data can also be screened based on contribution and importance, with sample data with greater contribution and importance selected for model training. Since the electronic nose sample values come from the electronic nose's metal oxide sensor, the final electronic nose sample values can be selected by screening effective sensors. Since the wavelengths of different bands in the hyperspectral sample image have a strong correlation with specific components and can be used to reveal information about these components, the final hyperspectral sample image can be selected by screening the effective wavelengths that are most representative of a specific component.
[0077] In the above scheme, by screening effective sensors and effective wavelengths, redundant and irrelevant data can be removed, noise and interference can be reduced, thereby improving the quality of the sample data set and enabling the model to learn the classification rules of traditional Chinese medicine more accurately.
[0078] In some embodiments, screening the plurality of gas sensors for valid sensors includes:
[0079] The importance of the characteristics of the medicinal materials in the samples corresponding to each gas sensor is calculated respectively.
[0080] If the importance exceeds a predetermined threshold corresponding to the feature, the gas sensor corresponding to the feature is determined to be a valid sensor.
[0081] Alternatively, screening for effective wavelengths in multiple spectral bands includes:
[0082] The importance of the characteristics of the sample Chinese medicine corresponding to each spectral band is calculated respectively.
[0083] If the importance exceeds a predetermined threshold corresponding to the feature, it is determined that the spectral band corresponding to the feature includes a valid wavelength.
[0084] The idea behind selecting effective sensors, or effective wavelengths, is to perform feature screening on sample data. As an implementation, each gas sensor in the electronic nose corresponds to a feature, and each spectral band in the hyperspectral image corresponds to a feature. The importance of each feature is calculated. When the importance of a feature exceeds a predetermined threshold, the sample data corresponding to that feature is selected for training the classification model; otherwise, it is deleted. The predetermined thresholds for electronic nose data and hyperspectral data are different, and the predetermined thresholds for different features are also different. The importance of each feature is calculated and screened one by one.
[0085] Feature selection can be achieved through algorithmic feature selection, but there is no single algorithm. Taking the Random Forest algorithm as an example, Random Forest uses out-of-bag (OOB) error to calculate the relative importance of feature variables and sort and filter them. OOB refers to instances that were not involved in the tree building process when building a single decision tree. The process of calculating the importance of feature f using Random Forest is as follows:
[0086] (a). Calculate the OOB error for the decision tree in the random forest, denoted as errOOB1.
[0087] (b) Randomly add noise interference to the features f of all OOB samples and calculate the out-of-bag error errOOB2 again.
[0088] (c) Assuming there are N trees in the random forest, the importance of feature f is agg(errOOB2-errOOB1) / N.
[0089] Based on the feature importance, the following steps can be used for feature selection:
[0090] (a). Calculate the importance of each feature and sort it in descending order.
[0091] (b) Determine the proportion to be eliminated, and eliminate the corresponding proportion of features based on feature importance to obtain a new feature set.
[0092] (c) Repeat the above process with the new feature set until m features remain.
[0093] As another implementation, feature screening can be achieved through feature dimensionality reduction. For example, data dimensionality can be reduced using principal component analysis, and principal components can be selected as features (ie, effective sensors, effective wavelengths).
[0094] The following example illustrates the effectiveness of feature filtering. For example, if 30 samples of traditional Chinese medicine are taken and the electronic nose has 14 metal oxide sensors, then there are 30 × 14 electronic nose sample values, and the hyperspectral image has 128 spectral bands, resulting in 30 × 128 hyperspectral sample image data. Without feature filtering, the sample data used for model training is 30 × (14 + 128). After feature filtering, the remaining electronic nose sample values may be 30 × 8, and the hyperspectral image sample data is 30 × 40, resulting in 30 × (8 + 40) sample data used for model training.
[0095] This application does not specifically limit the value of the predetermined threshold, and those skilled in the art can adjust it based on experience. For example, a cross-validation method can be used to find the optimal threshold. Specifically, the sample data set is combined into multiple groups of different training sets and test sets, and the training set is used to train the model. A sample in a training set may become a sample in the test set next time, which is the so-called "crossover". By reusing the data, a predetermined threshold is obtained. The predetermined threshold is obtained based on the data of the above-mentioned multiple training sets, and is essentially a threshold of the characteristics that can distinguish the target Chinese medicine. The predetermined thresholds for different target Chinese medicines are different. Even for the same target Chinese medicine, the predetermined threshold needs to be re-established for different purposes of the classification model.
[0096] In the above scheme, by calculating the importance of features and screening out important features, it can ensure that the classification model focuses on the information that contributes most to the classification results, reduce the interference of noise and irrelevant features, and help improve the prediction accuracy of the model.
[0097] In some embodiments, before training the machine learning model using the sample data set to obtain the classification model, the Chinese medicine classification method further includes:
[0098] The electronic nose sample values are preprocessed according to the ratio of the maximum sample value of the electronic nose to the baseline value of the electronic nose in clean air.
[0099] The maximum sample value of an electronic nose refers to the maximum value of the same gas sensor across multiple parallel measurements. For example, if the electronic nose has 14 gas sensors, 14 maximum sample values are determined across multiple parallel measurements. To account for interference from other gases in the detection environment, the following formula is used to preprocess the electronic nose sample values:
[0100]
[0101] Among them, H represents the electronic nose sample value after preprocessing, G represents the maximum sample value of the electronic nose, G1 represents the starting value of the gas sensor in the detection environment, and G0 represents the baseline value of the electronic nose in clean air.
[0102] In the above scheme, by performing a ratio processing on the maximum response value of the electronic nose and the baseline value in clean air, the baseline drift phenomenon that may occur in the sensor under different environmental conditions can be effectively eliminated.
[0103] In some embodiments, before training the machine learning model using the sample data set to obtain the classification model, the Chinese medicine classification method further includes:
[0104] The hyperspectral sample image is subjected to black and white correction to obtain an intermediate hyperspectral image.
[0105] Extract regions of interest based on intermediate hyperspectral images.
[0106] The final hyperspectral sample image is determined based on the region of interest.
[0107] In order to reduce the impact of uneven distribution of halogen linear light source and eliminate the noise caused by dark current of hyperspectral camera, it is necessary to perform black and white correction on the collected hyperspectral sample image. The following formula is used for black and white correction:
[0108]
[0109] Among them, R represents the intermediate hyperspectral image after black and white correction, R0 represents the original hyperspectral sample image obtained, and R W Represents a 100% white reflectance image, R D Represents a dark reflectance image of 0%.
[0110] Software is then used to extract a region of interest (ROI) from the intermediate hyperspectral image, and the average hyperspectral data of all pixels within the ROI is determined as the final hyperspectral sample image data. As an implementation, the intermediate hyperspectral image can be imported into ENVI 5.3 software, and the Region of Interest (ROI) tool in ENVI can be used to select a 100×100 pixel rectangular area as the ROI. When selecting the ROI, be careful to avoid the edge of the sample.
[0111] In this approach, black-white correction eliminates the effects of uneven light source intensity distribution and camera dark current on hyperspectral images, improving image quality and making subsequent feature extraction and analysis more accurate. By extracting regions of interest, we can further focus on key areas of the sample, extracting more representative features.
[0112] In some embodiments, if the sample dataset includes electronic nose sample values and hyperspectral sample images, before training the machine learning model using the sample dataset to obtain the classification model, the traditional Chinese medicine classification method further includes:
[0113] Normalize the electronic nose sample values and hyperspectral sample image data.
[0114] Since the dimensionality of electronic nose data is different from that of hyperspectral data, the data needs to be normalized. Normalization is essentially adjusting the scale of the data so that each feature can be compared at a unified scale. Normalization does not change the dimensionality of the data, but only eliminates the impact of different dimensions of different features, converting them into dimensionless form, so that each feature is within the same measurement range (for example, [0,1] or [-1,1]).
[0115] As an implementation method, the maximum and minimum method can be used for normalization. The specific formula is as follows:
[0116]
[0117] Among them, H n represents the normalized sample data, H represents the original sample data, and H min Represents the minimum value in the original sample data H, H max Represents the maximum value in the original sample data H.
[0118] In the above scheme, normalization can ensure that data from different sources have similar numerical ranges, thereby preventing some data from having an uneven impact on the classification model due to a large value range.
[0119] In some embodiments, after training the machine learning model using the sample data set to obtain the classification model, the method further includes:
[0120] The classification model was validated based on classification recognition accuracy and confusion matrix.
[0121] Classification recognition accuracy includes residual prediction deviation RPD, curve correlation coefficient R 2 wait.
[0122] The confusion matrix is a method for evaluating the prediction results of classification models in data analysis. Specific evaluation indicators include accuracy, sensitivity, and specificity. These indicators reflect the accuracy of classification from different aspects. Accuracy, sensitivity, and specificity are calculated using the following formula:
[0123]
[0124] Where TP is the number of true positive samples, TN is the number of true negative samples, FP is the number of false positive samples, and FN is the number of false negative samples. The confusion matrix has the true values in the top row, the predicted values in the columns, the correct classification results in the diagonal cells, the sensitivity and error rate at the bottom row, and the correct rate and false negative rate in the rightmost column.
[0125] The present application provides a computer program product, including computer program instructions. When the computer program instructions are read and executed by a processor, the methods provided by the above-mentioned method embodiments are executed.
[0126] The present application provides a computer-readable storage medium, including: a computer-readable storage medium storing computer instructions, wherein the computer instructions enable a computer to execute the methods provided by the above-mentioned method embodiments.
[0127] Figure 2 This is a schematic diagram of the electronic device structure provided in the embodiment of the present application, such as Figure 2 As shown, the electronic device includes: a processor 201, a memory 202 and a bus 203; wherein the processor 201 and the memory 202 communicate with each other through the bus 203; the memory 202 stores program instructions that can be executed by the processor 201, and the processor 201 calls the program instructions to execute the methods provided by the above-mentioned method embodiments.
[0128] The processor 201 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 201 can be a general-purpose processor, including a central processing unit (CPU), a micro control unit (MCU), a network processor (NP) or other conventional processors; it can also be a special-purpose processor, including a neural network processor (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. In addition, when there are multiple processors 201, some of them can be general-purpose processors and the other part can be special-purpose processors.
[0129] The memory 202 includes one or more (only one is shown in the figure), which may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The processor 201 and other possible components can access the memory 202 and read and / or write data therein.
[0130] In particular, one or more computer program instructions may be stored in the memory 202 , and the processor 201 may read and execute these computer program instructions to implement the weak password scanning behavior identification method provided in the embodiment of the present application.
[0131] Bus 203 includes one or more (only one is shown in the figure) and can be used to communicate directly or indirectly with other devices to exchange data. Bus 203 may include devices for wired and wireless communication, such as optical fibers, serial peripheral interface (SPI) modules, and integrated circuit buses (I2C). It may also include devices for wireless communication, such as Bluetooth modules, Wi-Fi modules, and mobile communication modules (e.g., 4G and 2G modules).
[0132] Understandably, Figure 2 The structure shown is only for illustration, and the electronic device may also include Figure 2 More or fewer components than shown, or with Figure 2 Different structures are shown. Figure 2 Each component shown in the figure can be implemented using hardware, software, or a combination thereof. The electronic device may be a physical device, such as a switch, router, server, or PC, or a virtual device, such as a virtual machine or virtualized container. Furthermore, the electronic device is not limited to a single device but may also be a combination of multiple devices or an integrated environment consisting of a large number of devices.
[0133] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. Those skilled in the art will appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for classifying traditional Chinese medicine, characterized in that: include: Obtaining electronic nose response values and / or hyperspectral images of target traditional Chinese medicine; Inputting the electronic nose response value and / or the hyperspectral image into at least two classification models to obtain classification results output by the at least two classification models, wherein the classification results characterize the properties of the target Chinese medicine corresponding to the target Chinese medicine, the classification models are determined based on a sample data set corresponding to the sample Chinese medicine, and the sample data set includes the electronic nose sample value and / or the hyperspectral sample image corresponding to the sample Chinese medicine; Based on the Dempster-Shafer theory and the classification results output by each classification model, the value of the BPA function corresponding to each classification model is calculated; According to the value of the BPA function and the DS fusion rule, the joint reliability corresponding to each target Chinese medicine property is obtained; Based on the joint reliability, the final target traditional Chinese medicine properties are determined according to the reliability rules.
2. The method according to claim 1, characterized in that The value of the BPA function is calculated using the following formula: Wherein, i is the i-th classification model, I is the number of the classification models, j is the j-th target Chinese medicine property, and m is ij is the value of the BPA function, representing the reliability value of the i-th classification model for the j-th target Chinese medicine property, and the P ij is the recognition accuracy of the i-th classification model for the j-th target Chinese medicine property calculated in advance on the test set, and the R i The value of is 1 or 0, 1 represents that the classification result of the target Chinese medicine output by the i-th classification model is the j-th target Chinese medicine property, and 0 represents that the classification result of the target Chinese medicine output by the i-th classification model is not the j-th target Chinese medicine property.
3. The method according to claim 1, characterized in that Before obtaining the electronic nose response value and / or hyperspectral image of the target traditional Chinese medicine, the traditional Chinese medicine classification method further includes: The machine learning model is trained using the following steps to obtain the classification model: Obtaining the sample data set corresponding to the sample traditional Chinese medicine; The machine learning model is trained using the sample data set to obtain the classification model.
4. The method according to claim 3, characterized in that The obtaining of the sample data set corresponding to the sample traditional Chinese medicine includes: If the sample data set includes the electronic nose sample value, obtaining the electronic nose sample value corresponding to the sample Chinese medicine, wherein the electronic nose sample value includes electronic nose sample values corresponding to a plurality of gas sensors respectively; screening effective sensors from the plurality of gas sensors, wherein the sample data set includes the electronic nose sample value corresponding to the effective sensor; If the sample data set includes the hyperspectral sample image, obtaining the hyperspectral sample image corresponding to the sample Chinese medicine, wherein the hyperspectral sample image includes hyperspectral sample images corresponding to a plurality of spectral bands; screening effective wavelengths in the plurality of spectral bands, wherein the sample data set includes the hyperspectral sample image corresponding to the effective wavelengths; If the sample data set includes the electronic nose sample values and the hyperspectral sample images, obtain the electronic nose sample values and the hyperspectral sample images corresponding to the sample Chinese medicine, wherein the electronic nose sample values include electronic nose sample values corresponding to multiple gas sensors, and the hyperspectral sample images include hyperspectral sample images corresponding to multiple spectral bands; screen effective sensors from the multiple gas sensors, and screen effective wavelengths from the multiple spectral bands, the sample data set including the electronic nose sample values corresponding to the effective sensors, and the hyperspectral sample images corresponding to the effective wavelengths.
5. The method according to claim 4, characterized in that The screening of effective sensors among the plurality of gas sensors comprises: respectively calculating the importance of the characteristics of the sample Chinese medicine corresponding to each of the gas sensors; If the importance exceeds a predetermined threshold corresponding to the feature, determining that the gas sensor corresponding to the feature is the valid sensor; Alternatively, the screening of effective wavelengths in the plurality of spectral bands comprises: Calculating the importance of the characteristics of the sample Chinese medicine corresponding to each of the spectral bands; If the importance exceeds a predetermined threshold corresponding to the feature, it is determined that the spectral band corresponding to the feature includes the effective wavelength.
6. The method according to any one of claims 3 to 5, characterized in that: Before the machine learning model is trained using the sample data set to obtain the classification model, the method further includes: The electronic nose sample value is preprocessed according to a ratio of a maximum sample value of the electronic nose to a baseline value of the electronic nose in clean air.
7. The method according to any one of claims 3 to 5, characterized in that: Before the machine learning model is trained using the sample data set to obtain the classification model, the method further includes: Performing black and white correction on the hyperspectral sample image to obtain an intermediate hyperspectral image; extracting a region of interest based on the intermediate hyperspectral image; A final hyperspectral sample image is determined according to the region of interest.
8. The method according to any one of claims 3 to 5, characterized in that: If the sample data set includes the electronic nose sample values and the hyperspectral sample images, before using the sample data set to train the machine learning model to obtain the classification model, the method further includes: The electronic nose sample values and the hyperspectral sample image data are normalized.
9. An electronic device, characterized in that: include: A processor, a memory, and a bus, wherein the processor and the memory communicate with each other via the bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method according to any one of claims 1 to 8 by calling the program instructions.
10. A computer-readable storage medium, characterized in that include: The computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method according to any one of claims 1 to 8.