Traditional Chinese medicine functional substance determination method and electronic equipment

By constructing a multi-label learning framework for traditional Chinese medicine (TCM) components and efficacy, and using a multi-label feature selection algorithm to automatically identify key components of TCM, the problem of low efficiency and high cost in determining the efficacy substances of TCM is solved, and efficient and accurate identification of efficacy substances of TCM is achieved.

CN121812002APending Publication Date: 2026-04-07JIANGSU KANION PHARMA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies cannot quickly, accurately, and cost-effectively identify the active ingredients in traditional Chinese medicine. Traditional experimental methods are time-consuming and costly, and single-label machine learning methods are not suitable for multi-functional labeling scenarios of traditional Chinese medicine.

Method used

A multi-label learning framework for TCM-component-efficacy is constructed. The multi-label relevance feature selection algorithm (LRDG) is used to automatically identify key components, select the components most relevant to efficacy, and establish a mapping table of "TCM-component-efficacy".

Benefits of technology

It improves the accuracy of predicting the efficacy relationship of traditional Chinese medicine, can interpretively output the key components that determine efficacy, and realizes the systematic discovery of the efficacy substances of traditional Chinese medicine, which is efficient and low-cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121812002A_ABST
    Figure CN121812002A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medicine, and discloses a traditional Chinese medicine functional substance determination method and electronic equipment, the method comprises the following steps: obtaining a plurality of training samples, one training sample corresponding to one traditional Chinese medicine and a plurality of functional labels of the traditional Chinese medicine, the training sample comprising a plurality of components of the traditional Chinese medicine; evaluating the contribution degree of each component in the training sample to each efficacy label through a multi-label correlation feature selection algorithm to obtain a corresponding contribution degree; for each training sample, screening key components according to the contribution degree of each component to each efficacy label; and obtaining the prediction efficacy of each training sample by using the key components and the efficacy labels corresponding to the plurality of training samples, and outputting the key components corresponding to each prediction efficacy. The method can efficiently and accurately determine the functional substances of the traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medicine, in particular to a traditional Chinese medicine efficacy substance determination method and an electronic device. BACKGROUND

[0002] Traditional Chinese medicine produces efficacy by the joint action of multiple chemical components, and there may be synergistic or antagonistic effects between different components. Therefore, quickly and accurately identifying the key components (efficacy substance basis) that determine efficacy is an important direction of modernization research of traditional Chinese medicine. The discovery of efficacy component substances of traditional Chinese medicine currently mainly relies on experimental methods, and the efficacy contribution of different components is verified through component separation, pharmacological experiments and other methods. Such experimental methods are often time-consuming, high-cost, and difficult to conduct systematic research on a large-scale traditional Chinese medicine library. In order to improve efficiency and reduce cost, related technologies also propose to use artificial intelligence and machine learning methods to discover efficacy substances in traditional Chinese medicine, such as using classification models to predict the efficacy of traditional Chinese medicine, or using network analysis methods to construct a traditional Chinese medicine component-target-disease network, but such methods are not suitable for efficacy substance discovery of traditional Chinese medicine. SUMMARY

[0003] The present application provides a traditional Chinese medicine efficacy substance determination method and an electronic device to solve the problem of being unable to quickly and accurately determine the efficacy substances of traditional Chinese medicine at low cost.

[0004] In a first aspect, the present application provides a traditional Chinese medicine efficacy substance determination method, which comprises: Obtaining a plurality of training samples, one of the training samples corresponding to a traditional Chinese medicine and a plurality of efficacy labels of the traditional Chinese medicine, the training sample comprising a plurality of components of the traditional Chinese medicine; Evaluating the contribution of each component to each efficacy label in the training sample by a multi-label correlation feature selection algorithm to obtain the corresponding contribution; For each training sample, screening key components according to the contribution of each component to each efficacy label; Using the key components and the efficacy labels corresponding to the plurality of training samples to obtain the predicted efficacy of each training sample, and outputting the key components corresponding to each predicted efficacy.

[0005] In a second aspect, the present application provides a traditional Chinese medicine efficacy substance determination device, which comprises: A training sample acquisition module is configured to obtain a plurality of training samples, one of the training samples corresponding to a traditional Chinese medicine and a plurality of efficacy labels of the traditional Chinese medicine, the training sample comprising a plurality of components of the traditional Chinese medicine; A contribution evaluation module is configured to evaluate the contribution of each component to each efficacy label in the training sample by a multi-label correlation feature selection algorithm to obtain the corresponding contribution. The key component screening module is used to screen key components for each training sample according to the contribution of each component to each efficacy label. The prediction module is used to obtain the predicted power of each training sample by using the key components and power labels corresponding to the multiple training samples, and output the key components corresponding to each predicted power.

[0006] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method for determining the efficacy substances of traditional Chinese medicine as described in the first aspect or any corresponding embodiment.

[0007] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the method for determining the efficacy substances of traditional Chinese medicine as described in the first aspect or any corresponding embodiment.

[0008] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the method for determining the efficacy substances of traditional Chinese medicine described in the first aspect or any corresponding embodiment.

[0009] The method and electronic device for determining the efficacy substances of traditional Chinese medicine (TCM) provided in this invention construct a multi-label learning framework of "TCM-component-efficacy" and automatically identify the key components that determine efficacy using a multi-label feature selection algorithm. This not only improves the accuracy of predicting the TCM-efficacy relationship but also provides interpretable output of the key components that determine efficacy, supporting the systematic discovery of TCM efficacy substances. Furthermore, this embodiment does not rely on experimental methods, making it both efficient and low-cost. Attached Figure Description

[0010] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0011] Figure 1 This is a schematic flowchart of the first method for determining the effective substances of traditional Chinese medicine according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the second process of the method for determining the active ingredients of traditional Chinese medicine according to an embodiment of the present invention; Figure 3This is a structural block diagram of the device for determining the efficacy substances of traditional Chinese medicine according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0014] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0015] Determining the efficacy substances of traditional Chinese medicine using experimental methods is costly, time-consuming, and labor-intensive. However, the efficacy of traditional Chinese medicine often corresponds to multiple efficacy labels (such as tonifying qi, promoting blood circulation, clearing heat, etc.). Therefore, single-label machine learning methods or other single-label calculation methods are not suitable for determining the efficacy substances of traditional Chinese medicine.

[0016] According to an embodiment of the present invention, a method for determining the active ingredients of traditional Chinese medicine is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0017] This embodiment provides a method for determining the active ingredients of traditional Chinese medicine, which can be used in various electronic devices, such as mobile terminal devices and desktop computer devices. Figure 1 This is a flowchart of a method for determining the active ingredients of traditional Chinese medicine according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain multiple training samples, each training sample corresponding to a Chinese herbal medicine and multiple efficacy tags of the Chinese herbal medicine, and the training sample includes multiple components of the Chinese herbal medicine.

[0018] Specifically, a training dataset consisting of multiple training samples can be represented as: ,in, Let n be a sample matrix of Chinese medicine and its components, where n represents the number of Chinese medicines and d represents the number of components. If the nth component is a sample matrix of Chinese medicine and its components, then the nth component is a sample matrix of Chinese medicine and its components. The Chinese medicine contains the first Each component, then Otherwise, it is 0; It is a Chinese medicine efficacy label matrix, where m represents the number of efficacy labels. Indicates the first Chinese medicine There is a first One function, Indicates the first Chinese medicine No first One function.

[0019] Step S102 involves evaluating the contribution of each component in the training samples to each efficacy label using a Label-Dependency based RelevantFeature Discovery for multi-label learning (LRDG) algorithm, thereby obtaining the corresponding contribution. Here, the components are used as features.

[0020] In this embodiment, to screen out the key components most relevant to efficacy from the original component space, a multi-label relevance feature selection algorithm is introduced to evaluate the contribution of each component to all efficacy labels. In different training samples, the contribution of the same component to the same efficacy label is considered the same.

[0021] Step S103: For each training sample, key components are selected according to their contribution to each efficacy label.

[0022] Step S104: Using the multiple training samples and the corresponding key components and power labels, obtain the predicted power of each training sample, and output the key components corresponding to each predicted power.

[0023] This embodiment utilizes the correlation data between traditional Chinese medicine (TCM) and its efficacy and components. TCM is used as a sample, components as features, and efficacy as a label. A multi-label feature selection algorithm (LRDG) is introduced to screen the component features, select the set of key components most relevant to efficacy from the original component space, and then perform efficacy prediction based on this set.

[0024] In this embodiment, efficacy prediction is primarily used as a means of verifying the effectiveness of feature (i.e., component) selection. That is, if the selected components can stably predict the efficacy of traditional Chinese medicine, it indicates that these selected components possess strong and highly discriminative efficacy-related information, and therefore can be considered as candidate material bases for the corresponding efficacy. Specifically, this embodiment uses a multi-label relevance feature selection algorithm to evaluate the contribution of each component in the training samples to each efficacy label, obtaining the corresponding contribution value. This ensures that the efficacy predicted by the key components selected based on contribution value is as consistent as possible with the efficacy label.

[0025] Therefore, the method for determining the efficacy substances of traditional Chinese medicine (TCM) provided in this embodiment is a method based on multi-label feature (i.e., component) selection. By constructing a multi-label learning framework of "TCM-component-efficacy," it automatically identifies the key components that determine efficacy using a multi-label feature selection algorithm. This not only improves the accuracy of predicting the TCM-efficacy relationship but also interpretably outputs the key components that determine efficacy, providing support for the systematic discovery of TCM efficacy substances. Furthermore, this embodiment does not rely on experimental methods, making it both efficient and low-cost.

[0026] In some optional implementations, step S102, namely, evaluating the contribution of each component in the training samples to each efficacy label using a multi-label relevance feature selection algorithm, includes: Step S1021: For each training sample, construct a target pseudo-label generation function based on the components of the training sample. The target pseudo-label generation function is a function with the contribution of the components in the training sample to the efficacy label as a variable. Step S1022: Construct a target function based on the target pseudo-label generation function and the power labels corresponding to the training samples; Step S1023: Optimize and adjust the contribution variable in the objective function to minimize the deviation between the target pseudo-label obtained based on the target pseudo-label generation function and the corresponding efficacy label, thereby obtaining the contribution of each component to each efficacy label.

[0027] Due to the Chinese medicine – efficacy label matrix Typically, this is represented by binary labels, which can only depict the presence or absence of a certain effect, lacking continuity and fine-grained semantic expression, and making it difficult to fully utilize the potential correlation between effects. Therefore, this embodiment uses real labels... With ingredients An optimizable transition variable, a pseudo-label matrix, is introduced between them. This makes the label information richer and more continuous, and preserves the correlation between labels, thus providing a more reliable supervision signal for subsequent feature selection.

[0028] In some optional implementations, the objective function includes a first deviation term and a second deviation term; The first deviation term is used to calculate the deviation between the target pseudo-label obtained by the target pseudo-label generation function and the pseudo-label to be optimized, and the second deviation term is used to calculate the deviation between the pseudo-label to be optimized and the efficacy label.

[0029] Accordingly, step S1023, which is to optimize and adjust the contribution variable in the objective function so as to minimize the deviation between the target pseudo-label obtained based on the target pseudo-label generation function and the corresponding efficacy label, and to obtain the contribution of each component to each efficacy label, specifically involves: The contribution variables and pseudo-labels to be optimized in the objective function are optimized and adjusted to minimize the value of the objective function, thereby obtaining the contribution of each component to each efficacy label.

[0030] Among them, such as Figure 2 As shown, the target pseudo-label generation function can be a linear mapping function. Specifically, a linear regression model can be constructed to obtain the target pseudo-label. In this embodiment, a traditional Chinese medicine-ingredient sample matrix is ​​used. The formula for obtaining the target pseudo-label matrix is ​​expressed as follows: ,in, This is the contribution matrix of the component in the training samples to the efficacy label.

[0031] The formula for the first deviation term can be specifically expressed as follows: ,in, The pseudo-label to be optimized.

[0032] The formula for the second deviation term can be specifically expressed as follows: ,in, The power labels for the training samples. The pseudo-label to be optimized, This represents the weight of the second bias term. The second bias term is used to preserve pseudo-labels. With real labels To ensure that pseudo-labels retain the main semantics of genuine efficacy labels.

[0033] In some alternative embodiments, the objective function further includes a manifold structure constraint term for the pseudo-label; The manifold structure constraint term of the pseudo-label is used to calculate the product of the power similarity and the pseudo-label bias for each pair of training samples, where the power similarity is the power label similarity between the two training samples and the pseudo-label bias is the bias between the pseudo-labels of the two training samples.

[0034] Specifically, the formula for the manifold structure constraint term of the pseudo-label is expressed as:

[0035] in, For the first The pseudo-label matrix to be optimized for each training sample. For the first The pseudo-label matrix to be optimized for each training sample. For the first The training sample and the first The power label similarity matrix of each training sample. It's an efficacy label. The Laplace matrix, ; yes diagonal matrix .

[0036] As mentioned above, in this embodiment, minimizing the objective function is required when calculating the contribution. Minimizing the objective function requires minimizing the manifold structure constraint term of the pseudo-labels. Therefore, the manifold structure constraint term of the pseudo-labels ensures that two training samples with similar real efficacy labels (corresponding to two Chinese herbal medicines) also have similar pseudo-labels in the low-dimensional space; that is, the manifold structure of the pseudo-labels is consistent with the real efficacy labels. Here, "low-dimensional space" specifically refers to the continuous semantic space constructed through pseudo-label learning, with a dimension equal to the efficacy number m. Through manifold regularization constraints, it preserves the sample similarity defined by the original efficacy labels, thus becoming a more interpretable and robust representation space suitable for subsequent feature selection and classification tasks.

[0037] In this embodiment, an efficacy label similarity matrix is ​​introduced simultaneously. and its Laplace matrix This ensures that the manifold structure of pseudo-labels remains consistent with that of genuine efficacy labels.

[0038] In summary, when the objective function includes the first bias term, the second bias term, and the manifold structure constraint term of the pseudo-label, the objective function can be expressed as:

[0039] in, This is a regularization term used to prevent overfitting and improve generalization ability.

[0040] In some optional implementations, the objective function further includes a latent representation learning term, which is used to calculate the difference between the component similarity (also known as the similarity of traditional Chinese medicine samples) between every two training samples and a first product, wherein the first product is the product of the pseudo-label matrix to be optimized and the transpose of the pseudo-label matrix to be optimized.

[0041] In this embodiment, although pseudo-labels can enhance the information of true efficacy labels, they still do not fully utilize the component similarity between training samples (traditional Chinese medicine examples). In the original component space... In this context, the similarity between instances is easily affected by noise and sparsity. Evaluating the relevance between features (components) and labels solely based on the original features (components) may yield unstable results. Therefore, this embodiment learns the latent relevance between instances to obtain a more robust low-dimensional representation. It can both preserve the similarity between samples and serve as a pseudo-label to optimize feature selection, thereby improving robustness.

[0042] Specifically, the Symmetric Nonnegative Matrix Factorization (SNMF) method can be used to transform the component similarity matrix of the training samples (traditional Chinese medicine examples). Decomposed into ,this It can both preserve the similarity between samples and serve as a pseudo-label to optimize feature selection, thereby improving robustness.

[0043] In this embodiment, the latent representation learning term can be represented as: .

[0044] In some optional implementations, the component similarity between two training samples is calculated using a hot kernel function.

[0045] In this embodiment, the component similarity between traditional Chinese medicines is calculated using a heat kernel function to obtain a component similarity matrix. :

[0046] in, is a hyperparameter of the heat kernel function; For defined k nearest neighbor, For defined k nearest neighbor, yes The nearest neighbor parameter.

[0047] In some optional implementations, the objective function further includes a dynamic graph constraint term, which is used to constrain the contribution using a dynamic graph Laplacian matrix, which is obtained based on the similarity of the transposes of the pseudo-label matrices corresponding to the two training samples.

[0048] In related technologies, a fixed Laplacian matrix is ​​often used to approximate the feature manifold structure when constraining feature selection. However, a fixed graph cannot accurately reflect the true manifold structure of the data, is prone to oversimplification, and leads to the accumulation of training errors. Therefore, this embodiment utilizes the dynamic information obtained from the pseudo-label / latent representation learning to construct a graph structure that adapts to the data (i.e., a dynamic graph Laplacian), thereby constraining the feature weight matrix (i.e., the contribution matrix). In this way, the feature selection results can better preserve the manifold structure of the data, improving robustness and discriminativeness.

[0049] In some optional implementations, the expression for the dynamic graph constraint term is:

[0050] in, For the first The contribution matrix of each component of each training sample. For the first The contribution matrix of each component of each training sample. For the first The training sample and the first The transpose of the pseudo-label matrix to be optimized for each training sample. The similarity matrix, for The dynamic graph Laplacian matrix, , for A diagonal matrix.

[0051] In this embodiment, a dynamic graph Laplacian matrix is ​​constructed using the low-dimensional manifold structure of pseudo-labels. The feature weights (i.e., contribution) are constrained by this matrix.

[0052] Specifically, no. The training sample and the first The transpose of the pseudo-label matrix to be optimized for each training sample. Similarity matrix It can be calculated using the heat kernel function:

[0053] In this embodiment, the dynamic graph constraint can be used to introduce a regularization term into the feature (component) selection optimization objective function. Therefore, by minimizing The algorithm learns to match the component importance pattern with the distribution of efficacy labels: if a feature (i.e., a component) is uniquely associated with a key efficacy label (such as "artemisinin" and "antimalarial"), its weight (i.e., contribution) is strengthened; if a feature (i.e., a component) is associated with multiple strongly related efficacy labels, its weight (i.e., contribution) is suppressed due to dynamic graph constraints, thus avoiding redundancy. This adaptive adjustment mechanism significantly improves the accuracy of feature selection, especially in scenarios where labels frequently co-occur.

[0054] After completing the design of the pseudo-labels, latent representation learning terms, and dynamic graph constraints, a unified optimization objective function can be constructed to collaboratively model and jointly optimize the three parts of information. The resulting objective function is:

[0055] Among them, the first item Ensure that features (i.e., components) can generate pseudo-labels after linear mapping; the second term Ensure that pseudo-tags closely resemble real tags while preserving their original semantics; the third item Maintaining sample similarity through latent representation learning; fourth item The manifold structure of features (i.e., components) and labels is preserved through dynamic graph constraints.

[0056] In this embodiment, the number of training samples (n) significantly affects the parameters. , , The accuracy of the model is affected. When the sample size is small, the model is prone to overfitting, and the parameter estimates may be unstable, especially the feature weight matrix (i.e., the contribution matrix of each component). The reliability of the data will decrease, leading to biases in the selection of key components. As the sample size increases, the constraints in the objective function (such as regression terms and manifold regularization) will receive more sufficient data support, improving the robustness of parameter estimation. and More accurately capturing the relationship between ingredients and efficacy, while the pseudo-label matrix This also allows for better consistency with the actual efficacy structure. Therefore, sufficient training samples are an important prerequisite for ensuring that the model accurately identifies the active ingredients and achieves strong generalization ability.

[0057] To screen key substances highly correlated with efficacy from the characteristics of traditional Chinese medicine (TCM) components, this method proposes an optimization scheme based on nonnegative matrix factorization (NMF) and alternating iterative updates for the aforementioned objective function. This scheme can progressively approach the optimal solution while ensuring nonnegativity constraints and output component weights that can be used for feature screening. Specifically, firstly, based on the TCM-component matrix and the TCM-efficacy label matrix, the similarity information between TCM samples is calculated, and a fixed prior matrix (such as the TCM component similarity matrix) is obtained. Efficacy Similarity Matrix (etc.). Then, the feature weight matrix (i.e., the contribution matrix) is... and pseudo-label matrix Initialize as a non-negative random matrix.

[0058] In some optional implementations, the optimization and adjustment of the contribution variable and the pseudo-label to be optimized in the objective function includes: Fix one of the contribution variable and the pseudo-label to be optimized, and adjust the other of the contribution variable and the pseudo-label to be optimized, and adjust them alternately until convergence.

[0059] Since the objective function is composed of the pseudo-label matrix Non-convex functions and eigenweight matrices Composed of convex functions, direct joint optimization is difficult. Therefore, this embodiment employs a block coordinate descent approach, fixing one variable while optimizing another, iteratively updating alternately until convergence: when the variable is fixed... Optimize This allows the pseudo-labels to closely approximate the true efficacy labels while maintaining the manifold structure between samples; when fixed Optimize This allows the feature weight matrix to maximize the differentiation of key components with different effects. Lagrange multipliers and KKT conditions are introduced during the update process to ensure that the matrix remains consistent after each iteration. The non-negativity constraint is always satisfied to ensure the interpretability of the results (i.e., component weights will not have negative values). During the iteration process, as the objective function gradually converges, a stable feature weight matrix (i.e., contribution matrix) can eventually be obtained. .

[0060] In summary, in this embodiment, the data on traditional Chinese medicine components are characterized by high dimensionality and sparsity (many types of components but few effective components per sample) and strong correlation among multiple labels (such as "anti-inflammatory" and "immune regulation" often co-occurring). Traditional feature selection methods are prone to feature redundancy due to neglecting label correlation. To address this, this embodiment introduces a multi-label feature selection algorithm (LRDG) based on latent representation learning and dynamic graph constraints. Through the coupling of pseudo-label learning, latent representation learning, and dynamic graph constraints, it achieves accurate mining of component-efficacy association patterns and key feature screening.

[0061] Furthermore, after obtaining the contribution matrix of each component of each training sample to each efficacy label by optimizing the above objective function, for each training sample, when selecting key components according to the contribution of each component to each efficacy label (i.e., step S103 above), the corresponding key components (also called core components) are selected for each efficacy label. The specific steps are as follows: For the first... One efficacy label Extract the feature weight matrix (i.e., the contribution matrix). The corresponding column vector ,in Indicates the first Components ( ) for the first One efficacy label ( The contribution of ). For example Figure 2 As shown, the weight values ​​in this column vector are sorted in descending order, and the column vector with the largest weight value is selected first. These are the key components that constitute this efficacy. Traverse all Each efficacy label is used as a reference, and the above process is repeated to ultimately output a structured "efficacy-key ingredient" mapping table. This table directly reveals the most relevant material basis for each efficacy.

[0062] Regarding step S104 above, which is the step of obtaining the predicted power of each training sample using the key components and power labels corresponding to the multiple training samples, and outputting the key components corresponding to each predicted power, it can specifically be: After completing the screening of key components for each training sample (one training sample corresponds to one Chinese herbal medicine), the key component features obtained from the screening are utilized. As input, combined with the traditional Chinese medicine-efficacy tag matrix The multi-label K-Nearest Neighbors (ML-KNN) algorithm is used as a classifier to predict the efficacy of traditional Chinese medicine. Specifically, for each training sample... First, calculate its Euclidean distance with other samples in the training set and select... nearest neighbors The predicted probability of each efficacy label is calculated based on prior and conditional probabilities. Subsequently, binary predicted labels are generated using probability normalization and thresholding strategies. This allows for the joint prediction of multiple efficacy tags. The final output includes the predicted efficacy and a ranking of the key components corresponding to each efficacy, forming a knowledge base of efficacy substances in traditional Chinese medicine.

[0063] In summary, this embodiment establishes a high-dimensional and systematic predictive model between the components and efficacy of traditional Chinese medicine. For the scenario of traditional Chinese medicine with a large number of components and limited efficacy, the feature selection algorithm can clearly identify the components that play a key role in specific efficacy, providing direct evidence for the study of pharmacodynamic mechanisms.

[0064] This embodiment also provides a device for determining the active ingredients of traditional Chinese medicine. This device is used to implement the above embodiments and preferred embodiments, and will not be repeated for details already described. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0065] This embodiment provides a device for determining the active ingredients of traditional Chinese medicine, such as... Figure 3 As shown, it includes: The training sample acquisition module 301 is used to acquire multiple training samples, each training sample corresponding to a Chinese herbal medicine and multiple efficacy tags of the Chinese herbal medicine, and the training sample includes multiple components of the Chinese herbal medicine; The contribution evaluation module 302 is used to evaluate the contribution of each component in the training sample to each efficacy label through a multi-label correlation feature selection algorithm, and obtain the corresponding contribution. The key component screening module 303 is used to screen key components for each training sample according to the contribution of each component to each efficacy label. The prediction module 304 is used to obtain the predicted power of each training sample by using the key components and power labels corresponding to the multiple training samples, and output the key components corresponding to each predicted power.

[0066] In some optional embodiments, the contribution evaluation module 302 includes: A pseudo-label generation function construction unit is used to construct a target pseudo-label generation function based on the components of the training samples for each training sample. The target pseudo-label generation function is a function with the contribution of the components in the training samples to the efficacy label as a variable. The objective function construction unit is used to construct an objective function based on the objective pseudo-label generation function and the power labels corresponding to the training samples; The optimization and adjustment unit is used to optimize and adjust the contribution variables in the objective function so that the deviation between the target pseudo-label obtained based on the target pseudo-label generation function and the corresponding efficacy label is minimized, thereby obtaining the contribution of each component to each efficacy label.

[0067] In some optional implementations, the objective function includes a first deviation term and a second deviation term; The first deviation term is used to calculate the deviation between the target pseudo-label obtained by the target pseudo-label generation function and the pseudo-label to be optimized, and the second deviation term is used to calculate the deviation between the pseudo-label to be optimized and the efficacy label; The optimization and adjustment of the contribution variable in the objective function to minimize the deviation between the target pseudo-label obtained based on the target pseudo-label generation function and the corresponding efficacy label, thereby obtaining the contribution of each component to each efficacy label, includes: The contribution variables and pseudo-labels to be optimized in the objective function are optimized and adjusted to minimize the value of the objective function, thereby obtaining the contribution of each component to each efficacy label.

[0068] In some optional implementations, the objective function may further include a manifold structure constraint term for the pseudo-label; The manifold structure constraint term of the pseudo-label is used to calculate the product of the power similarity and the pseudo-label bias for each pair of training samples, where the power similarity is the power label similarity between the two training samples and the pseudo-label bias is the bias between the pseudo-labels of the two training samples.

[0069] In some optional implementations, the objective function further includes a latent representation learning term, which is used to calculate the difference between the component similarity between every two training samples and a first product, wherein the first product is the product of the pseudo-label matrix to be optimized and the transpose of the pseudo-label matrix to be optimized.

[0070] In some optional implementations, the component similarity between two training samples is calculated using a hot kernel function.

[0071] In some optional implementations, the objective function further includes a dynamic graph constraint term, which is used to constrain the contribution using a dynamic graph Laplacian matrix, which is obtained based on the similarity of the transposes of the pseudo-label matrices corresponding to the two training samples.

[0072] In some optional implementations, the expression for the dynamic graph constraint term is:

[0073] in, For the first The contribution matrix of each component of each training sample. For the first The contribution matrix of each component of each training sample. For the first The training sample and the first The transpose of the pseudo-label matrix to be optimized for each training sample. The similarity matrix, for The dynamic graph Laplacian matrix, , for A diagonal matrix.

[0074] In some optional implementations, the optimization adjustment unit is specifically used to fix one of the contribution variable and the pseudo-label to be optimized, and adjust the other of the contribution variable and the pseudo-label to be optimized, and perform the adjustment alternately until convergence.

[0075] The device for determining the efficacy substances of traditional Chinese medicine provided in this embodiment of the invention can execute the method for determining the efficacy substances of traditional Chinese medicine provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0076] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0077] The following is a detailed reference. Figure 4 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from memory 408 into random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0078] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0079] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a memory 408, or installed from a ROM 402. When the computer program is executed by the processor 401, it performs the functions defined in the method for determining the efficacy substances of traditional Chinese medicine according to embodiments of the present invention.

[0080] Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0081] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the method for determining the efficacy substances of traditional Chinese medicine shown in the above embodiments is implemented.

[0082] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0083] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for determining the active ingredients of traditional Chinese medicine, characterized in that, The method includes: Multiple training samples are obtained, each training sample corresponding to a Chinese herbal medicine and multiple efficacy tags of the Chinese herbal medicine, and the training sample includes multiple components of the Chinese herbal medicine; The contribution of each component in the training sample to each efficacy label is evaluated using a multi-label relevance feature selection algorithm to obtain the corresponding contribution. For each of the training samples, key components are selected according to their contribution to each efficacy label; Using the key components and efficacy labels corresponding to the multiple training samples, the predicted efficacy of each training sample is obtained, and the key components corresponding to each predicted efficacy are output.

2. The method according to claim 1, characterized in that, The step of evaluating the contribution of each component in the training samples to each efficacy label using a multi-label relevance feature selection algorithm includes: For each training sample, a target pseudo-label generation function is constructed based on the components of the training samples. The target pseudo-label generation function is a function with the contribution of the components in the training samples to the efficacy label as a variable. Based on the target pseudo-label generation function and the power labels corresponding to the training samples, a target function is constructed; The contribution variable in the objective function is optimized and adjusted to minimize the deviation between the target pseudo-label obtained based on the target pseudo-label generation function and the corresponding efficacy label, thereby obtaining the contribution of each component to each efficacy label.

3. The method according to claim 2, characterized in that, The objective function includes a first deviation term and a second deviation term; The first deviation term is used to calculate the deviation between the target pseudo-label obtained by the target pseudo-label generation function and the pseudo-label to be optimized, and the second deviation term is used to calculate the deviation between the pseudo-label to be optimized and the efficacy label; The optimization and adjustment of the contribution variable in the objective function, so as to minimize the deviation between the target pseudo-label obtained based on the target pseudo-label generation function and the corresponding efficacy label, to obtain the contribution of each component to each efficacy label, includes: The contribution variables and pseudo-labels to be optimized in the objective function are optimized and adjusted to minimize the value of the objective function, thereby obtaining the contribution of each component to each efficacy label.

4. The method according to claim 2 or 3, characterized in that, The objective function also includes the manifold structure constraint term of the pseudo-label; The manifold structure constraint term of the pseudo-label is used to calculate the product of the power similarity and the pseudo-label bias for each pair of training samples, where the power similarity is the power label similarity between the two training samples and the pseudo-label bias is the bias between the pseudo-labels of the two training samples.

5. The method according to claim 3, characterized in that, The objective function further includes a latent representation learning term, which is used to calculate the difference between the component similarity between every two training samples and a first product, where the first product is the product of the pseudo-label matrix to be optimized and the transpose of the pseudo-label matrix to be optimized.

6. The method according to claim 5, characterized in that, The component similarity between two training samples is calculated using a hot kernel function.

7. The method according to claim 3, characterized in that, The objective function also includes a dynamic graph constraint term, which is used to constrain the contribution using a dynamic graph Laplacian matrix. The dynamic graph Laplacian matrix is ​​obtained based on the similarity of the transposes of the pseudo-label matrices corresponding to the two training samples.

8. The method according to claim 7, characterized in that, The expression for the dynamic graph constraint term is: in, For the first The contribution matrix of each component of each training sample. For the first The contribution matrix of each component of each training sample. For the first The training sample and the first The transpose of the pseudo-label matrix to be optimized for each training sample. The similarity matrix, for The dynamic graph Laplace matrix, , for A diagonal matrix.

9. The method according to claim 3, characterized in that, The optimization and adjustment of the contribution variable and the pseudo-label to be optimized in the objective function includes: Fix one of the contribution variable and the pseudo-label to be optimized, and adjust the other of the contribution variable and the pseudo-label to be optimized, and adjust them alternately until convergence.

10. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected and the memory stores computer instructions. The processor executes the computer instructions to perform the method for determining the efficacy substances of traditional Chinese medicine as described in any one of claims 1 to 9.