Core party mining method and system based on improved association and composite clustering

By improving the methods of association and composite clustering, and combining them with the constraint theory of traditional Chinese medicine, we screened strongly associated drug groups and performed fourth-order chain clustering. This solved the problems of low-frequency drug omission and one-way allocation in TCM big data, and achieved accurate capture of TCM prescription logic and clinical interpretability.

CN121997284APending Publication Date: 2026-05-08杨阳 +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
杨阳
Filing Date
2025-12-31
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for mining big data in traditional Chinese medicine rely excessively on high-frequency statistics, leading to the omission of effective drug combinations in low-frequency cases. Furthermore, drugs can only be classified into a single cluster category, failing to reflect multiple roles and causing combinations to lose their clinical or formulary significance.

Method used

A method based on improved association and composite clustering was adopted. Strongly associated drug groups were screened through four-degree association rule analysis. Combined with the theory of constraints in traditional Chinese medicine, a fourth-order chain composite clustering algorithm was used for clustering to form a three-level system of core-association-special.

Benefits of technology

It accurately captures low-frequency effective combinations, breaks through the limitations of unidirectional drug distribution, ensures that the results conform to the logic of traditional Chinese medicine prescription, improves the interpretability of core prescriptions and the flexibility of clinical addition and subtraction, and provides reliable technical support for the inheritance of prescription rules and the creation of new prescriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997284A_ABST
    Figure CN121997284A_ABST
Patent Text Reader

Abstract

The invention provides a core prescription mining method and system based on improved association and composite clustering, and relates to the technical field of data processing, and the method comprises the steps: carrying out the statistics of the frequency of traditional Chinese medicines in a standardized traditional Chinese medicine data set; calculating a four-degree association rule analysis index of the binary drug group according to the co-occurrence frequency of the binary drug group; in combination with a preset four-degree association rule analysis index threshold value, screening out a plurality of strongly associated binary drug groups from the binary drug groups; performing standardization and vectorization processing on each strong correlation binary drug group, and constructing structured data containing channel tropism attributes and efficacy attributes; clustering the structured data through a four-order chain type composite clustering algorithm, and determining an extended cluster structure; in combination with a traditional Chinese medicine constraint theory, correcting the extended cluster structure, and outputting a traditional Chinese medicine constraint clustering result; and according to a preset core drug screening rule, optimizing the traditional Chinese medicine constraint clustering result, and generating a hierarchical core prescription structure comprising a core layer, an association layer and a special layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a core mining method and system based on improved association and composite clustering. Background Technology

[0002] Traditional Chinese medicine (TCM), as a unique health resource in my country, relies heavily on the exploration of its formula compatibility rules for inheritance and innovation of TCM diagnostic and treatment experience. Formulas, as the core carrier of TCM's syndrome differentiation and treatment, embody the synergistic effects of their constituent drugs. Accurately identifying the effective core drug combinations (i.e., core formulas) within formulas is crucial for a deeper understanding of their formulation rules, for inheriting the diagnostic and treatment experience of generations of TCM physicians, and for assisting in the creation of new formulas.

[0003] This core formula mining method deeply integrates the core theories of traditional Chinese medicine (TCM) such as the compatibility logic of principal, assistant, adjuvant, and guide herbs, and the efficacy of herbs tropism, with the advantages of modern algorithms such as improved association rules and composite clustering. It breaks away from the superficial reliance of traditional methods on high-frequency statistics and single clustering. It accurately aligns with the holistic compatibility thinking of TCM, ensuring the theoretical rationality and clinical interpretability of core combinations, while overcoming the limitations of missing low-frequency effective drug pairs and unidirectional drug allocation. It accurately captures low-frequency effective combinations for specific syndromes and supports multi-role drug assignment across clusters. This provides solid support that is both professional and practical for the deep decoding of TCM prescription patterns, the optimization of clinical syndrome differentiation and medication, and the scientific development of innovative prescriptions.

[0004] However, existing technologies for mining core combinations of prescriptions based on big data in traditional Chinese medicine are mainly implemented through shallow frequency statistics and traditional clustering algorithms (such as K-means). Shallow frequency statistics only focus on high-frequency single herbs, resulting in an overly superficial analysis that fails to reveal the core relationships within the prescription structure. Traditional clustering algorithms (such as K-means) group prescriptions based on the similarity of drug characteristics, ignoring the actual strength of the relationships. This leads to effective combinations being mechanically split into different clustering results, causing the mined "combinations" to lose their actual clinical or prescription significance. Summary of the Invention

[0005] To address the technical problems of existing technologies that rely excessively on high-frequency statistics, leading to the omission of low-frequency effective drug pairs, and that drugs can only be assigned to a single cluster category, failing to reflect multiple roles and exhibiting a one-way drug allocation defect, this invention provides a core mining method and system based on improved association and composite clustering.

[0006] The technical solutions provided by the embodiments of the present invention are as follows:

[0007] The first aspect of this invention provides a core square mining method based on improved association and composite clustering, comprising:

[0008] S1: Collect the original data of the prescription.

[0009] S2: Preprocess the original prescription data to obtain a standardized Chinese medicine dataset.

[0010] S3: Statistically analyze the frequency of Chinese medicines in the standardized Chinese medicine dataset, including the frequency of single herbs and the co-occurrence frequency of binary drug groups.

[0011] S4: Calculate the four-degree association rule analysis index of the binary drug group based on the co-occurrence frequency of the binary drug group.

[0012] S5: Combining the four-degree association rule analysis indicators and the preset four-degree association rule analysis indicator thresholds, multiple strongly associated binary drug groups are screened from the binary drug groups.

[0013] S6: Standardize and vectorize each strongly correlated binary drug group to construct structured data containing meridian tropism and efficacy attributes.

[0014] S7: Using a fourth-order chain-like composite clustering algorithm, structured data is clustered to determine the extended cluster structure.

[0015] S8: Combining the theory of constraints in traditional Chinese medicine, the extended cluster structure is modified, and the results of clustering under the constraints of traditional Chinese medicine are output.

[0016] S9: Based on the preset core drug screening rules, optimize the TCM constrained clustering results to generate a hierarchical core formula structure including a core layer, a related layer, and a special layer.

[0017] A second aspect of this invention provides a core square mining system based on improved association and composite clustering, comprising:

[0018] processor;

[0019] A memory storing computer-readable instructions, which, when executed by the processor, implement the core mining method based on improved association and composite clustering as described in the first aspect.

[0020] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the core mining method based on improved association and composite clustering as described in the first aspect.

[0021] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0022] In this embodiment of the invention, to address the issue of "over-reliance on high-frequency statistics leading to the omission of low-frequency effective drug pairs," a four-order chain clustering approach is used to integrate high-frequency core drug pairs with low-frequency characteristic combinations. Simultaneously, four-degree association rules (support, confidence, etc.) are used to accurately screen strongly associated drug pairs, avoiding interference from meaningless high-frequency combinations and resolving the problem of missing low-frequency drug pairs. To address the issue of "unidirectional drug allocation failing to reflect the multi-role nature of 'drugs changing with the prescription,'" clustering is modified using dual constraints from Traditional Chinese Medicine (hard isolation contraindications, soft support across clusters) to overcome the limitations of unidirectional allocation. Finally, to address the issue of "insufficient integration with the theory of monarch, minister, assistant, and guide in Traditional Chinese Medicine," a three-level system of "core-association-special" is formed. This ensures that the results conform to the logic of Traditional Chinese Medicine prescriptions, improves the interpretability of core prescriptions, and retains the flexibility for clinical additions and subtractions, providing reliable technical support for the inheritance of prescription rules and the creation of new prescriptions. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating a core mining method based on improved association and composite clustering, provided for an embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram of the structure of a formula for osteoarthritis (OA) provided in an embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram illustrating the association between an ID and a corresponding Chinese medicine name, provided as an embodiment of the present invention.

[0027] Figure 4 This is a schematic diagram illustrating the frequency of different drugs appearing simultaneously, provided as an embodiment of the present invention.

[0028] Figure 5 This is a schematic diagram illustrating the correlation index of support, confidence, and lift of different drug combinations provided in an embodiment of the present invention.

[0029] Figure 6 This is a schematic diagram of the structure of a meridian tropism vector of traditional Chinese medicine provided in an embodiment of the present invention.

[0030] Figure 7 This is a schematic diagram of a 31-dimensional efficacy result of traditional Chinese medicine provided in an embodiment of the present invention.

[0031] Figure 8 This is a schematic diagram of the core mining system based on improved association and composite clustering, provided as an embodiment of the present invention. Detailed Implementation

[0032] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0033] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0034] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0035] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0036] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0037] Reference manual attached Figure 1 The diagram illustrates a flowchart of a core mining method based on improved association and composite clustering provided by an embodiment of the present invention.

[0038] This invention provides a core square mining method based on improved association and composite clustering. This method can be implemented by a core square mining device based on improved association and composite clustering, which can be a terminal or a server. The processing flow of the core square mining method based on improved association and composite clustering may include the following steps:

[0039] S1: Collect the original data of the prescription.

[0040] Among them, the original data of prescriptions refers to prescription information from the TCM literature knowledge analysis platform, classic prescription works (such as the "Dictionary of TCM Prescriptions"), literature from specific periods (such as medical books from the Ming and Qing Dynasties), and clinical prescriptions from famous veteran TCM doctors.

[0041] For example, the sources of original prescription data include the Traditional Chinese Medicine Literature Knowledge Analysis Platform, classic prescription works (such as the "Dictionary of Traditional Chinese Medicine Prescriptions" edited by Peng Huai-ren et al.), literature from specific periods (such as medical books from the Ming and Qing Dynasties), and clinical treatment prescriptions from famous veteran TCM doctors. For example, 2,858 prescriptions related to osteoarthritis (OA) can be exported from the Traditional Chinese Medicine Literature Knowledge Analysis Platform.

[0042] In this embodiment of the invention, the collection of original prescription data provides diverse and professional basic data for subsequent core prescription mining, ensuring that the data coverage is relevant to clinical practice.

[0043] S2: Preprocess the original prescription data to obtain a standardized Chinese medicine dataset.

[0044] Preprocessing refers to the process of splitting the original data into Chinese herbal medicine names, parsing compound drug names, deleting irrelevant characters, and standardizing the names. A standardized Chinese herbal medicine dataset refers to a collection of Chinese herbal medicine data with unified names and a regular structure (such as the 31,354 Chinese herbal medicine records related to OA). Specifically, the preprocessing process includes splitting the Chinese herbal medicine names, parsing compound drug names, deleting irrelevant characters, and standardizing the split Chinese herbal medicines. The standardization reference is mainly based on the *Pharmacopoeia of the People's Republic of China* (China Medical Science and Technology Press, 2020) compiled by the National Pharmacopoeia Commission, supplemented by *Chinese Materia Medica* (Shanghai Science and Technology Press, 1998) and Ling Yikui's *Traditional Chinese Materia Medica* (Shanghai Science and Technology Press, 2001). For example, the OA-related prescriptions can be split, standardized, and then organized into 31,354 Chinese herbal medicine records.

[0045] In this embodiment of the invention, preprocessing eliminates the problem of inconsistent names of Chinese medicines, avoids subsequent analysis bias caused by name confusion, and provides a unified data foundation for frequency statistics.

[0046] S3: Statistically analyze the frequency of Chinese medicines in the standardized Chinese medicine dataset, including the frequency of single herbs and the co-occurrence frequency of binary drug groups.

[0047] Among them, the frequency of a single herb refers to the frequency of a single herb appearing in all prescriptions, and the frequency of co-occurrence of two herbs in a binary drug group refers to the frequency of two herbs appearing simultaneously in all prescriptions.

[0048] For example, the co-occurrence frequency of Ligusticum chuanxiong and Angelica sinensis can be counted as 488, and the co-occurrence frequency of Angelica sinensis and Cinnamomum cassia can be counted as 482.

[0049] In this embodiment of the invention, quantifying the occurrence and co-occurrence of drugs provides key data support for subsequent calculation of association rule indicators, which is the basis for identifying drug associations.

[0050] S4: Calculate the four-degree association rule analysis index of the binary drug group based on the co-occurrence frequency of the binary drug group.

[0051] The four-dimensional association rule analysis index refers to a set of indicators that evaluate the effectiveness of drug associations from different dimensions: support measures the generality of the association, confidence assesses the strength of the rule, lift tests the relevance, and certainty supplements the reverse risk analysis. The collaborative calculation of multi-dimensional indicators avoids the bias of a single indicator, comprehensively assesses the value of drug associations, and provides an objective basis for screening strongly associated drug pairs.

[0052] It should be noted that in association rule mining, support is used to screen for general associations, confidence is used to assess the strength of the rule, lift is used to test the relevance, and confidence is used to supplement the reverse risk analysis. These four core indicators can be combined to evaluate the practical value and effectiveness of association rules from multiple dimensions such as general applicability, strength, relevance, and reverse risk.

[0053] Optionally, the four-dimensional association rule analysis metrics include support, confidence, lift, and certainty.

[0054] The specific formula for calculating support is as follows:

[0055]

[0056] in, This indicates the frequency with which drug A and drug B appear simultaneously in all prescriptions. This represents the frequency of drug A and drug B appearing simultaneously in the prescription dataset, and N represents the total number of prescriptions.

[0057] It should be noted that the support level can filter out drug associations that are generally representative, exclude occasional niche combinations, and ensure that the rules are statistically significant.

[0058] The specific formula for calculating the confidence level is as follows:

[0059]

[0060] in, This represents the conditional probability that drug B also occurs when drug A is present. This indicates the frequency of drug A appearing alone in the prescription dataset.

[0061] It should be noted that confidence level can assess A's predictive ability for B, and a confidence level of ≥50% can ensure that the association rule has practical effectiveness.

[0062] The specific formula for calculating lift is as follows:

[0063]

[0064] in, This indicates the degree to which the presence of drug A increases the probability of the presence of drug B. This indicates the level of support for the simultaneous occurrence of drug A and drug B. This indicates the level of support for drug A appearing alone. This indicates the level of support for drug B appearing alone.

[0065] It should be noted that the lift can test the positive correlation between A and B (e.g., ≥1.5 indicates a significant correlation), avoiding misjudging drugs with no actual association as having a strong association.

[0066] The formula for calculating confidence level is as follows:

[0067]

[0068] in, This indicates the reverse error risk that occurs if rule A occurs. This represents the conditional probability that drug B also occurs when drug A is present.

[0069] It should be noted that confidence level compensates for the neglect of reverse rules by confidence level. Achieving the target (e.g., ≥1.2) can further reduce the risk of errors in association rules and improve reliability.

[0070] In this embodiment of the invention, by calculating the four-dimensional association rule analysis indicators of support, confidence, lift and certainty, the association effectiveness of binary drug groups can be comprehensively evaluated from multiple dimensions such as association universality, rule strength, relevance and reverse risk. This avoids judgment bias caused by a single indicator and provides objective and reliable data support for the subsequent accurate screening of strongly associated binary drug groups.

[0071] S5: Combining the four-degree association rule analysis indicators and the preset four-degree association rule analysis indicator thresholds, multiple strongly associated binary drug groups are screened from the binary drug groups.

[0072] It should be noted that those skilled in the art can set the preset threshold values ​​of the four-degree association rule analysis indicators according to actual needs, and this invention does not limit this.

[0073] Specifically, the preset thresholds include support ≥5%, confidence ≥50%, lift ≥1.5, and certainty ≥1.2. By screening binary drug groups that meet all of the above thresholds, strongly correlated binary drug groups can be obtained. For example, 44 pairs of strongly correlated binary groups such as frankincense and myrrh, and chuanxiong and angelica can be screened.

[0074] In this embodiment of the invention, threshold screening can accurately capture strong synergistic drug pairs and eliminate high-frequency meaningless combinations, providing high-quality data for subsequent standardization and clustering.

[0075] S6: Standardize and vectorize each strongly correlated binary drug group to construct structured data containing meridian tropism and efficacy attributes.

[0076] Standardization and vectorization refer to the processing steps including frequency standardization, symmetric matrix construction, meridian tropism vectorization, and efficacy vectorization. Structured data refers to the regularized data that integrates drug co-occurrence characteristics and TCM attributes (meridian tropism, efficacy).

[0077] In one possible implementation, S6 specifically includes sub-steps S601 to S604:

[0078] S601: Construct a symmetric matrix for each strongly correlated binary drug group.

[0079] It should be noted that the core purpose of constructing a symmetric matrix for strongly correlated binary drug groups is to prepare for the subsequent merging of bidirectional duplicate drug pair records. Through the structural characteristics of the symmetric matrix, the recording format of bidirectional drug pairs can be standardized, avoiding data redundancy caused by different drug pair orders and ensuring the consistency of subsequent processing.

[0080] Among them, the symmetric matrix refers to the matrix used to unify the bidirectional drug pair records, which can map drug pairs such as "drug A-drug B" and "drug B-drug A" that only differ in order to the same matrix position.

[0081] It should be noted that the construction of symmetric matrices provides a structural basis for merging bidirectional duplicate drug pairs and reducing data redundancy, ensuring consistency in subsequent processing.

[0082] S602: Based on a symmetric matrix, drug pairs with bidirectional duplicate records are merged into a single record to obtain multiple strongly correlated binary pairs.

[0083] Among them, bidirectional duplicate drug pairs refer to drug pairs such as "myrrh-frankincense" and "frankincense-myrrh", which are essentially the same but differ only in order.

[0084] For example, based on a symmetric matrix, when merging bidirectional duplicate drug pairs into a single record, the two bidirectional duplicate records “Myrrh-Frankincense (205)” and “Frankincense-Myrrh (205)” can be merged into a single record “Myrrh-Frankincense”. After merging, a total of 43 strongly correlated binary pairs are obtained, which effectively reduces data redundancy.

[0085] It should be noted that this step can eliminate data redundancy in bidirectional drug pairs, ensuring that only a single valid record is retained for each type of drug pair, thereby improving data processing efficiency.

[0086] S603: Perform frequency standardization on each strongly correlated binary pair to determine multiple relative frequencies.

[0087] Frequency standardization refers to the process of converting the original co-occurrence frequency into a relative frequency. The relative frequency is the ratio of the frequency of a certain drug pair to the total frequency of all strongly associated binary pairs (e.g., the relative frequency of Ligusticum chuanxiong-Angelica sinensis is ≈0.05).

[0088] Specifically, frequency standardization is performed on each strongly correlated pair to determine multiple relative frequencies. This involves converting the original co-occurrence frequency of each strongly correlated pair into a relative frequency, with the conversion logic being "the frequency of a certain drug pair divided by the total frequency of all drug pairs". For example, the original co-occurrence frequency of Ligusticum chuanxiong-Angelica sinensis is 488, and the total frequency of all strongly correlated pairs is 9722. The calculated relative frequency of Ligusticum chuanxiong-Angelica sinensis is approximately 0.05, thus eliminating dimensional differences between different drug pairs.

[0089] It should be noted that frequency standardization can eliminate dimensional differences between different drug pairs and avoid analytical bias caused by different original frequency scales.

[0090] S604: Based on the relative frequencies, combined with the preset meridian numbering system and the preset Chinese medicine efficacy standard table, vectorized attributes are added to each Chinese medicine to obtain structured data. Among them, the vectorized attributes include meridian tropism vectorized attributes and efficacy vectorized attributes.

[0091] The pre-defined meridian numbering system assigns corresponding numbers to each of the 12 meridians: 1 corresponds to the liver, 2 to the heart, 3 to the spleen, 4 to the lungs, 5 to the kidneys, 6 to the pericardium, 7 to the gallbladder, 8 to the small intestine, 9 to the stomach, 10 to the large intestine, 11 to the bladder, and 12 to the triple burner. The meridian tropism vectorization attribute is based on this system, using 1 to indicate that a drug tropizes a particular meridian and 0 to indicate that a drug does not tropize a particular meridian, forming a 12-dimensional meridian tropism vector.

[0092] Specifically, meridian tropism similarity refers to the degree of similarity between two Chinese herbal medicines in their meridian tropism. The higher the similarity value, the more similar the meridian tropism between the two medicines. For example, Angelica sinensis (Dang Gui) tropizes the liver, heart, and spleen meridians, while Ligusticum chuanxiong (Chuan Xiong) tropizes the liver, gallbladder, and pericardium meridians; their meridian tropism similarity is approximately 0.333. The meridian tropism similarity between Angelica sinensis and Paeonia lactiflora (Bai Shao), which tropize the liver and spleen meridians, is approximately 0.817. The meridian tropism similarity between Boswellia carterii (Ru Xiang) and Commiphora myrrha (Mo Yao) is 1, indicating that their meridian tropism is highly consistent.

[0093] Furthermore, the pre-set standard table of Chinese medicine efficacy is set according to the Chinese Materia Medica Classification Method of the 10th edition of "Chinese Materia Medica". It is a 31-dimensional efficacy system, covering efficacy types such as relieving exterior syndromes, clearing heat, purging, dispelling wind and dampness, promoting diuresis and eliminating dampness, warming the interior, regulating qi, promoting digestion, expelling parasites, stopping bleeding, promoting blood circulation and removing blood stasis, resolving phlegm, relieving cough and asthma, calming the mind, opening the orifices, tonifying deficiency, astringing, inducing vomiting, external use and others. Based on this, a Chinese medicine-efficacy knowledge base is constructed to provide a unified standard for the vectorization of the efficacy of each Chinese medicine.

[0094] For example, when adding a vectorized efficacy attribute to each Chinese herbal medicine, a value needs to be assigned based on the medicine's performance in the 31-dimensional efficacy system: 1 for a corresponding efficacy and 0 for no corresponding efficacy. For instance, Angelica sinensis has the effects of nourishing blood, promoting blood circulation, relieving pain, and moistening the intestines, and its efficacy vector is [0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,1,0,0,0,0,0,0,1,0,0,0,0,0,0]. Ligusticum chuanxiong has the effects of promoting blood circulation, regulating qi, dispelling wind, and relieving pain, and its efficacy vector is [0,0,0,0,0,0,0,0,0,0,1,0,0,0,1,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0]. White peony root has the effects of nourishing blood and astringing yin, softening the liver and relieving pain, and calming liver yang. Its efficacy vector is [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,1,0,0,0,0,0].

[0095] Furthermore, the standardization and vectorization process employs a triple standardization strategy: A symmetric matrix is ​​constructed for strongly correlated binary drug pairs, merging bidirectional drug pair records into a single record. Drug pairs with repeated bidirectional records are merged based on the symmetric matrix; for example, myrrh-frankincense and frankincense-myrrh are merged into a single myrrh-frankincense record. Frequency standardization is performed on each strongly correlated binary pair, converting the original co-occurrence frequency into a relative frequency to eliminate dimensional differences. Combining a pre-defined meridian numbering system and a standard table of Chinese herbal medicine efficacy, vectorized meridian tropism attributes (such as 12-dimensional vectors for liver and heart meridians) and efficacy vectorized attributes (such as a 31-dimensional efficacy system vector) are added to each Chinese herbal medicine, ultimately resulting in structured data.

[0096] In this embodiment of the invention, by constructing structured data containing meridian tropism and efficacy attributes, the differences in prescription dosage dimensions and data redundancy are eliminated, so that the data has both statistical consistency and TCM professional semantics, and is suitable for subsequent clustering requirements.

[0097] S7: Using a fourth-order chain-like composite clustering algorithm, structured data is clustered to determine the extended cluster structure.

[0098] Among them, the fourth-order chain composite clustering algorithm refers to a four-stage clustering process that successively improves Apriori, Ward hierarchical clustering, improves DBSCAN and corrects for traditional Chinese medicine constraints. The extended cluster structure refers to a drug cluster that includes core drugs and related drugs (such as the extended cluster of blood-nourishing and blood-activating drugs).

[0099] In one possible implementation, S7 specifically includes sub-steps S701 to S705:

[0100] S701: By using the improved Apriori algorithm, based on dynamic support threshold and lift threshold, structured data is identified to obtain high-frequency strongly correlated drug pairs.

[0101] Among them, the improved Apriori algorithm refers to an algorithm variant that introduces a dynamic support threshold (such as 0.02) and lift threshold for screening. It uses relative frequency as a substitute for traditional metric and uses alternative support to further calculate and identify high-frequency drug pairs. High-frequency strongly correlated drug pairs refer to drug pairs that meet the support criteria and have a lift > 1 (such as Angelica sinensis-Ligusticum chuanxiong).

[0102] It should be noted that those skilled in the art can set the dynamic support threshold and lift threshold according to actual needs, and this invention does not limit them.

[0103] For example, after setting a dynamic support threshold, the support and lift of binary drug pairs can be calculated. Binary drug pairs with support ≥ dynamic support threshold and lift > 1 are identified as high-frequency strongly correlated drug pairs. For example, the support of Angelica sinensis-Ligusticum chuanxiong is 0.05 and the lift is 6.278, which meets the screening criteria and can be regarded as high-frequency strongly correlated drug pairs.

[0104] For example, when identifying high-frequency, strongly correlated drug pairs using the improved Apriori algorithm, dynamic support thresholds (e.g., 0.02) and lift thresholds (e.g., >1) need to be set. Taking a total drug pair frequency of 9722 after merging bidirectional drug pairs as an example, the co-occurrence frequency of Angelica sinensis and Ligusticum chuanxiong is 488, with a support of 488 / 9722=0.05 and a lift of 6.278>1, meeting the screening criteria and being identified as a high-frequency, strongly correlated drug pair. However, the support of Angelica sinensis and Paeonia lactiflora is 178 / 9722=0.018<0.02, and they are not clustered separately for now.

[0105] In one possible implementation, S701 specifically includes sub-steps S7011 to S7013:

[0106] S7011: Set the dynamic support threshold and lift threshold.

[0107] Among them, the dynamic support threshold refers to the minimum support standard set according to the data scale (e.g., 0.02 when the total drug pair frequency is 9722), and the lift threshold refers to the minimum standard used to screen positively correlated drug pairs (e.g., >1).

[0108] It should be noted that the threshold setting provides a clear execution standard for subsequent drug pair screening, ensuring the accuracy of identifying high-frequency, strongly correlated drug pairs.

[0109] S7012: Calculate the support and lift of a binary drug group based on dynamic support and lift thresholds.

[0110] Among them, the support is calculated based on "drug pair frequency / total drug pair frequency", and the lift is calculated based on "co-occurrence support / (product of individual support)".

[0111] It should be noted that the support and improvement of the binary drug group are calculated by constructing a distance matrix by integrating statistical features of the data (co-occurrence intensity) and TCM attributes (meridian similarity), so that the distance measurement between drugs is more in line with the TCM prescription logic.

[0112] S7013: Binary drug pairs whose support is greater than the dynamic support threshold and whose lift is greater than the lift threshold, or whose support is equal to the dynamic support threshold and whose lift is greater than the lift threshold, are identified as high-frequency strongly correlated drug pairs.

[0113] It should be noted that the improved Apriori algorithm can accurately identify high-frequency drug pairs with positive correlations, providing a core association basis for subsequent hierarchical clustering and excluding high-frequency drug pairs with no actual association.

[0114] S702: Based on high-frequency, strongly correlated drug pairs, a distance matrix is ​​constructed by fusing co-occurrence strength and meridian similarity using the Ward hierarchical clustering definition method.

[0115]

[0116] Where Distance represents the combined distance between drugs A and B, and α represents the weight of co-occurrence frequencies. Indicates the co-occurrence frequency of drugs A and B. This represents the maximum frequency of all drugs affecting the CCP. This indicates the similarity in meridian tropism between drugs A and B.

[0117] Among them, the "meridian similarity" in the fusion co-occurrence strength and meridian similarity refers to the degree of similarity between two Chinese herbal medicines in their meridian tropism. The higher the meridian similarity value, the more similar the meridians tropism between the two medicines. For example, Angelica sinensis tropizes the liver, heart, and spleen meridians, while Ligusticum chuanxiong tropizes the liver, gallbladder, and pericardium meridians, and their meridian similarity is approximately 0.333. By fusing this similarity with the co-occurrence strength, a comprehensive distance matrix between the medicines can be constructed, providing a basis for subsequent clustering.

[0118] S703: Based on the distance matrix and the principle of minimizing variance increment, neighboring drugs are layer-by-layer merged to form multiple initial core clusters:

[0119]

[0120] Where, n A n B n C Let d(A,C), d(B,C), and d(A,B) represent the number of drugs contained in clusters A, B, and C, respectively. Let d(A,C), d(B,C), and d(A,B) represent the original distances between clusters A and C, B and C, and A and B, respectively. Let d(AB,C) represent the distance between the merged new cluster AB and another cluster C.

[0121] Among them, the principle of minimizing variance increment refers to choosing the merging method that minimizes the variance increment within the cluster when merging clusters. The initial core cluster refers to the drug cluster with a clear function formed by merging layer by layer (such as blood-nourishing and blood-activating drugs, and yang-warming and meridian-clearing drugs).

[0122] Specifically, based on the principle of minimizing variance increment, an initial core cluster is formed. First, the overall distance between drugs is calculated, and then neighboring drugs are merged layer by layer. For example, the distance between Angelica sinensis and Ligusticum chuanxiong is 0.2668, so they are merged into a new cluster first. Then, the inter-cluster distance between this new cluster and Paeonia lactiflora is calculated, and further merging is performed to form a cluster containing Angelica sinensis, Ligusticum chuanxiong, and Paeonia lactiflora. This process is repeated to gradually expand the initial core cluster, ultimately forming: Initial Core Cluster A (blood-nourishing and blood-activating: Angelica sinensis, Ligusticum chuanxiong, Paeonia lactiflora, Myrrh, Frankincense), Initial Core Cluster B (warming yang and unblocking collaterals: Cinnamon, Aconite, Saposhnikovia divaricata, Notopterygium incisum, Angelica pubescens, Ephedra sinica, Ginseng, Prepared Licorice), and Initial Core Cluster C (kidney-tonifying and bone-strengthening: Achyranthes bidentata, Eucommia ulmoides, Rehmannia glutinosa, Dioscorea hypoglauca, Cistanche deserticola.

[0123] It should be noted that layer-by-layer combination of neighboring drugs to form multiple initial core clusters can create initial core clusters with focused functions and strong internal correlations, ensuring that the drugs within the clusters conform to the synergistic logic of TCM efficacy and avoiding the disordered grouping of traditional clustering.

[0124] S704: Calculate the average power vector of each cluster in each initial core cluster.

[0125] Among them, the average efficacy vector refers to the vector obtained by taking the average of the components of the efficacy vectors of all drugs in the cluster (such as the average vector of the blood-nourishing and blood-activating cluster [0,0,...0.667,...]), which can characterize the "functional profile" of the cluster.

[0126] It should be noted that calculating the average efficacy vector of each initial core cluster, which is the "functional profile" of the cluster, requires taking the average of the components of the efficacy vectors of all drugs within the cluster. For example, in the average efficacy vector of cluster C1={Angelica sinensis, Ligusticum chuanxiong, Paeonia lactiflora}, the "blood-tonifying and blood-activating" functional dimension has a strength of 0.667, becoming the dominant function of this cluster, which can intuitively reflect the core efficacy direction of the cluster.

[0127] S705: By improving the DBSCAN algorithm, based on each average efficacy vector, the functional similarity between low-frequency drug pairs in strongly associated binary drug groups and the initial core cluster is calculated, and low-frequency drug pairs with functional similarity greater than the functional similarity threshold are classified into associated drugs to determine the extended cluster structure.

[0128] Among them, the improved DBSCAN algorithm refers to the DBSCAN variant that introduces the TCM functional similarity radius (ε=0.6), low-frequency drug pairs refer to drug pairs that do not meet the high-frequency strong correlation standard (such as the Dipsacus asper related drug pairs), and associated drugs refer to low-frequency drugs that meet the functional similarity standard and are classified into the core cluster.

[0129] It should be noted that those skilled in the art can set the functional similarity threshold according to actual needs, and this invention does not limit it.

[0130] Furthermore, to determine the extended cluster structure by improving the DBSCAN algorithm, the algorithm parameters need to be set first (functional similarity radius ε=0.6, minimum neighbor number minPts=3). Then, the functional similarity between low-frequency drug pairs and the initial core cluster is calculated. For example, the functional similarity between Dipsacus asperoides and the initial core cluster C is approximately 0.686≥0.6, and it is classified as an associated drug in cluster C. However, the functional similarity between low-frequency drugs such as Gastrodia elata and Aucklandia lappa and each initial core cluster is <0.6, and their drug pairs (such as Gastrodia elata-Saposhnikovia divaricata, Dendrobium nobile-Achyranthes bidentata) are marked as "special combinations" and retained, ultimately forming an extended cluster structure containing core drugs and associated drugs.

[0131] In one possible implementation, S705 specifically includes sub-steps S7051 to S7054:

[0132] S7051: Set parameters for the improved DBSCAN algorithm.

[0133] S7052: By using the improved DBSCAN algorithm with parameter settings and combining it with the preset Chinese medicine efficacy standard table, the efficacy vectorization processing is performed on each drug and each initial core cluster in the low-frequency drug pair to obtain the low-frequency drug efficacy vector and the initial core cluster efficacy vector.

[0134] S7053: Using the cosine similarity formula, calculate the functional similarity between the low-frequency drug pairs and each initial core cluster based on the low-frequency drug efficacy vector and the initial core cluster efficacy vector:

[0135]

[0136] in, This represents the functional similarity between drugs A and B based on efficacy vectors, where n represents the total dimension of the efficacy vectors. i B represents the component value of the i-th dimension cluster effectiveness vector. i This represents the component value of the i-th dimension that requires cross-cluster drug efficacy vector.

[0137] S7054: Incorporate low-frequency drug pairs with functional similarity greater than the functional similarity threshold into the associated drugs to determine the extended cluster structure.

[0138] Specifically, parameters for the improved DBSCAN algorithm are set, including the TCM functional similarity radius (e.g., ε=0.6) and the minimum number of neighbors (e.g., minPts=3). Based on a pre-defined table of TCM efficacy standards, efficacy vectorization is performed on each drug in low-frequency drug pairs and each initial core cluster. Functional similarity is calculated using the cosine similarity formula. Low-frequency drug pairs with a functional similarity > 0.6 are classified as associated drugs. For example, the functional similarity between Dipsacus asperoides and the initial core cluster of kidney-tonifying and bone-strengthening drugs is 0.686 ≥ 0.6, and they can be classified as associated drugs in that cluster.

[0139] It should be noted that by improving the DBSCAN algorithm, low-frequency effective drug pairs can be integrated into the corresponding core cluster, realizing the synergistic integration of high-frequency and low-frequency combinations, and avoiding the defect of traditional clustering that ignores high-frequency effective combinations.

[0140] In this embodiment of the invention, the fourth-order chain-based composite clustering algorithm can integrate high-frequency core drug pairs with low-frequency effective combinations, while incorporating the constraints of traditional Chinese medicine theory, thus avoiding the problems of mechanical splitting or omission of effective combinations in traditional clustering.

[0141] S8: Combining the theory of constraints in traditional Chinese medicine, the extended cluster structure is modified, and the results of clustering under the constraints of traditional Chinese medicine are output.

[0142] Traditional Chinese medicine's theory of constraint includes both hard constraints and soft constraints.

[0143] Among them, the TCM constraint theory refers to the TCM theoretical system that includes the hard contraindications of the Eighteen Incompatibilities and Nineteen Antagonisms and the soft rules of drug cross-cluster allocation, and the TCM constraint clustering result refers to the final cluster after being modified by hard and soft constraints.

[0144] In one possible implementation, S8 specifically includes sub-steps S801 to S806:

[0145] S801: Based on preset rules for incompatibilities in traditional Chinese medicine combinations, perform hard constraint checks on the extended cluster structure.

[0146] Among them, the preset rules for the incompatibilities of traditional Chinese medicine refer to the "Eighteen Incompatibilities and Nineteen Antagonisms" rule in traditional Chinese medicine. For example, aconite should not be combined with pinellia, trichosanthes, or fritillaria. By traversing each cluster and checking whether the drugs in the cluster violate the rule, the hard constraint check can be completed.

[0147] For example, by traversing the drugs in extended clusters A, B, and C, we can check whether there are any combinations that violate the rule. After checking, extended clusters A (Angelica sinensis, Ligusticum chuanxiong, etc.), extended clusters B (Cinnamomum cassia, Aconitum carmichaelii, etc.), and extended clusters C (Achyranthes bidentata, Eucommia ulmoides, etc.) do not violate the contraindications and do not require correction.

[0148] S802: Based on the results of hard constraint checks, combined with soft constraints, correct drug combinations in the extended cluster structure that violate the rules of incompatibility of traditional Chinese medicine.

[0149] Among them, drug combinations that violate the rules of incompatibilities in traditional Chinese medicine refer to drug combinations that do not conform to the "Eighteen Incompatibilities and Nineteen Antagonisms" (such as aconite and pinellia in the same group).

[0150] It should be noted that hard constraint checks can correct clustering results with potential safety risks, further ensuring the clinical compliance and applicability of the extended cluster structure.

[0151] S803: Based on the modified extended cluster structure, calculate the proportion of each drug's primary meridian tropism:

[0152]

[0153] in, This indicates the proportion of the main meridians into which the drug is tropized. Indicates the frequency of primary meridian aspiration. This represents the total frequency of all meridian affinities.

[0154] Among them, the proportion of drugs that are paired with ...

[0155] It should be noted that the proportion of primary meridian tropism is used to determine whether the meridian tropism of a drug is dispersed. Its calculation requires statistical analysis of the compatibility frequency of the drug with the primary meridian tropism drug (primary meridian tropism frequency) and the total compatibility frequency of the drug with all meridian tropism drugs (total frequency of all meridian tropisms). It is obtained by multiplying the ratio of the primary meridian tropism frequency to the total frequency of all meridian tropisms by 100%. For example, the primary meridian tropism frequency of Notopterygium root is 369, and the total frequency of all meridian tropisms is 709, so the proportion of primary meridian tropism is approximately 52.0%.

[0156] S804: By combining the proportion of each primary meridian and the preset threshold for the proportion of primary meridian, candidate drugs with dispersed meridian tropism are screened out.

[0157] Among them, the preset threshold for the proportion of main meridian tropism refers to the minimum standard for judging whether the meridian tropism of a drug is dispersed (e.g., 60%), and the candidate drugs with dispersed meridian tropism refer to drugs whose proportion of main meridian tropism is lower than the threshold (e.g., the proportion of main meridian tropism of Qianghuo is ≈52.0%).

[0158] It should be noted that those skilled in the art can set the threshold value of the main meridian proportion according to actual needs, and this invention does not limit it.

[0159] For example, when calculating the proportion of primary meridian tropism for each drug, taking Rehmannia glutinosa (processed) as an example: Rehmannia glutinosa tropizes the liver and kidney meridians, and its total frequency of combination with drugs of the liver meridian is 475 times. The total frequency of combination with all meridians is 475 times, and the proportion of primary meridian tropism is 100% ≥ 60%, indicating no dispersion in meridian tropism. Taking Notopterygium incisum (Qianghuo) as an example: Notopterygium incisum tropizes the bladder and kidney meridians, and its frequency of combination with drugs of the bladder meridian is 369 times. The total frequency of combination with all meridians is 709 times, and the proportion of primary meridian tropism is approximately 52.0% < 60%, which is considered dispersion in meridian tropism, making it a candidate drug.

[0160] S805: Calculate the functional similarity between candidate drugs and each extended cluster, and assign candidate drugs with functional similarity greater than or equal to the functional similarity threshold to the corresponding extended clusters.

[0161] Among them, the functional similarity threshold refers to the minimum standard for judging that a drug is similar in efficacy to an extended cluster (such as 0.6), and cross-cluster assignment refers to classifying a drug into multiple extended clusters that meet the efficacy similarity standard.

[0162] Specifically, when calculating the functional similarity between candidate drugs and each extended cluster, taking the meridian-dispersed Notopterygium root as an example: the efficacy vector of Notopterygium root corresponds to the functions of "relieving exterior syndromes with pungent and warm properties and dispelling wind and dampness," and its functional similarity with extended clusters A, B, and C are 0, 0.442, and 0.171, respectively, all <0.6, which does not meet the cross-cluster assignment condition. Since none of the meridian-dispersed drugs reached the functional similarity threshold, the final extended cluster structure remains unchanged, and the output is the clustering result after correction by hard and soft constraints.

[0163] S806: Based on the allocation results, output the TCM constraint clustering results after correction by hard and soft constraints.

[0164] Among them, hard constraints refer to the inspection and correction of the "Eighteen Incompatibilities and Nineteen Antagonisms", soft constraints refer to the cross-cluster allocation of drugs that are tropism-based and functionally dispersed, and TCM constraint clustering results refer to the final drug clusters after double constraint correction (such as the corrected blood-nourishing and blood-activating cluster).

[0165] In this embodiment of the invention, the extended cluster structure is modified by combining the theory of constraints in traditional Chinese medicine. This modification ensures that the clustering results have both data objectivity and professional rationality in traditional Chinese medicine, avoids disconnection from the logic of traditional Chinese medicine compatibility, and improves the clinical applicability of the results.

[0166] S9: Based on the preset core drug screening rules, optimize the TCM constrained clustering results to generate a hierarchical core formula structure including a core layer, a related layer, and a special layer.

[0167] The pre-set core drug screening rules include pre-set dual threshold rules and pre-set rules for defining principal and assistant drugs;

[0168] Among them, the preset core drug screening rules refer to the dual threshold standards for screening core drugs (the number of high-frequency drug pairs involved is ≥3 and the proportion of main meridians is ≥60%). The hierarchical core formula structure refers to a three-level structure including the core layer (principal and assistant drugs), the related layer (adjuvant and guide drugs), and the special layer (clinical addition and subtraction drug pairs).

[0169] In one possible implementation, S9 specifically includes sub-steps S901 to S905:

[0170] S901: Based on a preset dual threshold rule, core drugs are selected from each drug cluster. The dual threshold rule is as follows: the number of high-frequency drug pairs involved by the drug is not less than the first threshold, and the proportion of the drug's main meridian tropism is not less than the second threshold.

[0171] Among the preset dual threshold rules, the first threshold is "the number of high-frequency drug pairs involved by the drug is ≥3", and the second threshold is "the proportion of the drug's main meridian tropism is ≥60%". Taking Angelica sinensis in cluster A as an example: Angelica sinensis participates in 10 high-frequency drug pairs such as Ligusticum chuanxiong-Angelica sinensis and Notopterygium incisum-Angelica sinensis, with a main meridian tropism proportion of approximately 85.7%, thus meeting both thresholds and being selected as a core drug.

[0172] Specifically, the first threshold can be set to 3 and the second threshold can be set to 60%. For example, if the number of high-frequency drug pairs involved by Angelica sinensis is 10≥3 and the proportion of the main angelica sinensis is about 85.7%≥60%, it can be screened as a core drug if it meets the double threshold rule.

[0173] S902: Based on the pre-defined rules for defining principal and assistant drugs, the principal and assistant drugs are identified among the core drugs to form the core layer.

[0174] Among them, the principal drug is the drug with the highest total frequency in its drug cluster and whose main meridian tropism ratio reaches the third threshold, and the assistant drug is the drug whose co-occurrence frequency with the principal drug reaches the fourth threshold and whose main meridian tropism ratio is not lower than the second threshold.

[0175] Among them, the third threshold refers to the minimum standard of the proportion of the principal drug in the meridian (such as 70%), the fourth threshold refers to the minimum standard of the co-occurrence frequency of the assistant drug and the principal drug (such as high co-occurrence frequency), and the core layer refers to the hierarchy composed of the principal drug and the assistant drugs (such as the core layer of the blood-nourishing and blood-activating cluster: Angelica sinensis is the principal drug, and Ligusticum chuanxiong and Paeonia lactiflora are the assistant drugs).

[0176] For example, when defining the principal and assistant herbs to form the core layer, taking cluster A as an example: the principal herb must meet the requirement of "highest total frequency within the cluster and primary meridian tropism percentage ≥ 70%". Angelica sinensis has a total frequency of 2440 (highest within the cluster) and a primary meridian tropism percentage of 85.7% (≥ 70%), and is therefore defined as the principal herb. The assistant herbs must meet the requirement of "high co-occurrence with the principal herb and primary meridian tropism percentage ≥ 60%". Ligusticum chuanxiong and Angelica sinensis have a co-occurrence frequency of 488 and a primary meridian tropism percentage of 100%, and Paeonia lactiflora and Angelica sinensis have a co-occurrence frequency of 282 and a primary meridian tropism percentage of 71%, and are both defined as assistant herbs, together forming the core layer.

[0177] S903: For non-core layer drugs, calculate the functional similarity with the drug cluster to which they belong, and group drugs with functional similarity not lower than the fifth threshold into related drugs of the drug cluster to form an association layer.

[0178] Among them, non-core drugs refer to drugs that have not been screened as core drugs (such as myrrh and frankincense), the fifth threshold refers to the lowest standard of functional similarity (such as 0.6), and the association layer refers to the hierarchy composed of adjuvant drugs (such as the blood-nourishing and blood-activating cluster association layer: myrrh and frankincense as adjuvants).

[0179] Specifically, when non-core drugs form an association layer, taking myrrh and frankincense in cluster A as examples: neither of them reached the core drug screening threshold, but their functional similarity with cluster A was >0.6, and they played an auxiliary role in promoting blood circulation and relieving pain. They were classified as associated drugs of cluster A, corresponding to adjuvant drugs in traditional Chinese medicine prescriptions, thus forming an association layer.

[0180] S904: Drug pairs with functional similarity below the fifth threshold and not classified into any drug cluster are retained as special combinations to form a special layer.

[0181] Among them, special combinations refer to drug pairs with a functional similarity of ≤0.6 and not classified into any cluster (such as Gastrodia elata-Saposhnikovia divaricata, Dendrobium nobile-Achyranthes bidentata), and special layers refer to the layers that retain special combinations for clinical addition and subtraction of drug combinations.

[0182] It should be noted that when retaining special combinations to form special layers, the retained "special combinations" (such as Gastrodia elata-Saposhnikovia divaricata, Dendrobium nobile-Achyranthes bidentata) can be used for clinical modification and combination. For example, when there is concurrent liver wind syndrome, the Gastrodia elata-Saposhnikovia divaricata combination can be added to the basic formula of cluster A. When OA belongs to liver and kidney yin deficiency syndrome, the Dendrobium nobile-Achyranthes bidentata combination can be added to the basic formula of cluster C. This conforms to the flexible formula composition concept of "differentiation of syndromes and treatment" in traditional Chinese medicine, and finally outputs a hierarchical core formula structure of "core layer (principal and assistant) - related layer (adjuvant and guide) - special layer (drug pair)".

[0183] For example, if the fifth threshold is set to 0.6, the functional similarity of Gastrodia elata-Saposhnikovia divaricata and Dendrobium nobile-Achyranthes bidentata are both ≤0.6 and they are not classified into any drug cluster. These two drug pairs can be retained as a special combination to form a special layer. This special combination can be used for clinical addition and subtraction of prescriptions. For example, when there is concurrent liver wind, the Gastrodia elata-Saposhnikovia divaricata combination can be added to the basic prescription for nourishing blood and promoting blood circulation.

[0184] S905: Outputs a hierarchical core structure including a core layer, related layers, and special layers.

[0185] Among them, the core layer corresponds to the principal and assistant herbs in traditional Chinese medicine, the related layer corresponds to the adjuvant and guiding herbs, and the special layer corresponds to the clinical addition and subtraction of herbs. Together, the three constitute the hierarchical core formula structure.

[0186] It should be noted that those skilled in the art can set the values ​​of the first threshold, the second threshold, the third threshold, the fourth threshold, and the fifth threshold according to actual needs, and this invention does not limit these settings.

[0187] In this embodiment of the invention, the results of TCM constrained clustering are optimized according to preset core drug screening rules, and the clustering results are transformed into core formula structures that conform to the TCM theory of monarch, minister, assistant and guide, thereby improving the interpretability and clinical application value of the results and providing direct support for the inheritance of formulas and the creation of new formulas.

[0188] Reference manual attached Figure 2 The diagram shows a structural schematic of a formula for osteoarthritis (OA) provided in an embodiment of the present invention.

[0189] It should be noted that the present invention can collect, for example... Figure 2 The diverse original prescription data provides rich and professional basic data support for subsequent standardized processing, frequency statistics and core prescription mining, ensuring that the mining results are consistent with the clinical practice of traditional Chinese medicine.

[0190] Reference manual attached Figure 3 This diagram illustrates an association between an ID and a corresponding Chinese medicine name, as provided in an embodiment of the present invention.

[0191] It should be noted that this invention can split and organize the collected prescription data into Chinese herbal medicine names to form a standardized Chinese herbal medicine data list as described above, providing basic data support for subsequent frequency statistics, vectorization processing, etc.

[0192] Reference manual attached Figure 4 This diagram illustrates the frequency of different drugs appearing simultaneously, as provided in an embodiment of the present invention.

[0193] It should be noted that this invention statistically analyzes the frequency of individual herbs and the co-occurrence frequency of binary drug groups in a standardized Chinese herbal medicine dataset, such as the co-occurrence frequency of Chuanxiong-Danggui and the frequency of each individual herb, etc. This lays a data foundation for subsequent calculation of four-degree association rule analysis indicators and accurate screening of strongly associated binary drug groups, ensuring the reliability of association analysis.

[0194] Reference manual attached Figure 5 This diagram illustrates a correlation index of support, confidence, and lift of different drug combinations provided in an embodiment of the present invention.

[0195] It should be noted that this invention uses four-dimensional association rule analysis indicators—support, confidence, lift, and certainty—such as the multi-dimensional indicators of myrrh-frankincense drug pairs mentioned above, to comprehensively evaluate the effectiveness of drug association and provide accurate basis for screening strongly associated binary drug groups.

[0196] Reference manual attached Figure 6The diagram shows a structural schematic of a meridian tropism vector for traditional Chinese medicine provided in an embodiment of the present invention.

[0197] It should be noted that this invention converts the meridian tropism of Chinese medicine into vector representations, such as the meridian tropism vectors of Angelica sinensis, Ligusticum chuanxiong, and Paeonia lactiflora mentioned above, thereby quantifying the meridian tropism attributes in Chinese medicine and providing a basis for judgment from a Chinese medicine perspective for subsequent cluster analysis that integrates meridian tropism similarity.

[0198] Reference manual attached Figure 7 This diagram illustrates a 31-dimensional efficacy result of traditional Chinese medicine provided by an embodiment of the present invention.

[0199] It should be noted that, based on the aforementioned 31-dimensional standard table of Chinese medicine efficacy, this invention can standardize and vectorize the efficacy of Chinese medicine, providing a unified efficacy dimension benchmark for subsequent functional similarity calculation, clustering, and core formula efficacy analysis.

[0200] Reference manual attached Figure 8 The diagram shows a schematic of the core mining system based on improved association and composite clustering provided by the present invention.

[0201] This invention also provides a core square mining system 20 based on improved association and composite clustering, applied to the aforementioned core square mining method based on improved association and composite clustering, comprising:

[0202] Processor 201.

[0203] The memory 202 stores computer-readable instructions, which, when executed by the processor 201, implement the core mining method based on improved association and composite clustering as described in the method embodiment.

[0204] The core square mining system 20 based on improved association and composite clustering provided by the present invention can execute the core square mining method based on improved association and composite clustering described above and achieve the same or similar technical effects. To avoid duplication, the present invention will not elaborate further.

[0205] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0206] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0207] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0208] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0209] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0210] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0211] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0212] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0213] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0214] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0215] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0216] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0217] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the core mining method based on improved association and composite clustering as described in the method embodiments.

[0218] The present invention provides a computer-readable storage medium that can implement the steps and effects of the core mining method based on improved association and composite clustering in the above-described method embodiments. To avoid repetition, the present invention will not elaborate further.

[0219] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0220] The following points need to be explained:

[0221] (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.

[0222] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0223] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0224] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A core idea mining method based on improved association and composite clustering, characterized in that, include: S1: Collect original data of the prescription; S2: Preprocess the original data of the prescription to obtain a standardized Chinese medicine dataset; S3: Statistically analyze the frequency of Chinese medicines in the standardized Chinese medicine dataset, wherein the frequency of Chinese medicines includes the frequency of single herbs and the co-occurrence frequency of binary drug groups; S4: Calculate the four-degree association rule analysis index of the binary drug group based on the co-occurrence frequency of the binary drug group; S5: Combining the four-degree association rule analysis index and the preset four-degree association rule analysis index threshold, select multiple strongly associated binary drug groups from the binary drug groups; S6: Standardize and vectorize each of the strongly correlated binary drug groups to construct structured data containing meridian tropism attributes and efficacy attributes; S7: The structured data is clustered using a fourth-order chain composite clustering algorithm to determine the extended cluster structure; S8: Based on the theory of constraints in traditional Chinese medicine, the extended cluster structure is modified, and the results of clustering under the constraints of traditional Chinese medicine are output. S9: Based on the preset core drug screening rules, optimize the TCM constrained clustering results to generate a hierarchical core formula structure including a core layer, an association layer, and a special layer.

2. The core idea mining method based on improved association and composite clustering according to claim 1, characterized in that, The four-dimensional association rule analysis indicators include support, confidence, lift, and certainty. The specific formula for calculating the support is as follows: ; in, This indicates the frequency with which drug A and drug B appear simultaneously in all prescriptions. This represents the frequency of drug A and drug B appearing simultaneously in the prescription dataset, and N represents the total number of prescriptions. The specific formula for calculating the confidence level is as follows: ; in, This represents the conditional probability that drug B also occurs when drug A is present. This indicates the frequency of drug A appearing alone in the prescription dataset; The specific formula for calculating the lift is as follows: ; in, This indicates the degree to which the presence of drug A increases the probability of the presence of drug B. This indicates the level of support for the simultaneous occurrence of drug A and drug B. This indicates the level of support for drug A appearing alone. This indicates the level of support for drug B appearing alone; The formula for calculating the degree of certainty is as follows: ; in, This indicates the reverse error risk that occurs if rule A occurs. This represents the conditional probability that drug B also occurs when drug A is present.

3. The core idea mining method based on improved association and composite clustering according to claim 1, characterized in that, S6 specifically includes: S601: Construct a symmetric matrix for each of the strongly correlated binary drug groups; S602: Based on the symmetric matrix, the drug pairs with bidirectional repeated records are merged into a single record to obtain multiple strongly correlated binary pairs; S603: Perform frequency standardization on each of the strongly correlated binary pairs to determine multiple relative frequencies; S604: Based on the relative frequencies of each herb, and in conjunction with the preset meridian numbering system and the preset Chinese herbal efficacy standard table, vectorized attributes are added to each herb to obtain the structured data. The vectorized attributes include meridian tropism vectorized attributes and efficacy vectorized attributes.

4. The core idea mining method based on improved association and composite clustering according to claim 1, characterized in that, Specifically, S7 includes: S701: Using the improved Apriori algorithm, based on dynamic support threshold and lift threshold, the structured data is identified to obtain high-frequency strongly correlated drug pairs; S702: Based on the aforementioned high-frequency strongly correlated drug pairs, a distance matrix is ​​constructed by fusing co-occurrence intensity and meridian similarity using the Ward hierarchical clustering definition method. ; Where Distance represents the combined distance between drugs A and B, and α represents the weight of co-occurrence frequencies. Indicates the co-occurrence frequency of drugs A and B. This represents the maximum frequency of all drugs affecting the CCP. This indicates the similarity in meridian tropism between drugs A and B; S703: Based on the distance matrix and the principle of minimizing variance increment, the neighboring drugs are layer-by-layer merged to form multiple initial core clusters: ; Where, n A n B n C Let d(A,C), d(B,C), and d(A,B) represent the number of drugs contained in clusters A, B, and C, respectively. Let d(A,C), d(B,C), and d(A,B) represent the original distances between clusters A and C, B and C, and A and B, respectively. Let d(AB,C) represent the distance between the merged new cluster AB and another cluster C. S704: Calculate the average power vector of each cluster in each of the initial core clusters; S705: By improving the DBSCAN algorithm, based on each of the average efficacy vectors, calculate the functional similarity between the low-frequency drug pairs in the strongly associated binary drug group and the initial core cluster, and classify the low-frequency drug pairs with functional similarity greater than the functional similarity threshold into the associated drugs to determine the extended cluster structure.

5. The core idea mining method based on improved association and composite clustering according to claim 4, characterized in that, Specifically, S701 includes: S7011: Set the dynamic support threshold and lift threshold; S7012: Calculate the support and lift of the binary drug group based on the dynamic support threshold and the lift threshold; S7013: The binary drug pair with the support greater than the dynamic support threshold and the lift greater than the lift threshold, or the support equal to the dynamic support threshold and the lift greater than the lift threshold, is determined as the high-frequency strongly correlated drug pair.

6. The core mining method based on improved association and composite clustering according to claim 4, characterized in that, S705 specifically includes: S7051: Set the parameters of the improved DBSCAN algorithm; S7052: Using the improved DBSCAN algorithm with parameter settings, combined with the preset Chinese medicine efficacy standard table, the efficacy vectorization processing is performed on each drug in the low-frequency drug pair and each initial core cluster to obtain the low-frequency drug efficacy vector and the initial core cluster efficacy vector. S7053: Using the cosine similarity formula, calculate the functional similarity between the low-frequency drug pairs and each of the initial core clusters, based on the low-frequency drug efficacy vector and the initial core cluster efficacy vector: ; in, This represents the functional similarity between drugs A and B based on efficacy vectors, where n represents the total dimension of the efficacy vectors. i B represents the component value of the i-th dimension cluster effectiveness vector. i This represents the component value of the i-th dimension that requires cross-cluster drug efficacy vector; S7054: Classify the low-frequency drug pairs with functional similarity greater than the functional similarity threshold into the associated drugs to determine the extended cluster structure.

7. The core idea mining method based on improved association and composite clustering according to claim 1, characterized in that, S8 specifically includes: The TCM constraint theory includes hard constraints and soft constraints; S801: Based on preset rules of incompatibility between traditional Chinese medicines, perform hard constraint checks on the extended cluster structure; S802: Based on the results of the hard constraint check, and in conjunction with the soft constraints, correct the drug combinations in the extended cluster structure that violate the rules of incompatibility of traditional Chinese medicine. S803: Based on the modified extended cluster structure, calculate the proportion of each drug's primary meridian tropism: ; in, This indicates the proportion of the main meridians into which the drug is tropized. Indicates the frequency of primary meridian aspiration. This represents the total frequency of all meridian affinities; S804: Combining the proportion of each of the main meridians and the preset threshold for the proportion of the main meridians, candidate drugs with dispersed meridian tropism are screened out; S805: Calculate the functional similarity between the candidate drug and each extended cluster, and assign candidate drugs with functional similarity greater than the functional similarity threshold and functional similarity equal to the functional similarity threshold to the corresponding extended clusters across clusters; S806: Based on the allocation results, output the TCM constraint clustering results after correction by the hard constraints and the soft constraints.

8. The core idea mining method based on improved association and composite clustering according to claim 1, characterized in that, S9 specifically includes: The preset core drug screening rules include preset dual threshold rules and preset rules for defining principal and assistant drugs; S901: Based on the preset dual threshold rule, core drugs are selected from each drug cluster, wherein the dual threshold rule is: the number of high-frequency drug pairs involved by the drug is not less than the first threshold, and the proportion of the drug's main meridian is not less than the second threshold. S902: According to the preset rules for defining principal and assistant drugs, the principal and assistant drugs are defined in the core drugs to form the core layer; S903: Calculate the functional similarity between non-core layer drugs and their respective drug clusters, and group drugs with functional similarity not lower than the fifth threshold into associated drugs of their respective drug clusters to form the associated layer; S904: Drug pairs with functional similarity below the fifth threshold and not classified into any drug cluster are retained as special combinations to form the special layer; S905: Output a hierarchical core structure including the core layer, the associated layer, and the special layer.

9. A core method mining system based on improved association and composite clustering, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the core mining method based on improved association and composite clustering as described in any one of claims 1 to 8.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the core mining method based on improved association and composite clustering as described in any one of claims 1 to 8.