Method for generating class parties and related products
By constructing a prescription network and merging similar nodes, highly reliable class prescriptions are generated, solving the problem that class prescription identification relies on doctors' subjective experience. This enables accurate, reliable, and systematic automatic generation of class prescriptions, improving the objectivity and reproducibility of TCM syndrome differentiation and treatment.
Patent Information
- Application Number
- CN202511040064.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, prescription identification mainly relies on doctors' subjective experience and judgment, lacking objective quantitative and standardized analysis methods, which makes it difficult to reproduce prescription identification results and unable to form a systematic knowledge system.
By constructing an initial prescription network, similar nodes are merged based on the efficacy attributes and compatibility information in the prescriptions to generate a target prescription network. A pre-set model is then used to mine highly reliable similar prescriptions, thereby realizing a computable model of TCM syndrome differentiation and treatment thinking.
It achieves accurate, reliable, and systematic automatic generation of prescriptions, breaks through the limitations of subjective experience, provides a calculable and verifiable objective framework, and improves the accuracy and efficiency of prescription selection and medication.
Smart Images

Figure CN120998526A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the field of traditional Chinese medicine formula identification technology. More specifically, this application relates to a method for generating formulas and related products. Background Technology
[0002] "Classified formulas" is an important concept in Traditional Chinese Medicine (TCM) formulary, referring to a collection of formulas that share similarities in drug composition, efficacy, indications, or compatibility. These formulas are derived from a core formula, such as a classic prescription, by adding or subtracting drugs and adjusting dosages to create a series of formulas targeting different conditions, constitutions, or concurrent symptoms. While these classified formulas may differ in the number and dosage of drugs, they share a core compatibility structure and treatment principles. Classified formulas hold significant value in TCM clinical practice and academic research. Clinically, doctors can quickly match formulas best suited to a patient's symptoms through comparison, improving the accuracy and efficiency of prescription selection. Academically, research on classified formulas helps to deepen the understanding of the compatibility principles and evolutionary patterns of TCM formulas, uncover the potential application value of classic formulas, and provide insights for new drug development and theoretical innovation.
[0003] However, current methods for identifying similar prescriptions mainly rely on doctors' subjective experience and judgment, lacking objective quantitative and standardized analytical methods. This makes it difficult to reproduce the results of similar prescription identification and prevents the formation of a systematic knowledge system.
[0004] In view of this, there is an urgent need to provide a method and related products for generating prescriptions, so as to transform the TCM syndrome differentiation and treatment thinking into a computable model and realize the automatic generation of accurate and reliable systematic prescriptions. Summary of the Invention
[0005] In order to at least address one or more of the technical problems mentioned above, this application proposes methods and related products for generating class squares in several aspects.
[0006] In a first aspect, this application provides a method for generating class prescriptions, comprising: acquiring prescription data to be analyzed, wherein the prescription data includes multiple prescriptions, the prescriptions including traditional Chinese medicines, efficacy attributes of the traditional Chinese medicines, and / or compatibility information of the traditional Chinese medicines; constructing an initial prescription network based on the efficacy attributes and / or compatibility information of each traditional Chinese medicine in the prescriptions; merging similar nodes in the initial prescription network to construct a target prescription network, wherein each node in the target prescription network is referred to as a target node, and the target node includes multiple similar nodes; and determining the class prescription corresponding to the target node based on the nodes included in the target node in the target prescription network.
[0007] In some embodiments, prior to obtaining the prescription data to be analyzed, the method includes: extracting raw prescription data from medical record documents, wherein the raw prescription data includes multiple prescriptions, the prescriptions including traditional Chinese medicine and / or compatibility information of the traditional Chinese medicine; extracting the efficacy attributes of each traditional Chinese medicine in the prescription from the pharmacopoeia; mapping the efficacy attributes into binary vectors; and associating the binary vectors with each traditional Chinese medicine in the prescription to obtain the prescription data to be analyzed.
[0008] In some embodiments, constructing an initial prescription network based on the efficacy attributes and / or compatibility information of each herb in the prescription includes: dividing the prescription data into first prescription data and second prescription data based on whether or not the compatibility information is included, wherein each prescription in the first prescription data includes the compatibility information, and each prescription in the second prescription data does not include the compatibility information; determining the similarity between prescriptions in the first prescription data based on the compatibility information of each herb in the prescription; determining the similarity between prescriptions in the second prescription data based on the efficacy attributes of each herb in the prescription; and constructing the initial prescription network by establishing undirected edge connections between prescription pairs whose similarity meets a preset condition, using the similarity as the edge weight.
[0009] In some embodiments, determining the similarity between prescriptions based on the compatibility information of each herb in the prescription in the first prescription data includes: obtaining the weights of the compatibility information; using the weights of the compatibility information of each herb in the prescription as the weights of the herb, constructing a weighted vector of the prescription, wherein the weighted vector is a set of weights of each herb in the prescription; and determining the similarity between each prescription based on the weighted vector of each prescription.
[0010] In some embodiments, determining the similarity between prescriptions based on the efficacy attributes of each herb in the prescription within the second prescription data includes: constructing an efficacy attribute set for each prescription based on the efficacy attributes corresponding to all herbs in the prescription, wherein the efficacy attribute set includes a subset of efficacy attributes corresponding to attribute types, and the subset of efficacy attributes records the efficacy attributes of each specific type corresponding to the attribute type and the frequency of occurrence of each specific type of efficacy attribute; and determining the similarity between prescriptions based on the efficacy attribute set of each prescription.
[0011] In some embodiments, the preset condition is that the similarity is greater than a preset similarity threshold.
[0012] In some embodiments, merging similar nodes in the initial prescription network to construct a target prescription network includes: traversing each node of the initial prescription network and determining the gain between the currently traversed node and its neighboring nodes based on the edge weights of the nodes; if the gain is greater than a preset gain, merging the currently traversed node into the neighboring node; after traversal, determining the sum of the edge weights between the nodes connected by each edge in the initial prescription network; updating the sum to the edge weights between the nodes connected by each edge; continuing to execute the steps of traversing each node of the initial prescription network and determining the gain between the currently traversed node and its neighboring nodes based on the edge weights and the number of edge connections of the nodes, until the gain between any node in the initial prescription network is less than or equal to the preset gain, thereby constructing the target prescription network.
[0013] In some embodiments, determining the class of prescription corresponding to the target node based on the nodes included in the target node in the target prescription network includes: determining the importance of each node included in the target node in the target prescription network; determining the relative risk of each Chinese medicine in each node included in the target node in the target prescription network relative to other target nodes in the target prescription network; and determining the class of prescription corresponding to the target node in the target prescription network based on the importance and the relative risk.
[0014] In some embodiments, determining the class prescription corresponding to the target node in the target prescription network based on the importance and the relative risk includes: selecting a first preset number of nodes from the target nodes to form a first prescription set based on the importance of each node included in the target node; selecting a second preset number of target Chinese herbs from each node in the first prescription set based on the relative risk of each of the Chinese herbs; and taking the intersection of the target Chinese herbs selected from each node as the class prescription of the target node.
[0015] In some embodiments, the method further includes: using the prescription data to be analyzed as input to a preset model, and using each target node in the target prescription network as a label, training the preset model; continuing to execute the step of obtaining the prescription data to be analyzed; inputting the prescription data into the preset model to obtain the target nodes of each prescription in the prescription data in the target prescription network; merging each prescription in the prescription data into the corresponding target node; continuing to execute the step of determining the class prescription corresponding to the target node based on the nodes included in the target prescription network, until the class prescription corresponding to the target node is obtained.
[0016] In some embodiments, the method further includes: determining the prescription to which each herb in the formula belongs; determining the user corresponding to the medical record document to which the prescription belongs; extracting original target information from all medical record documents of the user, wherein the original target information includes at least one of symptoms, efficacy, and disease; determining target information that was originally recorded but disappeared in subsequent records according to the chronological order of the original target information; constructing a data association matrix and a contingency table based on the herbs included in the formula and the target information; and performing a chi-square test based on the data association matrix and the contingency table to obtain the correlation strength between the formula and the target information.
[0017] In a second aspect, this disclosure provides a processing apparatus comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions, which, when loaded and executed by the processor, cause the processor to perform the method according to the foregoing first aspect and any embodiment thereof.
[0018] In a third aspect, this disclosure provides a computer-readable storage medium storing program instructions that, when loaded and executed by a processor, cause the processor to perform the method according to the foregoing first aspect and any of its embodiments.
[0019] Using the method for generating similar prescriptions provided above, this application embodiment constructs a prescription network with the characteristics of TCM syndrome differentiation thinking by combining the efficacy attributes and compatibility information in the prescription. Based on the prescription data in the prescription network nodes, highly reliable similar prescriptions with similar prescription structures, compatibility rules and treatment orientations are mined. This realizes the transformation of TCM syndrome differentiation and treatment thinking into a computable model, and thus can accurately, reliably and systematically generate similar prescriptions automatically, completing the systematic extraction from prescription data to similar prescriptions. Attached Figure Description
[0020] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, and the same or corresponding reference numerals denote the same or corresponding parts, wherein:
[0021] Figure 1 An exemplary flowchart of a method for generating class squares according to some embodiments of this application is shown;
[0022] Figure 2 An exemplary flowchart of a method for generating class squares according to some embodiments of this application is shown;
[0023] Figure 3 An exemplary flowchart of a method for generating class squares according to some embodiments of this application is shown;
[0024] Figure 4 An exemplary flowchart of a method for generating class squares according to some embodiments of this application is shown;
[0025] Figure 5 An exemplary flowchart of a method for generating class squares according to some embodiments of this application is shown;
[0026] Figure 6 An exemplary structural block diagram of a processing apparatus according to some embodiments of this application is shown. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] It should be understood that the terms "comprising" and "including" used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0029] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0030] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0031] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0032] Exemplary application scenarios
[0033] "Classified formulas" is an important concept in Traditional Chinese Medicine (TCM) formulary, referring to a collection of formulas that share similarities in drug composition, efficacy, indications, or compatibility. These formulas are derived from a core formula, such as a classic prescription, by adding or subtracting drugs and adjusting dosages to create a series of formulas targeting different conditions, constitutions, or concurrent symptoms. While these classified formulas may differ in the number and dosage of drugs, they share a core compatibility structure and treatment principles. Classified formulas hold significant value in TCM clinical practice and academic research. Clinically, doctors can quickly match formulas best suited to a patient's symptoms through comparison, improving the accuracy and efficiency of prescription selection. Academically, research on classified formulas helps to deepen the understanding of the compatibility principles and evolutionary patterns of TCM formulas, uncover the potential application value of classic formulas, and provide insights for new drug development and theoretical innovation.
[0034] However, current methods for identifying similar prescriptions mainly rely on doctors' subjective experience and judgment, lacking objective quantitative and standardized analytical methods. This makes it difficult to reproduce the results of similar prescription identification and prevents the formation of a systematic knowledge system.
[0035] Exemplary application scheme
[0036] In view of this, embodiments of this application provide a method for generating similar prescriptions. By combining the efficacy attributes and compatibility information in the prescription, a prescription network with the characteristics of TCM syndrome differentiation thinking is constructed. Based on the prescription data in the prescription network nodes, highly reliable similar prescriptions with similar prescription structures, compatibility rules and treatment orientations are mined. This realizes the transformation of TCM syndrome differentiation and treatment thinking into a computable model, and thus can accurately, reliably and systematically generate similar prescriptions automatically, completing the systematic extraction from prescription data to similar prescriptions.
[0037] The following describes various embodiments of a method for generating class squares provided in this application.
[0038] Figure 1 An exemplary flowchart of a method 100 for generating class squares according to some embodiments of this application is shown, which includes steps S101 to S104.
[0039] In step S101, the prescription data to be analyzed is obtained, wherein the prescription data includes multiple prescriptions, and the prescriptions include traditional Chinese medicine, the efficacy attributes of traditional Chinese medicine and / or the compatibility information of traditional Chinese medicine.
[0040] A group of similar prescriptions refers to a group of prescriptions that share similarities in their composition, compatibility rules, and therapeutic goals. It represents a systematic summary of prescriptions with homologous or analogous characteristics. Therefore, in some embodiments, the more prescription data acquired and the more comprehensive the clinical diagnostic scenarios covered, the better. Clinical diagnostic scenarios refer to the acquisition of prescription data covering different syndrome types of the same pathogenesis and different stages of the same disease (such as the initial, middle, and recovery stages).
[0041] In some embodiments, the attribute type of efficacy attributes includes at least one of efficacy, nature and flavor, and meridian tropism. Each of these attributes can be further subdivided into various specific types. Taking efficacy as an example, its efficacy attributes may include, but are not limited to, the following categories: exterior-releasing attributes such as inducing sweating to relieve exterior symptoms, dispelling wind and clearing heat, heat-clearing attributes such as clearing heat and detoxifying, clearing heat and drying dampness, tonifying attributes such as tonifying qi, tonifying blood, tonifying yang, tonifying yin, etc., and blood-activating and stasis-removing attributes such as activating blood and relieving pain, removing stasis and stopping bleeding, etc. Compatibility information includes at least one of principal, assistant, adjuvant, and guiding herbs. It can be understood that compatibility information refers to the combination relationship and functional positioning between drugs in a prescription. For example, a prescription may include Bupleurum (principal) and Astragalus (assistant).
[0042] In some embodiments, to obtain the prescription data to be analyzed, raw prescription data can be extracted from the user's medical record document. This raw prescription data includes multiple prescriptions, each containing traditional Chinese medicine (TCM) and / or its compatibility information. Then, the efficacy attributes of each TCM herb in the prescription are extracted from the pharmacopoeia. Next, the efficacy attributes are mapped to binary vectors. Finally, the binary vectors are associated with each TCM herb in the prescription to obtain the prescription data to be analyzed. The resulting prescription data includes each prescription containing TCM, its efficacy attributes, and / or its compatibility information.
[0043] It should be noted that the prescriptions in the original prescription data were personalized treatment plans issued by doctors based on clinical diagnostic information such as the user's symptoms, tongue appearance, and pulse, and in accordance with Traditional Chinese Medicine (TCM) theory. These prescriptions at least include Chinese herbal medicine components, but due to the influence of doctors' individual recording habits, there may be differences in whether or not information on the compatibility of Chinese herbs is recorded. Therefore, some prescriptions in the original prescription data include Chinese herbs and their compatibility information, while some prescriptions only include Chinese herbs without their compatibility information. Thus, the statement that prescriptions in the original prescription data include Chinese herbs and / or their compatibility information means that some prescriptions in the original prescription data include both Chinese herbs and their compatibility information, while others only include Chinese herbs without their compatibility information. Based on this, the statement that prescriptions in the final prescription data to be analyzed include Chinese herbs, their efficacy attributes, and / or their compatibility information means that some prescriptions in the final prescription data to be analyzed include Chinese herbs, their efficacy attributes, and their compatibility information, while others only include Chinese herbs and their efficacy attributes without their compatibility information.
[0044] A pharmacopoeia is a codex of drug quality standards formulated and promulgated by the state, which records the efficacy attributes of various traditional Chinese medicines. It should be noted that in this embodiment, the attribute dimensions of the efficacy attributes of each traditional Chinese medicine are consistent; that is, regardless of the type of traditional Chinese medicine, the type of efficacy attribute it is marked with is the same.
[0045] Taking efficacy attributes as an example, each Chinese herbal medicine needs to be labeled with the efficacy of "tonifying qi", "tonifying blood", "tonifying yin", and "tonifying yang". Based on this, if Chinese herbal medicine A has the efficacy of "tonifying qi" and "tonifying blood", but not the efficacy of "tonifying yin" and "tonifying yang", then the corresponding efficacy of "tonifying qi", "tonifying blood", "tonifying yin", and "tonifying yang" for Chinese herbal medicine A will be mapped to binary vectors "1", "1", "0", and "0" respectively, and these binary vectors will be associated with Chinese herbal medicine A.
[0046] Through the above operations, prescription data to be analyzed, including information on traditional Chinese medicine, efficacy attributes, and / or compatibility, can be obtained. Specifically, this prescription data can be represented as P = {P1, P2, ..., P...} n}, P n Let P represent the nth prescription, and P represent the nth prescription. n ={h1,h2,…,h n}, h n This represents the nth Chinese medicinal herb.
[0047] In some embodiments, the efficacy attributes and compatibility information of each Chinese herbal medicine can be embedded into the Chinese herbal medicine to obtain a Chinese herbal medicine characterization containing multidimensional features, so that h n ={s n q n r n y n}, where s n q n r n They represent the Chinese medicine h respectively n The efficacy, properties, and meridian tropism of [the substance / property]. n Indicates Chinese medicine h n Compatibility information.
[0048] In step S102, an initial prescription network is constructed based on the efficacy attributes and / or compatibility information of each Chinese herbal medicine in the prescription.
[0049] In this embodiment, based on the efficacy attributes and / or compatibility information of each Chinese herbal medicine in the prescription, the edges and edge weights used to construct the initial prescription network are determined. Then, using each prescription as a node, undirected edges are established to connect prescription pairs with edges, and edge weights are assigned to the edges, thus constructing the initial prescription network. Based on this, the connection relationships between the nodes of the initial prescription network reflect the association between prescriptions in terms of efficacy attributes and compatibility patterns, thereby providing a calculable and verifiable objective framework for prescription classification. This approach inherits the traditional Chinese medicine academic thought of "grouping prescriptions by category" while overcoming the limitations of subjective experience through prescription networks, enabling prescription classification generation to move from qualitative description to a more precise and systematic modern research paradigm.
[0050] In step S103, similar nodes in the initial prescription network are merged to construct a target prescription network, wherein each node in the target prescription network is called a target node, and the target node includes multiple similar nodes.
[0051] In this embodiment, to facilitate the distinction between nodes in the initial prescription network and the target prescription network, each node in the target prescription network is referred to as a target node. It can be understood that a target node includes multiple similar nodes (prescriptions).
[0052] In step S104, the class of prescription corresponding to the target node is determined based on the nodes included in the target node in the target prescription network.
[0053] In this embodiment, as described above, each node in the initial prescription network corresponds to one prescription, while class prescriptions are a systematic induction of multiple prescriptions with similarity or homology. Therefore, it is necessary to further merge multiple prescription nodes with similarity or homology in the initial prescription network to construct a target prescription network that can be used for class prescription generation, and then class prescriptions can be mined based on the nodes included in the target nodes in the target prescription network.
[0054] The above combination Figure 1 The present application provides a detailed description of a method 100 for generating prescription-like formulas in some embodiments. This method constructs a prescription network with the characteristics of traditional Chinese medicine (TCM) diagnostic thinking by combining the efficacy attributes and compatibility information in the prescription. Based on the prescription data in the prescription network nodes, it mines highly reliable prescription-like formulas with similar prescription structures, compatibility rules, and treatment orientations. This transforms TCM diagnostic and treatment thinking into a computable model, enabling the accurate, reliable, and systematic automatic generation of prescription-like formulas, thus completing the systematic extraction from prescription data to prescription-like formulas.
[0055] Figure 2 The method 200 for generating class squares shown in some embodiments of this application can be used as a specific implementation of step S102 in method 100 above. Therefore, the foregoing combined with Figure 1 The described features can be similarly applied here. For example... Figure 2 As shown, method 200 includes steps S201 to S204 as described below.
[0056] In step S201, the prescription data is divided into first prescription data and second prescription data according to whether or not it includes compatibility information. Each prescription in the first prescription data includes compatibility information, while each prescription in the second prescription data does not include compatibility information.
[0057] In this embodiment, as described above, some prescriptions in the prescription data to be analyzed include traditional Chinese medicine, efficacy attributes, and compatibility information, while others only include traditional Chinese medicine and efficacy attributes, without compatibility information. Therefore, the classification here is a clear division based on the presence or absence of compatibility information, ensuring consistency between the two types of data within their respective categories and providing a differentiated processing basis for the subsequent construction of the initial prescription network.
[0058] In step S202, the similarity between prescriptions is determined based on the compatibility information of each Chinese herbal medicine in the first prescription data.
[0059] In some embodiments, to determine the similarity between prescriptions in the first prescription data, the following method is used: First, the weights of the compatibility information are obtained. Then, the weights of the compatibility information of each Chinese herb in the prescription are used as the weights of the herbs to construct a weighted vector for the prescription, where the weighted vector is the set of weights of each Chinese herb in the prescription. For example, if a prescription includes Bupleurum (chief herb) and Astragalus (assistant herb), the weight of the chief herb is 'a', and the weight of the assistant herb is 'b'. Therefore, the weight of Bupleurum is 'a', and the weight of Astragalus is 'b', and the weighted vector of the prescription is {a, b}. Finally, the similarity between each prescription is determined based on the weighted vectors of each prescription.
[0060] In some embodiments, the similarity calculation between prescriptions follows the expression:
[0061]
[0062] Among them, S(P i ,P j ) indicates prescription P i and prescription P j The similarity between them, W i,k Indicates prescription P i The weighted vector W i The weight of the internal Chinese medicine k, W j,k Indicates prescription P j The weighted vector W j The weights of the internal Chinese medicine k, where m represents the weighting vector W. i and weighted vector W j The maximum number of elements in the dataset.
[0063] Therefore, the similarity between prescriptions can be calculated using the above expression and the weighted vector of the prescriptions.
[0064] In some embodiments, the weights of the principal, minister, assistant, and envoy in the compatibility information can be set to 2, 1, 0.6, and 0.4, respectively. Of course, this is not the only possible setting, and those skilled in the art can flexibly adjust the weights of the compatibility information based on the disclosure and teachings of this embodiment and their own professional experience.
[0065] In step S203, in the second prescription data, the similarity between prescriptions is determined based on the efficacy attributes of each Chinese herbal medicine in the prescription.
[0066] In some embodiments, to determine the similarity between prescriptions in the second prescription data, the following method is used: First, obtain a set of efficacy attributes for each prescription. This set of efficacy attributes includes a subset of efficacy attributes corresponding to each attribute type, and this subset records the frequency of occurrence of each specific type of efficacy attribute corresponding to the attribute type. Then, based on the set of efficacy attributes for each prescription, determine the similarity between the prescriptions.
[0067] In some embodiments, the similarity calculation between prescriptions follows the expression:
[0068]
[0069] Here, the set of efficacy attributes is represented as G. n ={s′ n ,q′ n ,r′ n}, s′ n q′ represents a subset of attributes whose type is efficacy. n r′ represents a subset of properties of type 'flavor' with corresponding functional properties. n This represents a subset of efficacy attributes whose attribute type is meridian tropism; similarity(G i G j ) indicates prescription G i and prescription G j The similarity between them; f(s′) i ,s′ j ) indicates prescription G i Attribute type s′ i and prescription G j Attribute type s′ j The similarity between them; f(q′) i ,q′ j ) indicates prescription G i Attribute type q′ i and prescription G j Attribute type q′ j The similarity between them; f(r′) i ,r′ j ) indicates prescription G i Attribute type r′ i and prescription G j Attribute type r′ j The similarity between them; softmax() represents the softmax function (also known as the activation function), used to smooth the values of the power attribute as a binary vector of 0.
[0070] In some embodiments, to obtain the efficacy attribute set of a prescription, the efficacy attributes of all Chinese herbs in the prescription corresponding to the same attribute type are combined to obtain a subset of efficacy attributes of the corresponding attribute type. Then, the subsets of efficacy attributes corresponding to all attribute types are combined to obtain the initial efficacy attribute set of the prescription. For example, if the attribute types include efficacy, nature and flavor, and meridian tropism, the obtained initial efficacy attribute set includes the subsets of efficacy attributes corresponding to the three attribute types: efficacy, nature and flavor, and meridian tropism. For example, the initial efficacy attribute set obtained for a certain prescription is {Efficacy subset: [Clearing heat and detoxifying: 1, Tonifying Qi: 1, Clearing heat and detoxifying: 0], Nature and flavor subset: [Sweet: 1, Bitter: 1, Sweet: 1, Cool: 0], Meridian tropism subset: [Lung meridian: 1, Spleen meridian: 1, Lung meridian: 1]}, where the values are binary vectors corresponding to the efficacy attributes. It should also be noted that the Chinese characters to the left of the binary vectors are only for visual display of the efficacy attributes corresponding to each binary vector; in reality, the data of the obtained initial efficacy attribute set consists only of binary vectors. After obtaining the initial set of efficacy attributes, the frequency of occurrence of each specific type of efficacy attribute in each subset of efficacy attributes is statistically analyzed to obtain the set of efficacy attributes. The frequency of an efficacy attribute can be understood as the frequency of its binary vector 1. For example, the set of efficacy attributes obtained for a prescription might be {Efficacy: {Clearing heat and detoxifying: 2, Tonifying Qi: 1}, Taste and nature: {Sweet: 2, Bitter: 1, Cool: 0}, Meridian tropism: {Lung Meridian: 2, Spleen Meridian: 0}}, where the values represent the frequency of occurrence of each specific type of efficacy attribute. It should be noted that the Chinese text to the left of the values (frequency of efficacy attributes) is only for visually displaying the corresponding efficacy attributes; the actual data in the obtained set of efficacy attributes only includes the frequency of occurrence of each specific type of efficacy attribute.
[0071] Finally, in step S204, prescriptions are used as nodes, and undirected edges are established between prescription pairs that meet the preset similarity conditions. The similarity is used as the edge weight to construct the initial prescription network.
[0072] In this embodiment, the preset condition is that the similarity is greater than a preset similarity threshold. Therefore, after obtaining the similarity between prescriptions in the prescription data to be analyzed, each prescription is used as a node, and undirected edges are established between prescription pairs with similarity greater than the preset similarity threshold. The similarity of the prescription pair is used as the edge weight, thereby constructing the initial prescription network.
[0073] In some embodiments, the preset similarity threshold is 0.7.
[0074] The above combination Figure 2The present application provides a detailed description of a method 200 for generating prescription-like structures based on some embodiments. This method determines the similarity between prescriptions by examining the compatibility relationships of the Chinese herbs in the first prescription data containing compatibility information. For the second prescription data not containing compatibility information, the similarity is determined by examining the efficacy attributes of the Chinese herbs in the prescriptions. Finally, using prescriptions as nodes, undirected edges are established between prescription pairs whose similarity meets preset conditions, and the similarity of these prescription pairs is used as the edge weight. This constructs an initial prescription network that comprehensively reflects the relationships between prescriptions and embodies the dialectical thinking of Traditional Chinese Medicine. This method simultaneously captures the similarity of prescriptions in two dimensions: compatibility rules and efficacy characteristics, providing a data foundation for subsequent prescription-like structure mining.
[0075] Figure 3 The present application shows a method 300 for generating class squares according to some embodiments, which can be used as a specific implementation of step S103 in method 100 above. Therefore, the foregoing combined with Figure 1 The described features can be similarly applied here. For example... Figure 3 As shown, method 300 includes steps S301 to S305 as described below.
[0076] In step S301, each node of the initial prescription network is traversed, and the modularity gain of the currently traversed node after merging with its neighboring nodes is determined based on the edge weights of the nodes.
[0077] In this embodiment, the node directly connected to the currently traversed node by an edge is its neighboring node. Modularity gain refers to the increase in the modularity of the initial prescription network after merging the currently traversed node with its neighboring nodes.
[0078] In some embodiments, the module gain calculation follows the following expression:
[0079]
[0080] Where ΔQ represents the modularity gain of the currently traversed node after merging with its neighboring nodes, ∑in represents the sum of the edge weights of all nodes included in the neighboring node, and k i,in ∑tot represents the sum of edge weights between the currently traversed node and all nodes included in its neighboring nodes, ∑tot represents the sum of edge weights between all nodes included in its neighboring nodes and all nodes in the initial prescription network, and k represents the sum of edge weights between the currently traversed node and all nodes included in its neighboring nodes. i This represents the sum of the edge weights of the currently traversed node and all nodes in the initial prescription network, where m represents half the sum of the edge weights of all nodes in the initial prescription network.
[0081] In step S302, if the modularity gain is greater than the preset gain, the currently traversed node is merged into the adjacent node.
[0082] In this embodiment, the preset gain can be 0. It can be understood that if the modularity gain is greater than the preset gain, it indicates that after the currently traversed node is merged into adjacent nodes, the modularity of the initial prescription network is improved; in other words, the currently traversed node is similar to all nodes included in its adjacent nodes.
[0083] In step S303, after the traversal is completed, the sum of the edge weights between the nodes connected by each edge in the initial prescription network is determined.
[0084] In this embodiment, after the traversal is complete, all nodes in the initial prescription network with similar efficacy attributes and similar compatibility patterns are merged. In other words, prescriptions with similar efficacy attributes and similar compatibility patterns are merged. At this point, the edge weights of each edge connection in the initial prescription network can no longer represent the similarity between two nodes, so they need to be updated. Specifically, the sum of edge weights between nodes connected by each edge refers to the sum of edge weights of all originally connected node pairs included in the two nodes connected by the edge. For example, in the current initial prescription network, nodes A and B are connected by an edge, where node A includes nodes A1 and A2, and node B includes nodes B1 and B2. Originally, nodes A1 and B1 were connected by an edge with an edge weight of a1, and nodes A2 and B2 were connected by an edge with an edge weight of b1. Therefore, the sum of edge weights is c = a1 + b1.
[0085] In step S304, the sum is updated to reflect the edge weights between the nodes connected by each edge. In this embodiment, it can be understood that the sum is the new edge weights between the nodes connected by the edge.
[0086] In step S305, the process of traversing each node of the initial prescription network and determining the gain between the currently traversed node and its neighboring nodes based on the node's edge weight and the number of edge connections is continued until the gain between any node in the initial prescription network is less than or equal to a preset gain, thereby constructing the target prescription network.
[0087] In this embodiment, after updating the edge weights of all nodes, the process continues to merge all similar nodes in the initial prescription network until the gain of any node is less than or equal to a preset gain, meaning further merging is no longer possible. This indicates that all similar nodes have been completely categorized, thus constructing the target prescription network. It can be understood that each node in the currently obtained target prescription network has the following characteristics: each node represents a merged group of similar prescriptions, aggregated from prescriptions with similar treatment strategies and drug compatibility patterns; the edge weights between nodes are iteratively updated, reflecting the overall correlation between the merged prescription groups, with higher values indicating a stronger frequency of combined use or synergistic effect. Therefore, this embodiment achieves the goal of providing a structured basis for subsequent prescription classification mining by constructing the target prescription network.
[0088] The above combination Figure 3 The present application provides a detailed description of a method 300 for generating similar prescriptions in some embodiments. This method constructs a target prescription network by merging nodes (prescriptions) with similar efficacy attributes (treating similar symptoms) and similar compatibility rules in an initial prescription network. This results in the aggregation of all nodes (prescriptions) with similar treatment strategies and similar drug compatibility rules within each target node of the target prescription network, providing a structured basis for subsequent prescription mining.
[0089] Figure 4 The method 400 for generating class squares shown in some embodiments of this application can be used as a specific implementation of step S104 in method 100 above. Therefore, the foregoing combined with Figure 1 The described features can be similarly applied here. For example... Figure 4 As shown, method 400 includes steps S401 to S403 as described below.
[0090] In step S401, the importance of each node included in the target node in the target prescription network is determined.
[0091] As discussed above, each target node in the target prescription network aggregates all nodes (prescriptions) with similar efficacy attributes and similar compatibility rules. These cannot be directly used as class prescriptions; further exploration is needed to identify class prescriptions with the same core efficacy and core compatibility rules from the nodes included in the target node. Therefore, in this embodiment, PR (PageRank) values are used to represent importance. Specifically, firstly, all nodes included in the target node are assigned the same initial PR value to ensure consistency in the evaluation starting point. Subsequently, the PR value of the target node is iteratively calculated using the PR value expression, dynamically updating the PR value of each target node. This process simulates the transmission and convergence of importance between nodes. After each iteration, the change in the PR value of the target node between two consecutive iterations is calculated, and then compared with a preset convergence threshold. If the change is less than the preset convergence threshold, the current PR value is taken as the importance of the node.
[0092] In some embodiments, the PR value is calculated according to the following expression:
[0093]
[0094] Among them, PR(D) i ) represents the target node D i The PR value after the current iteration, where d represents the damping coefficient, usually set to 0.85, (1-d) represents the probability of randomly jumping to any target node, and M... i This represents all pointers to the target node D. i The set of nodes, j∈M iRepresents target node D j Pointing to the target node set M i For any target node, PR(D) j ) represents the target node D j The PR value after the last iteration Indicates pointing to target node D j The sum of all edge weights, W ji Represents target node D i and target node D j The edge weight between the target node D is... i and target node D j If there is no edge connecting them, then W ji It is 0.
[0095] In step S402, the relative risk of each Chinese medicine in each node included in the target prescription network relative to other target nodes in the target prescription network is determined.
[0096] In this embodiment, the Relative Risk (RR) value is used to represent the relative risk, which measures the specificity of the traditional Chinese medicine (TCM) within the target node compared to the external target node. A higher RR value indicates that the TCM has significant characteristics within the target node. In some embodiments, the RR value is calculated according to the following expression:
[0097]
[0098] Among them, RR(h) i ) indicates Chinese medicine h i The RR value, c1_pos_count represents the total number of Chinese medicines h within the target node. i The total number, c2_pos_coun represents all Chinese medicines h outside the target node. i The total number of nodes (prescriptions) included in the target node, c1_count represents the total number of nodes (prescriptions) included in the target node, and c2_count represents the total number of nodes (prescriptions) included in the target node outside the target node.
[0099] In step S403, the class of prescription corresponding to the target node in the target prescription network is determined based on importance and relative risk.
[0100] In this embodiment, to mine the class prescriptions of the target node, the following method can be used: Based on the importance of each node included in the target node, a first preset number of nodes are selected from the target node to form a first prescription set; specifically, a first preset number of nodes with high importance are selected to form the first prescription set. For example, the top 5 nodes in terms of importance are selected from the target node to form the first prescription set. Then, based on the relative risk of each herb, a second preset number of target herbs are selected from each node in the first prescription set; specifically, a second preset number of target herbs with high relative risk are selected. For example, the top 5 herbs in terms of relative risk are selected from each node as target herbs. Finally, the intersection of the target herbs selected from each node is taken as the class prescription of the target node.
[0101] In some embodiments, the specific values of the first preset quantity and the second preset quantity mentioned above may be the same or different. The specific values of the first preset quantity and the second preset quantity can be determined according to the number of Chinese medicines included in the node (prescription). For example, if the node includes a large number of Chinese medicines, the specific values of the first preset quantity and the second preset quantity will be larger; if the node includes a small number of Chinese medicines, the specific values of the first preset quantity and the second preset quantity will be smaller. In some embodiments, the specific values of the first preset quantity and the second preset quantity mentioned above can also be determined by the doctor himself, and this embodiment does not impose specific restrictions on this.
[0102] The above combination Figure 4 The method 400 for generating class prescriptions in some embodiments of this application is described in detail. It selects class prescriptions for target nodes by jointly screening the importance of each node included in the target prescription network and the relative danger of each Chinese medicine in each node included in the target prescription network relative to other target nodes in the target prescription network. This not only provides a comprehensive solution for the structured mining of TCM class prescriptions, but also the mined class prescriptions have high credibility.
[0103] Figure 5 The present application illustrates a method 500 for generating class squares according to some embodiments, which can be used as an additional technical solution to the method 100 described above. Therefore, the preceding text is combined with... Figure 1 The described features can be similarly applied here. For example... Figure 5 As shown, method 500 includes steps S501 to S505 as described below.
[0104] In step S501, the prescription data to be analyzed is used as the input to the preset model, and each target node in the target prescription network is used as a label to train the preset model.
[0105] In this embodiment, the prescription data to be analyzed, which is used above to construct the target prescription network, is used as the input of the preset model. Each target node in the target prescription network is used as a label to train the preset model, so that the preset model can learn the classification pattern of the target prescription network. When there is new prescription data to be analyzed, the preset model can determine the target node to which each prescription in the prescription data belongs in the target prescription network, without having to go through the similarity calculation and similar node merging calculation mentioned above again, thereby greatly improving the efficiency of prescription generation.
[0106] In some embodiments, the preset model may be a random forest, XGboost (eXtreme Gradient Boosting), or other similar models. This embodiment does not impose any specific restrictions on this.
[0107] In step S502, the step of obtaining the prescription data to be analyzed continues. Then, in step S503, the prescription data is input into a preset model to obtain the target nodes of each prescription in the prescription data in the target prescription network. In step S504, each prescription in the prescription data is merged into its corresponding target node.
[0108] In this embodiment, new prescription data to be analyzed is input into a preset model, and the preset model outputs the target nodes of each prescription in the prescription data within the target prescription network. It can be understood that the efficacy attributes and compatibility rules of the prescriptions are similar to those of the prescriptions within the target nodes.
[0109] In step S505, the step of determining the class prescription corresponding to the target node based on the nodes included in the target prescription network continues until the class prescription corresponding to the target node is obtained. In this embodiment, the class prescription mining of the target node is consistent with the method 400 above, and will not be described again here.
[0110] In some embodiments, the newly acquired prescription data to be analyzed may contain prescriptions that are the same as those in the target prescription network. Therefore, in order to avoid redundancy and reduced efficiency in the target prescription network, prescriptions that are the same as those in the target prescription network can be removed from the prescription data before being input into the preset model.
[0111] In some embodiments, the newly acquired prescription data to be analyzed may contain prescriptions that do not match any of the target nodes in the target prescription network. Therefore, in step S503 above, after inputting the prescription data into the preset model, if at least one prescription does not have a target node in the target prescription network, then for these prescriptions without target nodes and the previously acquired prescription data to be analyzed, the step of constructing an initial prescription network based on the efficacy attributes and / or compatibility information of each Chinese herbal medicine in the prescription is performed. Then, the step of merging similar nodes in the initial prescription network to construct the target prescription network is performed, thereby reconstructing the target prescription network that includes the target nodes of these prescriptions.
[0112] In some embodiments, after generating the class prescription, the correlation strength between the class prescription and at least one of symptoms, efficacy, and disease can be further analyzed. Specifically, first, the prescription to which each Chinese herb in the class prescription belongs is determined. Then, the user corresponding to the medical record document to which the prescription belongs is further determined. Next, the original target information is extracted from all the user's medical record documents, whereby the original target information includes at least one of symptoms, efficacy, and disease. This original target information records at least one of symptoms, efficacy, and disease throughout the user's entire treatment process. Then, based on the chronological order of the original target information's recording time, the target information that was originally recorded but disappeared in subsequent records is determined. After this, a data association matrix and a contingency table can be constructed based on the Chinese herb and target information included in the class prescription. Then, a chi-square test is performed on the data recorded in the data association matrix and contingency table to obtain the p (Probability Value) value between the class prescription and the target information (symptoms, efficacy, disease), which represents the correlation strength between the class prescription and the target information. If the p-value is less than a preset threshold, the prescription is determined to have a correlation effect on the target information. In other words, the prescription is determined to have an improving effect on the corresponding symptoms in the target information, a matching effect with the corresponding efficacy (i.e., acting on the pathogenesis through the efficacy), and a therapeutic effect on the corresponding disease. If the p-value is greater than or equal to the preset threshold, the prescription is determined to have no correlation effect on the target information. As an example, the preset threshold can be set to 0.05.
[0113] It is understandable that the target information that was originally recorded but disappeared in subsequent records includes at least one of the following: symptoms that have been cured, efficacy that no longer needs to be used due to improvement in the pathogenesis, and diseases that have been cured.
[0114] The data association matrix uses the Chinese herbs included in the formula as rows and the target information (symptoms, efficacy, disease) as columns. Cells in the matrix indicate whether the Chinese herbal medicine and the target information (symptoms, efficacy, disease) co-occur, typically using "1" to indicate co-occurrence and "0" to indicate no co-occurrence, thus visually representing the association between the Chinese herbal medicine and the target information (symptoms, efficacy, disease). The contingency table uses the Chinese herbal medicines in the formula as rows and the target information (symptoms, efficacy, disease) as columns. Cells in the contingency table indicate the specific number of times the Chinese herbal medicine and the target information (symptoms, efficacy, disease) co-occur, reflecting the frequency of co-occurrence between the two.
[0115] The above combination Figure 5 The present application provides a detailed description of a method 500 for generating similar prescriptions in some embodiments. This method uses prescription data to be analyzed as input to a preset model and each target node in the target prescription network as a label to train the preset model. The preset model outputs the target nodes of each prescription in the target prescription network based on the newly obtained prescription data. At this point, each prescription can be directly merged into the corresponding target node, saving the time spent on recalculating similarity and merging similar nodes, thus significantly improving the efficiency of generating similar prescriptions.
[0116] To implement the method steps described above in conjunction with the accompanying drawings at the software and hardware level, embodiments of this application also provide a processing apparatus, which can be as follows: Figure 6 The processing device shown. Figure 6 An exemplary structural block diagram of the processing apparatus 60 according to an embodiment of this application is shown, such as... Figure 6 As shown, the processing device 60 of this application may include a processor 610 and a memory 620. The memory 620 stores an executable program, which the processor 610 can load and execute, enabling the processing device 60 to implement any of the method steps described above.
[0117] In one example scenario, processor 610 can be used to control memory 620. Further, processor 610 can be a central processing unit (CPU), application processor (AP), or similar integrated within processing device 60; while memory 620, as hardware implementing storage functions, can be read-only memory (ROM), dynamic RAM (DRAM), or similar.
[0118] This application also provides a computer-readable storage medium storing program instructions that, when executed by a processor of a processing device, cause the processor to perform the method steps described in any embodiment of this application.
[0119] While numerous embodiments of this application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and intent of this application. It should be understood that various alternatives to the embodiments of this application described herein may be employed in the practice of this application. The appended claims are intended to define the scope of protection of this application and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for generating class squares, characterized in that, include: Acquire prescription data to be analyzed, wherein the prescription data includes multiple prescriptions, and the prescriptions include traditional Chinese medicine, efficacy attributes of traditional Chinese medicine and / or compatibility information of traditional Chinese medicine; Based on the efficacy attributes and / or compatibility information of each Chinese herbal medicine in the prescription, an initial prescription network is constructed; Similar nodes in the initial prescription network are merged to construct a target prescription network, wherein each node in the target prescription network is called a target node, and the target node includes multiple similar nodes; Based on the nodes included in the target node in the target prescription network, the class of prescription corresponding to the target node is determined.
2. The method according to claim 1, characterized in that, The process of acquiring the prescription data to be analyzed includes: Extract raw prescription data from medical record documents, wherein the raw prescription data includes multiple prescriptions, and the prescriptions include traditional Chinese medicine and / or the compatibility information of the traditional Chinese medicine; Extract the efficacy properties of each Chinese herbal medicine in the prescription from the pharmacopoeia; Map the aforementioned efficacy attributes to binary vectors; The binary vector is associated with each of the Chinese herbs in the prescription to obtain the prescription data to be analyzed.
3. The method according to claim 1, characterized in that, The step of constructing an initial prescription network based on the efficacy attributes and / or compatibility information of each Chinese herbal medicine in the prescription includes: The prescription data is divided into first prescription data and second prescription data based on whether or not the compatibility information is included, wherein each prescription in the first prescription data includes the compatibility information, and each prescription in the second prescription data does not include the compatibility information. In the first prescription data, the similarity between the prescriptions is determined based on the compatibility information of each Chinese herbal medicine in the prescription; In the second prescription data, the similarity between each prescription is determined based on the efficacy attributes of each Chinese herbal medicine in the prescription; Using prescriptions as nodes, undirected edges are established to connect prescription pairs that meet preset similarity conditions, and the similarity is used as the edge weight to construct the initial prescription network.
4. The method according to claim 3, characterized in that, In the first prescription data, determining the similarity between prescriptions based on the compatibility information of each Chinese herbal medicine in the prescription includes: Obtain the weights of the matching information; The weights of the compatibility information of each Chinese herb in the prescription are used as the weights of the Chinese herbs to construct a weighted vector of the prescription, wherein the weighted vector is the set of weights of each Chinese herb in the prescription; The similarity between the prescriptions is determined based on the weighted vectors of each prescription.
5. The method according to claim 3, characterized in that, In the second prescription data, determining the similarity between the various prescriptions based on the efficacy attributes of each Chinese herbal medicine in the prescription includes: The efficacy attribute set of the prescription is constructed based on the efficacy attributes corresponding to all Chinese medicines in the prescription. The efficacy attribute set includes a subset of efficacy attributes corresponding to the attribute type. The subset of efficacy attributes records the frequency of occurrence of each specific type of efficacy attribute corresponding to the attribute type. The similarity between the prescriptions is determined based on the set of efficacy attributes of each prescription.
6. The method according to claim 3, characterized in that, The preset condition is that the similarity is greater than a preset similarity threshold.
7. The method according to claim 1, characterized in that, The step of merging similar nodes in the initial prescription network to construct the target prescription network includes: Traverse each node of the initial prescription network and determine the gain between the currently traversed node and its neighboring nodes based on the edge weights of the nodes; If the gain is greater than the preset gain, then the currently traversed node is merged into the adjacent node; After the traversal is completed, the sum of the edge weights between the nodes connected by each edge in the initial prescription network is determined; The sum is then updated to reflect the edge weights between the nodes connected by each edge. Continue executing the step of traversing each node of the initial prescription network and determining the gain between the currently traversed node and its neighboring nodes based on the node's edge weight and the number of edges connected to the node, until the gain between any node in the initial prescription network is less than or equal to a preset gain, thereby constructing the target prescription network.
8. The method according to claim 1, characterized in that, The determination of the class of prescription corresponding to the target node based on the nodes included in the target prescription network includes: Determine the importance of each node included in the target node in the target prescription network; Determine the relative risk of each Chinese herbal medicine in each node of the target node in the target prescription network relative to other target nodes in the target prescription network; Based on the importance and the relative risk, the class of prescription corresponding to the target node in the target prescription network is determined.
9. The method according to claim 8, characterized in that, The step of determining the class of prescription corresponding to the target node in the target prescription network based on the importance and the relative risk includes: Based on the importance of each node included in the target node, a first preset number of nodes are selected from the target node to form a first prescription set; Based on the relative risk of each of the Chinese herbal medicines, a second preset number of target Chinese herbal medicines are selected from each of the nodes in the first prescription set; The intersection of the target Chinese medicines selected from each of the nodes is taken as the class formula of the target node.
10. The method according to claim 1, characterized in that, The method further includes: The prescription data to be analyzed is used as input to the preset model, and each target node in the target prescription network is used as a label to train the preset model. Continue with the steps described above for obtaining the prescription data to be analyzed; The prescription data is input into the preset model to obtain the target node of each prescription in the prescription data in the target prescription network; Merge each prescription in the prescription data into the corresponding target node; Continue executing the step of determining the class prescription corresponding to the target node based on the nodes included in the target prescription network, until the class prescription corresponding to the target node is obtained.