Method and electronic device for checking drug interactions
By analyzing medical record data using the implicit Dirichlet distribution model, high-risk drug combinations are screened out, solving the problem of low efficiency in examining drug interactions in existing technologies and achieving rapid and accurate risk identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ACER INC
- Filing Date
- 2022-06-14
- Publication Date
- 2026-05-08
AI Technical Summary
Current technology makes it difficult to quickly and effectively identify high-risk drug combinations, which may lead to unexpected hospitalizations for patients.
The Latent Dirichlet Allocation (LDA) model was used to analyze medical record data. By calculating the odds ratio and score of drug combinations, high-risk drug combinations were screened out, and the data was processed and output using a processor and transceiver.
High-risk combinations were identified from a large number of medication combinations, and the risks were confirmed to originate from drug interactions, thus improving the efficiency and accuracy of the examination.
Smart Images

Figure CN116779094B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and electronic device for examining drug interactions. Background Technology
[0002] Patients often need to take multiple medications at the same time. These medications may interact, leading to serious adverse reactions and unexpected hospitalizations. To avoid this, it is necessary to examine the interactions of various medication combinations. However, the number of medication combinations is vast, making it highly inefficient to examine each combination individually. Therefore, developing a method to rapidly identify high-risk medication combinations is one of the goals pursued by those in the field. Summary of the Invention
[0003] This invention provides a method and electronic device for examining drug interactions, which can output high-risk drug combinations for user reference.
[0004] A method for examining drug interactions according to the present invention includes: obtaining a plurality of medical records, wherein at least one of the plurality of medical records indicates whether a patient taking a first drug combination has experienced a hospitalization event; generating a set of drug combinations based on the plurality of medical records, wherein the set of drug combinations includes a first drug combination, a second drug combination, and a third drug combination, wherein both the first drug combination and the second drug combination include a first drug, and both the first drug combination and the third drug combination include a second drug; generating a first ratio between the first drug combination and the hospitalization event, a second ratio between the second drug combination and the hospitalization event, and a third ratio between the third drug combination and the hospitalization event based on the plurality of medical records; generating a first score corresponding to the first drug based on the second ratio, wherein the first score is negatively correlated with the second ratio; generating a second score corresponding to the second drug based on the third ratio, wherein the second score is negatively correlated with the third ratio, wherein the first score is greater than or equal to the second score; and outputting the first drug combination in response to the first ratio being greater than a first threshold, the sum of the first score and the second score being greater than the second threshold, and the quotient of the first score and the second score being less than a third threshold.
[0005] In one embodiment of the present invention, the step of generating a first score corresponding to the first drug based on a second ratio includes: marking a second drug combination in response to the second ratio being greater than a risk threshold; generating a third score based on the marked second drug combination, wherein the third score is equal to the number of drug combinations in the drug combination set that contain the first drug but not the second drug and are marked, divided by the number of drug combinations in the drug combination set that contain the first drug but not the second drug; and calculating a first score based on the third score, wherein the sum of the first score and the third score is equal to one.
[0006] In one embodiment of the present invention, the step of generating a set of medication combinations based on multiple medical records includes: performing a screening process to generate a first unique set of medication combinations, comprising: generating K topic vectors containing a first topic vector based on multiple medical records and an implicit Dirichlet distribution model, where K is the number of first topics, the K topic vectors respectively correspond to K topics, the K topics contain a first topic corresponding to the first topic vector, and the first topic vector contains the probability distribution of all medication combinations; starting with the medication combination with the highest probability, selecting multiple important medication combinations from the first topic vectors to generate a first important set of medication combinations; determining a first unique set of medication combinations based on the first important set of medication combinations; and generating a set of medication combinations based on the first unique set of medication combinations.
[0007] In one embodiment of the present invention, the aforementioned K topics include a second topic, wherein the step of determining the first unique drug combination set based on the first important drug combination set includes: in response to the first important drug combination being included in the first important drug combination set corresponding to the first topic and the second important drug combination set corresponding to the second topic, deleting the first important drug combination from the first important drug combination set to generate the first unique drug combination set.
[0008] In one embodiment of the present invention, the step of generating a drug combination set based on a first unique drug combination set includes: repeatedly performing a screening process multiple times to generate a plurality of unique drug combination sets containing the first unique drug combination set; in response to the number of first drug combinations in the plurality of unique drug combination sets being greater than a quantity threshold, generating a first stable drug combination set corresponding to a first topic based on the first drug combinations; and generating a drug combination set based on the first stable drug combination set.
[0009] In one embodiment of the present invention, the step of generating a medication combination set based on a first stable medication combination set includes: generating multiple medical record vectors corresponding to multiple medical records based on multiple medical records and an implicit Dirichlet distribution model, wherein each of the multiple medical record vectors contains a probability distribution of K topics; determining a set of medical records corresponding to a first topic among the multiple medical records based on the probability distribution of the K topics; calculating the ratio of at least one medical record in the medical record set to the total medical record set, wherein the at least one medical record indicates at least one medication combination in the first stable medication combination set; and generating a medication combination set based on the first stable medication combination set in response to the ratio being greater than a ratio threshold, wherein the medication combination set contains multiple medication combinations in the first stable medication combination set.
[0010] In one embodiment of the present invention, the first medical record in the above-mentioned medical record set corresponds to a first probability distribution of K topics, wherein the step of determining the set of medical records corresponding to the first topic among multiple medical records according to the probability distribution of K topics includes: in response to the maximum probability in the first probability distribution corresponding to the first topic, determining that the first medical record corresponds to the first topic.
[0011] In one embodiment of the present invention, the above method further includes: generating a first index corresponding to a first number of topics and a second index corresponding to a second number of topics based on multiple medical records and an implicit Dirichlet distribution model; and comparing the first index and the second index to select the first number of topics from the first number of topics and the second number of topics as K.
[0012] In one embodiment of the present invention, the step of generating a first index corresponding to a first number of topics includes: generating K topic vectors based on multiple medical records, a latent Dirichlet distribution model, and a first number of topics; and calculating the average similarity of all 2-combinations of the K topic vectors as the first index.
[0013] In one embodiment of the present invention, the step of generating a first indicator corresponding to the number of first topics includes: generating multiple medical record vectors corresponding to the multiple medical records based on multiple medical records, an implicit Dirichlet distribution model, and the number of first topics, wherein each of the multiple medical record vectors contains a probability distribution of K topics; determining at least one medical record among the multiple medical records that corresponds to the first topic based on the probability distribution of the K topics; and calculating a ratio based on the number of at least one medical record and the total number of multiple medical records as the first indicator.
[0014] In one embodiment of the present invention, the step of determining at least one medical record corresponding to the first topic among multiple medical records based on the probability distribution of K topics includes: obtaining a first probability distribution of K topics corresponding to at least one medical record from multiple medical record vectors; and determining that at least one medical record corresponds to the first topic in response to the maximum probability in the first probability distribution corresponding to the first topic being greater than a probability threshold.
[0015] In one embodiment of the present invention, the step of generating a first index corresponding to the number of first topics includes: generating multiple medical record vectors corresponding to the multiple medical records based on multiple medical records, an implicit Dirichlet distribution model, and the number of first topics, wherein each of the multiple medical record vectors contains a probability distribution of K topics; dividing the multiple medical records into K groups based on the probability distribution of the K topics, wherein the K groups correspond to K topics respectively; calculating a first statistical value of the distance between the K groups; calculating a second statistical value of the distance within the K groups; and calculating the ratio of the first statistical value to the second statistical value as the first index.
[0016] In one embodiment of the present invention, the step of calculating the first statistical value of the distance between groups based on K groups includes: calculating multiple distances between K topic vectors; and adding the multiple distances together to obtain the first statistical value.
[0017] In one embodiment of the present invention, the aforementioned K groups include a first group and a second group, wherein the step of calculating a second statistical value of the distance within the groups based on the K groups includes: calculating multiple distances between multiple elements in the first group to generate a first group intra-distance sum corresponding to the first group; and adding the first group intra-distance sum corresponding to the first group to the second group intra-distance sum corresponding to the second group to obtain the second statistical value.
[0018] An electronic device for detecting drug interactions according to the present invention includes a processor and a transceiver. The processor is coupled to the transceiver and configured to perform: acquiring a plurality of medical records via the transceiver, wherein at least one of the plurality of medical records indicates whether a patient taking a first drug combination has experienced a hospitalization event; generating a set of drug combinations based on the plurality of medical records, wherein the set of drug combinations includes a first drug combination, a second drug combination, and a third drug combination, wherein both the first and second drug combinations contain a first drug, and both the first and third drug combinations contain a second drug; and generating a first ratio between the first drug combination and the hospitalization event, and a ratio between the second drug combination and the hospitalization event based on the plurality of medical records. The second ratio between events and the third ratio between the third medication combination and the hospitalization events; generating a first score corresponding to the first drug based on the second ratio, wherein the first score is negatively correlated with the second ratio; generating a second score corresponding to the second drug based on the third ratio, wherein the second score is negatively correlated with the third ratio, wherein the first score is greater than or equal to the second score; and outputting the first medication combination via transceiver in response to the first ratio being greater than a first threshold, the sum of the first score and the second score being greater than the second threshold, and the quotient of the first score and the second score being less than a third threshold.
[0019] Based on the above, the present invention can screen out high-risk drug combinations from a large number of drug combinations, and can confirm that the reason why the drug combination is high-risk does not come from the drugs in the drug combination themselves, but from the drug interactions. Attached Figure Description
[0020] Figure 1 A schematic diagram of an electronic device for detecting drug interactions is shown according to an embodiment of the present invention;
[0021] Figure 2 A schematic diagram illustrating the relationship between the average similarity between topics and the number of topics K is shown according to an embodiment of the present invention;
[0022] Figure 3A schematic diagram illustrating the relationship between the ratio of medical records with a particular theme and the number of themes K, according to an embodiment of the present invention;
[0023] Figure 4 A schematic diagram illustrating the relationship between clustering effectiveness index and the number of topics K is shown according to an embodiment of the present invention;
[0024] Figure 5 A schematic diagram illustrating the use of the elbow method to identify important drug combinations is shown in one embodiment of the present invention;
[0025] Figure 6 A flowchart of a method for examining drug interactions is shown according to an embodiment of the present invention.
[0026] Explanation of reference numerals in the attached figures
[0027] 100: Electronic devices;
[0028] 110: Processor;
[0029] 120: Storage medium;
[0030] 130: Transceiver;
[0031] 50: Curve;
[0032] S601, S602, S603, S604, S605, S606: Steps. Detailed Implementation
[0033] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same element references are used in the drawings and description to denote the same or similar parts.
[0034] Figure 1 A schematic diagram of an electronic device 100 for detecting drug interactions is shown according to an embodiment of the present invention. The electronic device 100 may include a processor 110, a storage medium 120, and a transceiver 130.
[0035] Processor 110 may be, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microcontroller (MCU), microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), graphics processing unit (GPU), image signal processor (ISP), image processing unit (IPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar components or combinations thereof. Processor 110 may be coupled to storage medium 120 and transceiver 130, and access and execute multiple modules and various applications stored in storage medium 120.
[0036] Storage medium 120 may be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar components or combinations thereof, for storing multiple modules or various applications that can be executed by processor 110.
[0037] Transceiver 130 transmits and receives signals wirelessly or via a wired connection. Transceiver 130 can also perform operations such as low-noise amplification, impedance matching, mixing, up or down frequency conversion, filtering, amplification, and similar functions.
[0038] Processor 110 can obtain N medical records corresponding to N patients via transceiver 130, where N is a positive integer. The medical records can indicate the medication combinations the patient has taken, where a combination may contain two drugs. For example, medication combination (α,β) may contain drug α and drug β. The medical records can also indicate whether the patient experienced an unexpected hospitalization event. Table 1 shows a schematic diagram of the N medical records. Taking medical record #1 as an example, medical record #1 indicates that the patient corresponding to medical record #1 has taken combinations of drugs A and B (A,B), drugs X and Y (X,Y), and drugs U and W (U,W). Medical record #1 also indicates that the patient corresponding to medical record #1 experienced an unexpected hospitalization event. That is, the medication combinations the patient previously took may have led to an unexpected hospitalization event. Taking medical record #N as an example, medical record #N indicates that the patient corresponding to medical record #N has taken combinations of drugs C and D (C,D) and drugs O and P (O,P). Medical record #N also indicates that the patient corresponding to medical record #N has not experienced any unexpected hospitalization events. In other words, the combination of medications the patient previously took did not result in any unexpected hospitalization events. Medication combinations on the medical record are recorded, for example, in the form of anatomical therapeutic chemical (ATC) codes.
[0039] Table 1
[0040] Medical records Medication combination Hospitalization incident #1 (A,B),(X,Y),(U,W) yes #2 (G,H) yes … … … #N (C,D),(O,P) no
[0041] In this embodiment, it is assumed that N medical records contain a total of M medication combinations, where M is a positive integer. Taking Table 1 as an example, the M medication combinations include at least the combinations (A,B), (X,Y), (U,W), (G,H), (C,D), and (O,P). The M medication combinations are the union of all medication combinations in the N cases.
[0042] Processor 110 can generate a set of medication combinations containing multiple medication combinations based on N medical records, and then select medication combinations with high-risk interactions from the set for user reference. First, processor 110 can analyze the N medical records using a Latent Dirichlet Allocation (LDA) model. The parameters of the LDA model can include the number of topics K, where K is a positive integer. K determines that the output of the LDA model is related to K topics. Processor 110 can first determine the value of the optimal number of topics Kopt.
[0043] Specifically, the output of the LDA model can be associated with topics and words. The processor 110 can input N medical records into the LDA model to generate K topic vectors corresponding to K topics (or medication patterns). Each topic vector can contain the probability distribution of all medication combinations (i.e., M medication combinations). Medication combinations are the words in the LDA model. The topic vector contains the probability distribution of all words. In other words, a topic vector can be a vector containing M probabilities, where each of the M probabilities corresponds to one of the M medication combinations.
[0044] Table 2 shows an example of K topic vectors. Taking the topic vector corresponding to topic #1 as an example, the topic vector may contain at least the probability value "0.20" corresponding to medication combination (A,B), the probability value "0.05" corresponding to medication combination (C,D), and the probability value "0.20" corresponding to medication combination (X,Y). The sum of all elements in the topic vector (i.e., the M probabilities) is equal to "1".
[0045] Table 2
[0046] serial number Medication combination Topic #1 Topic #2 … Topic #K #1 (A,B) 0.20 0.00 … 0.10 #2 (C,D) 0.05 0.10 … 0.20 … … … … … … #M (X,Y) 0.20 0.15 … 0.05
[0047] On the other hand, the LDA model can also generate N case vectors corresponding to the N case records. Each case vector can contain a probability distribution of K topics. In other words, a case vector can be a vector containing K probabilities, where each of the K probabilities corresponds to one of the K topics. Table 3 shows an example of N case vectors. Taking the case vector corresponding to case #1 as an example, the case vector can contain at least the probability value "0.20" corresponding to topic #1, the probability value "0.00" corresponding to topic #2, and the probability value "0.10" corresponding to topic #K. The sum of all elements (i.e., the K probabilities) in the case vector is equal to "1".
[0048] Table 3
[0049] Medical records Topic #1 Topic #2 … Topic #K #1 0.20 0.00 … 0.10 #2 0.05 0.10 … 0.20 … … … … … #N 0.00 0.05 … 0.15
[0050] The processor 110 can determine the optimal number of topics Kopt based on factors such as the similarity between topics, the ratio of cases with a biased topic (i.e., a biased medication pattern), and the grouping effectiveness index.
[0051] To identify different medication patterns, the greater the difference between topics, the better. That is, the lower the similarity between topics, the better. In one embodiment, processor 110 can compute all 2-combinations (a total of K topic vectors) of the K topic vectors. The average similarity of the three topics (Kopt) is used as an indicator to determine the optimal number of topics. The similarity can be, for example, cosine similarity or Jaccard similarity, but this disclosure is not limited to these. Taking Table 2 as an example, assume K equals "3". The processor 110 can calculate three similarities: the similarity between the topic vector [0.20 0.05 … 0.20] corresponding to topic #1 and the topic vector [0.00 0.10 … 0.15] corresponding to topic #2; the similarity between the topic vector [0.20 0.05 … 0.20] corresponding to topic #1 and the topic vector [0.10 0.20 … 0.05] corresponding to topic #3; and the similarity between the topic vector [0.00 0.10 … 0.15] corresponding to topic #2 and the topic vector [0.10 0.20 … 0.05] corresponding to topic #3. The average of these three similarities is then calculated to obtain the average similarity, as shown in Table 4.
[0052] Table 4
[0053]
[0054] Figure 2 A schematic diagram illustrating the relationship between the average similarity between topics and the number of topics K is shown in one embodiment of the present invention. Comparing the average similarity corresponding to each number of topics K reveals that the larger the value of K, the smaller the average similarity between topics. Therefore, if the average similarity is used as the criterion for determining the optimal number of topics Kopt, the processor 110 can select a larger value as the optimal number of topics Kopt.
[0055] To ensure that each distinct medical record (or patient) is categorized into a representative medication pattern, a higher ratio of medical records with a particular theme is desirable. In one embodiment, processor 110 may calculate a ratio based on the number of medical records corresponding to a specific theme to the total number of all medical records (i.e., N) as an indicator to determine the optimal number of themes, Kopt. Specifically, processor 110 may determine whether a medical record vector is biased towards a specific theme based on the probability distribution of K themes in the medical record vector. If the maximum probability in the probability distribution corresponds to a specific theme and the maximum probability is greater than a probability threshold, then processor 110 may determine that the medical record vector (or medical record) is biased towards the specific theme.
[0056] Table 5
[0057] theme Medical Record #1 Medical Record #2 Medical Record #3 Medical Record #4 Medical Record #5 #1 0.60 0.05 0.40 0.10 0.35 #2 0.30 0.70 0.25 0.20 0.30 #3 0.10 0.25 0.35 0.70 0.35
[0058] Table 5 provides an example of multiple medical record vectors, where N is assumed to be "5", K to be "3", and the probability threshold to be "0.50". Taking medical record #1 as an example, processor 110 can determine that medical record #1 is biased towards topic #1 in response to the maximum probability "0.60" in medical record #1 corresponding to topic #1 and being greater than the probability threshold "0.50". Taking medical record #3 as an example, processor 110 can determine that medical record #3 is not biased towards any topic in response to the maximum probability "0.40" in medical record #3 being less than or equal to "0.50". And so on, processor 110 can obtain the biased topic for each of all medical records based on the data in Table 5, as shown in Table 6.
[0059] Table 6
[0060] Medical Record #1 Medical Record #2 Medical Record #3 Medical Record #4 Medical Record #5 Emphasis on theme Topic #1 Topic #2 none Topic #3 none
[0061] After determining the theme emphasized in each medical record, processor 110 can calculate the ratio of the number of medical records corresponding to a specific theme to the total number of all medical records as an indicator. For example, in Table 6, the number of medical records corresponding to a specific theme (i.e., medical record #1, medical record #2, and medical record #4) is equal to "3" and the total number of all medical records N is equal to "5". Processor 110 can calculate the ratio "3 / 5" as an indicator to determine the optimal number of themes Kopt.
[0062] Figure 3 A schematic diagram illustrating the relationship between the ratio of medical records with a biased theme and the number of topics K is shown in one embodiment of the present invention. Comparing the ratios corresponding to each topic number K reveals that the smaller the value of K, the larger the ratio of medical records with a biased theme. Therefore, if the ratio of medical records with a biased theme is used as the indicator for determining the optimal number of topics Kopt, the processor 110 can select a smaller value as the optimal number of topics Kopt.
[0063] In one embodiment, processor 110 can use a clustering performance metric as an indicator to determine the optimal number of topics Kopt. First, processor 110 can assign medical records to groups of specific topics. Specifically, processor 110 can assign the medical records to one of K groups based on the probability distribution of K topics in the case vector corresponding to the medical record, where each of the K groups corresponds to one of the K topics. For example, processor 110 can assign medical records to the group corresponding to the topic with the highest probability in the medical record vector. Using Table 5 as an example, processor 110 can assign medical record #1 to the group corresponding to topic #1, medical record #2 to the group corresponding to topic #2, medical record #3 to the group corresponding to topic #1, and medical record #4 to the group corresponding to topic #3. If there are multiple maximum probabilities in the medical record vector, processor 110 can assign the medical records to one of the topics corresponding to the multiple maximum probabilities according to a default rule or randomly. Taking Table 5 as an example, the processor 110 can assign medical record #5 to one of the groups corresponding to topic #1 and topic #3 according to default rules or randomly.
[0064] After assigning N cases to groups, processor 110 can calculate a first statistic corresponding to the inter-group distance and a second statistic corresponding to the intra-group distance based on K groups. The grouping effectiveness index can be equal to the ratio of the first statistic to the second statistic.
[0065] The first statistic is, for example, the sum of multiple distances between K topic vectors. Specifically, the first statistic can be the sum of distances between all 2-groups of the K groups. For example, suppose K equals "3" and the K groups include group #1, group #2, and group #3. Processor 110 can calculate three distances: the distance between group #1 and group #2, the distance between group #1 and group #3, and the distance between group #2 and group #3, and calculate the sum of the three distances to obtain the first statistic. The distance can be calculated based on the distance between topic vectors. For example, the distance between group #1 corresponding to topic #1 and group #2 corresponding to topic #2 can be equal to the distance between two topic vectors: the topic vector corresponding to topic #1 (e.g., [0.20 0.05…0.20] in Table 2) and the topic vector corresponding to topic #2 (e.g., [0.00 0.10…0.15] in Table 2). The larger the distance between groups, the better the clustering performance. Therefore, the first statistical value can be directly proportional to the cluster effectiveness index.
[0066] The second statistical value is, for example, the sum of the K intra-group distances corresponding to K groups. The intra-group distance of each group can be equal to the sum of multiple distances between multiple elements in the group. More specifically, the intra-group distance corresponding to a group can be the sum of the distances of all 2-combinations of all elements in the group. Taking group #1 corresponding to topic #1 as an example, assuming that group #1 contains three elements: medical record #1, medical record #2, and case #3 (i.e., three out of N cases correspond to topic #1), processor 110 can calculate the distance between medical record #1 and medical record #2, the distance between medical record #1 and medical record #3, and the distance between medical record #2 and medical record #3, and add the three distances together to obtain the intra-group distance of group #1.
[0067] In one embodiment, processor 110 can vectorize the medication combinations recorded in the medical records to calculate the distance between medical records. Taking medical records #1 and #2 in Table 1 as examples, if processor 110 wants to calculate the distance between medical records #1 and #2, processor 110 can convert the medication combination "(A,B),(X,Y),(U,W)" of medical record #1 into a vector and convert the medication combination "(G,H)" of medical record #2 into another vector. Processor 110 can calculate the distance between the two vectors as the distance between medical records #1 and #2. The smaller the distance between elements within a group, the better the clustering performance. Therefore, the second statistic can be inversely proportional to the clustering performance index.
[0068] Figure 4 A schematic diagram illustrating the relationship between clustering performance index and the number of topics K is shown in one embodiment of the present invention. By comparing the ratios corresponding to the various number of topics K, it can be found that the larger the value of K, the larger the clustering performance index. Therefore, if the clustering performance index is used as the criterion for determining the optimal number of topics Kopt, the processor 110 can select a larger value as the optimal number of topics Kopt.
[0069] Processor 110 can be based on Figure 2 , Figure 3 and Figure 4 The optimal number of topics, Kopt, is determined. After determining the optimal number of topics, Kopt, the processor 110 can set the optimal number of topics, Kopt, as the number of topics, K in the LDA model parameters. Then, the processor 110 can generate K topic vectors corresponding to K topics (as shown in the example in Table 2) and N case vectors corresponding to N cases (as shown in the example in Table 3) based on the number of topics, K, the LDA model, and N cases.
[0070] Processor 110 can execute a screening process to generate a unique set of medication combinations for each of the K topics. Specifically, processor 110 can select one or more important medication combinations based on the probability distribution of all medication combinations (i.e., M medication combinations) in the topic vector to generate a set of important medication combinations corresponding to the topic vector. In one embodiment, processor 110 can use the elbow method to select multiple important medication combinations from the topic vector, starting with the medication combination with the highest probability, to generate a set of important medication combinations.
[0071] Table 7 shows an example of an important set of drug combinations. Taking topic #1 as an example, processor 110 can arrange the drug combinations according to their probability based on the topic vector corresponding to topic #1. Drug combinations with higher probabilities are placed earlier. Next, processor 110 can use the elbow method to find the inflection point of the arranged drug combinations and select the drug combinations placed before the inflection point as important drug combinations. Figure 5 According to an embodiment of the present invention, a schematic diagram is shown for identifying important drug combinations using the elbow method. Processor 110 can arrange drug combinations in a specific theme (e.g., theme #1) according to their probability to plot a curve 50. If the inflection point is the fourth drug combination (i.e., the drug combination with the fourth highest probability), processor 110 can select the top four drug combinations (i.e., the drug combinations with the top four highest probabilities) from the M drug combinations to generate a set of important drug combinations. In the theme vector of theme #1, drug combination (A,B) has the highest probability, drug combination (C,D) has the second highest probability, drug combination (E,F) has the third highest probability, and drug combination (G,H) has the fourth highest probability. If drug combination (G,H) corresponds to the inflection point, processor 110 can select the drug combinations arranged before drug combination (G,H) to generate a set of important drug combinations.
[0072] Table 7
[0073] theme Important drug combination collection #1 (A,B),(C,D),(E,F),(G,H) #2 (G,H),(I,J),(X,Y),(S;W) … … #K (S;W),(O,P),(Q,R)
[0074] After obtaining the K sets of important drug combinations corresponding to K themes, processor 110 can generate K sets of unique drug combinations corresponding to the K themes. Specifically, processor 110 can delete drug combinations that appear repeatedly in different sets of important drug combinations to generate unique drug combination sets. Taking Table 7 as an example, processor 110 can delete drug combination (G,H) from the important drug combination set of theme #1 and from the important drug combination set of theme #2 in response to drug combination (G,H) being included in both the important drug combination set corresponding to theme #1 and the important drug combination set corresponding to theme #2, thereby generating unique drug combination sets, as shown in Table 8. Table 8 can be the result of processor 110 performing the first screening process.
[0075] Table 8
[0076]
[0077] Because the LDA algorithm is probabilistic, the important and unique drug combinations selected in the aforementioned screening process may only appear by chance in this screening process. To ensure that the screened drug combinations are stable, the processor 110 can repeat the screening process multiple times. Specifically, the processor 110 can execute the screening process multiple times to generate multiple sets of unique drug combinations. In response to the number of a specific drug combination in the multiple sets of unique drug combinations exceeding a certain threshold, the processor 110 can generate a stable set of drug combinations based on the specific drug combination, wherein the stable set of drug combinations corresponds to the same theme as the specific drug combination.
[0078] Taking drug combination (A,B) from the unique drug combination set corresponding to theme #1 as an example, assume that processor 110 executes 10 screening processes with a quantity threshold of "6". If the screening processes in which drug combination (A,B) appears are as shown in Table 9, then processor 110 can determine that drug combination (A,B) is stable in response to the fact that the number of times drug combination (A,B) appears in the 10 screening processes (i.e., 7 times) is greater than the quantity threshold. Accordingly, processor 110 can generate a stable drug combination set corresponding to theme #1 based on drug combination (A,B), where the stable drug combination set corresponding to theme #1 is, for example, drug combinations (A,B), (C,D), and (E,F) in Table 8.
[0079] Table 9
[0080] Screening process #1 #2 #3 #4 #5 #6 #7 #8 #9 #10 Does (A,B) exist? yes yes yes no yes no yes yes no yes
[0081] Table 10 provides examples of stable medication combinations for each theme. After generating a set of K stable medication combinations corresponding to K themes, processor 110 can verify whether each stable medication combination set conforms to the medication patterns of a sufficient number of people. Specifically, processor 110 can determine the theme corresponding to a case based on the maximum probability in the probability distribution of the K themes in the case record vector. If the maximum probability in the case record vector corresponds to a specific theme, processor 110 can determine that the case corresponds to the specific theme. If the case record vector contains multiple maximum probabilities, processor 110 can determine, according to default rules or randomly, that a theme corresponding to one of the multiple maximum probabilities corresponds to a case. After completing the determination, each theme can correspond to a set of case records containing at least one case. For example, if 10 out of N cases correspond to theme #1, then the set of case records corresponding to theme #1 contains 10 cases.
[0082] Table 10
[0083] theme Stable medication combination collection #1 (A,B),(C,D),(E,F) #2 (I,J),(X,Y) … … #K (O,P),(Q,R)
[0084] Processor 110 can calculate the ratio of at least one case in a set of medical records, where the at least one case indicates at least one medication combination in a set of stable medication combinations. If the ratio is greater than a ratio threshold, it means that the number of medical records that conform to the theme (i.e., medication pattern) corresponding to the set of stable medication combinations is sufficient. Accordingly, processor 110 can generate a final set of medication combinations based on the set of stable medication combinations. For example, suppose the ratio threshold is 50%. If 60% of the medical records out of N cases contain at least one medication combination from the set of stable medication combinations corresponding to theme #1, it means that the number of samples (i.e., medical records) conforming to the medication pattern of theme #1 is sufficient. Therefore, processor 110 can generate a final set of medication combinations based on the set of stable medication combinations corresponding to theme #1. Table 11 shows an example of a set of medical records corresponding to theme #1. Conversely, if only 40% of the medical records out of N cases contain at least one medication combination from the set of stable medication combinations corresponding to theme #1, it means that the number of samples conforming to the medication pattern of theme #1 is insufficient. Therefore, processor 110 may not generate the final set of medication combinations based on the set of stable medication combinations corresponding to topic #1.
[0085] Assume that the medical record set corresponding to topic #1 contains at least medical records #10, #11, and #12 (i.e., the dominant topic of medical records #10, #11, and #12 is topic #1). Referring to Tables 10 and 11, since the medication combination (C,D) recorded in medical record #10 appears in the stable medication combination set of topic #1, processor 110 can determine that medical record #10 is one of the aforementioned at least one case. Since the medication combination (E,F) recorded in medical record #11 appears in the stable medication combination set of topic #1, processor 110 can determine that medical record #11 is one of the aforementioned at least one case. Since the medication combination (X,Y) recorded in medical record #12 does not appear in the stable medication combination set of topic #1, processor 110 can determine that medical record #12 is not one of the aforementioned at least one case.
[0086] Table 11
[0087]
[0088] Processor 110 can check, according to the steps described above, whether each of the K stable medication combination sets conforms to the medication pattern of a sufficient number of patients. If the stable medication combination set conforms to the medication pattern of a sufficient number of patients, processor 110 can retain the stable medication combination set. If the stable medication combination set does not conform to the medication pattern of a sufficient number of patients, processor 110 can delete the stable medication combination set. Accordingly, processor 110 can select k stable medication combination sets from the K stable medication combination sets corresponding to K topics, where k is a positive integer less than or equal to K. Processor 110 can take the union of the k stable medication combination sets to obtain the final medication combination set, wherein the medication combination set may contain multiple medication combinations. Each medication combination in the final medication combination set has characteristics such as high importance, high uniqueness, and high stability, and conforms to the medication pattern of a large number of patients.
[0089] After obtaining the final set of medication combinations, processor 110 can label the risk level of the medication combinations in the set. Specifically, processor 110 can calculate the odds ratio (OR) for a specific medication combination based on N cases, as shown in equation (1) and the confusion matrix in Table 12, where e1y1 represents the number of cases in the N cases that have taken the medication combination and had a hospitalization event, e1y0 represents the number of cases in the N cases that have taken the medication combination but had no hospitalization event, e0y1 represents the number of cases in the N cases that have not taken the medication combination but had a hospitalization event, and e0y0 represents the number of cases in the N cases that have not taken the medication combination and had no hospitalization event.
[0090]
[0091] Table 12
[0092] N cases Hospitalization occurred No hospitalizations have occurred. Previously took this medication combination <![CDATA[e1y1]]> <![CDATA[e1y0]]> This medication combination has not been taken. <![CDATA[e0y1]]> <![CDATA[e0y0]]>
[0093] After calculating the odds ratio for each medication combination in the set of medication combinations, processor 110 can assign a risk level to the medication combinations based on the odds ratio. If the odds ratio of a medication combination is greater than a risk threshold, it means that the medication combination is very likely to be the cause of the hospitalization event. Accordingly, processor 110 can mark the medication combination as high-risk. Conversely, if the odds ratio of a medication combination is less than or equal to the risk threshold, it means that the medication combination is less relevant to the occurrence of the hospitalization event. Accordingly, processor 110 can mark the medication combination as low-risk. Table 13 shows an example of risk level marking for medication combinations. Assuming the risk threshold is "1.3", processor 110 can mark medication combinations with an odds ratio greater than "1.3" as high-risk and medication combinations with an odds ratio less than or equal to "1.3" as low-risk.
[0094] Table 13
[0095] Medication combination Ratio mark (A,C) 1.5 High risk (A,D) 0.9 Low risk (A,E) 1.1 Low risk (A,F) 2.7 High risk (B,G) 2.7 High risk (B,H) 0.8 Low risk (B,I) 6.1 High risk (B,J) 3.0 High risk
[0096] After obtaining the label of each drug combination in the drug combination set, the processor 110 can generate a risk combination fraction (RCF) corresponding to a drug in the drug combination based on the label of the drug combination. Taking drug combination (α,β) as an example, in order to confirm that the combination of drug α in drug combination (α,β) with other drugs in the drug combination set (i.e., drugs other than drug β) is safe, the processor 110 can calculate the risk combination fraction (RCF) (or "third fraction") corresponding to drug α according to equation (2), where S(α,β) is the number of drug combinations in the drug combination set that contain drug α but not drug β, and S′(α,β) is the number of drug combinations in the drug combination set that contain drug α but not drug β and are marked as high risk.
[0097]
[0098] After obtaining the RCF of drug α, processor 110 can calculate the normal combination fraction (NCF) of drug α (or "first fraction" or "second fraction") according to equation (3). The higher the normal combination fraction of drug α, the lower the risk of drug α in combination with other drugs besides drug β. The NCF of drug α can be negatively correlated with the ratio of drug combinations that include drug α but do not include drug β. Taking Table 13 as an example, the NCF of drug A can be negatively correlated with the ratio of drug combinations (A,C), (A,D), (A,E), or (A,F).
[0099] NCF = 1 - RCF…(3)
[0100] Taking drug A in Table 13 as an example, assume that Table 13 contains all drug combinations except for drug combination (A,B) in the drug combination set. There are two drug combinations that contain drug A and are marked as high-risk, namely drug combination (A,C) and (A,F). Accordingly, processor 110 can calculate that the RCF of drug A is equal to "0.5" and the NCF of drug A is equal to "0.5" according to equations (2) and (3). Taking drug B in Table 13 as an example, there are three drug combinations that contain drug B and are marked as high-risk, namely drug combination (B,G), (B,I) and (B,J). Accordingly, processor 110 can calculate that the RCF of drug B is equal to "0.75" and the NCF of drug B is equal to "0.25" according to equations (2) and (3).
[0101] Assuming the NCF of drug α is greater than or equal to the NCF of drug β, after obtaining the NCFs of drug α and drug β, processor 110 can calculate the quotient (or ratio) Q(α,β) of the NCFs of drug α and drug β, as shown in equation (4), where NCF(α) is the NCF of drug α and NCF(β) is the NCF of drug β. The value of Q(α,β) is greater than or equal to 1, and the lower the value, the closer the risk of drug α in combination with other drugs besides drug β and the risk of drug β in combination with other drugs besides drug α are.
[0102]
[0103] Processor 110 can determine whether a specific drug combination has a high-risk interaction based on the following three conditions. Taking drug combination (α,β) as an example, if the ratio of drug combination (α,β) is greater than a first threshold, the sum of the NCF(α) of drug α and the NCF(β) of drug β (or the total NCF) is greater than a second threshold, and the quotient of the NCF(α) of drug α and the NCF(β) of drug β (or the NCF quotient) is less than a third threshold, then processor 110 can determine that drug combination (α,β) has a high-risk interaction. Processor 110 can output drug combination (α,β) through transceiver 130 for user reference.
[0104] Table 14 shows the relevant parameters of the ratio and NCF for multiple drug combinations. Assume the first threshold is "2.0", the second threshold is "1.2", and the third threshold is "1.8". Since the ratio of drug combination (E,F) is greater than the first threshold, the sum of NCF is greater than the second threshold, and the NCF quotient is less than the third threshold, processor 110 can determine that drug combination (E,F) fully meets the three conditions. Therefore, processor 110 can output drug combination (E,F). Since the ratio of drug combination (G,H) is less than the first threshold, processor 110 can determine that drug combination (G,H) does not fully meet the three conditions. Therefore, processor 110 may not output drug combination (G,H). Since the sum of NCF for drug combination (A,B) is less than the second threshold or the NCF quotient is greater than the third threshold, processor 110 can determine that drug combination (A,B) does not fully meet the three conditions. Therefore, processor 110 may not output drug combination (A,B). Since the NCF quotient of the drug combination (C,D) is greater than the third threshold, the processor 110 can determine that the drug combination (C,D) does not fully meet the three conditions. Therefore, the processor 110 may not output the drug combination (C,D).
[0105] Table 14
[0106]
[0107] Figure 6 A flowchart of a method for examining drug interactions is shown according to an embodiment of the present invention, wherein the method may be performed by, for example Figure 1The illustrated electronic device 100 is implemented. In step S601, multiple medical records are obtained, wherein at least one of the multiple medical records indicates whether a patient taking the first drug combination has experienced a hospitalization event. In step S602, a set of drug combinations is generated based on the multiple medical records, wherein the set of drug combinations includes a first drug combination, a second drug combination, and a third drug combination, wherein both the first and second drug combinations contain a first drug, and both the first and third drug combinations contain a second drug. In step S603, a first ratio between the first drug combination and the hospitalization event, a second ratio between the second drug combination and the hospitalization event, and a third ratio between the third drug combination and the hospitalization event are generated based on the multiple medical records. In step S604, a first score corresponding to the first drug is generated based on the second ratio, wherein the first score is negatively correlated with the second ratio. In step S605, a second score corresponding to the second drug is generated based on the third ratio, wherein the second score is negatively correlated with the third ratio, and wherein the first score is greater than or equal to the second score. In step S606, in response to the first ratio being greater than a first threshold, the sum of the first score and the second score being greater than a second threshold, and the quotient of the first score and the second score being less than a third threshold, the first medication combination is output.
[0108] In summary, the electronic device of this invention can analyze multiple medical records using a hidden Dirichlet distribution model, selecting the optimal number of medication combination groups based on factors such as the similarity of medication patterns, patient-preferred medication patterns, or grouping efficacy indicators. After grouping each medication combination to a specific medication pattern according to the optimal number, the electronic device can screen out the most representative medication combinations based on factors such as the importance, uniqueness, stability, and sample size of the medication combinations. If two safe drugs in a specific medication combination are prone to adverse interactions, information about that specific medication combination is output for the user's reference.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for examining drug interactions, characterized in that, include: Obtain multiple medical records, wherein at least one of the multiple medical records indicates whether a patient taking the first combination of medications has experienced a hospitalization event; A set of medication combinations is generated based on the multiple medical records, wherein the set of medication combinations includes a first medication combination, a second medication combination, and a third medication combination, wherein both the first medication combination and the second medication combination include a first drug, and both the first medication combination and the third medication combination include a second drug; A first ratio between the first medication combination and the hospitalization event, a second ratio between the second medication combination and the hospitalization event, and a third ratio between the third medication combination and the hospitalization event are generated based on the multiple medical records. A first score corresponding to the first drug is generated based on the second ratio, wherein the first score is negatively correlated with the second ratio; A second score corresponding to the second drug is generated based on the third ratio, wherein the second score is negatively correlated with the third ratio, and wherein the first score is greater than or equal to the second score; as well as In response to the first ratio being greater than a first threshold, the sum of the first score and the second score being greater than a second threshold, and the quotient of the first score and the second score being less than a third threshold, the first medication combination is output.
2. The method of claim 1, wherein the step of generating the first fraction corresponding to the first drug based on the second ratio comprises: The second medication combination is flagged in response to the second ratio being greater than a risk threshold; A third score is generated based on the marked second medication combination, wherein the third score is equal to the number of marked medication combinations in the set that include the first drug but do not contain the second drug, divided by the number of medication combinations in the set that include the first drug but do not contain the second drug; and The first score is calculated based on the third score, wherein the sum of the first score and the third score is equal to one.
3. The method according to claim 1, wherein the step of generating the medication combination set based on the plurality of medical records includes: A screening process was performed to generate a first unique set of drug combinations, including: Based on the multiple medical records and the implicit Dirichlet distribution model, K topic vectors are generated, including a first topic vector, where K is the number of first topics. The K topic vectors correspond to K topics, and the K topics include the first topic corresponding to the first topic vector. The first topic vector includes the probability distribution of all drug combinations. Starting with the drug combination with the highest probability, multiple important drug combinations are selected from the first topic vector to generate a first important drug combination set; and The first unique drug combination set is determined based on the first important drug combination set; and The drug combination set is generated based on the first unique drug combination set.
4. The method of claim 3, wherein the K topics include a second topic, wherein the step of determining the first unique drug combination set based on the first important drug combination set includes: In response to the first important drug combination being included in the first important drug combination set corresponding to the first topic and the second important drug combination set corresponding to the second topic, the first important drug combination is removed from the first important drug combination set to generate the first unique drug combination set.
5. The method of claim 3, wherein the step of generating the drug combination set based on the first unique drug combination set comprises: The screening process is repeated multiple times to generate multiple unique drug combination sets, including the first unique drug combination set; In response to the fact that the number of the first drug combination in the plurality of unique drug combination sets is greater than a quantity threshold, a first stable drug combination set corresponding to the first topic is generated based on the first drug combination; and The medication combination set is generated based on the first stable medication combination set.
6. The method of claim 5, wherein the step of generating the medication combination set based on the first stable medication combination set comprises: Based on the multiple medical records and the implicit Dirichlet distribution model, multiple medical record vectors are generated corresponding to the multiple medical records, wherein each of the multiple medical record vectors includes the probability distribution of the K topics; Based on the probability distribution of the K topics, determine the set of medical records corresponding to the first topic among the multiple medical records; Calculate the ratio of at least one medical record in the medical record set to the total medical record set, wherein the at least one medical record indicates at least one medication combination in the first stable medication combination set; and In response to the ratio being greater than a ratio threshold, the medication combination set is generated based on the first stable medication combination set, wherein the medication combination set includes multiple medication combinations from the first stable medication combination set.
7. The method according to claim 6, wherein the first medical record in the medical record set corresponds to the first probability distribution of the K topics, wherein the step of determining the set of medical records corresponding to the first topic among the plurality of medical records according to the probability distribution of the K topics includes: In response to the maximum probability in the first probability distribution corresponding to the first topic, it is determined that the first medical record corresponds to the first topic.
8. The method according to claim 3, further comprising: Based on the multiple medical records and the implicit Dirichlet distribution model, a first indicator corresponding to the number of the first topic and a second indicator corresponding to the number of the second topic are generated. as well as The first indicator is compared with the second indicator to select the first topic quantity as K from the first topic quantity and the second topic quantity.
9. The method of claim 8, wherein the step of generating the first index corresponding to the first number of topics comprises: The K topic vectors are generated based on the multiple medical records, the implicit Dirichlet distribution model, and the first number of topics. as well as The average similarity of all 2-combinations of the K topic vectors is calculated as the first metric.
10. The method of claim 8, wherein the step of generating the first index corresponding to the first number of topics comprises: Based on the multiple medical records, the implicit Dirichlet distribution model, and the first number of topics, multiple medical record vectors are generated, each corresponding to the multiple medical records, wherein each of the multiple medical record vectors includes the probability distribution of the K topics; Based on the probability distribution of the K topics, determine at least one medical record among the plurality of medical records that corresponds to the first topic; as well as The ratio is calculated based on the number of at least one medical record to the total number of the plurality of medical records as the first indicator.
11. The method of claim 10, wherein the step of determining the at least one medical record corresponding to the first topic among the plurality of medical records based on the probability distribution of the K topics comprises: A first probability distribution of the K topics corresponding to the at least one medical record is obtained from the plurality of medical record vectors; as well as In response to the maximum probability in the first probability distribution corresponding to the first topic and being greater than a probability threshold, it is determined that the at least one medical record corresponds to the first topic.
12. The method of claim 8, wherein the step of generating the first index corresponding to the first number of topics comprises: Based on the multiple medical records, the implicit Dirichlet distribution model, and the first number of topics, multiple medical record vectors are generated, each corresponding to the multiple medical records, wherein each of the multiple medical record vectors includes the probability distribution of the K topics; The multiple medical records are divided into K groups according to the probability distribution of the K topics, wherein the K groups correspond to the K topics respectively; Calculate the first statistical value of the inter-group distance based on the K groups; Calculate a second statistical value of the distance within each of the K groups; as well as The ratio of the first statistical value to the second statistical value is calculated as the first indicator.
13. The method of claim 12, wherein the step of calculating the first statistical value of the inter-group distance based on the K groups comprises: Calculate multiple distances between the K topic vectors; as well as The first statistical value is obtained by summing the multiple distances.
14. The method of claim 12, wherein the K groups include a first group and a second group, wherein the step of calculating the second statistical value of the distance within the groups based on the K groups includes: Calculate multiple distances between multiple elements in the first group to generate a sum of intra-group distances corresponding to the first group; as well as The second statistical value is obtained by adding the sum of the distances within the first group corresponding to the first group to the sum of the distances within the second group corresponding to the second group.
15. An electronic device for detecting drug interactions, characterized in that, include: transceiver; as well as A processor, coupled to the transceiver and configured to execute: Multiple medical records are obtained through the transceiver, wherein at least one of the multiple medical records indicates whether a patient taking the first combination of medications has experienced a hospitalization event; A set of medication combinations is generated based on the multiple medical records, wherein the set of medication combinations includes a first medication combination, a second medication combination, and a third medication combination, wherein both the first medication combination and the second medication combination include a first drug, and both the first medication combination and the third medication combination include a second drug; A first ratio between the first medication combination and the hospitalization event, a second ratio between the second medication combination and the hospitalization event, and a third ratio between the third medication combination and the hospitalization event are generated based on the multiple medical records. A first score corresponding to the first drug is generated based on the second ratio, wherein the first score is negatively correlated with the second ratio; A second score corresponding to the second drug is generated based on the third ratio, wherein the second score is negatively correlated with the third ratio, and wherein the first score is greater than or equal to the second score; as well as In response to the first ratio being greater than a first threshold, the sum of the first score and the second score being greater than a second threshold, and the quotient of the first score and the second score being less than a third threshold, the first medication combination is output through the transceiver.
Citation Information
Patent Citations
Diabetes-related biomarkers and methods of use thereof
CN102317786A
System and methods for the production of personalized drug products
CN103250176A