Drug Adverse Reaction Mining Method, System, Terminal and Medium Based on Label Propagation

Through the method of mining adverse drug reactions based on label transmission, the term characteristics in the drug relationship diagram and the online page are used, combined with the support vector machine and the label transmission algorithm, the problem of dependence on manual label data in the existing technology is solved, and efficient and precise mining of adverse drug reactions is achieved.

CN115083623BActive Publication Date: 2025-05-27KAIFENG CENT HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210713228.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-05-27
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

The prior art relies on a large amount of manual labeling data in mining adverse drug reactions, and requires each drug to have complete adverse reaction data, resulting in greater limitations in the method.

Method used

The drug adverse reaction mining method based on label transmission is adopted. By obtaining the drug category set, the drug relationship diagram is constructed, the anchor text in the network page is extracted as a term set, and a feature vector containing implicit themes, frequency and distance characteristics is established. The support vector machine is used to identify the adverse reaction candidate set, and the adverse reactions are mined from similar drugs in combination with the label transmission algorithm.

Benefits of technology

Without manual labeling of data, efficient mining of adverse reactions to drugs is reduced, and the accuracy and efficiency of obtaining adverse reactions of drugs is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115083623B_ABST
    Figure CN115083623B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, terminal and medium for mining adverse drug reactions based on label propagation, including: obtaining a drug set M = {m1,..., m i ,..., m n}, and based on a regular expression, obtaining a set C i of the categories to which the drug m i belongs from the attribute list corresponding to the drug m i ; constructing a drug relationship graph in a weighted form of drugs based on the drug set C i ; for the web page corresponding to the drug m i , extracting the anchor text in the page as a term set T i based on a regular expression; establishing a feature vector including implicit theme features, frequency features, and distance features based on the term ti in the term set T i ; obtaining a candidate set cADR i of "adverse drug reactions" of the drug m i from the term set T i based on the "adverse drug reaction" mining algorithm based on label propagation; mining to obtain an "adverse drug reaction" set ADR i from the candidate set cADR i of "adverse drug reactions" of the drug m i based on the drug relationship graph and label propagation. The present invention can mine adverse drug reactions without the need for manual data annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and natural language processing, and relates to a method, system, terminal and medium for mining adverse drug reactions based on label propagation. Background Art

[0002] Adverse drug reactions refer to harmful reactions that occur under normal dosage and usage and are unrelated to the purpose of medication. Adverse drug reactions are usually obtained through clinical trials. However, due to limitations in aspects such as the subjects under test and data collection methods, it is difficult to comprehensively obtain adverse reaction data. In this regard, a series of methods for mining adverse drug reactions have been proposed at home and abroad. Early methods mainly utilized models such as support vector machines and conditional random fields. In recent years, they have mainly been based on deep learning models such as bidirectional long short-term memory networks and BERT. However, most of these methods rely on a large amount of manually labeled data and require each drug to have complete adverse reaction data. This greatly limits the limitations of such methods. How to mine the adverse drug reactions without the need for manually labeled data is a problem that needs to be solved. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems in the prior art, and provide a method, system, terminal and medium for mining adverse drug reactions based on label propagation, which can mine the adverse drug reactions without the need for manually labeled data.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] The method for mining adverse drug reactions based on label propagation includes:

[0006] Obtain a drug set M = {m 1 ,..., m i ,..., m n}, and based on a regular expression, obtain the set C i of the categories to which the drug m i belongs from the attribute list corresponding to the drug m i ;

[0007] Based on the drug set C i , construct a drug relationship graph in a weighted form of drugs;

[0008] For the web page corresponding to the drug m i , extract the anchor text in the page as a term set T i based on a regular expression;

[0009] Based on the term ti in the term set T i , establish a feature vector including implicit theme features, frequency features, and distance features;

[0010] Based on the feature vectors of implicit topic features, frequency features, and distance features, construct an "adverse reaction" recognition algorithm based on support vector machines, starting from the term set T i to obtain the "adverse reaction" candidate set cADR of drug m i ; i

[0011] The "adverse reaction" mining algorithm based on the drug relationship graph and label propagation mines the "adverse reaction" set ADR from the "adverse reaction" candidate set cADR of drug m i ; i i .

[0012] A further improvement of the present invention lies in:

[0013] Based on regular expressions, obtain the set C of categories to which drug m i belongs from the attribute list corresponding to drug m i , specifically: i

[0014] Based on regular expressions RE1 and RE2, extract the set C of categories to which m i belongs from the Categories list at the bottom of the Wikipedia page corresponding to m i ; i

[0015] The regular expressions RE1 and RE2 are as follows:

[0016] RE1: normal-catlinks″>(.+)hidden-catlinks

[0017] RE2: <a href=″([^″]+)″\s+title=″[^″]+″>([^<]+)

[0018] Among them, RE1 is used to match the Category list in the form of this paragraph on the page, and RE2 is used to match a group of categories to which the drug belongs from the Category list.

[0019] Based on the drug set C i , construct a drug relationship graph in the form of weighted drugs, specifically: Based on the drug set {C i |i∈[1..n]}, construct a weighted graph form of the drug relationship graph G=(M, E, ψ), where indicates that drug m i and m j belong to at least one same category, is a function that assigns weights to each edge in E. ​​​​

[0020] Based on {C i |m i ∈M}, the steps to construct the drug relationship graph G=(M, E, ψ) are as follows:

[0021] Let ψ(m i , m j ) = |C i ∩C j |, where m i , m j ∈M and i≠j;

[0022] Let E be initially For and i≠j, if ψ(m i , m j )≠0, then E←E∪{e ij}, and set the weight of the edge e ij to ψ(m i , m j ).

[0023] For the web page corresponding to the drug m i , extract the anchor text in the page as the term set T i based on the regular expression; specifically:

[0024] For the drug m i , from the Wikipedia page corresponding to m i , use the regular expression RE3 to extract the anchor text in the page as the term set T i ;

[0025] RE3: <a href=″([^″]+)″\s+title=″([^″]+)″

[0026] Among them, RE3 is used to match the anchor text in the page as the term.

[0027] Based on the feature vectors of implicit topic features, frequency features, and distance features, construct a support vector machine-based "adverse reaction" recognition algorithm to obtain the "adverse reaction" candidate set cADR i of the drug m i ; specifically: i ;

[0028] Use the page set of the offline version of Wikipedia Medicine to infer the parameters of the unsupervised latent Dirichlet allocation model; the three hyperparameters are: the number of latent topics k; the "document-topic" distribution parameter α=(0.1,..., 0.1,..., 0.1), with a dimension of 10; the "topic-word" distribution parameter β=(1 / m,..., 1 / m,..., 1 / m), with a dimension of Represents the page set P M The number of all terms in ;

[0029] for Using the LDA model, we get the term t il Latent topic features of LDA il ; Statistics P M Contains the term t il Number of pages n il , the term t il Frequency characteristics for Statistical term t il With T i The average distance of other terms in is used as the distance feature and is defined as:

[0030]

[0031] Among them, L P Indicates the byte size of page p, dis(t il ,t,p) represents the term t il The closest character interval to t in page p. If the two characters do not appear at the same time in page p, then dis(t il ,t,p)=L P ;

[0032] Concatenating the above features, we get the term t il The eigenvector of For splicing operation;

[0033] Manually annotate several training samples containing the above feature vectors, train a Boolean classifier based on support vector machine, and distinguish the categories as "adverse reaction" and "non-adverse reaction" respectively; use the Boolean classifier to classify each T i Generate the corresponding "adverse reaction" candidate set cADR i .

[0034] "Adverse reaction" mining algorithm based on drug relationship graph and label propagation from drug m i The candidate set of "adverse reactions" cADR i Mining to get the "adverse reaction" set ADR i , specifically:

[0035] (1) Generate the corresponding adjacency matrix A based on the drug relationship graph G = (M, E, ψ) G =(a ij ) n×n ,in,

[0036]

[0037] Among them, a ij represents the similarity between drugs m i , where m j ∈M;

[0038] (2) According to the "adverse reaction" candidate set cADR i corresponding to T i (i ∈ [1..n]), generate a matrix describing the "drug - adverse reaction" bimodal data:

[0039] R 0 =(r ih ) n×q (3)

[0040] Among them,

[0041] q = |cADR| (4)

[0042]

[0043]

[0044] (3) Let t = 0, and iteratively execute the following process until ||R t+1 -R t || F < 0.01, output R = R t+1 , where ||·|| F is the F - norm:

[0045] 3.1, R t+1 = A G R t ; Multiply the adjacency matrix by the bimodal matrix to achieve one - time propagation and obtain a new bimodal matrix;

[0046] 3.2, Divide all elements in each row of the matrix R t+1 by the maximum value of that row to normalize R t+1 ;

[0047] 3.3, t ← t + 1;

[0048] (4) For each drug m i ∈M, according to the "adverse reaction" set ADR ih generated by R=(r n×q ), execute the following process: i

[0049] 4.1, Initialize ADR i as

[0050] 4.2, For If r ih>0.5, then ADR i ←ADR i ∪{t′ h};

[0051] Finally, for each drug m i The corresponding ADR i is the term set of the adverse drug reactions of this drug.

[0052] The adverse drug reaction mining system based on label propagation includes:

[0053] The first acquisition module, which is used to acquire the drug set M = {m 1 ,..., m i ,..., m n}, and based on the regular expression, obtain the set C i of the categories to which the drug m i belongs from the attribute list corresponding to the drug m i ;

[0054] The first construction module, which constructs a drug relationship graph in the weighted form of drugs based on the drug set C i ;

[0055] The extraction module, which is used to extract the anchor text in the web page corresponding to the drug m i as the term set T i based on the regular expression;

[0056] The second construction module, which establishes a feature vector including implicit topic features, frequency features, and distance features based on the terms ti in the term set T i ;

[0057] The second acquisition module, which constructs an "adverse reaction" recognition algorithm based on the support vector machine based on the feature vector of the implicit topic feature, frequency feature, and distance feature, and obtains the "adverse reaction" candidate set cADR i of the drug m i from the term set T i ;

[0058] The mining module, which mines the "adverse reaction" set ADR i from the "adverse reaction" candidate set cADR i of the drug m i based on the drug relationship graph and the "adverse reaction" mining algorithm of label propagation.

[0059] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0060] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0061] Compared with the prior art, the present invention has the following beneficial effects:

[0062] In the present invention, for each drug, a set of categories to which the drug belongs is extracted from the corresponding attribute list by using a regular expression; a drug relationship graph in the form of a weighted graph is constructed by using the category set. Anchor texts are extracted from the network page corresponding to each drug by using a regular expression as a set of terms related to the drug. For each term of each drug, a feature vector including implicit topic features, frequency features, and distance features is established; according to this vector, an "adverse reaction" recognition algorithm based on a support vector machine is used to obtain a candidate set of "adverse reactions" of the drug from the set of terms. An "adverse reaction" mining algorithm based on label propagation uses the drug relationship graph to generate a corresponding adjacency matrix and the candidate set of "adverse reactions" to generate a "drug - adverse reaction" bimodal matrix, and propagates the "adverse reactions" of similar drugs to each other; finally, a set of "adverse reactions" of each drug is obtained. The present invention can mine the adverse reactions of drugs without the need for manual annotation of data. Through label propagation of similar drugs in terms of "adverse reactions", the present invention can efficiently complete the task of mining drug adverse reactions only by using Wikipedia pages, reducing the dependence on manual annotation while ensuring the accuracy rate.

[0063] Further, the similarity of drugs is obtained by using the Categories list at the bottom of the Wikipedia page, providing a basis for label propagation.

[0064] Further, by using three features of implicit topic features, frequency features, and distance features and a classifier based on a support vector machine, a candidate set of "adverse reactions" is generated, greatly saving the candidate space. Description of the Drawings

[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0066] Figure 1Flowchart of the method for mining adverse drug reactions based on label propagation according to the present invention;

[0067] Figure 2 Structural diagram of the system for mining adverse drug reactions based on label propagation according to the present invention. Detailed implementation manners

[0068] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0069] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0070] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0071] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship when the product of the present invention is normally placed. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0072] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0073] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and defined, if the terms "set", "install", "connect", and "link" appear, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0074] The following further describes the present invention in detail with reference to the accompanying drawings:

[0075] See Figure 1 , the present invention discloses a method for mining adverse drug reactions based on label propagation, including:

[0076] STEP1: Construct a drug relationship graph

[0077] Obtain a drug set M = {m 1 ,..., m i ,..., m n}, and based on regular expressions, obtain the set C i of the categories to which the drug m i belongs from the attribute list corresponding to the drug m i ;

[0078] Let the drug set be M = {m 1 ,..., m i ,..., m n}, and the process of constructing the drug relationship graph G = (M, E, ψ) is as follows:

[0079] 1.1: Obtain the Categories list corresponding to the drug set M = {m 1 ,..., m i ,..., m n}. Based on the regular expression RE1, extract the Categories list from the bottom of the Wikipedia page corresponding to each drug m i ∈ M. For example, for the drug "entecavir", in the page "https: / / zh.wikipedia.org / wiki / entecavir", the following can be extracted: "Category: Bristol-Myers Squibb | Nucleoside analog reverse transcriptase inhibitors | Purines | Cyclopentyl compounds | World Health Organization essential medicines | Olefin derivatives".

[0080] RE1: normal-catlinks″>(.+)hidden-catlinks

[0081] 1.2: Obtain the category corresponding to the drug. For each drug m i ∈M, use the regular expression RE2 to extract a set C i consisting of a group of phrases from the corresponding Categories list of m i as the category to which m i belongs; for example, for the drug "Entecavir", the set of its belonging categories is {"Bristol-Myers Squibb", "Nucleoside analog reverse transcriptase inhibitor", "Purine", "Cyclopentyl compound", "WHO Model List of Essential Medicines", "Olefin derivative"}.

[0082] RE2: <a href="([^"]+)"\s+title="[^"]+">([^<]+)

[0083] Based on the drug set C i , construct a drug relationship graph in the weighted form of drugs.

[0084] 1.3: Use {C i |m i ∈M} to generate the drug relationship graph G=(M, E, ψ). G is a weighted graph indicating that drug m i and m j belong to at least one same category and ψ is a function that assigns weights to each edge in E, ψ(m i , m j ) = |C i ∩C j |, where m i , m j ∈M and i≠j. The generation process is as follows: Initialize E as For and i≠j, if ψ(m i , m j )≠0, then E←E∪{e ij}, and set the weight of edge e ij to ψ(m i , m j ). For example, the two drugs "Entecavir" and "Lamivudine" have two common categories "Nucleoside analog reverse transcriptase inhibitor" and "WHO Model List of Essential Medicines"; therefore, add an edge between the two drugs in the drug relationship graph G, and the weight is 2.

[0085] STEP2: Generate a term set related to drugs

[0086] For the web page corresponding to drug m i , extract the anchor text in the page as the term set T i based on the regular expression;

[0087] For each drug m i , from the Wikipedia page corresponding to m i , use the regular expression RE3 to match the anchor text in the page as the term set T i .

[0088] RE3: <a href="([^"]+)"\s+title="([^"]+)"

[0089] For example, terms such as "antiviral drug", "hepatitis B virus", "headache", and "nausea" are matched from the Wikipedia page of "entecavir".

[0090] STEP3: Generate the candidate set of "adverse reactions"

[0091] Based on the terms ti in the term set T i , establish a feature vector that includes implicit topic features, frequency features, and distance features.

[0092] Implicit topic features: Use the page set of the offline version of Wikipedia medicine to infer the parameters of the unsupervised latent Dirichlet distribution model; the three hyperparameters are: the number of implicit topics k = 10; the "document-topic" distribution parameter α = (0.1,..., 0.1,..., 0.1), with a dimension of 10; the "topic-word" distribution parameter β = (1 / m,..., 1 / m,..., 1 / m), with a dimension of denotes the number of all terms in the page set P M ; use the trained LDA model to obtain the implicit topic features of the term t il ∈T i . Each dimension represents the conditional probability of the term t on the corresponding topic. il

[0093] Frequency features: For count the number of pages n M in the Wikipedia page set P corresponding to the drug set M that contain the term t il , and the frequency feature of the term t il is defined as follows: il

[0094]

[0095] Distance features: Since terms related to "adverse reactions" are usually close to each other in the page, therefore, for count the distance between the term t il and T i ​​The average distance of other terms in it is used as a distance feature, which is defined as follows:

[0096]

[0097] In formula (2), L P represents the byte size of page p, and dis(t il , t, p) represents the closest character interval between term t il and t in page p. If they do not appear simultaneously in page p, it is stipulated that dis(t il , t, p) = L P .

[0098] Concatenate the above three features to obtain the feature vector il of term t The operation is concatenation.

[0099] Manually annotate 50 training samples containing the above feature vectors, and train a Boolean classifier based on a support vector machine. The discriminant categories are "adverse reaction" and "non - adverse reaction" respectively. Using this classifier, for each term t i in T il , determine whether it is an "adverse reaction", and generate the "adverse reaction" candidate set cADR i corresponding to each T i . Compared with T i obtained in STEP2, the scale of cADR i is reduced, which can improve the accuracy of "adverse reaction" mining and reduce the time and space overhead.

[0100] STEP4: Generate the "adverse reaction" set

[0101] Based on the feature vectors of implicit topic features, frequency features, and distance features, construct an "adverse reaction" recognition algorithm based on a support vector machine, and obtain the "adverse reaction" candidate set cADR i of drug m i . i

[0102] Mine the "adverse reaction" set ADR i from the "adverse reaction" candidate set cADR i of drug m i using the "adverse reaction" mining algorithm based on the drug relationship graph and label propagation.

[0103] Utilize the phenomenon that "if the drug categories are close, the adverse reactions also tend to be close", and propose an "adverse reaction" mining algorithm based on label propagation to spread the "adverse reactions" of similar drugs in M to each other, and obtain the drug m​i "Adverse Reaction" set ADR i , The process is as follows:

[0104] STEP4.1. Generate the adjacency matrix. Generate the corresponding adjacency matrix A according to the drug relationship graph G=(M, E, ψ) G =(a ij ) n×n , where a ij represents the similarity between drugs m i , m j ∈M, which is used to control the intensity of the mutual transmission of "adverse reactions" between the two, and is defined as follows:

[0105]

[0106] STEP4.2. Generate the "drug - adverse reaction" bimodal matrix. Generate the matrix i R i describing the "drug - adverse reaction" bimodal data according to the "adverse reaction" candidate set cADR

[0107] corresponding to T 0 =(r ih ) n×q (4)

[0108] where

[0109] q = |cADR| (5)

[0110]

[0111] r ih represents the probability that drug m i has the "adverse reaction" t' h , and is defined as follows:

[0112]

[0113] STEP4.3. Label propagation. The algorithm is as follows:

[0114] Input: adjacency matrix A G and "drug - adverse reaction" bimodal matrix R 0 ;

[0115] Output: "drug - adverse reaction" bimodal matrix R after label propagation;

[0116] STEP4.3.1. t = 0;

[0117] STEP4.3.2. R t+1 = A G Rt , perform one - time propagation by multiplying the adjacency matrix with the bimodal matrix to obtain a new bimodal matrix;

[0118] STEP4.3.3. Divide all elements in each row of matrix R t+1 by the maximum value of that row to normalize R t+1 ;

[0119] STEP4.3.4. If ||R t+1 - R t || F < 0.01, output R = R t+1 , and the algorithm ends;

[0120] STEP4.3.5. Otherwise, t ← t + 1, and return to STEP4.3.2) to continue execution;

[0121] where ||·|| F is the F - norm, and ||A|| F is the sum of the squares of all elements in matrix A.

[0122] STEP4.4 Determine the adverse drug reactions of each drug. For each drug m i ∈M, according to the "adverse drug reaction" set ADR ih generated by R = (r n×q ), perform the following process: i

[0123] S106.4.1. Initialize ADR i as

[0124] S106.4.2. For If r ih > 0.5, indicating that the probability that drug m i has the "adverse drug reaction" t′ h exceeds 0.5, then ADR i ← ADR i ∪{t′ h};

[0125] Finally, the ADR i corresponding to each drug m i is the term set of the adverse drug reactions of this drug.

[0126] See Figure 2 , the present invention discloses a system for mining adverse drug reactions based on label propagation, including:

[0127] The first acquisition module, which is used to acquire the drug set M = {m 1 ,..., mi ,..., m n}, and based on a regular expression, obtain the set C of the categories to which the drug m belongs from the corresponding attribute list of the drug m i ; i ; i ;

[0128] The first construction module, which constructs a drug relationship graph in the weighted form of drugs based on the drug set C i ;

[0129] The extraction module, which is used to extract the anchor text in the web page corresponding to the drug m as the term set T based on a regular expression i ; i ;

[0130] The second construction module, which establishes a feature vector including implicit theme features, frequency features, and distance features based on the term ti in the term set T i ;

[0131] The second acquisition module, which constructs an "adverse reaction" recognition algorithm based on a support vector machine based on the feature vector of implicit theme features, frequency features, and distance features, and obtains the "adverse reaction" candidate set cADR of the drug m from the term set T i ; i ; i ;

[0132] The mining module, which mines the "adverse reaction" set ADR from the "adverse reaction" candidate set cADR of the drug m based on the drug relationship graph and the "adverse reaction" mining algorithm of label propagation i ; i ; i .

[0133] The terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned various method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned various device embodiments are implemented.

[0134] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.

[0135] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0136] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0137] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and by invoking the data stored in the memory, the processor realizes various functions of the terminal device.

[0138] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-described various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, Read-Only Memory (ROM), Random Access Memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0139] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for mining adverse drug reactions based on label propagation, characterized in that, it includes: Obtain the drug set M = {m 1 ,..., m i ,..., m n}. Based on the regular expression, obtain the set C i of the categories to which the drug m i belongs from the attribute list corresponding to the drug m i ; Based on the regular expression, from drug m i Obtain the set C of the categories to which drug m i belongs from the corresponding attribute list i , specifically: Based on the regular expressions RE1 and RE2, extract the set Ci of the categories to which m i belongs from the Categories list at the bottom of the Wikipedia page corresponding to m i ; The regular expressions RE1 and RE2 are as follows: RE1: normal-catlinks″>(.+)hidden-catlinks RE2: <a href=″([^″]+)″\s+title=″[^″]+″>([^<]+) where RE1 is used to match the Category list in the form of this text paragraph on the page, and RE2 is used to match a group of categories to which the drug belongs from the Category list; Based on the drug set C i , construct a drug relationship graph in the weighted form of drugs For drug m i For the corresponding web page, extract the anchor text in the page based on regular expressions as the term set T i ; Based on the term set T i for the term t i , a feature vector including implicit topic features, frequency features, and distance features is established; Construct an "adverse reaction" recognition algorithm based on support vector machines using feature vectors based on implicit topic features, frequency features, and distance features, and obtain drug m from the term set Ti i 's "adverse reaction" candidate set cADR i ; "Adverse Reaction" Mining Algorithm Based on Drug Relationship Graph and Label Propagation, Mining the "Adverse Reaction" Set ADRi from the "Adverse Reaction" Candidate Set cADR of Drug mi i Mining the "Adverse Reaction" set ADRi.

2. The method for mining adverse drug reactions based on label propagation according to claim 1, characterized in that, Based on the drug set C i , construct a drug relationship graph in the form of weighted drugs, specifically: based on the drug set {C i | i ∈ [1..n]}, construct a drug relationship graph G = (M, E, ψ) in the form of a weighted graph, where represents drug m i and m j belong to at least one same category, is a function that assigns weights to each edge in E.

3. The method for mining adverse drug reactions based on label propagation according to claim 2, characterized in that, The steps of constructing the drug relationship graph G=(M, E, ψ) based on {C i |m i ∈M} are as follows: Let ψ(m i , m j ) = |C i ∩ C j |, where m i , m j ∈ M and i ≠ j; Let E be initially θ; for , and i ≠ j, if ψ(m i , m j ) ≠ 0, then E ← E ∪ {e ij}, and set the weight of edge e ij to ψ(m i , m j ).

4. The method for mining adverse drug reactions based on label propagation according to claim 3, characterized in that, The said drug m i For the corresponding web page, extract the anchor text in the page based on regular expressions as the term set T i Specifically: For drug m i , from the Wikipedia page corresponding to m i , use the regular expression RE3 to extract the anchor text in the page as the term set T i ; RE3: <a href=″([^″]+)″\s+title=″([^″]+)″ where RE3 is used to match the anchor text on the page as a term.

5. The method for mining adverse drug reactions based on label propagation according to claim 4, characterized in that, Based on the feature vector of implicit topic features, frequency features, and distance features, construct an "adverse reaction" recognition algorithm based on support vector machines, starting from the term set T i to obtain the "adverse reaction" candidate set cADR of drug m i ; specifically: i ​ Using the page set of the offline version of Wikipedia Medicine for parameter inference of the unsupervised Latent Dirichlet Allocation model; the three hyperparameters are: the number of latent topics k; the "document-topic" distribution parameter α = (0.1,..., 0.1,..., 0.1), with a dimension of 10; the "topic-word" distribution parameter β = (1 / m,..., 1 / m,..., 1 / m), where m is the number of all terms in the page set P M represented by M For : Obtain the latent topic feature LDA of term t il ; Count the number of pages n containing term t il in P M , the frequency feature of term t il ; For il , count the average distance between term t il and other terms in T as the distance feature, defined as: , count the average distance between term t il and other terms in T i as the distance feature, defined as: (1) Among them, L P represents the byte size of page p, and dis(t il , t, p) represents the closest character interval between term t il and t in page p. If they do not appear simultaneously in page p, it is stipulated that dis(t il , t, p) = L P ; Concatenate the above features to obtain the term t il The feature vector of is the concatenation operation; Manually annotate a number of training samples containing the above feature vectors, and train a Boolean classifier based on a support vector machine, with the discriminant categories being "adverse reaction" and "non-adverse reaction"; use the Boolean classifier to process each T i to generate a corresponding "adverse reaction" candidate set cADR i .

6. The method for mining adverse drug reactions based on label propagation according to claim 5, characterized in that, The "Adverse Reaction" mining algorithm based on the drug relationship graph and label propagation mines the "Adverse Reaction" set ADR i from the "Adverse Reaction" candidate set cADR i of drug m, specifically as follows: i , specifically: (1) Generate the corresponding adjacency matrix A according to the drug relationship graph G=(M, E, ψ). G =(a ij ) n×n , where (2) Among them, a ij represents the similarity between drugs m i , m j ∈ M; (2) According to T i The candidate set cADR of "adverse reactions" corresponding to (i ∈ [1..n]) i , generate a matrix describing the "drug - adverse reaction" bimodal data: R 0 =(r ih ) n×q (3) where, q=|cADR|(4) (5) (6) (3) Let \(t = 0\), and iteratively execute the following process until \(\left\|\mathbf{R}\right.\) t +1 -\(\mathbf{R}\) t \(\left\|\right.\) F \(< 0.01\), output \(\mathbf{R}=\mathbf{R}\) t +1 , \(\left\|\cdot\right\|\) F is the \(F\) - norm: 3.1, R t +1 = A G R t ; By multiplying the adjacency matrix with the bimodal matrix, one propagation is achieved to obtain a new bimodal matrix; 3.2, Divide all elements in each row of matrix R t+1 by the maximum value of that row to achieve the normalization of R t +1 ; 3.3, t←t + 1; (4) For each drug m i ∈ M, according to R = (r ih ) n×q Generate the "adverse reaction" set ADR i , perform the following process: 4.1, Initialize ADR i as θ, 4.2, for , if r ih > 0.5, then ADR i ← ADR i ∪ {t′ h}; Finally, for each drug m i the corresponding ADR i is the terminology set of the adverse drug reactions of this drug.

7. A system for mining adverse drug reactions based on label propagation, characterized in that, it includes: The first acquisition module, which is used to acquire a medicine set M = {m 1 ,..., m i ,..., m n}, and based on a regular expression, acquire a set C i of the categories to which the medicine m i belongs from the attribute list corresponding to the medicine m i ; The first construction module, which constructs a drug relationship graph in the weighted form of drugs based on the drug set C i , and constructs a drug relationship graph in the weighted form of drugs An extraction module, which is used to process drug m i The corresponding web page, and extract the anchor text in the page as the term set T based on the regular expression i ; The second construction module, which is based on the terms in the term set T i to establish a feature vector including implicit topic features, frequency features, and distance features i for the term t The second acquisition module, which constructs an "adverse reaction" recognition algorithm based on a support vector machine from the feature vectors of the latent topic feature, frequency feature, and distance feature, and obtains the term set T i to obtain the drug m i of the "adverse reaction" candidate set cADR i ; Mining module, which mines the "adverse reaction" set ADR i from the "adverse reaction" candidate set cADR i of drug m based on the drug relationship graph and the "adverse reaction" mining algorithm of label propagation i of drug m i by the mining module i .

8. A terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 - 6 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 - 6 are implemented.

Citation Information

Patent Citations

  • Adverse drug reaction mining method and system

    CN106055879A

  • A method for carrying out theme facet mining through a label propagation algorithm

    CN109815495A