A method for vectorizing medical text words
Through the optimization of GLOVE model and co-occurrence matrix, the heterogeneity problem of medical data sets is solved, efficient vectorization of medical texts is realized, and data mining effect is improved.
Patent Information
- Application Number
- CN202111185056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-10-12
AI Technical Summary
The heterogeneity of medical data sets makes it difficult to integrate and mine data, and the existing technology is difficult to effectively process medical texts with structured, semi-structured and unstructured data, and lacks efficient word vectorization methods.
The GLOVE model is adopted to establish a co-occurrence matrix and a loss function, combine global and detailed co-occurrence matrix calculations to perform word vectorization in medical texts, and use gradient descent to optimize word vectors and preserve contextual relationships.
It realizes efficient vectorization of medical texts. As the scale of the corpus increases, the performance increases monotonically, retaining the contextual connection of the text and improving the effect of data mining.
Smart Images

Figure CN114004225B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pre-training for natural language processing, and specifically provides a method for medical text word vectorization. Background Art
[0002] The promotion of medical and health big data has facilitated the research on medical data mining and knowledge discovery, and the amount of heterogeneous data from different data sources is huge. In the field of medical data, it can be clearly felt that the characteristics of medical data sets are data heterogeneity, that is, due to medical detection methods, the proportion of data visualization is relatively high. However, it also includes a part of structured data and most unstructured data. Therefore, medical data sets are typical heterogeneous data sets that coexist with unstructured data and structured data.
[0003] Therefore, it is now necessary to integrate, clean, and mine medical data. Summary of the Invention
[0004] In view of the above-mentioned deficiencies of the prior art, the present invention provides a medical text word vectorization method with strong practicability.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] A medical text word vectorization method prepares for subsequent vectorization by exploring the original medical text data to establish a thesaurus, and then performs medical data word vectorization through the GLOVE model;
[0007] The original medical text is divided into structured data, semi-structured data, and unstructured data. The structured data has fixed filling requirement data. The semi-structured data includes a part of electronic medical record data. There are fixed identifiers in the semi-structured data, and the content in the fixed identifiers may be empty. The unstructured data also includes a part of electronic medical record data. The unstructured data has no identifiers and is extracted according to knowledge.
[0008] Further, in the construction of the GLOVE model, a co-occurrence matrix is statistically calculated.
[0009] a. Let the co-occurrence matrix be X, and its element be Xij. The meaning of Xij is: in the entire corpus, the number of times that word i and word j appear in the same window;
[0010] b. Establish a vocabulary frequency for the entire corpus, and return a dictionary D->(a,f), which maps the string to the word ID and the word corpus frequency;
[0011] c. Establish a co-occurrence list for the given corpus.
[0012] Further, in the construction of the GLOVE model, the loss function of the model is as follows:
[0013]
[0014] Among them, vi and vj are the word vectors of word i and word j, bi and bj are two bias terms, f is a weight function, N is the size of the vocabulary, and the dimension of the co-occurrence matrix is N*N.
[0015] Furthermore, the derivation method of the GLOVE model is the ratio of two conditional probabilities:
[0016]
[0017] The sum of the elements in the i-th row of the matrix is obtained. The following conditional probability represents the probability that k appears in the context of word i
[0018]
[0019] It can be seen from the above formula that the relevance of these three words is related to the probability. The Ratio index is inversely proportional to the relevance of jk and directly proportional to the relevance of ik.
[0020] Furthermore, the co-occurrence matrix and the word vector can be the same, and have the following form:
[0021]
[0022] Among them, wi is the word vector and wk is the independent context word vector.
[0023] The information of F is represented in proportion in the vector space. Using the difference of vectors, the above equation is modified as:
[0024]
[0025] Next, the independent variable in the above formula is a vector, and the right side of the equation is a scalar. Take the dot product of the parameters on the left side:
[0026]
[0027] It can be obtained from the above formula:
[0028]
[0029] Take the logarithm of both sides of the equation to get:
[0030]
[0031] At this time, the rightmost log(Xi) has nothing to do with k, and the symmetry of both sides can be restored by adding a bias. Finally, the equation is obtained as follows:
[0032]
[0033] Perform weighted least squares regression on the above formula, convert the above formula into a least squares method problem, and obtain the final loss function as follows:
[0034]
[0035] Among them, the function F(X) should meet the following requirements:
[0036] a. It should be 0 and continuous at 0;
[0037] b. The function should not be decreasing;
[0038] c. F(X) should be relatively small when x is large;
[0039] Final selection:
[0040]
[0041] Furthermore, in the construction of the GLOVE model, it includes the construction of large text, the construction of detailed text, the construction of the global and detailed co-occurrence matrix, and matrix vectorization;
[0042] The construction of large text segments the large text. It only needs to divide the text into multiple sentences. The word segmentation method is to cut according to punctuation marks such as [$%&'()*+,---。,;:......], that is, T = [S1, S2, S3…Sn],
[0043] Keep the result S in the database for calculating word vectors later.
[0044] Furthermore, perform detailed segmentation on the above [S1, S2, S3…Sn] sentences. In this part, the Chinese stop word list is needed to segment the text as follows:
[0045] S = [X1, X2, X3…Xn]
[0046] Because the GLOVE model calculates word vectors based on global information, some context relationships will be missing. Therefore, it is necessary to calculate the co-occurrence matrix for both S and T texts, and take the co-occurrence matrix calculated for the T text:
[0047] Si = X Si = ∑ k X Sik
[0048] It is obtained that when calculating Xi generated by the S text, Si is added as follows:
[0049] Xi → Si Xi
[0050] In this way, some context connections can be retained.
[0051] Furthermore, in the construction of the global and detailed co-occurrence matrix,
[0052] Each clause has no repeated words: When decomposing the T text, mark each S text with C, so that the subsequent Xi co-occurrence matrix is marked and used for Si to calculate the corresponding sub-items:
[0053] T = [(S1, C1), (S2, C2), (S3, C3)…(Sn, Cn)]
[0054] Si = [(X1, Ci), (X2, Ci), (X3, Ci)…(Xn, Ci)]
[0055] In the case of repeated words: When decomposing the T text, mark each S text with C, so that the subsequent Xi co-occurrence matrix is marked and used for Si to calculate the corresponding sub-items. For sub-items containing repeated items, store them as CiCj and multiply by multiple Si
[0056] T = [(S1, C1), (S2, C2), (S3, C3)…(Sn, Cn)]
[0057] Si = [(X1, CiCj), (X2, Ci), (X3, CiCj)…(Xn, Ci)]
[0058] Furthermore, in the matrix vectorization,
[0059] a. Set input parameters
[0060] The input parameters are the co-occurrence matrix T, the corpus dictionary D, the output dimension, and the number of iterations;
[0061] The co-occurrence matrix T is calculated through the construction in the global and detailed co-occurrence matrix. The fixed window size is 10. The corpus dictionary D is calculated by establishing a word frequency. The return value is D->(a, f), which maps the sentence to the word ID and the corpus frequency of the word. The output dimension and the number of iterations are fixed values;
[0062] b. Construct the vector matrix
[0063] This matrix is 2V*d, where V is the size of the corpus vocabulary and d is the dimension of the word vector. Randomly initialize all elements in the range (-0.5, 0.5), and construct two word vectors for each word: one word as the main word and one word as the context word;
[0064] c. Construct the bias term:
[0065] Construct an array of size 2V through the above vectors, and the value range is (0.5, 0.5);
[0066] d. Gradient descent method:
[0067] Training is performed by the gradient descent method;
[0068] e. Output vector matrix
[0069] If the input co-occurrence matrix is not empty, the provided function will call the known matrix after each iteration, and then pass the parameters to the loss function, and finally return the calculated word vector matrix.
[0070] Compared with the prior art, a medical text word vectorization method of the present invention has the following prominent beneficial effects:
[0071] The present invention uses global information and local information extraction, provides multiple options for text vectorization, and its performance monotonically increases as the corpus size increases. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0073] Attached Figure 1 is a schematic flowchart of a medical text word vectorization method;
[0074] Attached Figure 2 is a flowchart of the original text splitting in a medical text word vectorization method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] In order to enable those skilled in the art to better understand the solution of the present invention, the following further detailed description of the present invention is provided in conjunction with specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0076] The following gives a best embodiment:
[0077] As Figure 1-2 shown, a medical text word vectorization method in this embodiment prepares for subsequent vectorization by exploring and establishing a word library for the original medical text data, and then performs medical data word vectorization through the GLOVE model.
[0078] Medical texts are divided into structured data, semi-structured data, and unstructured data.
[0079] Among them, most of the structured data are the data with fixed filling requirements in the hospital system design, such as outpatient visit data, patient personal information data, and disease diagnosis information data. Some of the semi-structured data contain part of the electronic medical record data, which have fixed identifiers, but the content in the identifiers may be empty. There is also a part of the hospital's electronic case data that is unstructured data, and this part of the data has no identifiers and needs to be extracted according to knowledge.
[0080] In the construction of the GLOVE model, the co-occurrence matrix is statistically calculated:
[0081] a. Let the co-occurrence matrix be X, and its element be Xij;
[0082] The meaning of the element Xij is: in the entire corpus, the number of times that word i and word j appear in a window together.
[0083] b. Establish a vocabulary frequency for the entire corpus. Return the dictionary D -> (a, f), which maps the string to the word ID and the word corpus frequency.
[0084] c. Establish a co-occurrence list for the given corpus.
[0085] The loss function of the model is as follows
[0086]
[0087] Among them, vi and vj are the word vectors of word i and word j, bi and bj are two bias terms, f is the weight function, and N is the size of the vocabulary (the dimension of the co-occurrence matrix is N * N).
[0088] The derivation method of the model is the ratio of two conditional probabilities
[0089]
[0090] What is obtained is the sum of a row of matrix i. The following conditional probability represents the probability that k appears in the context of word i
[0091]
[0092] It can be seen from the above formula that there is a relationship between the relevance of these three words and the probability. The Ratio index is inversely proportional to the relevance of jk and directly proportional to the relevance of ik.
[0093] Assume that these two indicators can be the same, then there is the following form:
[0094]
[0095] Among them, wi is the word vector, and wk is the independent context word vector:
[0096] The information of F is represented proportionally in the vector space. Using the difference of vectors, the above equation is modified to:
[0097]
[0098] Next, the independent variable in the above formula is a vector, and the right side of the equation is a scalar. Take the dot product of the parameters on the left side
[0099]
[0100] It can be obtained from the above formula that:
[0101]
[0102] Taking the logarithm of both sides of the equation gives
[0103]
[0104] At this time, the log(Xi) on the far right has nothing to do with k, and the symmetry of both sides can be restored by adding a bias.
[0105] Finally, the equation is obtained as follows
[0106]
[0107] Perform weighted least squares regression on the above formula, and convert the above formula into a least squares problem to obtain the final loss function as follows:
[0108]
[0109] Among them, the function F(X) should meet the following requirements:
[0110] a. It should be 0 and continuous at 0
[0111] b. The function should not be decreasing
[0112] c. F(X) should be relatively small when x is large
[0113] Finally, select:
[0114]
[0115] First, build a large text. Perform the first step of segmentation on the large text contained, and divide the text into multiple sentences without further refined distinction. The specific method is to cut according to punctuation marks such as [$%&'()*+,---。,;:......]. At this time, the electronic medical record text is as follows:
[0116] T = [S1, S2, S3…Sn]
[0117] Keep the result in the database for subsequent calculation of word vectors.
[0118] Then, perform the construction of detailed text:
[0119] Perform a detailed cut on the above sentences [S1, S2, S3... Sn]. The Chinese stop word list is needed to segment the text as follows:
[0120] S = [X1, X2, X3... Xn]
[0121] Since the GLOVE model calculates word vectors based on global information, some context relationships will be missing. Therefore, it is necessary to calculate the co-occurrence matrices for both S and T texts, and take the co-occurrence matrix calculated for the T text:
[0122] Si = X Si = ∑ k X Sik
[0123] It is obtained that when calculating Xi generated from the S text, Si is added as follows:
[0124] Xi → Si Xi
[0125] In this way, some context connections can be retained.
[0126] Construction of global and detailed collinear matrices:
[0127] Each clause has no repeated words: When decomposing the T text, mark each S text with C, so that the subsequent Xi co-occurrence matrix is marked and used to calculate the corresponding sub-items for Si.
[0128] T = [(S1, C1), (S2, C2), (S3, C3),..., (Sn, Cn)]
[0129] Si = [(X1, Ci), (X2, Ci), (X3, Ci),..., (Xn, Ci)]
[0130] There is a situation of repeated words: When decomposing the T text, mark each S text with C, so that the subsequent Xi co-occurrence matrix is marked and used to calculate the corresponding sub-items for Si. For the sub-items containing repeated items, store them as CiCj and multiply by multiple Si:
[0131] T = [(S1, C1), (S2, C2), (S3, C3),..., (Sn, Cn)]
[0132] Si = [(X1, CiCj), (X2, Ci), (X3, CiCj),..., (Xn, Ci)]
[0133] Matrix vectorization:
[0134] a. Set input parameters
[0135] The input parameters are the co-occurrence matrix T, the corpus dictionary D, the output dimension, and the number of iterations.
[0136] The co-occurrence matrix T is calculated by constructing the global and detailed co-occurrence matrices with a fixed window size of 10. The corpus dictionary D is calculated by establishing a word frequency, and the return value is D -> (a, f), which maps sentences to word IDs and corpus frequencies of words. The output dimension and the number of iterations are fixed values.
[0137] b. Construct a vector matrix
[0138] This matrix is (2V) * d, where V is the size of the corpus vocabulary and d is the dimension of the word vector. All elements are randomly initialized within the range (-0.5, 0.5). Two word vectors are constructed for each word: one word as the main (central) word and one word as the context word.
[0139] c. Construct the bias term:
[0140] An array of size 2V with a value range of (0.5, 0.5) is constructed through the above vectors.
[0141] d. Gradient descent method:
[0142] Training is performed using the gradient descent method.
[0143] e. Output the vector matrix:
[0144] If the input co-occurrence matrix is not empty, the provided function will call the known matrix after each iteration. Then, by passing the parameters to the loss function, the calculated word vector matrix is finally returned.
[0145] For example:
[0146] Co-occurrence matrix test with a window size of 10, an output dimension of 10, and 500 iterations.
[0147] Original sentence:
[0148] Usual health status: Poor health. Any history of infectious diseases? Denied. Name of infectious disease: Hepatitis.
[0149] Word frequency:
[0150] {'Health': (0, 1), 'Status': (1, 1), 'Health status': (2, 1), 'Poor': (3, 1), 'Any': (4, 1), 'Infectious disease': (5, 2), 'History': (6, 2), 'Denied': (7, 1), 'Name': (8, 1), 'Hepatitis': (9, 1)}
[0151] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any appropriate changes or substitutions made by any person of ordinary skill in the art that comply with the claims of a medical text word vectorization method of the present invention shall fall within the patent protection scope of the present invention.
[0152] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for vectorizing medical text words, characterized in that, Build a thesaurus for the original medical text data and vectorize the words in the medical text data through the GLOVE model; The original medical text is divided into structured data, semi-structured data, and unstructured data; In the construction of the GLOVE model, it includes the construction of large text segments, the construction of detailed text items, the construction of the global and detailed co-occurrence matrix, and matrix vectorization; the construction of large text segments is to divide the text into multiple sentences according to the sentence break symbol to obtain T = [S1, S2, S3, …, Sn]; perform detailed cutting on the above [S1, S2, S3, …, Sn] sentences to obtain S = [X1, X2, X3, …, Xn]; Because the GLOVE model calculates word vectors based on global information and will miss some relationships between contexts, the co-occurrence matrices of both texts S and T are calculated, and for the co-occurrence matrix calculated from text T, take: Si = X Si = ∑ k X Sik When calculating the co-occurrence matrix Xi generated by the S text, add Si to retain the connection between some contexts; In the construction of the global and detailed co-occurrence matrix, if there are no repeated words in each clause: when decomposing the T text, mark each S text with C, so that the subsequent co-occurrence matrix Xi is marked and used to calculate the corresponding sub-items for Si. The text after marking is expressed as: T’ = [(S1, C1), (S2, C2), (S3, C3), …, (Sn, Cn)] Si’ = [(X1, Ci), (X2, Ci), (X3, Ci), …, (Xn, Ci)] If there are repeated words in the clause: when decomposing the T text, mark each S text with C, so that the subsequent co-occurrence matrix Xi is marked and used to calculate the corresponding sub-items. For the sub-items containing repeated items, mark them as CiCj. The text after marking is expressed as: T” = [(S1, C1), (S2, C2), (S3, C3), …, (Sn, Cn)] Si” = [(X1, CiCj), (X2, Ci), (X3, CiCj), …, (Xn, Ci)]; In matrix vectorization, it includes: setting input parameters, where the input parameters are the co-occurrence matrix calculated in the construction of the global and detailed co-occurrence matrix; constructing a vector matrix; constructing a bias term; training through the gradient descent method; and outputting a vector matrix.
Citation Information
Patent Citations
A word vector representation learning method based on word pair asymmetric co-occurrence
CN109670171A
Real-time intelligent auxiliary ICD encoding system and method based on medical record
CN111462896A