A data augmentation method, device and equipment for clinical entity mapping
Through the data enhancement method of clinical entity mapping, multiple embedded models are used to calculate semantic relationship weights and perform data enhancement, which solves the accuracy and comprehensiveness of data statistics and analysis in clinical research, reduces labeling costs, and improves data quality and algorithm performance.
Patent Information
- Application Number
- CN202111538575.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In clinical research, it is difficult to accurately count and analyze the independent clinical data record specifications of different hospitals and departments in the prior art, resulting in inaccurate and incomplete statistical information, and the accuracy of the automatic mapping algorithm is low and the generalization ability is poor.
Provide a data enhancement method for clinical entity mapping. By obtaining the clinical entity word collection and manual labeling corpus, sampling and generating training sets, using word indexing model, word indexing model, word embedding model and word embedding model to calculate semantic relationship weights, select words with the highest semantic similarity for data enhancement, and form a training data set with data enhancement.
It effectively reduces the demand for manual labeling of data, reduces the production cost of data labeling, improves the quantity and quality of training data, and improves the accuracy and generalization ability of clinical entity mapping.
Smart Images

Figure CN114398894B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a data enhancement method, device, and equipment for clinical entity mapping. Background Art
[0002] In the process of clinical scientific research, doctors need to perform statistical analysis on clinical case information. However, the data sources of many electronic medical records are diverse, and each hospital or even department has an independent clinical data recording specification or mode. As a result, when statistically analyzing some key data (such as clinical entity information such as diseases, drugs, surgeries, symptoms, etc.), records that need to be concerned in scientific research cannot be queried from the database, and manual review and entity mapping are often required, ultimately leading to inaccurate and incomplete statistical information, heavy workload for doctors, low efficiency, and other problems.
[0003] In addition, because there are multiple standards for some key clinical entity information in China currently, when developing automatic mapping algorithms, multiple entity standards need to be independently annotated, which further exacerbates the production cost and quality problems of data annotation. Therefore, the accuracy of current these algorithms is relatively low, and the generalization ability is very poor, and they cannot be actually applied in clinical scientific research. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a data enhancement method, device, and equipment for clinical entity mapping.
[0005] The present disclosure provides a data enhancement method for clinical entity mapping, and the method includes:
[0006] Obtain a set of clinical entity words and an artificial annotation corpus; wherein, the set of clinical entity words includes: clinical entity words of an internal standard, and the artificial annotation corpus includes: original entity words that are not internal standards and their annotated first internal standard entity words;
[0007] Sample the artificial annotation corpus, and the obtained sampling result includes: a first training set and a development set;
[0008] According to a preset character index model and word index model, generate a first entity word list corresponding to the original entity words in the sampling result; wherein, the words in the first entity word list are words that satisfy character semantic similarity, word semantic similarity, and conform to the internal standard;
[0009] According to a character embedding model and a word embedding model, calculate the first semantic relationship weight scores between the original entity words and each word in the first entity word list;
[0010] Select the top several words with the highest semantic similarity from the first entity word list according to the first semantic relationship weight score to obtain the second internal standard entity words, and obtain the second training set based on the selected second internal standard entity words and the original entity words;
[0011] Select the first type of negative samples from the second internal standard entity words of the second training set; select the second type of negative samples from the clinical entity word set; select the positive samples from the first internal standard entity words of the manually annotated corpus;
[0012] Through random sampling and insertion, form triples with the positive samples, the first type of negative samples, and the second type of negative samples to obtain a data-augmented training data set.
[0013] The present disclosure provides a data augmentation device for clinical entity mapping, and the device includes:
[0014] A corpus acquisition module for acquiring a clinical entity word set and a manually annotated corpus; wherein, the clinical entity word set includes: internally standard clinical entity words, and the manually annotated corpus includes: non-internally standard original entity words and their annotated first internally standard entity words;
[0015] A sampling module for sampling the manually annotated corpus, and the sampling result includes: a first training set and a development set;
[0016] A list generation module for generating a first entity word list corresponding to the original entity words in the sampling result according to a preset character index model and word index model; wherein, the words in the first entity word list are words that satisfy character semantic similarity, word semantic similarity, and conform to internal standards;
[0017] A weight calculation module for calculating the first semantic relationship weight scores between the original entity words and each word in the first entity word list according to a character embedding model and a word embedding model;
[0018] A data selection module for selecting the top several words with the highest semantic similarity from the first entity word list according to the first semantic relationship weight score to obtain the second internal standard entity words, and obtaining the second training set based on the selected second internal standard entity words and the original entity words;
[0019] A sample selection module for selecting the first type of negative samples from the second internal standard entity words of the second training set; selecting the second type of negative samples from the clinical entity word set; selecting the positive samples from the first internal standard entity words of the manually annotated corpus;
[0020] A data augmentation module, which is used to form triples by randomly sampling and inserting the positive samples, the first type of negative samples, and the second type of negative samples, so as to obtain a data-augmented training data set.
[0021] The present disclosure provides an electronic device, which includes: a processor; a memory for storing executable instructions executable by the processor;
[0022] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above method.
[0023] The present disclosure provides a computer-readable storage medium, which stores a computer program for executing the above method.
[0024] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:
[0025] The embodiments of the present disclosure provide a data augmentation method, device, and equipment for clinical entity mapping, which can more efficiently mine the set of clinical entity words and utilize a small-scale manually annotated corpus, effectively reduce the demand for manually annotated data, perform data augmentation processing on a small number of original entity words, that is, according to the semantic relationships comprehensively calculated by the character index model, word index model, character embedding model, word embedding model, etc. between clinical entity words, a large number of positive and negative samples are mined and made to conform to a certain occurrence probability and order, thereby constituting a data-augmented training data set. The embodiments of the present disclosure can reduce the labor cost of data annotation and improve the quantity and quality of training data. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a flowchart of the data augmentation method for clinical entity mapping according to the embodiments of the present disclosure;
[0029] Figure 2 It is a schematic diagram of the model architecture of the character embedding model and the word embedding model according to the embodiments of the present disclosure;
[0030] Figure 3Block diagram of the data augmentation device for clinical entity mapping according to an embodiment of the present disclosure;
[0031] Figure 4 Schematic diagram of the structure of the electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0032] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0033] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0034] The embodiments of the present disclosure provide a data augmentation method, device and equipment for clinical entity mapping. The technical solution can effectively reduce the demand for manually labeled data, reduce the production cost of data annotation, perform data augmentation processing on a small number of labeled samples, and increase the quantity and quality of training data. For ease of understanding, the embodiments of the present disclosure are described below.
[0035] Figure 1 Flowchart of the data augmentation method for clinical entity mapping provided by an embodiment of the present disclosure. The method provided in this embodiment includes the following steps:
[0036] Step S1, obtain a set of clinical entity words and a manually labeled corpus. Among them, the set of clinical entity words includes: internally standard clinical entity words, and the manually labeled corpus includes: non-internally standard original entity words and their labeled first internally standard entity words.
[0037] The manner of obtaining the set of clinical entity words in this embodiment includes semantically integrating multiple specific standard databases of a certain clinical entity (such as surgery) to form a unified internal standard. Through semantic clustering algorithms and manual review of these different version standards, accurate mapping between multiple specific standards (i.e., non-internal standards) and internal standards can be achieved, and a complete set of internally standard clinical entity words and a word mapping set between non-internal standards and internal standards can be formed. The above semantic clustering algorithm refers to a word vector language model obtained by introducing external Internet big data statistical learning, semantically representing entity words of different standards, calculating the semantic distance between any two entity words, selecting word pairs with a semantic similarity score exceeding a preset threshold as candidate clustering results, and then through methods such as manual review, using the best words as standard words and other original words as mapping words corresponding to the respective standards.
[0038] The process of obtaining the manually annotated corpus in this embodiment includes:
[0039] Step S1.1: Obtain the first original entity words that are not internal standards from the user's historical annotation data. In actual clinical research, relevant staff such as doctor experts will perform a small amount of manual annotation and mapping for a specific standard to reduce the data workload of doctors; this annotation can provide the manual conversion of the original clinical entity words in the electronic medical record to a specific standard. In this case, this embodiment can obtain the first original entity words that are not internal standards from the recorded historical annotation data.
[0040] Step S1.2: Map the first original entity words to internal standard entity words that conform to the internal standard according to the preset set of word mappings between non-internal standards and internal standards. The above set of word mappings between non-internal standards and internal standards is obtained during the process of obtaining the set of clinical entity words.
[0041] Step S1.3: From the preset data, count the second original entity words whose occurrence frequency is higher than the preset frequency threshold. Perform data annotation on the second original entity words to obtain internal standard entity words that conform to the internal standard.
[0042] The preset database is generally a raw massive medical information database. In this embodiment, higher-frequency clinical entity words are statistically mined from the database and used as the second original entity words. Perform data annotation on the second original entity words, map the second original entity words to entity words that conform to the internal standard, and obtain internal standard entity words. This method avoids the problem of repeated annotation of multiple different standards.
[0043] Step S1.4: Use the first original entity words and their annotated internal standard entity words, and the second original entity words and their annotated internal standard entity words as the manually annotated corpus. The manually annotated corpus includes: the original entity words that are not internal standards composed of the first original entity words and the second original entity words, and the first internal standard entity words of the internal standard composed of the internal standard entity words annotated by the first original entity words and the second original entity words respectively; it can be understood that there is a one-to-one annotation relationship between the original entity words and the first internal standard entity words in the manually annotated corpus.
[0044] The annotation data in this embodiment includes two parts of data, namely historical annotation data and entity words that frequently appear in the database. The quantity does not need to be too much, which significantly reduces the pressure of manual annotation, and at the same time, the quality is relatively high. The high-quality manually annotated corpus thus formed can be used as the data source for subsequent data mining and algorithm training.
[0045] Step S2, sample the manually annotated corpus, and the sampling result includes: a first training set and a development set.
[0046] Randomly sample the manually annotated corpus. The sampling result can include a first training set and a development set. In addition, for applying to actual model training, the sampling result can also include a test set. For example, the number of corpus records in the first training set accounts for 70% of the total number of corpus records, the development set accounts for 15%, and the test set accounts for 15%. Each record consists of the original entity word and its corresponding first internal standard entity word.
[0047] Step S3, construct the clinical entity word set into a character index model and a word index model.
[0048] In this embodiment, for each original entity word in the manually annotated corpus, in order to generate a list of internal standard entity words that simultaneously satisfy the double semantic similarity of the character index model and the word index model, the clinical entity word set can be first constructed into a character index model and a word index model respectively. Among them, the character index model is used to represent entity words by using sparse character vectors, and the word index model is used to represent entity words by using sparse word vectors; the vector values of characters and words are calculated respectively according to the statistical mining of the characters and words in a large-scale corpus.
[0049] Specifically, the specific calculation method of the character index model is:
[0050] V ti =a 1 ·L ti +a 2 ·M ti +a 3 ·N ti (1)
[0051] Among them, V ti is the inverse document frequency comprehensive weight value of each character in the character index model, L ti , M ti , N ti are respectively the reciprocals of the frequencies of the character in the large-scale Internet corpus database, the massive medical informatization database, and various clinical entity word sets, and a 1 , a 2 , a 3 are their corresponding weight value coefficients respectively.
[0052] The specific calculation method of the word index model is:
[0053] W qi =b 1 ·L qi +b 2 ·Mqi +b 3 ·N qi (2)
[0054] Among them, W qi is the comprehensive inverse document frequency weight value of each word in the word index model, and L qi , M qi , N qi are respectively the reciprocals of the frequencies of the word appearing in the large-scale Internet corpus database, the massive medical informatization database, and various clinical entity word sets, while b 1 , b 2 , b 3 are respectively their corresponding weight value coefficients.
[0055] Perform the data selection operation shown in the following steps S4 - S6 on the sampling results to obtain the second training set:
[0056] Step S4, generate a first entity word list corresponding to the original entity words in the sampling results according to the preset character index model and word index model; among them, the words in the first entity word list are words that meet the requirements of character semantic similarity, word semantic similarity, and comply with internal standards.
[0057] Specific embodiments include: Step S4.1, generate a second entity word list corresponding to the original entity words in the manually annotated corpus according to the character index model; among them, the words in the second entity word list comply with internal standards, and the semantic relationship weight score between them and the original entity words meets the requirement of character semantic similarity.
[0058] The process of generating the second entity word list in this embodiment includes: First, calculate the comprehensive inverse document frequency weight value of each character in the character index model according to the above formula (1), and then obtain the character index model vector representations of the original entity words and each internal standard clinical entity word respectively in the following manner: According to the comprehensive inverse document frequency weight value of each character and the frequency value weight of the character appearing in the current clinical entity word, obtain the character index model vector representation of the current word, referring to the following formula (3):
[0059] S Z =[k 1 ·V t1 ,k 2 ·V t2 ,k 3 ·V t3 ,…,k i ·V ti ,…,k n ·V tn (3)
[0060] Among them, SZ is the character index model vector representation of the current clinical entity term Z, V t1 , V t2 , V t3 ,..., V ti ,..., V tn is the inverse document frequency comprehensive weight value of each character in the corresponding character index model calculated according to the above formula (1), k 1 , k 2 , k 3 ,..., k i ,..., k n is the frequency value weight of the corresponding character appearing in the current term.
[0061] As shown in formula (4), calculate the semantic relationship weight score D(A, B) based on the character index model vector representation between the original entity term A and any internal standard clinical entity term B:
[0062]
[0063] wherein, wherein S A and S B are both character index model vector representations calculated according to the above formula (1) and formula (3), V ai and V bi are respectively the weight value parameters of a series of character indexes corresponding to the original entity term A and the internal standard clinical entity term B, originating from k in the above formula (3) i ·V ti .
[0064] Select multiple internal standard clinical entity terms with the semantic relationship weight score based on the character index model vector representation higher than the preset score value to obtain the second entity term list. Specifically, after calculating the semantic relationship weight score between any two clinical entity terms through the above formula (4), select multiple entity terms with the semantic relationship weight score higher than the preset score value. The selected terms are semantically similar to the original entity term, and the selected multiple terms constitute the second entity term list.
[0065] Step S4.2, generate a third entity term list corresponding to the original entity term in the manually annotated corpus according to the word index model; wherein, the terms in the third entity term list meet the internal standards and the semantic relationship weight score between them and the original entity term satisfies word semantic similarity.
[0066] The method for generating the third entity term list in this embodiment is basically the same as the principle of the method for generating the second entity term list above, and some specific calculation methods involved in this process are referred to as follows.
[0067] The specific calculation method for realizing the semantic relationship between two entity words based on the word index model is as follows:
[0068] T Z =[k 1 ·W q1 ,k 2 ·W q2 ,k 3 ·W q3 ,…,k i ·W qi ,…,k n ·W qn (5)
[0069] Among them, T Z is the word index model vector representation of the current clinical entity word Z, and W q1 , W q2 , W q3 ,... W qi ,... W qn are the comprehensive inverse document frequency weight values of each word in the corresponding word index calculated according to the above formula (2), and k 1 , k 2 , k 3 ,... k i ,... k n are the frequency value weights of the corresponding word appearing in the current word.
[0070]
[0071] Among them, E(A,B) is the semantic relationship weight score of the original entity word A and any internal standard clinical entity word B based on the word index model vector representation, where T A and T B are both calculated according to the above formulas (2) and (5), and W ai and W bi are the weight value parameters of a series of word indexes corresponding to the original entity word A and the internal standard clinical entity word B respectively, which are derived from k i ·W qi in the above formula (5).
[0072] Step S4.3, according to the second entity word list and the third entity word list, obtain the first entity word list that simultaneously satisfies character semantic similarity and word semantic similarity.
[0073] In this embodiment, according to the second entity word list and the third entity word list, the weights are calculated by combining and referring to the following formula (7) to obtain the first entity word list corresponding to each original entity word, which meets the internal standard of double semantic similarity that both the character semantic similarity and the word semantic similarity are satisfied:
[0074] F(A,B) = p 1 ·D(A,B) + p 2 ·E(A,B) (7)
[0075] Among them, F(A,B) is the semantic relationship weight score of the double semantic similarity recall of the above entity words A and B, and D(A,B) and E(A,B) are respectively calculated by the above formula (4) and formula (6), p 1 and p 2 are respectively the semantic weight influence factors of the character index model and the word index model.
[0076] Step S5: Calculate the first semantic relationship weight scores between the original entity words and each word in the first entity word list according to the character embedding model and the word embedding model.
[0077] The first, second, and third entity word lists and the corresponding semantic relationship weight scores in the above step S4 are recalled and scored according to the vector index and retrieval method based on the character index model and the word index model, and the evaluated is the literal shallow semantic relationship of a clinical entity word; in contrast, at the current stage, the implicit semantic relationship of the character embedding model and the word embedding model can be utilized and the corresponding implicit semantic distance can be calculated.
[0078] In this embodiment, the character embedding model and the word embedding model are used to evaluate the best weight allocation, and at the same time, it is combined and calculated into the semantic relationship weight scores in step S4 to obtain the clinical entity words that meet the four semantic relationships of the character index model, the word index model, the character embedding model, and the word embedding model. Furthermore, the clinical entity words are sorted, and the first multiple (such as the first 1000) words with similar semantics are selected as the data set for subsequent further training.
[0079] For the character embedding model and the word embedding model, two similar context prediction model architectures are adopted, as Figure 2 shown. By comprehensively using the external large-scale Internet corpus database and the massive medical informatization database records, as well as various clinical entity word sets, it is mined. Taking the context vector representation of a character or word as the input layer, through the non-linear transformation and semantic information of the hidden layer, the probability result of each situation of the unknown word in the output layer is calculated, or taking the vector representation of a character or word as the input layer, through the non-linear transformation and semantic information of the hidden layer, the probability result of each situation of other context words in the output layer is calculated.
[0080] For a specific embodiment of calculating the first semantic relationship weight score between the original entity word and each word in the first entity word list according to the above character embedding model and word embedding model, refer to the following.
[0081] Step S5.1, referring to the following formulas (8)-(11), based on the character embedding model, calculate the first implicit semantic distance between the original entity word and each word in the first entity word list. The first implicit semantic distance can represent a new semantic weight of each entity word in the first entity word list.
[0082] The specific calculation method for generating the first implicit semantic distance based on the character embedding model is as follows:
[0083] M a =[m 1 ,m 2 ,m 3 ,…,m i ,…,m 300 (8)
[0084] Among them, M a is the implicit semantic vector representation of a certain character a in the character embedding model, the vector dimension is 300, and m 1 , m 2 , m 3 ,..., m i ,..., m 300 are the probability values of the corresponding character a in these 300 implicit semantic spaces.
[0085] P Z =[p 1 ,p 2 ,p 3 ,…,p i ,…,p 300 (9)
[0086]
[0087] Among them, P Z is the implicit semantic vector representation of the current clinical entity word Z based on the character embedding model, the vector dimension is 300, and p 1 , p 2 , p 3 ,..., p i ,..., p 300 are the probability values of the corresponding word Z in these 300 implicit semantic spaces. The word Z contains n characters, and the probability value m ij of the implicit semantic vector of each character in these 300 implicit semantic spaces is calculated by the above formula (8), k jis the influence factor reflected by a corresponding character j in the current word Z.
[0088]
[0089] Among them, G(A, B) is the first implicit semantic distance between the original entity word A and the internal standard clinical entity word B based on the character embedding model representation, where P A and P B are respectively calculated by the above formulas (9) and (10), and p ai and p bi are respectively the probability values of the implicit semantic vector representations corresponding to the two clinical entity words A and B in 300 implicit semantic spaces, which are derived from p i in the above formula (10).
[0090] Step S5.2: Based on the word embedding model, calculate the second implicit semantic distance between the original entity word and each word in the first entity word list based on words.
[0091] In this embodiment, the calculation method of the second implicit semantic distance is the same as the above calculation method of the first implicit semantic distance. The specific calculation method for generating the second implicit semantic distance based on the word embedding model is as follows:
[0092] N b =[n 1 , n 2 , n 3 ,…, n i ,…, n 600 (12)
[0093] Among them, N b is the implicit semantic vector representation of a certain word b in the word embedding model, the vector dimension is 600, and n 1 , n 2 , n 3 ,..., n i ,..., n 600 are the probability values of the corresponding word b in these 600 implicit semantic spaces.
[0094] Q Z =[q 1 , q 2 , q 3 ,…, q i ,…, q 600 (13)
[0095]
[0096] Among them, Q Zis the implicit semantic vector representation of the current clinical entity term Z based on the word embedding model, with a vector dimension of 600, q 1 , q 2 , q 3 ,..., q i ,..., q 600 are the probability values of the corresponding term Z in these 600 implicit semantic spaces. The term Z contains n words, and the probability values of the implicit semantic vectors of each word in these 600 implicit semantic spaces are n ij are calculated from the above formula (12), k j is the influence factor of a corresponding word j reflected in the current term Z.
[0097]
[0098] Among them, H(A,B) is the second implicit semantic distance between the original entity term A and the internal standard clinical entity term B based on the word embedding model representation, where Q A and Q B are calculated from the above formula (13) and formula (14) respectively, q ai and q bi are the probability values of the corresponding implicit semantic vector representations of the above entity terms A and B in 600 implicit semantic spaces, originating from q i .
[0099] Step S5.3, calculate the first semantic relationship weight score T(A,B) between the original entity term and each term in the first entity term list according to the first implicit semantic distance and the second implicit semantic distance.
[0100] Referring to the following formula (16), score the first entity term list according to the first implicit semantic distance and the second implicit semantic distance, so as to obtain entity terms that satisfy a total of four semantic relationships of the character index model, word index model, character embedding model, and word embedding model:
[0101] T(A,B) = q 1 ·G(A,B) + q 2 ·H(A,B) + q 3 ·F(A,B) (16)
[0102] Among them, T(A,B) is the semantic relationship weight score between the original entity term A and the internal standard clinical entity term B after comprehensive calculation by four semantic models. G(A,B), H(A,B), and F(A,B) are calculated from the above formula (11), formula (15), and formula (7) respectively, q 1 , q 2 and q3 They are respectively the semantic influence factors of the character embedding model, the word embedding model, and the dual semantic relationship in step (10).
[0103] Step S6: Select the top several words with the highest semantic similarity from the first entity word list according to the first semantic relationship weight score to obtain the second internal standard entity words, and obtain the second training set based on the selected second internal standard entity words and the original entity words.
[0104] Specifically, sort the words in the first entity word list according to the first semantic relationship weight score T(A, B) from high to low, select the top 1000 words as the second internal standard entity words according to the sorting result, and form the second training set with the original entity words.
[0105] Next, based on the second training set, this embodiment performs the data augmentation operations shown in the following steps S7 and S8 to obtain the data-augmented training data set.
[0106] Step S7: Select the first type of negative samples from the second internal standard entity words in the second training set; select the second type of negative samples from the clinical entity word set; select the positive samples from the first internal standard entity words in the manually annotated corpus.
[0107] Specifically, randomly select a certain number (such as 50) of the first type of negative samples from the second internal standard entity words in the second training set; the first type of negative samples and the original entity words have a high semantic similarity and are words that are relatively difficult to distinguish. Exclude the manually annotated internal standard entity words in the first training set and the development set. Finally, for each original entity word, the mining result of the first type of negative samples is obtained, and the specific content is a sequence of internal standard clinical entity words and the corresponding semantic relationship weight scores. The elements of this sequence U are composed as follows:
[0108] U = {[w a1 , k a1 , [w a2 , k a2 , [w a3 , k a3 , …, [w a50 , k a50} (17)
[0109] Among them, U refers to the word sequence of the current first type of negative samples, w a1 , w a2 , w a3 , …, w a50 are 50 randomly selected first type of negative samples, and k a1 , k a2 , k a3 , …, ka50 It is the comprehensive semantic relationship weight score T(A,B) between two entity words calculated according to formula (16).
[0110] Randomly select the same number of second-class negative samples from the clinical entity word set. It is easy to understand that the second-class negative samples and the original entity words have little semantic similarity and are relatively easy to distinguish words. Exclude the manually labeled internal standard entity words in the first training set and the development set, and exclude the second training set. Finally, for each original entity word, the second-class negative sample mining result is obtained. The specific content is an internal standard clinical entity word sequence and the corresponding semantic relationship weight score. The elements of this sequence V are composed as follows:
[0111] V = {[s b1 ,k b1 ,[s b2 ,k b2 ,[s b3 ,k b3 ,…,[s b50 ,k b50} (18)
[0112] Among them, V refers to the word sequence of the current second-class negative samples. s b1 , s b2 , s b3 ,... s b50 are randomly selected from the clinical entity word set in step S1, and k b1 , k b2 , k b3 ,... k b50 are the comprehensive semantic relationship weight scores T(A,B) between two entity words calculated according to formula (16).
[0113] Select positive samples from the first internal standard entity words in the manually labeled corpus.
[0114] Step S8, through random sampling and insertion, form triples with positive samples, first-class negative samples, and second-class negative samples to obtain an enhanced training data set.
[0115] Randomly insert the positive samples into the first-class negative samples and the second-class negative samples multiple times to satisfy the triples composed of one positive sample, one first-class negative sample, and one second-class negative sample. The currently inserted positive sample and its position sequence are the positive sample mining results. The specific content is an internal standard clinical entity word sequence and the corresponding semantic relationship weight score. The elements of this sequence W are composed as follows:
[0116] W = {[t c1 ,k c1,[t c2 ,k c2 ,[t c3 ,k c3 ,…,[t c50 ,k c50} (19)
[0117] Among them, W refers to the current positive sample word sequence, t c1 , t c2 , t c3 ,... t c50 are positive samples randomly inserted, and k c1 , k c2 , k c3 ,... k c50 are the comprehensive semantic relationship weight scores between two entity words, and the positive sample is defaulted to 1.
[0118] In this embodiment, positive samples, the first type of negative samples, and the second type of negative samples are combined to produce a new training data set that has been effectively processed multiple times and is data-augmented and complete. This training data set can contain the true matching semantics of the original entity words and the internal standard entity word set, as well as complex semantic relationships that are easy to distinguish and difficult to distinguish. This training data set Z consists of a series of triples:
[0119] Z = {z 1 , z 2 , z 3 ,…, z n ,…, z 50} (20)
[0120] Among them, each triple is generated by randomly sampling and inserting positive samples into the first type of negative samples and the second type of negative samples. The specific composition is:
[0121] z i = {[w ai , k ai , [s bi , k bi , [t ci , k ci} (21)
[0122] Among them, [w ai , k ai , [s bi , k bi , [t ci , k ci are the first type of negative samples, the second type of negative samples, and positive samples respectively.
[0123] The second training set and the data-augmented training data set obtained according to the above embodiments can be applied to the training process of the model.
[0124] In summary, the data augmentation method for clinical entity mapping provided by the embodiments of the present disclosure can more efficiently mine the clinical entity word set and utilize a small-scale manually annotated corpus, effectively reducing the demand for manually annotated data. Data augmentation processing is performed on a small number of original entity words, that is, a large number of positive and negative samples are mined according to the semantic relationships comprehensively calculated by the character index model, word index model, character embedding model, word embedding model, etc. between clinical entity words, and they are made to conform to a certain occurrence probability and order, thereby constituting a data-augmented training data set.
[0125] Compared with the traditional method, the data augmentation method of the present technical solution only needs to collect a relatively small amount of manually annotated data, and only needs the historical annotations made by expert doctors on a certain specific standard in their own clinical research. There is no need for expert doctors to specifically perform data annotation work, which minimizes the workload and complexity of expert doctors; in addition, external annotators annotate according to the unified internal standard and do not need to repeat the annotation in each different standard data task, which also greatly improves the work efficiency. This data augmentation solution can greatly expand the original manually annotated data and maximize the mining of various implicit semantic relationships in the clinical entity word set. Therefore, the present disclosure can improve the quantity and quality of the annotated data on the basis of reducing the manual annotation cost.
[0126] Referring to Figure 3 , the embodiments of the present disclosure provide a data augmentation device for clinical entity mapping, and the device includes:
[0127] A corpus acquisition module 302, configured to acquire a clinical entity word set and a manually annotated corpus; wherein, the clinical entity word set includes: clinical entity words of the internal standard, and the manually annotated corpus includes: original entity words of non-internal standard and their annotated first internal standard entity words;
[0128] A sampling module 304, configured to sample the manually annotated corpus, and the sampling result includes: a first training set and a development set;
[0129] A list generation module 306, configured to generate a first entity word list corresponding to the original entity words in the sampling result according to a preset character index model and word index model; wherein, the words in the first entity word list are words that satisfy character semantic similarity, word semantic similarity, and conform to the internal standard;
[0130] A weight calculation module 308, configured to calculate a first semantic relationship weight score between the original entity word and each word in the first entity word list according to a word embedding model and a word embedding model;
[0131] A data selection module 310, configured to select the top several words with the highest semantic similarity from the first entity word list according to the first semantic relationship weight score to obtain a second internal standard entity word, and obtain the second training set based on the selected second internal standard entity word and the original entity word;
[0132] A sample selection module 312, configured to select a first type of negative sample from the second internal standard entity words of the second training set; select a second type of negative sample from the clinical entity word set; select a positive sample from the first internal standard entity words of the manually annotated corpus;
[0133] A data augmentation module 314, configured to form triples by randomly sampling and inserting the positive sample, the first type of negative sample, and the second type of negative sample to obtain a data-augmented training data set.
[0134] The device provided in this embodiment has the same implementation principle and the same technical effects as those in the foregoing method embodiment. For a brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0135] Figure 4 It is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. As Figure 4 shown, the electronic device 400 includes one or more processors 401 and a memory 402.
[0136] The processor 401 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.
[0137] The memory 402 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may run the program instructions to implement the methods of the embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0138] In one example, the electronic device 400 may further include: an input device 403 and an output device 404, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0139] In addition, the input device 403 may further include, for example, a keyboard, a mouse, and so on.
[0140] The output device 404 may output various information to the outside, including the determined distance information, direction information, etc. The output device 404 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0141] Of course, for simplicity, Figure 4 only some of the components related to the present disclosure in the electronic device 400 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 400 may further include any other appropriate components.
[0142] Furthermore, this embodiment also provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the above-mentioned data enhancement method for clinical entity mapping.
[0143] A computer program product of a data enhancement method, device, electronic device, and medium for clinical entity mapping provided by an embodiment of the present disclosure includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated herein.
[0144] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0145] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data augmentation method for clinical entity mapping, characterized in that, the method includes: Obtain a set of clinical entity words and an artificially annotated corpus; wherein, the set of clinical entity words includes: internally standard clinical entity words, and the artificially annotated corpus includes: non-internally standard original entity words and their annotated first internally standard entity words; Sample the artificially annotated corpus, and the sampling result includes: a first training set and a development set; According to a preset character index model and word index model, generate a first entity word list corresponding to the original entity words in the sampling result; wherein, the words in the first entity word list are words that meet character semantic similarity and word semantic similarity and conform to internal standards; the character index model is used to represent entity words with sparse character vectors, and the word index model is used to represent entity words with sparse word vectors; According to the character embedding model and the word embedding model, calculate the first semantic relationship weight scores between the original entity words and each word in the first entity word list; Select the top multiple words with the highest semantic similarity from the first entity word list according to the first semantic relationship weight scores to obtain second internally standard entity words, and obtain a second training set based on the selected second internally standard entity words and the original entity words; Select the first type of negative samples from the second internally standard entity words in the second training set; select the second type of negative samples from the set of clinical entity words; select positive samples from the first internally standard entity words in the artificially annotated corpus; Through random sampling and insertion, form triples with the positive samples, the first type of negative samples, and the second type of negative samples to obtain a data-augmented training data set.
2. The method according to claim 1, characterized in that, the step of obtaining the artificially annotated corpus includes: Obtain the non-internally standard first original entity words from the user's historical annotation data; According to a preset set of word mappings between non-internally standard and internally standard, map the first original entity words to internally standard entity words that conform to internal standards; From the preset data, count the second original entity words whose occurrence frequency is higher than the preset frequency threshold; Perform data annotation on the second original entity words to obtain internally standard entity words that conform to internal standards; Use the first original entity words and their annotated internally standard entity words, and the second original entity words and their annotated internally standard entity words as the artificially annotated corpus.
3. The method according to claim 1, characterized in that, the method further includes: Construct the set of clinical entity words into the character index model and the word index model.
4. The method according to claim 3, characterized in that, the generating a first entity word list corresponding to the original entity words in the sampling result according to a preset character index model and word index model includes: Generate a second entity word list corresponding to the original entity words in the manually annotated corpus according to the character index model; wherein, the words in the second entity word list conform to internal standards, and the semantic relationship weight score between the words and the original entity words satisfies character semantic similarity; Generate a third entity word list corresponding to the original entity words in the manually annotated corpus according to the word index model; wherein, the words in the third entity word list conform to internal standards, and the semantic relationship weight score between the words and the original entity words satisfies word semantic similarity; Obtain a first entity word list that simultaneously satisfies character semantic similarity and word semantic similarity according to the second entity word list and the third entity word list.
5. The method according to claim 4, wherein, The generating a second entity word list corresponding to the original entity words in the manually annotated corpus according to the character index model includes: Calculate the inverse document frequency comprehensive weight value of each character in the character index model; Obtain the character index model vector representations of the original entity words and the clinical entity words of each internal standard in the following manner: According to the inverse document frequency comprehensive weight value of each character and the frequency value weight of the character in the current word, obtain the character index model vector representation of the current word; Calculate the semantic relationship weight score based on the character index model vector representation between the original entity word and any clinical entity word of the internal standard; Select multiple clinical entity words of the internal standard with the semantic relationship weight score based on the character index model vector representation higher than the preset score value to obtain the second entity word list.
6. The method according to claim 5, wherein, The calculating the semantic relationship weight score based on the character index model vector representation between the original entity word and any clinical entity word of the internal standard includes: Among them, D(A, B) is the semantic relationship weight score based on the character index model vector representation between the original entity word A and the internal standard clinical entity word B, and S A and S B are the character index model vector representations of the original entity word A and the clinical entity word B respectively.
7. The method according to claim 1, wherein, The calculating the first semantic relationship weight score between the original entity word and each word in the first entity word list according to the character embedding model and the word embedding model includes: Based on the character embedding model, calculate the first implicit semantic distance based on characters between the original entity word and each word in the first entity word list; Based on the word embedding model, calculate the second implicit semantic distance based on words between the original entity word and each word in the first entity word list; Calculate the first semantic relationship weight score between the original entity word and each word in the first entity word list according to the first implicit semantic distance and the second implicit semantic distance.
8. A data augmentation device for clinical entity mapping, wherein, The device includes: A corpus acquisition module for acquiring a set of clinical entity words and a manually annotated corpus; wherein, the set of clinical entity words includes: clinical entity words of internal standards, and the manually annotated corpus includes: original entity words that are not internal standards and their annotated first internal standard entity words; A sampling module for sampling the manually annotated corpus, and the sampling result includes: a first training set and a development set; A list generation module for generating a first entity word list corresponding to the original entity words in the sampling result according to a preset character index model and a word index model; wherein, the words in the first entity word list are words that satisfy character semantic similarity and word semantic similarity and conform to internal standards; the character index model is used to represent entity words with sparse character vectors, and the word index model is used to represent entity words with sparse word vectors; A weight calculation module for calculating a first semantic relationship weight score between the original entity words and each word in the first entity word list according to a character embedding model and a word embedding model; A data selection module for selecting the top several words with the highest semantic similarity from the first entity word list according to the first semantic relationship weight score to obtain second internal standard entity words, and obtaining a second training set based on the selected second internal standard entity words and the original entity words; A sample selection module for selecting a first type of negative sample from the second internal standard entity words in the second training set; selecting a second type of negative sample from the clinical entity word set; and selecting a positive sample from the first internal standard entity words in the manually annotated corpus; A data augmentation module for forming triples by randomly sampling and inserting the positive sample, the first type of negative sample, and the second type of negative sample to obtain a data-augmented training data set.
9. An electronic device, characterized in that, the electronic device includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-7 above.
10. A computer-readable storage medium, characterized in that, the storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1-7 above.
Citation Information
Patent Citations
Medical term automatic standardization system and method integrating self-supervision and active learning
CN113436698A
Clinical term standardization method and device, electronic equipment and storage medium
CN113593661A