Method, apparatus, device, and storage medium for medical terminology standardization
By combining a label prefix tree with a seq2seq model, the problems of low accuracy and efficiency in the standardization of medical terminology are solved, and efficient and comprehensive standardization of medical terminology is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies suffer from insufficient accuracy and high time costs in the standardization of medical terminology, especially when using sequence-to-sequence models for medical terminology standardization, which cannot guarantee the accuracy and efficiency of medical terminology standardization.
By constructing a label prefix tree, the corresponding standard terminology set is searched from the target classification label based on the semantic features of medical terms, avoiding probability prediction for each standard term, and combining it with a seq2seq model to standardize medical terms.
It improves the accuracy and efficiency of medical terminology standardization, ensures the comprehensiveness of terminology standardization, reduces the probability prediction of each standard term in the medical knowledge base, and increases processing speed.
Smart Images

Figure CN116012862B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of natural language processing, in particular to a medical terminology standardization method, device, equipment and storage medium. BACKGROUND
[0002] In the medical field, the clinical terminology standardization task has become an indispensable task in medical data statistics. For the same medical diagnosis, surgery, medicine, medical examination, test, symptom, etc., there are usually multiple different medical terms. Therefore, in order to ensure the standardization of medical data statistics, it is necessary to convert the original medical terms into unified standard terms, so that researchers can perform corresponding diagnostic analysis on the electronic medical records of each patient.
[0003] Under normal circumstances, the original medical terms can be analyzed by a pre-constructed sequence-to-sequence model to predict the character probability of each standard term in the knowledge base, so as to obtain the corresponding standard term of the medical term. However, considering that the knowledge base includes a large number of cumbersome standard terms, the sequence-to-sequence model is used to predict the character probability of each standard term to standardize the original medical term, which consumes a lot of time cost and cannot guarantee the accuracy of medical terminology standardization, so that medical terminology standardization has certain limitations. SUMMARY
[0004] The embodiments of the present application provide a medical terminology standardization method, device, equipment and storage medium, which ensures the accuracy of medical terminology standardization and improves the efficiency and comprehensiveness of medical terminology standardization.
[0005] In a first aspect, the embodiments of the present application provide a medical terminology standardization method, which comprises:
[0006] determining the semantic features of any medical term;
[0007] According to the semantic features, searching the standard term set corresponding to the medical term from the label prefix tree constructed under at least one target classification label to which the medical term belongs;
[0008] The label prefix tree is composed of a plurality of standard terms under the target classification label.
[0009] In a second aspect, the embodiments of the present application provide a medical terminology standardization device, which comprises:
[0010] The feature determination module is configured to determine the semantic features of any medical term;
[0011] The medical term standardization module is configured to search, according to the semantic feature, a standard term set corresponding to the medical term from a label prefix tree constructed under at least one target classification label to which the medical term belongs.
[0012] The label prefix tree is composed of a plurality of standard terms under the target classification label.
[0013] In a third aspect, an electronic device is provided, and the electronic device includes:
[0014] The processor and the memory are configured to store a computer program, and the processor is configured to invoke and run the computer program stored in the memory to execute the medical term standardization method provided in the first aspect.
[0015] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium is configured to store a computer program, and the computer program is configured to enable a computer to execute the medical term standardization method provided in the first aspect.
[0016] In a fifth aspect, a computer program product is provided, and the computer program product includes computer programs / instructions, and the computer programs / instructions are configured to enable a processor to implement the medical term standardization method provided in the first aspect.
[0017] The medical term standardization method, device, equipment, and storage medium provided in the embodiments of the present application are configured to, for any medical term, first determine the semantic feature of the medical term. Then, according to the semantic feature, search, from a label prefix tree constructed under at least one target classification label to which the medical term belongs, a standard term set corresponding to the medical term, avoid missing the standard term after standardization of the medical term, and thus ensure the comprehensiveness of medical term standardization. Moreover, without performing probability prediction between each standard term in a medical knowledge base and the medical term, on the basis of ensuring the accuracy of medical term standardization, the label prefix tree constructed under the target classification label is used to improve the efficiency of medical term standardization. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0019] Figure 1 A flowchart of a medical term standardization method according to an embodiment of the present application is shown.
[0020] Figure 2 A structural diagram of a sequence-to-sequence model shown in an embodiment of the present application;
[0021] Figure 3 An exemplary diagram of a label prefix tree under a classification label of brain surgery operation shown in an embodiment of the present application;
[0022] Figure 4 A flowchart of another method of medical terminology standardization shown in an embodiment of the present application;
[0023] Figure 5 A principle diagram of medical terminology standardization by a sequence-to-sequence model combined with a label prefix tree shown in an embodiment of the present application;
[0024] Figure 6 A method flowchart of medical terminology standardization under a current classification label shown in an embodiment of the present application;
[0025] Figure 7 A principle block diagram of a device for medical terminology standardization shown in an embodiment of the present application;
[0026] Figure 8 An exemplary block diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0028] It should be noted that the terms “first”, “second”, and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0029] In order to solve the problem of inaccurate term standardization and high time cost for any medical language, the embodiment of the application designs a new medical term standardization scheme. For each standard term under each classification label in the term knowledge base, a label prefix tree is constructed. Then, for any medical language, the corresponding standard term set of the medical language can be searched from the label prefix tree constructed under at least one target classification label to which the medical language belongs according to the semantic characteristics of the medical language, so as to ensure the comprehensiveness and accuracy of medical term standardization.
[0030] Figure 1 A flowchart of a medical term standardization method according to an embodiment of the application is shown. Referring to Figure 1 The method can include the following steps:
[0031] S110, determining the semantic characteristics of any medical language.
[0032] It is considered that in the medical field, the same medical diagnosis, surgery, medicine, medical examination, test, symptom, etc. will usually have multiple different expressions. Then, when different doctors write the corresponding diagnosis and treatment medical records for each patient, the actual situation of the patient's diagnosis, surgery, test, etc. will also be recorded in the medical language in accordance with their own habits.
[0033] Therefore, in order to ensure the standardization of medical data statistics, the application can first analyze the corresponding characteristics of any medical language to extract the semantic characteristics of the medical language after obtaining the medical language, so as to accurately standardize the medical language in the subsequent term.
[0034] As an optional implementation scheme in the application, it is considered that medical term standardization is mainly to transform any medical language into a standard term in the medical field, and both the medical language and the standard term are indefinite length text sequences. Then, medical term standardization can be considered as transforming a text sequence of indefinite length into another text sequence of indefinite length. Therefore, for medical term standardization, the application can realize it by training a sequence-to-sequence (Sequence-to-Sequence, abbreviated as seq2seq) model.
[0035] Next, the network structure of the seq2seq model is described as follows:
[0036] For example, Figure 2As shown, the seq2seq model in the present application can include two parts of an encoder and a decoder. The encoder and the decoder can each be a recurrent neural network (RNN) composed of multiple layers of chained recurrent units (RNN cells), and each recurrent unit can represent a time step in the evolution direction of a text sequence. In the present application, the encoder can be used to analyze the semantic features of medical terms, the input of the decoder can be the semantic features of medical terms, and the output can be the standard term set after the medical terms are processed by term standardization.
[0037] Therefore, for the semantic features of any medical term, the present application can use a large number of medical terms as training samples, and the standard terms corresponding to each medical term as sample labels to train a sequence-to-sequence model. Then, the text representation vector of the medical term is input into the encoder of the pre-constructed sequence-to-sequence model, and the semantic features of the medical term are output.
[0038] That is, in order to ensure the accuracy of the semantic features of medical terms, the present application first converts any medical term into a text representation vector. Then, the text representation vector of the medical term is encoded and compressed by the encoder of the trained seq2seq model to a vector of a specified length, i.e., the semantic features of the medical term are obtained.
[0039] Moreover, for the text representation vector of any medical term, the present application can input the medical term into the pre-constructed language representation model to output the text representation vector of the medical term. Through the pre-constructed language representation model, such as the bidirectional encoder representations from transformer (Bert) model, natural language processing is performed on any medical term to obtain the text representation vector of the medical term. The Bert model can not be fine-tuned, but directly use the official provided parameters for direct training to obtain the text representation vector of any medical term.
[0040] In the present application, the encoder and the decoder in the seq2seq model can be a bidirectional long short-term memory (LSTM) model.
[0041] Taking the medical term "bilateral deep brain stimulation implantation" as an example, the text representation vector of the medical term obtained by processing the medical term through the Bert model can be [m1, m2, …, m n ]. Then, as shown in FIG. 2, the semantic features of the medical term can be obtained by inputting the text representation vector of the medical term into the encoder of the pre-constructed seq2seq model. Figure 2As shown, the number of the recurrent units in the encoder of the seq2seq model is the same as the dimension of the text representation vector of the medical term, so that the feature value m i may be input into the i-th recurrent unit in the encoder. Then, the feature value m n in each dimension of the text representation vector is sequentially analyzed by the respective recurrent units in the encoder to obtain the last feature value m
[0042] In addition, the hidden state vector output by the last recurrent unit in the encoder may also be transformed to obtain the semantic feature of the medical term. Alternatively, all the hidden state vectors output by the respective recurrent units in the encoder may also be uniformly transformed to obtain the semantic feature of the medical term. The present application does not limit this.
[0043] Subsequently, the semantic feature output by the encoder for the medical term may be input into the decoder to generate the character output and the hidden state of the first recurrent unit in the decoder, and the character output and the hidden state of the first recurrent unit in the decoder may be input into the second recurrent unit in the decoder to generate the character output and the hidden state of the second recurrent unit, and so on, to obtain the standardized standard terms of the medical term.
[0044] S120, according to the semantic feature, searching the standard term set corresponding to the medical term from the label prefix tree constructed under at least one target classification label to which the medical term belongs.
[0045] For each standard term in the medical field, it is usually stored in a term knowledge base, such as the ICD-10 knowledge base under the International Classification of Diseases (ICD). The organization of each standard term in the term knowledge base is usually organized by operation type or site, for example, eye surgery can be composed of codes starting with 3.
[0046] Considering that when a medical term is standardized, it is usually necessary to predict the probability of the characters in all standard terms in the term knowledge base related to the medical term, which results in some characters that often appear but have no actual meaning as the result of standardization, reducing the accuracy of medical term standardization.
[0047] Therefore, the present application can select part of the standard terms related to the medical term from the term knowledge base, and then predict the probability of the characters in the part of the standard terms related to the medical term, to improve the efficiency and accuracy of medical term standardization.
[0048] In the present application, in order to ensure the efficiency of medical terminology standardization, the present application divides each standard terminology into multiple categories according to the organization mode of each standard terminology in the terminology knowledge base, and sets a corresponding classification label for each category, such as label1, label2, …, label N Then, in order to ensure the rapid search of standard terminology under each classification label during medical terminology standardization, the present application can construct a label prefix tree for each classification label by analyzing the character sequence represented by all standard terminologies under each classification label.
[0049] In the label prefix tree constructed under each classification label, the root node can be the classification label, the first character in each standard terminology under the classification label can be a child node of the root node, and the characters represented by each child node are different. Then, the next character of the first character in each standard terminology under the classification label is a child node of the node represented by the first character, and this cycle is repeated to construct the label prefix tree under the classification label. At this time, in the label prefix tree, the string passed through from each child node of the root node to each leaf node is the standard terminology under the classification label.
[0050] Suppose the classification label of brain surgery operation is label s There are multiple standard terminologies such as brain cortex adhesion lysis, brain repair, skull decompression, intracranial nerve stimulator implantation, intracranial nerve stimulator replacement, brain deep electrode implantation, and ring clamp replacement. Then, as shown in Figure 3 , the root node of the label prefix tree constructed under the label s represented by brain surgery operation can be label s , and the child nodes of the root node can be the first four characters: big, brain, skull, and ring. The child node of the node with the character “big” can be the next character “brain” of “big” in each standard terminology, and the child nodes of each node are continuously constructed until there is no child node.
[0051] In the present application, after determining the semantic features of medical language, the semantic features of medical language can be analyzed to determine the category to which the medical language belongs, thereby determining the corresponding target classification label.
[0052] It should be understood that since a certain medical language may be standardized into a combination of multiple standard terminologies after terminology standardization, the number of standard terminologies corresponding to the medical language can also be determined by analyzing the semantic features of any medical language, thereby determining multiple target classification labels.
[0053] Then, according to the upper and lower layer relationship of each node in the label prefix tree constructed under each target classification label, the probability prediction related to the medical language can be performed on the character represented by each node of each layer in the label prefix tree under the target classification label, so that the character with the highest probability in each layer node under each target classification label is searched out, and a corresponding string is formed, that is, the corresponding standard term of the medical language under the target classification label can be searched out. According to the above manner, the corresponding target standard term of the medical language under each target classification label can be searched out in each standard term in the label prefix tree under each target classification label, and a standard term set after the medical language is standardized is formed. The standard term set can include at least one standard term, and the number of standard terms is the same as the number of target classification labels.
[0054] The technical scheme provided by the embodiments of the present application can first determine the semantic feature of any medical language. Then, according to the semantic feature, the standard term set corresponding to the medical language is searched from the label prefix tree constructed under at least one target classification label to which the medical language belongs, so as to avoid the omission of the standard term after the medical language is standardized, thereby ensuring the comprehensiveness of medical terminology standardization. Moreover, without performing the probability prediction between each standard term in the medical knowledge base and the medical language, the efficiency of medical terminology standardization is improved on the basis of ensuring the accuracy of medical terminology standardization by means of the label prefix tree constructed under the target classification label.
[0055] As an optional implementation scheme in the present application, considering the transformation between the medical language and the standard term, the seq2seq model can be used to implement the transformation. Therefore, the semantic feature output by the encoder for any medical language can be processed by each cycle unit in the decoder of the seq2seq model in sequence, so as to output the standard term set corresponding to the medical language.
[0056] In order to ensure the accuracy of medical terminology standardization, the specific process of the medical language standardization based on the seq2seq model and combined with the label prefix tree under each classification label can be described in detail.
[0057] Figure 4 The flowchart of another medical terminology standardization method shown in the embodiments of the present application is shown in FIG. 4, which can include the following steps: Figure 4
[0058] S410, determining the semantic feature of any medical language.
[0059] S420, determining the initial classification label as the current classification label according to the semantic feature of the medical language.
[0060] By analyzing the semantic features of the medical term, it can be determined that the medical term obviously belongs to a certain classification label, and the initial classification label in the present application is obtained. Considering that the medical term can correspond to multiple standard terms, there can be multiple classification labels. Therefore, the present application can take the initial classification label as the current classification label to analyze the target standard term corresponding to the medical term under the current classification label. Then, by analyzing the feature state of the medical term when it is standardized under the initial classification label, other classification labels to which the medical term belongs can be further analyzed, so that the target standard term corresponding to the medical term under each target classification label is analyzed in turn.
[0061] As an optional implementation in the present application, when the medical term is standardized by the seq2seq model, the encoder in the seq2seq model can output the semantic features of the medical term. Then, as shown in Figure 5 , the semantic features of the medical term can be input into the first recurrent unit in the encoder of the seq2seq model, and the semantic features are processed by the first recurrent unit, i.e. the initial classification label is output.
[0062] Taking the medical term "bilateral deep brain stimulation implantation" as an example, it can be known that after the term standardization, the bilateral deep brain stimulation implantation can be a combination of "deep brain electrode implantation" and "cranial nerve stimulation pulse generator implantation", and both belong to the category of brain operation.
[0063] Then, as shown in Figure 5 , the initial classification label l1 output by the first recurrent unit in the encoder of the seq2seq model can be label s , as the current classification label in the present application.
[0064] S430, according to the current classification label and the hidden state vector under the current classification label, searching for the target standard term of the medical term under the current classification label from the label prefix tree constructed under the current classification label.
[0065] When the corresponding recurrent unit in the decoder of the seq2seq model outputs the current classification label, it will also output a hidden state vector to represent the hidden state features of the medical term after processing by each previous recurrent unit. Therefore, from the label prefix trees constructed under each classification label, a certain label prefix tree taking the current classification label as the root node can be found as the label prefix tree under the current classification label.
[0066] Then, through multiple cycle units after the cycle unit in the decoder of the seq2seq model for outputting the current classification label, the probability prediction of the character represented by each node of each layer in the label prefix tree under the current classification label is performed according to the upper and lower relationships of the characters represented by each node in the label prefix tree under the current classification label and the hidden state vector under the current classification label, and the character with the highest probability in each node of each layer is searched from the label prefix tree under the current classification label to form the target standard term of the medical term under the current classification label.
[0067] As an optional implementation in the present application, for the term standardization of the medical term under the current classification label, as shown in the following formula (3), the present application can be determined by the following steps: Figure 6
[0068] S61, the child nodes of the root node in the label prefix tree under the current classification label are taken as the current nodes.
[0069] Considering that the decoder in the seq2seq model will analyze the characters at each sequence position in turn according to the multiple cycle units in the evolution direction of the standard term, and each character in each standard term is represented from the child nodes of the root node to the leaf nodes according to the node hierarchy in the label prefix tree under the current classification label.
[0070] Then, after determining the current classification label each time, the first character of the medical term standardized under the current classification label is predicted, and the first character is recorded in each child node of the root node in the label prefix tree under the current classification label. Therefore, under each current classification label, the present application can first take each child node of the root node in the label prefix tree under the current classification label as the current node, so as to perform the probability prediction of the character represented by each current node.
[0071] S62, according to the current classification label and the hidden state vector under the current classification label, the probability of the character under the current node is predicted to determine the first character searched out of the medical term under the current classification label, and the first character is taken as the current character.
[0072] In the decoder of the seq2seq model, after outputting the current classification label and the hidden state vector under the current classification label by a certain loop unit, the current classification label and the hidden state vector under the current classification label are input into the next loop unit connected after the loop unit, and the characters under each child node of the node where the current classification label is located, that is, the characters under each current node, are found as candidate characters from the label prefix tree under the current classification label by the next loop unit according to the node where the current classification label is located. Then, by the next loop unit connected after the loop unit where the current classification label is located, the probability of each candidate character is predicted according to the hidden state vector under the current classification label, so that the first character of the medical term searched under the current classification label is the candidate character with the highest probability, and the first character is taken as the current character.
[0073] Taking the medical term "bilateral deep brain stimulation implantation" as an example, after the term standardization, the bilateral deep brain stimulation implantation can be a combination of "deep brain electrode implantation" and "cranial nerve stimulation pulse generator implantation", as shown in Figure 5 If the current classification label is the initial classification label, the input of the first loop unit in the decoder of the seq2seq model is the semantic feature output by the encoder and the start character "##". <s>"Then, by analyzing the semantic features through the first loop unit, the current classification label l1 will be output as label." s The hidden state vector under the current category label is input into the second loop unit. The second loop unit can find candidate characters for each current node based on the current category label l1, predict the probability of each candidate character, and output the candidate character with the highest probability as the first character.
[0074] S63: Take the child node of the node where the current character is located in the label prefix tree as the new current node.
[0075] After determining the current character, the next loop unit in the decoder of the seq2seq model will continue to analyze the next character. The next character will be located in the child nodes of the node where the current character is located in the label prefix tree under the current category label. Therefore, this application can take the child nodes of the node where the current character is located in the label prefix tree as the new current node.
[0076] S64. Based on the character search trajectory of the current character in the label prefix tree and the hidden state vector under the current character, predict the character probability under the new current node, determine the next character of the current character, and take the next character as the new current character.
[0077] In the decoder of the seq2seq model, after the current character is output by a certain recurrent unit, its hidden state vector is also output. The current character and its hidden state vector are then input into the next recurrent unit. The next recurrent unit, based on the characters output by the previous recurrent units, determines the character search trajectory of the current character within the label prefix. Then, following this search trajectory, it searches the label prefix tree for each character under the new current node, using them as candidate characters for the next character. Furthermore, the next recurrent unit, based on the hidden state vector of the current character, performs probability predictions for each candidate character related to the medical term, selecting the candidate character with the highest probability as the next character. This next character is then used as the new current character to continue predicting the next character, and this process is repeated to obtain the character sequence of the medical term searched under the current category label.
[0078] Taking the medical term "bilateral deep brain stimulation implantation" as an example, it can be seen that after standardization of terminology, "bilateral deep brain stimulation implantation" can be a combination of "deep brain electrode implantation" and "cranial nerve stimulation pulse generator implantation", both of which belong to the category of brain surgical procedures.
[0079] like Figure 5 As shown, assume that the current classification label is the initial classification label l1, which is the label of the brain surgery operation category s , then the first recurrent unit in the decoder of the seq2seq model will output the current classification label l1 and the first hidden state vector.
[0080] Assume that the current character is "深" in "脑深部电极植入术", denoted as to illustrate the determination process of the next character of the current character:
[0081] As Figure 5 shown, the first three recurrent units in the decoder of the seq2seq model will successively output the current classification label l1 represented by the initial classification label, the character "脑", and the character "深". At this time, the third recurrent unit will output the current character "深" and the third hidden state vector, and input them into the fourth recurrent unit. Then, through the fourth recurrent unit, it can be determined that the character search trajectory of the current character "深" in the label prefix tree is [label s , 脑, 深]. According to this character search trajectory [label s , 脑, 深], it can be determined that the new current nodes represented by each child node of the node where the current character "深" is located in the label prefix tree, and it is determined that the character "部" under the new current node is the candidate character. Then, the fourth can predict the probability of the candidate character "部" as 0.34 by analyzing the third hidden state vector, and the probabilities of other characters are 0. Therefore, the candidate character "部" can be used as the next character of the current character "深", so that the fourth recurrent unit can output the character "部" and the fourth hidden state vector. Furthermore, taking the character "部" as the new current character, following the above steps, the fifth recurrent unit continues to output the next character of the character "部", and so on in a loop until the target standard term "脑深部电极植入术" under the current classification label l1 is completely searched.
[0082] S65. Determine whether the next character is the next classification label. If not, continue to return to execute S63; if so, execute S66.
[0083] Since the decoder of the seq2seq model processes the standard terms corresponding to each target classification label to which the medical term belongs in sequence, when a certain recurrent unit in the decoder of the seq2seq model outputs the last character in the target standard term corresponding to the current classification label, the next recurrent unit will output the next classification label of the current classification label according to the last character input by the previous recurrent unit and the hidden state vector under the last character. And if the current classification label is the last classification label, then the next classification label output by the next recurrent unit is the blank character "< / s> It can be seen that the next classification label can include normal label characters and blank characters "", which are used to analyze the actual characters of the next classification label to determine whether the term standardization of the medical term under all target classification labels is completed.
[0084] Therefore, in order to accurately determine whether the decoder of the seq2seq model completes the term standardization of the medical term under the current classification label, when a next character is output by a certain loop unit in the decoder of the seq2seq model, it is first determined whether the next character is the next classification label. If the next character is the next classification label, it means that the decoder of the seq2seq model has completed the term standardization of the medical term under the current classification label, and the next classification label needs to be taken as the new current classification label to continue the corresponding term standardization of the medical term under the next classification label.
[0085] If the next character is not the next classification label, it means that the decoder of the seq2seq model has not completed the term standardization of the medical term under the current classification label, so the next character can be taken as the new current character, and it can continue to return to S63 to determine the new next character until the next character is the next classification label.
[0086] S66, the character sequence searched by the medical term under the current classification label is taken as the corresponding target standard term.
[0087] When the next character is the next classification label or a blank character, it indicates that the decoder of the seq2seq model has completed the term standardization of the medical term under the current classification label, and then the characters output from the multiple cycle units after a certain cycle unit for outputting the current classification label and before another cycle unit for outputting the next classification label or the blank character can be obtained, and the characters are combined to form a character sequence searched by the medical term under the current classification label. Then, the character sequence is taken as the target standard term of the medical term under the current classification label.
[0088] S440, predicting the next classification label as a new current classification label according to the end character in the target standard term and the hidden state vector under the end character.
[0089] After completing the term standardization of the medical term under the current classification label, the target standard term of the medical term under the current classification label can be obtained. At this time, the end character in the target standard term and the hidden state vector under the end character can be output by a certain cycle unit in the decoder of the seq2seq model and input into the next cycle unit. Then, the next cycle unit analyzes the hidden state vector under the end character and the end character, and outputs the next classification label of the current classification label, so that the next classification label is taken as a new current classification label to continue the corresponding term standardization of the medical term under the next classification label.
[0090] S450, judging whether the next classification label is a blank character, if not, returning to execute S430; if yes, executing S460.
[0091] In order to accurately judge whether the term standardization of the medical term under all target classification labels has been completed, the present application further judges whether the next classification label is a blank character after predicting the next classification label. If the next classification label is a blank character, it indicates that the decoder of the seq2seq model has completed the term standardization of the medical term under all target classification labels, and the target standard term of the medical term under each target classification label can be obtained.
[0092] If the next classification label is not a blank character but a normal label character, it indicates that the decoder of the seq2seq model has not completed the term standardization of the medical term under all target classification labels, so the next classification label is taken as a new current classification label, and the process returns to S430 to continue the target standard term of the medical term under the next current classification label. The cycle continues until the next classification label is a blank character.
[0093] S460, determining the standard term set corresponding to the medical term.
[0094] When the next classification label is a blank character, the target standard term of the medical term under each target classification label can be obtained through each cycle unit in the decoder of the seq2seq model. Then, the standard term set corresponding to the medical term can be obtained by combining the target standard terms of the medical term under each target classification label.
[0095] The technical scheme provided by the embodiments of the present application can first determine the semantic feature of any medical term. Then, according to the semantic feature, the standard term set corresponding to the medical term is searched from the label prefix tree constructed under at least one target classification label to which the medical term belongs, so as to avoid the omission of the standard term after the medical term standardization, thereby ensuring the comprehensiveness of the medical term standardization. Moreover, without performing the probability prediction between each standard term in the medical knowledge base and the medical term, the efficiency of the medical term standardization is improved on the basis of ensuring the accuracy of the medical term standardization by means of the label prefix tree constructed under the target classification label.
[0096] Figure 7 A principle block diagram of a device for medical term standardization is shown in the embodiments of the present application. As shown in the figure, the device 700 can include: Figure 7
[0097] The feature determination module 710 is configured to determine the semantic feature of any medical term.
[0098] The medical term standardization module 720 is configured to search the standard term set corresponding to the medical term from the label prefix tree constructed under at least one target classification label to which the medical term belongs according to the semantic feature.
[0099] The label prefix tree is composed of a plurality of standard terms under the target classification label.
[0100] In some implementable manners, the medical term standardization module 720 can be specifically configured to:
[0101] The initial classification label determination unit is configured to determine an initial classification label as a current classification label according to the semantic feature of the medical term.
[0102] The standard term searching unit is configured to search the target standard term of the medical term under the current classification label from the label prefix tree constructed under the current classification label according to the current classification label and the hidden state vector under the current classification label.
[0103] The standard term set determination unit is configured to predict a next classification label as a new current classification label according to a last character in the target standard term and a hidden state vector under the last character, and continue to return to the standard term searching unit to perform the standard term searching step until the next classification label is a blank character, so as to obtain the standard term set corresponding to the medical language.
[0104] In some implementations, the standard term searching unit can be specifically configured to:
[0105] Take a child node of a root node in a label prefix tree under the current classification label as a current node;
[0106] Predict a character probability under the current node according to the current classification label and a hidden state vector under the current classification label, so as to determine a first character searched out by the medical language under the current classification label, and take the first character as a current character;
[0107] Perform a character prediction step: take a child node of a node where the current character is located in the label prefix tree as a new current node;
[0108] Predict a character probability under the new current node according to a character search track of the current character in the label prefix tree and a hidden state vector under the current character, so as to determine a next character of the current character;
[0109] Take the next character as a new current character, and continue to return to perform the above-mentioned character prediction step until the next character is a next classification label, so as to take a character sequence searched out by the medical language under the current classification label as a corresponding target standard term.
[0110] In some implementations, the feature determination module 710 can be specifically configured to:
[0111] Input a text representation vector of the medical language into an encoder in a pre-constructed sequence-to-sequence model, and output a semantic feature of the medical language.
[0112] In some implementations, the medical term standardization apparatus 700 can further include:
[0113] A text representation module configured to input the medical language into a pre-constructed language representation model, and output a text representation vector of the medical language.
[0114] In some implementations, the sequence-to-sequence model includes an encoder and a decoder, an input of the decoder is the semantic feature, and an output of the decoder is a standard term set corresponding to the medical language.
[0115] In the embodiments of the present application, for any medical term, the semantic feature of the medical term is determined first. Then, according to the semantic feature, the standard terminology set corresponding to the medical term is searched from the label prefix tree constructed under at least one target classification label to which the medical term belongs, so as to avoid omission of the standard terminology after standardization of the medical term, thereby ensuring comprehensiveness of medical terminology standardization. Moreover, without performing probability prediction between each standard terminology in the medical knowledge base and the medical terminology, on the basis of ensuring accuracy of medical terminology standardization, the label prefix tree constructed under the target classification label is used to improve efficiency of medical terminology standardization.
[0116] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, the details are not described here. Specifically, Figure 7 The device 700 shown can perform any of the method embodiments of the present application, and the foregoing and other operations and / or functions of each module in the device 700 are respectively for realizing the corresponding processes in each method in the embodiments of the present application. For the sake of brevity, the details are not described here.
[0117] The device 700 of the embodiments of the present application is described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be realized by hardware, or by instructions in the form of software, or by a combination of hardware and software modules. Specifically, each step of the method embodiments in the embodiments of the present application can be completed by integrated logic circuits of hardware in the processor and / or instructions in the form of software, and the steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware code processor execution completion, or executed by a combination of hardware and software modules in the code processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiments.
[0118] Figure 8 The electronic device shown in the embodiments of the present application is a schematic block diagram.
[0119] As shown in Figure 8 The electronic device 800 can include:
[0120] The memory 810 is used to store computer programs and transmit the program codes to the processor 820. In other words, the processor 820 can call and run the computer programs from the memory 810 to realize the method in the embodiments of the present application.
[0121] For example, the processor 820 can be configured to perform the above-described method embodiments according to instructions in the computer program.
[0122] In some embodiments of the present application, the processor 820 can include but is not limited to:
[0123] A general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc.
[0124] In some embodiments of the present application, the memory 810 includes but is not limited to:
[0125] A volatile memory and / or a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a synch link DRAM (SLDRAM), and a Direct Rambus RAM (DR RAM).
[0126] In some embodiments of the present application, the computer program can be divided into one or more modules, which are stored in the memory 810 and executed by the processor 820 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the electronic device.
[0127] As shown in Figure 8 The electronic device can further include:
[0128] The transceiver 830 can be connected to the processor 820 or the memory 810.
[0129] The processor 820 can control the transceiver 830 to communicate with other devices, specifically, can send information or data to other devices, or receive information or data sent by other devices. The transceiver 830 can include a transmitter and a receiver. The transceiver 830 can further include an antenna, and the number of antennas can be one or more.
[0130] It should be understood that various components in the electronic device are connected through a bus system, wherein the bus system includes a data bus, a power supply bus, a control bus and a state signal bus in addition to a data bus.
[0131] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to execute the method of the above method embodiment. Alternatively, the present application embodiment also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to execute the method of the above method embodiment.
[0132] When implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage media can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired computer program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or a twisted pair, as examples, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0133] In one embodiment, the techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the software can be executed in a computer system, which can include one or more computers. The software can be stored on one or more computer readable media, such as a magnetic disk, optical disk, or solid state memory. The computer readable media can be distributed among one or more computer systems.
[0134] In several embodiments provided in the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the above-described device embodiments are merely illustrative, and the division of the modules is merely a logical function division. In actual implementation, another division manner can be used, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be indirect coupling or communication connection through some interfaces, devices, or modules, and can be electrical, mechanical, or other forms.
[0135] The modules illustrated as separate components may or may not be physically separate, and the components illustrated as modules may or may not be physical modules, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules in various embodiments of the present application can be integrated in one processing module, or each module can exist physically separately, or two or more modules can be integrated in one module.
[0136] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of medical terminology standardization, characterized by, The method comprises: determining the semantic feature of any medical term; searching the standard terminology set corresponding to the medical term from the label prefix tree constructed under at least one target classification label to which the medical term belongs according to the semantic feature; wherein the label prefix tree is composed of a plurality of standard terminologies under the target classification label; wherein the searching the standard terminology set corresponding to the medical term from the label prefix tree constructed under at least one target classification label to which the medical term belongs according to the semantic feature comprises: determining an initial classification label as a current classification label according to the semantic feature of the medical term; performing a standard terminology search step: searching the target standard terminology of the medical term under the current classification label from the label prefix tree constructed under the current classification label according to the current classification label and the hidden state vector under the current classification label; predicting a next classification label as a new current classification label according to the last character in the target standard terminology and the hidden state vector under the last character, and continuing to perform the above standard terminology search step until the next classification label is a blank character, thereby obtaining the standard terminology set corresponding to the medical term.
2. The method of claim 1, wherein, The searching the target standard terminology of the medical term under the current classification label from the label prefix tree constructed under the current classification label according to the current classification label and the hidden state vector under the current classification label comprises: taking the child node of the root node in the label prefix tree under the current classification label as a current node; predicting the character probability under the current node according to the current classification label and the hidden state vector under the current classification label, so as to determine the first character searched out by the medical term under the current classification label, and taking the first character as a current character; performing a character prediction step: taking the child node of the node where the current character is located in the label prefix tree as a new current node; predicting the character probability under the new current node according to the character search track of the current character in the label prefix tree and the hidden state vector under the current character, so as to determine the next character of the current character; taking the next character as a new current character, and continuing to perform the above character prediction step until the next character is a next classification label, thereby taking the character sequence searched out by the medical term under the current classification label as a corresponding target standard terminology.
3. The method of claim 1, wherein, The determining the semantic feature of any medical term comprises: inputting the text representation vector of the medical term into the encoder of the pre-constructed sequence-to-sequence model, and outputting the semantic feature of the medical term.
4. The method of claim 3, wherein, The method further comprises: inputting the medical term into the pre-constructed language representation model, and outputting the text representation vector of the medical term.
5. The method of claim 3, wherein, The sequence-to-sequence model comprises an encoder and a decoder, and the input of the decoder is the semantic feature and the output is the standard terminology set corresponding to the medical term.
6. An apparatus for medical terminology standardization, characterized by, The method comprises: a feature determination module configured to determine the semantic feature of any medical term; The medical term standardization module is configured to search for a standard term set corresponding to the medical term from a label prefix tree constructed under at least one target classification label to which the medical term belongs according to the semantic features of the medical term. The label prefix tree is composed of a plurality of standard terms under the target classification label. The medical term standardization module is specifically configured to: determine an initial classification label as a current classification label according to the semantic features of the medical term; perform a standard term search step of searching for a target standard term of the medical term under the current classification label from a label prefix tree constructed under the current classification label according to the current classification label and a hidden state vector under a last character in the target standard term; predict a next classification label as a new current classification label according to the last character in the target standard term and a hidden state vector under the last character, and continue to perform the standard term search step until the next classification label is a blank character, and obtain a standard term set corresponding to the medical term.
7. An electronic device, comprising: The medical term standardization method comprises: a processor and a memory, the memory being configured to store a computer program, and the processor being configured to invoke and run the computer program stored in the memory to execute the medical term standardization method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, A computer program for storing a computer program, which enables a computer to execute the medical term standardization method according to any one of claims 1-5.
9. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the medical term standardization method according to any one of claims 1-5.
Citation Information
Patent Citations
Symptom question and answer method, device and equipment based on knowledge graph and storage medium
CN115186068A