Method, apparatus, and computer device for standardizing model-based clinical terms
Through the model-based clinical term standardization method, multi-dimensional feature processing, deep convolutional neural network, attention model and twin network are used to solve the problem of low accuracy in clinical term standardization in the prior art, achieving higher accuracy and accuracy.
Patent Information
- Application Number
- CN202111013436.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-08-31
AI Technical Summary
The existing standardization methods for clinical terms cannot obtain accurate standard terms corresponding to clinical terms, and there is a problem of low accuracy.
Using a model-based clinical term standardization method, text features are generated by obtaining the clinical terms to be processed, multi-dimensional feature processing is performed to generate the text features, Embedding vectors are processed using deep convolutional neural networks and attention models, candidate sets are generated based on medical term databases, and similarity is calculated through twin networks to determine standard clinical terms.
It effectively improves the accuracy of standard clinical terms and standardization accuracy, and solves the problem of low accuracy of clinical terms labeling.
Smart Images

Figure CN113657109B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device and computer device for standardizing clinical terms based on a model. Background Art
[0002] Clinical terms refer to the professional terms used in the medical field to describe diagnoses, surgeries, drugs, examinations, tests, and symptoms. Clinical terms are essential components for clinical information systems to express medical information. However, the expressions of many clinical terms are usually not standardized, so it is necessary to standardize clinical terms for subsequent data applications. Existing clinical term standardization mainly uses traditional machine learning methods for processing. However, simply using traditional machine learning methods often fails to obtain accurate standard terms corresponding to clinical terms, resulting in a problem of low accuracy in clinical term annotation. Therefore, how to improve the accuracy of clinical term annotation has become an urgent problem to be solved currently. Summary of the Invention
[0003] The main purpose of this application is to provide a method, device, computer device and storage medium for standardizing clinical terms based on a model, aiming to solve the technical problem that existing clinical term standardization methods often fail to obtain accurate standard terms corresponding to clinical terms and have low accuracy.
[0004] This application proposes a method for standardizing clinical terms based on a model, and the method includes the steps of:
[0005] Obtain the clinical terms to be processed;
[0006] Perform multi-dimensional feature processing on the clinical terms to generate multiple text features corresponding to the clinical terms, and convert each of the text features into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features and surgical approach features, the surgical site feature refers to the feature corresponding to any open site in the skin or viscera, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations;
[0007] Process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector;
[0008] Generate a candidate set corresponding to the clinical terms based on a preset medical term library; wherein, the candidate set includes multiple candidate texts;
[0009] Input each of the candidate texts into a preset Transformer encoder network to generate multiple candidate text Embedding vectors respectively corresponding to the candidate texts;
[0010] Based on a preset siamese network, calculate the similarity between the specified Embedding vector and each of the candidate text Embedding vectors;
[0011] Based on all the similarities, determine the standard clinical term corresponding to the clinical term from all the candidate set texts.
[0012] Optionally, the step of performing multi-dimensional feature processing on the clinical term to generate multiple text features corresponding to the clinical term includes:
[0013] Perform word segmentation processing on the clinical term based on a preset word segmentation tool to obtain corresponding character features, word features, and part-of-speech features;
[0014] Obtain the pinyin feature corresponding to the clinical term based on a preset character-to-pinyin model;
[0015] Obtain the surgical site feature and surgical approach feature corresponding to the clinical term based on a preset medical thesaurus;
[0016] Use the character features, word features, part-of-speech features, pinyin features, surgical site features, and surgical approach features as the text features.
[0017] Optionally, the step of processing all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector includes:
[0018] Invoke the deep convolutional neural network and the attention model;
[0019] Input all the first Embedding vectors into the deep convolutional neural network, and obtain a corresponding plurality of second Embedding vectors obtained by the deep convolutional neural network after extracting feature information from each of the first Embedding vectors; wherein, input each of the first Embedding vectors into the convolutional layer of the deep convolutional neural network respectively, perform convolutional processing on each of the first Embedding vectors through the convolutional layer to obtain abstract feature Embedding vectors corresponding to each of the first Embedding vectors respectively, perform pooling processing on each of the abstract feature Embedding vectors in the convolutional layer through the pooling layer of the deep convolutional neural network to obtain processed abstract feature Embedding vectors corresponding to each of the abstract feature Embedding vectors respectively, and use the processed abstract feature Embedding vectors as the second Embedding vectors;
[0020] Input all the second Embedding vectors into the attention model, and output a multi-dimensional vector corresponding to the second Embedding vectors through the attention model;
[0021] Use the multi-dimensional vector as the specified Embedding vector.
[0022] Optionally, the step of generating a candidate set corresponding to the clinical term based on a preset medical term library includes:
[0023] Perform word segmentation on the clinical term using a preset word segmentation engine to obtain initial keywords;
[0024] Perform part-of-speech tagging on the initial keywords using an N-gram model to obtain the first keywords after tagging;
[0025] Delete the keywords with the part of speech of function words from the first keywords to obtain corresponding second keywords;
[0026] Call the medical term library;
[0027] Respectively obtain the number of occurrences of each of the second keywords in the medical term library, and screen out the third keywords whose number of occurrences is greater than a preset occurrence threshold from all the second keywords;
[0028] Obtain the first medical terms corresponding to the third keywords from the medical term library;
[0029] Judge whether there are the same medical terms among all the first medical terms;
[0030] If so, perform deduplication on the first medical term to obtain the corresponding second medical term;
[0031] Use all the second medical terms as the candidate set.
[0032] Optionally, the siamese network includes a first network and a second network. The step of calculating the similarity between the specified Embedding vector and each of the candidate text Embedding vectors based on the preset siamese network includes:
[0033] Input the specified Embedding vector into the first network to obtain the first vector data corresponding to the specified Embedding vector output by the first network;
[0034] Input the specified candidate text Embedding vector into the second network to obtain the second vector data corresponding to the specified candidate text Embedding vector output by the second network; where the specified candidate text Embedding is any one of all the candidate text Embeddings;
[0035] Calculate the cosine value of the vectors between the first vector data and the second vector data through the cosine similarity formula;
[0036] Use the cosine value of the vectors as the similarity between the first vector data and the second vector data.
[0037] Optionally, the step of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities includes:
[0038] Screen out the first similarity with the largest value from all the similarities;
[0039] Judge whether the number of the first similarities is equal to 1;
[0040] If so, obtain the first candidate text Embedding vector corresponding to the first similarity;
[0041] Obtain the first candidate text corresponding to the first candidate text Embedding vector from the candidate set;
[0042] Use the first candidate text as the standard clinical term.
[0043] Optionally, the step of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities includes:
[0044] Select the second similarity with the largest value from all the similarities;
[0045] Determine whether there is a third similarity; wherein, the difference between the second similarity and the third similarity is within a preset numerical range;
[0046] If there is the third similarity, obtain the second candidate text Embedding vectors corresponding one-to-one to each of the second similarities and the third similarity;
[0047] Obtain the second candidate texts corresponding one-to-one to each of the second candidate text Embedding vectors from the candidate set;
[0048] Receive the third candidate text selected by a preset user from all the second candidate texts;
[0049] Take the third candidate text as the standard clinical term.
[0050] This application also provides a standardization device for clinical terms based on a model, including:
[0051] An acquisition module, configured to acquire clinical terms to be processed;
[0052] A first generation module, configured to perform multi-dimensional feature processing on the clinical terms, generate multiple text features corresponding to the clinical terms, and respectively convert each of the text features into corresponding first Embedding vectors; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features and surgical approach features, the surgical site feature refers to the feature corresponding to any open site in the skin or viscera, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations;
[0053] A processing module, configured to process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector;
[0054] A second generation module, configured to generate a candidate set corresponding to the clinical terms based on a preset medical term library; wherein, the candidate set includes multiple candidate texts;
[0055] A third generation module, configured to input each of the candidate texts into a preset transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each of the candidate texts respectively;
[0056] A calculation module, configured to calculate the similarity between the specified Embedding vector and each of the candidate text Embedding vectors respectively based on a preset twin network;
[0057] A determination module, configured to determine a standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities.
[0058] This application also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above method are implemented.
[0059] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0060] The method, device, computer device, and storage medium for standardizing clinical terms based on a model provided in this application have the following beneficial effects:
[0061] The standardized method, device, computer device, and storage medium for clinical terms based on a model provided in this application, after obtaining the clinical terms to be processed, will first perform multi-dimensional feature processing and feature transformation processing on the clinical terms to obtain corresponding first Embedding vectors, and then process all the first Embedding vectors based on a preset model to obtain specified Embedding vectors. Then, a candidate set corresponding to the clinical terms is generated based on a preset medical term library, and each candidate text in the candidate set is input into a preset transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each candidate text. Subsequently, the similarity between the specified Embedding vector and each candidate text Embedding vector is calculated based on a preset siamese network. Finally, based on all the similarities, the standard clinical term corresponding to the clinical terms is determined from all the candidate set texts. After obtaining the clinical terms to be processed, this application generates the specified Embedding vector corresponding to the clinical terms through multi-dimensional feature processing and the processing of a preset deep convolutional neural network and attention model, and generates the candidate text Embedding vectors of the candidate set corresponding to the clinical terms through a preset medical term library and a transformer encoder network. Then, the similarity between the Embedding vector and the candidate text Embedding vector is calculated and analyzed based on a preset siamese network to screen out the final standard clinical term corresponding to the clinical terms, which can effectively improve the accuracy of the obtained standard clinical terms and the accuracy of standardizing clinical terms. Description of the Drawings
[0062] Figure 1 is a schematic flowchart of a method for standardizing clinical terms based on a model according to an embodiment of the present application;
[0063] Figure 2 is a schematic structural diagram of a device for standardizing clinical terms based on a model according to an embodiment of the present application;
[0064] Figure 3 is a schematic structural diagram of a computer device according to an embodiment of the present application.
[0065] The implementation, functional features, and advantages of the objectives of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0066] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0067] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.
[0068] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0069] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0070] Refer to Figure 1 , a method for standardizing clinical terms based on a model according to an embodiment of the present application includes:
[0071] S1: Acquire the clinical terms to be processed;
[0072] S2: Perform multi-dimensional feature processing on the clinical terms to generate multiple text features corresponding to the clinical terms, and respectively convert each text feature into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features, and surgical approach features, the surgical site feature refers to the feature corresponding to any open site in the skin or internal organs, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations;
[0073] S3: Process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector;
[0074] S4: Generate a candidate set corresponding to the clinical terms based on a preset medical term library; wherein, the candidate set includes multiple candidate texts;
[0075] S5: Input each of the candidate texts into a preset Transformer Encoder network to generate multiple candidate text Embedding vectors respectively corresponding to the candidate texts;
[0076] S6: Calculate the similarity between the specified Embedding vector and each of the candidate text Embedding vectors respectively based on a preset Siamese network;
[0077] S7: Determine the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities.
[0078] As described in the above steps S1 to S7, the execution subject of the method embodiment is a model-based clinical term standardization device. In practical applications, the above model-based clinical term standardization device can be implemented by a virtual device, such as software code, or by an entity device written or integrated with relevant execution code, and can perform human-computer interaction with users through means such as a keyboard, a mouse, a remote control, a touchpad or a voice control device. The model-based clinical term standardization device in this embodiment can effectively improve the accuracy of the obtained standard clinical terms, thereby improving the accuracy of standardizing clinical terms. Specifically, first, obtain the clinical terms to be processed. Among them, clinical terms refer to professional terms used in the medical field to describe diagnoses, surgeries, drugs, examinations, tests, symptoms. These terms are necessary components for clinical information systems to express medical information. However, when clinical terms are not processed by data standardization, they usually contain some non-standard data, such as medical aliases, synonyms, etc. If a unified standard cannot be achieved, these non-standardized clinical terms are difficult to be applied in subsequent medical applications, thus causing waste of data. Therefore, it is currently necessary to standardize clinical terms into corresponding standard terms to improve the usability of data.
[0079] Then, multi-dimensional feature processing is performed on the clinical terms to generate multiple text features corresponding to the clinical terms, and each of the text features is respectively converted into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features, and surgical approach features. The surgical site feature refers to the feature corresponding to any open site in the skin or viscera, or can also refer to the feature corresponding to the site generated by any opening in the skin or viscera for a specific medical purpose. The "open" site refers to the surgical part where medical staff directly physically contact the area of interest. The surgical site may include, but is not limited to, organs, muscles, ligaments, connective tissues, etc. The surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations. Additionally, multi-dimensional feature processing can be performed on the clinical terms based on a preset word segmentation tool, character-to-pinyin model, and medical thesaurus, thereby generating multiple text features corresponding to the clinical terms. Additionally, each of the text features can be converted into a corresponding first Embedding vector by using Multi-view encoders. Multi-view encoders are the multi-view attention network in the multi-view attention based denoising auto-encoder (MADAE) model. Based on the MADAE model, a normalization process for text features can be performed. The DAE itself is designed to learn more robust features. By introducing random noise in the Embedding layer of the network, it can be said that the input data is destroyed and the original data is reconstructed. Therefore, the features trained by it will be more robust. Multi-view encoders perform Embedding based on the original text after adding noise. According to its six vector representation results of characters, words, sounds, part-of-speech, surgical sites, and surgical approaches, an attention layer is passed through and then a decoder is used to return to the original input data. In addition, after obtaining the first Embedding vector, the length of the first Embedding vector can be further made to match the length of the character vector.
[0080] After obtaining the first Embedding vectors, all the first Embedding vectors are processed based on a preset deep convolutional neural network and a preset attention model to obtain specified Embedding vectors. Among them, the deep convolutional neural network can also be called the DCNN model, and the attention model can also be called the Self-attention model. In addition, the role of the DCNN model is to enable each feature to consider the information of other features; that is, Self-attention only considers the relationship between each type of feature and does not consider the information between features. DCNN can make up for this weakness, making each type of feature related to each other and obtaining the information between each type of feature. For example, corresponding to the Embedding of word features, after inputting the Embedding of word features into the deep convolutional neural network, the output is each word, an Embedding vector containing the information of other features. DCNN is dilated convolution. Compared with ordinary convolution, a dilation factor is added. For example, if we set the dilation factor to 2, then the receptive field is twice that of ordinary convolution, so that 6 multi-views can directly establish context associations and improve the model effect. For example: It's very hot today, and I want to buy ice cream. There is a strong context association between "hot" and "ice cream" here, and this association can be directly constructed by DCNN. The role of the Self-attention model is to enable each feature to consider its context information. For example, corresponding to the Embedding of word features, its context information can be obtained through the Self-attention model. For example, when inputting "left upper lobe resection", the output is each word, an Embedding containing context information. Specifically, all the first Embedding vectors can be first input into the deep convolutional neural network to obtain a corresponding number of second Embedding vectors obtained by the deep convolutional neural network extracting feature information from each of the first Embedding vectors. And all the second Embedding vectors are input into the attention model, and a multi-dimensional vector corresponding to the second Embedding vectors is output through the attention model. Furthermore, the multi-dimensional vector is used as the specified Embedding vector.
[0081] After that, a candidate set corresponding to the clinical term is generated based on a preset medical term library; wherein, the candidate set includes a plurality of candidate texts. Among them, after obtaining the clinical term, a part of the medical terms that meet the subsequent comparison requirements can be screened out from the medical term library based on the inverted index processing means as candidate texts. The specific implementation process will be further described in the subsequent specific embodiments and will not be elaborated here. After obtaining the candidate texts, each of the candidate texts is input into a preset Transformer encoder network to generate a plurality of candidate text Embedding vectors corresponding to each of the candidate texts respectively. Among them, the Transformer encoder network is the encoder network structure of the existing Transformer model (machine translation model). In addition, all candidate texts can be input into the Transformer encoder network to generate corresponding candidate text Embedding vectors by processing each candidate text through the Transformer encoder network to generate a specified Embedding vector based on the clinical term.
[0082] Subsequently, the similarity between the specified Embedding vector and each candidate text Embedding vector is calculated based on a preset siamese network. Among them, after inputting the specified Embedding vector and each candidate text Embedding vector into the siamese network, first vector data and corresponding multiple second vector data can be obtained correspondingly, and then the cosine similarity formula can be called to calculate the vector cosine value between the first vector data and the second vector data to quickly calculate the similarity between the specified Embedding vector and each candidate text Embedding vector. Finally, based on all the similarities, the standard clinical term corresponding to the clinical term is determined from all the candidate set texts. Among them, the determination process of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities is not limited. The candidate text corresponding to the first similarity with the largest value can be used as the standard clinical term corresponding to the clinical term after screening all the similarities. Or when there are other similarities with values close to the similarity with the largest similarity value, the standard clinical term corresponding to the clinical term is determined from the multiple candidate texts corresponding to the similarity with the largest value based on the selection of a preset user, such as an expert.
[0083] After obtaining the clinical terms to be processed in this embodiment, the clinical terms will first be subjected to multi-dimensional feature processing and feature transformation processing to obtain corresponding first Embedding vectors, and then all the first Embedding vectors will be processed based on a preset model to obtain specified Embedding vectors. After that, a candidate set corresponding to the clinical terms will be generated based on a preset medical term library, and each candidate text in the candidate set will be input into a preset Transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each candidate text. Subsequently, the similarity between the specified Embedding vector and each candidate text Embedding vector will be calculated respectively based on a preset siamese network. Finally, based on all the similarities, the standard clinical term corresponding to the clinical terms will be determined from all the candidate set texts. After obtaining the clinical terms to be processed in this embodiment, the specified Embedding vector corresponding to the clinical terms is generated through multi-dimensional feature processing and the processing of a preset deep convolutional neural network and attention model, and the candidate text Embedding vectors corresponding to the clinical terms are generated through a preset medical term library and Transformer encoder network. Then, the similarity between the Embedding vector and the candidate text Embedding vector is calculated and analyzed based on a preset siamese network to screen out the final standard clinical term corresponding to the clinical terms, which can effectively improve the accuracy of the obtained standard clinical terms and the accuracy of standardizing clinical terms.
[0084] Further, in an embodiment of the present application, the multi-dimensional feature processing of the clinical terms in step S2 to generate multiple text features corresponding to the clinical terms includes:
[0085] S200: Perform word segmentation processing on the clinical terms based on a preset word segmentation tool to obtain corresponding character features, word features, and part-of-speech features;
[0086] S201: Obtain the pinyin feature corresponding to the clinical terms based on a preset character-to-pinyin model;
[0087] S202: Obtain the surgical site feature and surgical approach feature corresponding to the clinical terms based on a preset medical word library;
[0088] S203: Use the character features, word features, part-of-speech features, pinyin features, surgical site features, and surgical approach features as the text features.
[0089] As described in the above steps S200 to S203, the step of performing multi-dimensional feature processing on the clinical term to generate multiple text features corresponding to the clinical term may specifically include: First, perform word segmentation processing on the clinical term based on a preset word segmentation tool to obtain corresponding character features, word features, and part-of-speech features. Among them, the word segmentation tool can be trained and generated using the HMM model (Hidden Markov Model), and the training and generation process of the word segmentation tool is implemented by existing technologies, which will not be elaborated here. For example, if the clinical term to be processed is "left upper lobectomy", after inputting the clinical term "left upper lobectomy" into the word segmentation tool, three text feature results corresponding to the characters, words, and medical part-of-speech of the clinical term can be output, which are: characters [left, upper, lung, lobe, cut, remove]; words [left upper, lung lobe, resection]; part-of-speech [nfw, njp, nss]. In addition, a part-of-speech model can also be pre-trained to extract part-of-speech features from the text. The Google word2vector method or other common word vector methods can be used to train it on a large-scale Chinese corpus, which will not be elaborated here. Then, obtain the pinyin feature corresponding to the clinical term based on a preset Chinese character to pinyin model. Among them, the Chinese character to pinyin model can be pre-trained. The Chinese character to pinyin model can be regarded as a mapping library between Chinese characters and pinyin, and has the function of matching Chinese characters and outputting pinyin. For example: after inputting the clinical term "left upper lobectomy" into the Chinese character to pinyin model, the output is: zuǒshàngfèiyèqiēchú. In addition, the Chinese character to pinyin model can be trained using the Google word2vector method or other common word vector methods on a large-scale Chinese corpus, which will not be elaborated here. After that, obtain the surgical site feature and surgical approach feature corresponding to the clinical term based on a preset medical term library. Among them, the medical term library is a medical library manually constructed by a medical team according to documents such as drug instructions, equipment instructions, ICD9 / ICD10, and public medical literature. The medical term library stores at least surgical terms, as well as surgical site descriptions and surgical approach descriptions related to the surgical terms. The application of the medical term library can be regarded as a query operation. According to the words obtained after word segmentation of the clinical term, the corresponding surgical site and surgical approach are found from the medical term library. For example: input "dorsoulnar incision of forearm" into the medical term library, and the output is: surgical site [ulna], surgical approach [forearm]. In addition, the order of the steps for obtaining each text feature is not limited to the order shown in this embodiment, and the acquisition and processing of each text feature can also be performed simultaneously. Finally, the character features, the word features, the part-of-speech features, the pinyin features, the surgical site features, and the surgical approach features are used as the text features.In this embodiment, by using a word segmentation tool, a character-to-pinyin model, and a medical thesaurus, the character features, word features, medical part-of-speech features, pinyin features, surgical site features, and surgical approach features corresponding to the clinical term can be quickly obtained and used as the text features, so that subsequently, by performing corresponding processing on the Embedding vectors corresponding to the text features, a specified Embedding vector for calculating the similarity with the candidate text Embedding vector can be accurately generated, and then based on the obtained similarity value, the standard clinical term corresponding to the clinical term can be accurately determined, thereby effectively improving the accuracy of the obtained standard clinical term.
[0090] Further, in an embodiment of the present application, the above step S3 includes:
[0091] S300: Call the deep convolutional neural network and the attention model;
[0092] S301: Input all the first Embedding vectors into the deep convolutional neural network, and obtain a corresponding plurality of second Embedding vectors obtained by the deep convolutional neural network extracting feature information from each of the first Embedding vectors; wherein, input each of the first Embedding vectors into the convolutional layer of the deep convolutional neural network respectively, perform convolutional processing on each of the first Embedding vectors through the convolutional layer to obtain abstract feature Embedding vectors corresponding to each of the first Embedding vectors respectively, perform pooling processing on each of the abstract feature Embedding vectors in the convolutional layer through the pooling layer of the deep convolutional neural network to obtain processed abstract feature Embedding vectors corresponding to each of the abstract feature Embedding vectors respectively, and use the processed abstract feature Embedding vectors as the second Embedding vectors;
[0093] S302: Input all the second Embedding vectors into the attention model, and output a multi-dimensional vector corresponding to the second Embedding vectors through the attention model;
[0094] S303: Use the multi-dimensional vector as the specified Embedding vector.
[0095] As described in the above steps S300 to S303, the step of processing all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector may specifically include: First, call the deep convolutional neural network and the attention model. Then, input all the first Embedding vectors into the deep convolutional neural network, and obtain a corresponding plurality of second Embedding vectors obtained by the deep convolutional neural network after extracting feature information from each of the first Embedding vectors. Among them, the process of the deep convolutional neural network obtaining the corresponding plurality of second Embedding vectors after extracting feature information from each of the first Embedding vectors may include: (1) Input each of the first Embedding vectors into the convolutional layer in the deep convolutional neural network respectively, and perform convolutional processing on each of the first Embedding vectors through the convolutional layer to obtain an abstract feature Embedding vector corresponding to each of the first Embedding vectors. Specifically, the convolutional processing includes: where represents the input first Embedding vector, f x,i represents the convolutional template, n represents the number of elements in the convolutional template, C x represents the abstract feature Embedding vector obtained after convolutional processing, represents the bias, and j represents the position block information. (2) Perform pooling processing on each of the abstract feature Embedding vectors in the convolutional layer through the pooling layer in the deep convolutional neural network to obtain a processed abstract feature Embedding vector corresponding to each of the abstract feature Embedding vectors, and use the processed abstract feature Embedding vector as the second Embedding vector. Specifically, the pooling processing includes: where represents the second Embedding vector obtained after performing pooling processing on the obtained C x , the F() function represents the ReLU function, represents the weight information, It represents the bias, and k represents the position block information. Additionally, the number of the second Embedding vectors is the same as that of the first Embedding vectors. The role of the deep convolutional neural network is to enable each feature to consider the information of other features; that is, in Self-attention, only the relationships between each type of feature are considered, without considering the information between features. DCNN can make up for this weakness, enabling each type of feature to be correlated with each other and obtaining the information between each type of feature. For example: for the Embedding corresponding to the word feature, after inputting the Embedding of the word feature into the deep convolutional neural network, the output is the Embedding vector of each word, containing the information of other features. Then, all the second Embedding vectors are input into the attention model, and a multi-dimensional vector corresponding to the second Embedding vectors is output through the attention model. Specifically, in the process of generating a multi-dimensional vector through the attention module, mainly by substituting each second Embedding vector into the Attention vector SelfAttentionY respectively and then concatenating to determine the vector Z, and then determining the multi-dimensional vector of the clinical term based on the vector Z and the Attention vector SelfAttentionZ. Among them, the specific process of determining the Attention vector SelfAttentionY for evaluating the weighted relationship of each second Embedding vector based on the output vector Y and the weight vector WeightingY includes: first determining the output vector Y, Y = WR, where R is the input vector, W is the fully connected parameter matrix, and the R is the input vector, that is, it contains each of the second Embedding vectors (R1, R2, R3, R4, R5, R6), R i , i = 1, 2, 3, 4, 5, 6, perform a one-step fully connected network operation on R i to obtain Y i , the W is a parameter that the algorithm needs to learn, which can be randomly generated during initialization and updated through the backpropagation algorithm during subsequent iterative processes. WR is a common matrix multiplication. Then, based on the output vector Y, the weight vector WeightY is determined. The weight vector WeightY is determined by the following formula: where the Y is the output vector, the Y T is the transpose of the matrix Y, and the d k is the dimension of the matrix YY T . Finally, the Attention vector SelfAttentionY is determined based on the weight vector WeightY. The Attention vector SelfAttentionY can be determined by the following formula: wherein, the said Y is the output vector, and the said Y T is the transpose of matrix Y, and the said d k is the dimension of matrix YY T For vector Y, a unified Attention vector is obtained by combining weight calculations. This vector can capture the relationships within a single vector. For example, in a technical text, it can determine which word has the greatest impact on the semantics of the technical text, as well as the mutual influence between words. Thus, it can correctly model the semantics of the technical text. Additionally, by segmenting the clinical terms to be processed into words or splitting them character by character, performing part-of-speech tagging on each word, converting each character in the sentence into pinyin, and extracting the surgical site and the surgery, six semantic features are extracted and six corresponding Embedding vectors are constructed. After vector alignment, for each vector in the six Embedding vectors, the weight between each component is learned using the deep learning Self-Attention mechanism. Among them, the above Attention function can be considered as a mapping from a query and a set of key-value pairs to an output. Among them, the query, keys, value, and output are all vectors. The output is the weighted sum of the values, and the weight assigned to each value is calculated by the function of the query and the corresponding key. The dimensions of the query and the key are d k and the dimension of the value is d v , calculate the dot product of the query and all keys, and then divide it by the application of the SoftmaL function to obtain the weight of the value. The formula is as follows: Softmax is the formula for the weight vector: In Self-Attention, Q = K = V. So, by first passing the vector R into a fully connected layer, the output is Y, and the weight vector can be calculated using the following formula: Furthermore, the process of determining the multi-dimensional vector corresponding to the second Embedding vector based on the six Embedding vectors and the Attention vector SelfAttentionY includes: substituting each second Embedding vector into the Attention vector SelfAttentionY respectively and then concatenating them to determine the vector Z. Then, based on the vector Z and the Attention vector SelfAttentionZ, the multi-dimensional vector is determined. Among them, the multi-dimensional vector D is determined by the following formula: The said Z is the concatenated vector, and the said Z T is the transpose of matrix Z, and the said d k is the matrix ZZ TThe dimension. Additionally, for the second Embedding vector R of each dimension, after calculating its SelfAttentionY, the second Embedding vectors are concatenated to obtain Z, and then the SelfAttention vector of this vector Z is calculated. Z = (S1, S2, S3, S4); The obtained multi-dimensional vector can capture the relationships (context information) among the six-dimensional Embedding vectors, thereby improving the data richness of the obtained specified Embedding vector. For example: When there is a misspelled word in a sentence, if only word / character vectors are used for feature extraction, feature loss may occur due to the misspelled word. However, by adding features of pinyin and part of speech, the weights of the features in these two dimensions may increase, and the semantics of the entire sentence can be corrected through the weights of these two vectors, thereby achieving the purpose of correctly modeling the sentence features. Finally, the multi-dimensional vector is used as the specified Embedding vector. In this embodiment, by using a deep convolutional neural network and a preset attention model, the specified Embedding vector can be quickly and intelligently output, which is beneficial for subsequently calculating the similarity between the specified Embedding vector and each of the candidate text Embedding vectors based on a preset siamese network, and then quickly and accurately determining the standard clinical term corresponding to the clinical term from all the candidates based on all the similarities, so as to achieve the accurate acquisition of the standard clinical term for the clinical term.
[0096] Further, in an embodiment of the present application, the above step S4 includes:
[0097] S400: Use a preset word segmentation engine to perform word segmentation on the clinical term to obtain initial keywords;
[0098] S401: Use an N-gram model to perform part-of-speech tagging on the initial keywords to obtain the first keywords after tagging;
[0099] S402: Delete the keywords with the part of speech of function words from the first keywords to obtain the corresponding second keywords;
[0100] S403: Call the medical term library;
[0101] S404: Respectively obtain the occurrence times of each of the second keywords in the medical term library, and screen out the third keywords whose occurrence times are greater than a preset occurrence times threshold from all the second keywords;
[0102] S405: Obtain the first medical terms corresponding to the third keywords from the medical term library;
[0103] S406: Determine whether there are identical medical terms among all the first medical terms;
[0104] S407: If so, perform duplicate removal processing on the first medical terms to obtain corresponding second medical terms;
[0105] S408: Use all the second medical terms as the candidate set.
[0106] As described in the above steps S400 to S408, the step of generating a candidate set corresponding to the clinical term based on a preset medical term library may specifically include: First, use a preset word segmentation engine to perform word segmentation on the clinical term to obtain initial keywords. Among them, the word segmentation engine can use an existing word segmentation engine, such as BosonNLP. Then use the N-gram model to perform part-of-speech tagging on the initial keywords to obtain the first keywords after tagging. Among them, the N-gram model is a language model (Language Model, LM). A language model is a probability-based discriminant model. Its input is a sentence (a sequence of word orders), and its output is the probability of this sentence, that is, the joint probability of these words (jointprobability). In addition, the N-gram model has the function of part-of-speech tagging. After obtaining the first keywords, delete the keywords with the part-of-speech of function words from the first keywords to obtain the corresponding second keywords. Then call the medical term library to respectively obtain the number of occurrences of each of the second keywords in the medical term library, and screen out the third keywords whose number of occurrences is greater than a preset occurrence threshold from all the second keywords. Among them, the medical term library can be a pre-created database storing surgical standard terms (ICD9). Usually, the medical term library stores a large number of surgical operation text. If the similarity is calculated between the clinical term to be processed and each surgical operation text in it, it will cause the model to take a relatively long time for one mapping. Therefore, before performing the similarity comparison, a recall will be performed according to the clinical term input by the user. Generally, the recall quantity is less than 5,000. The recall is implemented using an inverted index. For example: The clinical term input by the user is: "Left upper lobectomy", and the recall results are: [Lobectomy, Resection of left abdominal wall mass, Gingivectomy, etc.]. In addition, the process of obtaining the first surgical text corresponding to the second keyword whose number of occurrences is greater than the preset number threshold can be based on the inverted index. First, perform an inverted index on the second keyword to obtain the corresponding inverted index result. Then, based on the inverted index result, obtain the number of occurrences of each of the second keywords in the medical term library. The inverted index originated from the need to find records according to the values of attributes in practical applications. Each item in this index table includes an attribute value and the addresses of each record with this attribute value. Since it is not the record that determines the attribute value, but the attribute value that determines the position of the record, it is called an inverted index. The inverted index table is used to record which documents contain a certain word.Generally, there are many documents in a document collection that contain a certain word. Each document records information such as the document number, the number of times the word appears in this document, and the positions where the word appears in the document. Such information related to a document is called an inverted index entry. A series of inverted index entries containing this word form a list structure, which is the inverted list corresponding to a certain word. Through the inverted index method, a batch of candidate texts with relatively high matching possibilities can be obtained. Among these candidate texts, there are standard clinical terms corresponding to the clinical terms to be processed. Then, through subsequent similarity calculations, the standard clinical terms can be determined. In this embodiment, by using the inverted index method to screen out the candidate set, it is beneficial to screen out the most suitable standard clinical terms for the clinical terms to be processed from the candidate set, thereby improving the accuracy of clinical term standardization. In addition, the value of the preset occurrence threshold is not specifically limited and can be set according to actual needs. After obtaining the third keyword, the first medical term corresponding to the third keyword is obtained from the medical term library. Subsequently, it is determined whether there are the same medical terms among all the first medical terms. If there are the same medical terms, duplicate removal processing is performed on the first medical terms to obtain the corresponding second medical terms. Among them, the duplicate removal processing means that for multiple identical medical terms, only one of them is retained and the other medical terms are deleted. Finally, all the second medical terms are used as the candidate set. In this embodiment, after obtaining the clinical terms, it will first screen out some medical term texts that meet the subsequent comparison requirements from the preset medical term library based on inverted index processing as candidate texts, which is beneficial to accurately determine the final standard clinical terms based on the obtained candidate texts. Since only the data comparison processing between the specified Embedding vector corresponding to the clinical term and the Embedding vector of each candidate text needs to be performed subsequently, rather than comparing the specified Embedding vector corresponding to the clinical term with the Embedding vector corresponding to each medical term included in the medical term library, the time spent in generating the standard clinical terms can be effectively reduced, thereby improving the processing efficiency of generating the standard clinical terms.
[0107] Further, in an embodiment of the present application, the siamese network includes a first network and a second network. The above step S6 includes:
[0108] S600: Input the specified Embedding vector into the first network, and obtain the first vector data output by the first network corresponding to the specified Embedding vector;
[0109] S601: Input the Embedding vector of the specified candidate text into the second network, and obtain the second vector data output by the second network corresponding to the Embedding vector of the specified candidate text; wherein, the specified candidate text Embedding is any one of all the candidate text Embeddings;
[0110] S602: Calculate the cosine value of the vectors between the first vector data and the second vector data through the cosine similarity formula;
[0111] S603: Use the cosine value of the vectors as the similarity between the first vector data and the second vector data.
[0112] As described in the above steps S600 to S603, the siamese network includes a first network and a second network. The steps of calculating the similarity between the specified Embedding vector and each candidate text Embedding vector based on the preset siamese network may specifically include: First, input the specified Embedding vector into the first network, and obtain the first vector data output by the first network corresponding to the specified Embedding vector. And at the same time, input the specified candidate text Embedding vector into the second network, and obtain the second vector data output by the second network corresponding to the specified candidate text Embedding vector. Wherein, the specified candidate text Embedding is any one of all the candidate text Embeddings. In addition, the siamese network is a BiLSTM network, and the two networks (the first network and the second network) share the weight W. The relevant introduction to the LSTM network is as follows: The Long Short-Term Memory network, generally called LSTM, is a special type of RNN. LSTM learns long-term dependence information through the forget gate, input gate, and output gate. Forget gate: Select to forget some past information, f t =σ(W f *[h t-1 , x t +b f ); Input gate: Remember some current information and merge the memories of the past and the present, i t =σ(W i *[h t-1 , x t +b i ); Output gate: Output, o t= σ(W o [h t-1 , x t + b o ); h t = o t * tanh(C t ); Parameter explanation: f t : The output content of the forgetting gate action at the current moment; h t-1 : The output value of the previous moment's action; x t : The action input at the current moment; i t : The output content of the input gate at the current moment; The content of the cell state at the current moment; C t : The value for updating the cell state combined with the previous moment; o t : The output result of the output gate; h t : Combine the cell state and o t Output the result at the current moment. Then, calculate the cosine value of the vector between the first vector data and the second vector data through the cosine similarity formula. Among them, after inputting the specified Embedding vector into the first network to obtain the first vector data, and inputting the specified candidate text Embedding vector into the second network to obtain the second vector data, the cosine similarity formula can be used subsequently to calculate the similarity between the first vector data and the second vector data. The closer the result is to 1, the higher the similarity between the specified candidate text and the clinical term. In addition, the cosine similarity formula is an existing technology at present. More specifically, relevant literature can be referred to, which will not be elaborated here. Finally, use the cosine value of the vector as the similarity between the first vector data and the second vector data. In this embodiment, after generating the first vector data and the second vector data by using the siamese network, the cosine value of the vector between the first vector data and the second vector data can be called to quickly calculate the similarity between the specified Embedding vector and each candidate text Embedding vector, which is beneficial to subsequent determination of the standard clinical term corresponding to the clinical term from all the candidates based on all the similarities, so as to accurately determine the standard clinical term corresponding to the clinical term, and further improve the accuracy of the obtained standard clinical term.
[0113] Further, in an embodiment of the present application, the above step S7 includes:
[0114] S700: Screen out the first similarity with the largest value from all the similarities;
[0115] S701: Determine whether the number of the first similarities is equal to 1;
[0116] S702: If so, obtain the first candidate text Embedding vector corresponding to the first similarity;
[0117] S703: Obtain the first candidate text corresponding to the first candidate text Embedding vector from the candidate set;
[0118] S704: Take the first candidate text as the standard clinical term.
[0119] As described in the above steps S700 to S704, the step of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities may specifically include: First, screen out the first similarity with the largest value from all the similarities. Among them, after obtaining the similarities, these similarities can be sorted in descending order of value, and the similarity ranked first is used as the first similarity with the largest value. Then judge whether the number of the first similarities is equal to 1. If it is equal to 1, obtain the first candidate text Embedding vector corresponding to the first similarity. Then obtain the first candidate text corresponding to the first candidate text Embedding vector from the candidate set. Finally, take the first candidate text as the standard clinical term. In this embodiment, after obtaining the similarities, the candidate text corresponding to the first similarity with the largest value is taken as the standard clinical term after screening all the similarities, effectively ensuring the accuracy of the obtained standard clinical term and being beneficial to improving the accuracy of clinical term standardization.
[0120] Further, in an embodiment of the present application, the above step S7 includes:
[0121] S710: Screen out the second similarity with the largest value from all the similarities;
[0122] S711: Judge whether there is a third similarity; wherein, the difference between the second similarity and the third similarity is within a preset numerical range;
[0123] S712: If there is the third similarity, obtain the second candidate text Embedding vectors corresponding to each of the second similarity and the third similarity;
[0124] S713: Obtain the second candidate texts corresponding to each of the second candidate text Embedding vectors from the candidate set;
[0125] S714: Receive the third candidate text selected by the preset user from all the second candidate texts;
[0126] S715: Use the third candidate text as the standard clinical term.
[0127] As described in the above steps S710 to S715, the step of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities may specifically include: First, screen out the second similarity with the largest value from all the similarities. Among them, after obtaining the similarities, these similarities can be sorted in descending order of value, and the similarity ranked first is used as the second similarity with the largest value. Then, determine whether there is a third similarity. Among them, the difference between the second similarity and the third similarity is within a preset numerical range. In addition, the value of the preset numerical range is not limited and can be set according to actual needs. When the difference between the second similarity and the third similarity is within the preset numerical range, it can indicate that the values of the second similarity and the third similarity are quite close. If there is the third similarity, obtain the second candidate text Embedding vectors corresponding to each of the second similarities and the third similarity one by one. Then, obtain the second candidate texts corresponding to each of the second candidate text Embedding vectors from the candidate set. Subsequently, receive the third candidate text selected by the preset user from all the second candidate texts. Among them, the preset user can be an expert. Finally, use the third candidate text as the standard clinical term. In this embodiment, when there are other similarities with values close to the similarity with the largest similarity value, the candidate text corresponding to the similarity with the largest value is not directly used as the standard clinical term, but the final standard clinical term is determined from multiple candidate texts corresponding to the similarity with the largest value based on the selection of the preset user, that is, the experience of the preset user will be further referred to determine the final standard clinical term, thereby further improving the accuracy of the obtained standard clinical term and the accuracy of clinical term standardization.
[0128] The method for standardizing clinical terms based on a model in the embodiments of the present application can also be applied to the blockchain field, such as storing data such as the above standard clinical terms on the blockchain. By using the blockchain to store and manage the above standard clinical terms, the security and immutability of the above standard clinical terms can be effectively ensured.
[0129] The above blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. A blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0130] Referring to Figure 2 , in an embodiment of the present application, a standardization device for clinical terms based on a model is further provided, including:
[0131] An acquisition module 1, configured to acquire clinical terms to be processed;
[0132] A first generation module 2, configured to perform multi-dimensional feature processing on the clinical terms, generate multiple text features corresponding to the clinical terms, and respectively convert each text feature into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features, and surgical approach features. The surgical site feature refers to the feature corresponding to any open site in the skin or internal organs, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations;
[0133] A processing module 3, configured to process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector;
[0134] A second generation module 4, configured to generate a candidate set corresponding to the clinical terms based on a preset medical term library; wherein, the candidate set includes multiple candidate texts;
[0135] A third generation module 5, configured to input each candidate text into a preset Tansformer encoder network to generate multiple candidate text Embedding vectors respectively corresponding to each candidate text;
[0136] A calculation module 6, configured to calculate the similarity between the specified Embedding vector and each candidate text Embedding vector based on a preset siamese network;
[0137] A determination module 7, configured to determine a standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities.
[0138] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0139] Further, in an embodiment of the present application, the above-mentioned first generation module 2 includes:
[0140] A first processing unit, configured to perform word segmentation processing on the clinical term based on a preset word segmentation tool to obtain corresponding character features, word features, and part-of-speech features;
[0141] A first acquisition unit, configured to acquire a pinyin feature corresponding to the clinical term based on a preset character-to-pinyin model;
[0142] A second acquisition unit, configured to acquire a surgical site feature and a surgical approach feature corresponding to the clinical term based on a preset medical thesaurus;
[0143] A first determination unit, configured to use the character features, the word features, the part-of-speech features, the pinyin feature, the surgical site feature, and the surgical approach feature as the text features.
[0144] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0145] Further, in an embodiment of the present application, the above-mentioned processing module 3 includes:
[0146] A first calling unit, configured to call the deep convolutional neural network and the attention model;
[0147] A first input unit, configured to input all the first Embedding vectors into the deep convolutional neural network to obtain a plurality of corresponding second Embedding vectors obtained by the deep convolutional neural network after extracting feature information from each of the first Embedding vectors; wherein, each of the first Embedding vectors is respectively input into a convolutional layer in the deep convolutional neural network, and each of the first Embedding vectors is respectively subjected to convolutional processing by the convolutional layer to obtain an abstract feature Embedding vector corresponding to each of the first Embedding vectors, and each of the abstract feature Embedding vectors in the convolutional layer is respectively subjected to pooling processing by a pooling layer in the deep convolutional neural network to obtain a processed abstract feature Embedding vector corresponding to each of the abstract feature Embedding vectors, and the processed abstract feature Embedding vector is used as the second Embedding vector;
[0148] A second input unit, configured to input all the second Embedding vectors into the attention model, and output, through the attention model, multi-dimensional vectors corresponding to the second Embedding vectors;
[0149] A second determination unit, configured to use the multi-dimensional vectors as the specified Embedding vectors.
[0150] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0151] Further, in an embodiment of the present application, the above second generation module 4 includes:
[0152] A second processing unit, configured to perform word segmentation on the clinical terms by using a preset word segmentation engine to obtain initial keywords;
[0153] A tagging unit, configured to perform part-of-speech tagging on the initial keywords by using an N-gram model to obtain first keywords after tagging;
[0154] A deletion unit, configured to delete keywords with function word parts of speech from the first keywords to obtain corresponding second keywords;
[0155] A second calling unit, configured to call the medical terminology library;
[0156] A first screening unit, configured to respectively obtain the occurrence times of each of the second keywords in the medical terminology library, and screen out third keywords from all the second keywords, where the occurrence times are greater than a preset occurrence times threshold;
[0157] A third obtaining unit, configured to obtain first medical terms corresponding to the third keywords from the medical terminology library;
[0158] A first judgment unit, configured to judge whether there are identical medical terms among all the first medical terms;
[0159] A third processing unit, configured to, if so, perform duplicate removal processing on the first medical terms to obtain corresponding second medical terms;
[0160] A third determination unit, configured to use all the second medical terms as the candidate set.
[0161] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0162] Further, in an embodiment of the present application, the above calculation module 6 includes:
[0163] A fourth acquisition unit, configured to input the specified Embedding vector into the first network, and acquire first vector data output by the first network corresponding to the specified Embedding vector;
[0164] A fifth acquisition unit, configured to input a specified candidate text Embedding vector into the second network, and acquire second vector data output by the second network corresponding to the specified candidate text Embedding vector; wherein, the specified candidate text Embedding is any one of all the candidate text Embeddings;
[0165] A calculation unit, configured to calculate a vector cosine value between the first vector data and the second vector data through a cosine similarity formula;
[0166] A fourth determination unit, configured to use the vector cosine value as the similarity between the first vector data and the second vector data.
[0167] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0168] Further, in an embodiment of the present application, the above determination module 7 includes:
[0169] A second screening unit, configured to screen out a first similarity with the largest value from all the similarities;
[0170] A second judgment unit, configured to judge whether the number of the first similarities is equal to 1;
[0171] A sixth acquisition unit, configured to, if so, acquire a first candidate text Embedding vector corresponding to the first similarity;
[0172] A seventh acquisition unit, configured to acquire a first candidate text corresponding to the first candidate text Embedding vector from the candidate set;
[0173] A fifth determination unit, configured to use the first candidate text as the standard clinical term.
[0174] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0175] Further, in an embodiment of the present application, the above determination module 7 includes:
[0176] A third screening unit, configured to screen out a second similarity with the largest value from all the similarities;
[0177] A third determination unit, configured to determine whether there is a third similarity; wherein, a difference between the second similarity and the third similarity is within a preset value range;
[0178] An eighth acquisition unit, configured to, if there is the third similarity, acquire second candidate text Embedding vectors corresponding to each of the second similarities and the third similarity one by one;
[0179] A ninth acquisition unit, configured to acquire second candidate texts corresponding to each of the second candidate text Embedding vectors from the candidate set one by one;
[0180] A receiving unit, configured to receive a third candidate text selected by a preset user from all the second candidate texts;
[0181] A sixth determination unit, configured to use the third candidate text as the standard clinical term.
[0182] In this embodiment, the operations respectively performed by the above modules or units correspond one by one to the steps of the method for standardizing clinical terms based on a model in the foregoing embodiment, and will not be elaborated herein.
[0183] Referring to Figure 3 , in the embodiment of the present application, a computer device is further provided. The computer device may be a server, and its internal structure may be as Figure 3As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, an input device, and a database connected through a system bus. Among them, the processor designed for the computer device is used to provide computing and control capabilities. The memory of the computer device includes a storage medium and an internal memory. The storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the storage medium. The database of the computer device is used to store clinical terms, first Embedding vectors, specified Embedding vectors, candidate texts, candidate text Embedding vectors, similarity, and standard clinical terms. The network interface of the computer device is used to communicate with an external terminal through a network connection. The display screen of the computer device is an essential graphic and text output device in the computer, which is used to convert digital signals into optical signals so that text and graphics are displayed on the screen of the display screen. The input device of the computer device is the main device for information exchange between the computer and users or other devices, and is used to transmit data, instructions, and certain flag information into the computer. When the computer program is executed by the processor, it implements a method for standardizing clinical terms based on a model.
[0184] The above processor executes the steps of the above method for standardizing clinical terms based on a model:
[0185] Obtain the clinical term to be processed;
[0186] Perform multi-dimensional feature processing on the clinical term to generate multiple text features corresponding to the clinical term, and convert each text feature into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features, and surgical approach features. The surgical site feature refers to the feature corresponding to any open site in the skin or internal organs, and the surgical approach feature refers to the feature corresponding to the incision position for clinical surgical operations;
[0187] Process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector;
[0188] Generate a candidate set corresponding to the clinical term based on a preset medical term library; wherein, the candidate set includes multiple candidate texts;
[0189] Input each candidate text into a preset Transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each candidate text respectively;
[0190] Calculate the similarity between the specified Embedding vector and each of the candidate text Embedding vectors respectively based on a preset twin network;
[0191] Based on all the similarities, determine the standard clinical term corresponding to the clinical term from all the candidate set texts.
[0192] Those skilled in the art can understand that Figure 3 the structure shown in
[0193] One embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for standardizing clinical terms based on a model, specifically:
[0194] Obtain the clinical term to be processed;
[0195] Perform multi-dimensional feature processing on the clinical term to generate multiple text features corresponding to the clinical term, and convert each of the text features into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features and surgical approach features, the surgical site feature refers to the feature corresponding to any open site in the skin or internal organs, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations;
[0196] Process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector;
[0197] Generate a candidate set corresponding to the clinical term based on a preset medical term library; wherein, the candidate set includes multiple candidate texts;
[0198] Input each of the candidate texts into a preset Transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each of the candidate texts respectively;
[0199] Calculate the similarity between the specified Embedding vector and each of the candidate text Embedding vectors respectively based on a preset twin network;
[0200] Based on all the similarities, determine the standard clinical term corresponding to the clinical term from all the candidate set texts.
[0201] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0202] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for standardizing model-based clinical terms, characterized in that, Including: Obtain clinical terms to be processed; Perform multi-dimensional feature processing on the clinical terms to generate multiple text features corresponding to the clinical terms, and convert each of the text features into a corresponding first Embedding vector; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features and surgical approach features, the surgical site feature refers to the feature corresponding to any open site in the skin or viscera, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations; Process all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector; Generate a candidate set corresponding to the clinical terms based on a preset medical term library; wherein, the candidate set includes multiple candidate texts; Input each of the candidate texts into a preset transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each of the candidate texts; Calculate the similarity between the specified Embedding vector and each of the candidate text Embedding vectors respectively based on a preset siamese network; Determine the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities; The step of processing all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector includes: Invoke the deep convolutional neural network and the attention model; Input all the first Embedding vectors into the deep convolutional neural network to obtain multiple corresponding second Embedding vectors obtained by the deep convolutional neural network after extracting feature information from each of the first Embedding vectors; wherein, input each of the first Embedding vectors into the convolutional layer of the deep convolutional neural network, perform convolutional processing on each of the first Embedding vectors through the convolutional layer to obtain abstract feature Embedding vectors corresponding to each of the first Embedding vectors, perform pooling processing on each of the abstract feature Embedding vectors in the convolutional layer through the pooling layer of the deep convolutional neural network to obtain processed abstract feature Embedding vectors corresponding to each of the abstract feature Embedding vectors, and use the processed abstract feature Embedding vectors as the second Embedding vectors; Input all the second Embedding vectors into the attention model, and output a multi-dimensional vector corresponding to the second Embedding vector through the attention model; Use the multi-dimensional vector as the specified Embedding vector; The step of generating a candidate set corresponding to the clinical term based on a preset medical term library includes: Performing word segmentation on the clinical term using a preset word segmentation engine to obtain initial keywords; Performing part-of-speech tagging on the initial keywords using an N-gram model to obtain the first tagged keywords; Deleting the keywords with function words as their part of speech from the first keywords to obtain corresponding second keywords; Invoking the medical term library; Respectively obtaining the occurrence times of each of the second keywords in the medical term library, and screening out the third keywords with the occurrence times greater than a preset occurrence times threshold from all the second keywords; Obtaining the first medical terms corresponding to the third keywords from the medical term library; Judging whether there are the same medical terms among all the first medical terms; If so, performing deduplication processing on the first medical terms to obtain corresponding second medical terms; Taking all the second medical terms as the candidate set; The step of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities includes: Screening out the first similarity with the largest value from all the similarities; Judging whether the value of the first similarity is equal to 1; If so, obtaining the first candidate text Embedding vector corresponding to the first similarity; Obtaining the first candidate text corresponding to the first candidate text Embedding vector from the candidate set; Taking the first candidate text as the standard clinical term; The step of determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities includes: Screening out the second similarity with the largest value from all the similarities; Judging whether there is a third similarity; wherein, the difference between the second similarity and the third similarity is within a preset numerical range; If there is the third similarity, obtaining the second candidate text Embedding vectors corresponding to each of the second similarities and the third similarity one by one; Obtaining the second candidate texts corresponding to each of the second candidate text Embedding vectors from the candidate set one by one; Receiving the third candidate text selected by a preset user from all the second candidate texts; Taking the third candidate text as the standard clinical term.
2. The method for standardizing model-based clinical terms according to claim 1, characterized in that,The step of performing multi-dimensional feature processing on the clinical term to generate multiple text features corresponding to the clinical term includes: Performing word segmentation on the clinical term based on a preset word segmentation tool to obtain corresponding character features, word features and part-of-speech features; Obtaining the pinyin feature corresponding to the clinical term based on a preset character-to-pinyin model; Obtaining the surgical site feature and the surgical approach feature corresponding to the clinical term based on a preset medical word library; Taking the character features, the word features, the part-of-speech features, the pinyin features, the surgical site features and the surgical approach features as the text features.
3. The standardized method for model-based clinical terms according to claim 1, wherein, The Siamese network includes a first network and a second network. The step of calculating the similarity between the specified Embedding vector and each of the candidate text Embedding vectors based on the preset Siamese network includes: Input the specified Embedding vector into the first network to obtain the first vector data output by the first network corresponding to the specified Embedding vector; Input the specified candidate text Embedding vector into the second network to obtain the second vector data output by the second network corresponding to the specified candidate text Embedding vector; wherein, the specified candidate text Embedding is any one of all the candidate text Embeddings; Calculate the cosine value of the vectors between the first vector data and the second vector data through the cosine similarity formula; Use the cosine value of the vectors as the similarity between the first vector data and the second vector data.
4. A standardized device for model-based clinical terms, wherein, Includes: An acquisition module for acquiring clinical terms to be processed; A first generation module for performing multi-dimensional feature processing on the clinical terms to generate multiple text features corresponding to the clinical terms, and respectively converting each of the text features into corresponding first Embedding vectors; wherein, the text features include: character features, word features, part-of-speech features, pinyin features, surgical site features and surgical approach features, the surgical site feature refers to the feature corresponding to any open part in the skin or internal organs, and the surgical approach feature refers to the feature corresponding to the incision position for performing clinical surgical operations; A processing module for processing all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector; A second generation module for generating a candidate set corresponding to the clinical terms based on a preset medical term library; wherein, the candidate set includes multiple candidate texts; A third generation module for inputting each of the candidate texts into a preset transformer encoder network to generate multiple candidate text Embedding vectors corresponding to each of the candidate texts respectively; A calculation module for calculating the similarity between the specified Embedding vector and each of the candidate text Embedding vectors based on a preset Siamese network; A determination module for determining the standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities; The step of processing all the first Embedding vectors based on a preset deep convolutional neural network and a preset attention model to obtain a specified Embedding vector includes: Invoke the deep convolutional neural network and the attention model; Input all the first Embedding vectors into the deep convolutional neural network, and obtain a corresponding plurality of second Embedding vectors obtained by the deep convolutional neural network after extracting feature information from each of the first Embedding vectors; wherein, input each of the first Embedding vectors into the convolutional layer of the deep convolutional neural network respectively, perform convolutional processing on each of the first Embedding vectors through the convolutional layer to obtain abstract feature Embedding vectors corresponding to each of the first Embedding vectors respectively, perform pooling processing on each of the abstract feature Embedding vectors in the convolutional layer through the pooling layer of the deep convolutional neural network to obtain processed abstract feature Embedding vectors corresponding to each of the abstract feature Embedding vectors respectively, and use the processed abstract feature Embedding vectors as the second Embedding vectors; Input all the second Embedding vectors into the attention model, and output a multi-dimensional vector corresponding to the second Embedding vectors through the attention model; Use the multi-dimensional vector as the specified Embedding vector; The step of generating a candidate set corresponding to the clinical term based on a preset medical term library includes: Use a preset word segmentation engine to perform word segmentation on the clinical term to obtain initial keywords; Use an N-gram model to perform part-of-speech tagging on the initial keywords to obtain the first keywords after tagging; Delete the keywords with the part-of-speech of function words from the first keywords to obtain corresponding second keywords; Call the medical term library; Obtain the occurrence times of each of the second keywords in the medical term library respectively, and screen out the third keywords whose occurrence times are greater than a preset occurrence times threshold from all the second keywords; Obtain the first medical terms corresponding to the third keywords from the medical term library; Judge whether there are the same medical terms among all the first medical terms; If so, perform deduplication processing on the first medical terms to obtain corresponding second medical terms; Use all the second medical terms as the candidate set; The step of determining a standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities includes: Screen out the first similarity with the largest value from all the similarities; Judge whether the value of the first similarity is equal to 1; If so, obtain the first candidate text Embedding vector corresponding to the first similarity; Obtain the first candidate text corresponding to the first candidate text Embedding vector from the candidate set; Use the first candidate text as the standard clinical term; The step of determining a standard clinical term corresponding to the clinical term from all the candidate set texts based on all the similarities includes: Select the second similarity with the largest value from all the similarities; Determine whether there is a third similarity; wherein, the difference between the second similarity and the third similarity is within a preset numerical range; If there is the third similarity, obtain the second candidate text Embedding vectors corresponding one-to-one to each of the second similarities and the third similarity; Obtain the second candidate texts corresponding one-to-one to each of the second candidate text Embedding vectors from the candidate set; Receive the third candidate text selected by the preset user from all the second candidate texts; Use the third candidate text as the standard clinical term.
5. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.
6. A computer-readable storage medium, on which a computer program is stored, and characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Intention recognition method, apparatus, and device, and computer readable storage medium
WO2021143018A1
Bad term recognition method and device, electronic device, and storage medium
WO2021143020A1