Term standardization method and device, electronic equipment and storage medium
By extracting the terms to be standardized from the original text and using keyword search and diffusion models to generate the final standardized terms, communication barriers caused by inconsistency of terms are solved, and efficient sharing and standardization of medical information is achieved.
Patent Information
- Application Number
- CN202510837828.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-13
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Misunderstanding of intentions and communication barriers caused by inconsistency of terminology between different medical information systems and institutions affects the sharing of medical data and the quality of medical services.
By extracting the terms to be standardized from the original text, using keyword search matching to generate candidate normalized terms, and format control through diffusion models and large language models, the final standardized terms are finally screened to ensure the accuracy and normativeness of the terms.
It improves the vocabulary scale of term recall, enhances the flexibility and accuracy of term standardization, reduces information omissions, and promotes the interoperability and data sharing of medical information.
Smart Images

Figure CN120336511A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for term standardization. Background Art
[0002] Medical terms refer to specialized words and phrases used to describe the human body, diseases, diagnoses, treatments, and medical research. They scientifically define concepts, indicators, symptoms, etc. in the medical field, helping medical professionals accurately communicate and understand medical information, and having basic characteristics such as professionalism, scientificity, univocality, and systematicness.
[0003] In view of the different terms used in different scenarios, different ways, and different institutions during medical treatment, there may be misunderstandings in intention and communication barriers. For example, there are still cases of incorrect collocations of three common groups of related medical terms formed by the three characters "symptom, sign, syndrome", namely "disease condition, symptom; physical sign, syndrome; indication, contraindication" in many medical records, drug instructions, media news, and related publications.
[0004] Therefore, it is urgent to ensure the consistency and mutual understanding of medical semantic expressions between different medical information systems through medical term standardization, reduce medical errors caused by term confusion, promote the sharing of medical data across regions and institutions, support broader modern medical research and intelligent medical services, and improve the safety and quality of medical services. Summary of the Invention
[0005] The present invention provides a method, device, electronic device, and storage medium for term standardization to solve the defects of misunderstanding in intention and communication barriers caused by inconsistent terms in the prior art.
[0006] The present invention provides a method for term standardization, including: extracting terms to be standardized and keywords in the terms to be standardized from the original text; performing a retrieval match on the keywords, and obtaining candidate standardized terms based on the matching results; generating target standardized terms with the candidate standardized terms as conditional texts; screening the target standardized terms to obtain final standardized terms.
[0007] According to the method for term standardization provided by the present invention, the step of generating target standardized terms with the candidate standardized terms as conditional texts includes: generating target standardized terms that meet the format requirements with the candidate standardized terms as conditional texts based on a trained diffusion model.
[0008] According to the term standardization method provided by the present invention, based on the trained diffusion model, using the candidate standardized term as the conditional text, generating the target standardized term that meets the format requirements includes: Based on the diffusion model, mapping the candidate standardized term to the noise space to obtain a noise-added term, and performing conditional encoding on the noise-added term to obtain a noise-added term feature; Denosing the noise-added term feature to obtain a denoised term feature, and decoding the denoised term feature to obtain the target standardized term.
[0009] According to the term standardization method provided by the present invention, extracting the term to be standardized and the keywords in the term to be standardized from the original text includes: Based on a large language model, semantically understanding the original text and extracting the key information in the original text to obtain the term to be standardized; Performing phrase grouping on the term to be standardized to obtain the keywords in the term to be standardized.
[0010] According to the term standardization method provided by the present invention, retrieving and matching the keywords, and obtaining candidate standardized terms based on the matching results includes: Based on a term association network, retrieving and matching the keywords, and performing keyword conversion on the matched terms to obtain the converted keywords in the term to be standardized; Recombining the converted keywords in the term to be standardized to obtain the candidate standardized terms.
[0011] According to the term standardization method provided by the present invention, screening the target standardized term to obtain the final standardized term includes: Based on the correlation between the target standardized term and the reference entries, scoring and ranking the target standardized term, and screening the target standardized term based on the ranking result to obtain the final standardized term; The reference entries include at least one of the original text, the converted keywords, and the term to be standardized.
[0012] According to the term standardization method provided by the present invention, scoring and ranking the target standardized term based on the correlation between the target standardized term and the reference entries includes: Based on a large language model, determining the text semantic correlation between the target standardized term and the original text; Based on the inclusion degree between the target standardized term and the converted keywords, determining the inclusion degree correlation; Determine character relevance based on the relevance between each character of the target standardized term and each character of the conversion keyword; Determine term semantic relevance based on the semantic relevance between the target standardized term and the term to be standardized; Score and rank the target standardized term based on at least one of the text semantic relevance, the inclusion relevance, the character relevance, and the term semantic relevance.
[0013] The present invention also provides a term standardization device, including: A term extraction unit, configured to extract a term to be standardized and keywords in the term to be standardized from the original text; A retrieval matching unit, configured to perform retrieval matching on the keywords and obtain candidate standardized terms based on the matching results; A term generation unit, configured to generate a target standardized term with the candidate standardized term as conditional text; A term screening unit, configured to screen the target standardized term to obtain the final standardized term.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the term standardization method as described in any one of the above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the term standardization method as described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the term standardization method as described in any one of the above is implemented.
[0017] The term standardization method, device, electronic device, and storage medium provided by the present invention utilize keyword information retrieval, integrate the ability to generate conditional text, avoid omission of term information extraction, and increase the scale of the term recall library. By generating conditional control of the term format, it is possible to process complex and variable term forms, enhancing the flexibility of term standardization.
[0018] In addition, by refining the term standardization process, the present invention makes the term standardization process clearer. Through process staging, the latest and more external solutions or knowledge can be introduced in each stage, which helps to combine specific scenarios for application practice and optimization, improving the flexibility and accuracy of term standardization. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is one of the flow diagrams of the term standardization method provided by the present invention.
[0021] Figure 2 It is the flow diagram of the implementation manner of step 110 in the term standardization method provided by the present invention.
[0022] Figure 3 It is the flow diagram of the implementation manner of step 120 in the term standardization method provided by the present invention.
[0023] Figure 4 It is one of the flow diagrams of the implementation manner of step 130 in the term standardization method provided by the present invention.
[0024] Figure 5 It is the second flow diagram of the implementation manner of step 130 in the term standardization method provided by the present invention.
[0025] Figure 6 It is the second flow diagram of the term standardization method provided by the present invention.
[0026] Figure 7 It is the structural diagram of the term standardization device provided by the present invention.
[0027] Figure 8 It is the structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0028] To make the objectives, technical solutions, and advantages of the present invention clearer, the following clearly and completely describes the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0029] Medical terminology standardization refers to the process of unifying medical terms, aiming to ensure consistent and accurate communication among different medical systems and personnel. Standardization promotes information transmission, data sharing, medical research, and cross-regional health cooperation. The task of medical terminology standardization is to convert informal medical expressions into formal medical concepts. For example, "persistent nasal bleeding" is mapped to "epistaxis", and "lung abscess without pneumonia" is mapped to "lung abscess not accompanied by pneumonia".
[0030] Currently, based on the classification and summary of non-standard terms or phrases, to solve the problem of term standardization, attempts can be made from 10 perspectives, including exact matching, abbreviation restoration, language component conversion, numerical substitution, hyphen parsing, suffix completion, synonym replacement, stemming, compound term splitting, partial matching, etc. From an algorithmic perspective, existing methods for solving medical terminology standardization can be divided into three categories: rule- and classical machine learning-based methods, deep learning- and natural language understanding-based methods, and pre-trained model-based methods.
[0031] (1) Rule- and classical machine learning-based methods.
[0032] From the perspective of rule design, first analyze the formal characteristics of non-standard medical terms, classify medical terms or phrases according to sentence patterns, syntax, sentence types, etc., and design processing rules and methods based on 10 standardization perspectives. For example, by applying rules to extract the main body of a phrase or restore the original form of a phrase, establish a thesaurus, use rules to identify synonyms in medical terms, and thus map them to unified standard terms. For abbreviations in medical terms, expand them into full names according to rules.
[0033] In the application of classical machine learning, unlabeled datasets can be used, and unsupervised learning methods such as clustering and dimensionality reduction can be adopted to discover the correlations and patterns among terms, thereby achieving standardization. Supervised learning methods such as classification or regression can also be used with labeled datasets to train models and map non-standard terms to standard terms. Medical terms can also be converted into feature representations that can be processed by machine learning algorithms. For example, methods such as the bag-of-words model and TF-IDF can be used to perform term standardization.
[0034] (2) Deep learning- and natural language understanding-based methods.
[0035] With the development of deep learning algorithms and technological breakthroughs in natural language understanding, a common strategy is to map medical terms to a continuous real-valued vector space by training word vector models, such as the Word2Vec algorithm and the GloVe algorithm, in order to capture the semantic similarity between words. In addition, through the self-attention mechanism, Transformer can effectively capture the dependencies between different positions in medical terms, which helps to better understand the semantic associations between terms.
[0036] Methods based on deep learning and natural language understanding rely on initial word embeddings and cannot represent polysemy. Relying solely on discriminative models cannot obtain complete semantic information, and the feature inclusion of word vectors is not rich enough to meet complex and specialized language understanding scenarios.
[0037] (3) Methods based on pre-trained models.
[0038] As a pre-trained language model based on the Transformer architecture, the BERT (Bidirectional Encoder Representations from Transformers) model greatly improves the performance of natural language processing tasks by learning context information bidirectionally. In related technologies, the medical term standardization task is regarded as a translation task. In the first stage, a generative model is used to generate candidate entities. In the second stage, the BERT pre-trained model is used to rank the semantic similarity of candidate entities to obtain the final medical term standardization result. This scheme has achieved good results in the medical term standardization task. In another two-stage term standardization scheme, in the first stage, similarity recall is performed based on various traditional statistical methods such as Jaccard and TF-IDF. In the second stage, the RoBERTa-wwm-ext pre-trained language model is used to perform phrase pair text classification on the entities to be matched and candidate entities, and it is judged whether two medical terms can be aligned according to the classification results. This scheme has also achieved good results in practical applications.
[0039] However, due to data bias and model bias, the "hallucination" problem may occur, especially in the processing of rare terms, where the performance is not good.
[0040] In view of the above problems, an embodiment of the present invention proposes a term standardization method, in which the terms to be standardized and the keywords in the terms to be standardized are first extracted from the original text, the keywords are searched and matched, and candidate standardized terms are obtained based on the matching results; the candidate standardized terms are used as conditional texts to generate target standardized terms; the target standardized terms are screened to obtain the final standardized terms. The method provided by the embodiment of the present invention integrates the conditional text generation capability with the help of keyword information retrieval, avoids omissions in term information extraction, and increases the vocabulary size of term recall.
[0041] In addition, the embodiments of the present invention refine the terminology standardization process to make the terminology standardization process clearer. By dividing the process into stages, each stage can introduce the latest and more external solutions or knowledge, which helps to carry out application practice and optimization in combination with specific scenarios, thereby improving the flexibility and accuracy of terminology standardization.
[0042] The embodiments of the present invention can be applied to scenarios where terminology standardization is required, such as in the medical field, mechanical engineering field, scientific field, education field, etc. The execution subject of the method can be an electronic device such as a terminal device, a computer, a server, a server cluster, or a specially designed terminology standardization device, or a terminology standardization device set in the electronic device, and the terminology standardization device can be implemented by software, hardware, or a combination of the two.
[0043] Figure 1 is one of the flow charts of the terminology standardization method provided by the present invention, such as Figure 1 As shown, this embodiment takes medical terminology standardization as an example to describe the terminology standardization method. The method includes the following steps: Step 110 , extracting the terms to be standardized and the keywords in the terms to be standardized from the original text.
[0044] Specifically, the original text refers to the text containing the terms to be standardized, which may be non-standard, irregular or ambiguous expressions. The original text may be, for example, a question text asked by a patient, a medical record text, or a diagnosis text, etc., which is not specifically limited in the embodiments of the present invention.
[0045] Terms to be standardized refer to words or phrases in the original text that need to be converted or replaced into more standard and standardized expressions.
[0046] The extraction of the terms to be standardized can be obtained from the original text by rule-based methods, statistical methods, pre-trained language model-based methods, etc. For example, in the medical field, rules can be set to identify drug names, disease names, etc., or to identify important words or phrases in the text and use them as terms to be standardized.
[0047] Taking the original text as the user's question text as an example, the original text is "Hello, doctor. I recently had a brain infarction detected. I feel dizzy and my memory is a bit reduced. My heart sometimes beats irregularly. The doctor said it's not a problem with the heart rate. I want to know how serious these problems are? My blood pressure has been 150 / 95 recently. I don't know if this value is high? Also, are there any lifestyle suggestions that can help improve these symptoms? Do I need to take any medicine? Thank you!"
[0048] Obviously, the patient's expression is relatively non-standard, and they are used to using colloquial and common language forms, such as expressing "cerebral infarction" as the non-standard "cerebral infarction", miswriting "arrhythmia" as "heart rate irregularity", and even having typos. Although it is understandable for doctors to grasp the patient's key intentions, for artificial intelligence or intelligent medical Q&A models, it may lead to understanding deviations, incorrect medical advice answers, and affect user satisfaction. Therefore, it is necessary to standardize the terms in the original text.
[0049] Let the number of terms to be standardized in a user question (original text) be n, and the set represents the set of terms to be standardized, represents the nth term to be standardized. The goal is to convert all elements in the set into the corresponding standard terms, that is , represents the nth standard term.
[0050] Here, the keyword in the term to be standardized is the core part of the term to be standardized, usually a word with specific meaning or function. The keyword can be extracted from the term to be standardized through keyword extraction algorithms, domain dictionary matching, semantic analysis, etc.
[0051] For example, the term to be standardized is "multiple enlarged cervical lymph nodes", and the keywords in it include "neck", "lymph node", "multiple", and "enlarged".
[0052] Step 120, perform a retrieval match on the keyword and obtain candidate standardized terms based on the match result.
[0053] Specifically, after extracting the term to be standardized and its keyword from the original text, use the keyword to perform a retrieval match in a specific term library, dictionary, or database. This term library or database should contain a large number of standardized terms and their corresponding non-standard forms. Through the retrieval match, related ones can be found.
[0054] Candidate standardized terms refer to the possible standardized terms obtained based on the keyword retrieval match, and these terms need to be further screened and confirmed.
[0055] Step 130: Generate a target standardized term with the candidate standardized term as the conditional text.
[0056] In this step, considering that in the related art, a term standardization method based on a retrieval strategy is usually adopted, for terms not included in the library, it may be impossible to standardize them and it is difficult to handle complex and variable term forms. In the embodiments of the present invention, based on the retrieval strategy, a target standardized term is generated with the candidate standardized term as the conditional text. Through the generative technology, terms not included in the library can be processed, the scale of term recall is increased, and thus the coverage rate of term standardization is improved. In addition, the generated target standardized term is generated with the candidate standardized term as the conditional text. By generating conditional control of the term format, complex and variable term forms such as abbreviations, synonyms, and specific format requirements in the medical field can be processed, enhancing the flexibility of term standardization.
[0057] For the generation of the target standardized term, it can be implemented through certain rules or algorithms, such as spelling check, grammar analysis, synonym replacement, etc. It can also be implemented through a trained text generation model. The candidate standardized term is input into the trained text generation model as the conditional text, and the text sequence output by the text generation model, that is, the target standardized term, is obtained. The target standardized term is the expected standardized term generated based on the candidate standardized term through certain rules or algorithm processing.
[0058] The text generation model can be a sequence-to-sequence (Seq2Seq) model, a denoising diffusion model, a generative adversarial network, a Transformer model, etc. The embodiments of the present invention do not make specific limitations on this.
[0059] Step 140: Screen the target standardized term to obtain the final standardized term.
[0060] Specifically, after the target standardized term is generated, it needs to be further screened and confirmed. The screening process can be automatically screened using machine learning algorithms. The screening criteria can include the accuracy, normativity, general acceptance, and context adaptability of the term. Through screening, it can be ensured that the finally obtained standardized term is accurate, normative, and meets the requirements.
[0061] The method provided by the embodiments of the present invention integrates the ability to generate conditional text by means of keyword information retrieval, avoiding omission of term information extraction and improving the scale of the thesaurus for term recall. By generating conditional control of the term format, complex and variable term forms can be processed.
[0062] In addition, in the embodiments of the present invention, by refining the term standardization process, the term standardization process becomes clearer. Through the process of phasing, the latest and more external solutions or knowledge can be introduced in each stage, which helps to carry out application practice and optimization in combination with specific scenarios, and improves the flexibility and accuracy of term standardization.
[0063] Based on any of the above embodiments, Figure 2 is a schematic flowchart of the implementation manner of step 110 in the term standardization method provided by the present invention, as Figure 2 shown, extracting the terms to be standardized and the keywords in the terms to be standardized from the original text, that is, step 110 specifically includes: Step 111, based on a large language model, perform semantic understanding on the original text and extract key information in the original text to obtain the terms to be standardized; Step 112, phrase the terms to be standardized to obtain the keywords in the terms to be standardized.
[0064] Specifically, for the terms to be standardized and the keywords they contain, information extraction and term phrasing can be achieved through a large language model.
[0065] The large language model here can be a pre-trained large language model, for example, it can include BERT, RoBERTa, RoBERTa-wwm, iFlytek Spark Medical Model, etc.
[0066] Semantic understanding means that the model can understand the text content, so as to perform information extraction tasks more accurately. First, use prompt engineering techniques to construct a suitable Prompt, set the granularity of entity extraction, and require the large model to combine the context information of the original text (such as the question text of the user's question) to perform semantic understanding on the original text and supplement the implicit medical entities. Organize the extracted key information into a list of terms to be standardized.
[0067] On this basis, phrase the terms to be standardized. Phrasing can be achieved through a component decomposition model, such as methods like named entity recognition and conditional random fields, to phrase the content after extraction (it can also continue to use the large language model to complete this task, and at this time, a new Prompt needs to be designed according to the task objectives and requirements).
[0068] After phrasing, a set of keywords in the terms to be standardized is obtained, and this set of keywords can be represented as , where represents the th keyword after phrasing of the nth term to be standardized, represents the size of the keyword list obtained by phrasing each different term to be standardized.
[0069] The method provided by the embodiments of the present invention semantically understands the original text based on a large language model and extracts key information, and then performs phrase formation to obtain keywords. With the help of the deep semantic understanding ability of the large language model, it can accurately extract the terms to be standardized, thereby improving the accuracy and efficiency of information extraction.
[0070] Based on any of the above embodiments, Figure 3 is a schematic flowchart of the implementation manner of step 120 in the term standardization method provided by the present invention. As Figure 3 shown, retrieve and match the keywords, and obtain candidate standardized terms based on the matching results. That is, step 120 specifically includes: Step 121, based on the term association network, retrieve and match the keywords, and perform keyword conversion based on the matched terms to obtain the conversion keywords in the terms to be standardized; Step 122, reorganize the conversion keywords in the terms to be standardized to obtain candidate standardized terms.
[0071] Specifically, multiple keyword lists with different lengths can be obtained in step 110. Then, candidate standardized terms are obtained based on the keyword retrieval strategy.
[0072] Here, the term association network is a pre-constructed complex network structure, where nodes represent terms and edges represent the relationships between terms. Such relationships can be synonyms, near-synonyms, hyponyms, etc.
[0073] Use the constructed term association network (such as a medical term association network) to retrieve and match each group of keywords. If no matching standard term is retrieved, the original keyword is retained. If a standard term matching the keyword is retrieved, the keyword is replaced with the retrieved standard term. In this way, each group of keywords is converted into a new keyword list with almost no information loss. Here, the new keyword list is the conversion keyword.
[0074] Convert the keywords in the terms to be standardized, expressed as , to obtain the conversion keyword set . Among them, is the keyword set after phrase formation of the i-th term to be standardized, is the th keyword in , and is the j-th conversion keyword in the conversion keyword set
[0075] Subsequently, each newly generated keyword group, i.e., the transformed keyword, is reorganized. For example, "cervical lymph nodes#swelling", "multiple#cervical lymph nodes#swelling", etc., thus arriving at the set of candidate standardized terms. , which is regarded as a term or phrase with added noise. Among them, is the nth candidate standardized term in
[0076] Preferably, considering that the keywords obtained in step 110 may have problems such as typos, synonyms, abbreviations, numbers, hyphens, etc., which will affect the accuracy of keyword matching. Therefore, before step 121, a rule-based method can be first used to process the keywords, such as the edit distance algorithm, the synonym comparison table, the abbreviation comparison table, the number conversion algorithm, the character processing function, etc., to improve the accuracy of subsequent keyword retrieval and matching.
[0077] The method provided by the embodiment of the present invention performs keyword transformation through the results of keyword retrieval and matching, and reorganizes the transformed keywords to obtain candidate standardized terms, providing text conditions for subsequent term generation, making the text generation interpretable, and effectively overcoming the "hallucination" problem of the model.
[0078] Based on any of the above embodiments, taking the candidate standardized term as the conditional text, generating the target standardized term, that is, step 130 specifically includes: Step 131, based on the trained diffusion model, taking the candidate standardized term as the conditional text, generating the target standardized term that meets the format requirements.
[0079] Specifically, the generation of the target standardized term can be achieved through the trained diffusion model. The diffusion model is a generative model based on deep learning technology. It has undergone a large amount of data training, can learn the latent distribution of the data, and generate new data samples similar to the training data. The trained diffusion model has the ability to generate high-quality target standardized terms that meet the requirements based on the given conditional text, that is, the candidate standardized term.
[0080] Terms (such as clinical medical terms) are a specialized and scientific representation form, with strict disciplinary connotation information, so they are the result of highly condensed information. In the practice of natural language research, especially in the field of dense representation, a method worth learning from is the diffusion model. Because the diffusion model generates the target through a multi-step denoising method, and by introducing the idea of multi-step fusion of multivariate information in the multi-step diffusion process, the vector representation ability of the entire vector recall can be improved.
[0081] Based on this, in this embodiment, a text-driven form is used to generate target standardized terms or phrases through a diffusion model. That is, given an input text (i.e., the source sentence or phrase) , being the L-th character in , with the goal of maximizing the conditional probability , being the T-th character in . Specifically, the objective function can be expressed as , where represents the parameters of the diffusion model, represents the conditional probability of generating the target text y given the input text
[0082] Figure 4 is one of the flow diagrams of the implementation manner of step 130 in the term standardization method provided by the present invention. As Figure 4 shown, the source sentence or phrase can be a candidate standardized term. The candidate standardized term is input into the diffusion model of the encoder-decoder architecture to obtain the target standardized term.
[0083] Based on any of the above embodiments, based on the trained diffusion model, with the candidate standardized term as the conditional text, a target standardized term that meets the format requirements is generated, specifically including: Step 131-1, based on the diffusion model, map the candidate standardized term to the noise space to obtain a noisy term, and perform conditional encoding on the noisy term to obtain the noisy term feature; Step 131-2, denoise the noisy term feature to obtain the denoised term feature, and decode the denoised term feature to obtain the target standardized term.
[0084] Specifically, Figure 5 is the second flow diagram of the implementation manner of step 130 in the term standardization method provided by the present invention. As Figure 5 shown, the diffusion model includes an input layer, a model layer, and an output layer.
[0085] The input layer is the candidate standardized term obtained in step 120. For each candidate standard term, each part of it is randomly masked , and each position corresponds to a keyword. Among them, M is the mask, being
[0086] The main body of the model layer is a diffusion model, and its calculation method follows the general method of diffusion models, that is, the forward process represents gradually perturbing the data sample at the initial moment with random noise , and the reverse process relies on a denoising network to gradually remove the random noise until the desired data sample. Among them, is the data sample at the (t - 1)th moment, is the data sample at the tth moment. As Figure 5 shown, the data sample at the Tth moment is , the data sample at the tth moment is , the data sample at the (t - 1)th moment is , and the data sample at the initial moment is . It can be understood that Figure 5 each data sample in it is a candidate standard term
[0087] That is, based on the diffusion model, the candidate standardized terms are mapped to the noise space to obtain the noisy terms, and the noisy terms are conditionally encoded to obtain the noisy term features
[0088] Based on the reparameterization trick, sampling can be performed from , as follows: where
[0089] Among them, is the probability distribution from to , is the identity matrix, N is the Gaussian distribution , , is a noise scale. In this way, the model can be effectively optimized during the training process and high-quality data samples can be generated. During the inference process, the reverse process samples noise from the Gaussian distribution , and performs iterative denoising through until is obtained. Among them, is the data sample at the Tth moment
[0090] Next, denoising processing is performed on the noisy term features. This is usually achieved through a reverse diffusion process, in which noise is gradually removed from the noise features to restore some features of the original terms. Finally, the denoised features are decoded to obtain the target standardized terms. The decoder network can convert the denoised features back into a readable standardized term form
[0091] To adapt it to the task of standardizing clinical medical terms, a transformation matrix is designed , and there is used to perturb the data samples. The forward process is as follows:
[0092] Among them, and are both represented by a one-hot vector encoding. is a categorical distribution with respect to and . can be expressed as:
[0093] Among them, . By using Bayesian theory, the posterior probability is calculated as follows:
[0094] Among them, represents element-wise multiplication. After that, the objective loss function of the diffusion model can be calculated by the KL divergence (Kullback-Leibler divergence) between the cumulative posterior probability and each component of the reverse process .
[0095] The method provided by the embodiments of the present invention generates target standardized terms that meet the format requirements through a conditional diffusion model. Through a multi-step diffusion process, the idea of multi-step fusion of multivariate information is introduced, which can improve the vector representation ability of the entire vector recall and increase the coverage rate of term standardization. In particular, by using a discrete text generation diffusion model, a candidate standardized term is regarded as a noisy input short sentence, and by setting the form of the term to be generated, such as the form of "term + result", the adaptability of term standardization is improved.
[0096] Based on any of the above embodiments, screening the target standardized terms to obtain the final standardized terms, that is, step 140 specifically includes: Scoring and ranking the target standardized terms based on the correlation between the target standardized terms and the reference entries, and screening the target standardized terms based on the ranking results to obtain the final standardized terms; The reference entries include at least one of the original text, conversion keywords, and terms to be standardized.
[0097] Specifically, after obtaining the target standardized term, it is possible to evaluate whether the target standardized term is the final standardized term from at least one level of the original text, conversion keywords, and terms to be standardized.
[0098] Construct a relevance evaluation system for quantifying the relevance between the target standardized term and the reference entry. The evaluation system can be constructed based on statistical methods (such as correlation coefficients), machine learning algorithms (such as classifiers), or natural language processing techniques (such as semantic similarity calculation).
[0099] Using the relevance evaluation system, calculate the relevance score between each target standardized term and the reference entry. The higher the score, the stronger the relevance between the target standardized term and the reference entry, and the higher the possibility that the target standardized term can be used as the final standardized term; conversely, the lower the score, the weaker the relevance between the target standardized term and the reference entry, and the lower the possibility that the target standardized term can be used as the final standardized term.
[0100] According to the relevance scores, sort the target standardized terms. The sorting result can be used as the basis for subsequent screening and selection of the final standardized terms.
[0101] Based on any of the above embodiments, score and sort the target standardized terms based on the relevance between the target standardized term and the reference entry, specifically including: Based on the large language model, determine the text semantic relevance between the target standardized term and the original text; Based on the degree of inclusion between the target standardized term and the conversion keywords, determine the inclusion relevance; Based on the relevance between each character of the target standardized term and each character of the conversion keywords, determine the character relevance; Based on the semantic relevance between the target standardized term and the term to be standardized, determine the term semantic relevance; Based on at least one of the text semantic relevance, inclusion relevance, character relevance, and term semantic relevance, score and sort the target standardized terms.
[0102] Specifically, for the relevance between the target standardized term and the reference entry, it can be based on at least one of the text semantic relevance, inclusion relevance, character relevance, and term semantic relevance.
[0103] First is the text semantic relevance between the target standardized term T and the original text. Taking medical terms as an example, in order to enable the pre-trained model to make full use of prior knowledge, designed based on the idea of prompt engineering, design a Prompt, and use the iFlytek Spark Medical Large Model to determine the medical connotation, as follows: Suppose you are a medical expert. Given a medical term or phrase and a patient's question, determine whether the term or phrase is implied in the patient's question statement.
[0104] Please directly output "Yes" or "No".
[0105] Patient's question: Hello, doctor. I was recently diagnosed with cerebral infarction. I feel dizzy and my memory has decreased a bit. My heartbeat is sometimes irregular, and the doctor said it's not a rapid heart rate. I want to know how serious these problems are? My blood pressure has been 150 / 95 recently. I wonder if this value is high? Also, are there any lifestyle suggestions to help improve these symptoms? Do I need to take any medicine? Thank you! Medical term or phrase: Enlarged cervical lymph nodes Obviously, the output of this example is "No". That is, the text semantic relevance between the target standardized term and the original text is not within the threshold range, so this target standardized term is excluded.
[0106] For inclusion relevance, it can be determined based on the degree of inclusion between the target standardized term and the conversion keywords. Determine the degree of inclusion between the target standardized term T and the set of conversion keywords . Use the TF-IDF algorithm, which is an efficient algorithm for calculating feature weights and can be used to solve the problem of short text similarity. Here, it is used to calculate the relevance score between each conversion keyword and the target standardized term. First, calculate the weights of the conversion keywords in the generated list of target standardized terms, and then calculate the relevance scores between each keyword and each entry in the list of target standardized terms. If the set of conversion keywords has a relatively high relevance to a certain target term entry, then this term entry may be the final standardized term. This is the scoring from the perspective of keywords.
[0107] For character relevance, it can be determined based on the relevance between each character of the target standardized term and each character of the conversion keyword. The relevance is evaluated from the literal distance, and the Jaccard distance calculation method is used, that is, the Jaccard correlation coefficient. It mainly measures the similarity between individuals from the perspective of a single character, that is, literal measurement. Concatenate the set of conversion keywords in order to form an initial term to be standardized , and calculate the correlation coefficient between it and each character in the target standardized term .
[0108] In addition to considering from the semantic perspective, keyword weight perspective, and character-related perspective of the user question (original text), it is also necessary to consider the semantic relationship between the target standardized term and the term to be standardized. Here, RoBERTa-wwm-ext is used as the semantic judgment model. By training a binary classifier, it predicts whether the target standardized term and the term to be standardized can be mapped and gives a mapping confidence score.
[0109] Finally, after layers of screening, filtering, and calculation, after completing the scoring and ranking, the optimal K target standardized terms are selected for output.
[0110] The method provided by the embodiments of the present invention evaluates whether the target standardized term is the final standardized term from multiple levels, that is, from the perspective of the original text context, from the perspective of keyword coverage, from the perspective of character matching, and from the perspective of semantic similarity, effectively improving the screening accuracy of standardized terms. Especially for the characteristics of dense information in the medical field, it ensures the consistency of the user's expression intention.
[0111] Based on any of the above embodiments, Figure 6 is the second flow schematic diagram of the term standardization method provided by the present invention, as Figure 6 shown, the method includes: S1, Extract the term to be standardized and the keywords in the term to be standardized from the original text. Specifically, it includes: based on a large language model, perform semantic understanding on the original text and extract key information in the original text to obtain the term to be standardized; perform phrase grouping on the term to be standardized to obtain the keywords in the term to be standardized. Among them, the patient asks a question, enabling the large language model to obtain the original text of the question, and at the same time, the patient's condition can be recorded.
[0112] During the semantic understanding process, hidden entity words can also be supplemented. For example, the term to be standardized obtained can be "[neck] multiple enlarged lymph nodes". Through a component disassembling model, the term to be standardized can be disassembled. For example, "neck", "lymph nodes", "multiple", and "enlarged" can be obtained.
[0113] S2, Perform keyword retrieval and obtain candidate standardized terms based on the matching results. Specifically, it includes: based on the medical term association network, perform retrieval and matching on the keywords, perform keyword conversion based on the matched terms to obtain the converted keywords in the term to be standardized; recombine the converted keywords in the term to be standardized to obtain candidate standardized terms. For example, the candidate standardized terms can include "neck lymph nodes#enlarged", "multiple neck lymph nodes#enlarged", "lymph node enlargement#multiple in the neck", etc.
[0114] S3, using the candidate standardized terms as conditional text, generates the target standardized terms. The process of generating the target standardized terms is the process of conditional text generation. Specifically, it includes: based on the trained denoising diffusion model, mapping the candidate standardized terms to the noise space to obtain the noisy terms, and conditionally encoding the noisy terms to obtain the noisy term features; denoising the noisy term features to obtain the denoised term features, and decoding the denoised term features to obtain the target standardized terms. For example, the target standardized terms can be "# swollen cervical lymph nodes", "multiple swollen cervical lymph nodes", "multiple swollen cervical lymph nodes", etc.
[0115] S4, based on the candidate standardized terms and the terms to be standardized, the target standardized terms are screened to obtain the final standardized terms. Specifically, it includes: determining the text semantic relevance between the target standardized terms and the original text based on a large language model; determining the inclusion relevance based on the inclusion degree between the target standardized terms and the conversion keywords; determining the character relevance based on the correlation between each character of the target standardized term and each character of the conversion keyword; determining the term semantic relevance based on the semantic relevance between the target standardized term and the term to be standardized; scoring and sorting the target standardized terms based on at least one of text semantic relevance, inclusion relevance, character relevance and term semantic relevance.
[0116] Here, both text semantic relevance and term semantic relevance can be achieved through a semantic discriminant model. The semantic discriminant model is a BERT model, which is obtained by extracting an embedding vector (Embedding) and calculating the cosine similarity.
[0117] Both the inclusion correlation and the character correlation can be determined by calculating the literal distance, and the literal distance can be determined by using the Jaccard distance.
[0118] The target standardized terms are scored and comprehensively ranked to obtain the final standardized terms.
[0119] The method provided by the embodiment of the present invention utilizes the semantic understanding and key information extraction capabilities of a large language model, relies on keyword information retrieval, integrates the controlled text generation capabilities of a diffusion model, avoids omissions in the extraction of medical terminology information, realizes the mapping of core words first and then the reorganization, and uses a denoising diffusion model to generate text on the basis of extensive recall, and finally determines the semantic association. Compared with the original "key entity extraction + semantic discrimination model" model, it can effectively overcome the "hallucination" problem while increasing the vocabulary size of term recall through fine-grained conditional text generation, and standardizes terms based on a specific format, thereby improving the flexibility and accuracy of standardization.
[0120] The following describes the term standardization device provided by the present invention. The term standardization device described below can be correspondingly referred to the term standardization method described above.
[0121] Figure 7 is a schematic structural diagram of the term standardization device provided by the present invention. As Figure 7 shown, the device includes: A term extraction unit 710, configured to extract terms to be standardized and keywords in the terms to be standardized from the original text; A retrieval matching unit 720, configured to perform retrieval matching on the keywords and obtain candidate standardized terms based on the matching results; A term generation unit 730, configured to generate target standardized terms with the candidate standardized terms as conditional texts; A term screening unit 740, configured to screen the target standardized terms to obtain final standardized terms.
[0122] The device provided by the embodiment of the present invention, by means of keyword information retrieval and integrating the conditional text generation ability, avoids omission of term information extraction and improves the scale of the term recall library. By generating conditional control of the term format, it can thus handle complex and variable term forms.
[0123] In addition, by refining the term standardization process, the term standardization process becomes clearer. Through process stageization, the latest and more external solutions or knowledge can be introduced in each stage, which helps to combine specific scenarios for application practice and optimization, and improves the flexibility and accuracy of term standardization.
[0124] Based on any of the above embodiments, the term generation unit is specifically configured to: Based on the trained diffusion model, generate target standardized terms that meet the format requirements with the candidate standardized terms as conditional texts.
[0125] Based on any of the above embodiments, the term generation unit is specifically configured to: Based on the diffusion model, map the candidate standardized terms to the noise space to obtain noise-added terms, and perform conditional encoding on the noise-added terms to obtain noise-added term features; Denoise the noise-added term features to obtain denoised term features, and decode the denoised term features to obtain the target standardized terms.
[0126] Based on any of the above embodiments, the term extraction unit is specifically configured to: Based on a large language model, perform semantic understanding on the original text and extract key information in the original text to obtain the terms to be standardized; Phrase the to-be-standardized term to obtain the keywords in the to-be-standardized term.
[0127] Based on any of the above embodiments, the retrieval matching unit is specifically configured to: Based on the term association network, retrieve and match the keywords, and perform keyword conversion based on the matched terms to obtain the converted keywords in the to-be-standardized term; Recombine the converted keywords in the to-be-standardized term to obtain the candidate standardized term.
[0128] Based on any of the above embodiments, the term screening unit is specifically configured to: Based on the relevance between the target standardized term and the reference entries, score and rank the target standardized term, and screen the target standardized term based on the ranking result to obtain the final standardized term; The reference entries include at least one of the original text, the converted keywords, and the to-be-standardized term.
[0129] Based on any of the above embodiments, the term screening unit is specifically configured to: Based on the large language model, determine the text semantic relevance between the target standardized term and the original text; Based on the inclusion degree between the target standardized term and the converted keywords, determine the inclusion degree relevance; Based on the relevance between each character of the target standardized term and each character of the converted keywords, determine the character relevance; Based on the semantic relevance between the target standardized term and the to-be-standardized term, determine the term semantic relevance; Based on at least one of the text semantic relevance, the inclusion degree relevance, the character relevance, and the term semantic relevance, score and rank the target standardized term.
[0130] Figure 8 Illustrates a schematic diagram of the entity structure of an electronic device, such as Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute a term standardization method, which includes: extracting terms to be standardized and keywords in the terms to be standardized from the original text; performing a retrieval match on the keywords, and obtaining candidate standardized terms based on the match result; generating target standardized terms with the candidate standardized terms as conditional text; and screening the target standardized terms to obtain final standardized terms.
[0131] In addition, when the logic instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0132] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the term standardization method provided by the above-mentioned various methods. The method includes: extracting terms to be standardized and keywords in the terms to be standardized from the original text; performing a retrieval match on the keywords, and obtaining candidate standardized terms based on the match result; generating target standardized terms with the candidate standardized terms as conditional text; and screening the target standardized terms to obtain final standardized terms.
[0133] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the term standardization method provided by the above-mentioned various methods. The method includes: extracting terms to be standardized and keywords in the terms to be standardized from the original text; performing a retrieval match on the keywords, and obtaining candidate standardized terms based on the matching result; generating target standardized terms with the candidate standardized terms as conditional texts; and screening the target standardized terms to obtain final standardized terms.
[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for term standardization, characterized in that, Comprising: Extracting the terms to be standardized and the keywords in the terms to be standardized from the original text; Performing a retrieval match on the keywords, and obtaining candidate standardized terms based on the match results; Generating target standardized terms with the candidate standardized terms as conditional texts; Screening the target standardized terms to obtain the final standardized terms.
2. The method for term standardization according to claim 1, wherein The generating target standardized terms with the candidate standardized terms as conditional texts includes: Based on the trained diffusion model, generating target standardized terms that meet the format requirements with the candidate standardized terms as conditional texts.
3. The terminology standardization method according to claim 2, wherein The generating target standardized terms that meet the format requirements with the candidate standardized terms as conditional texts based on the trained diffusion model includes: Based on the diffusion model, mapping the candidate standardized terms to the noise space to obtain the noise-added terms, and performing conditional encoding on the noise-added terms to obtain the noise-added term features; Denosing the noise-added term features to obtain the denoised term features, and decoding the denoised term features to obtain the target standardized terms.
4. The terminology standardization method according to any one of claims 1 to 3, characterized in that, The extracting the terms to be standardized and the keywords in the terms to be standardized from the original text includes: Based on the large language model, semantically understanding the original text and extracting the key information in the original text to obtain the terms to be standardized; Performing phrase grouping on the terms to be standardized to obtain the keywords in the terms to be standardized.
5. The method for term standardization according to any one of claims 1 to 3, characterized in that The performing a retrieval match on the keywords and obtaining candidate standardized terms based on the match results includes: Based on the term association network, performing a retrieval match on the keywords, and performing keyword conversion on the matched terms to obtain the converted keywords in the terms to be standardized; Recombining the converted keywords in the terms to be standardized to obtain the candidate standardized terms.
6. The method for term standardization according to any one of claims 1 to 3, characterized in that The screening the target standardized terms to obtain the final standardized terms includes: Based on the correlation between the target standardized terms and the reference entries, scoring and ranking the target standardized terms, and screening the target standardized terms based on the ranking results to obtain the final standardized terms; The reference entries include at least one of the original text, the converted keywords, and the terms to be standardized.
7. The method for term standardization according to claim 6, characterized in that, The scoring and ranking the target standardized terms based on the correlation between the target standardized terms and the reference entries includes: Based on the large language model, determining the text semantic correlation between the target standardized terms and the original text; Based on the inclusion degree between the target standardized terms and the converted keywords, determining the inclusion degree correlation; Based on the correlation between each character of the target standardized terms and each character of the converted keywords, determining the character correlation; Based on the semantic correlation between the target standardized terms and the terms to be standardized, determining the term semantic correlation; Based on at least one of the text semantic correlation, the inclusion degree correlation, the character correlation, and the term semantic correlation, scoring and ranking the target standardized terms.
8. A terminology standardization device, characterized in that, Comprising: A term extraction unit, configured to extract terms to be standardized and keywords in the terms to be standardized from the original text; A retrieval and matching unit, configured to perform retrieval and matching on the keywords and obtain candidate standardized terms based on the matching results; A term generation unit, configured to generate target standardized terms with the candidate standardized terms as conditional texts; A term screening unit, configured to screen the target standardized terms to obtain final standardized terms.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, the term standardization method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the term standardization method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Medical term processing method and device, computer equipment and storage medium
CN114153995A
Searching method and device based on keywords and semantics, equipment and storage medium
CN115438166A
A Medical Terminology Retrieval Method and System Based on Medical Semantic Understanding
CN116804998A
Diversity controllable text generation method and device based on diffusion model
CN118551735A
Term standardization method, electronic device, storage medium and computer program product
CN118798207A
Cited By
Multi-modal term normalization-based vertical class demand conversion method, device and equipment
CN121118840A
Matching method and device for medical terms
CN121434403A
A method and apparatus for matching medical terms
CN121434403B