Term standardization method and device, electronic equipment and storage medium

By extracting medical terms and their keywords from the original text, performing search matching and diffusion model generation, and finally screening out standardized terms, solving the problem of misunderstanding of intentions and communication barriers caused by inconsistency of medical terms, and improving the flexibility and accuracy of term standardization.

CN120104770AInactive Publication Date: 2025-06-06IFLYTEK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510051152.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Misunderstanding of intentions and communication barriers caused by inconsistency of medical terms in the prior art affects the sharing of medical data and the development of modern medical research.

Method used

A term standardization method includes extracting the term to be standardized and its keywords from the original text, performing a search match to obtain candidate normalized terms, using a diffusion model to generate target normalized terms that meet the format requirements, and screening them to obtain the final normalized terms.

Benefits of technology

It has improved the vocabulary scale of term recall, enhanced the flexibility and accuracy of term standardization, reduced medical errors, promoted the sharing of medical data and the development of modern medical research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104770A_ABST
    Figure CN120104770A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a term standardization method and device, electronic equipment and a storage medium, and the method comprises the steps: extracting to-be-standardized terms and keywords in the to-be-standardized terms from an original text; performing retrieval matching on the keywords, and obtaining candidate standardized terms based on a matching result; generating a target standardized term by taking the candidate standardized term as a condition text; and screening the target standardized terms to obtain final standardized terms. According to the term standardization method and device, the electronic equipment and the storage medium provided by the embodiment of the invention, term information extraction omission is avoided by means of keyword information retrieval and integration condition text generation capability, and the term recall word bank scale is improved. The formats of the terms are controlled by generating conditions, so that complex and changeable term forms can be processed, and the flexibility of term standardization is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a terminology standardization method, device, electronic device and storage medium. Background Art

[0002] Medical terminology refers to specialized words and phrases used to describe the human body, diseases, diagnosis, treatment, and medical research. It scientifically defines definitions, concepts, indicators, symptoms, etc. in medical-related fields to help medical professionals accurately communicate and understand medical information. It has basic characteristics such as professionalism, scientificity, unambiguity, and systematicity.

[0003] Different terms are used in different scenarios, different methods and different institutions in the medical process, which may lead to misunderstanding of intentions and communication barriers. For example, the three commonly used groups of related medical terms "symptoms, symptoms; signs, syndromes; indications, contraindications" composed of the three words "symptoms, signs, syndromes" are still incorrectly matched in many medical records, drug instructions, media news and related publications.

[0004] Therefore, it is urgent to ensure consistency in medical semantic expression and mutual understanding between different medical information systems through medical terminology standardization, reduce medical errors caused by terminology confusion, promote cross-regional and cross-institutional medical data sharing, support more extensive modern medical research and smart medical services, and improve the safety and quality of medical services. Summary of the invention

[0005] The present invention provides a terminology standardization method, device, electronic device and storage medium, which are used to solve the defects of misunderstanding of intention and communication barriers caused by inconsistent terminology in the prior art.

[0006] The present invention provides a terminology standardization method, comprising: Extracting to-be-standardized terms and keywords in the to-be-standardized terms from the original text; Searching and matching the keywords, and obtaining candidate standardized terms based on the matching results; Using the candidate standardized terms as conditional text, generating target standardized terms; The target standardized terms are screened to obtain final standardized terms.

[0007] According to the term standardization method provided by the present invention, the step of generating a target standardized term using the candidate standardized term as a conditional text includes: Based on the trained diffusion model, the target standardized terms that meet the format requirements are generated with the candidate standardized terms as conditional texts.

[0008] According to the terminology standardization method provided by the present invention, the method of generating target standardized terms that meet format requirements based on the trained diffusion model and taking the candidate standardized terms as conditional texts includes: Based on the diffusion model, the candidate standardized terms are mapped to the noise space to obtain noisy terms, and the noisy terms are conditionally encoded to obtain noisy term features; The noisy term features are denoised to obtain denoised term features, and the denoised term features are decoded to obtain the target standardized terms.

[0009] According to the terminology standardization method provided by the present invention, the step of extracting the terms to be standardized and the keywords in the terms to be standardized from the original text includes: Based on a large language model, semantic understanding is performed on the original text and key information in the original text is extracted to obtain the term to be standardized; The to-be-standardized terms are phrased to obtain keywords in the to-be-standardized terms.

[0010] According to the term standardization method provided by the present invention, searching and matching the keywords and obtaining candidate standardized terms based on the matching results include: Based on the term association network, the keywords are searched and matched, and the keywords are converted based on the matched terms to obtain the converted keywords in the term to be standardized; The conversion keywords in the to-be-standardized terms are reorganized to obtain the candidate standardized terms.

[0011] According to the terminology standardization method provided by the present invention, the target standardized terms are screened to obtain final standardized terms, including: Scoring and ranking the target standardized terms based on the correlation between the target standardized terms and the reference terms, and screening the target standardized terms based on the ranking results to obtain the final standardized terms; The reference entry includes at least one of the original text, conversion keywords, and the term to be standardized.

[0012] According to the terminology standardization method provided by the present invention, the scoring and ranking of the target standardized terms based on the correlation between the target standardized terms and the reference terms includes: Determining text semantic relevance between the target standardized term and the original text based on a large language model; Determining a degree of inclusion relevance based on a degree of inclusion between the target standardized term and the conversion keyword; Determining character correlation based on correlation between each character of the target standardized term and each character of the conversion keyword; Determining term semantic relevance based on the semantic relevance between the target standardized term and the term to be standardized; The target standardized terms are scored and ranked based on at least one of the text semantic relevance, the inclusion relevance, the character relevance, and the term semantic relevance.

[0013] The present invention also provides a terminology standardization device, comprising: A term extraction unit, used for extracting the terms to be standardized and the keywords in the terms to be standardized from the original text; A search and matching unit, used to search and match the keywords, and obtain candidate standardized terms based on the matching results; A term generation unit, used to generate a target standardized term using the candidate standardized term as a conditional text; The term screening unit is used to screen the target standardized terms to obtain final standardized terms.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any of the above-mentioned terminology standardization methods is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the terminology standardization method described in any one of the above is implemented.

[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the terminology standardization method described above is implemented.

[0017] The terminology standardization method, device, electronic device and storage medium provided by the present invention integrate the conditional text generation capability with the help of keyword information retrieval, avoid omissions in terminology information extraction, and increase the vocabulary size of terminology recall. By controlling the format of terms through generation conditions, it is possible to handle complex and changeable terminology forms, thereby enhancing the flexibility of terminology standardization.

[0018] In addition, the present invention makes the terminology standardization process clearer by refining the process. By dividing the process into stages, each stage can introduce the latest and more external solutions or knowledge, which helps to carry out application practice and optimization in combination with specific scenarios, thereby improving the flexibility and accuracy of terminology standardization. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 It is one of the flow charts of the terminology standardization method provided by the present invention.

[0021] Figure 2 It is a flowchart diagram of an implementation of step 110 in the terminology standardization method provided by the present invention.

[0022] Figure 3 It is a flowchart of the implementation of step 120 in the terminology standardization method provided by the present invention.

[0023] Figure 4 It is one of the flowchart diagrams of the implementation of step 130 in the terminology standardization method provided by the present invention.

[0024] Figure 5 This is the second flowchart diagram of the implementation of step 130 in the terminology standardization method provided by the present invention.

[0025] Figure 6 This is the second flow chart of the terminology standardization method provided by the present invention.

[0026] Figure 7 It is a structural schematic diagram of the terminology standardization device provided by the present invention.

[0027] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] Medical terminology standardization refers to the process of unifying medical terminology, aiming to ensure consistent and accurate communication between different medical systems and personnel. Standardization promotes information transfer, data sharing, medical research and cross-regional health cooperation. The task of medical terminology standardization is to convert informal medical expressions into formal medical concepts. For example, "persistent nasal bleeding" is mapped to "nosebleeding" and "lung abscess without pneumonia" is mapped to "lung abscess without pneumonia".

[0030] At present, according to the classification and summary of non-standard terms or phrases, to solve the problem of terminology standardization, we can try from 10 angles, including exact matching, abbreviation restoration, language component conversion, numerical replacement, hyphen parsing, suffix completion, synonym replacement, stem extraction, compound term splitting, partial matching, etc. From the perspective of algorithms, the existing methods for solving medical terminology standardization can be divided into three categories: methods based on rules and classical machine learning, methods based on deep learning and natural language understanding, and methods based on pre-trained models.

[0031] (1) Rule-based and classical machine learning methods.

[0032] From the perspective of rule design, we first analyze the formal characteristics of non-standard medical terms, classify medical terms or phrases according to sentence patterns, syntax, sentence patterns, etc., and design processing rules and methods based on 10 standardization perspectives. For example, by applying rules to extract the main body of a phrase or restore the original form of a phrase, establish a synonym library, use rules to identify synonyms in medical terms, and map them to unified standard terms. For abbreviations in medical terms, expand them into full names according to rules.

[0033] In terms of classic machine learning applications, we can use unlabeled data sets and unsupervised learning methods such as clustering and dimensionality reduction to discover the correlations and patterns between terms, thereby achieving standardization. We can also use labeled data sets and train models through supervised learning methods such as classification or regression to map non-standard terms to standard terms. We can also convert medical terms into feature representations that can be processed by machine learning algorithms, such as using bag-of-words models, TF-IDF, and other methods to perform term standardization.

[0034] (2) Methods based on deep learning and natural language understanding.

[0035] With the development of deep learning algorithms and technological breakthroughs in natural language understanding, a common strategy is to map medical terms to a continuous real vector space by training word vector models, such as the Word2Vec algorithm and the GloVe algorithm, to capture the semantic similarity between words. In addition, through the self-attention mechanism, the Transformer can effectively capture the dependencies between different positions in medical terms, which helps to better understand the semantic associations between terms.

[0036] Methods based on deep learning and natural language understanding rely on initial word embeddings and cannot represent the polysemy of a word. Discriminative models alone cannot obtain complete semantic information, and the features contained in word vectors are not rich enough to meet complex and specialized language understanding scenarios.

[0037] (3) Methods based on pre-trained models.

[0038] The BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained language model based on the Transformer architecture. It greatly improves the performance of natural language processing tasks by learning context information in both directions. In related technologies, the medical terminology standardization task is regarded as a translation task. In the first stage, a generative model is used to generate candidate entities. In the second stage, the BERT pre-trained model is used to sort the candidate entities by semantic similarity to obtain the final medical terminology standardization results. This scheme has achieved good results in the medical terminology standardization task. In another two-stage terminology standardization scheme, the first stage performs similarity recall based on traditional Jaccard, TF-IDF and other statistical methods. In the second stage, the RoBERTa-wwm-ext pre-trained language model is used to perform phrase pair text classification on the matching entities and candidate entities, and the classification results are used to determine whether the two medical terms can be aligned. This scheme has also achieved good results in practical applications.

[0039] However, due to data bias and model bias, "hallucination" problems may occur, especially the processing of rare terms is not good.

[0040] In view of the above problems, an embodiment of the present invention proposes a term standardization method, in which the terms to be standardized and the keywords in the terms to be standardized are first extracted from the original text, the keywords are searched and matched, and candidate standardized terms are obtained based on the matching results; the candidate standardized terms are used as conditional texts to generate target standardized terms; the target standardized terms are screened to obtain the final standardized terms. The method provided by the embodiment of the present invention integrates the conditional text generation capability with the help of keyword information retrieval, avoids omissions in term information extraction, and increases the vocabulary size of term recall.

[0041] In addition, the embodiments of the present invention refine the terminology standardization process to make the terminology standardization process clearer. By dividing the process into stages, each stage can introduce the latest and more external solutions or knowledge, which helps to carry out application practice and optimization in combination with specific scenarios, thereby improving the flexibility and accuracy of terminology standardization.

[0042] The embodiments of the present invention can be applied to scenarios where terminology standardization is required, such as in the medical field, mechanical engineering field, scientific field, education field, etc. The execution subject of the method can be an electronic device such as a terminal device, a computer, a server, a server cluster, or a specially designed terminology standardization device, or a terminology standardization device set in the electronic device, and the terminology standardization device can be implemented by software, hardware, or a combination of the two.

[0043] Figure 1 is one of the flow charts of the terminology standardization method provided by the present invention, such as Figure 1 As shown, this embodiment takes medical terminology standardization as an example to describe the terminology standardization method. The method includes the following steps: Step 110 , extracting the terms to be standardized and the keywords in the terms to be standardized from the original text.

[0044] Specifically, the original text refers to the text containing the terms to be standardized, which may be non-standard, irregular or ambiguous expressions. The original text may be, for example, a question text asked by a patient, a medical record text, or a diagnosis text, etc., which is not specifically limited in the embodiments of the present invention.

[0045] Terms to be standardized refer to words or phrases in the original text that need to be converted or replaced into more standard and standardized expressions.

[0046] The extraction of the terms to be standardized can be obtained from the original text by rule-based methods, statistical methods, pre-trained language model-based methods, etc. For example, in the medical field, rules can be set to identify drug names, disease names, etc., or to identify important words or phrases in the text and use them as terms to be standardized.

[0047] Taking the original text as the user's question text as an example, the original text is "Hello doctor, I was recently diagnosed with cerebral infarction, and I feel dizzy and have a little memory loss. My heartbeat is sometimes irregular, and the doctor said that my heart rate is not urgent. I want to know if these problems are serious? My blood pressure is recently 150 / 95. I wonder if this value is high? In addition, are there any lifestyle suggestions that can help improve these symptoms? Do I need to take any medicine? Thank you!".

[0048] Obviously, patients' expressions are not standardized, and they are accustomed to using colloquialisms and colloquialisms. For example, they express "cerebral infarction" as the non-standard "cerebral infarction", and mistakenly write "arrhythmia" as "irregular heart rate", and even make typos. Although doctors can understand the key intentions of patients, artificial intelligence or smart medical question-and-answer models may lead to misunderstandings, incorrect answers to medical advice, and affect user satisfaction. Therefore, it is necessary to standardize the terms in the original text.

[0049] Assume that the number of terms to be standardized contained in a user question (original text) is n, and the set Represents the set of terms to be standardized, and the goal is to convert all elements in the set into corresponding standard terms, that is, .

[0050] Here, the keywords in the term to be standardized are the core parts of the term to be standardized, which are usually words with specific meanings or functions. Keywords can be extracted from the term to be standardized through keyword extraction algorithms, domain dictionary matching, semantic analysis, etc.

[0051] For example, the term to be standardized is “multiple swollen lymph nodes in the neck”, in which the key words include “neck”, “lymph nodes”, “multiple”, and “swollen”.

[0052] Step 120, search and match the keywords, and obtain candidate standardized terms based on the matching results.

[0053] Specifically, after extracting the terms to be standardized and their keywords from the original text, the keywords are used to search and match in a specific terminology library, dictionary or database. This terminology library or database should contain a large number of standardized terms and their corresponding non-standard forms. Through search and matching, relevant keywords can be found.

[0054] Candidate standardized terms refer to possible standardized terms obtained based on keyword search matching, which need to be further screened and confirmed.

[0055] Step 130 , using the candidate standardized terms as conditional text, generates target standardized terms.

[0056] In this step, considering that the related art usually adopts a term standardization method based on a retrieval strategy, it may not be possible to standardize the terms not included in the library and it is difficult to handle complex and changeable term forms. In the embodiment of the present invention, based on the retrieval strategy, the target standardized terms are generated with the candidate standardized terms as conditional text. Through generative technology, it is possible to process terms that are not included in the library, thereby improving the scale of term recall and thus improving the coverage of term standardization. In addition, the generated target standardized terms are generated with the candidate standardized terms as conditional text, and the format of the terms is controlled by generating conditions, so that it is possible to handle complex and changeable term forms, such as abbreviations, synonyms, and format requirements specific to the medical field, thereby enhancing the flexibility of term standardization.

[0057] The generation of target standardized terms can be achieved through certain rules or algorithms, such as spelling check, grammar analysis, synonym replacement, etc. It can also be achieved through a trained text generation model, where the candidate standardized terms are input as conditional text into the trained text generation model to obtain the text sequence output by the text generation model, i.e., the target standardized terms. The target standardized terms are the expected standardized terms generated based on the candidate standardized terms through certain rules or algorithm processing.

[0058] The text generation model may be a sequence-to-sequence (Seq2Seq) model, a denoising diffusion model, a generative adversarial network, a Transformer model, etc., which is not specifically limited in the embodiment of the present invention.

[0059] Step 140, screening the target standardized terms to obtain final standardized terms.

[0060] Specifically, after the target standardized term is generated, it needs to be further screened and confirmed. The screening process can be automatically screened using machine learning algorithms. The screening criteria can include the accuracy, standardization, general acceptance, and contextual adaptability of the term. Through screening, it can be ensured that the final standardized term is accurate, standardized, and meets the requirements.

[0061] The method provided by the embodiment of the present invention integrates the conditional text generation capability with the help of keyword information retrieval, avoids omissions in term information extraction, and increases the vocabulary size of term recall. The format of terms is controlled by generating conditions, so that complex and changeable term forms can be processed.

[0062] In addition, the embodiments of the present invention refine the terminology standardization process to make the terminology standardization process clearer. By dividing the process into stages, each stage can introduce the latest and more external solutions or knowledge, which helps to carry out application practice and optimization in combination with specific scenarios, thereby improving the flexibility and accuracy of terminology standardization.

[0063] Based on any of the above embodiments, Figure 2 is a flow chart of an implementation of step 110 in the terminology standardization method provided by the present invention, such as Figure 2 As shown, the terms to be standardized and the keywords in the terms to be standardized are extracted from the original text, that is, step 110 specifically includes: Step 111, based on the large language model, semantic understanding is performed on the original text and key information in the original text is extracted to obtain the term to be standardized; Step 112, grouping the terms to be standardized to obtain keywords in the terms to be standardized.

[0064] Specifically, for the terms to be standardized and the keywords they contain, information extraction and term grouping can be achieved through a large language model.

[0065] The large language model here can be a pre-trained large language model, for example, it can include BERT, RoBERTa, RoBERTa-wwm, iFlytek Spark Medical Large Model, etc.

[0066] Semantic understanding means that the model can understand the text content, so as to perform information extraction tasks more accurately. First, use prompt engineering techniques to build a suitable prompt, set the granularity of entity extraction, and require the large model to combine the contextual information of the original text (such as the question text asked by the user) to understand the original text semantically and supplement the implicit medical entities. The extracted key information is sorted into a list of terms to be standardized.

[0067] On this basis, the standardized terms are grouped into words. Grouping can be achieved through component decomposition models, such as named entity recognition, conditional random fields, etc., to group the extracted content into words (you can also continue to use the large language model to complete this task, in which case you need to design a new prompt based on the task objectives and requirements).

[0068] After the word grouping, the keyword set in the term to be standardized is obtained, which can be expressed as ,in Indicates the nth term to be standardized after the term is phrased. Keywords, Indicates the size of the keyword list obtained by phrase-grouping each different term to be standardized.

[0069] The method provided in the embodiment of the present invention performs semantic understanding of the original text and extracts key information based on a large language model, and then performs phraseology to obtain keywords. With the help of the deep semantic understanding ability of the large language model, the terms to be standardized can be accurately extracted, thereby improving the accuracy and efficiency of information extraction.

[0070] Based on any of the above embodiments, Figure 3 is a flow chart of an implementation of step 120 in the terminology standardization method provided by the present invention, such as Figure 3 As shown, the keywords are searched and matched, and candidate standardized terms are obtained based on the matching results, that is, step 120 specifically includes: Step 121, searching and matching keywords based on the term association network, performing keyword conversion based on the matched terms, and obtaining conversion keywords in the terms to be standardized; Step 122, reorganize the conversion keywords in the terms to be standardized to obtain candidate standardized terms.

[0071] Specifically, multiple groups of keyword lists of different lengths may be obtained in step 110. Candidate standardized terms are then obtained based on the keyword search strategy.

[0072] Here, the term association network is a pre-built complex network structure, in which nodes represent terms and edges represent the relationship between terms. Such relationship can be synonyms, antonyms, hyponyms, etc.

[0073] Use the constructed term association network (such as the medical term association network) to search and match each group of keywords. If no matching standard term is found, the original keyword is retained. If a standard term matching the keyword is found, the keyword is replaced with the matched standard term. In this way, each group of keywords is converted into a new keyword list with almost no information loss. Here, the new keyword list is the converted keyword.

[0074] Convert the keywords in the term to be standardized into , get the conversion keyword set .

[0075] Then, each newly generated keyword group, i.e., the conversion keyword, is reorganized, such as "cervical lymph node # enlargement", "multiple #cervical lymph node # enlargement", etc., so that the candidate standardized term set is obtained. , treating it as a term or phrase with added noise.

[0076] Preferably, considering that the keywords obtained in step 110 may have typos, synonyms, abbreviations, numbers, hyphens, etc., which may affect the accuracy of keyword matching. Therefore, before step 121, the keywords may be processed using a rule-based method, such as an edit distance algorithm, a synonym comparison table, an abbreviation comparison table, a number conversion algorithm, a character processing function, etc., to improve the accuracy of subsequent keyword search matching.

[0077] The method provided by the embodiment of the present invention performs keyword conversion through the results of keyword retrieval and matching, and reorganizes the converted keywords to obtain candidate standardized terms, and then provides text conditions for subsequent term generation, making the text generation explainable, effectively overcoming the "hallucination" problem of the model.

[0078] Based on any of the above embodiments, the target standardized term is generated by taking the candidate standardized term as the conditional text, that is, step 130 specifically includes: Step 131 , based on the trained diffusion model, the candidate standardized terms are used as conditional texts to generate target standardized terms that meet the format requirements.

[0079] Specifically, the generation of target standardized terms can be achieved through a trained diffusion model. The diffusion model is a generative model based on deep learning technology. It has been trained with a large amount of data and can learn the potential distribution of data and generate new data samples similar to the training data. The trained diffusion model has the ability to generate high-quality target standardized terms that meet the requirements based on given conditional text, i.e., candidate standardized terms.

[0080] Terminology (such as clinical medical terminology) is a specialized and scientific representation with strict disciplinary connotation information, and is therefore a highly condensed result of information. In the practice of natural language research, especially in the field of dense representation, a method worth learning from is the diffusion model. Because the diffusion model generates targets through multi-step denoising, by introducing the idea of ​​multi-step fusion of multivariate information through the diffusion multi-step process, the vector representation ability of the entire vector recall can be improved.

[0081] Based on this, in this embodiment, a text-driven form is used to generate target standardized terms or phrases in a targeted manner through a diffusion model. That is, given a source sentence or phrase , to maximize the conditional probability For the target, generate the target sentence or phrase Specifically, the objective function can be expressed as ,in represents the parameters of the diffusion model, Represents a given input text The conditional probability of generating the target text y. Obviously, this is a conditional text generation mode from sequence to sequence, and the encoder-decoder architecture can be used.

[0082] Figure 4 is one of the flow charts of the implementation of step 130 in the terminology standardization method provided by the present invention, such as Figure 4As shown, the source sentence or phrase can be a candidate standardized term, and the candidate standardized term is input into the diffusion model of the encoder-decoder architecture to obtain the target standardized term.

[0083] Based on any of the above embodiments, based on the trained diffusion model, the target standardized term that meets the format requirements is generated with the candidate standardized term as the conditional text, specifically including: Step 131 - 1 , based on the diffusion model, mapping the candidate standardized terms to the noise space to obtain the noisy terms, and conditionally encoding the noisy terms to obtain the noisy term features; Step 131 - 2 , denoising the noisy term features to obtain denoised term features, and decoding the denoised term features to obtain target standardized terms.

[0084] Specifically, Figure 5 FIG. 2 is a flow chart of the implementation of step 130 in the terminology standardization method provided by the present invention. Figure 5 As shown, the diffusion model includes an input layer, a model layer and an output layer.

[0085] The input layer is the candidate standardized terms obtained in step 120. For each candidate standardized term, each part of it is randomly masked , each position corresponds to a keyword.

[0086] The main body of the model layer is a diffusion model, and the calculation method follows the general method of the diffusion model, that is, the forward process represents the step-by-step perturbation of data samples with random noise , the reverse process Relying on a denoising network Step by step, random noise is removed until the desired data sample is obtained.

[0087] That is, based on the diffusion model, the candidate standardized terms are mapped to the noise space to obtain the noisy terms, and the noisy terms are conditionally encoded to obtain the noisy term features.

[0088] Based on the reparameterization technique, we can Sampling any intermediate latent variable , as follows: in , , is a noise scale. In this way, the model can be effectively optimized during training and high-quality data samples can be generated. During inference, the reverse process is from a Gaussian distribution Sampling noise, through Perform iterative denoising until you get .

[0089] Next, the noisy term features are denoised. This is usually achieved through a back-diffusion process, where noise is gradually removed from the noisy features to recover some features of the original terms. Finally, the denoised features are decoded to obtain the target normalized terms. The decoder network is able to convert the denoised features back into a readable normalized term form.

[0090] In order to adapt it to the task of clinical medical terminology standardization, a transformation matrix is ​​designed To perturb the data 1,...,K, the forward process is as follows: Where x represents a one-hot vector encoding, is a categorical distribution about x. It can be expressed as: in, By using Bayesian theory, the posterior probability The calculation is as follows: in, represents element-wise multiplication. Afterwards, the objective loss function of the diffusion model can be accumulated by and The KL divergence (Kullback-Leibler divergence) between each component is calculated.

[0091] The method provided in the embodiment of the present invention generates target standardized terms that meet format requirements through a conditional diffusion model, and introduces the idea of ​​multi-step fusion of multivariate information through a diffusion multi-step process, which can improve the vector representation ability of the entire vector recall and improve the coverage of term standardization. In particular, the discrete text generation diffusion model is used to treat a candidate standardized term as a noisy input short sentence, and the adaptability of term standardization is improved by setting the form of the term to be generated, such as the form of "term + result".

[0092] Based on any of the above embodiments, the target standardized terms are screened to obtain final standardized terms, that is, step 140 specifically includes: Based on the correlation between the target standardized terms and the reference terms, the target standardized terms are scored and sorted, and based on the sorting results, the target standardized terms are screened to obtain the final standardized terms; The reference entry includes at least one of original text, conversion keywords, and terms to be standardized.

[0093] Specifically, after obtaining the target standardized term, it can be evaluated whether the target standardized term is the final standardized term from at least one level of the original text, the converted keywords, and the term to be standardized.

[0094] Construct a relevance evaluation system to quantify the relevance between the target standardized term and the reference term. The evaluation system can be constructed based on statistical methods (such as correlation coefficient), machine learning algorithms (such as classifiers), or natural language processing techniques (such as semantic similarity calculation).

[0095] The relevance evaluation system is used to calculate the relevance score between each target standardized term and the reference term. The higher the score, the stronger the relevance between the target standardized term and the reference term, and the higher the possibility that the target standardized term can be used as the final standardized term; conversely, the lower the score, the weaker the relevance between the target standardized term and the reference term, and the lower the possibility that the target standardized term can be used as the final standardized term.

[0096] The target standardized terms are sorted according to the relevance scores. The sorting results can be used as the basis for subsequent screening and selection of the final standardized terms.

[0097] Based on any of the above embodiments, scoring and ranking the target standardized terms based on the relevance between the target standardized terms and the reference terms specifically includes: Determine the textual semantic relevance between the target standardized term and the original text based on a large language model; Determine the degree of containment correlation based on the degree of containment between the target standardized term and the conversion keyword; determining character relevance based on a relevance between each character of the target standardized term and each character of the conversion keyword; Determining the semantic relevance of terms based on the semantic relevance between the target standardized terms and the terms to be standardized; The target standardized terms are scored and ranked based on at least one of text semantic relevance, inclusion relevance, character relevance, and term semantic relevance.

[0098] Specifically, the correlation between the target standardized term and the reference term may be based on at least one of text semantic correlation, inclusion correlation, character correlation, and term semantic correlation.

[0099] The first is the text semantic relevance between the target standardized term T and the original text. Taking medical terms as an example, in order to enable the pre-trained model to fully utilize prior knowledge, Prompt is designed based on the prompt word engineering concept, and the iFlytek Spark Medical Big Model is used to determine the medical connotation, as shown below: Suppose you are a medical expert. Given a medical term or phrase and a patient question, please determine whether the term or phrase is implied in the patient's question.

[0100] Please just input "yes" or "no".

[0101] Patient Question: Hello, doctor. I was recently diagnosed with cerebral infarction. I feel dizzy and have a little memory loss. My heartbeat is sometimes irregular, and the doctor said that my heart rate is not urgent. I want to know if these problems are serious? My blood pressure is 150 / 95 recently. I wonder if this value is high? In addition, are there any lifestyle suggestions that can help improve these symptoms? Do I need to take any medicine? Thank you! Medical term or phrase: Swollen lymph nodes in the neck Obviously, the output of this example is "No". That is, the text semantic relevance between the target standardized term and the original text is not within the threshold range, so the target standardized term is excluded.

[0102] For the inclusion relevance, it can be determined based on the inclusion degree between the target standardized term and the conversion keyword. The TF-IDF algorithm is used, which is an efficient algorithm for calculating feature weights. It can be used to solve the problem of short text similarity. Here, it is used to calculate the relevance score between each conversion keyword and the target standardized term. First, the weight of the conversion keyword in the generated target standardized term list is calculated, and then the relevance score of each keyword and each entry in the target standardized term list is calculated. If the keyword set If the correlation with a target term is relatively high, then the term is likely to be the final standardized term. This is the score from the keyword perspective.

[0103] For character relevance, it can be determined based on the correlation between each character of the target standardized term and each character of the conversion keyword. The Jaccard distance calculation method, i.e., the Jaccard correlation coefficient, is used to evaluate the relevance from the literal distance. It mainly measures the similarity between individuals from the perspective of a single word, i.e., the literal measurement. Splice them together in order to form an initial standard word , calculate it and the target normalized term The correlation coefficient between the characters in .

[0104] In addition to considering the semantics of user questions (original text), keyword weights, and character correlation, it is also necessary to consider the semantic relationship between the target standardized term and the term to be standardized. RoBERTa-wwm-ext is used as a semantic judgment model. By training a binary discriminator, it predicts whether the target standardized term and the term to be standardized can be mapped and gives a mapping credibility score.

[0105] Finally, after layers of screening, filtering and calculation, scoring and sorting, the optimal K target standardized terms are selected for output.

[0106] The method provided by the embodiment of the present invention evaluates whether the target standardized term is the final standardized term from multiple aspects, namely, from the perspective of the original text context, from the perspective of keyword coverage, from the perspective of character matching, and from the perspective of semantic similarity, effectively improving the screening accuracy of standardized terms. In particular, in view of the dense information in the medical field, the consistency of the user's expression intention is ensured.

[0107] Based on any of the above embodiments, Figure 6 This is the second flow chart of the terminology standardization method provided by the present invention, such as Figure 6 As shown, the method includes: S1, extracting the terms to be standardized and the keywords in the terms to be standardized from the original text. Specifically, it includes: based on a large language model, semantically understanding the original text and extracting key information from the original text to obtain the terms to be standardized; grouping the terms to be standardized to obtain the keywords in the terms to be standardized.

[0108] S2, search and match the keywords, and obtain candidate standardized terms based on the matching results. Specifically, it includes: searching and matching the keywords based on the term association network, converting the keywords based on the matched terms, and obtaining the converted keywords in the terms to be standardized; reorganizing the converted keywords in the terms to be standardized to obtain candidate standardized terms.

[0109] S3, using the candidate standardized terms as conditional text, generates the target standardized terms. Specifically, it includes: based on the trained diffusion model, mapping the candidate standardized terms to the noise space to obtain the noisy terms, and conditionally encoding the noisy terms to obtain the noisy term features; denoising the noisy term features to obtain the denoised term features, and decoding the denoised term features to obtain the target standardized terms.

[0110] S4, screening the target standardized terms to obtain the final standardized terms. Specifically, it includes: determining the text semantic relevance between the target standardized terms and the original text based on a large language model; determining the inclusion relevance based on the inclusion degree between the target standardized terms and the conversion keywords; determining the character relevance based on the correlation between each character of the target standardized terms and each character of the conversion keywords; determining the term semantic relevance based on the semantic relevance between the target standardized terms and the terms to be standardized; scoring and ranking the target standardized terms based on at least one of text semantic relevance, inclusion relevance, character relevance and term semantic relevance.

[0111] The method provided by the embodiment of the present invention utilizes the semantic understanding and key information extraction capabilities of a large language model, relies on keyword information retrieval, integrates the controlled text generation capabilities of a diffusion model, avoids omissions in the extraction of medical terminology information, realizes the mapping of core words first and then the reorganization, and uses a denoising diffusion model to generate text on the basis of extensive recall, and finally determines the semantic association. Compared with the original "key entity extraction + semantic discrimination model" model, it can effectively overcome the "hallucination" problem while increasing the vocabulary size of term recall through fine-grained conditional text generation, and standardizes terms based on a specific format, thereby improving the flexibility and accuracy of standardization.

[0112] The terminology standardization device provided by the present invention is described below. The terminology standardization device described below and the terminology standardization method described above can be referenced to each other.

[0113] Figure 7 is a schematic diagram of the structure of the terminology standardization device provided by the present invention, such as Figure 7 As shown, the device comprises: A term extraction unit 710, used to extract the terms to be standardized and the keywords in the terms to be standardized from the original text; A search and matching unit 720, configured to search and match the keywords and obtain candidate standardized terms based on the matching results; A term generation unit 730, configured to generate a target standardized term using the candidate standardized term as a conditional text; The term screening unit 740 is used to screen the target standardized terms to obtain final standardized terms.

[0114] The device provided by the embodiment of the present invention integrates the conditional text generation capability with the help of keyword information retrieval, avoids omissions in term information extraction, and increases the vocabulary size of term recall. The format of terms is controlled by generating conditions, so that complex and changeable term forms can be processed.

[0115] In addition, by refining the terminology standardization process, the terminology standardization process becomes clearer. By dividing the process into stages, each stage can introduce the latest and more external solutions or knowledge, which helps to combine application practice and optimization with specific scenarios, thereby improving the flexibility and accuracy of terminology standardization.

[0116] Based on any of the above embodiments, the term generation unit is specifically used for: Based on the trained diffusion model, the target standardized terms that meet the format requirements are generated with the candidate standardized terms as conditional texts.

[0117] Based on any of the above embodiments, the term generation unit is specifically used for: Based on the diffusion model, the candidate standardized terms are mapped to the noise space to obtain noisy terms, and the noisy terms are conditionally encoded to obtain noisy term features; The noisy term features are denoised to obtain denoised term features, and the denoised term features are decoded to obtain the target standardized terms.

[0118] Based on any of the above embodiments, the term extraction unit is specifically used for: Based on a large language model, semantic understanding is performed on the original text and key information in the original text is extracted to obtain the term to be standardized; The to-be-standardized terms are phrased to obtain keywords in the to-be-standardized terms.

[0119] Based on any of the above embodiments, the retrieval and matching unit is specifically used for: Based on the term association network, the keywords are searched and matched, and the keywords are converted based on the matched terms to obtain the converted keywords in the term to be standardized; The conversion keywords in the to-be-standardized terms are reorganized to obtain the candidate standardized terms.

[0120] Based on any of the above embodiments, the term screening unit is specifically used for: Scoring and ranking the target standardized terms based on the correlation between the target standardized terms and the reference terms, and screening the target standardized terms based on the ranking results to obtain the final standardized terms; The reference entry includes at least one of the original text, conversion keywords, and the term to be standardized.

[0121] Based on any of the above embodiments, the term screening unit is specifically used for: Determining text semantic relevance between the target standardized term and the original text based on a large language model; Determining a degree of inclusion relevance based on a degree of inclusion between the target standardized term and the conversion keyword; Determining character correlation based on correlation between each character of the target standardized term and each character of the conversion keyword; Determining term semantic relevance based on the semantic relevance between the target standardized term and the term to be standardized; The target standardized terms are scored and ranked based on at least one of the text semantic relevance, the inclusion relevance, the character relevance, and the term semantic relevance.

[0122] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the term standardization method, which includes: extracting the terms to be standardized and the keywords in the terms to be standardized from the original text; searching and matching the keywords, and obtaining candidate standardized terms based on the matching results; using the candidate standardized terms as conditional texts to generate target standardized terms; screening the target standardized terms to obtain final standardized terms.

[0123] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0124] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the terminology standardization method provided by the above-mentioned methods, which includes: extracting the terms to be standardized and the keywords in the terms to be standardized from the original text; searching and matching the keywords, and obtaining candidate standardized terms based on the matching results; generating target standardized terms with the candidate standardized terms as conditional text; and screening the target standardized terms to obtain the final standardized terms.

[0125] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the terminology standardization method provided by the above-mentioned methods, the method comprising: extracting the terms to be standardized and the keywords in the terms to be standardized from the original text; searching and matching the keywords, and obtaining candidate standardized terms based on the matching results; generating target standardized terms using the candidate standardized terms as conditional texts; and screening the target standardized terms to obtain final standardized terms.

[0126] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A terminology standardization method, characterized in that: include: Extracting to-be-standardized terms and keywords in the to-be-standardized terms from the original text; Searching and matching the keywords, and obtaining candidate standardized terms based on the matching results; Using the candidate standardized terms as conditional text, generating target standardized terms; The target standardized terms are screened to obtain final standardized terms.

2. The terminology standardization method according to claim 1, characterized in that: The step of generating a target standardized term using the candidate standardized term as a conditional text includes: Based on the trained diffusion model, the target standardized terms that meet the format requirements are generated with the candidate standardized terms as conditional texts.

3. The terminology standardization method according to claim 2, characterized in that: The method of generating target standardized terms that meet format requirements based on the trained diffusion model and taking the candidate standardized terms as conditional texts includes: Based on the diffusion model, the candidate standardized terms are mapped to the noise space to obtain noisy terms, and the noisy terms are conditionally encoded to obtain noisy term features; The noisy term features are denoised to obtain denoised term features, and the denoised term features are decoded to obtain the target standardized terms.

4. The terminology standardization method according to any one of claims 1 to 3, characterized in that: The step of extracting the terms to be standardized and the keywords in the terms to be standardized from the original text includes: Based on a large language model, semantic understanding is performed on the original text and key information in the original text is extracted to obtain the term to be standardized; The to-be-standardized terms are phrased to obtain keywords in the to-be-standardized terms.

5. The terminology standardization method according to any one of claims 1 to 3, characterized in that: The searching and matching of the keywords and obtaining candidate standardized terms based on the matching results include: Based on the term association network, the keywords are searched and matched, and the keywords are converted based on the matched terms to obtain the converted keywords in the term to be standardized; The conversion keywords in the to-be-standardized terms are reorganized to obtain the candidate standardized terms.

6. The terminology standardization method according to any one of claims 1 to 3, characterized in that: The step of screening the target standardized terms to obtain final standardized terms includes: Scoring and ranking the target standardized terms based on the correlation between the target standardized terms and the reference terms, and screening the target standardized terms based on the ranking results to obtain the final standardized terms; The reference entry includes at least one of the original text, conversion keywords, and the term to be standardized.

7. The terminology standardization method according to claim 6, characterized in that: Scoring and ranking the target standardized terms based on the correlation between the target standardized terms and the reference terms includes: Determining text semantic relevance between the target standardized term and the original text based on a large language model; Determining a degree of inclusion relevance based on a degree of inclusion between the target standardized term and the conversion keyword; Determining character correlation based on correlation between each character of the target standardized term and each character of the conversion keyword; Determining term semantic relevance based on the semantic relevance between the target standardized term and the term to be standardized; The target standardized terms are scored and ranked based on at least one of the text semantic relevance, the inclusion relevance, the character relevance, and the term semantic relevance.

8. A terminology standardization device, characterized in that: include: A term extraction unit, used for extracting the terms to be standardized and the keywords in the terms to be standardized from the original text; A search and matching unit, used to search and match the keywords, and obtain candidate standardized terms based on the matching results; A term generation unit, used to generate a target standardized term using the candidate standardized term as a conditional text; The term screening unit is used to screen the target standardized terms to obtain final standardized terms.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the terminology standardization method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the terminology standardization method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Matching method and device for medical terms

    CN121434403A