A method, device and electronic device for identifying synonyms of traditional Chinese medicine
By constructing a Chinese medicine language dictionary and fine-tuning the large language model, synonyms and antonyms in Chinese medicine texts are solved, and the processing accuracy and depth are improved.
Patent Information
- Application Number
- CN202411000983.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-07-24
Smart Images

Figure CN118780279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent traditional Chinese medical technology, and in particular to a method, device and electronic equipment for identifying synonyms of traditional Chinese medicine. Background Art
[0002] In the context of the rapid development of artificial intelligence and modern medical technology, as a form of traditional medicine, the digital processing and application of TCM literature resources and clinical knowledge are particularly important. Traditional TCM texts contain a wealth of medical theories, clinical cases, and drug prescriptions. However, due to their unique terminology and professional expressions, the automated processing of TCM texts faces great challenges. At present, although there are a variety of Chinese natural language processing tools and corpora, these resources are mostly targeted at modern Chinese texts and lack support for TCM professional terminology. In response to this problem, some studies have attempted to build a dedicated TCM language resource library and use natural language processing technology to analyze and process TCM texts. For example, by building a TCM term dictionary and a classification system for symptoms and treatment methods, these methods have improved the accuracy and efficiency of TCM text processing to a certain extent. However, these early attempts usually rely on small data sets and limited technical means, and the understanding and application of TCM terminology are still relatively limited.
[0003] In addition, with the development of deep learning technology, large pre-trained language models such as BERT and GPT have performed well in a variety of natural language processing tasks. These models can capture the nuances and deep semantics of language through pre-training on large-scale corpora, providing a new technical approach for the processing of TCM texts. Nevertheless, these models usually lack understanding of specific fields and need to be fine-tuned in a targeted manner to better adapt to the language characteristics and task requirements of professional fields.
[0004] In the prior art, the patent document with publication number CN112632970A discloses a similarity scoring algorithm combining subject synonyms and word vectors. Among them, although the Word2Vec model used can effectively generate word vectors, its processing ability for new words, ancient words, dialects and low-frequency words is weak, and it may not be able to effectively capture the semantic features of these words. In addition, the Word2Vec model has limitations in processing contextual semantics, especially in understanding the multiple meanings of words in different contexts. It cannot meet the processing needs of traditional Chinese medicine texts. In addition, the algorithm relies on existing synonym libraries, which may lead to inaccurate recognition of synonyms. These thesaurus are usually statically defined, difficult to adapt to the dynamic changes of language and professional needs in specific fields, and may lead to inaccurate synonym matching, especially in professional fields. In addition, the algorithm lacks sufficient flexibility and scalability in design, for example, it is not adaptable enough when processing interdisciplinary texts or complex scenarios that require the combination of multiple semantic information. It is impossible to segment or punctuate ancient texts, and it also does not have the ability to understand long sentences composed of multiple words.
[0005] The patent document with publication number CN117195910A discloses a Chinese medical synonym discrimination method based on large model information enhancement. The scheme needs to classify the medical terms in the Chinese medical data set to form a set of medical term sample pairs, including positive sample pairs and negative sample pairs. The collection and classification of negative sample pairs require a lot of manual participation and data annotation, which increases the complexity and time cost of data preparation. In addition, the scheme uses a large model (such as ChatGPT) to enhance the data, then generates an enhanced data set, and then uses these enhanced data for model training and synonym discrimination. This method relies on multiple data processing and enhancement steps, which increases the computational complexity and processing time, and also requires additional computing resources to process and store enhanced data. In addition, when distinguishing synonyms, the scheme relies on using a large model to generate word meaning information for words, and then distinguishes by calculating the similarity of this information. Although this method can improve the accuracy of discrimination, it increases the amount of calculation and processing time, especially when facing large-scale data, which may significantly affect the processing efficiency. At the same time, the solution mainly focuses on synonym discrimination in modern Chinese medical texts, and does not specifically optimize or emphasize the processing of ancient books, classical Chinese, and long sentences. This may lead to certain limitations and deficiencies in processing complex historical documents and long sentences.
[0006] Therefore, although the existing technology has made some progress in medical text processing, the descriptions of medical terms vary greatly due to many factors such as the age, faction, and region of traditional Chinese medicine. Therefore, it is necessary to further develop a more accurate and efficient technical solution to fully explore and utilize traditional Chinese medicine text resources, identify the descriptions of synonyms or near synonyms, and accurately understand the expressions of different doctors in order to support the modernization and internationalization of medical care. Summary of the invention
[0007] In order to solve the problems existing in the prior art, the present invention provides the following technical solutions.
[0008] The first aspect of the present invention provides a method for identifying synonyms of traditional Chinese medicine, comprising:
[0009] S101, constructing a TCM language dictionary;
[0010] S102, using the TCM language dictionary, using a contrastive learning method and an InfoNCE loss function to fine-tune the large language model to obtain a fine-tuned large language model;
[0011] S103, using the fine-tuned large language model to vectorize each word in the TCM language dictionary to generate a high-dimensional vector space;
[0012] S104: Identify synonyms and antonyms of each word in the high-dimensional vector space.
[0013] Preferably, step S104 is followed by step S105: adding the recognition results of synonyms and antonyms to the TCM language dictionary; and step S106, repeatedly executing steps S102-S105 until the performance of the fine-tuned large language model meets the preset requirements.
[0014] Preferably, the constructing of the TCM language dictionary comprises: collecting and arranging commonly used terms and expressions in TCM professional literature, medical books and actual clinical cases to construct the TCM language dictionary.
[0015] Preferably, the step of vectorizing each word in the TCM language dictionary using the fine-tuned large language model comprises:
[0016] A special marker is added after each word in the TCM language dictionary to construct the corresponding input word;
[0017] Pass the constructed input vocabulary to the fine-tuned large language model, extract the vector of the special token, and output the contextual representation corresponding to the input vocabulary;
[0018] The vectors of special tokens are considered as semantic representations of the input vocabulary.
[0019] Preferably, identifying the synonyms and antonyms of each word in the high-dimensional vector space comprises:
[0020] For each word, the distance between it and other words is calculated in the high-dimensional vector space; the first several words with the closest distance are selected as synonyms; and the first several words with the farthest distance are selected as antonyms.
[0021] A second aspect of the present invention provides a device for identifying synonyms of traditional Chinese medicine, comprising:
[0022] Dictionary building module, used to build a TCM language dictionary;
[0023] A fine-tuning module, used to use the TCM language dictionary, a contrastive learning method and an InfoNCE loss function to fine-tune the large language model to obtain a fine-tuned large language model;
[0024] A vectorization module, used to vectorize each word in the TCM language dictionary using the fine-tuned large language model to generate a high-dimensional vector space;
[0025] The recognition module is used to recognize the synonyms and antonyms of each word in the high-dimensional vector space.
[0026] Preferably, the device for identifying synonyms in traditional Chinese medicine provided by the present invention also includes a feedback module: used to add the recognition results of synonyms and antonyms to the traditional Chinese medicine language dictionary; an optimization module, used to repeatedly call the fine-tuning module, the vectorization module, the recognition module and the feedback module until the performance of the large language model after fine-tuning meets the preset requirements.
[0027] A third aspect of the present invention provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method for identifying synonyms of traditional Chinese medicine as described in the first aspect.
[0028] The fourth aspect of the present invention provides an electronic device, comprising a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute the method for identifying synonyms of traditional Chinese medicine as described in the first aspect.
[0029] The beneficial effects of the present invention are as follows: the present invention provides a method, device and electronic device for identifying synonyms of traditional Chinese medicine. In the method, a traditional Chinese medicine language dictionary is first constructed; then, the traditional Chinese medicine language dictionary is used to fine-tune the large language model using a contrastive learning method and an InfoNCE loss function to obtain a fine-tuned large language model; then, each word in the traditional Chinese medicine language dictionary is vectorized using the fine-tuned large language model to generate a high-dimensional vector space; finally, in the high-dimensional vector space, the synonyms and antonyms of each word are identified. The above method is used to improve the accuracy of term processing, effectively identify doctor semantics, enhance the depth of semantic understanding, improve the recognition ability of synonyms and antonyms, increase the processing ability of complex sentences, and improve the processing ability of traditional Chinese medicine texts, providing a good foundation for the modernization and international development of traditional Chinese medicine. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a schematic diagram of the process of the method for identifying synonyms of traditional Chinese medicine according to the present invention;
[0031] Figure 2 This is a schematic diagram of the functional structure of the device for identifying synonyms in traditional Chinese medicine according to the present invention. DETAILED DESCRIPTION
[0032] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0033] The method provided by the present invention can be implemented in the following terminal environment, and the terminal may include one or more of the following components: a processor, a memory, and a display screen. The memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method described in the following embodiment.
[0034] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts of the entire terminal, and executes various functions of the terminal and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory.
[0035] The memory may include random access memory (RAM) or read-only memory (ROM). The memory may be used to store instructions, programs, codes, code sets or instructions.
[0036] The display screen is used to display the user interface of each application.
[0037] In addition, those skilled in the art can understand that the structure of the above terminal does not constitute a limitation on the terminal, and the terminal may include more or fewer components, or combine certain components, or arrange the components differently. For example, the terminal also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, and a power supply, which will not be described in detail here.
[0038] Embodiment 1
[0039] like Figure 1 As shown, an embodiment of the present invention provides a method for identifying synonyms of traditional Chinese medicine, including: S101, constructing a traditional Chinese medicine language dictionary; S102, using the traditional Chinese medicine language dictionary, using a contrastive learning method and an InfoNCE loss function to fine-tune a large language model to obtain a fine-tuned large language model; S103, using the fine-tuned large language model to vectorize each word in the traditional Chinese medicine language dictionary to generate a high-dimensional vector space; S104, in the high-dimensional vector space, identifying the synonyms and antonyms of each word.
[0040] Step S101 is executed to collect and organize commonly used terms and expressions in TCM professional literature, medical books and actual clinical cases to build a small but sufficient TCM language dictionary. This dictionary will serve as the basis for subsequent text processing to assist in identifying and understanding TCM terms.
[0041] Executing step S102 includes the following aspects:
[0042] 1. Selection of fine-tuning data: In the embodiments of the present invention, a smaller amount of data is used compared to the pre-training stage of the large language model to improve the performance of the large language model in specific tasks or fields.
[0043] 2. Adjustment of model structure: Freeze some key neurons or neural network layers so that the parameters of these parts remain unchanged during fine-tuning, and only optimize the parameters of the remaining parts.
[0044] 3. Optimization settings: Use a smaller learning rate for fine-tuning to prevent destroying the knowledge learned in the pre-training phase. Use the stochastic gradient descent (SGD) optimizer to optimize the objective function.
[0045] 4. Contrastive learning and InfoNCE loss function optimization: Synonymous pairs are used as positive samples for contrastive learning, and all other words in the same batch are used as negative samples. The InfoNCE loss function is calculated and optimized to minimize the objective function value.
[0046] In the embodiment of the present invention, the contrastive learning method and the InfoNCE loss function are used to fine-tune the dictionary of the large language model, so that the model can better capture the semantic relationship between TCM terms and improve the understanding and processing capabilities of TCM texts.
[0047] Executing step S103, using the fine-tuned large language model to vectorize each word in the traditional Chinese medicine language dictionary, can include: adding a special tag after each word in the traditional Chinese medicine language dictionary to construct a corresponding input vocabulary; passing the constructed input vocabulary to the fine-tuned large language model, extracting the vector of the special tag, and outputting the context representation corresponding to the input vocabulary; and regarding the vector of the special tag as the semantic representation of the input vocabulary.
[0048] This step can capture the semantic relationship between words through deep learning algorithms and generate high-quality word vectors.
[0049] Executing step S104, in the high-dimensional vector space, identifying the synonyms and antonyms of each word includes: for each word, in the high-dimensional vector space, calculating the distance between the word and other words; selecting the first several words with the closest distance as synonyms; and selecting the first several words with the farthest distance as antonyms.
[0050] This step is used to find the synonyms and antonyms of each word in the vector space, so as to deeply analyze the TCM text and understand the meaning, similarity and opposition. This function is very important for in-depth analysis of TCM text.
[0051] In a preferred embodiment of the present invention, after step S104, step S105 may be further included: adding the recognition results of synonyms and antonyms to the TCM language dictionary; and step S106, repeatedly executing steps S102-S105 until the performance of the fine-tuned large language model meets the preset requirements.
[0052] In one embodiment of the present invention, after fine-tuning the model, a sampling test for actual application can be performed. Through specific TCM case texts, the accuracy of the model in identifying synonyms and antonyms of TCM vocabulary and its effectiveness in actual TCM text processing are tested. The scoring mechanism will compare the output of the model with the standard answers of experts to evaluate the performance of the model. Based on the results of the sampling scoring, feedback from users and experts is collected to further adjust and optimize the model. This may include readjusting the parameters of the model, expanding or simplifying the dictionary content, or improving vectorization and search algorithms. The model is trained again to improve its accuracy and robustness in TCM text processing.
[0053] In practical applications, the method provided by the present invention can be used to process long sentences and ancient texts of traditional Chinese medicine texts. During the processing, for long sentences of traditional Chinese medicine texts containing multiple words, accurate word segmentation processing is first required. This step can ensure that the subsequent vectorization can accurately reflect the semantics and context of each word. Converting the word segmentation results into a mathematically processable vector form is particularly critical for capturing the deep semantics of the vocabulary. In order to understand the semantics from the perspective of the whole sentence, it is necessary to combine the vectors of a single word into a vector representation of the entire sentence. A sequence model can be used: the vector is further processed using a sequence processing model such as LSTM or Transformer. These models can consider the sequential relationship between words and better capture the overall semantics of the sentence. For some specific applications, it may be necessary to retain more original information. At this time, it is possible to choose to directly splice the vectors together, although this will result in a significant increase in dimensionality. The generated sentence vectors are used for specific application tasks, such as sentiment analysis, text classification, or semantic search. The efficiency and accuracy of the system in processing complex texts are ensured.
[0054] The method provided by the present invention has the following significant advantages in TCM text processing and analysis:
[0055] 1. Improved accuracy of term processing and effective recognition of doctor semantics: Through a language dictionary built specifically for the field of TCM, the present invention can more accurately recognize and parse TCM professional terms. This customized dictionary contains unique expressions and complex terms in the field of TCM, so it can effectively reduce misunderstandings and misinterpretations when processing professional texts.
[0056] 2. Enhanced depth of semantic understanding: Using a high-performance large language model for vocabulary vectorization enables the vector representation of each TCM vocabulary to capture deeper semantic information. Compared with traditional vectorization methods, it can better understand the use of vocabulary in specific TCM contexts, thereby improving the accuracy and effectiveness of text analysis.
[0057] 3. Improved recognition of synonyms and antonyms: The large language model is fine-tuned in all parameters using the contrastive learning method and the InfoNCE loss function. The fine-tuned large language model can significantly improve the quality of the TCM vocabulary vectorization process. This optimization ensures that the model can better distinguish between similar words and antonyms while maintaining the characteristics of the vocabulary itself, thereby improving the effectiveness and reliability of the model in practical applications. Therefore, the fine-tuned model can not only find accurate synonyms in the vectorization space, but also effectively identify antonyms. This is often difficult to achieve in traditional synonym libraries and simple model processing, especially in corpora that are not clearly marked. The present invention provides an effective solution strategy.
[0058] 4. Increased processing capabilities for complex sentences: For complex TCM texts, especially long sentences containing multiple words, the Tokenizer performs word segmentation and vectorization and then combines them to generate a comprehensive word vector for the entire sentence. This not only improves the processing depth of the text, but also makes it more accurate and comprehensive to extract information from long sentences.
[0059] The solution provided by the present invention significantly improves the performance of TCM text processing by combining advanced NLP technology and optimization specifically for the field of TCM, including improving accuracy, enhancing the depth of semantic understanding, optimizing word vector generation, improving the recognition ability of synonyms and antonyms, and increasing the processing ability of complex sentences. These advantages make the solution provided by the present invention show high practical value and technical innovation in the automated processing and analysis of TCM texts.
[0060] Embodiment 2
[0061] like Figure 2 As shown, another aspect of the present invention also includes a functional module architecture that is completely consistent with the aforementioned method flow, that is, an embodiment of the present invention also provides a device for identifying synonyms of traditional Chinese medicine, including: a dictionary construction module 201, used to construct a traditional Chinese medicine language dictionary; a fine-tuning module 202, used to use the traditional Chinese medicine language dictionary, adopt a contrastive learning method and an InfoNCE loss function, fine-tune the large language model, and obtain a fine-tuned large language model; a vectorization module 203, used to use the fine-tuned large language model to vectorize each word in the traditional Chinese medicine language dictionary to generate a high-dimensional vector space; an identification module 204, used to identify the synonyms and antonyms of each word in the high-dimensional vector space.
[0062] The device for identifying synonyms in traditional Chinese medicine provided by the present invention also includes a feedback module 205: used to add the recognition results of synonyms and antonyms to the traditional Chinese medicine language dictionary; an optimization module 206, used to repeatedly call the fine-tuning module, the vectorization module, the recognition module and the feedback module until the performance of the large language model after fine-tuning meets the preset requirements.
[0063] Furthermore, in the dictionary construction module, the construction of the TCM language dictionary includes: collecting and organizing commonly used terms and expressions in TCM professional literature, medical books and actual clinical cases to construct a TCM language dictionary.
[0064] Furthermore, in the vectorization module, the use of the fine-tuned large language model to vectorize each word in the traditional Chinese medicine language dictionary includes: adding a special tag after each word in the traditional Chinese medicine language dictionary to construct a corresponding input vocabulary; passing the constructed input vocabulary to the fine-tuned large language model, extracting the vector of the special tag, and outputting the context representation corresponding to the input vocabulary; and treating the vector of the special tag as the semantic representation of the input vocabulary.
[0065] Furthermore, in the recognition module, identifying the synonyms and antonyms of each word in the high-dimensional vector space includes: for each word, calculating the distance between it and other words in the high-dimensional vector space; selecting the first several words with the closest distance as synonyms; and selecting the first several words with the farthest distance as antonyms.
[0066] The device can be implemented by the method for identifying synonyms of traditional Chinese medicine provided in the above-mentioned embodiment 1. The specific implementation method can be found in the description of embodiment 1 and will not be repeated here.
[0067] The present invention also provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method for identifying synonyms of traditional Chinese medicine as described in the first embodiment.
[0068] The present invention also provides an electronic device, comprising a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute the method for identifying synonyms of traditional Chinese medicine as described in Example 1.
[0069] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for identifying synonyms of traditional Chinese medicine, characterized in that: include: S101, constructing a TCM language dictionary; S102, using the TCM language dictionary, using a contrastive learning method and an InfoNCE loss function to fine-tune the large language model to obtain a fine-tuned large language model; wherein the synonym pairs are used as positive samples for contrastive learning; the InfoNCE loss function is calculated and optimized to minimize the objective function value; S103, using the fine-tuned large language model to vectorize each word in the TCM language dictionary to generate a high-dimensional vector space; using the fine-tuned large language model to vectorize each word in the TCM language dictionary includes: adding a special tag after each word in the TCM language dictionary to construct a corresponding input vocabulary; passing the constructed input vocabulary to the fine-tuned large language model, extracting the vector of the special tag, and outputting the context representation corresponding to the input vocabulary; considering the vector of the special tag as the semantic representation of the input vocabulary; S104, in the high-dimensional vector space, identifying the synonyms and antonyms of each word includes: for each word, in the high-dimensional vector space, calculating the distance between the word and other words; selecting the first several words with the closest distance as synonyms; selecting the first several words with the farthest distance as antonyms; The step S104 also includes a step S105: adding the recognition results of synonyms and antonyms to the TCM language dictionary; and a step S106, repeatedly executing steps S102-S105 until the performance of the fine-tuned large language model meets the preset requirements.
2. The method for identifying synonyms of traditional Chinese medicine according to claim 1, characterized in that: The construction of the TCM language dictionary includes: collecting and arranging commonly used terms and expressions in TCM professional literature, medical books and actual clinical cases to construct the TCM language dictionary.
3. A device for identifying synonyms in traditional Chinese medicine, characterized in that: include: Dictionary building module, used to build a TCM language dictionary; A fine-tuning module is used to use the TCM language dictionary, a contrastive learning method and an InfoNCE loss function to fine-tune the large language model to obtain a fine-tuned large language model; wherein the synonym pairs are used as positive samples for contrastive learning; the InfoNCE loss function is calculated and optimized to minimize the objective function value; A vectorization module is used to vectorize each word in the TCM language dictionary using the fine-tuned large language model to generate a high-dimensional vector space; the vectorization of each word in the TCM language dictionary using the fine-tuned large language model includes: adding a special tag after each word in the TCM language dictionary to construct a corresponding input vocabulary; passing the constructed input vocabulary to the fine-tuned large language model, extracting the vector of the special tag, and outputting the context representation corresponding to the input vocabulary; treating the vector of the special tag as the semantic representation of the input vocabulary; The recognition module is used to recognize the synonyms and antonyms of each word in the high-dimensional vector space, including: for each word, in the high-dimensional vector space, calculating the distance between it and other words; selecting the first several words with the closest distance as synonyms; selecting the first several words with the farthest distance as antonyms; It also includes a feedback module: used for adding the recognition results of synonyms and antonyms to the Chinese medicine language dictionary; The optimization module is used to repeatedly call the fine-tuning module, the vectorization module, the recognition module and the feedback module until the performance of the large language model after fine-tuning meets the preset requirements.
4. A memory, characterized in that: A plurality of instructions are stored, and the instructions are used to implement the method for identifying synonyms of traditional Chinese medicine as described in any one of claims 1-2.
5. An electronic device, characterized in that: It includes a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute the method for identifying synonyms of traditional Chinese medicine as described in any one of claims 1-2.
Citation Information
Patent Citations
Similarity scoring algorithm combining subject synonyms and word vectors
CN112632970A
Chinese medical synonym discrimination method based on large model information enhancement
CN117195910A
Synonym screening method and system
CN107451126A
Online shopping assistant construction method for live marketing commodity quality perception analysis
CN116703509A
Training / application method of representation learning model, equipment and medium
CN118051774A