Traditional Chinese medicine term translation method and system based on natural language processing technology

Through the translation method of traditional Chinese medicine terminology based on natural language processing technology, traditional Chinese medicine terminology can be collected and updated in real time, deeply understand semantic changes, and accurately handle ambiguity, the problem of inability to update in real time and insufficient semantic understanding in the existing technology is solved, and efficient and accurate translation of traditional Chinese medicine terminology is achieved.

CN119990156AInactive Publication Date: 2025-05-13SAINS NEW MEDICAL COLLEGE OF GUANGXI UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065937.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology cannot update traditional Chinese medicine terms in real time, lack of semantic understanding, cannot accurately judge semantic changes, lack effective ambiguity analysis and processing mechanisms, and has low modularity, so it cannot flexibly respond to the translation needs of different types of traditional Chinese medicine terms.

Method used

The traditional Chinese medicine term translation method based on natural language processing technology is adopted, and the term text information in medical literature is collected in real time, pre-processing and semantic order arrangement is carried out, the degree of semantic change is judged, cluster analysis and feature extraction is carried out, and the term translation model and ambiguity analysis model are used to generate accurate translation results.

Benefits of technology

Real-time collection and update of traditional Chinese medicine terms, deeply understand semantic changes, accurately handle ambiguity, improve the accuracy and efficiency of translation, and meet the translation needs of different types of traditional Chinese medicine terms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990156A_ABST
    Figure CN119990156A_ABST
Patent Text Reader

Abstract

The invention provides a traditional Chinese medicine term translation method and system based on a natural language processing technology, and the method comprises the steps: collecting traditional Chinese medicine term text information in real time, preprocessing the text information, and converting the text information into a text fragment sequence; judging the semantic change degree of the next text fragment, if the semantic change degree is greater than a set value, performing clustering analysis to determine a semantic division type, and extracting feature information to form a matrix; the corresponding information analysis unit analyzes the matrix to obtain a translation result; if ambiguity exists, analyzing an ambiguous text fragment sequence to obtain a result, and generating accurate translation information; and obtaining user information and sending accurate translation information. According to the method, traditional Chinese medicine terms can be accurately and efficiently translated, the ambiguity problem is effectively solved, the translation quality is improved, and powerful support is provided for internationalized propagation of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and more specifically, to a method and system for translating traditional Chinese medicine terms based on natural language processing technology. Background Art

[0002] As the influence of traditional Chinese medicine continues to expand around the world, the accurate translation of traditional Chinese medicine terms is crucial to promoting the dissemination of traditional Chinese medicine culture and international exchanges. At present, the translation of traditional Chinese medicine terms mainly relies on manual translation and simple machine translation. Although manual translation can ensure high translation quality, it is inefficient and difficult to meet the translation needs of a large number of documents. When dealing with traditional Chinese medicine terms, the existing machine translation technology often results in inaccurate translation results and ambiguity due to the particularity and complexity of the terms, which cannot meet professional needs.

[0003] Existing machine translation technologies are usually based on general language models and lack an understanding of the professionalism and cultural connotations of TCM terminology. TCM terminology not only contains a wealth of medical knowledge, but also contains profound philosophical thoughts and cultural backgrounds. For example, Qi is a core concept in TCM, involving many aspects such as life activities and pathological changes, but existing translation technologies often have difficulty in accurately conveying its connotations. In addition, there are a large number of ancient Chinese expressions and professional terms in TCM literature. The translation of these terms requires deep professional knowledge and language skills, and existing technologies are obviously insufficient in this regard.

[0004] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: First, it is impossible to update and collect new terms in traditional Chinese medicine literature in real time, resulting in delayed translation content; second, the semantic understanding of traditional Chinese medicine terms is not deep enough, and it is impossible to accurately judge semantic changes, thus affecting the accuracy of translation; third, there is a lack of effective ambiguous analysis and processing mechanism, and when encountering polysemous terms, accurate translation results cannot be given; fourth, the existing translation system has a low degree of modularity and cannot flexibly respond to different types of traditional Chinese medicine term translation needs. Summary of the invention

[0005] The present invention provides a method and system for translating traditional Chinese medicine terms based on natural language processing technology.

[0006] In a first aspect of the present invention, a method for translating traditional Chinese medicine terms based on natural language processing technology is provided, comprising:

[0007] Real-time collection of TCM terminology text information in TCM literature;

[0008] Preprocessing the text information of traditional Chinese medicine terminology to obtain preprocessed text; and converting the preprocessed text into a sequence of text fragments arranged in a semantic order;

[0009] Based on the previous text segment, determine whether the semantic change degree of the next text segment is greater than a set degree value;

[0010] If yes, cluster analysis is performed on the next text segment to determine the semantic division type of the next text segment, and feature extraction model is used to extract feature information in the next text segment to form a feature information matrix;

[0011] According to the semantic division type, the corresponding information analysis unit in the term translation model is used to analyze the feature information matrix to obtain the term translation result;

[0012] When there is ambiguous information in the term translation result, an ambiguous text segment sequence of a set text length starting from the next text segment is obtained, and the ambiguous text segment sequence is analyzed using an ambiguity analysis model to obtain an ambiguity analysis result;

[0013] Generate accurate translation information based on the ambiguous analysis result; and obtain user information that needs translation of traditional Chinese medicine terms within a set range; and send accurate translation information to the user through the user information.

[0014] Furthermore, the method of judging whether the semantic change degree of the subsequent text segment is greater than a set degree value based on the previous text segment includes: converting the text segment into a vocabulary information set of corresponding parts of speech through part-of-speech tagging, and storing them in semantic order; obtaining the vocabulary information set of the previous text segment and the vocabulary information set of the subsequent text segment; calculating the semantic change degree value between the vocabulary information set of the subsequent text segment and the vocabulary information set of the previous text segment, and the calculation formula of the semantic change degree value is:

[0015]

[0016] in,

[0017] w i,后 is the frequency weight of the i-th word in the vocabulary information set of the next text segment,

[0018] w i,前 is the frequency weight of the ith word in the vocabulary information set of the previous text segment, and n is the number of words in the vocabulary information set; it is determined whether the semantic change degree value is greater than the set degree value.

[0019] Furthermore, the clustering analysis of the latter text segment to determine the semantic division type of the latter text segment includes: converting the latter text segment into an initial clustering data set via a preset converter; and analyzing the initial clustering data set using a clustering model constructed based on historical Chinese medicine literature data to obtain the semantic division type.

[0020] Furthermore, the extraction process of the feature extraction model includes: dividing the latter text segment into multiple irregular sub-segments; according to the semantic division type, screening out the sub-segments that meet the corresponding semantic division type from each sub-segment to obtain multiple key analysis sub-segments; using a feature extractor to extract feature information of each key analysis sub-segment to form multiple feature information matrices.

[0021] Furthermore, the processing process of the term translation model includes: inputting each feature information matrix into the corresponding information analysis unit according to the semantic division type; each information analysis unit judges whether the feature information matrix meets the set standards of the preset term knowledge base according to its own preset term knowledge base, and when it does not meet the set standards, obtains ambiguous sub-segments and obtains term translation results containing ambiguous sub-segment information.

[0022] Furthermore, the setting standards of the preset terminology knowledge base include a standard setting interval of each element in the feature information matrix and a set correlation between the element and surrounding elements.

[0023] Furthermore, the analysis process of the ambiguous analysis model includes: selecting segments containing ambiguous sub-segments from the ambiguous text segment sequence according to the ambiguous sub-segment information to obtain ambiguous segments; when the number of ambiguous segments exceeds a preset number, analyzing each ambiguous segment based on a preset ambiguous information library corresponding to the semantic division type to determine a first membership degree of each ambiguous segment and a second membership degree corresponding to the number of ambiguous segments;

[0024] The weighted average method is used to process the first membership and the second membership, and the average membership is calculated. The level corresponding to the average membership is used as the ambiguity level; the ambiguity analysis result is formed according to the semantic division type, ambiguity level and each ambiguous fragment.

[0025] Furthermore, the screening process for obtaining ambiguous fragments includes: based on the feature information matrix of the ambiguous sub-fragments, determining whether the feature information matrix converted from the ambiguous text fragment sequence contains a feature information matrix whose similarity exceeds a set similarity; if so, marking the ambiguous text fragment sequence as an ambiguous fragment.

[0026] Furthermore, the information analysis unit at least includes a traditional Chinese medicine analysis unit, a prescription analysis unit, a traditional Chinese medicine theory analysis unit, and a traditional Chinese medicine clinical application analysis unit.

[0027] In a second aspect of the present invention, a Chinese medicine term translation system based on natural language processing technology is provided, comprising:

[0028] The TCM terminology translation system comprises a data collection module, a data preprocessing module, a terminology translation module, an ambiguity analysis module, a data storage module and a translation output module; the TCM terminology translation system is used to implement the TCM terminology translation method through the mutual cooperation between the modules.

[0029] According to the above-mentioned embodiments of the present invention, at least the following beneficial effects are achieved: the method for translating TCM terminology of the present invention can collect terminology text information in TCM literature in real time, and can timely reflect the latest developments and terminology changes in the field of TCM through preprocessing and semantic order arrangement. During the translation process, the method can judge the degree of semantic change of the subsequent text segment based on the previous text segment. When the semantic change is significant, the semantic division type can be accurately determined and a feature information matrix can be formed through cluster analysis and feature extraction, and then the information analysis unit in the terminology translation model can be used for in-depth analysis to obtain accurate terminology translation results. In addition, when the translation result is ambiguous, the method can obtain a sequence of ambiguous text segments of a set length, and perform a detailed analysis through an ambiguous analysis model, and finally generate accurate translation information, effectively solving the ambiguity problem in the translation of TCM terminology.

[0030] Furthermore, the present invention also provides a TCM terminology translation system based on natural language processing technology, which includes a data acquisition module, a data preprocessing module, a terminology translation module, an ambiguity analysis module, a data storage module and a translation output module. The mutual cooperation between the modules can not only improve the translation efficiency, but also ensure the translation quality and meet the needs of different users for TCM terminology translation. The design of the system makes the TCM terminology translation work more intelligent and automated, and provides strong technical support for the international dissemination and academic exchanges of TCM. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which:

[0032] Figure 1 A flowchart of a method for translating Chinese medicine terms based on natural language processing technology provided by an embodiment of the present invention;

[0033] Figure 2 A schematic diagram of the structure of a traditional Chinese medicine term translation system based on natural language processing technology provided by an embodiment of the present invention;

[0034] Figure 3 The schematic diagram schematically shows the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.

[0036] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present invention may be implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0037] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0038] Reference below Figure 1 , Figure 1 A flowchart of a method for translating Chinese medicine terms based on natural language processing technology is provided in accordance with an embodiment of the present invention. Figure 1 As shown, a method 100 for translating Chinese medicine terms based on natural language processing technology includes:

[0039] Step 101, collecting text information of traditional Chinese medicine terms in traditional Chinese medicine literature in real time;

[0040] Step 102, preprocessing the TCM terminology text information to obtain a preprocessed text; and converting the preprocessed text into a sequence of text segments arranged in a semantic order;

[0041] Step 103, based on the previous text segment, determining whether the semantic change degree of the next text segment is greater than a set degree value;

[0042] Step 104: if yes, cluster analysis is performed on the next text segment to determine the semantic division type of the next text segment, and feature information in the next text segment is extracted using a feature extraction model to form a feature information matrix;

[0043] Step 105, according to the semantic division type, using the corresponding information analysis unit in the term translation model to analyze the feature information matrix to obtain the term translation result;

[0044] Step 106, when there is ambiguous information in the term translation result, an ambiguous text segment sequence of a set text length starting from the next text segment is obtained, and the ambiguous text segment sequence is analyzed using an ambiguity analysis model to obtain an ambiguity analysis result;

[0045] Step 107, generating accurate translation information according to the ambiguous analysis result; and obtaining user information requiring translation of traditional Chinese medicine terms within a set range; and sending accurate translation information to the user through the user information.

[0046] It should be noted that the present invention proposes a method for translating Chinese medicine terms based on natural language processing technology. The method first collects Chinese medicine terminology text information in Chinese medicine literature in real time. The Chinese medicine terminology text information here refers to professional terms and related descriptions in the field of Chinese medicine. These terms and descriptions are the basis for the dissemination and communication of Chinese medicine knowledge. Then, the collected Chinese medicine terminology text information is preprocessed. The purpose of preprocessing is to remove noise and redundant information in the text and improve the quality and readability of the text. The preprocessed text is converted into a sequence of text fragments arranged in a semantic order. The advantage of this is that the semantic structure and logical relationship in the text can be better captured, which facilitates subsequent semantic analysis and translation.

[0047] Specifically, the preprocessing process includes steps such as text cleaning, word segmentation, and part-of-speech tagging. Text cleaning mainly removes punctuation marks, special characters, and irrelevant information from the text. Word segmentation is to divide a continuous text string into independent vocabulary units, and part-of-speech tagging is to mark the part of speech of each vocabulary, such as noun, verb, adjective, etc. After these steps are completed, the text is divided into multiple text segments and arranged in semantic order to form a text segment sequence. When judging the degree of semantic change of the latter text segment, a threshold can be set. For example, when the semantic change degree value exceeds 0.5, it is considered that the semantics has changed significantly. The calculation of the semantic change degree value can be based on the word frequency weight of the vocabulary, and the semantic change can be quantified by comparing the word frequency weight difference of the same vocabulary in the previous and next text segments. When a semantic change is detected, the clustering analysis and feature extraction process are triggered. The clustering analysis can use the K-means algorithm to divide the text segments into different semantic categories, and the feature extraction can use the TF-IDF algorithm to extract the key feature vocabulary in each text segment to form a feature information matrix.

[0048] Preferably, in order to improve the accuracy of the judgment of the degree of semantic change, more semantic features can be introduced, such as part-of-speech sequences, dependency syntactic relations, etc. In cluster analysis, in addition to using the K-means algorithm, other clustering algorithms, such as hierarchical clustering, DBSCAN, etc., can also be tried to adapt to different types of text data. In the feature extraction link, in addition to the TF-IDF algorithm, word embedding technology, such as Word2Vec or BERT, can also be combined to extract richer semantic features. In addition, in the term translation model, special information analysis units can be constructed for different types of traditional Chinese medicine terminology, such as Chinese medicine names, prescription compositions, Chinese medicine theoretical concepts, etc., to improve the pertinence and accuracy of translation.

[0049] In some embodiments, judging whether the semantic change degree of the subsequent text segment is greater than a set degree value based on the previous text segment includes: converting the text segment into a vocabulary information set of corresponding parts of speech through part-of-speech tagging, and storing them in semantic order; obtaining the vocabulary information set of the previous text segment and the vocabulary information set of the subsequent text segment; calculating the semantic change degree value between the vocabulary information set of the subsequent text segment and the vocabulary information set of the previous text segment, and the calculation formula of the semantic change degree value is:

[0050]

[0051] in,

[0052] w i,后 is the frequency weight of the i-th word in the vocabulary information set of the next text segment,

[0053] w i,前 is the frequency weight of the ith word in the vocabulary information set of the previous text segment, and n is the number of words in the vocabulary information set; it is determined whether the semantic change degree value is greater than the set degree value.

[0054] It should be noted that the process of judging whether the degree of semantic change of the subsequent text segment mentioned in the present invention is greater than the set degree value, which is achieved by converting the text segment into a vocabulary information set of corresponding parts of speech through part-of-speech tagging, and storing it in semantic order. Part-of-speech tagging here refers to the grammatical classification of each word in the text, such as nouns, verbs, adjectives, etc., and the vocabulary information set refers to the vocabulary set obtained after part-of-speech tagging, which contains the vocabulary and its corresponding part-of-speech information. By calculating the semantic change degree value between the vocabulary information set of the subsequent text segment and the vocabulary information set of the previous text segment, the semantic difference between the text segments can be quantitatively evaluated, thereby judging whether the semantic change is significant.

[0055] More specifically, in practical applications, the word frequency weight can be determined by counting the frequency of a word in a text segment. For example, the more times a word appears in a text segment, the greater its word frequency weight. The set degree value can be adjusted according to actual needs. For example, for some translation tasks that are more sensitive to semantic changes, the set degree value can be set lower, such as 0.3; and for some tasks that are not so strict on semantic changes, the set degree value can be set higher, such as 0.7.

[0056] Preferably, in order to improve the accuracy of the judgment of the degree of semantic change, more semantic features can be considered to be introduced, such as the semantic similarity of words, context information, etc. When calculating the value of the degree of semantic change, in addition to using the word frequency weight, the semantic vector of the word can also be combined to measure the semantic relevance between words by calculating the cosine similarity between the semantic vectors.

[0057] Furthermore, a sliding window technique can be used to consider the local context information of words in a text segment to more accurately capture semantic changes. For example, a sliding window of size 5 can be set to calculate the word frequency weight and semantic vector of the words in the window to obtain a more refined semantic change degree value. At the same time, in order to adapt to different types of traditional Chinese medicine texts, the set degree value can be dynamically adjusted, and the size of the set degree value can be automatically adjusted according to factors such as the theme, style and complexity of the text to achieve more flexible and accurate semantic change judgment.

[0058] In some embodiments, the clustering analysis of the subsequent text segment to determine the semantic division type of the subsequent text segment includes: converting the subsequent text segment into an initial clustering data set via a preset converter; and analyzing the initial clustering data set using a clustering model constructed based on historical Chinese medicine literature data to obtain the semantic division type.

[0059] It should be noted that the process of clustering the latter text segment to determine its semantic division type in the present invention is achieved by converting the latter text segment into an initial clustering data set through a preset converter, and then analyzing it using a clustering model constructed based on historical Chinese medicine literature data. The preset converter here refers to a tool or algorithm that converts a text segment into a data format suitable for cluster analysis, and the initial clustering data set refers to a data set obtained after the conversion, which contains information such as feature vectors of the text segment. Through clustering model analysis, text segments can be classified into different semantic categories, thereby determining their semantic division types, which is crucial for subsequent feature extraction and translation analysis.

[0060] Specifically, the preset converter can use text vectorization technology, such as TF-IDF (Term Frequency-Inverse Document Frequency) or Word2Vec and other algorithms to convert text fragments into high-dimensional feature vectors. For example, when using the TF-IDF algorithm, the term frequency (TF) and inverse document frequency (IDF) of each word in the text fragment can be calculated to obtain the weight of each word, and then the weights of all words are combined into a feature vector. The clustering model can be trained based on historical Chinese medicine literature data, which can include a large number of Chinese medicine terms and their contextual information. During the training process, clustering algorithms such as K-means, hierarchical clustering or DBSCAN can be used to divide the text fragments into different semantic categories according to their feature vectors. For example, the number of cluster centers K in the K-means algorithm can be set to 10, indicating that the text fragments are divided into 10 main semantic categories.

[0061] Preferably, in order to improve the accuracy and efficiency of clustering analysis, the preset converter and clustering model can be optimized. In the preset converter, in addition to using TF-IDF or Word2Vec algorithms, advanced language models such as BERT (Bidirectional Encoder Representations from Transformers) can also be combined to obtain richer semantic features. In terms of clustering models, more complex clustering algorithms such as spectral clustering or Gaussian mixture model (GMM) can be used, which can better handle nonlinearly distributed data.

[0062] Furthermore, semi-supervised learning or unsupervised learning methods can be introduced to use a small amount of annotated TCM literature data to guide the clustering process and improve the accuracy and interpretability of clustering. For example, a small amount of annotated text fragments can be used as seed data, and the unannotated text fragments can be divided into corresponding semantic categories through a semi-supervised clustering algorithm. At the same time, in order to adapt to TCM literature data of different scales and complexities, the parameters of the clustering algorithm can be dynamically adjusted, such as the K value in the K-means algorithm, to automatically select the optimal number of cluster centers according to the distribution and characteristics of the data.

[0063] In some embodiments, the extraction process of the feature extraction model includes: dividing the latter text segment into multiple irregular sub-segments; according to the semantic division type, screening out the sub-segments that meet the corresponding semantic division type from each sub-segment to obtain multiple key analysis sub-segments; using a feature extractor to extract feature information of each key analysis sub-segment to form multiple feature information matrices.

[0064] It should be noted that the extraction process of the feature extraction model mentioned in the present invention involves dividing the latter text segment into multiple irregular sub-segments, and screening out the sub-segments that meet the corresponding semantic division type according to the semantic division type, thereby forming key analysis sub-segments. The irregular sub-segments here refer to smaller sub-units in the text segment divided according to certain rules or random methods, and these sub-segments may contain one or more words. The key analysis sub-segments refer to the sub-segments with representative and key information screened out under the guidance of the semantic division type, which are of great significance for subsequent feature extraction and translation analysis.

[0065] Specifically, the segmentation of text segments can be performed in a variety of ways, such as based on the co-occurrence frequency of words, grammatical structure or semantic relevance. During the segmentation process, a minimum segment length parameter, such as 3 words, can be set to ensure that the sub-segments have sufficient information. At the same time, a series of screening rules can be defined according to the semantic segmentation type. For example, if the semantic segmentation type is the name of a traditional Chinese medicine, the screening rules can include the word part (such as noun), specific prefixes or suffixes (such as grass, root, etc.), and contextual words related to traditional Chinese medicine (such as heat-clearing, detoxification, etc.). Through these rules, sub-segments that meet the corresponding semantic segmentation type are screened out from each sub-segment to obtain multiple key analysis sub-segments. The feature extractor can use traditional machine learning methods, such as support vector machines (SVM) or decision trees, or deep learning methods, such as convolutional neural networks (CNN) or recurrent neural networks (RNN), to extract feature information of each key analysis sub-segment to form multiple feature information matrices.

[0066] Preferably, in order to improve the accuracy and efficiency of feature extraction, the feature extraction model can be optimized. When segmenting text fragments, in addition to fixed rules, dependency syntactic analysis in natural language processing can also be introduced to divide sub-fragments according to the dependency relationship between words, so that the semantic structure can be better preserved. When screening key analysis sub-fragments, semantic similarity calculation can be combined, such as using Word2Vec or BERT model to calculate the semantic similarity between sub-fragments and semantic division types, and setting a similarity threshold, such as 0.7. Only when the similarity between the sub-fragment and the semantic division type exceeds the threshold, it is regarded as a key analysis sub-fragment. In terms of feature extraction, a multimodal feature extraction method can be used to combine text features and knowledge graph features to obtain more comprehensive semantic information. For example, for the name of traditional Chinese medicine, in addition to extracting lexical features in the text, it can also combine attribute information (such as efficacy, origin, etc.) in the knowledge graph of traditional Chinese medicine to form a richer feature information matrix. In addition, an attention mechanism can be introduced to give higher weights to important words in key analysis sub-fragments to highlight their importance in the feature information matrix.

[0067] In some embodiments, the processing process of the term translation model includes: inputting each feature information matrix into the corresponding information analysis unit according to the semantic division type; each information analysis unit judges whether the feature information matrix meets the set standards of the preset term knowledge base based on its own preset term knowledge base, and when it does not meet the set standards, obtains ambiguous sub-segments and obtains term translation results containing ambiguous sub-segment information.

[0068] It should be noted that the processing process of the term translation model in the present invention involves inputting each feature information matrix into the corresponding information analysis unit according to the semantic division type. The information analysis unit here refers to a module specially designed to analyze the feature information of a specific type of traditional Chinese medicine term, each of which contains a preset term knowledge base for judging whether the feature information matrix meets the preset standard. When the feature information matrix does not meet the preset standard, the model can identify the ambiguous sub-segment and generate a term translation result containing the ambiguous sub-segment information, which is crucial for processing polysemous terms and improving translation accuracy.

[0069] Specifically, the information analysis unit may include a Chinese medicine analysis unit, a prescription analysis unit, a Chinese medicine theory analysis unit, and a Chinese medicine clinical application analysis unit. The preset terminology knowledge base of each unit contains parameters such as the standard setting interval of each element in the feature information matrix and the set correlation between the element and the surrounding elements. For example, the knowledge base of the Chinese medicine analysis unit may contain the standard word frequency interval of the Chinese medicine name and the correlation threshold with the efficacy vocabulary. In practical applications, these parameters can be set based on the statistical analysis of Chinese medicine literature, such as determining the standard interval and correlation threshold by analyzing the frequency of occurrence and contextual association of specific terms in a large number of Chinese medicine literature.

[0070] Preferably, in order to further improve the accuracy and robustness of the term translation model, the information analysis unit and the preset term knowledge base can be optimized. In the information analysis unit, deep learning technology can be introduced, such as using a long short-term memory network (LSTM) or a Transformer model to analyze the feature information matrix to better capture the long-term dependencies in the sequence data. At the same time, the preset term knowledge base can be dynamically updated to adjust the standard setting interval and correlation threshold according to the latest research results and literature data of traditional Chinese medicine.

[0071] Furthermore, an integrated learning method can be used to integrate the judgment results of multiple information analysis units to improve the accuracy of ambiguous recognition. For example, when multiple units have inconsistent judgments on the same feature information matrix, the final translation result can be determined by a voting mechanism or a weighted average method. A user feedback mechanism can also be introduced to adjust the preset standards according to the user's translation satisfaction to achieve continuous optimization of the model.

[0072] In some embodiments, the setting standards of the preset terminology knowledge base include a standard setting interval of each element in the feature information matrix and a set correlation between the element and surrounding elements.

[0073] It should be noted that the setting standard of the preset terminology knowledge base mentioned in the present invention refers to a series of quantitative indicators used to evaluate whether the feature information matrix meets the requirements of specific Chinese medicine term translation. These indicators include the standard setting interval of each element in the feature information matrix and the set correlation between the element and the surrounding elements. The standard setting interval here refers to the reasonable numerical range that each feature element (such as word frequency, part of speech weight, etc.) should maintain during the translation process, and the set correlation refers to the semantic or grammatical correlation strength between the feature elements. These correlations can help judge the correctness and consistency of the terms in the context.

[0074] Specifically, the standard setting interval can be determined based on the statistical distribution of TCM terms in a large number of documents. For example, for a commonly used TCM name, its frequency of appearance in the text may have a common range, such as 5 to 15 times per thousand words. The correlation can be set by analyzing the co-occurrence frequency of the term with surrounding words. For example, if a TCM term often co-occurs with specific efficacy words (such as clearing heat and detoxification), the correlation between the two words can be set to a higher level, such as above 0.8. The setting of these parameters can be accomplished through data mining and analysis of a large number of TCM literature to ensure that they can accurately reflect the actual usage of TCM terms.

[0075] Preferably, in order to improve the accuracy and adaptability of the preset terminology knowledge base, these set standards can be updated regularly. The update can be done through an automated data mining process that monitors new trends and changes in TCM literature. For example, as new TCM research results are published, the frequency of use and contextual association of certain TCM terms may change.

[0076] Furthermore, an expert system can be introduced to allow TCM experts to review and adjust the set standards in the preset terminology knowledge base to ensure that they meet the latest TCM academic standards and clinical practices. It is also possible to consider using machine learning algorithms, such as reinforcement learning, to allow the system to automatically adjust the set standards based on user feedback and translation results to achieve self-optimization and evolution of the knowledge base.

[0077] In some embodiments, the analysis process of the ambiguity analysis model includes: selecting segments containing ambiguous sub-segments from the ambiguous text segment sequence according to the ambiguous sub-segment information to obtain ambiguous segments; when the number of ambiguous segments exceeds a preset number, analyzing each ambiguous segment based on a preset ambiguous information library corresponding to the semantic division type to determine a first membership degree of each ambiguous segment and a second membership degree corresponding to the number of ambiguous segments;

[0078] The weighted average method is used to process the first membership and the second membership, and the average membership is calculated. The level corresponding to the average membership is used as the ambiguity level; the ambiguity analysis result is formed according to the semantic division type, ambiguity level and each ambiguous fragment.

[0079] It should be noted that the analysis process of the ambiguous analysis model in the present invention aims to screen out segments containing ambiguous sub-segments from the ambiguous text segment sequence by analyzing the ambiguous sub-segment information, and determine the first membership and second membership of each ambiguous segment. The first membership here refers to the similarity between the ambiguous segment and a specific ambiguous category in the preset ambiguous information library, and the second membership refers to the membership corresponding to the number of ambiguous segments, which is used to measure the degree of ambiguity. By processing these two memberships by weighted average method, the average membership can be calculated, and then the ambiguity level can be determined, and finally the ambiguity analysis result is formed, which provides a basis for accurate translation.

[0080] Specifically, the ambiguous analysis model first selects the segments containing ambiguous sub-segments from the ambiguous text segment sequence according to the ambiguous sub-segment information. This process can be implemented by string matching or feature-based classification algorithm. For example, if the ambiguous sub-segment is a polysemous word, the model can search for all text segments containing the word. When the number of ambiguous segments selected exceeds the preset number, the model will analyze each ambiguous segment based on the preset ambiguous information library corresponding to the semantic division type. The preset ambiguous information library contains different ambiguous categories and their feature descriptions. The model determines the first membership of each ambiguous segment by calculating the feature similarity. At the same time, according to the number of ambiguous segments, the model can set a quantitative membership function, such as a linear or nonlinear function, to determine the second membership. The weights in the weighted average method can be adjusted according to actual needs. For example, if the first membership is considered more important, a higher weight can be given.

[0081] Preferably, in order to improve the accuracy and reliability of the ambiguity analysis, the ambiguity analysis model can be further optimized. When screening ambiguous fragments, in addition to simple string matching, a semantic understanding algorithm, such as BERT-based semantic similarity calculation, can be introduced to more accurately identify ambiguous fragments. When determining the membership, a more complex membership calculation method, such as fuzzy logic or probability model, can be used to better handle uncertainty and ambiguity.

[0082] Furthermore, a user feedback mechanism can be introduced to adjust the parameters and weights of membership calculation according to the user's satisfaction with the ambiguous analysis results. For example, if the user frequently corrects a certain type of ambiguous analysis result, the system can automatically adjust the relevant parameters to reduce the occurrence of similar errors. It is also possible to consider using deep learning methods, such as recurrent neural networks (RNNs) or their variants LSTM, to learn the sequence features of ambiguous fragments, so as to more accurately predict the ambiguity level.

[0083] In some embodiments, the screening process for obtaining ambiguous fragments includes: based on the feature information matrix of the ambiguous sub-fragments, determining whether the feature information matrix converted from the ambiguous text fragment sequence contains a feature information matrix whose similarity exceeds a set similarity; if so, marking the ambiguous text fragment sequence as an ambiguous fragment.

[0084] It should be noted that the screening process for obtaining ambiguous segments mentioned in the present invention is based on the feature information matrix of the ambiguous sub-segments, and determines whether the feature information matrix converted from the ambiguous text segment sequence contains a feature information matrix whose similarity exceeds the set similarity. The feature information matrix here refers to a matrix containing multi-dimensional information such as vocabulary features, part-of-speech features, and semantic features extracted from the text segment, and the set similarity is a threshold value for determining whether two feature information matrices are similar enough. When the similarity exceeds this threshold value, it can be considered that the corresponding text segment is ambiguous.

[0085] Specifically, the similarity calculation of the feature information matrix can adopt a variety of methods, such as cosine similarity, Jaccard similarity or Euclidean distance, etc. For example, when using cosine similarity, the dot product of the two feature information matrices can be calculated and divided by their respective moduli to obtain a similarity value between 0 and 1. The set similarity can be adjusted according to the ambiguity tolerance in actual applications. For example, if you want to identify ambiguity more strictly, you can set the set similarity to 0.8 or higher.

[0086] More specifically, the feature information matrix can include multiple features, such as word frequency, part of speech, contextual vocabulary, etc. The weights of these features can be adjusted according to the characteristics of TCM terminology to improve the accuracy of similarity calculation.

[0087] Preferably, in order to improve the accuracy and efficiency of ambiguous segment screening, the screening process can be further optimized. When calculating the similarity, more semantic features, such as word embedding vectors or semantic role annotations, can be introduced to more comprehensively capture the semantic information of the text segment. At the same time, machine learning methods, such as support vector machines (SVM) or random forests, can be used to train a classifier to automatically identify feature information matrices whose similarity exceeds a set threshold.

[0088] Furthermore, we can consider using deep learning methods, such as convolutional neural networks (CNN) or recurrent neural networks (RNN), to learn high-level feature representations of feature information matrices to more accurately judge similarity. We can also introduce a dynamic threshold adjustment mechanism to automatically adjust the set similarity according to the length, complexity, or domain specificity of the text fragment to adapt to different text analysis scenarios.

[0089] In some embodiments, the information analysis unit includes at least a traditional Chinese medicine analysis unit, a prescription analysis unit, a traditional Chinese medicine theory analysis unit, and a traditional Chinese medicine clinical application analysis unit.

[0090] It should be noted that the information analysis unit mentioned in the present invention at least includes a Chinese medicine analysis unit, a prescription analysis unit, a TCM theory analysis unit, and a TCM clinical application analysis unit. These analysis units are the core components of the term translation model and are specially designed to handle different types of TCM terms. The Chinese medicine analysis unit mainly targets terms such as the name of the Chinese medicine, efficacy, and medicinal properties; the prescription analysis unit focuses on terms such as prescription composition, incompatibility, usage and dosage; the TCM theory analysis unit covers terms such as basic theories of TCM, diagnostic methods, and treatment principles; and the TCM clinical application analysis unit involves terms such as clinical symptoms, treatment methods, and efficacy evaluation. These units analyze the feature information matrix through their respective preset term knowledge bases to ensure the accuracy and professionalism of the translation.

[0091] Specifically, the preset terminology knowledge base of each information analysis unit contains information related to the unit, such as a term list, grammatical rules, semantic associations, etc. For example, the knowledge base of the Chinese medicine analysis unit may contain information such as the Latin name, alias, and efficacy classification of Chinese medicine. The knowledge base of the prescription analysis unit may include the constituent drugs, dosage range, decoction method, etc. of the classic prescription. The construction of these knowledge bases can be completed by analyzing a large amount of Chinese medicine literature and organizing expert knowledge.

[0092] More specifically, in terms of parameter settings, different sensitivity thresholds can be set for each unit. For example, for the translation of Chinese medicine names, higher accuracy may be required, so a lower threshold can be set; while for some descriptive terms, the threshold can be appropriately relaxed. In addition, each unit can use different analysis algorithms, such as rule-based algorithms, machine learning algorithms, or deep learning algorithms, to adapt to the characteristics of different terms.

[0093] Preferably, in order to further improve the performance and adaptability of the information analysis unit, each unit can be refined and optimized. For example, in the Chinese medicine analysis unit, the analysis of pharmacodynamic components can be introduced to assist in the translation of terms related to pharmacodynamics by analyzing the content and effects of specific components in Chinese medicine. In the prescription analysis unit, the compatibility contraindications and usage and dosage information of the prescription can be updated in combination with the results of modern pharmacological research. In the TCM theory analysis unit, the interpretation of TCM philosophical thoughts can be introduced to help translate abstract concepts in TCM theory. In the TCM clinical application analysis unit, the translation of symptoms and treatment methods can be optimized in combination with clinical case data.

[0094] Furthermore, a dynamic update mechanism can be established to regularly update the knowledge base and analysis algorithms of each unit based on the latest research results and clinical practice data of traditional Chinese medicine, ensuring that the translation results always meet the latest academic standards and clinical needs.

[0095] The above-mentioned embodiments of the present invention have the following beneficial effects: the method for translating Chinese medicine terms based on natural language processing technology described in the present invention can collect terminology text information in Chinese medicine literature in real time, and ensure the timeliness and accuracy of the translation content. By preprocessing and arranging the text fragment sequence in semantic order, the method can effectively organize and prepare data, laying a solid foundation for subsequent translation analysis. When judging the degree of semantic change, it can accurately identify the significant change of semantics according to the set degree value, thereby triggering the clustering analysis and feature extraction process, which can not only determine the semantic division type of the text fragment, but also form a feature information matrix through the feature extraction model, provide detailed data support for the term translation model, and then obtain accurate term translation results. When there is ambiguity in the translation result, the method can obtain and analyze the ambiguous text fragment sequence, use the ambiguity analysis model to deeply analyze, and finally generate accurate translation information, effectively solve the ambiguity problem in the translation of Chinese medicine terms, and improve the reliability and professionalism of the translation.

[0096] Furthermore, the translation method can quantitatively evaluate the semantic differences between text fragments through a specific calculation formula for the degree of semantic change, providing a clear quantitative standard for the judgment of semantic changes. The detailed description of the cluster analysis and feature extraction model makes the establishment of semantic division types and the extraction of feature information more scientific and systematic, and improves the automation and intelligence level of the translation process. The processing process of the term translation model and the setting standard of the preset term knowledge base provide double guarantees for the accuracy and professionalism of the translation results. The analysis process of the ambiguity analysis model can accurately determine the ambiguity level through the calculation and processing of the degree of membership, and further optimize the translation results. In addition, the diversified settings of the information analysis unit cover multiple aspects such as traditional Chinese medicine, prescriptions, traditional Chinese medicine theory and clinical applications, so that the translation system can comprehensively and deeply process different types of traditional Chinese medicine terms to meet diverse translation needs. These detailed claims together constitute a comprehensive, efficient and accurate translation method for traditional Chinese medicine terms, providing strong technical support for the international dissemination of traditional Chinese medicine terms.

[0097] like Figure 2 As shown, in some embodiments, a TCM term translation system 200 based on natural language processing technology includes:

[0098] The TCM terminology translation system comprises a data collection module 201, a data preprocessing module 202, a terminology translation module 203, an ambiguity analysis module 204, a data storage module 205 and a translation output module 206; the TCM terminology translation system is used to implement the TCM terminology translation method through the mutual cooperation between the modules.

[0099] It is understandable that the modules and references recorded in the TCM terminology translation system 200 based on natural language processing technology Figure 1 Therefore, the operations, features and beneficial effects described above for the TCM terminology translation method based on natural language processing technology are also applicable to the TCM terminology translation system 200 based on natural language processing technology and the modules contained therein, and will not be repeated here.

[0100] Reference below Figure 3 , which shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic devices in some embodiments of the present invention may include but are not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0101] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0102] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0103] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.

[0104] The above descriptions are only some preferred embodiments of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention to form a technical solution.

Claims

1. A method for translating traditional Chinese medicine terms based on natural language processing technology, characterized in that: include: Real-time collection of TCM terminology text information in TCM literature; Preprocessing the text information of traditional Chinese medicine terminology to obtain preprocessed text; and converting the preprocessed text into a sequence of text fragments arranged in semantic order; Based on the previous text segment, determine whether the semantic change degree of the next text segment is greater than a set degree value; If yes, cluster analysis is performed on the next text segment to determine the semantic division type of the next text segment, and feature extraction model is used to extract feature information in the next text segment to form a feature information matrix; According to the semantic division type, the corresponding information analysis unit in the term translation model is used to analyze the feature information matrix to obtain the term translation result; When there is ambiguous information in the term translation result, an ambiguous text segment sequence of a set text length starting from the next text segment is obtained, and the ambiguous text segment sequence is analyzed using an ambiguity analysis model to obtain an ambiguity analysis result; Generate accurate translation information based on the results of ambiguity analysis; and obtain user information that requires translation of traditional Chinese medicine terms within a set range; Through the user information, accurate translation information is sent to the user.

2. The method for translating Chinese medicine terms according to claim 1, characterized in that: The method of judging whether the semantic change degree of the subsequent text segment is greater than a set value based on the previous text segment includes: converting the text segment into a vocabulary information set of corresponding parts of speech through part-of-speech tagging, and storing them in semantic order; obtaining the vocabulary information set of the previous text segment and the vocabulary information set of the subsequent text segment; calculating the semantic change degree value between the vocabulary information set of the subsequent text segment and the vocabulary information set of the previous text segment, and the calculation formula of the semantic change degree value is: in, w i,后 is the frequency weight of the i-th word in the vocabulary information set of the next text segment, w i,前 is the frequency weight of the ith word in the vocabulary information set of the previous text segment, and n is the number of words in the vocabulary information set; it is determined whether the semantic change degree value is greater than the set degree value.

3. The method for translating Chinese medicine terms according to claim 1, characterized in that: The clustering analysis of the latter text segment to determine the semantic division type of the latter text segment includes: converting the latter text segment into an initial clustering data set through a preset converter; and analyzing the initial clustering data set using a clustering model constructed based on historical Chinese medicine literature data to obtain the semantic division type.

4. The method for translating Chinese medicine terms according to claim 1, characterized in that: The extraction process of the feature extraction model includes: dividing the latter text segment into multiple irregular sub-segments; according to the semantic division type, screening out the sub-segments that meet the corresponding semantic division type from each sub-segment to obtain multiple key analysis sub-segments; using a feature extractor to extract feature information of each key analysis sub-segment to form multiple feature information matrices.

5. The method for translating Chinese medicine terms according to claim 4, characterized in that: The processing process of the term translation model includes: inputting each feature information matrix into the corresponding information analysis unit according to the semantic division type; each information analysis unit judges whether the feature information matrix meets the set standard of the preset term knowledge base according to its own preset term knowledge base, and when it does not meet the set standard, obtains ambiguous sub-segments and obtains term translation results containing ambiguous sub-segment information.

6. The method for translating Chinese medicine terms according to claim 5, characterized in that: The setting standards of the preset terminology knowledge base include the standard setting interval of each element in the feature information matrix and the setting correlation between the element and the surrounding elements.

7. The method for translating Chinese medicine terms according to claim 5, characterized in that: The analysis process of the ambiguous analysis model includes: selecting segments containing ambiguous sub-segments from the ambiguous text segment sequence according to the ambiguous sub-segment information to obtain ambiguous segments; when the number of ambiguous segments exceeds a preset number, analyzing each ambiguous segment based on a preset ambiguous information library corresponding to the semantic division type to determine a first membership degree of each ambiguous segment and a second membership degree corresponding to the number of ambiguous segments; The weighted average method is used to process the first membership and the second membership, and the average membership is calculated. The level corresponding to the average membership is used as the ambiguity level; the ambiguity analysis result is formed according to the semantic division type, ambiguity level and each ambiguous fragment.

8. The method for translating Chinese medicine terms according to claim 7, characterized in that: The screening process for obtaining ambiguous segments includes: based on the feature information matrix of the ambiguous sub-segments, determining whether the feature information matrix converted from the ambiguous text segment sequence contains a feature information matrix whose similarity exceeds a set similarity; if so, marking the ambiguous text segment sequence as an ambiguous segment.

9. The method for translating Chinese medicine terms according to claim 1, characterized in that: The information analysis unit at least includes a traditional Chinese medicine analysis unit, a prescription analysis unit, a traditional Chinese medicine theory analysis unit, and a traditional Chinese medicine clinical application analysis unit.

10. A Chinese medicine terminology translation system based on natural language processing technology, characterized in that: The Chinese medicine terminology translation system includes a data acquisition module, a data preprocessing module, a terminology translation module, an ambiguity analysis module, a data storage module and a translation output module; the Chinese medicine terminology translation system is used to implement the Chinese medicine terminology translation method as described in any one of claims 1-9 through the mutual cooperation between the modules.