Traditional Chinese medicine text relation extraction method based on prompt learning
By introducing knowledge correction template construction and relationship auxiliary vectors in the text relationship extraction of traditional Chinese medicine, the problems of incomplete external knowledge base and high computational complexity in the existing technology are solved, efficient and accurate extraction of traditional Chinese medicine text relationships are achieved, and the process of informatization of traditional Chinese medicine is promoted.
Patent Information
- Application Number
- CN202510207333.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology has problems such as incomplete external knowledge base, uneven data distribution, large amount of calculations generated by prompt learning templates, and inconsistent length of relationship labels in traditional Chinese medicine text relationship extraction, which affects the accuracy and efficiency of relationship extraction.
A template construction model based on knowledge correction is proposed, which reduces manual overhead and reduces the computational amount of template generation by introducing the knowledge supplement vector of entities. At the same time, a relational auxiliary vector is introduced to reduce the complexity of manual answer word construction through the generalization of answer word vector space, and enhance the generalization ability of the model.
It realizes automatic labeling of entities and relationships in traditional Chinese medicine texts, improves labeling efficiency and accuracy, reduces the time and error of manual labeling, and promotes the process of informatization of traditional Chinese medicine.
Smart Images

Figure CN120218031A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of prediction models, and specifically relates to a method for extracting Chinese medicine text relationships based on prompt learning. Background Art
[0002] The extraction of entities such as symptoms, Chinese medicines, prescriptions, treatment methods, and causes of diseases in traditional Chinese medicine medical records, as well as their mutual relationships, is an important basis for information retrieval, knowledge graph construction, and intelligent question answering. The relationships and entities in traditional Chinese medicine medical record texts are densely distributed. Manual annotation is not only time-consuming but also prone to mislabeling and missing labels. For example, in the original disease chapter of Treatise on Febrile Diseases and Synopsis of Prescriptions of the Golden Chamber, there are more than a thousand entities and more than four thousand relationships in a text of six thousand characters. It is difficult and time-consuming to establish accurate mappings for each entity and relationship manually. Relationship extraction (RE) technology can automatically identify the relationships between entities in specific texts from a set of predefined relevant relationships, which can greatly reduce the manual annotation work.
[0003] Regarding the relationship extraction problem, researchers have proposed various solutions. Chen et al. introduced external knowledge to form an open database for retrieval, and thereby predicted the relationships between entities (Relation Extraction as Open-book Examination: Retrieval-enhanced Prompt Tuning). Hu et al. used a convolutional neural network to perform multi-level convolution on different components of the text to capture the context meaning (A hierarchical convolutional model for biomedical relation extraction).
[0004] Prompt learning pre-training, including template engineering and oral expression engineering, aims to find the best templates and answer spaces. Prompt learning can transform downstream tasks into text generation tasks by adding "prompt information" to the input without significantly changing the structure and parameters of the pre-trained language model. The specific method is to endow the pre-trained model with artificial rules so that the model can better understand human instructions and thus better apply the pre-trained model to downstream tasks. Recently, researchers have also found that pre-trained language models can also use prompt tuning to obtain better performance in few-shot learning tasks. Zhao et al. combined prompt learning with the nearest neighbor algorithm to select similar semantic labels through the nearest neighbor algorithm (Biomedical document relation extraction with prompt learning and KNN). Lin et al. jointly trained prompt learning with graph convolutional neural networks to capture multimodal information (Multimodal learning on graphs for disease relation extraction). III. Defects of the Background Art
[0005] In the methods of introducing external knowledge for relation extraction, some external knowledge bases are incomplete, resulting in the inability to retrieve task-related information, especially the external knowledge in many professional fields. Due to the low adaptability between the external knowledge base and the dataset, the accuracy of the relation extraction results is affected.
[0006] In the case of few-shot datasets, due to the uneven data distribution, the model has poor relation recognition performance under long-tail data and often introduces additional noise, reducing the generalization degree, thus affecting the performance of long-tail relation extraction.
[0007] Although prompt learning has achieved certain results in text classification tasks, there are still many challenges. On the one hand, determining appropriate prompt templates for relation extraction requires domain expertise, and automatically constructing high-performance prompts using input entities usually requires additional computational costs; on the other hand, when the lengths of relation labels are different, the computational complexity of the label word search process is very high, usually exponential in the number of categories, and it is not easy to find a suitable target label word in the vocabulary to represent a specific relation label. In addition, there is rich semantic knowledge between relation labels, and there are structural knowledge meanings between relation triples. Previous solutions lack the assistance of the type of the entity itself for relation prediction, which not only makes the model search in the entire relation domain during training, resulting in slow model training, but also easily leads to incorrect relation predictions. Summary of the Invention
[0008] In view of the above problems, the present invention proposes a traditional Chinese medicine text relationship extraction method based on prompt learning. First, the present invention proposes a template construction mode based on knowledge correction. By introducing a knowledge supplement vector for entities on the basis of a static prompt template, the manual overhead is reduced, and the computational complexity of template generation is greatly reduced. Secondly, we propose a relationship auxiliary vector. By roughly determining the answer word vector space based on the knowledge supplement vector, the complexity of manual answer word construction is reduced, and the generalization ability is enhanced.
[0009] The technical solution of the present invention is as follows:
[0010] A traditional Chinese medicine text relationship extraction method based on prompt learning, comprising the following steps:
[0011] S1. Process the traditional Chinese medicine text to obtain first data and second data, specifically:
[0012] Split the traditional Chinese medicine text into sentences, expressed as:
[0013] T = Tokenize(X) = {t1, t2, …, t m}
[0014] where X is the input traditional Chinese medicine text sentence, t i is the i-th segmented phrase after segmentation, 1 ≤ i ≤ m; then perform serialization marking to obtain the first data expressed as:
[0015]
[0016] where [CLS] and [SEP] represent the start symbol and truncation symbol of the sentence, [E1] and [E2] represent the first entity and the second entity, 1 < u < m, 1 < v < m and u ≠ v;
[0017] Map the entities in the first data in the following manner:
[0018] X know = T know (E)
[0019] T know (·) = E is [MASK]
[0020] where T know (·) is a template function, E represents an entity, and [MASK] represents a masked text; encode X know using the BERT language model to obtain the hidden vector h know of [MASK], defined as the knowledge supplement vector of the entity:
[0021] h know= BERT(X know )
[0022] Embed the knowledge supplement vector tags [K1] and [K2] into the template function T prompt (·) to obtain a new template function T p ′ rompt (·), and the template function T prompt (·) is used to map a sentence into a prompt input. Use T p ′ rompt (·) to map the first data to obtain the second data:
[0023] The relationship with [K2]E2[K2] is [MASK]
[0024] Among them, [K1] and [K2] respectively represent the knowledge supplement vector tags of the first entity and the second entity, and the corresponding knowledge supplement vectors are defined as and
[0025] S2. Use the first data and the second data to train along two paths, and finally conduct joint training to obtain a relationship prediction vector, specifically:
[0026] In the first training path, correct the relationship of the second data, and pre-judge the relationship between two entities. The template is expressed as:
[0027] X res = T res (·) = the relationship between [K1] and [K2] is [MASK]
[0028] Among them, X res represents the input text generated by the template function, that is, the input text in the actual relationship input stage. Since the filling vector is inserted after the subsequent BERT encoding is completed, the actual text is not directly inserted at this time, so the input text is directly equal to the output result of the template function. T res (·) represents the template function, which involves the position and quantity of the subsequent inserted filling vector, and is pre-coded through the BERT language model:
[0029] A res = EBRT(X res )
[0030] Use the vectors and to fill the tags [K1] and [K2]:
[0031]
[0032] Among them, Fill res (·) represents the use of and to fill the tags [K1] and [K2] in the template;
[0033] Use the BERT language model for training to obtain the relationship auxiliary vector:
[0034] W res = BERT(B res )
[0035] Among them, W res represents the prediction vector of the first path for the [MASK] position, that is, the relationship auxiliary vector;
[0036] In the second training path, splice the first data and the second data, which is expressed as:
[0037]
[0038] Among them, [,] represents the splicing operation; perform pre-encoding through the BERT language model:
[0039] A plm = BERT(Input)
[0040] Use the vectors and to fill the tags [K1] and [K2]:
[0041]
[0042] Among them, Fill plm (·) represents the exclusive use of and to fill [K1] and [K2];
[0043] Use the BERT language model for training to obtain the intermediate hidden layer vector:
[0044] W plm = BERT(B plm )
[0045] Among them, W plm represents the prediction vector of the second path for the [MASK] position;
[0046] Utilize W res and W plm for joint training:
[0047] W o = w res W res+w plm W plm
[0048] Among them, w res and w plm represent learnable weights, and W o is the output vector of the final hidden layer, that is, the relationship prediction vector;
[0049] S3. Obtain the final predicted label through the following formula:
[0050]
[0051] Among them, v is the relationship word in the relationship word set V, and the relationship word set V represents the total set of all possible predicted output entity relationships of the model;
[0052] Then, map and unify the output result through the following mapping function:
[0053] p(y|x) = p([MASK] = f(y)|W o )
[0054] Among them, f(y) represents the mapping function, which maps the predicted word to the specific category y. The definition of f(y) is as follows:
[0055]
[0056] Among them, C i ∈{C1, C2, …, C n}, C i is the i-th sub-word under the mentioned set W to which the class y belongs. E(·) represents the word embedding function, and n is the total number of sub-words included under the class y. Finally, the obtained specific category y is the relationship between the input entities E1 and E2 considered by the model in the input text;
[0057] S4. Use the existing traditional Chinese medicine text database and train through the methods of S1 - S3. After obtaining the trained model, input the traditional Chinese medicine text for which the relationship needs to be extracted into the trained model for relationship extraction.
[0058] The beneficial effects of the present invention are as follows:
[0059] 1) Strengthened automatic annotation of entities and relationships: The present invention uses a prompt-based pre-trained language model and combines the knowledge in the field of traditional Chinese medicine to automatically identify the entities and the relationships between entities in traditional Chinese medicine texts, avoiding the traditional cumbersome manual annotation process.
[0060] 2) Improved annotation efficiency and accuracy: By automating the annotation process and leveraging a pre-trained language model with prompts, the efficiency and accuracy of traditional Chinese medicine text annotation have been improved, saving human resources and time costs while reducing human errors during the annotation process.
[0061] 3) Promoted the informatization process of traditional Chinese medicine: By achieving automatic annotation of traditional Chinese medicine texts, it helps to promote the informatization process of traditional Chinese medicine and provides support and guarantee for traditional Chinese medicine clinical practice, medical research, education, etc. Description of the Drawings
[0062] Figure 1 It is a model architecture diagram of the present invention.
[0063] Figure 2 It is an architecture diagram of the relationship correction module.
[0064] Figure 3 It is an execution flowchart of the present invention. Detailed Implementation Manner
[0065] The following combines the drawings to describe in detail the technical principles and solutions of the present invention:
[0066] The present invention provides a pre-trained model with prompts (PLM). By introducing relevant information of entities to guide the generation of prompt learning guidance templates. In particular, learnable entity classes are used to guide the generation of learnable relationship words, and knowledge constraints in the field of traditional Chinese medicine are used to jointly improve their representations. By fine-tuning the PLM using additional prompt learning, the rich knowledge distributed in the PLM can be further stimulated, so as to better serve downstream tasks such as relationship extraction. Finally, a fast and accurate model for constructing relationships between entities in traditional Chinese medicine texts is established to achieve a fast and efficient establishment of relationships between traditional Chinese medicine entities and reduce the cumbersome manual annotation process.
[0067] As Figure 1 shown, the model architecture of the present invention is divided into three parts, namely an input module, a training module, and a prediction module. The specific execution process is as Figure 3 shown, and the following will be described in detail respectively.
[0068] The input module is used to perform data preprocessing and obtain the data required for the subsequent training module through a knowledge supplementation module. Among them, the method of data preprocessing is: perform semantic analysis on the traditional Chinese medicine literature corpus, and use a tokenizer to tokenize each sentence that needs to perform relationship extraction. By splitting the original input text into sentences, entities that need to perform relationship judgment can be identified. Assume the original input sentence is X = [x1, x2,..., x n , x i is the i-th character.
[0069] T = Tokenize(X) = {t1, t2, …, t m}
[0070] where ti i is the i-th tokenized phrase after tokenization.
[0071] After that, the tokenized sentences are serially marked, and the specific formula is as follows:
[0072]
[0073] The serial marking function S(X) forms a new, tagged sentence by identifying the entity positions of X in the original sentence. [CLS] and [SEP] represent the start symbol and truncation symbol of the sentence. [E1] and [E2] represent the entity 1 label and entity 2 label.
[0074] For example, for the original text "Those who are dizzy and restless at night cannot sleep soundly. Treat them with Wendan Decoction.", after tokenization by the tokenizer, the tokenization result of ["dizzy", "restless at night", "those", ".", "Wendan Decoction", "treat", "them", "."] is output. After that, the tokenized result text is serially marked to generate the serial marking result ["[CLS]", "dizzy", "[E1]", "restless at night", "[E1]", "those", ".", "[E2]", "Wendan Decoction", "[E2]", "treat", "them", ".", "[SEP]"]
[0075] At the same time, for each sentence instance X, use the template function to map it to the prompt input:
[0076] X prompt = T prompt (X)
[0077] T prompt (·) = The relationship between E1 and E2 is [MASK]
[0078] where the template function T prompt (X) involves the position and quantity of adding additional words. E1 and E2 represent entity 1 and entity 2, and [MASK] represents the masked text, which needs to be predicted when the model outputs. For example, for the text "Those who are dizzy and restless at night cannot sleep soundly. Treat them with Wendan Decoction.", the template text "The relationship between restless at night and Wendan Decoction is [MASK]" will be generated.
[0079] The knowledge supplementation module is to bridge the gap between the pre-training task and the downstream task. To better predict the entity relationship, learnable entity classes are used to guide the generation of learnable relationship words, and the knowledge in the traditional Chinese medicine field is used to jointly improve its representation by introducing knowledge into the above template T prompt(·) method to increase the semantic information contained in the template. Specifically, first set the following template to extract the knowledge supplementation vector. X know = T know (E)
[0080] T know (·) = E is [MASK]
[0081] Among them, the template function T know (E) involves the position and quantity of adding additional words. E represents an entity.
[0082] For example, let the entity E be Wendan Decoction, and map it to X know = Bupleurum chinense DC. is [MASK]. Then, we can encode X know through the BERT language model to obtain the hidden vector h know of [MASK], that is, the knowledge supplementation vector of this entity, which can be expressed as:
[0083] h know = BERT(X know )
[0084] After that, by embedding the knowledge supplementation vector labels [K1] and [K2] into the original template function T prompt (·), it becomes a new template function to guide the relationship prediction to be carried out within a certain range, specifically as follows.
[0085]
[0086] Among them, E1 and E2 represent entity 1 and entity 2, and [K1] and [K2] represent the knowledge supplementation vector labels of entity 1 and entity 2. The corresponding knowledge supplementation vectors and will be used for filling later.
[0087] For example, the original T prompt template text is "The relationship between insomnia and Wendan Decoction is [MASK]". After passing through the knowledge supplementation module, the original template will be enhanced to "[SEP][K1]insomnia[K1] and [K2]Wendan Decoction[K2]'s relationship is [MASK]", where [K1] and [K2] will be filled with vector representations in the next module, which has a different meaning from [MASK].
[0088] After investigation, it shows that the possible relationships between entity types are limited. Especially in the relationship prediction of traditional Chinese medicine texts, this relationship constraint phenomenon will be more obvious. Therefore, this method of further constraining relationships through entity types is effective.
[0089] The training module conducts training based on the obtained data, specifically including two training lines:
[0090] The first training line has a relationship correction module. As Figure 2 shown, it performs relationship prediction on the knowledge supplement vector generated from the previous text through prompt learning to generate a relationship auxiliary vector. Using this vector for constraint guidance, it conducts a prior relationship verification, increasing the accuracy of relationship prediction by pre-judging the possible relationship between two entities. Specifically, for entity 1, Wendan Decoction, and entity 2, restless sleep at night, their knowledge supplement vectors are obtained through the knowledge supplement module in step 2 and Set the template as follows:
[0091] X res = T res (·) = The relationship between [K1] and [K2] is [MASK]
[0092] where [K1] and [K2] represent the knowledge supplement vector labels of entity 1 and entity 2.
[0093] Perform pre-encoding through the BERT language model:
[0094] A res = BERT(X res )
[0095] After pre-encoding is completed, specific vectors are used to fill the labels. For this example, we use the knowledge supplement vectors of entity 1, Wendan Decoction, and entity 2, restless sleep at night and for relevant filling.
[0096]
[0097] Among them, Fill res (·) means using and to fill the labels [K1] and [K2] in the template.
[0098] Finally, use the BERT language model for training to obtain the final relationship auxiliary vector:
[0099] W res = BERT(B res )
[0100] W resRepresents the final predicted vector for the [MASK] position, that is, the relation auxiliary vector. By extracting the output vector of the prompt learning hidden layer, without using a projection function to specifically map such a broad vector as the relation auxiliary vector into a relation. Since the description of words can be ambiguous, that is, there may be many different mentions for a specific word. Mapping into a specific relation may lead to incomplete mapping due to the complexity of the relationship between entities itself or the complexity of the corresponding mentions of the relation, affecting the prediction accuracy. Therefore, a continuous type embedding vector is selected instead of a discrete version. Then, this hidden layer vector is used as the entity virtual type embedding template to guide relation prediction.
[0101] The second route utilizes input concatenation by concatenating the result of data preprocessing and the prompt template containing the knowledge correction module:
[0102]
[0103] Where [,] represents the concatenation operation. For example, for the initial input text "Those who are dizzy and restless at night cannot sleep soundly. Treat with Wendan Decoction.", after data preprocessing and the knowledge supplementation module, the following text is obtained: "[CLS]Dizzy [E1]Restless at night and cannot sleep soundly [E1]. [E2]Wendan Decoction [E2]is the treatment. [SEP][K1]The relationship between [K1]Restless at night [K1] and [K2]Wendan Decoction [K2] is [MASK]".
[0104] Perform pre - encoding through the BERT language model.
[0105] A plm = BERT(Input)
[0106] After pre - encoding, specific vectors will be used to fill the labels. In this example, the knowledge supplementation vectors of entity 1 Wendan Decoction and entity 2 Restless at night and are used for relevant filling.
[0107]
[0108] Where, Fill plm (·) means only using and to fill specific labels in the template, that is, [K1] and [K2].
[0109] PLM training: By inputting the text obtained from preprocessing and prompt engineering into the BERT pre - trained model for training, the intermediate hidden layer vectors are obtained for subsequent joint training.
[0110] W plm = BERT(B plm )
[0111] W plm represents the final predicted vector for the [MASK] position.
[0112] Finally, joint training is carried out. By jointly training the two hidden layer vectors obtained from BERT pre-training and the relation correction module, the final relation prediction vector is obtained. The prior traditional Chinese medicine relation obtained in the relation correction module, that is, the hidden layer vector W of [MASK] res and the hidden layer vector W trained by PLM plm are jointly trained.
[0113] W o = w res W res + w plm W plm
[0114] where w res and w plm represent learnable weights. W o is the final hidden layer output vector.
[0115] The prediction module is used to obtain the prediction result, specifically:
[0116] Since the final output of the relation prediction is a probability distribution based on the vocabulary, the directly outputted words do not correspond one-to-one with the category labels of the task. Its output has diversity and there may be synonyms and near-synonyms output, and it may also output words irrelevant to the task in the vocabulary, which can be uniformly processed through answer word mapping. At the same time, the number of relations and entity types of traditional Chinese medicine entities is small, so clusters will be formed among entity classes. Vectors within the same cluster will be highly aggregated and similar, while the gap between different clusters will be large. Specifically, by retrieving the used corpus, all relation types between two entities can be determined. By aggregating all sub-mentions under the three relations with the highest occurrence frequencies between these entity types, answer word mapping can be carried out.
[0117]
[0118] where W = {C1, C2, …, C n}, C i is the i-th sub-word in the set of mentions W belonging to the class y. E(·) represents the word embedding function. For example, there may be three relation types with the highest occurrence frequencies such as {cause, treat, assist} between traditional Chinese medicine type entities and disease type entities. By retrieving the used corpus, for the mentions of the relation type y (treat), we may have W = {cure, heal, rescue} these three. By aggregating the word embedding vectors of all mentions, the mapping vector of this relation type can be obtained, denoted as f(y) represents the mapping function that maps the predicted word to the specific category y.
[0119] Finally, the final predicted label is generated by calculating and selecting the search with the highest probability output:
[0120]
[0121] where W o is the output vector of the final hidden layer, and v is the relational word in the relational word set V. Through the model, the word at the above [MASK] position can be predicted as a certain word in the vocabulary, but the word predicted at the masked position is not always consistent with the relational label word. Therefore, the label word needs to be transformed. For example, for the text "Those who are dizzy and restless at night cannot sleep. Treat them with Wendan Decoction.", predict the relationship between being restless at night and Wendan Decoction. The possible results predicted by the model are treatment, cure, application, etc. This is because the prompt-based pre-training model predicts the probability of the next word, and there are likely to be many answers, but these answers can all be classified into a large category. Therefore, we need to use a mapping function to map and unify the model output results, that is, map {treatment, cure, application} to "treatment".
[0122] p(y|x) = p([MASK] = f(y)|W o )
[0123] f(y) represents the mapping function that maps the predicted word to the specific category y.
Claims
1. A TCM text relationship extraction method based on prompt learning, characterized in that: The following steps are involved: S1. Processing the TCM text to obtain first data and second data, specifically: Sentence splitting of TCM text is expressed as: T=Tokenize(X)={t1,t2,…,t m } Where X is the input TCM text sentence, t i is the i-th segmentation phrase after segmentation, 1≤i≤m; then serialize and mark to get the first data It is expressed as: Wherein, [CLS] and [SEP] represent the start symbol and truncation symbol of the sentence, [E1] and [E2] represent the first entity and the second entity, 1<u<m, 1<v<m and u≠v; The first data The entities in are mapped as follows: X know =T know (E) T know (·)=E is [MASK] Among them, T know (·) is a template function, E represents entity, [MASK] represents mask text; know Encode the BERT language model and get the hidden vector h of [MASK] know , defined as the knowledge supplement vector of the entity: T know =BERT(X know ) In the template function T prompt (·) embeds the knowledge supplement vector labels [K1] and [K2] to obtain a new template function T p ′ rompt (·), the template function T prompt (·) is used to map sentences into prompt inputs; using T p ′ rompt (·) for the first data Mapping is performed to obtain the second data: Among them, [K1] and [K2] represent the knowledge supplement vector labels of the first entity and the second entity respectively, and the corresponding knowledge supplement vector is defined as and S2, using the first data and the second data to perform training in two paths, and finally performing joint training to obtain a relationship prediction vector, specifically: In the first training path, the relationship between the second data is corrected to pre-judge the relationship between the two entities. The template is expressed as: X res =T res (·)=The relationship between [K1] and [K2] is [MASK] Among them, X res Indicates that through the template function T res (·) Generated input text is pre-encoded by the BERT language model: A res =BERT(X res ) Using Vectors and Fill the labels [K1] and [K2]: Among them, Fill res (·) indicates use and Fill in the tags [K1] and [K2] in the template; Use the BERT language model for training to obtain the relation auxiliary vector: W res =BERT(B res ) Among them, W res Represents the prediction vector of the first path for the [MASK] position, i.e., the relation auxiliary vector; In the second training path, the first data and the second data are concatenated, which is expressed as: Among them, [,] represents the concatenation operation; pre-encoding is performed through the BERT language model: A plm =BERT(Input) Using Vectors and Fill the labels [K1] and [K2]: Among them, Fill plm (·) indicates that only and Fill [K1] and [K2]; Use the BERT language model for training to obtain the intermediate hidden layer vector: W plm =BERT(B plm ) Among them, W plm Represents the prediction vector of the second path for the [MASK] position; Using W res and W plm Conduct joint training: IN o =in res IN res +in plm IN plm Among them, w res and w plm represents the learnable weight, W o The final hidden layer output vector is the relationship prediction vector; S3. The final predicted label is calculated using the following formula: Among them, v_rel is the relation word in the relation word set V, and the relation word set V represents the total set of entity relations that may be predicted and output by the model; Then use the following mapping function to unify the output results: p(y|x)=p([MASK]=f(y)|W o ) Among them, f(y) represents the mapping function, which maps the predicted word to the specific category y. The definition of f(y) is as follows: Among them, C i ∈{C1,C2,…,C n }, C i is the i-th sub-word in the set of mentions of the class y, E(·) represents the word embedding function, and n is the total number of sub-words contained in the class y. Finally, the specific class y obtained is the relationship between the input entities E1 and E2 in the input text as considered by the model. S4. Using the existing TCM text database, train using the methods of S1-S3. After obtaining a trained model, input the TCM text that needs to extract relationships into the trained model for relationship extraction.