Text data subject-object association relation extraction method based on medical field features

Optimizing the subject-object association relationship extraction of medical text through Star-Transformer and Bi-LSTM models, solving the problems of insufficient subject-object semantic interaction and poor generalization of long texts in the existing methods, achieving higher entity and relationship recognition accuracy and lower computational complexity.

CN120296107APending Publication Date: 2025-07-11BEIJING NORMAL UNIV AT ZHUHAI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510349082.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing medical relationship extraction methods ignore semantic interactions between subjects and objects in medical texts and poor generalization when processing long texts, resulting in low accuracy in entity recognition and relationship prediction.

Method used

Using Star-Transformer-based method, the information interaction between the subject and object is enhanced through the star topology structure and the Bi-LSTM model, the pre-trained model BERT is used for encoding, and combining the self-attention mechanism and residual connections is optimized to improve the semantic representation ability of medical texts.

Benefits of technology

Effectively capture potential correlation information between subject and object, improve the recognition accuracy of medical entities and relationships, reduce the computational complexity and resource consumption of long text processing, and enhance the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296107A_ABST
    Figure CN120296107A_ABST
Patent Text Reader

Abstract

The invention provides a text data subject-object association relationship extraction method based on medical field features, which comprises the following steps: extracting feature information from a medical text to be extracted by using an encoder, and encoding context information in the medical text to obtain an encoding vector of the medical text; constructing a medical subject marker, and marking a starting position and an ending position of a subject entity in the medical text according to the coding vector; optimizing a medical subject marker by adopting a likelihood function, and obtaining a subject entity list in the medical text; constructing a medical relationship marker, modeling the medical relationship of each subject entity in the subject entity list, performing semantic enhancement through a star topology structure, identifying objects related to the subject entities, and extracting relationship features; and optimizing a medical relationship marker by adopting a likelihood function, and completing extraction of the subject-object association relationship of the text data based on the medical domain characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for extracting the association relationship between the main and object entities of text data, and in particular to a method for extracting the association relationship between the main and object entities of text data based on the characteristics of the medical field. Background Art

[0002] The information provided in this section is only background information related to the present disclosure, and it is not necessarily prior art.

[0003] As an important subtask of medical information extraction, relation extraction aims to extract entity pairs and their semantic relationships from medical texts (such as CT reports, MRI image descriptions, electronic medical records) to achieve the automatic construction of medical knowledge. The core of medical relation extraction is the (S, R, O) triple, where S (subject) represents a medical entity, such as a disease, an examination method, a drug, O (object) represents another medical entity, and R (relation) represents the semantic connection between them, such as "examination - diagnosis", "drug - indication", "symptom - complication". In downstream tasks, extracting entities such as diseases, examination methods, imaging results, treatment plans, etc. and their relationships from unstructured medical texts is a key step in applications such as medical knowledge graph construction, clinical decision support, and intelligent consultation. Early medical relation extraction methods adopted a pipeline approach, first identifying medical entities and then predicting relationships based on these entities. However, this method has the following problems:

[0004] (1) Entity recognition and relation prediction are completely separated, ignoring the information interaction between the two.

[0005] (2) Error propagation problem: Errors in the previous entity recognition will affect subsequent relation prediction, reducing the overall extraction accuracy.

[0006] To solve the error propagation problem, medical entity - relation joint extraction models have become a research hotspot. The joint extraction method enhances the information interaction between entity recognition and relation extraction by sharing parameters, reduces error accumulation, and improves the accuracy of relation prediction. However, although such methods have made some progress, there are still two main problems in the field of medical texts:

[0007] One is that the medical semantic interaction between the main and object entities is ignored:

[0008] (1) Existing methods mainly focus on context semantics to optimize medical entity recognition, but ignore the internal association between the main and object entities in the medical context.

[0009] (2) In medical texts, there are often implicit semantic connections in relationships such as disease-examination, symptom-disease, and drug-indication. If these connections cannot be accurately modeled, it may lead to failure in relationship extraction. For example, in an electronic medical record, "The patient's CT shows low-density shadows in the brain" can infer that the CT result is related to cerebral infarction. Similarly, "MRI indicates abnormal signals in the myelin sheath" may imply multiple sclerosis. These interactions of medical semantic information are crucial for the relationship extraction task.

[0010] Second, the generalization ability of the model is poor when dealing with long texts:

[0011] (1) In recent years, some studies have adopted a two-dimensional matrix marking method to annotate entity boundaries through an n×n matrix and used it for relationship prediction. Although this method effectively alleviates the problem of error propagation and can handle multiple relationship overlaps (such as the same entity pair can have multiple relationships, such as "examination-diagnosis" and "examination-abnormal finding"), its generalization ability is poor on long-text medical corpora (such as imaging reports and electronic medical record records). The reason for this problem is that: the dimension of the matrix changes with the sentence length n, and its computational complexity is O(n 2 ), resulting in an exponential growth of computational resources and affecting the efficiency and scalability of the model.

[0012] (2) In CT / MRI imaging reports and electronic medical records, the sentence length varies greatly and the density of medical entities is low. The matrix method is prone to sparsity, affecting the generalization performance of the model.

[0013] It should be noted that the information disclosed in the above background technology section is only used to strengthen the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0014] Objective of the Invention: The technical problem to be solved by the present invention is to provide a method for extracting the association relationship between the main body and the object of text data based on the characteristics of the medical field in view of the deficiencies of the prior art.

[0015] To solve the above technical problem, the present invention discloses a method for extracting the association relationship between the main body and the object of text data based on the characteristics of the medical field, including:

[0016] Step 1, use an encoder to extract feature information from the medical text to be extracted and encode the context information in the medical text to obtain the encoded vector of the medical text;

[0017] Step 2, construct a medical subject marker to mark the start position and end position of the subject entity in the medical text according to the encoded vector;

[0018] Step 3: Optimize the medical entity tagger using the likelihood function, and use the optimized medical entity tagger to obtain the list of entity entities in the medical text;

[0019] Step 4: Construct a medical relation tagger to model the medical relations of each entity entity in the entity entity list, perform semantic enhancement through a star topology structure, further identify the objects related to the entity entity, and extract relation features;

[0020] Step 5: Optimize the medical relation tagger using the likelihood function, and use the optimized medical relation tagger to complete the extraction of the main - object association relationship of the text data based on the medical domain features.

[0021] Furthermore, obtaining the encoded vector of the medical text in Step 1 includes:

[0022] Step 1 - 1: Assume that the medical text to be extracted is in the form of a character sequence:

[0023] S = {s1, s2, …, s n}

[0024] where s n represents the n - th character in the character sequence S; extract the feature information X from the character sequence S using the encoder, which is expressed as follows:

[0025] X = SW s

[0026] where W s is the character embedding matrix used to map each character to a low - dimensional dense vector representation;

[0027] Step 1 - 2: Use the pre - trained model BERT to encode the context information of the medical text to obtain the encoded vector h N , which is expressed as follows:

[0028] h N = BERT(S)

[0029] where BERT() is the encoding process of the pre - trained model BERT, which is a stack composed of N identical Transformer blocks.

[0030] Furthermore, the encoding of the pre - trained model BERT in Step 1 - 2 is specifically expressed as follows:

[0031] h0 = SW s + W p

[0032] h α = Transformer(h α-1), α ∈ [1, N]

[0033] Among them, h0 represents the initial feature representation of the input medical text, that is, the text feature vector after character embedding and position embedding, h α represents the medical text feature representation of the α-th layer, and Transformer() represents the Transformer block.

[0034] Furthermore, the marking of the start position and end position of the main entity in the medical text described in step 2 includes:

[0035] Step 2-1, use the Bi-LSTM model to construct a medical entity marker;

[0036] Step 2-2, input the encoded vector h N of the medical text into the medical entity marker to obtain the medical main entity feature representation H N processed by the Bi-LSTM model, which is expressed as follows:

[0037] H N = BiLSTM(h N )

[0038] Among them, BiLSTM() represents the processing process of the Bi-LSTM model;

[0039] Step 2-3, denote the feature representation of the i-th character in the medical main entity feature representation H N as H N [i], and calculate the probabilities and

[0040] of each character in the medical text being the start and end of the main entity and Compare the values of

[0041] with a threshold. The u-th character position exceeding the threshold is marked as 1, otherwise marked as 0, to achieve the marking of the start position and end position of the main entity in the medical text. and Specifically, the calculation of the probabilities

[0042]

[0043] of each character in the medical text being the start and end of the main entity is as follows: and represent the probabilities of the u-th character being the start and end of the main entity, W start and W endThe trainable weight parameters b for determining whether each character in a medical text is the start or end of a subject entity start and b end are the corresponding bias terms, and σ is the Sigmoid activation function that maps the input value to the (0, 1) probability space.

[0044] Furthermore, the likelihood function is used to optimize the medical subject tagger in step 3, including:

[0045] Step 3-1, design the likelihood function p θ (y|S);

[0046] Step 3-2, adjust the trainable parameters W start and W end in the medical subject tagger through the above likelihood function, so that the probabilities and predicted by the medical subject tagger for the start and end of the subject entity are closer to the true labels Specifically as follows:

[0047] For each character i in the medical text, if its true label is the start or end position of the subject, that is then maximize the probability that the medical subject tagger predicts it as 1 otherwise maximize the probability that the model predicts it as 0

[0048] Step 3-3, according to step 3-2, optimize the parameters in the medical subject tagger through gradient descent, obtain the start and end position tags of the subject entity, and adopt the closest start-end pair matching principle to determine the entity span according to the results of the start and end position tags, expressed as follows:

[0049]

[0050] where e represents the subject entity, and S[i start :i end represents the text from the start character position i start to the end character position i end Among them, the start character position i start satisfies The end character position i end satisfies

[0051] Step 3-3, according to all the subject entities e obtained in step 3-3, obtain the subject entity list E, expressed as follows:

[0052] E = {e1e1,…,e m}

[0053] Among them, e m represents the m-th subject entity.

[0054] Furthermore, the likelihood function p θ (y|S) described in step 3-1 is as follows:

[0055]

[0056] Among them, n is the length of the medical text S, is the true label where the start or end position of the subject entity is the i-th character. Here, q is the tag type, representing the start position start s or the end position end s , and the parameter θ = {w start , b start , W end , b end}.

[0057] Furthermore, the construction of the medical relation tagger described in step 4 includes:

[0058] Step 4-1: Perform independent attention calculations on the encoded vector h N of the medical text through the self-attention mechanism, and integrate the results of multiple attention heads, specifically as follows:

[0059] MulAtt = concat(z1, z2,..., z h ) · W O

[0060] Among them, MulAtt is the output of splicing multiple attention heads, concat(·) represents vector splicing, W O is a learnable parameter, and z h represents the attention output of the h-th head, specifically as follows:

[0061]

[0062] Among them, and are learnable parameters, and Att() is to calculate a single attention head, specifically as follows:

[0063]

[0064] Among them, Q is the query vector, representing the information to be focused on currently, K is the key vector, representing all possible information to be queried, V is the value vector, representing the information content related to the key, and d k represents the dimension of the query vector and the key vector, where:

[0065] K = X eW K , V = X e W V

[0066] Among them, X e is the representation of medical entity features, that is, the feature representation corresponding to the position of the subject entity e in the feature information X;

[0067] Step 4-2: Adopt a star topology to establish the relationship of the subject object, specifically as follows:

[0068] Take each subject entity in the subject entity list E as the relay node set S t , and the object entity as the satellite node set H t , and calculate through the self-attention mechanism in Step 4-1, specifically as follows:

[0069]

[0070] Among them, represents the context information of the r-th node in the t-th round, e r represents the medical entity embedding, represents the state of the r-th satellite node in the t-th round, specifically as follows:

[0071]

[0072] Among them, MulAtt represents the multi-head attention mechanism in Step 4-1;

[0073] Update the relay node S t , specifically as follows:

[0074] S t = MulAtt(s t-1 , [s t-1 ; H t , [s t-1 ; H t )

[0075] Step 4-3: Adopt a residual connection to simplify the star topology, that is, after calculating the subject and object representations S t and H t in Step 4-2, establish a residual connection between the information of the previous layer loop and the output of the current attention, and update the satellite node, specifically as follows:

[0076] H’ = LayerNorm(σ(net l ))

[0077] Among them, σ is the ReLU activation function, H’ represents the updated satellite node, LayerNorm(·) is the normalization operation, net lis the candidate hidden state of the current layer, specifically as follows:

[0078] net l = (w1[MulAtt(h f , C f , C f )] + b1 + net l-1 )

[0079] where w1 and b1 are learnable parameters, h f represents the hidden state of the f-th satellite node, C f represents its context information, the first C f is used as the key K and value V for calculating the attention distribution, and the second C f is used as the value V to provide information weighted calculation, and net l-1 is the candidate hidden state of the previous layer;

[0080] Step 4-4: Through the star topology structure simplified in Step 4-3, obtain the final satellite node representation and relay node representation, that is, the sentence vector representation and entity representation of the medical text, and predict the head and tail markers of the object;

[0081] Step 4-5: Use the average value of the global character-level representation as the main entity representation, that is, is represented as the average value of all character sequence vector representations included in the k-th subject.

[0082] Furthermore, the prediction of the head and tail markers of the object described in Step 4-4 is specifically as follows:

[0083]

[0084] where, and represent the probabilities of the sequence features of the i-th character as the start and end positions of the object; represents the encoded representation vector of the k-th entity detected in the encoder; and are learnable parameters, and are biases.

[0085] Furthermore, the likelihood function used to optimize the medical relation tagger described in Step 5 includes:

[0086] Design the likelihood function Specifically as follows:

[0087]

[0088] where n is the length of the medical text. For each character i in the medical text, if its true label is the start or end position of the subject, i.e., then maximize the probability that the medical relation tagger predicts it as 1 otherwise maximize the probability that the medical relation tagger predicts it as 0 is the binary tag for the start position of the object of the i-th character in the medical text S, representing the tag for the end position of the object of the i-th character; for the empty object o φ , which applies to all tags i; the tag parameters

[0089] Advantageous effects:

[0090] 1. In the present invention, each character in the sentence is used as a satellite node, and the medical entity representation is used as a relay node. This structure collects and sends information from the satellite nodes through the relay node, enabling effective information interaction between the subject and the object and capturing potential associated information between the subject and the object. For example, in a CT report, the implicit relationship between "low-density shadow" and "cerebral infarction" can be strengthened through modeling with a star topology.

[0091] 2. The present invention enhances the medical semantic representation of entities and sentences. Through the information interaction of the star structure, relying solely on local context for entity representation is avoided, and the global representation ability of medical entities is improved. This is mainly manifested in enhancing the reasoning ability of medical semantic relationships such as "disease-examination" and "symptom-treatment".

[0092] 3. While ensuring the inherent characteristics of the topology structure, the present invention uses residual connections to reduce its depth and complexity. To enable ST-ICRE to have a certain generalization ability, only two sets of vector representations are used for the boundary marking of entities, with a complexity of O(2n), effectively avoiding the problem of resource waste generated when processing long sentences. For the possible error propagation problem of this marking method, a pre-trained model combined with a bidirectional long short-term memory network (Bi-LSTM) is used to improve the accuracy of the marking, hoping to alleviate the impact of this problem on subsequent relationship prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] The following further specific description of the present invention is made in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0094] Figure 1 is a schematic diagram of the overall model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0095] The overall idea of the present invention is as follows:

[0096] To address the problems of insufficient information interaction between the subject and object and poor model generalization in processing long sentences in the medical text relation extraction task, the present invention proposes a relation extraction method based on the intrinsic correlation between the subject and object of Star-Transformer (Star-Transformer Intrinsic Correlation Relation Extraction Between Subject and Object, STIC-ICRE). In medical texts such as CT image reports, MRI diagnostic descriptions, and electronic medical records, there are often implicit medical semantic associations between the subject and object entities, but traditional methods are difficult to model. For this reason, ST-ICRE constructs a star topology structure. First, each character in the sentence is used as a satellite node, and the medical entity representation is used as a relay node. This structure collects and sends information to the satellite nodes through the relay node, enabling effective information interaction between the subject and object and capturing potential associated information between the subject and object. For example, in a CT report, the implicit relationship between "low-density shadow" and "cerebral infarction" can be strengthened through the star topology structure modeling. Second, the medical semantic representation of entities and sentences is enhanced. Through the information interaction of the star structure, relying solely on local context for entity representation is avoided, and the global representation ability of medical entities is improved. This is mainly manifested in enhancing the reasoning ability of medical semantic relations such as "disease-examination" and "symptom-treatment". In addition, while ensuring the inherent characteristics of the topology structure, residual connections are used to reduce its depth and complexity. To enable ST-ICRE to have a certain generalization ability, only two sets of vector representations are used for the boundary marking of entities, with a complexity of O(2n), effectively avoiding the problem of resource waste generated when processing long sentences. For the possible error propagation problem of this marking method, a pre-trained model combined with a bidirectional long short-term memory network (Bi-LSTM) is used to improve the accuracy of marking, in order to alleviate the impact of this problem on subsequent relation prediction.

[0097] To achieve the above object, the specific technical solution of the present invention is as follows: A method for extracting the associated relationship between the subject and object of text data based on the characteristics of the medical field, including:

[0098] Step 1, use an encoder to extract feature information X={x1, x2,..., x n} from the character sequence S={s1, s2,..., s n} of the medical text (such as CT image report, MRI diagnostic description, electronic medical record), where X = SW s , and use the pre-trained model BERT to encode the context information of the medical text to obtain the encoded vector h N。The BERT used in the present invention is a multi-layer bidirectional Transformer model based on medical language representation, which can fully learn the deep semantic representation of medical texts and is especially suitable for medical term modeling in long-text medical records (such as electronic medical records) and imaging reports (such as CT and MRI diagnoses). Compared with traditional methods, BERT enhances the ability to recognize complex medical entities by jointly conditioning on the left and right contexts of each character. Specifically, BERT consists of a stack of N identical Transformer blocks, each of which contains a self-attention mechanism and a feed-forward neural network, and can accurately capture the associations between entities such as diseases, examination methods, and treatment plans in the context of medical texts. When processing CT / MRI imaging reports, BERT can capture the potential relationship between "low-density shadow" and "cerebral infarction" through long-distance dependence modeling, and can accurately analyze the causal relationship between "diabetes" and "diabetic nephropathy" in electronic medical records. The specific operations are shown in Formulas 1 to 3 to ensure the construction of accurate semantic representations in medical texts and provide optimized input information for subsequent subject-object entity recognition and relationship extraction.

[0099] h0 = SW s +W p (1)

[0100] h α = Transformer(h α-1 ), α ∈ [1, N] (2)

[0101] h N = BERT(S) (3)

[0102] Among them, h0 represents the initial feature representation of the input medical text, that is, the text feature vector after character embedding and position embedding. S represents the character sequence representation of the medical text, which is composed of a character index matrix encoded by one-hot. For example, "hypertension" will be mapped to a one-hot vector. W s represents the character embedding matrix, which is used to map each medical character to a low-dimensional dense vector representation. For example, "blood" may have a semantic vector similar to "heart". W p represents the position embedding matrix, which represents the position information of medical characters in the text to ensure that the model can recognize the order of medical terms (such as "blood pressure" is different from "pressure blood"). h αDenote the medical text feature representation of the α-th layer, which integrates global context information, enabling the model to understand the associations among diseases, symptoms, examinations, and treatments. Transformer() represents the Transformer module, which further captures the deep semantics of medical entities and relationships through the multi-head self-attention mechanism and the feed-forward neural network. N represents the number of Transformer blocks, and S is the character input sequence.

[0103] Step 2, construct a medical subject tagger to use the medical text encoding vector h generated by the BERT encoder N as the input of the Bi-LSTM, combining the deep context modeling ability of BERT and the sequence feature extraction ability of Bi-LSTM to optimize the recognition performance of complex entities (such as diseases, examination methods, treatment methods) in medical texts. In medical texts such as CT, MRI reports, and electronic medical records, entities often have the characteristics of cross-sentence, strong implicit relevance, and strong context dependence. Therefore, it may be difficult to accurately capture the boundary information of entities only relying on BERT. Bi-LSTM can capture the dependencies of medical entities in both forward and backward directions through bidirectional recursive calculation, enhancing the ability to identify entity boundaries. To label the subjects in medical texts, the present invention adopts a binary tagging mechanism for entity recognition: each character is assigned a binary tag (0 / 1), and the start or end position of the entity is marked as 1, and the rest are marked as 0. This Subject tagger is applicable to the boundary detection of medical terms (such as "hypertension", "abnormal MRI signal"), examination results (such as "pulmonary nodule"), and disease diagnoses (such as "acute myocardial infarction"). The specific operation method of the Subject tagger is shown in Formulas 4 to 6 to ensure accurate extraction of subject entities in the long-dependency environment of medical texts, providing high-quality input for subsequent relationship prediction and entity matching.

[0104] H N = BiLSTM(h N ) (4)

[0105]

[0106] where h N is the medical text encoding vector, representing the initial medical text feature representation output by BERT, and H N is the medical entity feature representation processed by the Bi-LSTM layer, and represent the probabilities that the i-th character is the start and end positions of the subject. If the probability exceeds the threshold of 0.5, the corresponding tag is 1, otherwise it is 0. H N [i] represents the feature representation of the i-th character in the medical entity, and W start 、W endTrainable weight parameters b for determining whether each character in a medical text is the start or end of a subject entity syart and b end are the corresponding bias terms, and σ is the Sigmoid activation function, which maps the output value to the (0, 1) probability space.

[0107] Step 3: Optimize the medical subject tagger obtained in Step 2 using the likelihood function. Specifically, adjust the trainable parameters W start and W end of the model through the likelihood function so that the predicted start / end probabilities and are closer to the true tags The model optimizes these parameters through gradient descent to more accurately identify the start and end positions of entities in long-text medical corpora (such as electronic medical records and imaging reports), thereby improving the ability to detect the integrity of medical entities. See Equation 7. For entity start and end tags, the principle of closest start-end pair matching is adopted, and the start and end position tags are used to determine any detected span. At the same time, to match the end tag for a given start tag, the model does not consider the tags before the given tag position. Since in medical reports and electronic medical records, entities such as diseases, symptoms, and examination results often have natural continuity, if the start and end positions are correctly detected, this matching strategy can maintain the integrity of any entity span. In addition, based on the start / end tags, a list of subject entities E = {e1, e2,..., e m} is generated, and the principle of closest start-end pair matching is used to determine the entity span:

[0108]

[0109] where n is the length of the medical text. For each character i in the medical text, if its true label is the start or end position of the subject (i.e., where q represents the "tag type", i.e., start s or end s ), then maximize the probability that the medical subject tagger predicts it as 1 Otherwise, maximize the probability that the model predicts it as 0 is the binary tag for the start position of the subject corresponding to the i-th character in S, represents the subject end position tag. The parameter θ = {W start ,, b start , W end , b end}. p θ(y|S) represents the probability that the model predicts the correct entity start / end tag sequence y for the given input text S.

[0110] Step 4: Model each possible medical relationship, construct a relationship-specific tagger, and identify the relationship between the subject and the corresponding object. The list of subject entities E = {e1, e2, …, e m} While adapting to the characteristics of Chinese medical data, it enhances Chinese medical semantics through a star structure. This tagger consists of a group of relationship-specific object taggers with the same structure as the subject tagger. For all possible medical relationships, these taggers simultaneously identify the corresponding medical objects for each detected medical subject. Different from the subject tagger that decodes the encoded vector, the relationship-specific object tagger not only decodes the encoded vector but also, in view of the characteristics of Chinese medical texts, adopts a star structure to strengthen the semantic relevance of the entity-relationship pair to improve the accuracy of medical relationship extraction. The detailed operation of the relationship-specific object tagger on the encoded vector of each character is as follows.

[0111] Furthermore, the specific steps of Step 4 are as follows:

[0112] Step 4.1: Enhance the attention weight of the subject entity to ensure correct attention allocation. The Transformer uses attention heads to implement self-attention separately for the input sequence h N and integrates the results of each attention head together, which is called multi-head attention. Based on the list of subject entities E = {e1, e2, …, e m}, when calculating the query vector Q, higher weights are assigned to entity-related information to enhance the attention allocation of medical subject entities. For medical relationships, when calculating attention, a Medical Entity Dictionary (MED) is introduced, and the weights of medical terms are increased when calculating Q, K, and V. The formula for enhancing the correlation of medical proper nouns is shown in Formulas 8 to 11.

[0113]

[0114] K = X e W K , V = X e W V (9)

[0115] MulAtt = concat(z1, z2, …, z h ) · W O (10)

[0116]

[0117] Among them, Q is the query vector, representing the information to be focused on currently, K is the key vector, representing all possible information to be queried, V is the value vector, representing the information content related to the key, and K T represents the transpose of matrix K, such that QK T performs a dot product calculation. d k represents the dimension of the query vector and the key vector, which is used for normalization to prevent excessive numerical values, and z j represents the attention output of the j-th head, and W K 、W V 、W O 、 are learnable parameters, concat(·) represents vector concatenation, MulAtt is the output of concatenating multiple attention heads, and Att() is the calculation of a single attention head. X e is obtained by extracting the text features X corresponding to the list E of medical entities (such as "lung nodule", "MRI image") identified in step 3. Here, E is the E in the previous text, and X is the X in the previous text.

[0118] Step 4.2, use Star-Transformer to establish the subject-object relationship. The Star-Transformer topological structure consists of 1 relay node and several satellite nodes. The state of the r-th satellite node represents the feature of the i-th character in the medical text sequence, that is, the representation vector of the medical entity or term. The relay node acts as a virtual hub, responsible for collecting and distributing information from all satellite nodes, ensuring the global information dissemination of the medical text, and improving the association degree of distant medical entities. Star-Transformer proposes a time-step cyclic update method. Each satellite node is initialized with an input vector, the relay node is initialized as the average of all text sequence features, and the multi-head attention mechanism is used for node update.

[0119] In view of the characteristics of medical texts, the present invention proposes a topological structure optimization method, enabling the model to obtain remote dependency information while ensuring its connection with medical texts and other medical entities by taking the representation of the entire medical entity span as the relay node, so as to enhance the semantic representation and obtain the potential connection between entities. When constructing the Star-Transformer structure, the identified subject entity E is used as the relay node S t , and the object entity is used as the satellite node H t , for remote relationship modeling. The state update node of each satellite is based on its adjacent nodes, including the previous round representation of the previous node, the previous round representation of the current node, the previous round representation of the next node, and the previous round representation of the current node and the relay node. Calculate the medical relationship features through the attention in step 4.1, and its update process is shown in formulas 12 and 13.

[0120]

[0121] Among them, represents the context information of the r-th node, represents the state of the r-th satellite node in the t-th round, e r represents the medical entity embedding, and MulAtt represents the multi-head attention mechanism, which calculates the attention weights between the current node and the context and updates the state. The relay node s t is updated by the information H of all satellite nodes t and the state s of the previous round t-1 . The relay node S t is updated by the information H of all satellite nodes t and the state s of the previous round t-1 , as shown in Formula 14.

[0122] S t = MulAtt(s t-1 , [s t-1 ; H t , [s t-1 ; H t ) (14)

[0123] Step 4.3: Optimize the information transmission in long texts to prevent distortion in relationship modeling. The Star-Transformer topology adopts a cyclic update mechanism to continuously optimize the information representation of relay nodes and satellite nodes. However, in medical text processing, data such as electronic medical records usually contains long texts. When the number of nodes is too large, the consumption of computing resources by Multi-HeadAttention in the outer loop is too large, and overfitting is likely to occur when the number of loops is large. Therefore, to alleviate the above problems, a residual connection method is used for simplification. While retaining the inherent characteristics of satellite nodes and relay nodes, the depth and complexity of the Star-Transformer topology are reduced at the same time. The calculation method is to establish a residual connection between the information of the previous layer loop and the output of the current attention after calculating the subject-object representations S t , H t in Step 4.2. The specific calculation is shown in Formulas 15 and 16.

[0124] net l = (w1[MulAtt(h f , C f , C f )] + b1 + net l-1 ) (15)

[0125] H i = LayerNorm(σ(net l )) (16)

[0126] Among them, w1 and b1 are learnable parameters, and h f represents the hidden state of the f-th satellite node, and C f represents its context information. The first C f is used as the key K and value V for calculating the attention distribution, and the second C f is used as the value V for providing information weighted calculation. net l-1 is the candidate hidden state of the previous layer, σ is the ReLU activation function, and H i represents the updated satellite node. LayerNorm(·) is the normalization operation.

[0127] Step 4.4, the final satellite node representation and the relay node representation vector obtained by the Star-Transformer. Among them, the relay node representation S t provides the global semantic information of the entire medical text, and the satellite node representation H t captures the feature information of medical entities. These representations are used to predict the start and end positions of the object. The specific calculation method is shown in Formulas 17 and 18.

[0128]

[0129] Among them, and represent the sequence features of the i-th character as the probabilities of the start and end positions of the object. represents the encoded representation vector of the k-th entity detected in the encoding module. is a learnable parameter, is the bias.

[0130] In Chinese medical text data, entities usually contain multiple characters, and using head and tail vectors to represent them is likely to cause information loss. Therefore, to conform to this characteristic, is represented as the average value of the vector representations of all character sequences included in the k-th subject. For each subject, the same decoding process is repeatedly applied to determine the entity span of the medical object. The relation-specific object tagger optimizes the following likelihood function to determine the span of the object corresponding to the given sentence representation and the subject, as shown in Formula 19.

[0131]

[0132] Among them, n is the length of the medical text. For each character i in the medical text, if its true label is the start or end position of the subject (i.e., ), then maximize the probability that the model predicts it as 1 otherwise maximize the probability that the model predicts it as 0 A binary flag indicating the start position of the object corresponding to the i-th character in S Indicates the end position flag of the object corresponding to the i-th character. For the "empty" object o φ , Applies to all labels i. The label parameter

[0133] Example:

[0134] The following combines the accompanying drawings and examples to further describe in detail the specific implementation manners of the present invention. The following examples are used to illustrate the present invention but are not used to limit the scope of the present invention.

[0135] As Figure 1 shown, it is a schematic diagram of a model for an extraction method of the main-object association relationship of text data based on medical field features provided by an embodiment of the present invention.

[0136] Step 1: Extract medical feature information from the input medical text, and perform context encoding using a pre-trained medical BERT model to learn the deep semantic representation of medical terms. The specific operations are as follows:

[0137] 1. Character indexing: Convert the input medical text into an index representation of characters and represent it using One-hot encoding to obtain an index matrix S, that is, the character sequence S of the medical text = {s1, s2, …, s n}

[0138] 2. Medical character embedding: Use the medical character embedding matrix W s and the position embedding matrix W p to embed the character index and calculate the vector representation of each character.

[0139] 3. Multi-layer Transformer semantic enhancement: Through the results of multi-layer Transformer (a total of N layers), obtain the hidden state representation h α of the medical entity layer by layer, referring to the following formula:

[0140] h0 = SW s + W p

[0141] h α = Transformer(h α-1 ), α ∈ [1, N]

[0142] Step 2: Based on the hidden state h α output by the medical BERT encoder, further extract the medical feature representation required for downstream tasks. Further extract semantic information through the model and complete the specific calculations required for the tasks.

[0143] Step 3: Initialization of Medical Entity Tagger and Input Preparation:

[0144] 1. Input Preparation: First, input the character sequence S = {s1, s2, …, s n} of the input sentence into the BERT encoder for feature extraction to obtain the encoded vector h N = BERT(S).

[0145] 2. Bi-LSTM Layer Processing: Use the medical feature representation h N output by BERT as the input and pass it to the Bi-LSTM model to obtain the bidirectional feature representation H N of each character = BiLSTM(h N ).

[0146] Step 4: Binary Classification Prediction of Medical Entity Tagger

[0147] 1. Prediction of the Starting Position of the Entity: For each character position i, use a binary classifier to predict whether it is the starting position of a medical entity. Calculate the probability through the following formula:

[0148]

[0149] 2. Prediction of the Ending Position of the Entity: Similarly, predict whether each character position i is the ending position of a medical entity and calculate the probability:

[0150]

[0151] Example Text: "The patient's CT image shows a pulmonary nodule, and the doctor recommends follow-up." Mark this text as shown in Table 1:

[0152] Table 1 Text Marking Table

[0153] character suffer patient C T image image show show lung part node node , doctor doctor suggest suggestion follow up 。 mark 0 0 0 0 0 0 0 0 1 1 1 1 0 0 0 0 0 0 0 0

[0154] "Pulmonary nodule" is correctly identified as the entity, and the characters from "lung" to "nodule" are marked as 1, and the remaining characters are correctly marked as 0.

[0155] Step 5: Optimization of the Loss Function of the Medical Entity Tagger

[0156] 1. Optimization of the Loss Function: Use the likelihood function to optimize the model parameters, with the goal of maximizing the prediction correctness of the starting and ending position markings of medical entities. The specific optimization loss function:

[0157]

[0158] Example Text: "The patient's MRI image shows a pulmonary nodule.", Mark the above text with its start and end markers as shown in Table 2:

[0159] Table 2 Starting and Ending Marker Table

[0160] character suffer patient M R I image image show show lung part node node 。 start mark 0 0 1 0 0 0 0 0 0 1 0 0 0 0 end mark 0 0 0 0 1 0 0 0 0 0 0 0 1 0

[0161] "MRI" is correctly recognized as the main body, with the starting marker at "M" and the ending marker at "I". "Lung nodule" is correctly recognized as the main body, with the starting marker at "lung" and the ending marker at "nodule". All other characters are non-main body parts, and the starting / ending markers are both 0.

[0162] Step 6: Medical Entity Recognition by the Relationship-Specific Tagger

[0163] 1. Object Tagger: Through the relationship-specific tagger, identify the corresponding medical objects for each detected medical main body. The structure of this tagger is similar to that of the main body tagger, but it focuses on the recognition of relationships between medical entities, such as disease-symptom, drug-indication and other medical relationships.

[0164] Step 7: Initialization of the Star-Transformer Module

[0165] 1. Initialization of the Star-Transformer Topology: Initialize the relay node and satellite nodes of the Star-Transformer. The relay node transmits information through the feature information of all satellite nodes to ensure the global information dissemination of medical texts.

[0166] 2. Node State Update: Each satellite node updates its state according to the following formula See the following formula:

[0167]

[0168] Example text (input sequence): "The patient's MRI image shows a lung nodule." In the example text, "MRI image" may be the examination method, and "lung nodule" may be the examination result. The present invention hopes to use the Star-Transformer to enhance their semantic association. Assume that the text is tokenized and 5 satellite nodes are initialized, as shown in Table 3:

[0169] Table 3 Initialized Node Table

[0170]

[0171] Assume that the vector of each medical term is 2-dimensional (in actual situations, it may be 768-dimensional, such as the vector output by BERT).

[0172] Relay Node S 0 Initialized as the average of all satellite nodes:

[0173]

[0174] Hypothesis: The additional entity information vector e r (such as medical dictionary enhanced information) is e r =(0.05, 0.05). The nodes of the context relationship take the default values (boundary conditions): for h1 (patient), there is no previous node on its left, and it can be filled with a zero vector. Then we have:

[0175]

[0176] Step 8: Relay node information update

[0177] 1. Relay node update: Update the relay node S through the following formula t , where the information is determined by the current states H of all satellite nodes t and the state s of the previous round t-1 . See the following formula:

[0178] S t = MulAtt(s t-1 , [s t-1 ; H t , [s t-1 ; H t )

[0179] Step 9: Residual connection optimization

[0180] 1. Residual connection introduction: Simplify the network structure through the residual connection and avoid overfitting. Connect the output of the previous layer and the multi-head attention result of the current layer through the following formula:

[0181] net l = (w1[MulAtt(h f , C f , C f )] + b1 + net l-1 )

[0182] H i = LayerNorm(σ(net l ))

[0183] Step 10: Prediction of the head and tail markers of medical objects

[0184] 1. Object head and tail prediction: Use the representations of satellite nodes and relay nodes output by Star-Transformer to predict the start and end positions of the object. Calculate through the following formula:

[0185]

[0186] Step 11: Optimization of the loss function of the relationship-specific object tagger

[0187] Optimizing the loss function of the object tagger: The following likelihood function is used to optimize the parameters of the relationship-specific object tagger to improve the accuracy of medical entity relationship recognition:

[0188]

[0189] To illustrate the effect of the present invention, as shown in Table 4, it is a comparison of the effects of whether to use the method for extracting the subject-object association relationship of text data based on medical domain features;

[0190] Table 4 Validation results of the effectiveness of different methods

[0191] model Prec. / % Rec. / % F1 / % the present invention 64.4 51.9 57.3 other models 54.5 49.8 52.1 models without Bi-LSTM 64.2 50.9 56.8 models without ST-Res 55.6 50.6 53.0 models without both Bi-LSTM and ST-Res 54.5 49.8 52.1

[0192] As can be seen from Table 4, the accuracy rate when using the method for extracting the subject-object association relationship of text data based on medical domain features is 57.3%, which is 5.2% higher than other methods. Thus, it can be seen that the method for extracting the subject-object association relationship of text data based on medical domain features, by designing a topological structure, can not only capture the potential association information between the subject and object of text data with medical domain features, but also enhance the representation of the entire sentence, which is beneficial to the improvement of the prediction accuracy. In addition, in terms of ensuring the generalization of the model, ST-ICRE uses an entity tagging component with a linear complexity level, which alleviates the problem of resource occupation when dealing with long sentences.

[0193] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the invention of a method for extracting the subject-object association relationship of text data based on medical domain features provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0194] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in the storage medium, including several instructions to enable a device (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) containing a data processing unit to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0195] The present invention provides an idea and method for extracting the association relationship between the main and object of text data based on the characteristics of the medical field. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

Claims

1. A method for extracting the association relationship between the subject and object of text data based on the characteristics of the medical field, characterized in that, Including: Step 1: Use an encoder to extract feature information from the medical text to be extracted, encode the context information in the medical text, and obtain the encoded vector of the medical text; Step 2: Construct a medical entity tagger, and based on the encoded vector, mark the start position and end position of the entity in the medical text; Step 3: Optimize the medical entity tagger using a likelihood function, and use the optimized medical entity tagger to obtain the list of entities in the medical text; Step 4: Construct a medical relation tagger, model the medical relations of each entity in the entity list, perform semantic enhancement through a star topology structure, further identify the objects related to the entity, and extract relation features; Step 5: Optimize the medical relation tagger using a likelihood function, and use the optimized medical relation tagger to complete the extraction of the subject-object association relationship of the text data based on the characteristics of the medical field.

2. The method for extracting the association relationship between the subject and object of text data based on the characteristics of the medical field according to claim 1, wherein The obtaining of the encoded vector of the medical text described in Step 1 includes: Step 1-1: Assume that the medical text to be extracted is in the form of a character sequence: S = {s1, s2, …, s n} where s n represents the nth character in the character sequence S; the encoder is used to extract the feature information X from the character sequence S, which is expressed as follows: X = SW s Among them, W s is a character embedding matrix used to map each character to a low-dimensional dense vector representation; Step 1-2, use the pre-trained model BERT to encode the context information of the medical text to obtain the encoded vector h of the medical text N , which is expressed as follows: h N = BERT(S) Among them, BERT() is the encoding process of the pre-trained model BERT, which is a stack composed of N identical Transformer blocks.

3. The method for extracting the association relationship between the main and object of text data based on the characteristics of the medical field according to claim 2, characterized in that The encoding of the pre-trained model BERT described in Step 1-2 is specifically represented as follows: h0 = SW s +W p h α = Transformer(h α-1 ), α ∈ [1, N] Among them, h0 represents the initial feature representation of the input medical text, that is, the text feature vector after character embedding and position embedding, h α represents the medical text feature representation of the α-th layer, and Transformer() represents the Transformer block.

4. A method for extracting the association relationship between the subject and object of text data based on the characteristics of the medical field according to claim 3, wherein The marking of the start position and end position of the entity in the medical text described in Step 2 includes: Step 2-1: Use a Bi-LSTM model to construct a medical entity tagger; Step 2-2: Input the encoded vector h of the medical text N into the medical entity tagger to obtain the feature representation H of the medical entity after being processed by the Bi-LSTM model N which is expressed as follows: H N = BilSTM(h N ) Among them, BiLSTM() represents the processing process of the Bi-LSTM model; Step 2-3, denote the feature representation of the $i$-th character in the medical subject entity feature representation $H$ N as $H$ N [$i$], and calculate the probabilities of each character in the medical text being the start and end of the subject entity and Step 2-4, compare the values of and with the threshold value. Mark the position of the i-th character that exceeds the threshold value as 1, otherwise mark it as 0, so as to mark the start position and end position of the main entity in the medical text.

5. A method for extracting the association relationship between the subject and object of text data based on the characteristics of the medical field according to claim 4, wherein Calculating the probabilities of the start and end of each character in the medical text as the main entity as described in Step 2-3 and Specifically as follows: Among them, and represent the probabilities of the start and end of the i-th character as the main entity, where W start and W end are trainable weight parameters for determining whether each character in the medical text is the start or end of the main entity, and b start and b end are the corresponding bias terms, and σ is the Sigmoid activation function that maps the input value to the (0, 1) probability space.

6. The method for extracting the subject-object association relationship of text data based on the characteristics of the medical field according to claim 5, wherein The optimization of the medical entity tagger using a likelihood function described in Step 3 includes: Step 3-1, design the likelihood function p θ (y|S); Step 3-2: Adjust the trainable parameter W in the medical entity tagger through the above likelihood function start and W end such that the probabilities of the start and end of the entity predicted by the medical entity tagger and are closer to the true labels Specifically as follows: For each character i in the medical text, if its true label is the start or end position of the entity, i.e., then maximize the probability that the medical entity tagger predicts it as 1 otherwise maximize the probability that the model predicts it as 0, 1 - Step 3-3: According to Step 3-2, optimize the parameters in the medical entity tagger through gradient descent, obtain the start and end position marks of the entity, and adopt the closest start-end pair matching principle to determine the entity span according to the results of the start and end position marks, which is expressed as follows: Among them, e represents the main entity, S[i start :i end represents the text from the starting character position i start to the ending character position i end , where the starting character position i start satisfies and the ending character position i end satisfies Step 3-3. According to all the main entities e obtained in Step 3-3, obtain the main entity list E, which is expressed as follows: E = {e1e1, …, e m} Among them, e m represents the m-th subject entity.

7. A method for extracting the subject-object association relationship of text data based on the characteristics of the medical field according to claim 6, characterized in that The likelihood function p θ (y|S) described in step 3-1 is as follows: where n is the length of the medical text S, is the true label where the start or end position of the main entity is the i-th character, where q is the tag type indicating the start position start s or the end position end s , and the parameter θ = {W start , b start , W end , v end}.

8. A method for extracting the subject-object association relationship of text data based on the characteristics of the medical field according to claim 7, wherein, The construction of the medical relation tagger described in Step 4 includes: Step 4-1, perform independent attention calculation on the encoded vector h of the medical text through the self-attention mechanism, and integrate the results of multiple attention heads, specifically as follows: N Perform independent attention calculations on the encoded vector h of the medical text through the self-attention mechanism, and integrate the results of multiple attention heads, specifically as follows: MulAtt = concat(z1, z2, …, z h ) · W O Among them, MulAtt is the output of concatenating multiple attention heads, concat(·) represents vector concatenation, and W O is a learnable parameter, and z h represents the attention output of the h-th head, which is specifically as follows: Among them, and are learnable parameters, and Att() is used to calculate a single attention head, specifically as follows: Among them, Q is the query vector, representing the information to be focused on currently, K is the key vector, representing all possible information to be queried, V is the value vector, representing the information content related to the key, and d k represents the dimensions of the query vector and the key vector, where: K = X e W K , V = X e W V Among them, X e is the representation of medical entity features, that is, the feature representation corresponding to the position of the subject entity e in the feature information X; Step 4-2: Adopt a star topology structure to establish the subject-object relationship, specifically as follows: Take each subject entity in the list E of subject entities as the relay node set S t , and take the object entity as the satellite node set H t , and perform calculations through the self-attention mechanism in step 4-1, specifically as follows: Among them, represents the context information of the r-th node in the t-th round, e r represents the medical entity embedding, represents the state of the r-th satellite node in the t-th round, specifically as follows: Among them, MulAtt represents the multi-head attention mechanism in Step 4-1; Update relay node S t , as follows: S t = MulAtt(s t-1 , [s t-1 ; H t , [s t-1 ; H t ) Step 4-3, simplify the star topology using residual connections, that is, for the subject and object representations S t and H t calculated in Step 4-2, establish a residual connection between the information of the previous layer loop and the output of the current attention, and update the satellite node, specifically as follows: H’ = LayerNorm(σ(net l )) where σ is the ReLU activation function, H’ represents the updated satellite node, LayerNorm(·) is the normalization operation, and net l is the candidate hidden state of the current layer, as follows: net l =(w1[MulAtt(h f ,C f ,C f )]+b1+net l-1 ) Among them, w1 and b1 are learnable parameters, h f represents the hidden state of the f-th satellite node, C f represents its context information. The first C f is used as the key K and value V for calculating the attention distribution. The second C f is used as the value V to provide information weighted calculation, net l-1 is the candidate hidden state of the previous layer; Step 4-4: Through the simplified star topology structure in Step 4-3, obtain the final satellite node representation and relay node representation, that is, the sentence vector representation and entity representation of the medical text, and predict the head and tail marks of the object; Step 4-5, use the average value of the global character-level representation as the main entity representation, that is, is represented as the average value of all character sequence vector representations included in the k-th main body.

9. A method for extracting the association relationship between the main and object of text data based on the characteristics of the medical field according to claim 8, characterized in that The prediction of the head and tail marks of the object described in Step 4-4 is specifically as follows: Among them, and represents the probability of the sequence feature of the i-th character as the start and end positions of the object; represents the encoded representation vector of the k-th entity detected in the encoder; and are learnable parameters, and are biases.

10. A method for extracting the association relationship between the main and object of text data based on the characteristics of the medical field according to claim 8, wherein, The optimization of the medical relation tagger using a likelihood function described in Step 5 includes: Design likelihood function The details are as follows: Among them, n is the length of the medical text. For each character i in the medical text, if its true label is the start or end position of the subject, that is then maximize the probability that the medical relation tagger predicts it as 1 otherwise maximize the probability that the medical relation tagger predicts it as 0 is the binary tag of the start position of the object of the i-th character in the medical text S, indicating the object end position tag of the i-th character; for the empty object o φ , is applicable to all labels i; the label parameter

Citation Information

Cited By

  • Information extraction method, device and equipment for medical text and medium

    CN121528577A