A cross-specialty text structuring method based on language models

Through a cross-secondary text structured method based on language model, text structured tasks in electronic medical records are converted into generative question-and-answer tasks, and field values are generated using template terms separation and decoder, which solves the transferability and flexibility of cross-secondary tasks in the existing technology, and improves the accuracy and efficiency of text structure.

CN113408276BActive Publication Date: 2025-07-18EAST CHINA UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110678708.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-18
Publication Date
2025-07-18
Estimated Expiration
2041-06-18

AI Technical Summary

Technical Problem

The existing text structured methods have problems in the electronic medical records, different extraction goals for different applications, poor algorithm generalization capabilities, and relying on large-scale artificial corpus annotation, making it difficult to achieve the transferability and flexibility of cross-specialist tasks.

Method used

The cross-college text structure method based on language model is adopted to treat medical record text as knowledge, and the field names are transformed into a problem. The model is asked to answer the corresponding field values, convert three types of text structured tasks into generative question-and-answer tasks, and process them using template term separator and decoder to generate corresponding field values through Self Attention, Query Cross Attention, and Text Cross Attention.

Benefits of technology

It realizes the transferability of cross-specialist tasks and the flexibility of different types of tasks, improves the accuracy and efficiency of text structure, and reduces the dependence on manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113408276B_ABST
    Figure CN113408276B_ABST
Patent Text Reader

Abstract

The present invention proposes a cross-specialty text structuring method based on a language model. For the task of structuring electronic medical record texts, from the perspective of extraction targets, it can be divided into three categories: classification type, text span type, and generation type. The present invention uses an end-to-end text structuring method, regards the medical record text as knowledge, constructs questions after making certain transformations to the field names, allows the model to answer the corresponding field values, and converts the three types of text structuring tasks into generative question-and-answer tasks, so that the algorithm has the transferability of cross-specialty tasks and the flexibility to solve different types of tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text structuring technology, and particularly to a language model-based text structuring method for electronic medical records. Background Art

[0002] With the continuous advancement of medical informatization, the Electronic Health Record (EHR) of patients has become increasingly perfect. EHR is a digital record centered on personal health, healthcare, and treatment, which collects and stores patients' health information and medical visit information in a digital form.

[0003] In EHR, a large number of medical documents exist in text form. The text of the electronic medical record contains the patient's symptoms, examination results, as well as the diagnosis and description of the treatment process made by doctors based on basic data such as symptoms and physical and chemical indicators. Compared with well-structured medication records, test results, etc., clinical texts record the judgment basis of doctors and the effect tracking of various medical treatment behaviors, and these important information are stored in unstructured information and cannot be understood and processed by computers.

[0004] Existing text structuring, such as named entity recognition or relation extraction algorithms based on models like BERT, CRF, LSTM pre-trained in the general domain, all face challenges such as inaccurate problem definition, different extraction targets for different applications, poor algorithm generalization ability, and dependence on large-scale manual corpus annotation. Summary of the Invention

[0005] Aiming at the deficiencies in existing text structuring solutions, the present invention provides a cross-specialty text structuring method based on a language model, which uses an end-to-end text structuring method. Regarding the medical record text as knowledge, after transforming the field names, constructing them into questions, and allowing the model to answer the corresponding field values, the three types of text structuring tasks are converted into generative question-and-answer tasks, unifying different types of structuring tasks, and enabling the algorithm to have the transferability of cross-specialty tasks and the flexibility to solve different types.

[0006] The present invention adopts the following technical solutions:

[0007] 1. A cross-specialty text structuring method based on a language model, the main idea of which includes the following steps:

[0008] S1: After converting the question into a question-and-answer form, use a template term separator, take the medical knowledge base as a dictionary, match the corresponding term part from the medical text S, replace it, and generate a text template and a set of professional terms.

[0009] S2: Input the text template and professional terms, and obtain a vector representation after fusion;

[0010] S2: The decoder part, based on the medical record text representation and the field name representation, successively uses Self Attention, Query Cross Attention, and Text Cross Attention to capture the previous context information, field information, and medical record text information, and generates the corresponding field values through a gating unit.

[0011] 2. Specifically, in the step S1, the method for constructing the text template and professional terms includes the following steps:

[0012] S11: The input is the medical record text S Doc and the field name S key , and the output is the field value S value . Different processing methods are determined according to the field type. Specifically, the tasks can be divided into the following three categories:

[0013] (1) Classification type: It means that the candidate values of the field to be extracted are limited. For example, for the field of "gastric cancer surgery technique", its candidate values are: {"open surgery", "laparoscopic assistance", "total laparoscopy", "thoracoabdominal combined"};

[0014] (2) Text span type: It means that the field to be extracted is a section of text in the medical record. For example, for "upper resection margin distance", its field value is a corresponding span in the text;

[0015] (3) Generation type: It means that the field to be extracted has neither candidate values nor is it a section of text in the medical record, and needs to be regenerated from start to end. For example, for "symptom list", its field value is the normalized symptom words in the chief complaint text.

[0016] S12: If it is a classification type task, then enumerate the corresponding candidate values and convert it into a cloze problem; build a trie tree according to the candidate value set and convert it into a problem of generating a restricted candidate word list;

[0017] S13: If it is a text span type task, then enumerate all possible text spans and convert it into a cloze problem. The decoding process introduces a gating mechanism, and through additional neurons, calculates whether a word in the original text should be selected currently, and converts it into a problem of directly generating a word in the original text;

[0018] S14: If it is a generation type task, the Beam Search algorithm is introduced in the decoding process, and by expanding a certain amount of search space, it is converted into the corresponding generation problem;

[0019] S15: Using the trie tree matching algorithm, with the medical knowledge base KG as the dictionary, match the corresponding term part from the medical text S, and then replace it to generate the text template S pattern and the professional term set SKG 。

[0020] 3. Specifically, in step S2, constructing the language model includes the following steps:

[0021] S21: Construct a transferable language model based on the electronic medical record text, with the input being the medical record text D in and the medical knowledge graph KG. The output in the pre-training stage is the loss of the downstream task, and the output in the fine-tuning stage is the fused vector representation E l+1 ;

[0022] S22: Use the pre-trained language model, input the text template and professional terms, and output the fused vector representation E l+1 , where the template term encoder uses Patten Attention and KG Cross Attention to capture the context semantic information of the template and the association information between the template and the knowledge base in sequence, and then uses the FNN layer to perform a non-linear transformation on it to obtain the fused vector representation E l+1 ; The specific formula is as follows,

[0023] SelfAttention(X) = ln(mult_head h=12 (X, X, X, MASK) + X)

[0024] KGCrossAttention(X, K) = ln(mult_head h=12 (X, K, K, MASK) + X)

[0025] E l+1 = FFN(KGCrossAttention(Self Attention(E l ), K, MASK))

[0026] E 1 = layer_normal(add([x i s_max , [p i s_max ))

[0027]

[0028]

[0029] Among them, E 1 is the initial vector, originating from [x i obtained by mapping the text X through word vectors s_max and the corresponding position encoding [p i s_max ​​​, K is the set of professional terms S KG each word in the vector representation of, MASK is the mask matrix, and the number of layers l ∈ {x|1 ≤ x ≤ 12}.

[0030] 4. Specifically, in the step S3, the algorithm of the decoder includes the following steps:

[0031] S31: The decoder captures the information of the previous context, field information, and medical record text information in sequence according to the medical record text representation and the field name representation using SelfAttention, Query Cross Attention, and Text Cross Attention respectively, and generates the corresponding field value S through the gating unit value . The specific formula is as follows.

[0032] SelfAttention(X) = ln(mult_head h=12 (X, X, X, Tri) + X)

[0033]

[0034]

[0035] D l+1 = FFN(TextCrossAttention(QueryCrossAttention(SelfAttention(D l ), H N ), E))

[0036] D 1 = layer_normal(add([x i s_max , [p i s_max ))

[0037]

[0038]

[0039]

[0040] S32: Use the decoding algorithm based on the trie tree. When , it means that y i selects a word from the vocabulary; when , it means that y i selects a word from the medical text. The loss function of the final model is as follows. ​​

[0041] Description of the Drawings

[0042] After reading the specific embodiments of the present invention with reference to the drawings, various aspects of the present invention will be more clearly understood. Among them,

[0043] Figure 1 A schematic diagram showing a template and term separation generative extraction model based on a medical knowledge base according to an embodiment of the present invention. Detailed Description of the Invention

[0044] In order to make the technical content disclosed in this application more detailed and complete, reference may be made to the accompanying drawings and the following various specific embodiments of the present invention. The same reference numerals in the drawings represent the same or similar components. However, those of ordinary skill in the art should understand that the embodiments provided below are not used to limit the scope covered by the present invention. In addition, the drawings are only schematically illustrated and are not drawn according to their original sizes.

[0045] The following further describes in detail the specific embodiments of various aspects of the present invention with reference to the drawings.

[0046] Figure 1 A schematic diagram showing a template and term separation generative extraction model based on a medical knowledge base according to an embodiment of the present invention.

[0047] The main idea of the template and term separation generative extraction model based on a medical knowledge base includes the following steps:

[0048] S1: A template term separator, using a medical knowledge base as a dictionary, matches the corresponding term part from the medical text S, replaces it, and generates a text template and a set of professional terms;

[0049] S2: Input the text template and professional terms, and output the fused vector representation;

[0050] S3: According to the medical record text representation and the field name representation, the decoder sequentially uses Self Attention, QueryCross Attention, and Text Cross Attention to capture the above context information, field information, and medical record text information, and generates the corresponding field value through a gating unit. Further, generating the text template and the set of professional terms in step S1 is a process of constructing a question and answer. The specific method includes the following steps:

[0051] S11: The input is the medical record text S Doc and the field name S key , and the output is the field value S value, different processing methods are determined according to the field type. Specifically, the tasks can be divided into the following three categories:

[0052] (1) Classification type: It means that the candidate values of the fields to be extracted are a finite number of fields. For example, for the field of "gastric cancer surgery techniques", its candidate values are: {"open surgery", "laparoscopic assistance", "total laparoscopy", "thoracoabdominal combined"};

[0053] (2) Text span type: It means that the field to be extracted is a segment of text in the medical record. For example, for "upper resection margin distance", its field value is a corresponding span in the text;

[0054] (3) Generation type: It means that the field to be extracted has neither candidate values nor is it a segment of text in the medical record, and needs to be regenerated from beginning to end. For example, for "symptom list", its field value is the normalized symptom words in the chief complaint text.

[0055] S12: If it is a classification type task, enumerate the corresponding candidate values and convert it into a cloze problem; build a trie tree based on the candidate value set and convert it into a generation problem of restricting the candidate word list;

[0056] If it is a text span type task, enumerate all possible text spans and convert it into a cloze problem. The gating mechanism is introduced in the decoding process, and an additional neuron is used to calculate whether a word in the original text should be selected currently, and it is converted into a generation problem that can directly generate a word in the original text;

[0057] If it is a generation type task, the Beam Search algorithm is introduced in the decoding process, and by expanding a certain amount of search space, it is converted into the corresponding generation problem;

[0058] S13: Use the trie tree matching algorithm, with the medical knowledge base KG as the dictionary, match the corresponding term part from the medical text S, and then replace it to generate the text template S pattern with the professional term set S KG .

[0059] For example, for the gastric cancer surgery process text: "Exploration: liver, gallbladder, pancreas, spleen, small and large intestines are normal, the lesion is located in the anterior wall of the lesser curvature of the stomach", it can be decomposed into the template: "Exploration [body organ] [abnormal condition], the lesion is located in [body organ]" and the terms: "liver", "gallbladder", "pancreas", "spleen", "small and large intestines", "anterior wall of the lesser curvature of the stomach" and "normal" combination.

[0060] 1. The method for obtaining vector representation includes the following steps:

[0061] S21: Construct a transferable language model based on the electronic medical record text, and its input is the medical record text D inAnd a medical knowledge graph KG. The output in the pre-training stage is the loss for downstream tasks, and the output in the fine-tuning stage is the fused vector representation E l+1 。

[0062] Compared with RNN-type models, using pre-trained language models can capture more context information of text, and can achieve better performance when facing question-and-answer tasks after being processed by S1. After pre-training on electronic medical record text, it can be better transferred between different specialties.

[0063] S22: Use the pre-trained language model, input the text template and professional terms, and output the fused vector representation E l+1 。The template term encoder uses Patten Attention and KG Cross Attention to capture the context semantic information of the template and the association information between the template and the knowledge base in sequence, and then uses the FNN layer to perform a non-linear transformation on it to obtain the fused vector representation E l+1 。The specific formula is as follows, where the number of layers l ∈ {x|1 ≤ x ≤ 12}.

[0064] SelfAttention(X) = ln(mult_head h=12 (X,X,X,MASK)+X)

[0065] KGCrossAttention(X,K) = ln(mult_head h=12 (X,K,K,MASK)+X)

[0066] El +1 = FFN(KGCrossAttention(Self Attention(E l ),K,MASK))

[0067] E 1 = layer_normal(add([x i s_max ,[pi] s_max ))

[0068]

[0069]

[0070] Among them, E 1 is the initial vector, which comes from [x i obtained by mapping the text X through word vectors s_max and the corresponding position encoding [p i s_max 。K is the set of professional terms S​​KG The vector representation of each word in MASK is a mask matrix that can control the attention range of each word and is used here in KG Cross Attention to make the template only focus on the vector representation of the corresponding replacement position.

[0071] 2. In the step S3, the decoder includes the following steps:

[0072] S31: The decoder successively uses SelfAttention, Query Cross Attention, and Text Cross Attention to capture the above - context information, field information, and medical record text information based on the medical record text representation and the field name representation to generate the corresponding field value S through a gating unit. value . The specific formula is as follows.

[0073] SelfAttention(X) = ln(mult_head h=12 (X,X,X,Tri)+X)

[0074]

[0075]

[0076] D l+1 = FFN(TextCrossAttention(QueryCrossAttention(SelfAttention(D l ),H N ),E))

[0077] D 1 = layer_normal(add([x i s_max ,[p i s_max ))

[0078]

[0079]

[0080]

[0081] S32: When , it means that y i selects a word from the vocabulary; when , it means that y i selects a word from the medical text. The loss function of the final model is as follows. ​​

[0082]

[0083] To solve the problem that values outside the range of candidate words appear in the generative decoding process during actual use, for this situation, the present invention designs a decoding algorithm based on a trie tree to ensure that the generated candidate values must be within the corresponding range. The algorithm is as follows.

[0084]

Claims

1. A cross-specialty text structuring method based on a language model, characterized in that It includes the following steps: S1: Using a template term separator and taking a medical knowledge base as a dictionary, match the corresponding term part from the medical text S, replace it, and generate a text template and a collection of professional terms; S2: Input the text template and professional terms, and obtain a vector representation after fusion; S3: On the basis of using a language model, the decoder successively adopts Self Attention, Query Cross Attention, and Text Cross Attention to capture context information, field information, and medical record text information according to the medical record text representation and the field name representation, and generates corresponding field values through a gating unit; In the step S1, the method for constructing a text template and a collection of professional terms includes the following steps: S11: The input is the medical record text S Doc , the field name S key and the field value S value , and the output is the field value S value , and three task types are determined according to the field type, and the task types include classification tasks, text span tasks, and generation tasks; S12: According to different types of tasks, transform the tasks, specifically: If it is a classification task, enumerate the corresponding candidate values and convert it into a cloze problem; build a trie tree according to the candidate value set and convert it into a problem of generating a restricted candidate word list; If it is a text span task, enumerate all possible text spans, convert it into a cloze problem, introduce a gating mechanism in the decoding process, calculate through an additional neuron whether a word in the original text should be selected at present, and convert it into a generation problem that can directly generate a word in the original text; If it is a generation task, introduce the Beam Search algorithm in the decoding process, and convert it into a corresponding generation problem by expanding a certain amount of search space; S13: Using the trie matching algorithm, with the medical knowledge base KG as the dictionary, match the corresponding term part from the medical text S, and then replace it to generate the text template S pattern with the professional term set S kG .

2. The cross-specialty text structuring method based on a language model according to claim 1, wherein: In the step S2, the method for obtaining a vector representation includes the following steps: S21: Construct a transferable language model based on electronic medical record texts, with the input being medical record text D in and medical knowledge graph KG. The output in the pre-training stage is the loss for downstream tasks, and the output in the fine-tuning stage is the fused vector representation E l+1 ; S22: Use the pre-trained language model, input the text template and professional terms, and output the fused vector representation E l +1 , where the template term encoder uses PattenAttention and KG CrossAttention to capture the context semantic information of the template and the association information between the template and the knowledge base in sequence, and then uses the FNN layer to perform a non-linear transformation on it to obtain the fused vector representation E l+1 ; The specific formula is as follows SelfAttention(X) = ln(mult_head h=12 (X, X, X, MASK) + X) KGCrossAttention(X, K) = ln(mult_head h=12 (X, K, K, MASK) + X) E l+1 = FFN(KGCrossAttentton(SelfAttentton(E l ), K, MASK)) E 1 = layer_normal(add([x i s_max , [p i s_max ))​​ Among them, E 1 is the initial vector, originating from the word vector mapping of text X [x i s_max and the corresponding position encoding [p i s_max , K is the vector representation of each word kG in the set S of professional terms , MASK is the mask matrix, and the number of layers l ∈ {x|1 ≤ x ≤ 12}.​​ 3. The cross-specialty text structuring method based on a language model according to claim 2, characterized in that: In the step S3, the decoder construction process includes the following steps: S31: The decoder, based on the medical record text representation and the field name representation successively uses Self Attention, Query Cross Attention, and Text Cross Attention to capture context information, field information, and medical record text information, and generates the corresponding field value S through a gating unit value , and its specific formula is as follows: SelfAttention(X) = ln(mult_head h=12 (X, X, X, Tri) + X) D l+1 = FFN(TextCrossAttention(QueryCrossAttention(SelfAttentton(D l ), H N ), E)) D 1 = layer_normal(add([x i s_max , [p i s_max ))​​ S32: When , it means that yi selects a word from the vocabulary; when , it means that yi selects a word from the medical text; the loss function of the final model is:

Citation Information

Patent Citations

  • Clinical terms mining method, device, electronic apparatus and computer-readable medium

    CN109522338A