A Multi-Granularity Entity Recognition Method for Pathological Text Naming

By using a multi-grained entity recognition method in pathological text, and using the Bert model and central surrogate words/words for entity recognition, the problem of low efficiency and accuracy of naming entity recognition in pathological text is solved, especially in the case of few sample data, which significantly improves the generalization ability and entity recognition effect of the model.

CN115587595BActive Publication Date: 2025-06-24CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211380333.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-06-24
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

In pathological texts, it is difficult for the prior art to effectively identify named entities, especially in the case of small sample data, which leads to insufficient generalization capabilities of the model and low entity recognition efficiency and accuracy.

Method used

The multi-grained entity recognition method is used to segment the pathological text through word-grainedness and word-grainedness, encode it using the Bert model, and entity recognition is carried out using the central surrogate word and central surrogate word, and KEloss and CEloss loss functions are constructed to optimize the model.

Benefits of technology

It improves the accuracy of entity recognition in pathological text, alleviates the problem of insufficient generalization capabilities of model under small sample data, reduces the gap between pre-training and downstream tasks, and improves the effect of named entity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115587595B_ABST
    Figure CN115587595B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of natural language processing, and particularly relates to a multi-granularity entity recognition method for pathological text naming. The method includes: obtaining pathological text information, and segmenting the pathological text at the character granularity and the word granularity; performing random mask masking and vector initialization on the segmented text, and using two parameter-sharing Bert models to encode the text after random mask masking and vector initialization; presetting a central replacement word and a central replacement character for each entity of each category; using KL loss and CE loss to construct loss functions for the character granularity and the word granularity, optimizing the CE loss for calculating the loss of the replaced character granularity, and optimizing the KE loss for calculating the loss of the replaced word granularity to obtain the entity recognition result. The present invention constructs templates through the character granularity and the word granularity for prediction, can accurately identify and extract the entities of the pathological text, and has a good entity recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a multi-granularity entity recognition method for pathological text naming. Background Art

[0002] Named entity recognition is to structurally process the entities contained in the text into an organizational form like a table. The input to the named entity recognition system is the original text, and the output is entities in a fixed format; the entities are extracted from various documents and then integrated together in a unified form. The named entity recognition technology does not attempt to comprehensively understand the entire document, but only analyzes the part of the document that contains relevant entities.

[0003] Natural language processing is an important direction in the fields of computer science and artificial intelligence; natural language processing is to achieve natural language communication between humans and machines, and the research in this field will involve natural language, that is, the language used by people in daily life.

[0004] The task of named entity recognition for pathological text can be defined as a medical-related field, so the training corpus is also related to medicine. The continued pre-training of the field is to perform relevant field training on the original bert model to make the trained model more adaptable to the relevant field and also improve the performance of downstream tasks.

[0005] In order to enable the model to fit the provided training data as much as possible, corresponding metrics are needed to evaluate its fitting degree, and the function used is called the loss function. When the value of the loss function decreases, the fitting degree of the model is better.

[0006] In recent years, natural language processing technology (NLP) has developed rapidly and has been applied in various fields, including pathological artificial intelligence. In traditional clinical diagnosis, doctors want to understand the pathological status of patients by extracting information from pathological texts themselves, which not only consumes a lot of energy but also has low efficiency. If NLP technology can accurately label the entities that doctors are concerned about, it can greatly improve the efficiency of doctors. Moreover, the extracted data can also be used as scientific research data, and researchers can mine pathological information such as multi-relations through pathological texts. And in today's pathological data environment, the problem of few samples is often faced, so named entity recognition for few-sample pathological texts has become a very urgent task nowadays. Summary of the Invention

[0007] To solve the above technical problems, the present invention proposes a multi-granularity entity recognition method for pathological text naming, including:

[0008] S1: Obtain pathological text information and segment the pathological text according to character granularity and word granularity;

[0009] S2: Randomly mask the segmented text and initialize the vectors. Use a Bert model with shared parameters to encode the text after random masking and vector initialization to obtain the character encoding sequence of the pathological text data;

[0010] S3: Preset a central replacement word and a central replacement character for each entity of each category;

[0011] S4: Predict the masked words and characters in the character encoding sequence. Replace the masked entity characters with the preset central replacement characters of the corresponding category, replace the masked entity words with the preset central replacement words of the corresponding category, and replace non-entity words and characters with the original words and characters;

[0012] S5: Use KEloss and CEloss to construct loss functions for character-level and word-level. CEloss optimizes by calculating the loss for the replaced character-level, and KEloss optimizes by calculating the loss for the replaced word-level to obtain the optimized entity recognition result.

[0013] Preferably, each pathological text data in the dataset is segmented according to character-level and word-level, including:

[0014] For each pathological text in the dataset, segment it by individual characters, and at the same time use the jieba segmentation tool to segment it by each word.

[0015] Preferably, randomly mask the segmented text and initialize the vectors, including:

[0016] Randomly select words and characters from the text segmented at character-level and word-level at a masking rate of 15%, mask the selected words and characters, and use the Bert pre-trained model to initialize the vectors of all words and characters for the text after random masking.

[0017] Preferably, encode the text, including:

[0018] Truncate the text with a length greater than 512, and pad the text with a length less than 512 with 0 to obtain a character encoding sequence with a length of 512 for all.

[0019] Preferably, preset a central replacement word and a central replacement character for each entity of each category, expressed as:

[0020]

[0021] Among them, z c represents the central replacement word or central replacement character of the entity word, represents the Euclidean distance sum from the entity word or character vector w i to other entity word or character vectors, Wc The set of entity words or characters representing category c, argmin represents the value of w when taking the minimum i of w ik represents the vector value of the k-th dimension of the i-th entity word or character vector, w jk represents the vector value of the k-th dimension of the j-th entity word or character vector, and n represents the dimension of the entity word or character vector.

[0022] Preferably, the joint loss function constructed using KL loss and CE loss for character-level and word-level includes:

[0023] CELoss optimizes the loss calculation for character-level:

[0024]

[0025] KE loss optimizes the loss calculation for word-level:

[0026]

[0027] Among them, CELoss [MASK] represents the loss function of the character-level mask, p(y|x) represents the possible probability distribution at the mask position, p(y|x) = p([MASK] = Z c |V'), p([MASK] represents the prediction of the mask position, Z c represents the replacement word of the entity character, V' represents the original text of the non-entity masked, y represents the label, X represents the sample set, x represents the original entity character, |X| represents the number of samples, KELoss represents the word-level divergence loss function, r represents the subtraction of the first word vector and the last word vector of the central replacement word, d(h,t) represents the subtraction of the first word vector and the last word vector of the positive example word, d(h,t) = ||h - t||2, n represents the number of negative example words, h represents the first word of the positive example word, t represents the last word of the positive example word, h i ' represents the first word of the i-th negative example word, t i ' represents the last word of the i-th negative example word.

[0028] Advantageous effects of the present invention:

[0029] The present invention uses two different granularities, character-level and word-level, to enrich more pathological text information features, and uses the central replacement word and the central replacement character for named entity recognition, improving the accuracy of entity recognition.

[0030] The present invention aims to alleviate the problem of insufficient labeled data in pathological texts and improve the generalization ability of the model in the case of few samples. In the pre-training stage of the present invention, word-level and character-level granularities are used to enrich more pathological text information features from different granularities. When performing continued pre-training, central replacement words and central replacement characters are used to perform the training task of named entity recognition, reducing the gap between the pre-training stage and the downstream task and effectively improving the effect of named entity recognition of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a framework diagram of a multi-granularity entity recognition method for pathological text naming according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] A multi-granularity entity recognition method for pathological text naming, as Figure 1 shown, includes:

[0034] S1: Obtain pathological text information, and segment the pathological text according to word-level and character-level granularities;

[0035] S2: Perform random mask masking and vector initialization on the segmented text, and use two parameter-sharing Bert models to encode the text after random mask masking and vector initialization to obtain the character encoding sequence of the pathological text data;

[0036] S3: Preset central replacement words and central replacement characters for each entity of each category;

[0037] S4: Perform prediction of masked words and characters on the character encoding sequence, replace the masked entity characters with the preset central replacement characters of the corresponding category, replace the masked entity words with the preset central replacement words of the corresponding category, and replace the non-entity words and characters with the original words and characters;

[0038] S5: Use KEloss and CEloss to construct loss functions for word-level and character-level granularities. CEloss optimizes the loss calculation for the replaced word-level granularity, and KE loss optimizes the loss calculation for the replaced character-level granularity to obtain the optimized entity recognition result.

[0039] Preferably, each pathological text data in the dataset is segmented according to word-level and character-level granularities, including:

[0040] For each piece of pathological text in the dataset, perform word segmentation on a single-character basis and, at the same time, use the Jieba word segmentation tool to perform word segmentation on a per-word basis.

[0041] Preferably, perform random mask masking and vector initialization on the segmented text, including:

[0042] Randomly select words and characters from the text segmented at the character level and word level at a masking rate of 15%, mask out the selected words and characters, and use the Bert pre-trained model to perform vector initialization on all words and characters in the text after random mask masking.

[0043] Preferably, encode the text, including:

[0044] Truncate the text with a length greater than 512 and pad the text with a length less than 512 with 0 to obtain a character encoding sequence with a length of 512 for all.

[0045] Preferably, use the Universal Sentence Encoder semantic model to calculate the entity vectors in the training set, and use the Euclidean distance to find the central substitute word for each category, with the sum of the distances from this word to other words being the smallest:

[0046]

[0047] Among them, z c represents the central substitute word or character of the entity word, represents the entity word or word vector w i the sum of the Euclidean distances to other entity words or word vectors, W c represents the set of entity words or characters of category c, and argmin represents the value of w when taking the minimum, w i represents the i-th entity word or the vector value of the k-th dimension of the word vector, w ik represents the vector value of the k-th dimension of the j-th entity word or word vector, and n represents the dimension of the entity word or word vector. jk c

[0048] Preferably, use the KE loss and CE loss to construct a joint loss function for the character level and word level, including:

[0049] The CEloss optimizes the calculation of the loss for the character level:

[0050] p(y|x) = p([MASK] = Z c |V')

[0051]

[0052] KE loss optimizes the loss calculation for word granularity:

[0053]

[0054] d(h,t) = ||h - t|| p

[0055] Among them, CELoss [MASK] represents the loss function of the character granularity mask, p(y|x) represents the possible probability distribution of the mask position, p(y|x) = p([MASK] = Z c |V'), [MASK] represents predicting the mask position of the mask, Z c represents the replacement word of the entity character, V' represents the original text masked by non-entities, y represents the label, X represents the sample set, x represents the original entity character, |X| represents the number of samples, KELoss represents the word granularity divergence loss function, r represents the subtraction of the first character vector and the last character vector of the central replacement word, d(h,t) represents the subtraction of the first character vector and the last character vector of the positive example word, n represents the number of negative example words, h represents the first character of the positive example word, t represents the last character of the positive example word, h′ i represents the first character of the i-th negative example word, t′ i represents the last character of the i-th negative example word.

[0056] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-granularity entity recognition method for pathological text naming, characterized in that, Including: S1: Obtain pathological text information and segment the pathological text at the character granularity and word granularity. S2: Perform random mask masking and vector initialization on the segmented text, and use a Bert model with shared parameters to encode the text after random mask masking and vector initialization to obtain the character encoding sequence of the pathological text data. S3: Preset a central replacement word and a central replacement character for each entity of each category. Preset a central replacement word and a central replacement character for each entity of each category, expressed as: Among them, z c represents the central substitute word or character of the entity word, represents the Euclidean distance sum from the entity word or word vector w i to other entity words or word vectors, W c represents the set of entity words or characters of category c, and argmin represents the value of w when it is the minimum, i w ik represents the vector value of the k-th dimension of the i-th entity word or word vector, w jk represents the vector value of the k-th dimension of the j-th entity word or word vector, and n represents the dimension of the entity word or word vector; S4: Predict the masked words and characters in the character encoding sequence. The masked entity characters are replaced with the preset central replacement characters of the corresponding category, the masked entity words are replaced with the preset central replacement words of the corresponding category, and the non-entity words and characters are replaced with the original words and characters. S5: Use KEloss and CEloss to construct loss functions for the character granularity and word granularity. CEloss optimizes the loss calculation for the replaced character granularity, and KE loss optimizes the loss calculation for the replaced word granularity to obtain the optimized entity recognition result. The joint loss function constructed by using KEloss and CEloss for the character granularity and word granularity includes: CEloss optimizes the loss calculation for the character granularity: KE loss optimizes the loss calculation for the word granularity: Among them, CELoss [MASK] represents the loss function of character-level masks. p(y|x) represents the possible probability distribution of the mask positions, p(y|x) = p([MASK]=Z c |V'), [MASK] represents the prediction of the mask positions, Z c represents the replacement word of the entity word, V' represents the original word of the non-entity that is masked, y represents the label, X represents the sample set, x represents the original entity word, |X| represents the number of samples, KELoss represents the word-level divergence loss function, r represents the subtraction of the first word vector and the last word vector of the central replacement word, d(h,t) represents the subtraction of the first word vector and the last word vector of the positive example word, d(h,t) = ||h - t||2, n represents the number of negative example words, h represents the first word of the positive example word, t represents the last word of the positive example word, h i ' represents the first word of the i-th negative example word, t i ' represents the last word of the i-th negative example word.

2. The multi-granularity entity recognition method for pathological text naming according to claim 1, wherein, Segment each pathological text data in the dataset at the character granularity and word granularity, including: Segment each pathological text in the dataset into single characters, and at the same time use the jieba word segmentation tool to segment each word.

3. A multi-granularity entity recognition method for pathological text naming according to claim 1, characterized in that, Perform random mask masking and vector initialization on the segmented text, including: Randomly select words and characters from the text segmented at the character granularity and word granularity at a masking rate of 15%, mask the selected words and characters, and use the Bert pre-trained model to perform vector initialization on all words and characters in the randomly masked text.

4. The multi-granularity entity recognition method for pathological text naming according to claim 1, wherein, Encode the text, including: Truncate the text with a length greater than 512, and pad the text with a length less than 512 with 0 to obtain a character encoding sequence with a length of 512 for all.

Citation Information

Patent Citations

  • Entity name normalization system and method and computer readable medium

    CN112613318A

  • Entity identification method and system based on multi-granularity feature fusion and uncertain denoising

    CN113627172A