Discontinuous and overlapping named entity recognition system

By converting text into directed graph structures and using dependencies to enhance the representation, the problem of low accuracy of non-continuous and overlapping entity recognition in the prior art is solved, and accurate recognition and speed improvement are achieved.

CN120235153APending Publication Date: 2025-07-01SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311848889.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify discontinuous and overlapping named entities, with low recognition accuracy and slow training and reasoning speed.

Method used

Using one-dimensional text representation unit, dependency injection unit and two-dimensional text representation unit, the non-continuous and overlapping entities are decoded using a multi-layer perceptron by converting the input sentence pattern into a directed graph structure.

Benefits of technology

Unambiguous recognition of discontinuous and overlapping entities is achieved, which improves recognition accuracy and speeds up training inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235153A_ABST
    Figure CN120235153A_ABST
Patent Text Reader

Abstract

A discontinuous and overlapped named entity recognition system comprises a one-dimensional text representation unit, a dependency relationship injection unit, a two-dimensional text representation unit and an extraction decoding unit, and adopts two-dimensional text representation and uses a character-character relationship to extract discontinuous and overlapped named entities. After the input sentence pattern is converted into the directed graph, the overlapped entities are represented in an unambiguous mode, discontinuous and overlapped entities can be effectively recognized, the recognition accuracy is obviously improved, the reasoning speed is obviously improved through training, and the method has wide application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of natural language processing, specifically a system for extracting discontinuous and overlapping named entities from text. Background Art

[0002] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. Named entity recognition (NER) refers to the recognition and extraction of entities with specific meanings in an article, which includes two subtasks: the recognition of continuous entities and the recognition of discontinuous entities. The current named entity recognition methods mainly focus on the recognition of continuous entities. The present invention realizes a method for recognizing discontinuous and overlapping entities. Summary of the Invention

[0003] Aiming at the defects of the prior art that it cannot recognize discontinuous and overlapping entities, there are ambiguities in the representation of discontinuous and overlapping entities, the recognition accuracy is relatively low, and the training and inference speeds are relatively slow, the present invention provides a discontinuous and overlapping named entity recognition system. The system adopts a two-dimensional text representation and uses word-word relationships to extract discontinuous and overlapping named entities. After converting the input sentence pattern into a directed graph, it can represent overlapping entities unambiguously and can effectively recognize discontinuous and overlapping entities. The recognition accuracy is significantly improved, and the training and inference speeds are also significantly improved, and it has a wide range of application scenarios.

[0004] The present invention is realized by the following technical solutions:

[0005] The present invention relates to a discontinuous and overlapping entity recognition system, including: a one-dimensional text representation unit, a dependency injection unit, a two-dimensional text representation unit, and an extraction and decoding unit. Among them: the one-dimensional text representation unit processes the input text content through the pre-trained language model Bert to obtain a vectorized one-dimensional text representation; the dependency injection unit parses the dependency relationship of the text content through an open-source dependency relationship judgment tool according to the input text content, and then uses a graph convolutional neural network to process the obtained dependency relationship parsing result, and finally merges it into the vectorized one-dimensional text representation to enhance the one-dimensional text representation; the two-dimensional text representation unit uses a conditional layer normalization method to convert the one-dimensional text representation with enhanced dependency relationship into a two-dimensional text representation representing word-word relationships; the extraction and decoding unit uses a multi-layer perceptron to obtain a two-dimensional word-word relationship directed graph structure according to the two-dimensional text representation, and then uses a backtracking method to decode the discontinuous and overlapping entities appearing in the directed graph. Technical Effects

[0006] The present invention uses the structure of a directed graph to represent discontinuous and overlapping entities, and enhances text representation using dependency relationships. Compared with the prior art, the present invention realizes the recognition of discontinuous and overlapping entities through the directed graph structure; there is no ambiguity in the representation of discontinuous and overlapping entities; the enhancement of text representation by dependency relationships improves the accuracy of recognizing complex discontinuous and overlapping entities. Description of the Drawings

[0007] Figure 1 It is a flowchart of the present invention;

[0008] Figure 2 It is a system structure diagram of the present invention;

[0009] Figure 3 It is the dependency graph in the dependency injection unit;

[0010] Figure 4 It is the dependency link matrix in the dependency injection unit;

[0011] Figure 5 It is a schematic diagram of the effect of the embodiment. Detailed Embodiment

[0012] As Figure 1 shown, this embodiment relates to a method for identifying discontinuous and overlapping named entities based on the above system, including:

[0013] Step 1) Obtain the one-dimensional representation of the text to be recognized through Bert, specifically including: input the original sentence x to be extracted into the pre-trained language model Bert to obtain the one-dimensional vectorized text representation H, that is, in the form of a matrix where: the original sentence x to be extracted = x1, x2, …, x N , each letter represents a character, d h represents the dimension of the text representation, which is consistent with the dimension of the pre-trained language model Bert, and N is the length of the sentence.

[0014] Step 2) Inject dependency relationships, specifically including using Core NLP to annotate the dependency relationships of the original text to obtain a tree-shaped dependency relationship, converting it into the form of an adjacency matrix, and then aggregating the neighbors of the nodes and updating the node representations, specifically including:

[0015] 2.1) Text annotation, using the Standford Core NLP open-source package to annotate the dependency relationships of the original text x to obtain the dependency relationship result in the form of a tree structure as Figure 3 shown.

[0016] 2.2) Use as Figure 4The adjacent matrix W shown represents the dependency relationship result of the tree structure, W ∈ R N×N , where N is the length of the original text.

[0017] 2.3) Calculate based on the one-dimensional text representation H obtained in step 1) and the adjacent matrix W obtained in step 2.2 and then input it into the graph neural network to obtain the text representation injected with dependency relationships, where σ is the sigmoid function, l = {0, 1, 2, 3, 4}, and N i is the number of nodes connected to node i in the adjacent matrix.

[0018] The graph neural network mentioned above is implemented by using, but not limited to, the technology described by Dai, X., etc. in "An Effective Transition-based Model for Discontinuous NER" (In Proceedings of the ACL, 5860–5870).

[0019] In step 3), the text injected with dependency relationships obtained in step 2 is converted into a matrix-style two-dimensional text representation by using conditional layer normalization, that is, the vector representations of the i-th word and the j-th word are multiplied point by point to obtain the two-dimensional text representation V. Specifically: the relationship between the i-th word and the j-th word where: h i , h j is the text representation injected with dependency relationships obtained in step 2; γ ij = W α h i + b α , λ ij = W β h i + b β , and are both learnable parameter matrices, and finally obtain

[0020] In step 4), according to the two-dimensional text representation obtained in step 3, classify through a multi-layer perceptron to obtain Figure 5 the two-dimensional word-word relationship directed graph structure shown, that is, after the CSO label, traverse all paths in the two-dimensional word-word relationship directed graph structure by using the depth-first search method to obtain the target discontinuous and overlapping entities.

[0021] The multi-layer perceptron mentioned above includes two fully connected neural networks, and ReLU is used as the activation function in the middle.

[0022] The CSO tags mentioned above refer to: Combine, which means that two words can be concatenated; Span, which means that the position points from the last word of the entity to the first word, indicating the range of a complete entity interval; O, which means that there is no relationship between two words and they will not appear consecutively in any entity.

[0023] Through specific actual experiments, the CADEC dataset was used. This dataset originated from a forum AskaPatient where patients discussed their treatment experiences. The entity type of this dataset is adverse drug reactions (ADE), containing 7,597 sentences and 6,318 named entities, among which 675 entities are discontinuous, accounting for approximately 10.6% of all entities.

[0024] Compare the present invention with the Bert-CRF which is used as a baseline model and only has the ability to recognize continuous entities, the sequence annotation model proposed by Tang et al. (2018), the model proposed by Li et al. (2021a) which first identifies entity intervals and then judges the relationship between entity intervals; the model proposed by Wang and Lu (2019) which uses hypergraphs to handle discontinuous and overlapping entities; the method proposed by Dai et al. (2020) which identifies discontinuous and overlapping entities through transfer methods; and use the accuracy, recall rate, and F1 score of the extracted entities as evaluation metrics.

[0025] Among the above-mentioned evaluation metrics, the accuracy rate (P) indicates the proportion of entities predicted that are in the gold standard, and the recall rate (R) is the proportion of entities in the gold standard that are successfully predicted. F1 is obtained according to the accuracy rate and recall rate by the formula as follows.

[0026] The experimental results are as follows:

[0027] Through specific experiments, in the discontinuous and overlapping entity dataset CADEC, for the entity extraction training task, compared with the classical Bert-CRF named entity recognition method, the F1 has increased by 11.62%, and compared with the previous best discontinuous and overlapping named entity recognition Span-base method, the F1 value has increased by 3.62%. This shows that the accuracy of the present invention in recognizing discontinuous and overlapping entities has been significantly improved.

[0028] Furthermore, a specific discontinuous and overlapping entity recognition scenario is used to demonstrate the technical effect of the present invention. The original sentence is "Joint and Muscle Pain / Stiffness"

[0029] There are 4 entities in the original sentence: "Joint Pain", "Muscle Pain", "Joint Stiffness", and "Muscle Stiffness". Among the above methods, only the present invention completely extracts all the entities. The Bert-CRF model cannot recognize discontinuous entities, so this method only recognizes two entities. Moreover, the first entity is a combination of two entities, and the second entity is incomplete, both of which do not meet the extraction requirements. Due to the ambiguity in decoding of the Sequence Labeling method, the two extracted entities are both incorrect and semantically incorrect. The Transition-base method successfully extracts two entities but cannot handle the case where "Joint" and "Muscle" modify "Stiffness", so it cannot extract the latter two entities. The Span-base method successfully extracts three entities but fails to connect "Muscle" and "Stiffness" together, missing the fourth entity.

[0030] The following table shows the results of the dependency relationship for this sentence. We find that the dependency relationship clearly indicates that "Joint" and "Muscle" modify "Stiffness", which is crucial for the present invention to extract the correct results. Without the help of the dependency relationship, the extraction results of this method would be "Joint Pain", "Muscle Pain", and "Stiffness", missing two entities.

[0031] In summary, compared with the existing methods, the present invention can unambiguously represent discontinuous and overlapping entities through a directed graph structure, without the problem of representation ambiguity that occurred in the previous methods. At the same time, the dependency relationship is reasonably used and injected into the model structure, effectively improving the accuracy of extracting discontinuous and overlapping entities in complex sentences.

[0032] The above specific implementation can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.

Claims

1. A discontinuous and overlapping entity recognition system, characterized in that, It includes: A one-dimensional text representation unit, a dependency injection unit, a two-dimensional text representation unit, and an extraction and decoding unit, where: The one-dimensional text representation unit processes the input text content through the pre-trained language model Bert to obtain a vectorized one-dimensional text representation; The dependency injection unit parses the dependency relationship of the text content through an open-source dependency relationship judgment tool according to the input text content, then processes the obtained dependency relationship parsing result using a graph convolutional neural network, and finally merges it into the vectorized one-dimensional text representation to achieve the enhancement of the one-dimensional text representation; The two-dimensional text representation unit uses the conditional layer normalization method to convert the dependency-enhanced one-dimensional text representation into a two-dimensional text representation representing the character-character relationship; The extraction and decoding unit obtains a two-dimensional character-character relationship directed graph structure using a multi-layer perceptron according to the two-dimensional text representation, and then decodes the discontinuous and overlapping entities appearing in the directed graph using the backtracking method.

2. A method for identifying discontinuous and overlapping entities based on the system described in claim 1, characterized in that, It includes: Step 1) Obtain the one-dimensional representation of the text to be recognized through Bert, specifically including: input the original sentence x to be extracted into the pre-trained language model Bert to obtain the one-dimensional vectorized text representation H, that is, in the form of a matrix where: the original sentence x to be extracted = x1, x2, …, x N , each letter represents a character, d h represents the dimension of the text representation, which is consistent with the dimension of the pre-trained language model Bert, and N is the length of the sentence; Step 2) Inject the dependency relationship, specifically including using Core NLP to annotate the dependency relationship of the original text to obtain a tree-shaped dependency relationship, converting it into the form of an adjacency matrix, and then aggregating the neighbors of the nodes and updating the node representation; Step 3) Convert the text with injected dependency relationships obtained in Step 2 into a matrix - style two - dimensional text representation using conditional layer normalization, that is, multiply the vector representations of the i - th word and the j - th word point - by - point to obtain the two - dimensional text representation V. Specifically: the relationship between the i - th word and the j - th word where: h i , h j is the text representation with injected dependency relationships obtained in Step 2; γ ij = W α h i + b α , λ ij = W β h i + b β , and are both learnable parameter matrices, and finally obtain Step 4) According to the two-dimensional text representation obtained in Step 3, perform classification through a multi-layer perceptron to obtain a two-dimensional character-character relationship directed graph structure, that is, after the CSO label, traverse all paths in the two-dimensional character-character relationship directed graph structure through depth-first search to obtain the target discontinuous and overlapping entities.

3. The non - continuous and overlapping entity recognition method according to claim 2, characterized in that The specific steps of Step 2 include: 2.1) Text annotation, using the Standford Core NLP open-source package to annotate the dependency relationship of the original text x to obtain a dependency relationship result with a tree structure as shown in Figure 3; 2.2) Represent the dependency relationship result of the tree structure using the adjacency matrix W shown in Figure 4, where W ∈ R N×N , and N is the length of the original text; 2.3) Calculate based on the one-dimensional text representation H obtained in step 1) and the adjacency matrix W obtained in step 2.2 and then input it into the graph neural network to obtain the text representation with injected dependency relationships, where σ is the sigmoid function, l = {0, 1, 2, 3, 4}, and N i is the number of nodes connected to node i in the adjacency matrix.

4. The non - continuous and overlapping entity recognition method according to claim 2, characterized in that, The multi-layer perceptron includes two fully connected neural networks, with ReLU used as the activation function in the middle.

5. The non - continuous and overlapping entity recognition method according to claim 2, characterized in that, The CSO label refers to: Combine, that is, two characters can be connected together before and after; Span, that is, the position points from the end character of the entity to the start character, indicating the range of a complete entity interval; O, that is, it indicates that there is no relationship between two characters and they will not appear continuously in any entity.