Nested Named Entity Recognition Method, System, Electronic Device and Readable Medium

By text tagging and clustering of the corpus, combined with the named entity recognition model enhanced by adaptive data, the problem of insufficient nesting of named entities and training samples is solved, and the accuracy of named entity recognition and model training effect is improved.

CN110956042BActive Publication Date: 2025-06-24INFORMATION SCI RES INST OF CETC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911291456.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-16
Publication Date
2025-06-24
Estimated Expiration
2039-12-16

AI Technical Summary

Technical Problem

Existing named entity recognition technology is difficult to effectively deal with the problem of insufficient nesting of named entities and training samples, resulting in increased recognition complexity and poor model training effect.

Method used

By labeling and clustering the corpus based on preset text marking methods and clustering methods, clustering sets are generated, and a named entity recognition model with adaptive data augmentation is used for identification, the degree of data augmentation is gradually improved to improve the model training effect.

Benefits of technology

This method can effectively reduce the impact of named entity nesting on recognition effect, adapt to nested named entity recognition tasks under insufficient training samples, and improve the accuracy of named entity recognition and model training effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110956042B_ABST
    Figure CN110956042B_ABST
Patent Text Reader

Abstract

The nested named entity recognition method, system, electronic device and readable medium of the present invention include: marking each text in the corpus based on a preset text marking method to obtain a marking set, where the marking set includes the text and the corresponding named entities, and at least one text corresponds to multiple named entities; based on a preset clustering method, clustering the marking set according to each named entity to obtain a cluster set, where the cluster set includes the text and the named entity uniquely corresponding to the text; based on a preset named entity recognition model with adaptive data augmentation, respectively recognizing the named entities in each cluster set. The nested named entity recognition problem is transformed into a non-nested named entity recognition problem, reducing the impact of named entity nesting on the recognition effect; gradually increasing the data augmentation degree according to the training effect, controlling the data augmentation usage intensity at the optimal level, and improving the training effect to adapt to the nested named entity recognition task under the condition of insufficient samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of named entity recognition, and particularly relates to a nested named entity recognition method, a nested named entity recognition system, an electronic device and a computer-readable storage medium. Background Art

[0002] Named entity recognition (NER) is one of the basic research contents of natural language processing, and its task is to identify language chunks in text. Named entity recognition often faces the problems of named entity nesting and insufficient training samples in practical applications.

[0003] Named entity nesting makes it impossible to establish a one-to-one relationship between text and entity labels. For example, "Bethune Medical College" is an organizational name entity, while "Bethune" is a personal name entity. Therefore, during the text marking process, "Bethune" has two labels. The multi-label problem increases the complexity of named entity recognition and makes existing mature named entity recognition methods unable to be directly used.

[0004] Insufficient training samples are a common problem faced by entity recognition tasks. Constructing a training sample dataset for named entity recognition in a professional field is a time-consuming process that requires people with professional knowledge to perform data annotation. Therefore, it is difficult to form a large dataset. Data augmentation is an important method to solve the problem of insufficient training samples. By using an automated method to construct new samples based on the original dataset, the training effect of the model can be enhanced. Therefore, studying nested named entity recognition under the condition of insufficient training samples is of great significance for the practical application of named entity recognition. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art, and provides a nested named entity recognition method, a nested named entity recognition system, an electronic device and a computer-readable storage medium.

[0006] The first aspect of the present invention provides a nested named entity recognition method, including the following steps:

[0007] Mark each text in the corpus based on a preset text marking method to obtain a marking set, the marking set includes the text and the corresponding named entities, and at least one of the texts corresponds to multiple named entities;

[0008] Based on a preset clustering method, cluster the marking set according to each named entity to obtain a cluster set, the cluster set includes the text and the named entity uniquely corresponding to the text;

[0009] Based on a preset named entity recognition model with adaptive data augmentation, recognize the named entities in each of the cluster sets respectively.

[0010] Optionally, the step of clustering the tag set according to each named entity based on a preset clustering method includes:

[0011] Preset a clustering result metric function;

[0012] Based on the clustering result metric function, use a hierarchical clustering method to cluster the tag set according to each named entity to obtain a cluster set.

[0013] Optionally, the step of presetting a clustering result metric function includes:

[0014] Assume the corpus is [w1, w2, …, w n , where w i represents the i-th text in the corpus, use T i to represent the tag set of w i , use t a to represent the named entity a, and establish the following relational expression (1) for the indicator function of the named entity a for the i-th character:

[0015]

[0016] The relevance between the named entity t a and the named entity t b is defined as the following relational expression (2):

[0017] E a,b = ∑ 所有语料库 ∑ i f(t a , i)f(t b , i) (2);

[0018] E represents the distance matrix between named entities;

[0019] Let C represent the cluster set, C i represent the i-th cluster, and the internal distance of C i is the distance between the internal named entities of C i , and the calculation method is as follows in relational expression (3):

[0020]

[0021] max(E a,b ) represents the maximum value of the elements in E, and the distance between C i and C j is the distance between the named entities of the two clusters, and the calculation method is as follows in relational expression (4):

[0022]

[0023] Based on relational expression (3), relational expression (4), and according to the objective requirements of clustering, the clustering result metric function is obtained as the following relational expression (5):

[0024] g total = α(∑ i,j g out (C i , C j ) - ∑ i g in (C i )) - (1 - α)|C| / c (5);

[0025] |C| represents the number of clusters, c represents the constant number of types of named entities, and α is the weight parameter.

[0026] Optionally, based on the clustering result metric function, adopting a hierarchical clustering method to cluster the token set according to each named entity to obtain a cluster set, includes:[[]]

[0027] S110. Divide each named entity in the token set into a cluster;

[0028] S120. Randomly select two clusters;

[0029] S130. Merge the two randomly selected clusters, and determine whether g total decreases. If so, execute step S120. If not, execute step S140;

[0030] S140. Determine whether the increment of g total in several consecutive rounds of iteration is less than 0 or |C| = 1. If so, stop the iteration and return the clustering result to obtain the cluster set; if not, execute step S120.

[0031] Optionally, the steps of training the named entity recognition model with adaptive data augmentation specifically include:[[]]

[0032] S210. Scan the initial training sample corpus and initialize D a and S a , where D a represents the set of words contained in named entity a in the initial training sample corpus, and S a represents the set of statement numbers containing named entity a in the initial training sample corpus;

[0033] S220. Control the amount M of data augmentation according to the current round of iteration a(t), perform data augmentation on the initial training sample corpus;

[0034] S230. Use the BiLSTM-CRF recognition model to train on the training sample corpus enhanced in the current round of iteration, and obtain the training result R of each named entity a on the validation set a (t);

[0035] S240. Determine whether there is an R for all named entities a in the current round of iteration a (t) < R a (t - 1). If so, stop the iteration and end the training. If not, calculate the data augmentation control amount for the next round of iteration and execute step S220.

[0036] Optionally, the data augmentation according to the data augmentation degree control amount M a (t) for the initial training sample corpus includes:

[0037] For each named entity a in turn, randomly select M a samples in S a (t). For each of these samples, randomly select words in D a to replace them, and add the newly formed samples to the initial training sample corpus.

[0038] Optionally, the data augmentation degree control amount M a (t) adopts the following calculation formula:

[0039]

[0040] t represents the iteration order number, and R a (t) represents the F1 value of entity type a on the validation set after the end of the t-th round of training.

[0041] The second aspect of the present invention provides a nested named entity recognition system, including:

[0042] A marking module for marking each text in the corpus based on a preset text marking method to obtain a marking set, the marking set including the text and the corresponding named entities, and at least one of the texts corresponding to multiple named entities;

[0043] A clustering module for clustering the marking set according to each named entity based on a preset clustering method to obtain a cluster set, the cluster set including the text and the named entity uniquely corresponding to the text;

[0044] A data augmentation and recognition module, configured to recognize named entities in each of the cluster sets respectively based on a preset named entity recognition model with adaptive data augmentation.

[0045] The third aspect of the present invention provides an electronic device, including:

[0046] One or more processors;

[0047] A storage unit, configured to store one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the nested named entity method provided in the first aspect of the present invention.

[0048] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored,

[0049] When the computer program is executed by a processor, it can implement the nested named entity method provided in the first aspect of the present invention.

[0050] The nested named entity recognition method, system, electronic device and readable medium according to the embodiments of the present invention include: marking each text in a corpus based on a preset text marking method to obtain a marking set, where the marking set includes the text and the corresponding named entities, and at least one text corresponds to multiple named entities; clustering the marking set according to each named entity based on a preset clustering method to obtain a cluster set, where the cluster set includes the text and the named entity uniquely corresponding to the text; recognizing the named entities in each cluster set respectively based on a preset named entity recognition model with adaptive data augmentation. The nested named entity recognition method of the present invention clusters named entities, divides entities with nested relationships into different clusters, and completes named entity recognition in different clusters, converting nested named entities into non-nested named entities. Compared with existing multi-level recognition models, it can avoid error conduction and reduce the impact of named entity nesting on the recognition effect. In addition, through multiple rounds of iteration, the degree of data augmentation is gradually improved according to the training effect of the recognition model, so as to control the usage intensity of data augmentation at the optimal level, and further improve the training effect of the named entity recognition model, and can adapt to the nested named entity recognition task under the condition of insufficient training samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic flowchart of a nested named entity recognition method according to the first embodiment of the present invention;

[0052] Figure 2 is Figure 1 the overall flowchart in the nested named entity recognition method in

[0053] Figure 3 isFigure 1 Schematic diagram of the clustering process of the nested named entity recognition method in

[0054] Figure 4 is Figure 1 Schematic diagram of the data augmentation and recognition process of the nested named entity recognition method in

[0055] Figure 5 Block diagram showing the composition of a nested named entity recognition system according to the second embodiment of the present invention. Detailed implementation manners

[0056] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0057] As Figure 1 and Figure 2 shown, a nested named entity recognition method includes the following steps:

[0058] Mark each text in the corpus based on a preset text marking method to obtain a marking set, the marking set includes the text and the corresponding named entities, and at least one text corresponds to multiple named entities;

[0059] Based on a preset clustering method, cluster the marking set according to each named entity to obtain a cluster set, the cluster set includes the text and the named entity uniquely corresponding to the text;

[0060] Based on a preset named entity recognition model with adaptive data augmentation, recognize the named entities in each cluster set respectively.

[0061] Through the above steps, the nested named entity recognition method of the present invention clusters named entities, divides entities with nested relationships into different clusters, and completes named entity recognition in different clusters, transforming the nested named entity recognition problem into a non-nested named entity recognition problem. Compared with the existing multi-level recognition models, it can avoid error propagation and reduce the impact of named entity nesting on the recognition effect. In addition, through multiple rounds of iteration, the degree of data augmentation is gradually increased according to the training effect of the recognition model, so as to control the usage intensity of data augmentation at the optimal level, thereby improving the training effect of the named entity recognition model and being able to adapt to the nested named entity recognition task under the condition of insufficient training samples.

[0062] Specifically, in the text tagging step, the corpus is processed into a form convenient for subsequent processing. In the embodiments of the present invention, the "BMEOS" method is adopted to tag the text in units of characters. For example, in the sentence "Playing table tennis is beneficial to the body", "Playing table tennis" is of the Sport entity type. At this time, the tagging result of the sentence is: "Hit / B_Sport Ping / M_Sport Pang / M_Sport Qiu / E_sport is / O good / O for / O the / O body / O". For nested named entities, they are tagged in a set manner. For example, in the sentence "Bethune Medical College", "Bethune Medical College" is an "Organization entity" and "Bethune" is a "Name entity", then the tagging result of the sentence is:

[0063] "Bai / B_Organization#B_Name Qiu / M_Organization#M_Name En / M_Organization#E_Name Medical / M_Organization College / E_Organization". At this time, there is a one-to-many relationship between the text and the tags. For example, in the above example, the tagging set of the character "En" is {"M_Organization", "E_Name"}.

[0064] The goal of named entity clustering is to divide entities into several clusters, where the distance between entities within each cluster is as small as possible, the distance between entities belonging to different clusters is as large as possible, and the number of clusters is as small as possible. The input of named entity clustering in the embodiments of the present invention is the result of text tagging, and the clustering result metric function is obtained by the following steps:

[0065] Suppose one of the sentences in the corpus is [w1, w2,..., w n , where w i represents the i-th character of the sentence, and T i represents the tagging set of w i . Let t a represent the named entity a, and establish the following relationship formula (1) for the indicator function of the named entity a for the i-th character:

[0066]

[0067] The named entity t a and the named entity t b are defined as the following relationship formula (2):

[0068] E a,b =∑ 所有语句 ∑ i f(t a , i) f(t b , i) (2);

[0069] E represents the distance matrix between named entities;

[0070] Let C denote the set of clusters, C i represents the i-th cluster, C i The internal distance is C i The distance between internal named entities is calculated as follows:

[0071]

[0072] max(E a,b ) represents the maximum value of the elements in E, C i With C j The distance between them is the named entity distance between two clusters, which is calculated as follows:

[0073]

[0074] The goal of clustering is to make all g in (C i ) is as small as possible, all g out (C i ,C j ) is as large as possible. Based on this, based on equations (3), (4) and the target requirements of clustering, the clustering result metric function is obtained, as shown in equation (5):

[0075] g total =α(∑ i,j g out (C i ,C j )-∑ i g in (C i ))-(1-α)|C| / c (5);

[0076] Where |C| represents the number of clusters, c represents the number constant of named entity types, and α is the weight parameter. The clustering result hopes that the number of clusters is as small as possible, so in g total A regularization term -(1-α)|C| / c is added to the end of the calculation.

[0077] Based on the above clustering result measurement method, a hierarchical clustering method is used to cluster entities. The clustering in the embodiment of the present invention is a clustering method based on random merging and splitting, whose input is the distance matrix E and whose output is the clustering result. In the clustering process, each entity is first placed in a separate cluster, and then two entities are selected that can make g total The two clusters that do not decrease are merged, and then it is determined whether to proceed to the next step of clustering. If the total number of clusters is 1 or gtotal Stop clustering if there is no increase.

[0078] As Figure 3 shown, specifically, based on the clustering result metric function, a hierarchical clustering method is used to cluster the tag set according to each named entity to obtain a cluster set, including:

[0079] Step S110: Divide each named entity in the tag set into a cluster;

[0080] Step S120: Randomly select two clusters;

[0081] Step S130: Merge the two randomly selected clusters, and judge whether g total decreases. If so, execute Step S120. If not, execute Step S140;

[0082] Step S140: Judge whether the g total increment for several consecutive rounds of iteration is less than 0 or |C| = 1. If so, stop the iteration and return the clustering result to obtain the cluster set. If not, execute Step S120.

[0083] It should be noted that the embodiment of the present invention adopts a hierarchical clustering method, and other clustering methods can also be adopted, which is specifically determined according to application requirements. Compared with before clustering, the number of nested entities belonging to the same cluster after clustering is greatly reduced. For entities within the same cluster, a one-to-one relationship between text and entity tags can be established, and a single model can be used in the cluster. After the named entities are clustered, entities in different clusters can respectively adopt the following named entity recognition method based on adaptive data augmentation to use separate models. Since the models used by entities in different clusters are not related to each other, training and recognition can be completed in parallel.

[0084] In the named entity recognition process based on adaptive data augmentation in the embodiment of the present invention, the input is a training sample data set and a validation sample data set, including the text tagging results in the above steps, and the output is a trained BiLSTM-CRF named entity recognition model.

[0085] Let R a (t) represent the F1 value of type-a entities on the validation set after the end of the t-th round of training. Let M a (t) represent the data augmentation degree control amount of type-a entities, that is, the number of samples of type-a entities increased by data augmentation before the t-th training.

[0086] In the named entity recognition process based on adaptive data augmentation, first scan the training sample database to obtain D a and S a of all entities a, where D aDenote the set of words included in entity type a in the training samples as S a Denote the set of sentence numbers containing entity type a in the training set. Then, control the amount M a (t) to complete data augmentation. After that, use the BiLSTM-CRF recognition model to train on the training set and obtain R a (t) on the validation sample dataset. Finally, determine whether to stop the model training. When there is no increase in the training results of all entities, the entire training process ends; otherwise, calculate the data augmentation amount for the (t + 1)-th round of training. The intensity M a (t) of data augmentation is inversely proportional to the training result R a (t), and is proportional to the increase in the training result R a (t) - R a (t - 1) and M a (t - 1); when there is no increase in the training result, i.e., R a (t) - R a (t - 1) < 0, then stop the data augmentation for entity type a.

[0087] As Figure 4 shown, specifically, the steps for training the named entity recognition model with adaptive data augmentation specifically include:

[0088] Step S210: Scan the initial training sample corpus and initialize D a and S a , where D a denotes the set of words included in named entity a in the initial training sample corpus, and S a denotes the set of sentence numbers containing named entity a in the initial training sample corpus;

[0089] Step S220: Perform data augmentation on the initial training sample corpus according to the data augmentation intensity control amount M a (t) for the current round of iteration;

[0090] Step S230: Use the BiLSTM-CRF recognition model to train on the training sample corpus enhanced in the current round of iteration and obtain the training results R a (t) of each named entity a on the validation set;

[0091] Step S240: Determine whether there exists R a (t) < R a (t - 1) for all named entities a in the current round of iteration. If so, stop the iteration and the training ends. If not, calculate the data augmentation control amount for the next round of iteration and execute Step S220.

[0092] Specifically, according to the data augmentation degree control quantity M a (t), perform data augmentation on the initial training sample corpus, including:

[0093] For each named entity a in turn, randomly select M a samples in S a (t). For each sample among them, randomly select words in D a for replacement, and add the newly formed samples to the initial training sample corpus.

[0094] Specifically, the data augmentation degree control quantity M a (t) in the embodiments of the present invention adopts the following calculation formula:

[0095]

[0096] where t represents the iteration order number, and R a (t) represents the F1 value of entity type a on the validation set after the end of the t-th round of training.

[0097] It should be noted that the BiLSTM-CRF model is adopted in the embodiments of the present invention, and other named entity recognition models can also be adopted, which can be specifically selected according to application requirements.

[0098] The named entity recognition method based on adaptive data augmentation in the embodiments of the present invention adopts a multi-round iteration method, gradually improves the degree of data augmentation according to the training effect of the recognition model, thereby controlling the usage intensity of data augmentation at the optimal level, and further improving the training effect of the recognition model.

[0099] As Figure 5 shown, a nested named entity recognition system 100 is provided in the second aspect of the present invention. This system is based on the nested named entity recognition method provided by the present invention. For specific reference, please refer to the previous records and will not be elaborated here. The nested named entity recognition system 100 includes:

[0100] A tagging module 110, configured to tag each text in the corpus based on a preset text tagging method to obtain a tag set. The tag set includes texts and corresponding named entities, and at least one text corresponds to multiple named entities;

[0101] A clustering module 120, configured to cluster the tag set according to each named entity based on a preset clustering method to obtain a cluster set. The cluster set includes texts and the named entities uniquely corresponding to the texts;

[0102] A data augmentation and recognition module 130, configured to recognize the named entities in each cluster set based on a preset named entity recognition model with adaptive data augmentation.

[0103] A third aspect of the present invention provides an electronic device, comprising:

[0104] One or more processors;

[0105] A storage unit for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the nested named entity method provided by the present invention.

[0106] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored,

[0107] and the computer program, when executed by a processor, can implement the nested named entity method provided by the present invention.

[0108] Wherein, the computer-readable medium may be included in the device, equipment, system of the present invention, or may exist alone.

[0109] Wherein, the computer-readable storage medium may be any tangible medium that contains or stores a program, and it may be an electrical, magnetic, optical, electromagnetic, infrared, semiconductor system, device, equipment. More specific examples include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, an optical fiber, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0110] Wherein, the computer-readable storage medium may also include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code, and specific examples thereof include, but are not limited to, electromagnetic signals, optical signals, or any suitable combination thereof.

[0111] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention, and the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered within the protection scope of the present invention.

Claims

1. A nested named entity recognition method, characterized in that, Including the following steps: Based on a preset text marking method, each text in the corpus is marked to obtain a marking set, the marking set includes the text and the corresponding named entities, and at least one of the texts corresponds to multiple named entities; Based on a preset clustering method, the marking set is clustered according to each named entity to obtain a cluster set, the cluster set includes the text and the named entity uniquely corresponding to the text; Based on a preset named entity recognition model with adaptive data augmentation, the named entities in each cluster set are respectively recognized; The step of clustering the marking set according to each named entity based on a preset clustering method to obtain a cluster set includes: presetting a clustering result metric function; based on the clustering result metric function, using a hierarchical clustering method to cluster the marking set according to each named entity to obtain a cluster set; It further includes the step of training the named entity recognition model with adaptive data augmentation, specifically including: S210. Scan the initial training sample corpus and initialize D a and S a , where D a represents the set of words included in named entity a in the initial training sample corpus, and S a represents the set of statement numbers in the initial training sample corpus that contain named entity a; S220. Control the quantity M a (t) according to the data augmentation degree of the current round of iteration, and perform data augmentation on the initial training sample corpus, specifically: for each named entity a in turn, randomly select M a (t) samples in S a . For each of these samples, randomly select words in D a for replacement, and add the newly formed samples to the initial training sample corpus; S230. Use the BiLSTM-CRF recognition model to train on the training sample corpus after being iteratively enhanced in the current round, and obtain the training result R(t) of each named entity a on the validation set. a (t); S240. Determine whether there exists an R for all named entities a in the current round of iteration a (t) < R a (t - 1). If so, stop the iteration and end the training. If not, calculate the data augmentation control amount for the next round of iteration and execute step S220; Data augmentation degree control quantity M a (t) adopts the following calculation formula: t represents the order number of the iteration round, R a (t) represents the F1 value of class a entities on the validation set after the end of the t-th round of training.

2. The nested named entity recognition method according to claim 1, wherein The presetting of the clustering result metric function includes: Suppose the corpus is [w1, w2,..., w n , where w i represents the i-th text of the corpus, and T i represents the set of tags of w i , and t a represents the named entity a. The following relational expression (1) for the indicator function of the named entity a with respect to the i-th character is established: Named entity t a The relevance to the named entity t b is defined by the following relational expression (2): E a,b = ∑ 所有语料库 ∑ i f(t a , i) f(t b , i) (2); E represents the distance matrix between named entities; Let \(C\) denote the set of clusters, \(C\) i denote the \(i\)-th cluster, \(C\) i The internal distance of is \(C\) i The distance between internal named entities, calculated as follows in relation (3): max(E a,b ) represents the maximum value of the elements in E, C i The distance between j and C is the named entity distance between two clusters, and the calculation method is as follows in relation (4): Based on relationship (3), relationship (4) and according to the target requirements of clustering, the clustering result metric function is obtained, as shown in the following relationship (5): g total = α(∑ i,j g out (C i , C j ) - ∑ i g in (C i )) - (1 - α)|C| / c(5); |C| represents the number of clusters, c represents the number constant of the types of named entities, and α is a weight parameter.

3. The nested named entity recognition method according to claim 2, wherein The step of using a hierarchical clustering method to cluster the marking set according to each named entity based on the clustering result metric function to obtain a cluster set includes: S110: Divide each named entity in the marking set into a cluster; S120: Randomly select two clusters; S130. Merge the two randomly selected clusters and determine whether g total decreases. If so, execute step S120; if not, execute step S140. S140. Determine whether the g total increment in several consecutive rounds of iteration is less than 0 or |C| = 1. If so, stop the iteration and return the clustering result to obtain the cluster set; if not, execute step S120.

4. A nested named entity recognition system, characterized in that, The nested named entity method according to any one of claims 1 to 3 includes: A marking module, configured to mark each text in the corpus based on a preset text marking method to obtain a marking set, the marking set includes the text and the corresponding named entities, and at least one of the texts corresponds to multiple named entities; A clustering module, configured to cluster the marking set according to each named entity based on a preset clustering method to obtain a cluster set, the cluster set includes the text and the named entity uniquely corresponding to the text; A data augmentation and recognition module, configured to respectively recognize the named entities in each cluster set based on a preset named entity recognition model with adaptive data augmentation.

5. An electronic device, characterized in that, Including: One or more processors; A storage unit, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the nested named entity method according to any one of claims 1 to 3.

6. A computer-readable storage medium, on which a computer program is stored, characterized in that When the computer program is executed by a processor, it can implement the nested named entity method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Named entity relation extraction and construction method based on deep learning

    CN104199972A

  • Techniques for similarity analysis and data enrichment using knowledge sources

    CN106687952A