Unique expression extracting method, unique expression extracting apparatus, and program
The method enhances named entity extraction accuracy by using a reproduction model and external knowledge base to generate and combine candidate entities, addressing the challenges of short texts with errors in current models.
Patent Information
- Application Number
- JP2024231493
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing neural network-based named entity recognition models struggle with accurate fine-grained discrimination in short texts with spelling mistakes or typos, leading to incorrect entity recognition.
A method involving a named entity reproduction model to generate candidates, search an external knowledge base for additional information, and combine this with the original text to enhance the accuracy of named entity extraction.
Improves the accuracy of fine-grained named entity extraction by introducing additional context and information from the knowledge base, overcoming the limitations of current models in handling short texts with errors.
Smart Images

Figure 2025107160000001_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of machine learning and natural language processing (NLP, Natural Language Processing), and specifically relates to a method, apparatus, and storage medium for extracting named entities.
Background Art
[0002] In recent years, neural network-based named entity recognition (NER) models have been used for automatic extraction of semantic features of large-scale texts, and good effects have already been obtained. However, in actual applications, when fine-grained discrimination of expressions is required, situations may occur where the text to be discriminated is short (without sufficient context) and there are spelling mistakes or typos in the text to be discriminated. In such situations, it is difficult for popular pre-trained transformers to perform accurate discrimination either.
[0003] For example, in the case of the sentence "The term'meter maid' spread to the cute Rita of the Beatles, and this male singer was charmed by a female traffic warden", in the sentence, "the Beatles" is often recognized as expression types such as "book", "other products", "band", etc., and "cute Rita" is often recognized as expression types such as "book", "other products", "musical work", etc. In fine-grained named entity extraction, it is more desirable to recognize "the Beatles" as "band" and "cute Rita" as "musical work" in the above sentence. Therefore, for the above sentence, it is difficult for the current named entity extraction model to accurately recognize the named entities in it in a fine-grained manner. Therefore, there is an urgent need for a named entity extraction method that can improve the accuracy rate in complex (fine-grained) named entity extraction scenarios.
Summary of the Invention
Problems to be Solved by the Invention
[0004] At least one embodiment of the present invention aims to provide a named entity extraction method, apparatus, and storage medium that can improve the accuracy in a complex (narrow) named entity extraction scenario.
Means for Solving the Problems
[0005] To solve the above problems, a first aspect of the present invention is a named entity extraction method executed by a named entity extraction apparatus, including: using a named entity reproduction model that reproduces any type of named entity in a text to reproduce named entity candidates in a first text to be extracted; searching a knowledge base based on the named entity candidates in the first text and obtaining related information of the named entity candidates in the first text, where the knowledge base includes related information of a plurality of named entities, and the related information of the named entity includes a representation name, a representation type, and representation description information of the named entity; using the first text as the original text, using the related information of the named entity candidates in the first text as additional text, and combining the first text and the related information of the named entity candidates in the first text to obtain a second text; inputting the second text into a named entity extraction model that extracts named entities from the original text in the input text, and extracting named entities from the original text in the second text by the named entity extraction model.
[0006] Preferably, the method further includes: obtaining first training data in which only named entities are labeled and the corresponding representation types are not labeled, using the named entity reproduction model that reproduces any type of named entity in the text to reproduce the named entity candidates in the first text to be extracted; and training the named entity reproduction model using the first training data with the reproduction rate of the named entity as an optimization index for model training.
[0007] Also, preferably, the step of obtaining the first training data in which the specific expression is labeled and the expression type is not labeled is a step of obtaining the second training data having specific expression label information, where the specific expression label information includes a specific expression and an expression type, and based on the specific expression label information, identifying expression segments and non-expression segments in the second training data, and deleting the specific expression label information in the second training data, labeling the expression segments in the second training data as specific expressions, and labeling the non-expression segments as non-specific expressions to obtain the first training data.
[0008] Also, preferably, before searching the knowledge base based on the specific expression candidates in the first text, obtain knowledge data, clean and organize the knowledge data to obtain formatted data including the expression name, expression type, and expression description information of the specific expression, and use a search engine framework to construct a knowledge base that supports exact match search and fuzzy search based on the expression word based on the formatted data.
[0009] Also, preferably, the step of searching the knowledge base based on the specific expression candidates in the first text and obtaining the related information of the specific expression candidates in the first text includes any one of: using the specific expression candidates in the first text as data to perform exact match matching with the expression name in the knowledge base to obtain the first related information of the specific expression candidates in the first text; using the specific expression candidates in the first text as data to perform fuzzy matching with the pronunciation of the expression name in the knowledge base to obtain the second related information of the specific expression candidates in the first text; and using the specific expression candidates in the first text as data to perform fuzzy matching with the expression name having the same length and a similarity greater than a predetermined threshold in the knowledge base to obtain the third related information of the specific expression candidates in the first text.
[0010] Also preferably, before reproducing the candidate named entities in the first text to be extracted using the named entity reproduction model, the method obtains third training data including named entity label information, and uses the named entity reproduction model to reproduce the candidate named entities in the third training data; searches the knowledge base based on the candidate named entities in the third training data to obtain related information of the candidate named entities in the third training data; uses the third training data as the original text and the related information of the candidate named entities in the third training data as additional text, combines the third training data and the related information of the candidate named entities in the third training data to obtain fourth training data; and further includes the step of training the named entity extraction model using the fourth training data.
[0011] A second aspect of the present invention provides a named entity extraction apparatus, including: a first reproduction module that reproduces candidate named entities in a first text to be extracted using a named entity reproduction model that reproduces any type of named entity in text; a first search module that searches a knowledge base based on the candidate named entities in the first text to obtain related information of the candidate named entities in the first text, where the knowledge base includes related information of a plurality of named entities, and the related information of the named entities includes a name of the named entity, a type of the named entity, and description information of the named entity; a first combination module that uses the first text as the original text, uses the related information of the candidate named entities in the first text as additional text, and combines the first text and the related information of the candidate named entities in the first text to obtain a second text; and an extraction module that inputs the second text to a named entity extraction model that extracts named entities from the original text in the input text, and extracts named entities from the original text in the second text by the named entity extraction model.
[0012] Preferably, the apparatus further includes a first acquisition module that acquires first training data in which only unique expressions are labeled and corresponding expression types are not labeled, and a first training module that trains a unique expression reproduction model using the first training data with the reproduction rate of the unique expression as an optimization index for model training.
[0013] Preferably, the first acquisition module further acquires second training data having unique expression label information, the unique expression label information includes a unique expression and an expression type, and based on the unique expression label information, expression segments and non-expression segments in the second training data are identified. The unique expression label information in the second training data is deleted, the expression segments in the second training data are labeled as unique expressions, and the non-expression segments are labeled as non-unique expressions to obtain first training data.
[0014] Preferably, before the apparatus searches the knowledge base based on the unique expression candidates in the first text, the apparatus acquires knowledge data, cleans and organizes the knowledge data to obtain formatted data including the expression name, expression type, and expression description information of the unique expression, and uses a search engine framework to construct a knowledge base that supports exact match search and fuzzy search based on the expression word based on the formatted data.
[0015] Preferably, the first search module uses, as data, candidate named entities in the first text to perform an exact match with expression names in the knowledge base to obtain first related information on the candidate named entities in the first text, uses, as data, candidate named entities in the first text to perform a fuzzy match with the pronunciations of expression names in the knowledge base to obtain second related information on the candidate named entities in the first text, and uses, as data, candidate named entities in the first text to perform a fuzzy match with expression names in the knowledge base that have the same length and a similarity greater than a predetermined threshold to obtain third related information on the candidate named entities in the first text. The knowledge base is searched based on the candidate named entities in the first text by at least any one of the above methods to obtain related information on the candidate named entities in the first text.
[0016] Preferably, before reproducing, using the named entity reproduction model, candidate named entities in the first text to be extracted, the apparatus further includes: a second acquisition module that acquires third training data including named entity label information; a second reproduction module that reproduces candidate named entities in the third training data using the named entity reproduction model; a second search module that searches the knowledge base based on candidate named entities in the third training data to obtain related information on the candidate named entities in the third training data; a second combination module that combines the third training data with the related information on the candidate named entities in the third training data, with the third training data as the original text and the related information on the candidate named entities in the third training data as additional text, to obtain fourth training data; and a second training module that trains the named entity extraction model using the fourth training data.
[0017] A third aspect of the present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the above named entity extraction method are realized.
Advantages of the Invention
[0018] Compared with the prior art, the named entity extraction method and apparatus provided by the embodiments of the present invention reproduce named entity candidates in the original text using a named entity reproduction model, search an external knowledge base using the named entity candidates, obtain additional information related to the named entity candidates, and combine it with the original text to obtain a combined text including the additional information. In this way, when performing named entity extraction using the combined text, since more additional information is introduced, the accuracy of fine-grained named entity extraction can be improved.
Brief Description of the Drawings
[0019] Hereinafter, by describing the embodiments in detail, the advantages and effects of the present invention will become clear to those skilled in the art. The drawings are only used for the purpose of showing the preferred embodiments and do not limit the present invention. Throughout the drawings, the same members are denoted by the same reference numerals.
Figure 1
Figure 2
Figure 3
Figure 4
Modes for Carrying Out the Invention
[0020] Hereinafter, in order to more clearly illustrate the problems to be solved by the present invention, the configuration and effects of the present invention, specific embodiments will be described in detail with reference to the accompanying drawings. In the following description, the specific details of the configuration and elements are only for helping a comprehensive understanding of the embodiments of the present invention. Therefore, it is obvious to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. For the sake of clarity and brevity, the description of known functions and configurations will be omitted.
[0021] The phrase "one embodiment" or "an embodiment" referred to throughout the specification means that a specific feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the phrases "in one embodiment" or "in an embodiment" described in various places in the specification do not necessarily refer to the same embodiment. Additionally, these specific features, configurations, or characteristics can be appropriately combined in any manner in one or more embodiments. The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged in appropriate situations. Therefore, the embodiments of the present invention described herein may be implemented in an order other than, for example, those illustrated or described herein. Furthermore, the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to the explicitly listed steps or units, and may include other steps or units not explicitly listed or inherent to such a process, method, product, or apparatus. "And / or" in the specification and claims means at least one of the connected objects.
[0022] In each embodiment of the present invention, the magnitude of the numbers of the following processes does not mean the order of execution before or after. The execution order of each process is determined by its function and internal logic, and does not limit the implementation process of the embodiments of the present invention in any way.
[0023] The following description is merely illustrative and does not limit the scope of the claims, the scope of application, or the configuration. Changes can be made without departing from the spirit and scope of the present invention with respect to the functions and arrangements of the elements considered. Appropriate omissions, substitutions, or additions can be made to the procedures and elements of various examples. For example, the methods described can be executed in an order different from the order described, and various steps can be added, omitted, or combined for execution. In addition, the features described in some examples may be combined in other examples.
[0024] External knowledge data is useful for extracting specific expressions in order to provide important boundary information, expression classification information, and context information. Embodiments of the present invention provide a method for extracting specific expressions, and by introducing external knowledge data into the extraction of specific expressions, the accuracy of extracting fine-grained specific expressions is improved. As shown in FIG. 1, the method for extracting specific expressions of the present invention includes the following.
[0025] In step 11, a specific expression reproduction model that reproduces any type of specific expression existing in the text is used to reproduce the specific expression candidates in the first text to be identified.
[0026] The specific expression reproduction model means, for example, a model that evaluates the extraction result by a specific expression extraction model with recall. Here, the first text for which specific expression extraction is performed is input into the specific expression reproduction model, and the specific expression candidates in the first text reproduced by the specific expression reproduction model are obtained. The specific expression reproduction model reproduces any type of specific expression in the input text. That is, the specific expression reproduction model reproduces all types of specific expressions regardless of the type of expression.
[0027] In the embodiment of the present application, first, the named entity reproduction model is trained before step 11. Then, in step 11, the named entity reproduction model is used to reproduce the named entities in the first text. Specifically, the first training data in which only the named entities are marked and the expression type is not marked is used to train the named entity reproduction model with the reproduction rate of the named entities as the optimization index for model training.
[0028] Hereinafter, the training of the named entity reproduction model will be described in detail.
[0029] (1) Obtain the first training data in which only the named entities are marked and the corresponding expression type is not marked.
[0030] Usually, the annotation information of the training data used for named entity extraction includes information such as named entities and expression types. Second training data is obtained, which is training data with named entity annotation information including named entities and expression types. Then, based on the named entity annotation information, the expression segments and non-expression segments in the second training data are identified. Here, a segment may be a character, a subword, or a word.
[0031] Then, the named entity annotation information in the second training data is deleted, the expression segments in the second training data are marked as named entities, and the non-expression segments are marked as non-named entities to form the first training data. For example, in the case of an expression tag, the expression tag is set to 1 for the expression segment and 0 for the non-expression segment. Here, when the value of the expression tag is 1, it indicates that the segment is a named entity, and when the value is 0, it indicates that the segment is not a named entity.
[0032] (2) Use the first training data to train the named entity reproduction model with the reproduction rate of the named entities as the optimization index for model training.
[0033] Train the named entity reproduction model using the first training data. In model training, use the reproduction rate of named entities as the evaluation metric for the model (e.g., the named entity extraction model). That is, use the reproduction rate as the optimization metric for the model to reproduce as many correct named entities as possible. Preferably, the named entity reproduction model according to the embodiments of the present invention may use a named entity extraction model based on spans, but is not limited thereto. Other named entity extraction models may also be used.
[0034] In step 12, search a knowledge base containing relevant information of a plurality of named entities including the expression name, expression type, and expression description information based on the named entity candidates in the first text, and obtain the relevant information of the named entity candidates in the first text.
[0035] Here, an embodiment of the present invention constructs a knowledge base in advance. The knowledge base includes related information of a plurality of specific expressions. The related information of the specific expression includes the expression name of the specific expression, the expression type to which the specific expression belongs, and the expression description information of the specific expression, etc. For example, based on the knowledge data, a knowledge base for retrieval is constructed. The knowledge data may be external knowledge data, for example, data in databases such as Wikipedia, Baidu Encyclopedia, geographical name dictionaries, or data in other similar databases. Specifically, the knowledge data is obtained, the knowledge data is cleaned and sorted, and formatted data including the expression name, expression type, and expression description information of the specific expression, etc. is obtained. The expression description information is information for interpreting and explaining the specific expression. Then, using a search engine framework, a knowledge base is constructed based on the formatted data. The search engine framework may be a general search engine framework such as Elasticsearch, Lucence, Solr, or a similar search engine framework. The knowledge base supports exact match search and fuzzy search based on the expression word. For example, in the case of a Chinese specific expression, the expression name in the knowledge base sets two segments: "original specific expression" and "pronunciation form of the expression name". Among them, the "original specific expression" segment is the Chinese name of the expression name, and the "pronunciation form of the expression name" segment is the pronunciation of the expression name (for example, Chinese pinyin). In the case of an English specific expression, the "original specific expression" segment is the English name of the expression name, and the "pronunciation form of the expression name" segment is the same as the "original specific expression" segment. Also, an embodiment of the present invention may add and set an index for each segment in order to improve the search speed.
[0036] In this way, in step 12, the knowledge base is retrieved based on the specific expression candidates in the first text by at least one of the following methods, and the related information of the specific expression candidates in the first text is obtained.
[0037] (1) Using the named entity candidate in the first text as data, accurately match the expression name in the knowledge base to obtain first related information of the named entity candidate in the first text. Typically, the first related information includes the expression type and expression description information of the named entity candidate.
[0038] (2) Using the named entity candidate in the first text as data, fuzzy matching is performed with the pronunciation of the expression name in the knowledge base to obtain second related information of the named entity candidate in the first text. Typically, the second related information includes the expression type and expression description information of the named entity candidate. Here, by fuzzy matching of the pronunciation, if the named entity candidate has variant characters, the exact type of the named entity candidate can be identified. Variant characters refer to the writing of one character into another character. For example, the Chinese word "籃ボール場" can be written as "藍ボール場", where "藍" is a variant character.
[0039] (3) Using the named entity candidate in the first text as data, fuzzy matching is performed with expression names in the knowledge base that have the same length and a similarity greater than a predetermined threshold to obtain third related information of the named entity candidate in the first text. Typically, the third related information includes the expression type and performance description information of the named entity candidate. Here, by matching the length and similarity as described above, it is possible to identify the correct type of the named entity candidate when a typo exists in the named entity candidate. A typo refers to writing a character as a character that does not exist. For example, by writing the right half of the character "猴" as "危", a character that does not exist in the original is created.
[0040] By the above method, at least one expression type and at least one expression description information are obtained for each named entity candidate in the first text. As a result, a set of expression types corresponding to the named entity candidate is obtained from the at least one expression type, and a set of expression description information corresponding to the named entity candidate is obtained from the at least one expression description information.
[0041] In step 13, using the first text as the original text, adding the relevant information of the candidate proper expressions in the first text as additional text, and combining the first text and the relevant information of the candidate proper expressions in the first text to form a second text.
[0042] Here, there are multiple ways to combine the first text and the relevant information of the candidate proper expressions in the first text. For example, by sequentially combining the first text, the expression type to which the candidate proper expression in the first text belongs, and the expression description information of the candidate proper expression, a second text is formed. Also, for example, in the embodiment of the present invention, a second text is formed by filling the relevant information of the candidate proper expressions in the first text into a preset template. For example, in the case of Chinese, the form of the template is as follows.
[0043] [CLS]The first text ++ E1 is the set of expression types corresponding to E1. E1: The set of expression description information corresponding to E1. + E2 is the set of expression types corresponding to E2. E2: The expression description information corresponding to E2. + … + En is the set of expressions corresponding to En. En: The set of expression description information corresponding to En. [SEP].
[0044] In the above template, the underlined part represents the content to be filled. Ei represents the i-th candidate proper expression in the first text. Here, it is assumed that there are n candidate proper expressions in the first text. "[CLS]" is the start identifier of the second text, "[SEP]" is the text end identifier of the second text, and "" is the delimiter for separating the original text (here, the first text) and the additional text (here, the relevant information of the candidate proper expressions in the first text) in the input text.
[0045] Let the character sequence of the first text be \((x1, x2, \ldots, xn)\). Here, \(xi\) represents the \(i\)-th character of the first text. Let the two named entity candidate representations reproduced from the first text by the named entity reproduction model be representation \(E1\) and representation \(E2\) (assuming \(Ei\) is the sequence span of \(xm \ldots xm + t\)). As a result of searching for \(E1\) and \(E2\) in the external knowledge base, the set of representation types of \(E1\) is \((T1 - 1, T1 - 2, \ldots, T1 - n)\), the set of representation types of \(E2\) is \((T2 - 1, T2 - 2, \ldots, T2 - n)\), the set of representation description information of \(E1\) is \((D1 - 1, D1 - 2, \ldots, D1 - n)\), and the set of representation description information of \(E2\) is \((D2 - 1, D2 - 2, \ldots, D2 - n)\). Among them, \(T1 - i\) represents the \(i\)-th representation type of \(E1\) searched in the external knowledge base, and \(T2 - i\) represents the \(i\)-th representation type of \(E2\) searched in the external knowledge base. \(D1 - i\) represents the \(i\)-th representation description information of \(E1\) searched in the external knowledge base, and \(D2 - i\) represents the \(i\)-th representation information description of \(E2\) searched in the external knowledge base. In the above example, it is set that both the number of representation types of the searched information and the number of representation description information are \(n\). The second text generated based on the above template is as follows.
[0046] [CLS] \((x1, x2, \ldots, xn)\) + + \(E1\) is \((T1 - 1, T1 - 2, \ldots, T1 - n)\). \(E1:(D1 - 1, D1 - 2, \ldots, D1 - n)\). \(E2\) is \((T2 - 1, T2 - 2, \ldots, T2 - n)\). \(E2:(D2 - 1, D2 - 2, \ldots, D2 - n)\).[SEP].
[0047] Here, \([CLS]\) is the start identifier of the input text, \([SEP]\) is the end identifier of the input text, and is the delimiter for separating the original text and the additional text.
[0048] Next, an example of the second text generated by the above plate is shown.
[0049] Suppose the first text is: "The phrase'meter maid' spread among the cute Rita of the Beatles, and within that, this male singer was charmed by a female traffic warden." Suppose the named entity candidates obtained by the named entity reproduction model are "meter", "maid", "the Beatles", "cute Rita", and "Rita".
[0050] As expression candidates, input "meter", "maid", "the Beatles", "cute Rita", and "Rita" into the knowledge base for searching. The expression types of the searched "meter" are "book" and "other product". Also, the expression types of "maid" are "work" and "other product", and the expression description of "maid" is "a female employee does housework at the employer's home". The expression types of "the Beatles" are "book", "other product", and "band", and the expression description of "the Beatles" is "a famous British rock band, commonly known as the Beatles". The expression types of "cute Rita" are "work", "other product", and "music work". The expression description of "cute Rita" is "an original song written and composed by Lennon and McCartney, first recorded by the Beatles". The expression types of "Rita" are "artist", "other occupation", "work", "sports manager", "other product", and "sports player". The expression descriptions of "Rita" are "a female name", "a professional wrestler and singer in the United States", "surname", and "a pop singer and actor in Israel".
[0051] The second text formed by connecting these data according to the above template is as follows.
[0052] The term "meter maid" spread to the cute Rita of the Beatles, and within this context, this male singer was charmed by the female traffic warden. Meter is a work, other product. Maid is a book, other product. Maid: A female employee does housework at the employer's home. The Beatles are a book, other product, band. The Beatles: A famous British rock band, commonly known as the Beatles. Cute Rita is a book, other product, music work. Cute Rita: An original song written and composed by Lennon and McCartney, first recorded by the Beatles. Rita is an artist, other occupation, writing, sports manager, other product, sports player. Rita: A female name, an American professional wrestler and singer, surname, an Israeli pop singer and actor. In addition, the embodiments of the present invention do not specifically limit the method of combining the first text and the related information of the candidate proper expressions in the first text. The combined second text only needs to include the first text and the related information of the candidate proper expressions in the first text. That is, if the related information of the candidate proper expressions in the first text is introduced into the subsequent proper expression extraction, the purpose of improving the accuracy of the proper expression extraction can be achieved.
[0053] In step 14, input the second text into the proper expression extraction model that identifies proper expressions from the original text in the input text, and use the proper expression extraction model to identify proper expressions from the original text in the second text.
[0054] In the embodiments of the present invention, before step 14, a proper expression extraction model for extracting proper expressions from the original text in the input text is trained in advance. Then, in step 14, input the second text into the proper expression extraction model. The proper expression extraction model extracts proper expressions from the original text (i.e., the first text) in the input text (i.e., the second text). Thus, the proper expressions in the first text are extracted by the proper expression extraction model. Here, the proper expressions specifically include information such as expression names and expression types.
[0055] Continuing with the previous example for explanation. Using the "The phrase'meter maid' spread to the cute Rita of the Beatles, and within it, this male singer was charmed by a female traffic warden. Meter is a work, other product. Maid is a book, other product. Maid: A female employee does housework at the employer's home. The Beatles is a book, other product, band. The Beatles: A famous British rock band, commonly known as the Beatles. Cute Rita is a book, other product, music work. Cute Rita: An original song written and composed by Lennon and McCartney, first recorded by the Beatles. Rita is an artist, other occupation, writing, sports manager, other product, sports player. Rita: A female name, an American professional wrestler and singer, surname, an Israeli pop singer and actor." as the second text and inputting it into the named entity extraction model, from the original text "The phrase'meter maid' spread to the cute Rita of the Beatles, and within it, this male singer was charmed by a female traffic warden" in the second text, named entities are extracted by the named entity extraction model. In this example, as named entities, "The Beatles: band; Cute Rita: music work" are extracted.
[0056] In an embodiment of the present invention, the named entity extraction model is obtained by training using the fourth training data. The fourth training data is generated by combining the third training data as the original text and the related information of the named entity candidates in the third training data as the additional text. The third training data is training data of marked named entities and expression types.
[0057] The training of the named entity extraction model is described below. As shown in FIG. 2, the named entity extraction model is trained as follows.
[0058] In step 21, the third training data including named entity marking information is obtained, and using the named entity reproduction model, the named entity candidates in the third training data are reproduced.
[0059] In step 22, search for candidate named entities in the third training data in the knowledge base, and obtain the related information of the candidate named entities in the third training data. For the specific search method, refer to the above description.
[0060] In step 23, use the third training data as the original text, and use the related information of the candidate named entities in the third training data as additional text to combine the third training data and the related information of the candidate named entities in the third training data to form the fourth training data. For the specific combination method, refer to the above description.
[0061] In step 24, train the named entity extraction model using the fourth training data.
[0062] Here, based on a pre-training model (such as models like Bert / Albert / Roberta / Ernie, etc.), a named entity extraction model based on sequence tagging or a named entity extraction model based on spans can be constructed. Specifically, the named entity extraction model first generates a vector representation of the input text by the pre-training model. Here, the input text is the fourth training data. Then, it classifies the expressions for the original text in the input text (i.e., performs named entity extraction).
[0063] For example, input the fourth training data into a pre-training model (such as Bert / Albert / Roberta / Ernie, etc.) to obtain the corresponding sequence vector representation of the text (eCLS, e0, e1, e2, …, es, e0′, e1′, e2′, …, eSEP). Here, eCLS is the start vector representation of the fourth training data, eSEP is the end vector representation of the fourth training data, es is the vector representation of the delimiter between the original text and the additional text, and the other vector representations (for example, ei and ei′) are the vector representations of the tokens in the original text and the additional text in the fourth training data. After passing through the pre-training model, from the attention characteristics of the pre-training model, the tokens in the original text part have learned the additional text part. Next, use a general named entity extraction algorithm to map from the vector to the representation type in the token representation of the original text. To obtain the named entities in the original text, for example, an algorithm based on the representation span may be used, or a sequence representation method (BIO) may be used.
[0064] When training the named entity extraction model using the fourth training data, train with the macro-averaged F1 score (macro F1) of all named entity types as the optimization metric for model training to obtain the named entity extraction model.
[0065] Note that in the embodiments of the present invention, the second training data and the third training data may be the same training data or different training data, without particular limitation.
[0066] According to the embodiments of the present invention, through the above steps, the entity expression reproduction model reproduces the entity expression candidates in the original text, uses the entity expression candidates to search an external knowledge base, obtains additional information related to the entity expression candidates, and combines it with the original text, thereby obtaining a combined text including the additional information. In this way, when performing entity expression extraction using the combined text, since more additional information is introduced, the accuracy of fine-grained entity expression extraction can be improved. The method described above simply and effectively introduces external knowledge and uses a general search engine framework, so it is applicable to different forms of external knowledge. In addition, in the above method, a general entity expression extraction model is used, and its architecture is easily implemented.
[0067] Based on the above method, embodiments of the present invention further provide an apparatus for implementing the above method. As shown in FIG. 3, embodiments of the present invention provide an entity expression extraction apparatus including the following modules.
[0068] The first reproduction module 301 reproduces the entity expression candidates in the first text to be extracted by using an entity expression reproduction model that reproduces any type of entity expression in the text.
[0069] The first search module 302 searches a knowledge base having related information of a plurality of entity expressions including expression names, expression types, and expression description information based on the entity expression candidates in the first text, and obtains the related information of the entity expression candidates in the first text.
[0070] The first combination module 303 forms a second text by using the first text as the original text, using the related information of the entity expression candidates in the first text as additional text, and combining the first text and the related information of the entity expression candidates in the first text.
[0071] The extraction module 304 inputs the second text into a named entity extraction model that extracts named entities from the original text in the input text, and extracts named entities from the original text in the second text by the named entity extraction model.
[0072] According to the above template, the embodiments of the present invention can introduce more additional information when extracting named entities using the concatenated text, so as to improve the accuracy of thin named entity extraction.
[0073] Preferably, the above device further includes a first acquisition module that acquires first training data in which only named entities are marked and the corresponding expression types are not marked, and a first training module that trains a named entity reproduction model using the first training data with the recall rate of named entities as an optimization index for model training.
[0074] Preferably, the first acquisition module further acquires second training data having named entity marking information including named entities and expression types, identifies expression segments and non-expression segments in the second training data based on the named entity marking information, deletes the named entity marking information in the second training data, marks the expression segments in the second training data as named entities, and marks the non-expression segments as non-named entities to form first training data.
[0075] Preferably, the device further includes a construction module that, before searching for named entity candidates in the first text in the knowledge base, acquires knowledge data, cleans and organizes the knowledge data to form formatted data including the expression name, expression type, and expression description information of named entities, and constructs a knowledge base that supports exact match search and fuzzy search based on expression words based on the formatted data using a search engine framework.
[0076] Preferably, the first search module uses the candidate named entities in the first text as data, performs an exact match with the expression names in the knowledge base, and obtains the first related information of the candidate named entities in the first text; uses the candidate named entities in the first text as data, performs a fuzzy match with the pronunciations of the expression names in the knowledge base, and obtains the second related information of the candidate named entities in the first text; uses the candidate named entities in the first text as data, performs a fuzzy match with the expression names of the same length and a similarity greater than a predetermined threshold in the knowledge base, and obtains the third related information of the candidate named entities in the first text; Based on at least one of the above methods, searches a knowledge base containing related information of a plurality of named entities including expression names, expression types, and expression description information based on the candidate named entities in the first text, and obtains the related information of the candidate named entities in the first text.
[0077] Preferably, the apparatus uses the named entity reproduction model to obtain third training data including the named entity marking information before reproducing the candidate named entities in the first text to be extracted; uses the named entity reproduction model to reproduce the candidate named entities in the third training data; searches the knowledge base based on the candidate named entities in the third training data, and obtains the related information of the candidate named entities in the third training data; uses the third training data as the original text, uses the related information of the candidate named entities in the third training data as additional text, combines the third training data and the related information of the candidate named entities in the third training data to form fourth training data; A second training module for training the named entity extraction model using the fourth training data, is further included.
[0078] Here, each device / system provided by the above embodiments corresponds to the named entity extraction method described above. Any implementation method of each embodiment is applicable to the device embodiments described above, and the same technical effects can be achieved. The device provided by the embodiments of the present invention can implement the steps of the method implemented by the above method embodiments and achieve the same technical effects. Here, specific descriptions of the same parts and effects as those of the method embodiments of the present embodiment are omitted.
[0079] FIG. 4 shows a block diagram showing the hardware configuration of a named entity extraction device according to an embodiment of the present invention. As shown in FIG. 4, the named entity extraction device 400 includes a processor 402 and a memory 404 in which computer program instructions are stored.
[0080] By the computer program instructions being executed by the processor 402, steps are implemented including: reproducing named entity candidates in a first text using a named entity reproduction model that reproduces any type of named entity in text; searching a knowledge base having associated information of a plurality of named entities including expression names, expression types, and expression description information based on the named entity candidates in the first text, and obtaining the associated information of the named entity candidates in the first text; using the first text as the original text, using the associated information of the named entity candidates in the first text as additional text, and combining the first text and the associated information of the named entity candidates in the first text to form a second text; inputting the second text into a named entity extraction model that extracts named entities from the original text in the input text, and extracting named entities from the original text in the second text by the named entity extraction model.
[0081] Here, each device / system provided by the above embodiments corresponds to the above-described named entity extraction method, and any implementation manner of each embodiment can be applied to the device embodiments above, and the same technical effects can be achieved. The device provided by the embodiments of the present invention can implement the method steps of the method implemented by the above embodiments, and has the same technical effects. Here, specific descriptions of the same parts and effects as those of the method embodiments of this embodiment are omitted.
[0082] Furthermore, as shown in FIG. 4, the named entity extraction device 400 includes a network interface 401, an input device 403, a hard disk 405, and a display device 406.
[0083] The above interfaces and devices are interconnected with each other via a bus architecture. The bus architecture is a bus and bridges including any number of interconnections. Specifically, one or more central processing units (CPUs) represented by the processor 402 and / or graphics processing units (GPUs), and various circuits of one or more memories represented by the memory 404 are connected. The bus architecture is connected to various other circuits such as peripheral devices, voltage regulators, and power management circuits. The bus architecture is used to communicatively connect between these devices. The bus architecture includes a data bus, a power bus, a control bus, and a status signal bus in addition to the data bus. Since these are all well-known technologies in the field of the present invention, detailed descriptions are omitted.
[0084] The network interface 401 is connected to a network (such as the Internet, a local area network, etc.), receives external knowledge data via the network, and stores the received data in the hard disk 405.
[0085] The input device 403 receives various commands input by the user and sends them to the processor 402 for execution. The input device 403 includes a keyboard or a click device (for example, a mouse, a trackball, a touch panel, or a touch screen, etc.).
[0086] The display device 406 displays the results of the commands executed by the processor 402. For example, it displays the progress of model training and the like.
[0087] The memory 404 stores programs and data necessary for the execution of the operating system, and data such as intermediate results in the calculations by the processor 402.
[0088] In an embodiment of the present invention, the memory 404 is a volatile memory or a non-volatile memory, or includes both volatile and non-volatile memories. Among them, the non-volatile memory is a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory is a random access memory (RAM) used as an external cache. The memory 404 of the devices and methods described in the description of this specification includes these memories and any other suitable types of memories, but is not limited thereto.
[0089] In some embodiments, the memory 404 stores the operating system 4041 and the application program 4042 as executable modules or data structures, or subsets thereof, or extended sets thereof.
[0090] Among them, the operating system 4041 includes various system programs, such as a framework layer, a core library layer, and a driver layer, and is used to implement various core operations and hardware-based tasks. The application program 4042 includes various application programs, such as a web browser (Browser), etc., and is for implementing various application operations. The program for executing the method according to this embodiment is included in the application program 4042.
[0091] The method according to an embodiment of the present invention is applied to, or implemented by, a processor 402. The processor 402 is a type of integrated circuit chip having a function of processing signals. In the implementation process, each step of the method is implemented by an integrated logic circuit of hardware in the processor 402 or a command in the form of software. The processor 402 is a general-purpose processor, a digital signal processing device (DSP), an application-specific integrated circuit (ASIC), a ready-made programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute each method, step and logic box disclosed in the embodiments of the present invention. The general-purpose processor is a microprocessor or any general processor, etc. Each step of the method according to an embodiment of the present invention may be implemented by being executed by a decoder that is hardware, or may be implemented by a combination of hardware and software that can be performed in the decoder. The software module is stored in a storage medium mature in this field, such as random memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. From a memory 404 provided with the storage medium in which this software is stored, the processor 402 reads information and realizes the steps of the above method according to the hardware.
[0092] The embodiments described above are implemented by hardware, software, firmware, middleware, microcode, or a combination thereof. Among them, regarding the hardware implementation, the processing unit is implemented by one or more application-specific integrated circuits (ASICs), digital signal processing processors (DSPs), digital signal processing devices (DSPDs), programmable logic circuits (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units that execute the functions of the present invention, or a combination thereof.
[0093] Regarding the implementation of software, the above technology is implemented by modules (such as processes, functions, etc.) that implement the functions described above. The software code is stored in memory and executed by a processor. Note that the memory is implemented either inside or outside the processor.
[0094] Specifically, when a computer program is executed by the processor 402, before the step of reproducing candidate specific expressions in the first text using a specific expression reproduction model that reproduces any type of specific expression in the text, a step of obtaining first training data in which only specific expressions are marked and the corresponding expression types are not marked; and a step of training the specific expression reproduction model using the first training data with the reproduction rate of the specific expression as an optimization index for model training are further implemented.
[0095] Also, specifically, when a computer program is executed by the processor 402, a step of obtaining training data including specific expression marking information including specific expressions and expression types as second training data; a step of identifying expression segments and non-expression segments in the second training data based on the specific expression marking information; and a step of deleting the specific expression marking information in the second training data, marking the expression segments in the second training data as specific expressions, and marking the non-expression segments as non-specific expressions to form first training data are further implemented.
[0096] Also, specifically, when the computer program is executed by the processor 402, before searching for candidate specific expressions in the first text in the knowledge base, a step of obtaining knowledge data, cleaning and organizing the knowledge data, and forming formatted data including the expression name, expression type, and expression description information of the specific expression; and a step of constructing a knowledge base that supports exact match search and fuzzy search based on the expression word based on the formatted data using a search engine framework are further implemented.
[0097] Specifically, in order to search a knowledge base including related information of a plurality of specific expressions including an expression name, an expression type, and expression description information based on the specific expression candidates in the first text and obtain the related information of the specific expression candidates in the first text when the computer program is executed by the processor 402, any one of the following is adopted: using the specific expression candidates in the first text as data, performing exact matching with the expression name in the knowledge base to obtain first related information of the specific expression candidates in the first text; using the specific expression candidates in the first text as data, performing fuzzy matching with the pronunciation of the expression name in the knowledge base to obtain second related information of the specific expression candidates in the first text; and using the specific expression candidates in the first text as data, performing fuzzy matching with an expression name having the same length and a similarity greater than a predetermined threshold in the knowledge base to obtain third related information of the specific expression candidates in the first text.
[0098] Furthermore, specifically, when the computer program is executed by the processor 402, before reproducing the specific expression candidates in the first text to be extracted using the specific expression reproduction model, third training data including the specific expression marking information is obtained, and the steps of reproducing the specific expression candidates in the third training data using the specific expression reproduction model, searching the knowledge base based on the specific expression candidates in the third training data, and obtaining the related information of the specific expression candidates in the third training data are performed. Then, using the third training data as the original text and the related information of the specific expression candidates in the third training data as additional text, the third training data and the related information of the specific expression candidates in the third training data are combined to form fourth training data, and the specific expression extraction model is trained using the fourth training data.
[0099] Note that the above device according to the embodiment of the present invention can implement all the method steps implemented by the above method embodiment and achieve the same technical effect. Therefore, specific descriptions of the same parts and effects as those of the method embodiment of this example are omitted here.
[0100] Some embodiments of the present invention provide a computer-readable storage medium storing a program, which, when executed by a processor, uses a unique expression reproduction model for reproducing any type of unique expression in text to reproduce a unique expression candidate in a first text to be extracted, searches a knowledge base having related information of a plurality of unique expressions including an expression name, an expression type, and expression description information based on the unique expression candidate in the first text, and obtains related information of the unique expression candidate in the first text; forms a second text by combining the first text as an original text and the related information of the unique expression candidate in the first text as additional text; and inputs the second text to a unique expression extraction model for extracting a unique expression from the original text in the input text, and extracts a unique expression from the original text in the second text by the unique expression extraction model.
[0101] When the above program is executed by a processor, all implementation manners in the unique expression extraction method are implemented and achieve the same technical effect. To avoid repetition, the description is omitted here.
[0102] Those skilled in the technical field of the present invention can easily conceive that the units and algorithm steps of each example described in the embodiments disclosed above can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software depends on the specific application and design constraints of the invention. Those skilled in the art can implement the above functions in a manner corresponding to a specific application, but should not exceed the scope of the present invention.
[0103] Also, for the sake of convenience and conciseness in description, regarding the specific working processes of the above system, device, and unit, it is obvious to those skilled in the art that the corresponding processes in the above-described embodiments can be referred to, so detailed descriptions are omitted.
[0104] It can be easily conceived that the methods and devices disclosed in multiple embodiments of the present invention can also be realized in other forms. For example, the devices described above are only schematic. For example, the division of the above-described units is only an example of the allocation of logical functions, and other division methods may be adopted during actual implementation. For example, multiple units or modules can be combined, or aggregated into another system, or some functions can be omitted, or not executed. It should be noted that the above-indicated or disclosed mutual connections, direct connections, or communicable connections are connections via an interface. The indirect connections or communicable connections between devices or units may be electrical, mechanical, or other forms of connections.
[0105] The unit described as the separation member may or may not be physically separated. The member shown as a unit may or may not be a physical unit. That is, it may be in the same place or distributed on multiple network units. Select some or all of the units according to actual needs to achieve the purpose of the embodiments of the present invention.
[0106] It should be noted that each functional unit according to the embodiments of the present invention may be aggregated into one processing unit, may be physically independent, or may be aggregated into one unit with two or more.
[0107] When the above-described function is implemented in the form of a software functional unit and sold or used as an independent product, the function can be stored in a computer-readable storage medium. Based on such an understanding, the essence of the technical concept of the present invention, or the part that contributes to the prior art, or the part of the technical concept can be embodied in the form of a software product. This computer software product is stored in a storage medium, includes instructions, and causes a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or some of the steps of the methods described in each embodiment of the present invention. The above-described storage medium includes various media that can store program codes, such as a USB memory, a removable hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0108] The above description is not overly similar to the specific implementation manners of the present invention and does not limit the scope of protection of the present invention. Changes or substitutions that can be easily conceived by those skilled in the art within the scope disclosed by the present invention are included in the scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of the claims.
Claims
1. A method for extracting specific expressions executed by a specific expression extraction device, reproducing candidate specific expressions in a first text to be extracted using a specific expression reproduction model that reproduces specific expressions of any type in the text; searching a knowledge base based on the candidate specific expressions in the first text and obtaining related information of the candidate specific expressions in the first text, where the knowledge base includes related information of a plurality of specific expressions, and the related information of the specific expressions includes an expression name, an expression type, and expression description information of the specific expression; using the first text as the original text, using the related information of the candidate specific expressions in the first text as additional text, and obtaining a second text by combining the first text and the related information of the candidate specific expressions in the first text; inputting the second text into a specific expression extraction model that extracts specific expressions from the original text in the input text, and extracting specific expressions from the original text in the second text by the specific expression extraction model. The specific expression extraction method is characterized by including the above steps.
2. Before reproducing the candidate specific expressions in the first text to be extracted using a specific expression reproduction model that reproduces specific expressions of any type in the text, obtaining first training data in which only specific expressions are labeled and the corresponding expression types are not labeled; further including training a specific expression reproduction model using the first training data with the reproduction rate of specific expressions as an optimization index for model training. The specific expression extraction method according to claim 1 is characterized by including the above steps.
3. The step of obtaining first training data in which the specific expressions are labeled and the expression types are not labeled includes: obtaining second training data having specific expression label information, where the specific expression label information includes specific expressions and expression types; identifying expression segments and non-expression segments in the second training data based on the specific expression label information. Deleting the named entity label information in the second training data, labeling the expression segments in the second training data as named entities, and labeling the non-expression segments as non-named entities to obtain first training data; and the method for extracting named entities according to claim 2, characterized by including this step.
4. Before searching the knowledge base based on the named entity candidates in the first text, obtaining knowledge data, cleaning and sorting the knowledge data, and obtaining formatted data including the expression name, expression type, and expression description information of the named entity; and further including the step of using a search engine framework to build a knowledge base that supports exact matching search and fuzzy search based on expression words based on the formatted data, and the method for extracting named entities according to claim 1, characterized by including this step.
5. The step of searching the knowledge base based on the named entity candidates in the first text and obtaining the related information of the named entity candidates in the first text is using the named entity candidates in the first text as data to perform exact matching with the expression name in the knowledge base to obtain the first related information of the named entity candidates in the first text; and using the named entity candidates in the first text as data to perform fuzzy matching with the pronunciation of the expression name in the knowledge base to obtain the second related information of the named entity candidates in the first text; and using the named entity candidates in the first text as data to perform fuzzy matching with the expression name having the same length and a similarity greater than a predetermined threshold in the knowledge base, and the method for extracting named entities according to claim 4, characterized by including any one of obtaining the third related information of the named entity candidates in the first text.
6. Before using the named entity reproduction model to reproduce the named entity candidates in the first text to be extracted, obtaining third training data including named entity label information, and using the named entity reproduction model to reproduce the named entity candidates in the third training data; and searching the knowledge base based on the named entity candidates in the third training data and obtaining the related information of the named entity candidates in the third training data; and Using the third training data as the original text and the related information of the candidate named entities in the third training data as additional text, combining the third training data and the related information of the candidate named entities in the third training data to obtain the fourth training data; further comprising training the named entity extraction model using the fourth training data, wherein the named entity extraction method according to claim 2 is characterized by the above. **Claim 7** A first reproduction module that reproduces candidate named entities in the first text to be extracted using a named entity reproduction model that reproduces any type of named entity in text; A first search module that searches a knowledge base based on candidate named entities in the first text and obtains related information of the candidate named entities in the first text, wherein the knowledge base includes related information of a plurality of named entities, and the related information of the named entity includes the expression name, expression type, and expression description information of the named entity; A first combination module that uses the first text as the original text, the related information of the candidate named entities in the first text as additional text, and combines the first text and the related information of the candidate named entities in the first text to obtain a second text; An extraction module that inputs the second text into a named entity extraction model that extracts named entities from the original text in the input text, and extracts named entities from the original text in the second text by the named entity extraction model, wherein the named entity extraction device is characterized by the above. **Claim 8** A first acquisition module that acquires first training data in which only named entities are labeled and the corresponding expression types are not labeled; further comprising a first training module that trains a named entity reproduction model using the first training data with the reproduction rate of named entities as an optimization index for model training, wherein the named entity extraction device according to claim 7 is characterized by the above. **Claim 9** The first acquisition module further acquires second training data having named entity label information, wherein the named entity label information includes named entities and expression types; identifies expression segments and non-expression segments in the second training data based on the named entity label information; The apparatus for extracting named entities according to claim 8, characterized in that the apparatus obtains first training data by deleting named entity label information in the second training data, labeling expression segments in the second training data as named entities, and labeling non-expression segments as non-named entities.
10. Before searching the knowledge base based on named entity candidates in the first text, the apparatus further includes a construction module that obtains knowledge data, cleans and arranges the knowledge data to obtain formatted data including expression names, expression types, and expression description information of named entities, and constructs a knowledge base that supports exact match search and fuzzy search based on expression words based on the formatted data using a search engine framework.
11. The first search module uses named entity candidates in the first text as data to perform exact match matching with expression names in the knowledge base to obtain first related information of named entity candidates in the first text; uses named entity candidates in the first text as data to perform fuzzy matching with the pronunciations of expression names in the knowledge base to obtain second related information of named entity candidates in the first text; uses named entity candidates in the first text as data to perform fuzzy matching with expression names having the same length and a similarity greater than a predetermined threshold in the knowledge base, and obtains third related information of named entity candidates in the first text, and searches the knowledge base based on named entity candidates in the first text to obtain related information of named entity candidates in the first text. The apparatus for extracting named entities according to claim 10, characterized in that the apparatus searches the knowledge base based on named entity candidates in the first text to obtain related information of named entity candidates in the first text.
12. Before reproducing named entity candidates in the first text to be extracted using the named entity reproduction model, the apparatus further includes a second acquisition module that obtains third training data including named entity label information, a second reproduction module that reproduces named entity candidates in the third training data using the named entity reproduction model, and a second search module that searches the knowledge base based on named entity candidates in the third training data to obtain related information of named entity candidates in the third training data. a second reproduction module that reproduces named entity candidates in the third training data using the named entity reproduction model; a second search module that searches the knowledge base based on named entity candidates in the third training data to obtain related information of named entity candidates in the third training data. Using the third training data as the original text and the related information of the named entity candidate in the third training data as the additional text, a second concatenation module that concatenates the third training data and the related information of the named entity candidate in the third training data to obtain fourth training data, A second training module that trains the named entity extraction model using the fourth training data, the named entity extraction device according to claim 8 or 9, further comprising this.
13. A program that causes a processor to execute the named entity extraction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge processing method and device, electronic equipment and storage medium
CN116910250A
Language processing device, machine learning method, estimation method, and program
JP2023181819A