An English tweet named entity extraction method and device based on subjective and objective word lists
By constructing subjective and objective vocabulary list and tree-shaped parent-child structure, the problem of subjective words affecting naming entity recognition in English tweets is solved, and efficient and accurate extraction of naming entities is achieved.
Patent Information
- Application Number
- CN202211458427.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-11-21
AI Technical Summary
There are a large number of subjective words in English tweets that affect the recognition performance of named entities, and it is difficult for the existing technology to effectively extract named entities.
Construct a subjective and objective vocabulary list, extract noun phrases and noun clauses through a grammatical dependence analysis model, construct a tree-shaped parent-child structure for naming entity extraction, use multi-word list to filter invalid information, and perform recursive noun extraction.
It improves the accuracy of naming entity recognition, effectively filters subjective words and invalid information in informal English tweets, and improves the effect of naming entity recognition.
Smart Images

Figure CN116127971B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and particularly relates to a method and device for extracting English tweet named entities based on subjective and objective word lists. Background Art
[0002] With the rapid development of the Internet and the information industry, a large amount of text data is continuously generated. How to efficiently obtain useful information from the large amount of text data has become a current research hotspot, and information extraction technology has emerged as the times require. Named entity recognition is a subtask of information extraction, and its purpose is to extract specified entities from a large amount of text data. In the field of natural language processing applications, named entity recognition is a basic task for many natural language processing applications such as information retrieval, machine translation, and sentiment analysis. Therefore, the research on it has important significance and value.
[0003] Currently, the research techniques for named entity recognition are mainly divided into four different methods: rule-based methods, unsupervised learning methods, traditional supervised machine learning methods, and deep learning-based methods. However, due to the informality of tweets, there are a large number of subjective words in them, which will affect the performance of named entity (i.e., noun phrase) recognition. Summary of the Invention
[0004] In view of the above analysis, the present invention aims to provide a method and device for extracting English tweet named entities based on subjective and objective word lists; to solve the problem that a large number of subjective words in existing English tweets affect the subsequent named entity recognition performance.
[0005] The object of the present invention is mainly achieved through the following technical solutions:
[0006] On the one hand, the present invention provides a method for extracting English tweet named entities based on subjective and objective word lists, including the following steps:
[0007] Obtain English texts in multiple fields and construct a corpus;
[0008] Perform word segmentation and word frequency statistics on the texts in the corpus, and construct a subjective word list through screening;
[0009] Preprocess the English tweet to be recognized to obtain a standard tweet text;
[0010] Use a syntactic dependency analysis model to extract all noun phrases and noun clauses in the standard tweet text, and preprocess the noun phrases and noun clauses based on the subjective word list to construct a set NP p ;
[0011] Based on the set NP pConstruct a tree - shaped parent - child structure for the noun phrases and noun clauses in it and perform named entity extraction to obtain the named entity recognition result of the English tweet.
[0012] Further, the constructing of the tree - shaped parent - child structure and performing named entity extraction includes:
[0013] Based on the inclusion relationship of the noun phrases and noun clauses in the set NP p Take the noun phrase or noun clause containing at least one noun phrase as the parent string, and the noun phrases contained in the parent string as the child strings to construct a tree - shaped parent - child structure;
[0014] Based on the possessive structure of the nouns of each parent string, extract its core noun and save it to the named entity set
[0015] Remove all child strings from each parent string and merge the remaining content into a new string The set formed by all child strings is denoted as cp;
[0016] If the string meets the corresponding preset conditions, save the string to the named entity set Otherwise, based on the string Re - use the syntactic dependency parsing model to extract all noun phrases and noun clauses, and save the extracted noun phrases and noun clauses to the named entity set Merge the child string set cp into the named entity set and re - construct the tree - shaped parent - child structure to perform named entity extraction to obtain the named entity recognition result.
[0017] Further, re - construct the tree - shaped parent - child structure multiple times to perform named entity extraction to obtain the named entity recognition result; among them, the named entity set obtained by the j - th re - construction of the tree - shaped parent - child structure for named entity extraction is denoted as where j is an integer greater than 1; the named entity set obtained by the (j - 1) - th re - construction of the tree - shaped parent - child structure for named entity extraction is denoted as When the difference between the set and the set is an empty set, output the set as the named entity recognition result of the English tweet.
[0018] Further, the preset conditions corresponding to the string include: the string has only one word except for the contained child strings, and this word is a preposition or a conjunction, and the string The number of substrings is 2, the absolute value of the difference in the lengths of the two substrings is not greater than 2, and the lengths of both substrings are not greater than 3.
[0019] Further, extracting the core nouns based on the possessive noun structures of each parent string includes:
[0020] Based on the possessive noun structure, starting from the single quote (') and traversing to the left until the first capitalized word after the previous single quote; for XX YY's and XX YY'z structures, extracting the content between the first capitalized word and YY in the XXYY's or XX YY'z structure to obtain the possessive noun core noun; for XX YYs' and XX YYz' structures, extracting the content between the first capitalized word and YYs or YYz in the XX YYs' or XX YYz' structure to obtain the possessive noun core noun, where X and Y represent any words.
[0021] Further, obtaining English texts in multiple fields and constructing a corpus includes:
[0022] Obtaining named entities in multiple fields and merging them to obtain an initial entity list L init ;
[0023] Obtaining the entity ID corresponding to each entity in the initial entity list L init to obtain an entity ID list
[0024] Obtaining the entity IDs of other entities that have an existence relationship with each entity ID in the entity ID list and the entity IDs of other entities of the same category as each entity in the entity ID list and merging them to obtain an entity ID list
[0025] Using to replace Obtaining the other entities that have an existence relationship with each entity ID in the entity ID list and the entity IDs of other entities of the same category as each entity in the entity ID list and merging them to obtain an entity ID list
[0026] Obtaining the abstracts of the entities corresponding to all entity IDs in it, and constructing the corpus based on the obtained abstracts.
[0027] Further, constructing a subjective word list includes:
[0028] Tokenize the abstracts in the corpus and count the word frequencies, remove words with a word frequency less than the first preset threshold and a word length of 1 to obtain the objective word list Objws;
[0029] Obtain the word list Trws by getting nouns, adverbs, adjectives, numbers, prepositions, conjunctions, and articles in the ECDICT dictionary with a word frequency ranking less than the second preset threshold and a Collins star rating greater than 1;
[0030] Remove the words in the word list Objws from the word list Trws to obtain the subjective word list Subjws.
[0031] Furthermore, preprocess the noun phrases based on the subjective word list to construct a noun phrase set NP p ; including: traversing all noun phrases, removing the leading articles, quantifiers, and stop words therein; and removing the subjective words therein based on the subjective word list to obtain the preprocessed noun phrases to obtain the noun phrase set NP p .
[0032] Furthermore, for the leading articles, quantifiers, and subjective words, use the corresponding regular expressions to remove them respectively;
[0033] For the stop words, use the Aho-Corasick automaton to remove them; the Aho-Corasick automaton is constructed through a pre-built stop word list.
[0034] On the other hand, the present invention also provides a computer device, including at least one processor and at least one memory communicatively connected to the processor;
[0035] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the English tweet named entity extraction method involved in the present invention.
[0036] Advantages of the technical solution:
[0037] 1. The present invention aims at informal English tweet texts. Considering that there are a large number of subjective words in tweets, which affect downstream tasks of natural language processing, it uses texts on knowledge-based websites such as Wikipedia that have been revised by multiple parties and have standardized word usage to construct two word lists, respectively representing subjective words and objective words, which is convenient for subsequent corresponding content filtering.
[0038] 2. The present invention aims at the problem that the noun phrases extracted by the Transformer-based syntactic dependency analysis model may include some unnecessary information (such as invalid conjunctions, prepositions, etc.), or there are various noun phrases and noun clauses. It filters out invalid information through multiple word lists, and at the same time constructs a tree structure on the identified noun phrase set for recursive noun extraction, improving the accuracy of named entity recognition.
[0039] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the written description, claims, as well as the drawings. Description of the Drawings
[0040] The drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals denote the same components.
[0041] Figure 1 It is a flowchart of the method for extracting named entities in English tweets according to the embodiments of the present invention. Detailed Embodiments
[0042] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings constitute a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0043] A method for extracting named entities in English tweets based on subjective and objective word lists in this embodiment is as Figure 1 shown and includes the following steps:
[0044] Step S1: Obtain English texts in multiple fields and construct a corpus;
[0045] Compared with informal tweet texts, the texts on knowledge websites such as Wikipedia are relatively formal texts that have been revised by multiple parties and have standardized vocabulary, which is convenient for subsequent training of deep models. At the same time, considering that a large number of subjective words may be used in tweets, affecting downstream tasks of natural language processing, two word lists will be constructed to represent subjective words and objective words respectively, which is convenient for subsequent corresponding content filtering.
[0046] Specifically, when constructing the corpus, first obtain named entities in multiple fields and merge them to obtain an initial entity list L init ; Specifically, entities from different fields can be specified empirically, including fields such as politics, agriculture, information technology, religion, film and television entertainment, video games, education, medicine, etc., or several entities that are most concerned by the public recently can be listed based on Google Trends.
[0047] Further obtain the entity ID corresponding to each entity in the initial entity list L init to obtain an entity ID list Preferably, the WikiPageID corresponding to each entity can be queried using the SPARQL interface of DBpedia to form an entity ID list
[0048] Further use the SPARQL interface of DBpedia to obtain other entity IDs that have relationships with the entities corresponding to each entity ID in the entity ID list (for example, "capital" and another related entity "Beijing" that have relationships with the entity "China"), and the entity ID list The entity IDs corresponding to other entities of the same category as each entity in the entity ID list are merged to obtain an entity ID list Specifically, the category can be the Wiki category to which each entity is assigned, which is an attribute of the Wiki entity and was originally named Category in Wikipedia
[0049] With Replace To obtain the entity IDs of other entities that have relationships with the entities corresponding to each entity ID in the entity ID list and the entity ID list The entity IDs corresponding to other entities of the same category as each entity in the entity ID list are merged to obtain an entity ID list Through the method of this embodiment, the constructed entity ID list involves approximately 130,000 entities
[0050] Further obtain The abstracts of the entity pages corresponding to all entity IDs in the entity ID list, and a corpus is constructed based on all the obtained abstracts
[0051] Step S2: Tokenize the text in the corpus and perform word frequency statistics, and construct a subjective word list through screening
[0052] Specifically, after constructing the corpus, tokenize the abstracts in the corpus and count the word frequencies, and remove the words with a word frequency less than the first preset threshold and a word length of 1 to obtain an objective word list Objws; specifically, in this embodiment, the first threshold is set to 300, and the words with a word frequency less than 300 and a word length equal to 1 are removed to form the objective word list Objws
[0053] Further obtain nouns, adverbs, adjectives, numbers, prepositions, conjunctions, and articles in the ECDICT dictionary whose word frequency rankings are less than the second preset threshold and whose Collins star rating is greater than 1 to obtain a word list Trws; in this embodiment, the second threshold is set to 10,000, that is, the word list Trws contains common words with a word frequency ranking less than 10,000 and a Collins star rating greater than 1
[0054] Remove the words in the word list Objws from the word list Trws, and the remaining words are mostly words expressing subjective feelings such as nice and beautiful, and a subjective word list Subjws is constructed
[0055] Step S3: Preprocess the English tweet to be recognized to obtain a standard tweet text;
[0056] Specifically, the tweet can be preprocessed by existing methods such as text standardization, text tokenization, text cleaning, and word standardization. For the text preprocessing of English tweets, operations such as regularization, removal, and restoration can also be performed by identifying morphemes in specific formats.
[0057] In this embodiment, according to the characteristics of English tweets, when preprocessing, the English tweet is respectively subjected to tweet semantic standardization, informal morpheme standardization, correction of non-standard capitalized words, extraction of sentence-end value information, and clause extraction to obtain a regular standard tweet text. Specifically, the tweet preprocessing can be performed through the following steps:
[0058] Step S301: Tweet semantic standardization, including: restoring tag semantics, extracting secondary morphemes, restoring special semantic punctuation, and separating word-end punctuation:
[0059] (1) Restoring tag semantics:
[0060] The tag in the tweet refers to the text segment starting with "#", such as "#BlackLivesMatter". Its characteristic is that the words in the tag are usually merged, without spaces and punctuation. To address the problem of poor readability of tags in tweets, the "tag semantic restoration method based on the composite strategy of generalized camel case naming method, dictionary, and greedy search algorithm" is used to restore the tag semantics. That is, if the tag conforms to the "Generalized Camel Case" (abbreviated as GCC), the tag semantic restoration is performed according to its naming rules; otherwise, the "matching method integrating dictionary and greedy algorithm" is used for composite strategy tag semantic restoration.
[0061] Specifically, first, extract the tags in the tweet and form a tag set. Determine whether each tag conforms to the generalized camel case naming method. If it conforms, use regular expressions for semantic restoration and output the semantic restoration result of the tag. If it does not conform, perform semantic restoration through the matching method integrating dictionary and greedy algorithm.
[0062] Among them, the camel case naming method is a method for writing phrases without spaces and punctuation marks, and its applicable scope is extended to phrases including numbers, serial numbers, and multiple consecutive capitalized words. Its grammar is shown in Table 1;
[0063] Table 1 Generalized Camel Case Naming Method Grammar Table
[0064]
[0065] In this embodiment, the following regular expression is used to match the syntax rules of all generalized camel algorithms, and the semantic restoration of tags that conform to the "generalized camel naming method" is performed.
[0066] (?P <content>[A-Z]?[a-z]+|[A-Z]+(?!=[A-Z][a-z]+|$|\d)|\d+(?:st|nd|rd|th)?|(?<=[a-z])[A-Z][a-z]+?(?=[^a-z]|$))。
[0067] Furthermore, the matching method integrating the dictionary and the greedy algorithm includes:
[0068] First, query words through the ECDICT dictionary. If no hit result is found in the ECDICT dictionary, then decompose them using WordNinja based on the greedy search algorithm. Specifically:
[0069] First, construct a fuzzy query statement, that is, "SELECT * FROM stardict WHERE sw Like hashtag", where "hashtag" is the target label, and query whether this label exists in the ECDICT dictionary;
[0070] If the current label cannot match any ECDICT records, then use WordNinja for label semantic reduction, that is, an algorithm for maximum greedy matching on the known word list;
[0071] If the current label can match a word in the ECDICT dictionary, then directly query the ECDICT dictionary to determine whether the current label consists of only one word. If so, return the word; otherwise, use the "heuristic Gestalt pattern matching algorithm" to calculate the similarity between the current label and all retrieved ECDICT records, and then return the most similar phrase. Specifically, if the current label consists of only one word, the corresponding word will be matched; if the current label is a phrase, there will be no matching result. Specifically, there is no space between tweet labels. For example, if the current label is "SameAs", although there is an entry "same as" in ECDICT, since there is no space between "Same" and "As" in the current label, it will not be matched; similarly, if the current label has only one word, such as "same", the corresponding entry will be matched.
[0072] (2) Extract secondary morphemes:
[0073] The secondary morphemes in tweets refer to morphemes such as emoticons, kaomoji, and modal particles that express the author's mood or expression habits, as well as the syntax "@username" indicating the mention of a certain user;
[0074] Among them, since "@username" is a tweet syntax and its structure does not provide semantic information; therefore, for the "@username" syntax, a random string is generated, starting with an uppercase letter and followed by 5 random lowercase letters, as the nickname of the mentioned user;
[0075] In addition, tweets crawled from Twitter need to go through HTML encoding, first converted to UTF8 or other encoding formats, that is, decoded, or called transcoded. However, partial failure may occur during the conversion of some HTML texts, so processing is required. For the texts with decoding errors in tweets, the replacement content is determined according to the decoding table and replaced; for example, "&" needs to be replaced with "and", and for invalid morphemes such as "RT" indicating a reply at the beginning, multimedia links, email addresses, etc., corresponding regular expressions are used to remove them. Here, invalid morphemes refer to language elements in English tweets that cannot provide valid information, contain errors, or are non-English.
[0076] (3) Restore special semantic punctuation, including:
[0077] Remove the paired square brackets "[]"; restore the equal sign "=" to "equals to";
[0078] For the tilde "~" or hyphen "-" followed by a person's name, obtain the tilde and hyphen that may be followed by a person's name according to the following regular expression:
[0079] (?<=(?P <ispunc>[^\s])){0,1}(?P <pres>\s{0,})(?P [-\~])(?=(?P <lasts>\s{0,})(?P <name>(([A-Z][^\s]+\s?)+));
[0080] Then make a special judgment on the hyphen: If the punctuation being judged currently is a hyphen and there is no space before or after it, and the previous element is non-punctuation, then it is regarded as a word-forming hyphen rather than a hyphen with special semantics.
[0081] Finally, if the punctuation (referring to the tilde and hyphen) is followed by a person's name, the restored content also needs to consider the context morphemes of the tilde and hyphen, as shown in Table 2.
[0082] Table 2 Restoration content table corresponding to the relative positions of the tilde or hyphen and other morphemes in the sentence
[0083]
[0084]
[0085] For the tilde "~" adjacent to a number, obtain the tildes that may be followed by numbers according to the following regular expression:
[0086] (?P <prefigure>(?:\d|_NUM)){0,1}\s{0,}(?P <target>\~)\s{0,}(?=(?P <lastfigure>(?:\d|_NUM));
[0087] Among them, _NUM represents possible numerical words.
[0088] If there is no number before the tilde and a number follows it, replace it with "approximately"; if there are numbers before and after, replace it with "to".
[0089] (4) Separate word-terminal punctuation:
[0090] Word-terminal punctuation refers to the punctuation at the beginning and end of a word. For the convenience of subsequent clause extraction, add spaces on both sides of the punctuation (word-terminal punctuation) located before and after the word. In addition, since the possessive form of plural nouns in English (such as "sisters'" in "My sisters' hair") will be matched by the aforementioned rules, the single quotes (') in the words "z'" and "s'" that may represent the possessive form of plural nouns are replaced with "@#" for placeholder to prevent them from being split, and then replace "@#" back with the original single quote (') after all punctuation processing is completed.
[0091] After semantic standardization through the method of this embodiment, the semantic information of tweet tags, secondary morphemes, special semantic punctuation, and word-terminal punctuation can be restored, avoiding the problem that morphemes in non-standard forms that may contain semantics are filtered out, resulting in incomplete tweet content obtained in subsequent processing.
[0092] Step S302: Standardize informal morphemes. In this embodiment, a multi-source English word list and BERT are used together to standardize informal morphemes.
[0093] This step mainly processes four types of abbreviations: common abbreviations with end punctuation (including three categories: geography, date, and physiology or medicine), Latin abbreviations, preposition abbreviations, and combinations of "person (pronoun) - modal verb or auxiliary verb". Among them, except for the "common abbreviations with end punctuation" which are constructed through the ECDICT dictionary, the remaining abbreviation types are collected and sorted through Wikipedia; in addition, all matches are completed through "free element regular expressions".
[0094] Specifically, a free element is an element in a free state that conforms to one of the following characteristics:
[0095] a) There is a space before the element and the end of the sentence after the element;
[0096] b) The element is at the beginning of the sentence and there is a space after the element;
[0097] c) There are spaces before and after the element;
[0098] d) The sentence only contains the target element.
[0099] The regular expression for free elements is: (?<=\\s)k$|^k(?=\\s)|(?<=\\s)k(?=\\s)|^k$.
[0100] Specifically, during the restoration of "preposition abbreviations", the free "2" and "4" will be restored to "to" and "for". However, in some contexts, "2" and "4" do not represent prepositional parts of speech but purely indicate quantities. To avoid the above incorrect restoration problem, in this embodiment, the BERT model is used to determine whether the free "2" and "4" in the given text need to be restored to the prepositions "to" and "for", that is, converting this problem into a binary classification problem of "whether to restore", which is abbreviated as the "24 restoration problem".
[0101] First, use the clause extraction algorithm based on a double-stack structure to split all the abstracts in Abs into sentences. At the same time, only retain the sentences containing free "2", "4", "to", or "for", and replace the "2", "4", "to", or "for" in them with the [MASK] label. Also, add the [CLS] and [SEP] labels at the beginning and end of the sentence. Denote the processing result as PData.
[0102] Based on the "BERT base model (cased)", add a fully connected layer with a binary classification SoftMax structure;
[0103] Fine-tune the pre-trained model. First, divide PData into training, validation, and test sets according to the ratio of "6:2:2"; freeze the BERT weights; input the embedding vector of [MASK] output by BERT into the fully connected layer with a binary classification SoftMax structure and then calculate the binary cross-entropy. After loss iteration, obtain the 24 restoration model.
[0104] Input the English tweet to be recognized into the trained 24 restoration model, determine and restore "2" and "4", and obtain the standard representations of "2" and "4".
[0105] Step S303: Correction of non-standard capitalized words based on word morphology;
[0106] First, extract a large-scale text from Wikipedia as the formal text dataset; at the same time, to enhance the model's generalization ability for informal texts, on the basis of the formal text dataset, also merge the IMDB movie review dataset containing a large amount of informal texts to expand the training set's processing ability for informal texts.
[0107] Further, obtain the morpheme discrete embedding; specifically, in this embodiment, referring to the fact that when humans judge the correct form of a certain word in a sentence, they mainly consider factors such as the case of the word and its adjacent words and the common part-of-speech of the word itself, this embodiment uses the following features as the embedding vectors of each morpheme in the tweet:
[0108] Dimension 1: The id of the morpheme in ECDICT (starting from 1, 0 indicates an out-of-vocabulary word)
[0109] Dimension 2: The form of the morpheme, including: only contains numbers, all lowercase, all uppercase, first letter uppercase, partial uppercase in the middle of the word, only contains end punctuation, only contains paired punctuation, only contains connecting punctuation, only contains other punctuation, various mixtures, and empty placeholders.
[0110] Dimensions 3, 4, and 5: The top three part-of-speech frequencies of the morpheme in ECDICT (0 indicates that the value is empty). When the part-of-speech of this element cannot make up 3 types, such as "apple" only has the part-of-speech of noun, then the latter two dimensions are empty.
[0111] Among them, common words refer to the first 8000 words in the contemporary corpus word frequency order provided in ECDICT or words with a Collins star rating of at least 0.
[0112] Morphemes refer to various language elements such as words (the classification criteria can be part-of-speech, word frequency, etc.), punctuation marks, and numbers. Connecting punctuation generally indicates the context connection relationship in the text, including the ampersand "&", hyphen "-", colon ":", equal sign "=", underscore "_", or sign "|", tilde "~"; other punctuation refers to other punctuation in the ASCII punctuation symbol set except "end punctuation", "paired punctuation", and "connecting punctuation".
[0113] Furthermore, since completely lowercase words are the main components in normal English texts, and most of the existing training corpora are long formal texts or semi-formal texts, if the original corpus data is directly used as the model input, there are more normal texts in the input corpus, which will cause the model to tend to classify the target as normal texts, resulting in label imbalance. Therefore, this embodiment uses a composite sampling method that combines radius sliding window downsampling and negative sampling for sampling to obtain the perceptron training data set, denoted as Sam = {s1, s2, …, s n}.
[0114] Specifically, downsampling includes: using words with an irregular capitalized word style as the core morphemes, with a radius of WR (set to 1 in this embodiment), and introducing their neighboring morphemes as training corpus. Among them, the styles of irregular capitalized words include words in all caps, capitalized first letters, and partially capitalized words in the middle. For example, "He was ORDERED to leave Russia.", with "ordered" as the center and a radius of 1, the obtained training corpus is "was ORDERED to". Among them, if there are not enough morphemes within the WR radius of the core morpheme, an empty string is used instead; in addition, if the ranges between the clauses intercepted with multiple core morphemes overlap with each other, only the head and tail two terminal ranges of all overlapping ranges are considered, and all overlapping ranges in the middle are merged. For example, in "I really LOVE to SWIM in river", "LOVE" and "SWIM" should be used as the centers, with a radius of 1, and the two obtained training corpora are "really LOVE to" and "to SWIM in". In this example, since "to" appears repeatedly in the two corpora, the two corpora are merged into one corpus: "really LOVE to SWIM in".
[0115] Negative sampling includes: constructing negative examples on each clause obtained by downsampling to enhance the model's fitting ability for negative examples, specifically including:
[0116] Directly retain the original sentence without modification; or,
[0117] Randomly convert a completely lowercase word in the clause to all caps according to a certain quantity (the default value is 30% of the word count of the current input sentence); or,
[0118] Convert all completely lowercase words to all caps or capitalized first letters, and the conversion target depends on the style of the capitalized letter-containing word adjacent to the current completely lowercase word, and the opposite or the same style is randomly selected with a probability of 50% for conversion.
[0119] Furthermore, divide the perceptron training data set into training, validation, and test sets according to the ratio of "6:2:2"; then construct a classification multi-layer perceptron and train it with binary cross-entropy as the loss function. Among them, the multi-layer perceptron includes an Embedding layer, a hidden layer, and an output layer:
[0120] Embedding layer: used to convert the discrete features of each morpheme in each clause in the perceptron training data set Sam into continuous vector features. For the training clause s i ∈Sam, for each morpheme {w1, w2,..., w n} ∈ s i perform embedding (i.e., v j = Embedding(w j ), j ∈ [1, n]) and concatenate the head and tail to obtain the current clause s i 's hidden representation (where represents the vector concatenation operation);
[0121] Hidden layer: It consists of 5 fully connected layers. Each layer (denoted as l) uses ELU as the activation function, and increases the sensitivity of neurons to negative values by gradually increasing the slope of the negative half-axis. At the same time, Dropout is added according to the dropout rate d for regularization. The formal definition of the hidden representation of each layer is:
[0122] H l+1 = Dropout(ELU(WH l + b, 0.05 * l), d)
[0123] Output layer: Used to output the binary classification probability.
[0124] After this step, the ungrammatical capitalized words in the English tweet to be recognized are converted into the form of standard all - lowercase or capitalized first letter.
[0125] Step S304: Extract the end - value information of English tweets based on one - dimensional conjugate cellular automata;
[0126] Split the tweet text to be recognized processed by steps S301 - S303 into three segments: the start tag HS, the mixture of natural text and tags MI in the middle of the sentence, and the end tag HE; among them, the first word of MI is the next non - tag text of HS, and the last word of MI is the previous non - tag text of HE. After obtaining the three segmented data, progress forward and backward from MI.
[0127] Specifically, for the one - dimensional evolution from right to left from MI to HS, denote the current cell index as HS index , and the neighbors of each cell are adjacent words in HS. Set the forward evolution rule set SR, and the state set SS = {truncated, retained}. The specific content of the forward evolution rule set SR is as follows (mutually in an "or" relationship):
[0128] (1) The last word or the last phrase of HS is a conjunction; (Use the AC automaton constructed based on conjunctions to match the longest conjunction)
[0129] (2) The last word of MI is a preposition;
[0130] (3) The first word of MI starts with a lowercase letter;
[0131] (4) Through ECDICT, we searched for the two words adjacent to HS and MI, and found that neither of them was a common word;
[0132] (5) Merge the two adjacent words HS and MI and query whether there is a corresponding compound word through ECDICT.
[0133] Among them, the connecting word refers to the set of prepositions (phrases), conjunctions (phrases), modal verbs and auxiliary verbs. The connecting word is obtained by matching the longest connecting word using the AC automaton built based on the connecting word. For example, building an AC automaton in the set {as, as well as} can ensure that when matching "Mary as well as John", "as well as" is matched instead of "as".
[0134] When the automaton evolves under the constraints of the forward evolution rule set SR, the "truncation" and "retention" states in the state set SS can be specifically expressed as follows: if the current cell does not satisfy SR, HS is truncated to HS[HS index :] and terminate the automaton. If the forward evolution rule set SR is satisfied, HS in HS index The corresponding element remains at the front position of MI.
[0135] Furthermore, in the one-dimensional evolution from MI to HE from left to right, the backward evolution rule set ER is in a conjugate relationship with the forward evolution rule set SR, and its contents are as follows (they are in an "or" relationship with each other):
[0136] (1) The last word of MI is a connecting word;
[0137] (2) The first word of HE is a preposition;
[0138] (3) The first word of HE begins with a lowercase letter and the last word of MI is not a sentence-ending punctuation;
[0139] (4) Through ECDICT, we search for the two words adjacent to HE and MI, and neither of them is a common word;
[0140] (5) Merge the two adjacent words HE and MI and query whether there is a corresponding compound word through ECDICT.
[0141] By evolving from MI to HS and HE through automata, the words that satisfy the forward evolution rule set SR and the backward evolution rule set ER are retained to obtain the sentence end value information extraction results.
[0142] Step S305: Clause extraction based on the double stack structure, including:
[0143] If the last element of the sentence (which can be a letter or punctuation mark) is not the end punctuation mark, add ". " to the end of the sentence; initialize two stack structures S for storing punctuation marks in the sentence pair and S ter respectively represent "the left paired punctuation marks that have been traversed" and "the end punctuation marks that have been traversed";
[0144] Then traverse all the elements (including words, punctuation marks, etc.) obtained by splitting the current sentence by spaces:
[0145] If the current element is an end punctuation mark, then compare the index values of the last elements in S pair and S ter to determine whether the starting position of the current clause truncation should start from the previous end punctuation mark or from after the previous paired symbol; then record the current index in S ter as the candidate starting index for the next end punctuation mark;
[0146] If the current S ter is not empty and the current element is the right paired punctuation mark of the last element in S ter , then truncate the clause from the previous paired punctuation mark to the end of the current element; then determine whether the previous end symbol is within the range of the current paired symbol pair. If so, pop the last elements of S pair and S ter , otherwise only pop the last element of S pair ;
[0147] If the current element is a left paired punctuation mark, then push the index of the current element into S pair ;
[0148] Among them, when truncating the clause, replace the clause with a random nickname to avoid repeated matching.
[0149] After the above processing, if there are still unpaired punctuation marks in the sentence, they are directly removed.
[0150] After the aforementioned preprocessing process, the standard tweet text is obtained.
[0151] Step S4: Use the syntactic dependency analysis model to extract all noun phrases and noun clauses in the standard tweet text, preprocess the noun phrases and noun clauses based on the subjective word list, and construct the set NP p ;
[0152] Specifically, first input the standard tweet text (which can include the clauses in the original tweet text and the complete tags before execution 0) into the Transformer-based syntactic dependency analysis model trained with a large corpus to analyze the part of speech of each word and its part of speech in the sentence; according to the syntactic dependency analysis results, extract the noun phrases or noun clauses Np involved in the tweet i , where \(i\in\{1,2,\ldots,n\}\), \(n\) is the number of noun phrases and nominal clauses, and the set composed of all the extracted noun phrases and nominal clauses is denoted as \(NP = \{Np_1,Np_2,\ldots,Np\)\( i \(\ldots,Np\)\( n \}\). Among them, the noun phrases and nominal clauses extracted by the syntactic dependency analysis model may include long noun phrases containing multiple named entities or nominal clauses containing multiple noun phrases.
[0153] Furthermore, based on the subjective word list, preprocess the noun phrases and nominal clauses in the set \(NP\) to construct the set \(NP\)\( p ; including: traverse all noun phrases and nominal clauses, remove the leading articles, quantifiers and stop words; and based on the subjective word list, remove the subjective words in them to obtain the preprocessed set \(NP\)\( p \).
[0154] Specifically, for the leading articles: match whether the word at the beginning of the phrase is "a", "an" or "the", if so, remove it;
[0155] For the quantifiers: remove them after matching with the following regular expression, \((?i:a|an|\d+|that|this|these|those|_NUM)[a - z]\w+\s(?i:of)\s\);
[0156] where \(_NUM\) represents possible numerical words, such as:
[0157] three|quarter|three|quarters|two|thirds|one|two|first|last|three|next|million|four|five|second|six|third|billion|hundred|thousand|seven|eight|ten|nine|dozen|fourth|twenty|fifth|thirty|fifteen|fifty|twelve|sixth|forty|seventh|eleven|eighth|zero|twentieth|ninth|nineteenth|trillion|sixteen|eighteen|fourteen|sixty|thirteen|seventeen|eighty|tenth|nineteen|seventy|eighteenth|ninety|seventeenth|sixteenth|fourteenth|twelfth|fifteenth|eleventh|thirteenth|hundredth|fiftieth|thirtieth|fortieth|sixtieth|seventieth|eightieth|ninetieth。
[0158] For subjective words: Based on the subjective word list Subjws, use the following free element regular expression to remove all subjective words existing in the phrase;
[0159] (?<=\\s)k$|^k(?=\\s)|(?<=\\s)k(?=\\s)|^k$;
[0160] Among them, k can be replaced with any piece of free text to be matched, such as "hello", "#", etc.
[0161] For stop words: Since there are cases where words are included in stop words, such as "you" and "yourself", an AC automaton is constructed based on the stop word list, and the longest matching item in the noun phrase is removed. For example, if both "you" and "yourself" are matched, then "yourself" is removed. In this embodiment, according to the characteristics of English tweets and the tasks of this embodiment, the constructed stop word list includes:
[0162] it, its, itself, i, me, my, myself, we, us, our, ours, ourselves, ye, u, you, your, ur, yours, urs, yourselves, yourself, thyself, thine, thee, thou, he, him, his, himself, she, her, herself, they, them, their, theirselves, diz, this, that, these, those, there, here, thing, other, another, others, some, something, someday, somewhere, somehow, someone, somebody, sometime, sometimes, somewhat, every, everything, everywhere, everyone, everybody, everyday, any, anything, anyone, anybody, anyway, anymore, no, nothing, none, null, nowhere, nobody, just, mere, each, many, much, more, most, who, whom, what, where, which, how, whoever, whomever, whosoever, whomsoever, yes, such, not, do, don, dont, does, doesn, doesnt, did, didn, didnt, true, false, is, are, was, were, be, have, has, had, having。
[0163] Step S5: Based on the noun phrases and nominal clauses in the set NP p construct a tree-shaped parent-child structure and perform named entity extraction to obtain the named entity recognition result of the English tweet.
[0164] Since the set NP p may contain relatively long clauses such as nominal clauses, and a noun phrase may contain multiple noun phrases at the same time, it is necessary to establish a tree-shaped parent-child relationship between the phrases on the output noun phrase and nominal clause set, and then perform hierarchical structured analysis.
[0165] Specifically, based on the inclusion relationship of noun phrases, a noun phrase and a noun clause containing at least one noun phrase are used as a parent string, and the noun phrases contained in the parent string are used as child strings to construct a tree-like parent-child structure; that is, using the set NP p All noun phrases and noun clauses in the tree form parent-child structure NP t If a noun phrase has multiple parent strings, the longest parent string is used as its parent node, and all NP t The set of all parent strings in is FP t ={fp1,fp2,…,fp n }. Save noun phrases that do not have a parent string to the named entity collection
[0166] Based on the possessive structure of each parent string, extract the core noun and save it to the named entity set Specifically, the following regular expression is used to identify NP: t The possessive structure of each parent string in the , and extract its core noun,
[0167] ([A-Z0-9]\w+\s)*[\w^('s|'z|z'|s')]*(([s|z](?='))|(?='s|'z));
[0168] That is, based on the possessive structure of the noun, start from "'", traverse to the left to the first word with the first letter capitalized after the previous "'"; for the XX YY's and XX YY'z structures, extract the content between the first word with the first letter capitalized and YY in the XXYY's or XX YY'z structure to obtain the possessive core noun of the noun; for the XX YYs' and XX YYz' structures, extract the content between the first word with the first letter capitalized and YYs or YYz in the XX YYs' or XX YYz' structure to obtain the possessive core noun of the noun, and save the obtained possessive core noun to the named entity set Where X and Y represent arbitrary words.
[0169] Furthermore, the free element regular expression is used to remove all substrings in each parent string and merge the remaining content into a new string. The set formed by all substrings is recorded as cp. In practical applications, this embodiment sorts the substrings by length and removes them in descending order of length, so as to preferentially remove substrings with large information content.
[0170] If the string Satisfy the corresponding preset conditions and convert the string Save to named entity collection Otherwise, based on the string Reuse the syntactic dependency analysis model to extract all noun phrases and nominal clauses, and save the extracted noun phrases and nominal clauses to the named entity set Finally, incorporate the substring set cp into the named entity set in it
[0171] Using the named entity set to replace the set NP p , reconstruct the tree-shaped parent-child structure for named entity extraction; preferably, the tree-shaped parent-child structure can be reconstructed multiple times for named entity extraction to obtain the named entity recognition result; among them, the named entity set obtained by the jth reconstruction of the tree-shaped parent-child structure for named entity extraction is denoted as where j is an integer greater than 1; the named entity set obtained by the (j - 1)th reconstruction of the tree-shaped parent-child structure for named entity extraction is denoted as When the set and the set has an empty difference set, output the set as the named entity recognition result of the English tweet
[0172] Among them, the preset conditions for the string include: the string has only one word except for the included substrings, and this word is a preposition or a conjunction, and the number of substrings of this parent string is 2, the absolute value of the difference in the lengths of the two substrings is less than or equal to 2 and the lengths of both substrings are less than or equal to 3; if the string meets the preset conditions, then save the string to the named entity set Preferably, if the string does not meet the corresponding preset conditions, it can be determined whether the string meets: this is a non-empty string and all morphemes in this are not all invalid information (i.e., prepositions, conjunctions, punctuation marks, subjective words, stop words, and articles). If it meets, reuse the syntactic dependency analysis model to extract noun phrases and nominal clauses, and save the extracted noun phrases and nominal clauses to the named entity set to reconstruct the tree-shaped parent-child structure and perform named entity extraction; if it does not meet, the string
[0173] In summary, the method for extracting English tweet named entities based on subjective and objective word lists provided by the embodiments of the present invention uses texts on knowledge-based websites such as Wikipedia that have been revised by multiple parties and have standardized word usage to construct two word lists, respectively representing subjective words and objective words, so as to solve the problem that a large number of subjective words in informal English tweet texts affect downstream tasks of natural language processing. And for the problem that the noun phrases extracted by the grammar dependency analysis model based on Transformer may include some unnecessary information (such as invalid conjunctions, prepositions, etc.), or there are various noun phrases and nominal clauses, the invalid information is filtered through multiple word lists, and at the same time, a tree structure is constructed on the set of identified noun phrases for recursive noun extraction, improving the accuracy of named entity recognition.
[0174] Another embodiment of the present invention provides a computer device, including at least one processor and at least one memory communicatively connected to the processor; the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for extracting English tweet named entities in the foregoing embodiments.
[0175] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.
[0176] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.< / lastfigure> < / target> < / prefigure> < / name> < / lasts> < / pres> < / ispunc> < / content>
Claims
1. An English tweet named entity extraction method based on subjective and objective word lists, characterized in that, It includes the following steps: Obtain English texts in multiple fields and construct a corpus; Perform word segmentation and word frequency statistics on the texts in the corpus, and construct a subjective word list through screening; Preprocess the English tweet to be recognized to obtain a standard tweet text; Extract all noun phrases and nominal clauses in the standard tweet text using a syntactic dependency analysis model, preprocess the noun phrases and nominal clauses based on the subjective word list, and construct a set NP p ; Based on the noun phrases and nominal clauses in the set NP p construct a tree-shaped parent-child structure and perform named entity extraction to obtain the named entity recognition result of the English tweet; The construction of the tree-shaped parent-child structure and the named entity extraction includes: Based on the inclusion relationship of the noun phrases and nominal clauses in the set NP p A tree-shaped parent-child structure is constructed with a noun phrase or nominal clause containing at least one noun phrase as the parent string, and the noun phrases included in the parent string as the child strings; Based on the possessive structure of nouns for each parent string, extract its core nouns and save them to the named entity set Remove all substrings from each parent string and merge the remaining content into a new string The set formed by all substrings is denoted as cp; If the said string meets the corresponding preset condition, save the string to the named entity set Otherwise, based on the said string re - use the syntactic dependency analysis model to extract all noun phrases and nominal clauses, and save the extracted noun phrases and nominal clauses to the named entity set Merge the substring set cp into the named entity set and reconstruct the tree - shaped parent - child structure for named entity extraction to obtain the named entity recognition result.
2. The method for extracting named entities from English tweets according to claim 1, wherein Perform named entity extraction by reconstructing the tree - shaped parent - child structure multiple times to obtain the named entity recognition result. Among them, when performing named entity extraction by reconstructing the tree - shaped parent - child structure for the j - th time, the obtained named entity set is denoted as where j is an integer greater than 1. When performing named entity extraction by reconstructing the tree - shaped parent - child structure for the (j - 1)-th time, the obtained named entity set is denoted as When the set and the set the difference between them is an empty set, then output the set as the named entity recognition result of the English tweet.
3. The method for extracting English tweet named entities according to claim 1, wherein The string The corresponding preset conditions include: the string has only one word other than the included substring, and the word is a preposition or a conjunction, and the string has 2 substrings, the absolute value of the difference in the lengths of the two substrings is not greater than 2, and the lengths of the two substrings are both not greater than 3.
4. The method for extracting English tweet named entities according to claim 1, wherein Based on the possessive structure of nouns in each parent string, extract its core noun, including: Based on the possessive structure of nouns, starting from "'”", traverse to the left until the first capitalized word after the previous "'”"; for XX YY's and XX YY'z structures, extract the content between the first capitalized word and YY in the XX YY's or XX YY'z structure to obtain the possessive core noun; for XX YYs' and XX YYz' structures, extract the content between the first capitalized word and YYs or YYz in the XX YYs' or XX YYz' structure to obtain the possessive core noun, where X and Y represent any words.
5. The method for extracting English tweet named entities according to claim 1, characterized in that, The obtaining of English texts in multiple fields and the construction of a corpus includes: Obtain named entities in multiple domains and merge them to obtain the initial entity list L init ; Obtain the initial entity list L init for the entity ID corresponding to each entity in it, to obtain an entity ID list Obtain the list of entity IDs The other entity IDs associated with the entities corresponding to each entity ID in the list of entity IDs The entity IDs corresponding to the other entities of the same category as each entity, and merge them to obtain a list of entity IDs Replace with to obtain the other entity IDs associated with the entities corresponding to each entity ID in the entity ID list, and the entity ID list the entity IDs corresponding to the other entities of the same category as each entity in and merge them to obtain the entity ID list Obtain Obtain the abstracts of the entities corresponding to all entity IDs in i∈{0,1,2}, and construct the corpus based on the obtained abstracts.
6. The method for extracting English tweet named entities according to claim 1 or 5, characterized in that, The construction of the subjective word list includes: Perform word segmentation on the abstract in the corpus and count the word frequency, and remove words with a word frequency less than the first preset threshold and a word length of 1 to obtain an objective word list Objws; Obtain nouns, adverbs, adjectives, numbers, prepositions, conjunctions, and articles in the ECDICT dictionary with a word frequency ranking less than the second preset threshold and a Collins star rating greater than 1 to obtain a word list Trws; Remove the words in the word list Objws from the word list Trws to obtain a subjective word list Subjws.
7. The method for extracting English tweet named entities according to claim 1, characterized in that Preprocess the noun phrase based on the subjective vocabulary list to construct a noun phrase set NP p ; The method includes: traversing all noun phrases, removing the leading articles, quantifiers and stop words therein; and removing the subjective words therein based on the subjective vocabulary list to obtain the preprocessed noun phrases, thereby obtaining the noun phrase set NP p .
8. The method for named entity extraction of English tweets according to claim 7, wherein For word-ending articles, quantifiers, and subjective words, remove them respectively using corresponding regular expressions; For stop words, use an AC automaton to remove them; the AC automaton is constructed through a pre-built stop word list.
9. A computer device, characterized in that, It includes at least one processor and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for named entity extraction of English tweets based on subjective and objective word lists according to any one of claims 1-8.
Citation Information
Patent Citations
Entity identification method based on Weibo emotion
CN105335352A
Relation extraction method and device based on transfer dependency relation and structure assistant
CN110119510A