A Named Entity Recognition Method and Electronic Device Based on Word Representation Features

The integration of word2vec and LSTM features with a cascaded CRF model addresses the challenges of Chinese microblog entity recognition by capturing long-range contextual information, enhancing recognition accuracy.

CN114077838BActive Publication Date: 2025-07-15NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010825717.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-17
Publication Date
2025-07-15
Estimated Expiration
2040-08-17

AI Technical Summary

Technical Problem

The existing naming entity recognition methods are difficult to formulate appropriate recognition standards in Weibo texts, and miss information and lack long-term dependence on information. Especially in Chinese corpus, there is a problem of many interfering words, and the traditional methods are poor in robustness and portability.

Method used

Combining word embeddings trained by word2vec and word representations trained by LSTM, naming entities are identified through a cascade conditional random field model.

Benefits of technology

It improves the accuracy of Weibo named entity recognition, makes full use of long-distance context information, and identifies different types of named entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077838B_ABST
    Figure CN114077838B_ABST
Patent Text Reader

Abstract

The present invention provides a named entity recognition method and an electronic device based on word representation features, including: performing word segmentation on a text to be detected to obtain basic features of each word; forming each word into a word sequence, encoding each word, and extracting word embedding features of the encoding result; generating a word vector sequence according to a set weight and a set theme of the word sequence, and extracting word representation features of the word vector sequence; inputting the basic features, the word embedding features, and the word representation features into an entity recognition model to obtain named entities in the text to be detected. The present invention adopts word embeddings trained by word2vec and word representations trained by LSTM, captures the long-term dependencies of sentences, and fully utilizes long-distance context information to recognize named entities, which has better improvement compared with traditional models and improves the accuracy of recognizing named entities in microblogs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular, to a named entity recognition method and an electronic device based on word representation features. Background Art

[0002] With the development of the Internet, social network services such as Twitter, Tencent Weibo, and Sina Weibo have gradually emerged. Users are not only viewers of information but also broadcasters of information. The Internet has transformed from an information publishing platform into an interactive communication platform. Considering the characteristics of microblog text, such as being short, easy to publish, easy to read, convenient to share, and quickly spread, a large amount of information supported by microblogs has important value.

[0003] On the microblog platform, users talk about various things, such as sports, news, products, etc. Users repost the content they want to share on the microblog, comment on the content they are interested in on the microblog, and give likes to them. Therefore, identifying named entities from a large number of microblog posts is the basis and prerequisite for realizing public opinion supervision and business intelligence.

[0004] Currently, the entity recognition methods used in traditional Chinese corpora are still used to identify named entities from microblogs. However, these methods have problems such as difficulty in formulating appropriate recognition criteria, omission, and lack of consideration of context information. The most important thing is that these methods only consider the words in the context window and do not consider the long-term dependent information in the sentence, while the recognition of microblog named entities includes attributes such as person names, location names, organization names, dates, times, compound institution names, etc. Compared with traditional text corpora, microblog text contains too many interfering words, including emojis, popular emoticons, URLs, etc. At the same time, due to the complex characteristics of Chinese sentences, the recognition of named entities in Chinese media text is more difficult than that in English.

[0005] Like most natural language processing technologies, named entity recognition methods are mainly divided into two categories: rule-based methods and statistical-based methods. Earlier named entity recognition methods mostly used the method of manually constructing a finite state machine to match patterns and strings. However, rule-based methods lack robustness and portability. For the text in each new field, rules need to be updated to maintain optimal performance, which requires a large amount of specialized knowledge and manpower, and the cost is often very high.

[0006] The statistical-based methods mainly include the Hidden Markov Model (HMM) method, decision tree method, and so on. In the evaluation of these methods, the performance of HMM is generally considered to be relatively good. The main reason is that it can better capture the characteristic phenomena and positions of named entities. Moreover, due to the efficiency of the classic Viterbi algorithm in obtaining the optimal state sequence, HMM is increasingly frequently used in this field. However, since the probability knowledge obtained by statistical-based methods always lags behind the reliability of the professional knowledge of human experts, and some knowledge acquisition requires the experience of experts, the performance of statistical-based systems is lower than that of rule-based systems.

[0007] Chinese Patent Application CN109902307A discloses a named entity recognition method, a training method and device for a named entity recognition model. However, it is completely different in using LSTM as the first network layer of the entity recognition model, with fewer features used, resulting in inaccurate named entities. Summary of the Invention

[0008] To solve the above problems, the present invention provides a named entity recognition method and an electronic device based on word representation features. By combining the basic features of each word, the word embeddings trained by word2vec and the word representations trained by LSTM, the purpose of fusing context information is achieved, and named entities are accurately and efficiently recognized.

[0009] To achieve the above object, the technical solution of the present invention is as follows:

[0010] A named entity recognition method based on word representation features, the steps of which include:

[0011] 1) Segment the text to be detected to obtain the basic features of each word;

[0012] 2) Form each word into a word sequence, and encode each word to extract the word embedding features of the encoding result;

[0013] 3) Generate a word vector sequence according to the set weights and set topics of the word sequence, and extract the word representation features of the word vector sequence;

[0014] 4) Input the basic features, word embedding features and word representation features into an entity recognition model to obtain the named entities in the text to be detected;

[0015] Among them, the entity recognition model is obtained through the following steps:

[0016] a) Collect a number of sample texts to obtain a corpus;

[0017] b) Obtain the sample basic features, sample word embedding features and sample word representation features of each sample text in the corpus;

[0018] c) Input the sample basic features, sample word embedding features, and sample word representation features of each sample text into a cascaded conditional random field model for training to obtain an entity recognition model.

[0019] Furthermore, the text to be detected includes Chinese microblogs.

[0020] Furthermore, the basic features include word features, part-of-speech features, letter features, and digit features.

[0021] Furthermore, extract the word embedding features of the encoding result through the skip-gram model of word2vec.

[0022] Furthermore, input the word vector sequence into a recurrent neural network to extract the word representation features of the word vector sequence.

[0023] Furthermore, the recurrent neural network includes a long short-term memory network.

[0024] Furthermore, the bottommost conditional random field model of the entity recognition model outputs simple named entities, and other conditional random field models output combined complex named entities.

[0025] Furthermore, the simple named entities include geographical names and personal names; the combined complex named entities include organization names and company names.

[0026] A storage medium stores a computer program, wherein the computer program is configured to execute the above-mentioned method when running.

[0027] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer to execute the above-mentioned method.

[0028] Compared with the prior art, the advantages of the present invention are as follows:

[0029] 1) The word embeddings trained by word2vec and the word representations trained by LSTM are adopted to capture the long-term dependencies of sentences, and the long-distance context information is fully utilized to identify named entities.

[0030] 2) Different features are integrated into the cascaded conditional random field to identify different named entities, which has better improvement compared with the traditional model and improves the accuracy of microblog named entity recognition. Description of the Drawings

[0031] Figure 1 Schematic structural diagram of the cascaded conditional random field model.

[0032] Figure 2Schematic diagram of the structure of the LSTM network.

[0033] Figure 3 Flowchart of the named entity recognition method according to an embodiment of the present invention. Detailed implementation manners

[0034] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the following further details the method and steps for analyzing the microblog sentiment tendency based on sentiment object recognition and sentiment rules according to the present invention in conjunction with the accompanying drawings.

[0035] The method for identifying Chinese microblog named entities based on word representation features according to the present invention will, in this part, explain the long short-term memory network (LSTM), word2vec and cascaded conditional random fields, and propose a hybrid tagging architecture that adds the features trained by LSTM and word2vec to the cascaded CRF model to improve the effect of microblog named entity recognition. The long short-term memory network (LSTM) was invented by Jürgen Schmidhuber in 1997. Word2vec is a word vector training model open-sourced by Google and has two training modes, Skip Gram and CBOW (continuous bags of words). Among them, Skip Gram predicts the context based on the target word, and CBOW predicts the target word based on the context. Finally, some parameters of the model are used as word vectors. The cascaded conditional random field model is a serial combination of two conditional random field models.

[0036] According to the first aspect of the present invention, first, the conditional random field (CRF) is a typical named entity recognition model, and CRF is superior to the maximum entropy Markov model (MEMM) and the hidden Markov model (HMM). The conditional random field was proposed by Lafferty J in 2001. This is a framework for establishing a probability model to segment and label sequence data. Named entity recognition is actually a sequence tagging problem. For an input sentence o = o1, o2,..., o n , it is regarded as an observable word sequence. For the output state sequence s = s1, s2,..., s n , it corresponds to the labels assigned to the words in the input sequence X. Each element in the sequence S corresponds to a label I, and the label I is limited to a finite set of labels of length k. The definition of the probability of S given the input sequence O is as follows:

[0037]

[0038] where t k is defined on the margin of the feature function, called the transfer characteristic, and depends on the previous position and the current position. w lDefined on the nodes of the feature function, this function is called the state feature and depends on the current position. r k and u k are the learning weights of each feature function. z(o) is the normalization factor of the state sequence.

[0039] Named entities are sometimes relatively complex. To solve the problem of named entity recognition in complex situations, a cascaded conditional random field model is adopted to identify microblog named entities. The cascaded conditional random field model is constructed by using multiple simple superimposed models and a linear combination of ways that span these simple models. The coupling degree between layers of the cascaded conditional random field is very low, and each layer can be trained and modeled separately. The cascaded conditional random field model is as Figure 1 shown. The bottom-end CRF model can identify other simple entities such as geographical names and personal names, and then pass the results to the high-level model and support the decision-making of the high-level model to identify complex combined complex named entities such as organization names and company names. The wrong labels generated by the low-end model can be adjusted and corrected to a certain extent in the high-level model, thereby improving the effect of identifying the complex structure of named entities.

[0040] According to the second aspect of the present invention, the context information contained in the word representation can, to some extent, make up for the lack of the missing microblog context information, so as to better provide natural language processing tasks for microblogs. So far, many methods have used word embeddings to improve named entity recognition systems, and it is considered that word embeddings can represent each word as a vector with multiple topics based on different weights. The formula for the word representation trained by word2vec is as follows:

[0041] word = {v i |v i = (r1, r2, r3...r k ), 0 ≤ i ≤ N}

[0042] where, represents the v i represents the word vector of the i-th word, γ i represents the weight of the k-th dimension, and the vocabulary length is N. Training word embeddings for Chinese texts is not easy. First, Chinese words need to be segmented instead of directly training Chinese characters. In this method, the skip-gram model and negative sampling of word2vec are used to pre-train the words, and the trained word embeddings will be used as new features added to the cascaded CRF model.

[0043] According to the third aspect of the present invention, a Recurrent Neural Network (RNN) is a special deep neural network architecture. Different from the Feedforward Neural Network (FNN), it contains recurrent connections between the neurons in the hidden layer, which enables the RNN to potentially have a relatively deep depth network, thereby allowing the network to effectively handle the dependencies of the input sequence. In practice, the RNN cannot learn long-term dependencies, so it faces the problem of vanishing gradients during training. The Long Short-Term Memory (LSTM) network is a special recurrent neural network structure. The network structure of the LSTM cell is as shown in Figure 2 and was invented by Jürgen Schmidhuber in 1997. The Long Short-Term Memory Neural Network (LSTM) is a type of time-recurrent neural network, suitable for processing and predicting important events with relatively long intervals and delays in time series. Three gates are placed in a cell, namely the input gate, the forget gate, and the output gate. When an information enters the LSTM network, it can be judged whether it is useful according to the algorithm, and only the information that meets the algorithm certification will be retained, while the unqualified information will be forgotten through the forget gate. It solves the problem caused by vanishing gradients by adding three specific storage units in the network and proves that it can capture long-range dependencies. The formulas are as follows:

[0044] Input gate

[0045] Forget gate

[0046] Node

[0047] Output gate

[0048] Node output where the activation function f represents the control gate, and the activation functions g and h represent the input and output of the cell respectively. w represents the weight, represents the node information at the current moment, t represents the current moment, represents the forget gate, l represents the input gate, w represents the output gate, c represents the node, I represents that the network has I input units, H hidden units, C output units, b is the hidden unit information, and s is the output unit information.

[0049] Given an input sequence x = x1, x2,..., x n containing n words, each word generates a multi-dimensional word vector by word2vec, and this model will return a sequence b = b1, b2,..., b n, each word in this sequence contains the information of the previous words of the current word in the sentence. Similar to word embeddings, the output of this model - the word representation can be added as a new feature to the cascaded CRF entity recognition model.

[0050] The flowchart of the named entity recognition method of the present invention is as Figure 3 shown. First, a named entity recognition model is trained based on the Weibo corpus, and then after the preprocessing of Weibo data, the extracted features are input into the cascaded conditional random field for model training. Secondly, the Weibo data to be extracted is preprocessed and feature extracted, and finally, it is input into the trained entity recognition model to complete the Weibo named entity recognition.

[0051] Feature selection, which directly affects the effectiveness of the model, is extremely important for the named entity recognition model. Therefore, the features for identifying Weibo named entities are selected as follows:

[0052] 1. Word feature: the current word in the corpus;

[0053] 2. Part-of-speech feature: the part of speech of the current word;

[0054] 3. Letter feature: whether the current word contains letters or not;

[0055] 4. Number feature: whether the current word contains numbers;

[0056] 5. Word embedding feature: To find the most suitable window size and vector dimension, several experiments are conducted by adjusting the window size and vector dimension to determine the best experimental parameters, and then the vector dimension is finally set to 100. Finally, the word embedding data is added as a new feature to the model;

[0057] 6. Word representation feature: The output of the word representation trained by the LSTM algorithm is added as a new feature to the model.

[0058] The cascaded random field model adopted by this method consists of two parts: the lower-layer CRF model uses the above 1-4 basic features to identify simple named entities; the upper-layer CRF model uses the above 1-4 features and new functions, including the 5th feature, the 6th feature and the simple named entity features generated by the lower-layer model to identify and extract complex named entities.

[0059] The following is the comparison of entity recognition performance after adding new features in the present invention:

[0060] Table 1 Performance effect diagram of the entity recognition model

[0061]

[0062] It should be noted and understood that various modifications and improvements can be made to the present invention described in detail above without departing from the spirit and scope of the present invention as claimed. Therefore, the scope of the claimed technical solution is not limited by any specific exemplary teachings given.

Claims

1. A named entity recognition method based on word representation features, the steps of which include: 1) Segment the text to be detected to obtain the basic features of each word; wherein, the basic features include word features, part-of-speech features, letter features, and digital features; 2) Form each word into a word sequence, and encode each word to extract the word embedding features of the encoding result; 3) Generate a word vector sequence according to the set weights and set themes of the word sequence, and input the word vector sequence into a recurrent neural network to extract the word representation features of the word vector sequence; 4) Input the basic features, word embedding features, and word representation features into an entity recognition model to obtain the named entities in the text to be detected; the entity recognition model is composed of a serial combination of two conditional random field models. The low-level conditional random field model uses word features, part-of-speech features, letter features, and digital features to identify simple named entities, and the high-level conditional random field model uses word features, part-of-speech features, letter features, digital features, word embedding features, word representation features, and the features of the simple named entities for the recognition and extraction of combined complex named entities; Wherein, the entity recognition model is obtained through the following steps: a) Collect a number of sample texts to obtain a corpus; b) Obtain the sample basic features, sample word embedding features, and sample word representation features of each sample text in the corpus; c) Input the sample basic features, sample word embedding features, and sample word representation features of each sample text into a cascaded conditional random field model for training to obtain an entity recognition model.

2. The method according to claim 1, characterized in that, The text to be detected includes Chinese microblogs.

3. The method according to claim 1, characterized in that Extract the word embedding features of the encoding result through the skip-gram model of word2vec.

4. The method according to claim 1, wherein The recurrent neural network includes a long short-term memory network.

5. The method according to claim 1, wherein Simple named entities include: geographical names and personal names; combined complex named entities include: organization names and company names.

6. A storage medium, in which a computer program is stored, wherein, The computer program is set to execute the method according to any one of claims 1-5 when running.

7. An electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is set to run the computer program to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Named entity recognition method and named entity recognition model training method and device

    CN109902307A