Entity recognition method, model training method and device

By using deep learning language model, lexical association model and bidirectional hidden state extraction model in entity recognition method, combined with the understanding of grammar, semantics and lexical levels, the problem of low accuracy of long entity recognition is solved, and more accurate long entity recognition is achieved.

CN114036935BActive Publication Date: 2025-06-06BEIJING KINGSOFT DIGITAL ENTERTAINMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111316183.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-13
Filing Date
2021-11-08
Publication Date
2025-06-06
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

The prior art has low recognition accuracy in long entity recognition, making it difficult to accurately identify location and organizational entities in long entities.

Method used

An entity recognition method is adopted. By obtaining the text to be recognized and the entity recognition model that is pre-trained, the text to be recognized is input into the language model, the lexical association model and the two-way hidden state extraction model based on deep learning, multiple hidden states are extracted and spliced, and input to the classification layer for classification recognition to improve the recognition accuracy of long entities.

Benefits of technology

By combining the understanding of grammatical, semantic and lexical levels, the recognition accuracy of long entities is improved, and the location and organizational entities in the text can be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036935B_ABST
    Figure CN114036935B_ABST
Patent Text Reader

Abstract

The present application provides an entity recognition method, a model training method and a device, wherein the entity recognition method comprises: obtaining a text to be recognized and a pre-trained entity recognition model, inputting the text to be recognized into the first sub-model and the second sub-model of the entity recognition model respectively, obtaining a first word feature vector and a second word feature vector, and then inputting the first word feature vector and the second word feature vector into the third sub-model of the entity recognition model, extracting the bidirectional hidden state of the third sub-model, obtaining a plurality of first hidden states and a plurality of second hidden states, then splicing the plurality of first hidden states and the plurality of second hidden states to obtain the spliced ​​hidden states, inputting the spliced ​​hidden states into the classification layer of the entity recognition model, and obtaining the entity recognition result of the text to be recognized through classification recognition of the classification layer. During entity recognition, both grammatical semantics and lexical grammar are considered, thereby improving the recognition accuracy of long entities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence of computer technology, and in particular to an entity recognition method. The present application also relates to an entity recognition model training method, an entity recognition device, an entity recognition model training device, a computing device, and a computer-readable storage medium. Background Art

[0002] Artificial Intelligence (AI) refers to the ability of engineered (i.e. designed and manufactured) systems to perceive the environment, as well as the ability to acquire, process, apply and represent knowledge. The development status of key technologies in the field of artificial intelligence, including machine learning, knowledge graphs, natural language processing, computer vision, human-computer interaction, biometrics, virtual reality / augmented reality and other key technologies. Natural Language Processing (NLP) refers to the use of computers to process the form, sound, meaning and other information of natural language, that is, the operation and processing of input, output, recognition, analysis, understanding and generation of characters, words, sentences and passages. Natural Language Processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language.

[0003] Entity recognition, also known as named entity recognition (NER), refers to the recognition of entities with specific meanings in text, mainly including names of people, places, institutions, proper nouns, etc. Currently, NER is the most numerous and difficult task in language analysis. At the same time, NER is also an indispensable component of various natural language processing (NLP) technologies such as information extraction, information retrieval, machine translation, and question-answering systems.

[0004] Deep learning is the process of learning the inherent laws and representation levels of sample data. The information obtained in the learning process is of great help in the interpretation of data such as text, images and sounds. Its ultimate goal is to enable machines to have analytical learning capabilities like humans and to recognize data such as text, images and sounds. In the current NER task, deep learning methods are usually used. Specifically, the text to be recognized is input into the pre-trained entity recognition model, and the entity recognition result of the text to be recognized is obtained through the operation of the entity recognition model. The entity recognition model consists of a bidirectional encoder representation (BERT) model, a long short-term memory (LSTM) model and a conditional random field (CRF) layer. The BERT model is a model structure that uses the attention mechanism to implement pre-training or re-training tasks. It has a strong understanding ability at the grammatical and semantic levels and has a high recognition accuracy for short entities.

[0005] However, since the BERT model has strong understanding capabilities only at the grammatical and semantic levels, the recognition accuracy will decrease for long entities. For example, the long entity "Stanford University, California" is easily recognized as an organization (ORG) entity, but it should actually be recognized as a location (LOC) entity "California" and an ORG entity "Stanford University". Therefore, how to improve the recognition accuracy of long entities has become a technical problem that needs to be solved urgently. Summary of the invention

[0006] In view of this, the embodiment of the present application provides an entity recognition method to solve the technical defects existing in the prior art. The embodiment of the present application also provides an entity recognition model training method, an entity recognition device, an entity recognition model training device, a computing device, and a computer-readable storage medium.

[0007] According to a first aspect of an embodiment of the present application, there is provided an entity recognition method, comprising:

[0008] Obtaining a text to be recognized and a pre-trained entity recognition model, wherein the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0009] Input the text to be recognized into the first sub-model and the second sub-model respectively to obtain the first word feature vector and the second word feature vector;

[0010] Inputting the first word feature vector and the second word feature vector into the third sub-model, and extracting the bidirectional hidden state of the third sub-model to obtain a plurality of first hidden states and a plurality of second hidden states;

[0011] Concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state;

[0012] The concatenated hidden states are input into the classification layer, and the entity recognition result of the text to be recognized is obtained through classification and recognition by the classification layer.

[0013] According to a second aspect of an embodiment of the present application, a method for training an entity recognition model is provided, comprising:

[0014] Obtaining a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0015] Extract training text from the training set, and input the training text into the first sub-model and the second sub-model respectively to obtain the first word feature vector and the second word feature vector;

[0016] Inputting the first word feature vector and the second word feature vector into the third sub-model, and extracting the bidirectional hidden state of the third sub-model to obtain a plurality of first hidden states and a plurality of second hidden states;

[0017] Concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state;

[0018] The concatenated hidden states are input into the classification layer, and the entity prediction results of the training text are obtained through classification and recognition by the classification layer;

[0019] Compare the entity prediction results with the entity annotation information carried by the training text to obtain the difference value;

[0020] If the difference value is greater than a preset threshold, the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer are adjusted, and the step of extracting training text from the training set is returned to execute until the training stop condition is reached, the training is stopped, and the network model that has completed the training is determined to be an entity recognition model.

[0021] According to a third aspect of an embodiment of the present application, there is provided an entity identification device, including:

[0022] A first acquisition module is configured to acquire a text to be recognized and a pre-trained entity recognition model, wherein the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0023] The first language analysis module is configured to input the text to be recognized into the first sub-model and the second sub-model respectively to obtain a first word feature vector and a second word feature vector;

[0024] A first hidden state extraction module is configured to input the first word feature vector and the second word feature vector into the third sub-model, and obtain a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model;

[0025] A first splicing module is configured to splice the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state;

[0026] The recognition module is configured to input the concatenated hidden state into the classification layer, and obtain the entity recognition result of the text to be recognized through classification and recognition by the classification layer.

[0027] According to a fourth aspect of an embodiment of the present application, there is provided an entity recognition model training device, comprising:

[0028] A second acquisition module is configured to acquire a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0029] The second language analysis module is configured to extract training text from the training set and input the training text into the first sub-model and the second sub-model respectively to obtain a first word feature vector and a second word feature vector;

[0030] A second hidden state extraction module is configured to input the first word feature vector and the second word feature vector into the third sub-model, and obtain a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model;

[0031] A second splicing module is configured to splice the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state;

[0032] The prediction module is configured to input the concatenated hidden states into the classification layer, and obtain the entity prediction result of the training text through classification recognition by the classification layer;

[0033] A comparison module is configured to compare the entity prediction result with the entity annotation information carried by the training text to obtain a difference value;

[0034] The adjustment module is configured to adjust the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer if the difference value is greater than a preset threshold, and return to execute the step of extracting training text from the training set until the training stop condition is reached, stop training, and determine that the network model that has completed the training is an entity recognition model.

[0035] According to a fifth aspect of an embodiment of the present application, there is provided a computing device, including: a memory and a processor;

[0036] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the method provided in the first aspect or the second aspect of the embodiment of the present application.

[0037] According to the sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores computer instructions, and when the instructions are executed by a processor, the steps of the method provided by the first aspect or the second aspect of the embodiments of the present application are implemented.

[0038] The entity recognition method provided by the present application obtains a text to be recognized and a pre-trained entity recognition model, inputs the text to be recognized into a first sub-model and a second sub-model of the entity recognition model respectively, obtains a first word feature vector and a second word feature vector, then inputs the first word feature vector and the second word feature vector into a third sub-model of the entity recognition model, extracts a bidirectional hidden state of the third sub-model, obtains a plurality of first hidden states and a plurality of second hidden states, then concatenates the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state, inputs the concatenated hidden state into a classification layer of the entity recognition model, and obtains an entity recognition result of the text to be recognized through classification and recognition by the classification layer.

[0039] The first sub-model is a language model based on deep learning, which has strong understanding ability at the grammatical and semantic levels. The second sub-model is a lexical association model, which has strong coordination ability at the lexical level. The output results of the first sub-model and the second sub-model are used as the common input of the third sub-model. After the bidirectional hidden state extraction and hidden state splicing of the third sub-model, the hidden state of the input classification layer not only has strong understanding ability at the grammatical and semantic levels, but also has strong coordination ability at the lexical level. Therefore, when performing entity recognition, entity recognition can be performed both from the grammatical and semantic level and from the lexical level, thereby improving the recognition accuracy of long entities. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flow chart of an entity recognition method provided by an embodiment of the present application;

[0041] Figure 2 is a transfer processing flow chart of the third sub-model provided in one embodiment of the present application;

[0042] Figure 3 is a processing flow chart of an entity recognition method provided by an embodiment of the present application;

[0043] Figure 4 It is a flowchart of an entity recognition model training method provided in one embodiment of the present application;

[0044] Figure 5 It is a structural schematic diagram of an entity identification device provided in one embodiment of the present application;

[0045] Figure 6 It is a structural schematic diagram of an entity recognition model training device provided by an embodiment of the present application;

[0046] Figure 7 It is a structural block diagram of a computing device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0047] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0048] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms of "a", "said" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.

[0049] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.

[0050] First, the terms involved in one or more embodiments of the present invention are explained.

[0051] NER: A technology that uses a model to infer the various words in the input text. The words here can be nouns, verbs, etc. In general, the output of NER technology is words that represent entities with specific meanings. In addition, the output of NER technology can be words in various languages ​​such as Chinese and English.

[0052] CRF: Improves the accuracy of the final predicted label through constraints. These CRF constraints can be automatically learned by the CRF layer through which the training data passes.

[0053] BERT: Uses the Attention mechanism to implement the model structure of pre-training or re-training tasks.

[0054] LSTM: Used to solve the problems of gradient explosion and gradient disappearance during long sequence training.

[0055] Word2Vec (Word to Vector) model: A model that uses a given input text to predict the context.

[0056] LSTM output gating: The output content is the state information of the current hidden layer, and the gating structure (including forget gate, input gate and output gate) controls how much of the current state is visible to the outside world.

[0057] In the present application, an entity recognition method is provided. The present application also relates to an entity recognition model training method, an entity recognition device, an entity recognition model training device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0058] Figure 1 A flow chart of an entity recognition method provided according to an embodiment of the present application is shown, which specifically includes the following steps:

[0059] Step S102: Obtain the text to be recognized and a pre-trained entity recognition model, wherein the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model.

[0060] In the embodiment of the present application, the execution subject of the entity recognition method can be a smart device such as a mobile phone, a portable computer, a personal computer, etc. that can execute the entity recognition function.

[0061] The text to be recognized is the text that needs to be recognized as an entity, which can also be called corpus. It can be a sentence, a paragraph or an article. In order to ensure the recognition accuracy, the text to be recognized is usually a sentence. The text to be recognized can be manually input by the user on the above-mentioned smart device, or it can be obtained from the database by the above-mentioned smart device. The entity recognition model is a pre-trained end-to-end deep learning model. The acquired text to be recognized is input into the entity recognition model. After the internal operation of the entity recognition model, the entity recognition result of the text to be recognized can be directly obtained. The entity recognition result includes all entities that can be recognized in the text to be recognized.

[0062] In the embodiment of the present application, the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer. The first sub-model is a language model based on deep learning, which has strong understanding ability at the grammatical and semantic levels, such as BERT model, RoBERTa model, AlBert model, etc.; the second sub-model is a lexical association model, such as Word2Vec and other models with strong emphasis and ability in lexical terms; the third sub-model is a bidirectional hidden state extraction model, such as a bidirectional LSTM and other models with bidirectional hidden state extraction functions; the classification layer can be a network layer with classification function, such as a CRF layer based on probability distribution results, etc.

[0063] In one implementation of the embodiment of the present application, the first sub-model is a language model based on deep learning that has strong comprehension capabilities at the grammatical and semantic levels, and the BERT model can be executed concurrently, while extracting the relational features of each word in the text, and can extract relational features at multiple different levels, thereby more comprehensively reflecting the text semantics, so the first sub-model is selected as the BERT model; the second sub-model is a grammatical association model, and the Word2Vec model considers the context and has fewer dimensions when performing grammatical association, so in order to ensure the effect of grammatical association and faster processing speed, the second sub-model is selected as the Word2Vec model; the third sub-model is a bidirectional hidden state extraction model, and the bidirectional LSTM model, as a typical bidirectional hidden state extraction model, has convenient sequence modeling and long-term memory functions, so the third sub-model is selected as the bidirectional LSTM model; the classification layer is a network layer with classification function, and after selecting the bidirectional LSTM model, the CRF layer is generally selected as the classification layer, the CRF layer can add some constraints to the final predicted label to ensure that the predicted label is legal, and these constraints can be automatically learned by the CRF layer during the training data process. In order to improve the classification accuracy, the classification layer is selected as the CRF layer.

[0064] Of course, based on the traditional neural network model, a Softmax layer can also be set before the CRF layer. The Softmax layer is provided with a Softmax function (also known as a normalized exponential function). Softmax is a generalization of the binary classification function sigmoid in multi-classification. The purpose is to display the results of multi-classification in the form of probability to provide probabilistic input to the CRF layer. The specific calculation process of the Softmax function is well known in the art and will not be repeated here. CRF combines the characteristics of the maximum entropy model and the hidden Markov model. It is an undirected graph model. In recent years, it has achieved good results in sequence labeling tasks such as word segmentation, part-of-speech tagging, and named entity recognition. CRF is a typical discriminant model, and its joint probability can be written in the form of a combination of several potential functions, among which the most commonly used is the linear chain conditional random field.

[0065] The entity recognition model is pre-trained, and the specific training process can be performed by the above-mentioned smart device itself, or by other computing devices with model training functions. As an implementation method of the embodiment of the present application, the above-mentioned smart device performs the training process of the entity recognition model by itself as follows:

[0066] The first step is to obtain a training set and an initial network model, wherein the training set includes multiple training texts, each training text carries entity annotation information, and the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer. The first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model.

[0067] When training an entity recognition model, it is necessary to first obtain a training set including a massive amount of training text. Generally, the training set can be obtained by receiving a massive amount of training text input manually to form a training set, or by reading a massive amount of training text from other data acquisition devices or databases to form a training set.

[0068] The training texts in the obtained training set are generally manually annotated, and the specific annotated information is the entities in the training text, that is, the training text in the training set carries entity annotation information, wherein the training text may be a sentence, a short sentence, a paragraph, an article, etc. For example, if the training text is "The chief designer of Bridge A is Zhang XX", then the entity annotation information of the training text may include "Bridge A" and "Zhang XX".

[0069] The initial network model may be composed of a manually selected first sub-model, a second sub-model, a third sub-model and a classification layer.

[0070] In the second step, training text is extracted from the training set, and the training text is input into the first sub-model and the second sub-model respectively to obtain the first word feature vector and the second word feature vector.

[0071] During training, a training text can be arbitrarily extracted from the training set and input into the network model for a training iteration.

[0072] In the third step, the first word feature vector and the second word feature vector are input into the third sub-model, and a plurality of first hidden states and a plurality of second hidden states are obtained through bidirectional hidden state extraction of the third sub-model.

[0073] In the fourth step, the plurality of first hidden states and the plurality of second hidden states are concatenated to obtain a concatenated hidden state.

[0074] The concatenation method may be to extract one first hidden state and one second hidden state each time, and then connect the two hidden states together in the order of extraction, or to extract all the first hidden states and the second hidden states, and then connect the hidden states together in the order of extraction. The first hidden state and the second hidden state may be in the form of single data or in the form of vectors, and the concatenation method may be to arrange multiple data together in sequence to form a vector, or to arrange multiple vectors together in sequence to form a matrix.

[0075] The fifth step is to input the concatenated hidden state into the classification layer, and obtain the entity prediction result of the training text through classification and recognition by the classification layer.

[0076] The sixth step is to compare the entity prediction results with the entity annotation information carried by the training text to obtain the difference value.

[0077] In the seventh step, if the difference value is greater than the preset threshold, the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer are adjusted, and the second to seventh steps are returned to execute until the training stop condition is reached, the training is stopped, and the trained network model is determined to be an entity recognition model. Among them, the training stop condition is: the difference value is less than or equal to the preset threshold, or the number of loop iterations reaches a predetermined number of times.

[0078] The training process of the entity recognition model will be described in detail in subsequent training examples and will not be described in detail here.

[0079] Step S104: input the text to be recognized into the first sub-model and the second sub-model respectively to obtain the first word feature vector and the second word feature vector.

[0080] The acquired text to be recognized is input into the first sub-model and the second sub-model respectively. The first sub-model is a language model based on deep learning, and the output is the first word feature vector based on grammar and semantics. The second sub-model is a lexical association model, and the output is the second word feature vector based on lexical structure.

[0081] Step S106: input the first word feature vector and the second word feature vector into the third sub-model, and extract the bidirectional hidden state of the third sub-model to obtain a plurality of first hidden states and a plurality of second hidden states.

[0082] After the first word feature vector and the second word feature vector are obtained by the first sub-model and the second sub-model, the first word feature vector and the second word feature vector are input into the third sub-model. The third sub-model is a bidirectional hidden state extraction model, which can extract multiple hidden states for the input word feature vector. The difference between the bidirectional hidden state extraction model and the unidirectional hidden state extraction model is that the unidirectional hidden state extraction model can only be transmitted in one direction, that is, the hidden state is extracted in sequence from front to back, while the bidirectional hidden state extraction model can be transmitted in two directions, that is, the hidden state is extracted in sequence from front to back and from back to front. For the embodiment of the present application, in order to integrate the functions of the first sub-model and the second sub-model, it is necessary to extract the hidden states of the first word feature vector and the second word feature vector output by the first sub-model and the second sub-model respectively, and then fuse them. Therefore, in the embodiment of the present application, a bidirectional hidden state extraction model is selected to extract the hidden state. For the first word feature vector, the hidden state is extracted in order from front to back, and for the second word feature vector, the hidden state is extracted in order from back to front; or, for the first word feature vector, the hidden state is extracted in order from back to front, and for the second word feature vector, the hidden state is extracted in order from front to back. In this way, multiple first hidden states and multiple second hidden states can be extracted.

[0083] In one implementation of the embodiment of the present application, the third sub-model includes multiple hidden layers. Accordingly, S106 can be specifically implemented by the following steps: inputting the first word feature vector into the third sub-model, extracting hidden states from the order of each hidden layer in the third sub-model from the front to the back, and obtaining multiple first hidden states; inputting the second word feature vector into the third sub-model, extracting hidden states from the order of each hidden layer in the third sub-model from the back to the front, and obtaining multiple second hidden states.

[0084] The steps of extracting the hidden state of the first word feature vector and the steps of extracting the hidden state of the second word feature vector can be performed synchronously or in steps, and are not specifically limited here. As a more optimal implementation method, the steps of extracting the hidden state of the two word feature vectors are performed synchronously, that is, each hidden state is extracted synchronously for the first word feature vector and the second word feature vector, that is, when extracting the hidden state, a synchronization locking operation needs to be set, that is, after the hidden state is synchronously extracted, the operation of the hidden state extracted once needs to be completed before the next hidden state is synchronously extracted.

[0085] In the embodiment of the present application, the third sub-model includes multiple hidden layers, each hidden layer is used to extract the hidden state, and is divided into a forward channel and a reverse channel in the third sub-model. Each channel includes multiple hidden layers. The hidden layers in the forward channel and the hidden layers in the reverse channel can be the same or different. The transmission method of the third sub-model is as follows: Figure 2 As shown, A 0 , A 1 , A 2 , …, A n-1 is the hidden layer of the forward channel, A 0 ', ..., A n-2 '、A n-1 ' is the hidden layer of the reverse channel, S is the feature vector of the first word, S' is the feature vector of the second word, h 0 、h 1 、h 2 ,…,h n-1 is the first hidden state, h 0 '、h 1 '、h 2 ', ..., h n-1 ' is the second hidden state. It can be seen that in the forward channel, h i With h i-1 Related; on the reverse channel, h i 'with h i+1 'related.

[0086] In one implementation of the embodiment of the present application, the first word feature vector is input into the third sub-model, and the hidden states are extracted from the hidden layers in the third sub-model in order from the front to the back to obtain the step of obtaining multiple first hidden states. Specifically, the step can be implemented by the following steps:

[0087] According to the order of the hidden layers in the third sub-model from front to back, the first word feature vector is input into the first hidden layer, and the first hidden state is obtained after calculation by the first hidden layer;

[0088] The first hidden state is input into the second hidden layer, and the second hidden state is obtained after calculation by the second hidden layer;

[0089] A weighted operation is performed on a preset number of first hidden states that have been calculated before the i-th hidden layer to obtain a weighted result, and the weighted result is input into the i-th hidden layer. After calculation of the i-th hidden layer, the i-th first hidden state is obtained, wherein i is a positive integer greater than 2 and less than or equal to n, and n is the total number of hidden layers in the third sub-model.

[0090] When actually extracting the hidden state, the first hidden layer is the first hidden state directly extracted based on the first word feature vector. The extraction method can be to extract by using function calculation, for example, F(S), where F(S) is the hidden state extraction function, and the first hidden state h is obtained. 1 ; The second hidden layer can extract the hidden state based on the output of the first hidden layer, that is, F(S), and the extraction method is also to use the function calculation method to extract, that is, F(h 1 ); In order to ensure the model's continuous understanding of the semantics of long texts, starting from the third hidden layer, the output of the previous hidden layer should not be considered alone, but the output of the previous hidden layers should be considered comprehensively, that is, the preset number of first hidden states calculated before the i-th hidden layer should be weighted to obtain a weighted result. The weighted operation here can be to directly calculate the average value or to assign different weights. In practical applications, since the hidden layer closer to the distance has a greater impact on the current hidden layer, the weight should be assigned according to the distance. The closer the distance, the greater the weight assigned. The preset number can be set according to the length of the text. The longer the text, the larger the preset number can be set. In addition, according to the number of layers of the current hidden layer, the preset number can also be dynamically adjusted. In general, the preset number is an integer greater than or equal to 2. For example, the preset number corresponding to the third hidden layer is set to 2, the preset number corresponding to the fourth hidden layer is set to 3, and so on. Taking the preset number as 2 as an example, the weight of the output of a closer hidden layer can be assigned to 0.8 and the weight of the other can be set to 0.2. Then the i-th first hidden state of the current i-th hidden layer output is: h i =F(0.8*F(h i-1 )+0.2*F(h i-2 )).

[0091] In one implementation of the embodiment of the present application, the second word feature vector is input into the third sub-model, and the hidden states are respectively extracted from the hidden layers in the third sub-model in order from back to front to obtain a plurality of second hidden states. Specifically, the steps can be implemented by the following steps:

[0092] According to the order of the hidden layers in the third sub-model from back to front, the second word feature vector is input into the nth hidden layer, and the first second hidden state is obtained through calculation of the nth hidden layer, where the nth hidden layer is the last hidden layer in the third sub-model;

[0093] Input the first second hidden state into the n-1th hidden layer, and obtain the second second hidden state after calculation by the n-1th hidden layer;

[0094] A weighted operation is performed on a preset number of second hidden states calculated after the jth hidden layer to obtain a weighted result, and the weighted result is input into the jth hidden layer. After calculation in the jth hidden layer, the n-(j-1)th second hidden state is obtained, where j is a positive integer greater than or equal to 1 and less than n-1.

[0095] When actually extracting the hidden state, the nth hidden layer is the second hidden state directly extracted based on the second word feature vector. The extraction method can be to extract by using a function calculation method, such as F(S'), where F(S') is the hidden state extraction function, and the obtained value is the nth second hidden state h 1 '; The n-1th hidden layer can extract the hidden state based on the output of the nth hidden layer, that is, F(S'), and the extraction method is also to use the function calculation method to extract, that is, F(h 1 '); In order to ensure the model's continuous understanding of the semantics of long texts, starting from the n-2 hidden layer, the output of the next hidden layer should not be considered alone, but the output of the next hidden layers should be considered comprehensively, that is, the preset number of second hidden states calculated after the jth hidden layer should be weighted to obtain a weighted result. The weighted operation here can be to directly calculate the average value or to assign different weights. In practical applications, since the closer the hidden layer is, the greater the influence on the current hidden layer, the weight should be assigned according to the distance. The closer the distance is, the greater the weight assigned. The preset number can be set according to the length of the text. The longer the text is, the larger the preset number can be set. In addition, according to the number of layers of the current hidden layer, the preset number can also be dynamically adjusted. In general, the preset number is an integer greater than or equal to 2. For example, the preset number corresponding to the n-2 hidden layer is set to 2, the preset number corresponding to the n-3 hidden layer is set to 3, and so on. Taking the preset number as 2 as an example, the weight of the output of the nearest hidden layer can be assigned as 0.8 and the weight of the other hidden layer can be set as 0.2. Then the second hidden state of the current j-th hidden layer output is: h j '=F(0.8*F(h j-1 ')+0.2*F(h j-2 ')).

[0096] Step S108: concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state.

[0097] A splicing operation needs to be performed on the extracted multiple first hidden states and multiple second hidden states. The splicing method can be to extract one first hidden state and one second hidden state each time and then connect the two hidden states together, or to extract all the first hidden states and second hidden states and then connect the hidden states together.

[0098] In one implementation of the embodiment of the present application, S108 can be specifically implemented by the following steps: according to the extraction order of the first hidden state and the second hidden state, the second hidden state extracted in the same order is spliced ​​after the first hidden state to obtain a plurality of spliced ​​hidden states.

[0099] In a more preferred implementation of the embodiment of the present application, the extraction of the first hidden state and the extraction of the second hidden state are performed synchronously, that is, the first second hidden state is extracted while the first first hidden state is extracted, and then the first second hidden state is directly connected to the first first hidden state to form the first hidden state after splicing. Similarly, the i-th second hidden state is extracted while the i-th first hidden state is extracted, and then the i-th second hidden state is directly connected to the i-th first hidden state to form the i-th hidden state after splicing.

[0100] Step S110: input the concatenated hidden state into the classification layer, and obtain the entity recognition result of the text to be recognized through classification recognition by the classification layer.

[0101] The classification layer is a classification layer based on probability distribution results. It can classify words in the text and obtain the probability of the word in the category. The probability represents the accuracy of classification recognition. The classification layer uses the CRF layer, so that the greater the probability, the more accurate the recognition result, ensuring the recognition accuracy. The entity recognition result of the text to be recognized can be obtained through the classification recognition of the classification layer.

[0102] By using the embodiment of the present application, by obtaining the text to be recognized and the pre-trained entity recognition model, the text to be recognized is input into the first sub-model and the second sub-model of the entity recognition model respectively, and the first word feature vector and the second word feature vector are obtained, and then the first word feature vector and the second word feature vector are input into the third sub-model of the entity recognition model, and the bidirectional hidden state extraction of the third sub-model is performed to obtain multiple first hidden states and multiple second hidden states, and then the multiple first hidden states and the multiple second hidden states are spliced ​​to obtain the spliced ​​hidden state, and the spliced ​​hidden state is input into the classification layer of the entity recognition model, and the entity recognition result of the text to be recognized is obtained through classification recognition of the classification layer. The first sub-model is a language model based on deep learning, which has strong understanding ability at the grammatical and semantic levels, and the second sub-model is a lexical association model, which has strong coordination ability at the lexical level. The output results of the first sub-model and the second sub-model are used as the common input of the third sub-model, and the bidirectional hidden state extraction and hidden state splicing of the third sub-model make the hidden state of the input classification layer not only have strong understanding ability at the grammatical and semantic levels, but also have strong coordination ability at the lexical level. Therefore, when performing entity recognition, entity recognition can be performed both from a grammatical and semantic perspective and from a lexical perspective, thereby improving the recognition accuracy of long entities.

[0103] For ease of understanding, the entity recognition method provided by this application is introduced below with reference to specific examples. Figure 3 A processing flow chart of an entity recognition method provided by an embodiment of the present application is shown, which specifically includes the following steps:

[0104] The first step is to input the acquired corpus into the BERT model and Word2Vec model respectively.

[0105] In the second step, the output result of the BERT model is input into the left model of the bidirectional LSTM (the forward channel in the above embodiment), and the hidden states h0', h1', h2', etc. are extracted through the hidden layers such as LSTML0, LSTML1, LSTML2 of the left model; the output result of the Word2Vec model is input into the right model of the bidirectional LSTM (the reverse channel in the above embodiment), and the hidden states h0", h1", h2", etc. are extracted through the hidden layers such as LSTMR0, LSTMR1, LSTM R2 of the right model.

[0106] Among them, h2'=F(0.8*F(h1')+0.2*F(h0')), h2″=F(0.8*F(h0″)+0.2*F(h1″)).

[0107] The third step is to perform synchronization lock when extracting hidden states, for example, when extracting h0' and h0", and to concatenate h0' and h0". After concatenating to obtain h0, the operation of the next hidden state will be continued.

[0108] The fourth step is to input the concatenated h0, h1 and h2 into the Softmax layer to calculate the normalized exponential function to obtain the normalized probability data.

[0109] In the fifth step, the normalized probability data is input into the CRF layer, which classifies the data based on the probability distribution to obtain the final entity recognition result.

[0110] In this embodiment, the output results of the BERT model and the Word2Vec model are used as the common input of the bidirectional LSTM model, while inheriting the model understanding ability of the BERT model at the grammatical and semantic levels and the lexical coordination ability of Word2vec. In addition, this embodiment dynamically optimizes the calculation of multiple computing nodes of the bidirectional LSTM model (each hidden state extracted is called a computing node), that is, considering the calculation results of the two computing nodes before the computing node, to achieve the model's continuous understanding of the semantics of long texts.

[0111] Figure 4 A flow chart of an entity recognition model training method provided in an embodiment of the present application is shown, and the method specifically includes the following steps.

[0112] Step S402, obtaining a training set and an initial network model, wherein the training set includes multiple training texts, each training text carries entity annotation information, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model.

[0113] When training an entity recognition model, it is necessary to first obtain a training set including a massive amount of training text. Generally, the training set can be obtained by receiving a massive amount of training text input manually to form a training set, or by reading a massive amount of training text from other data acquisition devices or databases to form a training set.

[0114] The training texts in the obtained training set are generally manually annotated, and the specific annotated information is the entities in the training text, that is, the training text in the training set carries entity annotation information, wherein the training text may be a sentence, a short sentence, a paragraph, an article, etc. For example, if the training text is "The chief designer of Bridge A is Zhang XX", then the entity annotation information of the training text may include "Bridge A" and "Zhang XX".

[0115] The initial network model may be composed of a manually selected first sub-model, a second sub-model, a third sub-model and a classification layer. The first sub-model is a language model based on deep learning, which has strong comprehension capabilities at the grammatical and semantic levels, such as the BERT model, the RoBERTa model, the AlBert model, etc. The second sub-model is a lexical association model, such as Word2Vec and other models with strong lexical emphasis and capabilities; the third sub-model is a bidirectional hidden state extraction model, such as a bidirectional LSTM and other models with bidirectional hidden state extraction functions; the classification layer may be a network layer with classification functions, such as a CRF layer based on probability distribution results, etc.

[0116] Step S404: extract training text from the training set, and input the training text into the first sub-model and the second sub-model respectively to obtain a first word feature vector and a second word feature vector.

[0117] During training, a training text can be arbitrarily extracted from the training set and input into the network model for a training iteration. The extracted training text is input into the first sub-model and the second sub-model respectively. The first sub-model is a language model based on deep learning, and the output is the first word feature vector based on grammar and semantics. The second sub-model is a lexical association model, and the output is the second word feature vector based on lexical.

[0118] Step S406: input the first word feature vector and the second word feature vector into the third sub-model, and extract the bidirectional hidden state of the third sub-model to obtain a plurality of first hidden states and a plurality of second hidden states.

[0119] After the first word feature vector and the second word feature vector are obtained by the first sub-model and the second sub-model, the first word feature vector and the second word feature vector are input into the third sub-model. The third sub-model is a bidirectional hidden state extraction model, and multiple hidden states can be extracted for the input word feature vector. For the embodiment of the present application, in order to integrate the functions of the first sub-model and the second sub-model, it is necessary to extract the hidden states of the first word feature vector and the second word feature vector output by the first sub-model and the second sub-model respectively, and then fuse them. Therefore, in the embodiment of the present application, a bidirectional hidden state extraction model is selected to extract the hidden state. For the first word feature vector, the hidden state is extracted in sequence from front to back, and for the second word feature vector, the hidden state is extracted in sequence from back to front; or, for the first word feature vector, the hidden state is extracted in sequence from back to front, and for the second word feature vector, the hidden state is extracted in sequence from front to back. In this way, multiple first hidden states and multiple second hidden states can be extracted.

[0120] In one implementation of the embodiment of the present application, the third sub-model includes multiple hidden layers. Accordingly, S406 can be specifically implemented by the following steps: inputting the first word feature vector into the third sub-model, extracting hidden states from the order of each hidden layer in the third sub-model from the front to the back, and obtaining multiple first hidden states; inputting the second word feature vector into the third sub-model, extracting hidden states from the order of each hidden layer in the third sub-model from the back to the front, and obtaining multiple second hidden states.

[0121] The steps of extracting the hidden state of the first word feature vector and the steps of extracting the hidden state of the second word feature vector can be performed synchronously or in steps, and are not specifically limited here. As a more optimal implementation method, the steps of extracting the hidden state of the two word feature vectors are performed synchronously, that is, each hidden state is extracted synchronously for the first word feature vector and the second word feature vector, that is, when extracting the hidden state, a synchronization locking operation needs to be set, that is, after the hidden state is synchronously extracted, the operation of the hidden state extracted once needs to be completed before the next hidden state is synchronously extracted.

[0122] In one implementation of the embodiment of the present application, the first word feature vector is input into the third sub-model, and the hidden states are extracted from the hidden layers in the third sub-model in order from the front to the back to obtain the step of obtaining multiple first hidden states. Specifically, the step can be implemented by the following steps:

[0123] According to the order of the hidden layers in the third sub-model from front to back, the first word feature vector is input into the first hidden layer, and the first hidden state is obtained after calculation by the first hidden layer;

[0124] The first hidden state is input into the second hidden layer, and the second hidden state is obtained after calculation by the second hidden layer;

[0125] A weighted operation is performed on a preset number of first hidden states that have been calculated before the i-th hidden layer to obtain a weighted result, and the weighted result is input into the i-th hidden layer. After calculation of the i-th hidden layer, the i-th first hidden state is obtained, wherein i is a positive integer greater than 2 and less than or equal to n, and n is the total number of hidden layers in the third sub-model.

[0126] When actually extracting the hidden state, the first hidden layer is the first hidden state directly extracted based on the first word feature vector. The extraction method can be to extract by using function calculation, for example, F(S), where F(S) is the hidden state extraction function, and the first hidden state h is obtained. 1 ; The second hidden layer can extract the hidden state based on the output of the first hidden layer, that is, F(S), and the extraction method is also to use the function calculation method to extract, that is, F(h1 ); In order to ensure the model's continuous understanding of the semantics of long texts, starting from the third hidden layer, the output of the previous hidden layer should not be considered alone, but the output of the previous hidden layers should be considered comprehensively, that is, the preset number of first hidden states calculated before the i-th hidden layer should be weighted to obtain a weighted result. The weighted operation here can be to directly calculate the average value or to assign different weights. In practical applications, since the hidden layer closer to the distance has a greater impact on the current hidden layer, the weight should be assigned according to the distance. The closer the distance, the greater the weight assigned. The preset number can be set according to the length of the text. The longer the text, the larger the preset number can be set. In addition, according to the number of layers of the current hidden layer, the preset number can also be dynamically adjusted. In general, the preset number is an integer greater than or equal to 2. For example, the preset number corresponding to the third hidden layer is set to 2, the preset number corresponding to the fourth hidden layer is set to 3, and so on. Taking the preset number as 2 as an example, the weight of the output of a closer hidden layer can be assigned to 0.8 and the weight of the other can be set to 0.2. Then the i-th first hidden state of the current i-th hidden layer output is: h i =F(0.8*F(h i-1 )+0.2*F(h i-2 )).

[0127] In one implementation of the embodiment of the present application, the second word feature vector is input into the third sub-model, and the hidden states are respectively extracted from the hidden layers in the third sub-model in order from back to front to obtain a plurality of second hidden states. Specifically, the steps can be implemented by the following steps:

[0128] According to the order of the hidden layers in the third sub-model from back to front, the second word feature vector is input into the nth hidden layer, and the first second hidden state is obtained through calculation of the nth hidden layer, where the nth hidden layer is the last hidden layer in the third sub-model;

[0129] Input the first second hidden state into the n-1th hidden layer, and obtain the second second hidden state after calculation by the n-1th hidden layer;

[0130] A weighted operation is performed on a preset number of second hidden states calculated after the jth hidden layer to obtain a weighted result, and the weighted result is input into the jth hidden layer. After calculation in the jth hidden layer, the n-(j-1)th second hidden state is obtained, where j is a positive integer greater than or equal to 1 and less than n-1.

[0131] When actually extracting the hidden state, the nth hidden layer is the second hidden state directly extracted based on the second word feature vector. The extraction method can be to extract by using a function calculation method, such as F(S'), where F(S') is the hidden state extraction function, and the obtained value is the nth second hidden state h 1 '; The n-1th hidden layer can extract the hidden state based on the output of the nth hidden layer, that is, F(S'), and the extraction method is also to use the function calculation method to extract, that is, F(h 1 '); In order to ensure the model's continuous understanding of the semantics of long texts, starting from the n-2 hidden layer, the output of the next hidden layer should not be considered alone, but the output of the next hidden layers should be considered comprehensively, that is, the preset number of second hidden states calculated after the jth hidden layer should be weighted to obtain a weighted result. The weighted operation here can be to directly calculate the average value or to assign different weights. In practical applications, since the closer the hidden layer is, the greater the influence on the current hidden layer, the weight should be assigned according to the distance. The closer the distance is, the greater the weight assigned. The preset number can be set according to the length of the text. The longer the text is, the larger the preset number can be set. In addition, according to the number of layers of the current hidden layer, the preset number can also be dynamically adjusted. In general, the preset number is an integer greater than or equal to 2. For example, the preset number corresponding to the n-2 hidden layer is set to 2, the preset number corresponding to the n-3 hidden layer is set to 3, and so on. Taking the preset number as 2 as an example, the weight of the output of the nearest hidden layer can be assigned as 0.8 and the weight of the other hidden layer can be set as 0.2. Then the second hidden state of the current j-th hidden layer output is: h j '=F(0.8*F(h j-1 ')+0.2*F(h j-2 ')).

[0132] Step S408, concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state.

[0133] A splicing operation needs to be performed on the extracted multiple first hidden states and multiple second hidden states. The splicing method can be to extract one first hidden state and one second hidden state each time and then connect the two hidden states together, or to extract all the first hidden states and second hidden states and then connect the hidden states together.

[0134] In one implementation of the embodiment of the present application, S408 can be specifically implemented by the following steps: according to the extraction order of the first hidden state and the second hidden state, the second hidden state extracted in the same order is spliced ​​after the first hidden state to obtain a plurality of spliced ​​hidden states.

[0135] In a more preferred implementation of the embodiment of the present application, the extraction of the first hidden state and the extraction of the second hidden state are performed synchronously, that is, the first second hidden state is extracted while the first first hidden state is extracted, and then the first second hidden state is directly connected to the first first hidden state to form the first hidden state after splicing. Similarly, the i-th second hidden state is extracted while the i-th first hidden state is extracted, and then the i-th second hidden state is directly connected to the i-th first hidden state to form the i-th hidden state after splicing.

[0136] Step S410, input the concatenated hidden state into the classification layer, and obtain the entity prediction result of the training text through classification and recognition by the classification layer.

[0137] The classification layer is a classification layer based on probability distribution results. It can classify words in the text and obtain the probability of the word in the category. The probability represents the accuracy of classification recognition. The classification layer uses the CRF layer, so that the greater the probability, the more accurate the recognition result, ensuring the recognition accuracy. The entity prediction results of the training text can be obtained through the classification recognition of the classification layer.

[0138] Step S412, comparing the entity prediction result with the entity annotation information carried by the training text to obtain a difference value.

[0139] After obtaining the entity prediction result, compare the entity prediction result with the entity annotation information carried by the training text to obtain the difference value between the two.

[0140] Step S414, determine whether the difference value is less than or equal to a preset threshold, or whether the number of loop iterations reaches a predetermined number. If not, execute step S416, and if so, execute step S418.

[0141] Step S416, adjust the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer, and return to execute step S404.

[0142] Step S418, stop training, and determine that the trained network model is an entity recognition model.

[0143] The model parameters can be adjusted based on the obtained difference value. The goal of the adjustment is to make the difference value less than or equal to the preset threshold, or the number of loop iterations reaches a predetermined number. Through multiple loop iteration training, the recognition accuracy of the entity recognition model can be guaranteed.

[0144] Applying the embodiment of the present application, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer. The first sub-model is a language model based on deep learning, which has strong comprehension ability at the grammatical and semantic levels. The second sub-model is a lexical association model, which has strong lexical coordination ability. The output results of the first sub-model and the second sub-model are used as the common input of the third sub-model. After the bidirectional hidden state extraction and hidden state splicing of the third sub-model, the hidden state of the input classification layer not only has strong comprehension ability at the grammatical and semantic levels, but also has strong lexical coordination ability. After cyclic iterative training, the obtained entity recognition model has a certain entity recognition accuracy, and the trained entity recognition model can perform entity recognition from both the grammatical and semantic aspects and the lexical aspects, thereby improving the recognition accuracy of long entities.

[0145] Corresponding to the above-mentioned entity recognition method embodiment, Figure 5 A schematic diagram of the structure of an entity recognition device provided in an embodiment of the present application is shown, and the entity recognition device includes:

[0146] A first acquisition module 510 is configured to acquire a text to be recognized and a pre-trained entity recognition model, wherein the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0147] The first language analysis module 520 is configured to input the text to be recognized into the first sub-model and the second sub-model respectively to obtain a first word feature vector and a second word feature vector;

[0148] The first hidden state extraction module 530 is configured to input the first word feature vector and the second word feature vector into the third sub-model, and obtain a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model;

[0149] A first splicing module 540 is configured to splice the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state;

[0150] The recognition module 550 is configured to input the concatenated hidden states into the classification layer, and obtain the entity recognition result of the text to be recognized through classification and recognition by the classification layer.

[0151] Optionally, the device further includes a training module configured to:

[0152] Obtaining a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0153] Extract training text from the training set, and input the training text into the first sub-model and the second sub-model respectively to obtain the first word feature vector and the second word feature vector;

[0154] Inputting the first word feature vector and the second word feature vector into the third sub-model, and extracting the bidirectional hidden state of the third sub-model to obtain a plurality of first hidden states and a plurality of second hidden states;

[0155] Concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state;

[0156] The concatenated hidden states are input into the classification layer, and the entity prediction results of the training text are obtained through classification and recognition by the classification layer;

[0157] Compare the entity prediction results with the entity annotation information carried by the training text to obtain the difference value;

[0158] If the difference value is greater than a preset threshold, the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer are adjusted, and the step of extracting training text from the training set is returned to execute until the training stop condition is reached, the training is stopped, and the network model that has completed the training is determined to be an entity recognition model.

[0159] Optionally, the third sub-model includes multiple hidden layers; the first hidden state extraction module 530 is further configured to: input the first word feature vector into the third sub-model, extract the hidden states through each hidden layer in the third sub-model from the front to the back, and obtain multiple first hidden states; input the second word feature vector into the third sub-model, extract the hidden states through each hidden layer in the third sub-model from the back to the front, and obtain multiple second hidden states.

[0160] Optionally, the first hidden state extraction module 530 is further configured as follows: in the order of the hidden layers in the third sub-model from front to back, the first word feature vector is input into the first hidden layer, and the first first hidden state is obtained through calculation by the first hidden layer; the first first hidden state is input into the second hidden layer, and the second first hidden state is obtained through calculation by the second hidden layer; a weighted operation is performed on a preset number of first hidden states calculated before the i-th hidden layer to obtain a weighted result, and the weighted result is input into the i-th hidden layer, and the i-th first hidden state is obtained through calculation by the i-th hidden layer, wherein i is a positive integer greater than 2 and less than or equal to n, and n is the total number of hidden layers in the third sub-model.

[0161] Optionally, the first hidden state extraction module 530 is further configured as follows: according to the order of the hidden layers in the third sub-model from back to front, the second word feature vector is input into the nth hidden layer, and after calculation of the nth hidden layer, the first second hidden state is obtained, wherein the nth hidden layer is the last hidden layer in the third sub-model; the first second hidden state is input into the n-1th hidden layer, and after calculation of the n-1th hidden layer, the second second hidden state is obtained; a weighted operation is performed on a preset number of second hidden states calculated after the jth hidden layer to obtain a weighted result, and the weighted result is input into the jth hidden layer, and after calculation of the jth hidden layer, the n-(j-1)th second hidden state is obtained, wherein j is a positive integer greater than or equal to 1 and less than n-1.

[0162] Optionally, the first concatenation module 540 is further configured to: according to the extraction order of the first hidden state and the second hidden state, concatenate the second hidden state extracted in the same order after the first hidden state to obtain a plurality of concatenated hidden states.

[0163] Optionally, the first sub-model is a BERT model, the second sub-model is a Word2Vec model, the third sub-model is a bidirectional LSTM model, and the classification layer is a CRF layer.

[0164] By using the embodiment of the present application, by obtaining the text to be recognized and the pre-trained entity recognition model, the text to be recognized is input into the first sub-model and the second sub-model of the entity recognition model respectively, and the first word feature vector and the second word feature vector are obtained, and then the first word feature vector and the second word feature vector are input into the third sub-model of the entity recognition model, and the bidirectional hidden state extraction of the third sub-model is performed to obtain multiple first hidden states and multiple second hidden states, and then the multiple first hidden states and the multiple second hidden states are spliced ​​to obtain the spliced ​​hidden state, and the spliced ​​hidden state is input into the classification layer of the entity recognition model, and the entity recognition result of the text to be recognized is obtained through classification recognition of the classification layer. The first sub-model is a language model based on deep learning, which has strong understanding ability at the grammatical and semantic levels, and the second sub-model is a lexical association model, which has strong coordination ability at the lexical level. The output results of the first sub-model and the second sub-model are used as the common input of the third sub-model, and the bidirectional hidden state extraction and hidden state splicing of the third sub-model make the hidden state of the input classification layer not only have strong understanding ability at the grammatical and semantic levels, but also have strong coordination ability at the lexical level. Therefore, when performing entity recognition, entity recognition can be performed both from a grammatical and semantic perspective and from a lexical perspective, thereby improving the recognition accuracy of long entities.

[0165] The above is a schematic scheme of an entity recognition device of this embodiment. It should be noted that the technical scheme of the entity recognition device and the technical scheme of the above-mentioned entity recognition method belong to the same concept, and the details not described in detail in the technical scheme of the entity recognition device can be referred to the description of the technical scheme of the above-mentioned entity recognition method.

[0166] Corresponding to the above-mentioned entity recognition model training method embodiment, Figure 6 A schematic diagram of the structure of an entity recognition model training device provided in an embodiment of the present application is shown, and the entity recognition model training device includes:

[0167] The second acquisition module 610 is configured to acquire a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model;

[0168] The second language analysis module 620 is configured to extract training text from the training set and input the training text into the first sub-model and the second sub-model respectively to obtain a first word feature vector and a second word feature vector;

[0169] The second hidden state extraction module 630 is configured to input the first word feature vector and the second word feature vector into the third sub-model, and obtain a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model;

[0170] A second splicing module 640 is configured to splice the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state;

[0171] The prediction module 650 is configured to input the concatenated hidden state into the classification layer, and obtain the entity prediction result of the training text through classification recognition by the classification layer;

[0172] A comparison module 660 is configured to compare the entity prediction result with the entity annotation information carried by the training text to obtain a difference value;

[0173] The adjustment module 670 is configured to adjust the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer if the difference value is greater than a preset threshold, and return to execute the step of extracting training text from the training set until the training stop condition is reached, stop training, and determine that the network model that has completed the training is an entity recognition model.

[0174] Optionally, the third sub-model includes multiple hidden layers; the second hidden state extraction module 630 is further configured to: input the first word feature vector into the third sub-model, extract the hidden states through each hidden layer in the third sub-model from the front to the back, and obtain multiple first hidden states; input the second word feature vector into the third sub-model, extract the hidden states through each hidden layer in the third sub-model from the back to the front, and obtain multiple second hidden states.

[0175] Optionally, the second hidden state extraction module 630 is further configured as follows: in the order of the hidden layers in the third sub-model from front to back, the first word feature vector is input into the first hidden layer, and the first first hidden state is obtained through calculation by the first hidden layer; the first first hidden state is input into the second hidden layer, and the second first hidden state is obtained through calculation by the second hidden layer; a weighted operation is performed on a preset number of first hidden states calculated before the i-th hidden layer to obtain a weighted result, and the weighted result is input into the i-th hidden layer, and the i-th first hidden state is obtained through calculation by the i-th hidden layer, wherein i is a positive integer greater than 2 and less than or equal to n, and n is the total number of hidden layers in the third sub-model.

[0176] Optionally, the second hidden state extraction module 630 is further configured as follows: according to the order of the hidden layers in the third sub-model from back to front, the second word feature vector is input into the nth hidden layer, and after calculation of the nth hidden layer, the first second hidden state is obtained, wherein the nth hidden layer is the last hidden layer in the third sub-model; the first second hidden state is input into the n-1th hidden layer, and after calculation of the n-1th hidden layer, the second second hidden state is obtained; a weighted operation is performed on a preset number of second hidden states calculated after the jth hidden layer to obtain a weighted result, and the weighted result is input into the jth hidden layer, and after calculation of the jth hidden layer, the n-(j-1)th second hidden state is obtained, wherein j is a positive integer greater than or equal to 1 and less than n-1.

[0177] Applying the embodiment of the present application, the network model includes a first sub-model, a second sub-model, a third sub-model and a classification layer. The first sub-model is a language model based on deep learning, which has strong comprehension ability at the grammatical and semantic levels. The second sub-model is a lexical association model, which has strong lexical coordination ability. The output results of the first sub-model and the second sub-model are used as the common input of the third sub-model. After the bidirectional hidden state extraction and hidden state splicing of the third sub-model, the hidden state of the input classification layer not only has strong comprehension ability at the grammatical and semantic levels, but also has strong lexical coordination ability. After cyclic iterative training, the obtained entity recognition model has a certain entity recognition accuracy, and the trained entity recognition model can perform entity recognition from both the grammatical and semantic aspects and the lexical aspects, thereby improving the recognition accuracy of long entities.

[0178] The above is a schematic scheme of an entity recognition model training device of this embodiment. It should be noted that the technical scheme of the entity recognition model training device and the technical scheme of the above-mentioned entity recognition model training method belong to the same concept, and the details not described in detail in the technical scheme of the entity recognition model training device can be referred to the description of the technical scheme of the above-mentioned entity recognition model training method.

[0179] It should be noted that the components in the device should be understood as the functional modules that must be established to implement the steps of the program flow or the steps of the method, and the functional modules are not actually divided or separated. The device defined by such a group of functional modules should be understood as the functional module framework that mainly implements the solution through the computer program recorded in the specification, and should not be understood as a physical device that mainly implements the solution through hardware.

[0180] Figure 7The structure block diagram of a computing device 700 provided according to an embodiment of the present application is shown. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and the database 750 is used to store data.

[0181] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world microwave interconnection access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0182] In one embodiment of the present application, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 7 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.

[0183] The computing device 700 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. The computing device 700 may also be a mobile or stationary server.

[0184] Among them, the processor 720 is used to execute the following computer-executable instructions, and when the processor 720 executes the computer-executable instructions, the steps of the above-mentioned entity recognition method or entity recognition model training method are implemented.

[0185] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned entity recognition method and entity recognition model training method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned entity recognition method and entity recognition model training method.

[0186] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the steps of the entity recognition method or entity recognition model training method as described above.

[0187] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned entity recognition method and entity recognition model training method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned entity recognition method and entity recognition model training method.

[0188] An embodiment of the present application discloses a chip storing computer instructions, which, when executed by a processor, implement the steps of the entity recognition method or entity recognition model training method as described above.

[0189] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0190] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0191] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0192] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0193] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the present application. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can understand and use the present application well. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A method for entity recognition, It is characterized in that include: Obtaining a text to be recognized and a pre-trained entity recognition model, wherein the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, wherein the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model; Inputting the text to be recognized into the first sub-model and the second sub-model respectively, obtaining a first word feature vector based on grammar and semantics and a second word feature vector based on morphology; Inputting the first word feature vector and the second word feature vector into the third sub-model, extracting the hidden state of the third sub-model in a bidirectional manner, and obtaining a plurality of first hidden states and a plurality of second hidden states, wherein the third sub-model includes a plurality of hidden layers; inputting the first word feature vector and the second word feature vector into the third sub-model, extracting the hidden state of the third sub-model in a bidirectional manner, and obtaining a plurality of first hidden states and a plurality of second hidden states, comprises: inputting the first word feature vector into the third sub-model, extracting the hidden state of each hidden layer in the third sub-model from the front to the back, and obtaining the plurality of first hidden states; inputting the second word feature vector into the third sub-model, extracting the hidden state of each hidden layer in the third sub-model from the back to the front, and obtaining the plurality of second hidden states; concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state; The concatenated hidden state is input into the classification layer, and the entity recognition result of the text to be recognized is obtained through classification and recognition by the classification layer.

2. The entity recognition method according to claim 1, It is characterized in that Before the step of obtaining the text to be recognized and the pre-trained entity recognition model, the method further includes: Obtaining a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, and the network model includes a first sub-model, a second sub-model, a third sub-model, and a classification layer, wherein the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model; Extract training text from the training set, and input the training text into the first sub-model and the second sub-model respectively, to obtain a first word feature vector based on grammar and semantics and a second word feature vector based on morphology; Inputting the first word feature vector and the second word feature vector into the third sub-model, extracting the hidden state of the third sub-model in a bidirectional manner, and obtaining a plurality of first hidden states and a plurality of second hidden states, wherein the third sub-model includes a plurality of hidden layers; inputting the first word feature vector and the second word feature vector into the third sub-model, extracting the hidden state of the third sub-model in a bidirectional manner, and obtaining a plurality of first hidden states and a plurality of second hidden states, comprises: inputting the first word feature vector into the third sub-model, extracting the hidden state of each hidden layer in the third sub-model from the front to the back, and obtaining the plurality of first hidden states; inputting the second word feature vector into the third sub-model, extracting the hidden state of each hidden layer in the third sub-model from the back to the front, and obtaining the plurality of second hidden states; concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state; The concatenated hidden state is input into the classification layer, and the entity prediction result of the training text is obtained through classification and recognition by the classification layer; Comparing the entity prediction result with the entity annotation information carried by the training text to obtain a difference value; If the difference value is greater than a preset threshold, the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer are adjusted, and the step of extracting training text from the training set is returned to be executed until the training stop condition is reached, the training is stopped, and it is determined that the network model that has completed the training is an entity recognition model.

3. The entity recognition method according to claim 1 or 2, It is characterized in that The step of inputting the first word feature vector into the third sub-model, extracting hidden states from the front to the back order of each hidden layer in the third sub-model, and obtaining a plurality of first hidden states includes: According to the order of the hidden layers in the third sub-model from front to back, the first word feature vector is input into the first hidden layer, and the first hidden state is obtained through calculation by the first hidden layer; Inputting the first hidden state into the second hidden layer, and obtaining the second first hidden state through calculation by the second hidden layer; A weighted operation is performed on a preset number of first hidden states that have been calculated before the i-th hidden layer to obtain a weighted result, and the weighted result is input into the i-th hidden layer. After calculation of the i-th hidden layer, the i-th first hidden state is obtained, wherein i is a positive integer greater than 2 and less than or equal to n, and n is the total number of hidden layers in the third sub-model.

4. The entity recognition method according to claim 1 or 2, It is characterized in that The step of inputting the second word feature vector into the third sub-model, extracting hidden states from the back to the front of each hidden layer in the third sub-model, and obtaining a plurality of second hidden states includes: According to the order of the hidden layers in the third sub-model from back to front, the second word feature vector is input into the nth hidden layer, and the first second hidden state is obtained through calculation of the nth hidden layer, wherein the nth hidden layer is the last hidden layer in the third sub-model; Inputting the first second hidden state into the n-1th hidden layer, and obtaining the second second hidden state through calculation by the n-1th hidden layer; A weighted operation is performed on a preset number of second hidden states calculated after the jth hidden layer to obtain a weighted result, and the weighted result is input into the jth hidden layer. After calculation of the jth hidden layer, an n-(j-1)th second hidden state is obtained, where j is a positive integer greater than or equal to 1 and less than n-1.

5. The entity recognition method according to claim 1, It is characterized in that The step of splicing the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state includes: According to the extraction order of the first hidden state and the second hidden state, the second hidden state extracted in the same order is spliced ​​after the first hidden state to obtain a plurality of spliced ​​hidden states.

6. The entity recognition method according to claim 1, It is characterized in that The first sub-model is a BERT model, the second sub-model is a Word2Vec model, the third sub-model is a bidirectional LSTM model, and the classification layer is a CRF layer.

7. A method for training an entity recognition model, It is characterized in that include: Obtaining a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, and the network model includes a first sub-model, a second sub-model, a third sub-model, and a classification layer, wherein the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model; Extract training text from the training set, and input the training text into the first sub-model and the second sub-model respectively, to obtain a first word feature vector based on grammar and semantics and a second word feature vector based on morphology; Inputting the first word feature vector and the second word feature vector into the third sub-model, extracting the hidden state of the third sub-model in a bidirectional manner, and obtaining a plurality of first hidden states and a plurality of second hidden states, wherein the third sub-model includes a plurality of hidden layers; inputting the first word feature vector and the second word feature vector into the third sub-model, extracting the hidden state of the third sub-model in a bidirectional manner, and obtaining a plurality of first hidden states and a plurality of second hidden states, comprises: inputting the first word feature vector into the third sub-model, extracting the hidden state of each hidden layer in the third sub-model from the front to the back, and obtaining the plurality of first hidden states; inputting the second word feature vector into the third sub-model, extracting the hidden state of each hidden layer in the third sub-model from the back to the front, and obtaining the plurality of second hidden states; concatenating the plurality of first hidden states and the plurality of second hidden states to obtain a concatenated hidden state; The concatenated hidden state is input into the classification layer, and the entity prediction result of the training text is obtained through classification and recognition by the classification layer; Comparing the entity prediction result with the entity annotation information carried by the training text to obtain a difference value; If the difference value is greater than a preset threshold, the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer are adjusted, and the step of extracting training text from the training set is returned to be executed until the training stop condition is reached, the training is stopped, and it is determined that the network model that has completed the training is an entity recognition model.

8. The entity recognition model training method according to claim 7, It is characterized in that The step of inputting the first word feature vector into the third sub-model, extracting hidden states from the front to the back order of each hidden layer in the third sub-model, and obtaining a plurality of first hidden states includes: According to the order of the hidden layers in the third sub-model from front to back, the first word feature vector is input into the first hidden layer, and the first hidden state is obtained through calculation by the first hidden layer; Inputting the first hidden state into the second hidden layer, and obtaining the second first hidden state through calculation by the second hidden layer; A weighted operation is performed on a preset number of first hidden states that have been calculated before the i-th hidden layer to obtain a weighted result, and the weighted result is input into the i-th hidden layer. After calculation of the i-th hidden layer, the i-th first hidden state is obtained, wherein i is a positive integer greater than 2 and less than or equal to n, and n is the total number of hidden layers in the third sub-model.

9. The entity recognition model training method according to claim 7, It is characterized in that The step of inputting the second word feature vector into the third sub-model, extracting hidden states from the back to the front of each hidden layer in the third sub-model, and obtaining a plurality of second hidden states includes: According to the order of the hidden layers in the third sub-model from back to front, the second word feature vector is input into the nth hidden layer, and the first second hidden state is obtained through calculation of the nth hidden layer, wherein the nth hidden layer is the last hidden layer in the third sub-model; Inputting the first second hidden state into the n-1th hidden layer, and obtaining the second second hidden state through calculation by the n-1th hidden layer; A weighted operation is performed on a preset number of second hidden states calculated after the jth hidden layer to obtain a weighted result, and the weighted result is input into the jth hidden layer. After calculation of the jth hidden layer, an n-(j-1)th second hidden state is obtained, where j is a positive integer greater than or equal to 1 and less than n-1.

10. An entity recognition device, It is characterized in that include: A first acquisition module is configured to acquire a text to be recognized and a pre-trained entity recognition model, wherein the entity recognition model includes a first sub-model, a second sub-model, a third sub-model and a classification layer, wherein the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model; A first language analysis module is configured to input the text to be recognized into the first sub-model and the second sub-model respectively, to obtain a first word feature vector based on grammar and semantics and a second word feature vector based on morphology; The first hidden state extraction module is configured to input the first word feature vector and the second word feature vector into the third sub-model, and obtain a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model, wherein the third sub-model includes a plurality of hidden layers; the inputting the first word feature vector and the second word feature vector into the third sub-model, and obtaining a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model, comprises: inputting the first word feature vector into the third sub-model, and extracting hidden states from each hidden layer in the third sub-model in a forward-to-backward order, to obtain the plurality of first hidden states; inputting the second word feature vector into the third sub-model, and extracting hidden states from each hidden layer in the third sub-model in a backward-to-backward order, to obtain the plurality of second hidden states; A first splicing module is configured to splice the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state; The recognition module is configured to input the concatenated hidden state into the classification layer, and obtain the entity recognition result of the text to be recognized through classification and recognition by the classification layer.

11. An entity recognition model training device, It is characterized in that include: A second acquisition module is configured to acquire a training set and an initial network model, wherein the training set includes a plurality of training texts, each training text carries entity annotation information, and the network model includes a first sub-model, a second sub-model, a third sub-model, and a classification layer, wherein the first sub-model is a language model based on deep learning, the second sub-model is a lexical association model, and the third sub-model is a bidirectional hidden state extraction model; A second language analysis module is configured to extract training text from the training set, and input the training text into the first sub-model and the second sub-model respectively, to obtain a first word feature vector based on grammar and semantics and a second word feature vector based on morphology; The second hidden state extraction module is configured to input the first word feature vector and the second word feature vector into the third sub-model, and obtain a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model, wherein the third sub-model includes a plurality of hidden layers; the inputting the first word feature vector and the second word feature vector into the third sub-model, and obtaining a plurality of first hidden states and a plurality of second hidden states through bidirectional hidden state extraction of the third sub-model, comprises: inputting the first word feature vector into the third sub-model, and extracting hidden states from each hidden layer in the third sub-model in a forward-to-backward order, to obtain the plurality of first hidden states; inputting the second word feature vector into the third sub-model, and extracting hidden states from each hidden layer in the third sub-model in a backward-to-backward order, to obtain the plurality of second hidden states; A second splicing module is configured to splice the plurality of first hidden states and the plurality of second hidden states to obtain a spliced ​​hidden state; A prediction module is configured to input the concatenated hidden state into the classification layer, and obtain an entity prediction result of the training text through classification and recognition by the classification layer; A comparison module is configured to compare the entity prediction result with the entity annotation information carried by the training text to obtain a difference value; The adjustment module is configured to adjust the model parameters of the first sub-model, the second sub-model, the third sub-model and the classification layer if the difference value is greater than a preset threshold, and return to execute the step of extracting training text from the training set until the training stop condition is reached, stop training, and determine that the network model that has completed the training is an entity recognition model.

12. A computing device, It is characterized in that include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the method described in any one of claims 1 to 6 or any one of 7 to 9.

13. A computer-readable storage medium storing computer instructions, It is characterized in that When the instruction is executed by a processor, the steps of the method described in any one of claims 1 to 6 or any one of claims 7 to 9 are implemented.

Citation Information

Patent Citations

  • Named entity recognition method, electronic device and storage medium

    CN110287479A

  • Bad text detection method and device based on Bi-LSTM

    CN110321554A