Model training method and device based on natural language processing
By adopting multi-task learning in the natural language processing model and optimizing the loss functions of the first model and the second model, the problem of model overfitting is solved and the accuracy of named entity recognition is improved.
Patent Information
- Application Number
- CN202010293248.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-04-15
AI Technical Summary
In natural language processing, during the training process, due to the multi-meaning understanding of the natural language in the training sample, the model is prone to overfitting and the accuracy rate decreases.
A natural language processing model is constructed using multi-task learning, including a first model for named entity recognition and a second model for named entity recognition, and optimized by combining the loss functions of the first model and the second model.
The accuracy of the model's identification of named entities is improved, overfitting is avoided, reference information is provided through the recognition of the second model, and the recognition ability of the first model is enhanced.
Smart Images

Figure CN113536790B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a model training method and device based on natural language processing. Background Art
[0002] Natural Language Processing (NLP) is a discipline that studies language issues in human-computer interaction. It studies various theories and methods that enable effective communication between humans and computers using natural language, and consists of two main technical fields: natural language understanding and natural language generation. NLP technology is based on technologies and resources such as big data, knowledge graphs, machine learning, and linguistics, and can form specific application systems such as machine translation, deep question-answering, and dialogue systems, thereby serving various practical businesses and products.
[0003] Mathematical models based on natural language processing, such as hidden Markov models, maximum entropy models, conditional random fields, etc., need to be trained before being applied to actual scenarios to ensure the accuracy of the model output results. However, in the process of training the model, due to the understanding of multiple meanings of a word in the natural language of the training samples, the model is prone to overfitting, resulting in a decrease in accuracy. For example, when performing named entity recognition, using samples with polysemous words for training (People's Hospital, which can be both an organization name and a place name) can easily lead to overfitting of these noun features in the model, resulting in a decrease in model accuracy. Summary of the invention
[0004] In view of the above problems, the present invention proposes a model training method and device based on natural language processing, the main purpose of which is to improve the accuracy of model recognition.
[0005] In order to achieve the above object, the present invention mainly provides the following technical solutions:
[0006] On the one hand, the present invention provides a model training method based on natural language processing, which specifically includes:
[0007] Constructing a natural language processing model, wherein the natural language processing model includes a first model and a second model, wherein the first model is used to identify named entities, and the second model is used to identify leading words of named entities;
[0008] The natural language processing model is trained by adopting a multi-task learning method to obtain a loss function of the natural language processing model;
[0009] Optimization is performed according to the loss function of the natural language processing model to obtain an optimized natural language processing model.
[0010] Preferably, the first model adopts a BiLSTM-CRF model, whose structure includes:
[0011] An input layer, used to convert an input sentence into a vector sequence, wherein the input sentence includes a plurality of words;
[0012] A BiLSTM layer, used to combine the vector sequence converted by the input layer with context information to generate a corresponding feature vector;
[0013] A fully connected layer is used to receive the feature vector generated by the BiLSTM layer and calculate the distribution probability of the output label corresponding to each word on all labels;
[0014] The CRF layer is used to determine the output sequence of all labels according to the distribution probability output by the fully connected layer according to a preset rule.
[0015] Preferably, converting the input sentence into a vector sequence comprises:
[0016] Splitting the input sentence into word sequences;
[0017] The vector representation of each word is obtained to obtain the vector sequence of the input sentence.
[0018] Preferably, the BiLSTM layer includes multiple LSTM units, each LSTM unit is used to output a feature vector with a fixed length corresponding to a word vector in the vector sequence, wherein the multiple LSTM units have an association relationship corresponding to the arrangement order of the word vectors in the vector sequence.
[0019] Preferably, the structure of the second model includes: an input layer and a classification layer;
[0020] An input layer, used to obtain a vector representation of each word in an input sentence and a feature vector generated by the BiLSTM layer, and concatenate the vector representation of a specific word and the corresponding vector of the specific word in the feature vector to obtain a concatenated vector;
[0021] The classification layer is used to determine the distribution probability that the next word adjacent to the specific word is a named entity according to the concatenated vector by using a fully connected neural network.
[0022] Preferably, the loss function of the natural language processing model is obtained in the following manner:
[0023] Set weights for the loss function of the first model and the loss function of the second model respectively;
[0024] The loss function of the natural language processing model is calculated according to the loss function of the first model and the loss function of the second model after setting the weights.
[0025] Preferably, the loss function of the first model adopts a CRF likelihood function; and / or,
[0026] The loss function of the second model adopts a cross entropy loss function, and the cross entropy loss function is the cross entropy of the predicted distribution probability and the actual distribution probability of the named entity by the second model.
[0027] On the other hand, the present invention provides a model training device based on natural language processing, specifically comprising:
[0028] A setting unit, used to construct a natural language processing model, wherein the natural language processing model includes a first model and a second model, wherein the first model is used to identify named entities, and the second model is used to identify leading words of named entities;
[0029] A training unit, configured to train the natural language processing model constructed by the setting unit by adopting a multi-task learning method to obtain a loss function of the natural language processing model;
[0030] The optimization unit is used to optimize the loss function of the natural language processing model obtained by the training unit to obtain an optimized natural language processing model.
[0031] Preferably, the first model adopts a BiLSTM-CRF model, whose structure includes:
[0032] An input layer, used to convert an input sentence into a vector sequence, wherein the input sentence includes a plurality of words;
[0033] A BiLSTM layer, used to combine the vector sequence converted by the input layer with context information to generate a corresponding feature vector;
[0034] A fully connected layer is used to receive the feature vector generated by the BiLSTM layer and calculate the distribution probability of the output label corresponding to each word on all labels;
[0035] The CRF layer is used to determine the output sequence of all labels according to the distribution probability output by the fully connected layer according to a preset rule.
[0036] Preferably, the input layer is specifically used to: split the input sentence into a word sequence; obtain a vector representation of each word to obtain a vector sequence of the input sentence.
[0037] Preferably, the BiLSTM layer includes multiple LSTM units, each LSTM unit is used to output a feature vector with a fixed length corresponding to a word vector in the vector sequence, wherein the multiple LSTM units have an association relationship corresponding to the arrangement order of the word vectors in the vector sequence.
[0038] Preferably, the structure of the second model includes: an input layer and a classification layer;
[0039] An input layer, used to obtain a vector representation of each word in an input sentence and a feature vector generated by the BiLSTM layer, and concatenate the vector representation of a specific word and the corresponding vector of the specific word in the feature vector to obtain a concatenated vector;
[0040] The classification layer is used to determine the distribution probability that the next word adjacent to the specific word is a named entity according to the concatenated vector by using a fully connected neural network.
[0041] Preferably, the loss function of the natural language processing model is obtained in the following manner:
[0042] Set weights for the loss function of the first model and the loss function of the second model respectively;
[0043] The loss function of the natural language processing model is calculated according to the loss function of the first model and the loss function of the second model after setting the weights.
[0044] Preferably, the loss function of the first model adopts a CRF likelihood function; and / or,
[0045] The loss function of the second model adopts a cross entropy loss function, and the cross entropy loss function is the cross entropy of the predicted distribution probability and the actual distribution probability of the named entity by the second model.
[0046] On the other hand, the present invention provides a processor, which is used to run a program, wherein the program executes the above-mentioned model training method based on natural language processing when running.
[0047] By means of the above technical scheme, the present invention provides a model training method and device based on natural language processing. By constructing a model structure having a first model and a second model, the second model can recognize the leading words of a named entity, thereby providing reference information for the first model to recognize the named entity, so that the constructed natural language processing model has a higher accuracy in recognizing named entities and avoids the occurrence of overfitting. When training the natural language processing model, the loss function of the model is obtained by training using a multi-task learning method, that is, the optimal solution of the model loss function is obtained based on the combined optimization of the loss functions of the first model and the second model, so that the trained natural language processing model can perform a comprehensive analysis based on the recognition of the leading words of the named entity when recognizing the named entity, thereby improving the accuracy of the model in recognizing the named entity.
[0048] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0050] Figure 1 A flow chart of a model training method based on natural language processing proposed in an embodiment of the present invention is shown;
[0051] Figure 2 A structural block diagram of a natural language processing model proposed in an embodiment of the present invention is shown;
[0052] Figure 3 A block diagram of a model training device based on natural language processing proposed in an embodiment of the present invention is shown;
[0053] Figure 4 A block diagram showing the composition of another model training device based on natural language processing proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.
[0055] Natural language processing, that is, achieving natural language communication between humans and machines, or achieving natural language understanding and natural language generation, is very difficult. The fundamental reason for the difficulty is the various ambiguities or polysemy that exist widely at all levels of natural language texts and dialogues. A Chinese text is formally a string of Chinese characters (including punctuation marks, etc.). Characters can be combined into words, words can be combined into phrases, phrases can be combined into sentences, and then some sentences can be combined into paragraphs, sections, chapters, and articles. Whether at the various levels mentioned above: characters (characters), words, phrases, sentences, paragraphs, etc., or in the transition from the next level to the next level, there are ambiguities and polysemy, that is, a string of characters that is the same in form can be understood as different word strings, phrase strings, etc., and have different meanings in different scenes or different contexts.
[0056] Natural language processing is a model for studying language ability and language application. It is implemented by establishing a computer algorithm framework, and is finally used in various practical systems through training, improvement, and evaluation. The scenarios in which natural language processing models are applied include information retrieval, machine translation, document classification, information extraction, text mining, etc.
[0057] The embodiment of the present invention provides a model training method based on natural language processing. The main application scenario of the model is text content analysis, especially the recognition of named entities in text. The specific steps of this method are as follows Figure 1 As shown, the method includes:
[0058] Step 101: construct a natural language processing model, which includes a first model and a second model.
[0059] The first model is used to identify named entities in the text, and the second model is used to identify leading words of the named entities.
[0060] In this step, the leading word recognized by the second model refers to a word or a word segment before the named entity class word segment. That is to say, in the expression of natural language, the word or word following a leading word is most likely a named entity. It can be seen that in the natural language processing model constructed in this step, the purpose of the second model recognizing the leading word is to provide reference information for the first model to recognize the named entity, and the reference information is determined based on the word segmentation association relationship in the text.
[0061] Step 102: Train the natural language processing model using a multi-task learning method to obtain a loss function of the natural language processing model.
[0062] The training process of the model is to input samples with labeled information into the model, compare the output results with the actual labeled information of the samples, and adjust the relevant parameters of the model so that the output of the model is closer to the actual results, that is, the model has accurate prediction results. To this end, when building the model, a loss function will be determined for the model, and the loss function will be optimized using training samples to obtain its optimal solution, that is, to make the output of the model closest to the actual results.
[0063] In this step, the loss function of the natural language processing model can be regarded as a combination of the loss function of the first model and the loss function of the second model, and the specific combination method is not limited to weighted summation, averaging, maximum value, etc. The process of finding the optimal solution for the loss function of the natural language processing model in this step is to find the optimal solution for the combination of the loss function of the first model and the loss function of the second model, so that not only the characteristics of the input text content but also the text content associated with the text content are considered in the output result, thereby avoiding overfitting of the model. To this end, this step adopts a multi-task learning method, and sets the loss function of the first model and the loss function of the second model as two different tasks for combined training.
[0064] Step 103: Optimize according to the loss function of the natural language processing model to obtain an optimized natural language processing model.
[0065] The process of optimizing the loss function of the natural language processing model in this step is the process of obtaining the optimal solution through multi-task learning in the above step. Since multi-task learning has been widely used in the field of natural language processing, this step does not provide a detailed description of the solution process.
[0066] Through the description of the above embodiments, the model training method based on natural language processing provided by the present invention mainly constructs a model structure with a first model and a second model, so that the second model can recognize the leading words of the named entity, and can provide reference information for the first model to recognize the named entity. When training the natural language processing model, the loss function of the model is also composed of a combination of the loss functions corresponding to the first model and the second model, that is, the optimal solution of the model loss function is obtained based on the combined optimization of the loss functions of the first model and the second model, so that the trained natural language processing model can perform a comprehensive analysis based on the leading words to determine the corresponding named entity when recognizing named entities in the text, thereby improving the accuracy of the model output and avoiding the phenomenon of model overfitting.
[0067] Further, for Figure 1 The natural language processing model described in the following details the specific process of building and training the model when it is applied to the named entity recognition task scenario:
[0068] First of all, for building a named entity recognition model, the named entity recognition task is a basic task in natural language processing, which refers to identifying named referents from text, paving the way for tasks such as relationship extraction. Generally speaking, named entity recognition mainly refers to identifying three types of named entities in text: names of people, places, and organizational structures. At present, the more popular method in the field of named entity recognition is to convert the named entity recognition problem into a sequence labeling problem, and then solve it through sequence labeling methods. The general solutions for sequence labeling are: hidden Markov model HMM or conditional random field CRF or BiLSTM-CRF or BiLSTM-maximum entropy. The first two are statistical learning methods, and the latter two are neural network methods.
[0069] In the embodiment of the present invention, the named entity recognition model is constructed on the basis of the BiLSTM-CRF model. Specifically, the BiLSTM-CRF model is used as the first model, and its structure can be mainly divided into four layers, namely: input layer, BiLSTM layer, fully connected layer, and CRF layer.
[0070] The input layer is used to convert the sentences in the input text into a vector sequence, that is, to represent the sentences in the input text as a word vector sequence or a character vector sequence. In this embodiment, the input layer splits the input sentence into a character sequence, and then obtains the vector representation of each character by looking up a table, thereby obtaining the vector sequence of the entire sentence. The table to be looked up is a preset character vector comparison table, which records the mapping relationship between characters and vectors.
[0071] The BiLSTM layer is used to combine the vector sequence converted by the input layer with the context information to generate the corresponding feature vector. The BiLSTM layer is composed of multiple LSTM units, each of which is used to output a feature vector with a fixed length corresponding to a word vector in the vector sequence. The feature vector is determined based on the characteristics of the previous and next word vectors of the word vector, that is, the feature vector can be regarded as a feature vector corresponding to each word that incorporates the context information. Therefore, there is an association relationship between multiple LSTM units corresponding to the arrangement order of the word vectors in the vector sequence.
[0072] The fully connected layer is used to receive the feature vector generated by the BiLSTM layer and calculate the distribution probability of the output label corresponding to each word on all labels. For example, assuming there are two entity types: Person and Organization, if the BIO annotation system is used, five entity labels will be obtained: B-Person (the beginning of the person segmentation), I-Person (the middle of the person segmentation), B-Organization (the beginning of the organization segmentation), I-Organization (the middle of the organization segmentation), O (does not belong to any category), and the fully connected layer will output the distribution probability of each word in the corresponding label, such as B-Person (1.5), I-Person (0.9), B-Organization (0.1), I-Organization (0.08), O (0.05).
[0073] The CRF layer is used to determine the output sequence of all labels according to the distribution probability output by the fully connected layer according to the preset rules. That is, the rationality of the output sequence is analyzed according to the preset rules to obtain a reasonable prediction result and realize the labeling of the named entity segmentation in the sentence. In other words, the CRF layer searches for the global optimal output sequence according to the probability distribution of each word, that is, the words are combined and the labels of the named entities are marked on the segmentation composed of the words.
[0074] The above is a structural description of the BiLSTM-CRF model, which is a commonly used model for named entity recognition. Its specific implementation principle will not be described in detail in this embodiment. For the named entity recognition model required in this embodiment, in addition to using the BiLSTM-CRF model as the first model, a second model needs to be constructed. The second model is used to determine the distribution probability of the next word adjacent to the input word vector as a named entity by analyzing the input word vector. The specific structure of the second model includes: an input layer and a classification layer.
[0075] The input layer is used to obtain the vectors of each word in the sentence of the input text. Specifically, the input data of the output layer comes from the output of the input layer of the first model and the BiLSTM layer, that is, the first vector output by the input layer in the first model is concatenated with the second vector output by the BiLSTM layer, wherein the first vector is the vector representation of each word in the input sentence, and the second vector is the vector corresponding to the specific word corresponding to the first vector in the feature vector. For example, the first vector output by the input layer is [1,2], and the corresponding second vector in the feature vector output by the BiLSTM layer is [3,4]. Then the input layer of the second model concatenates these two vectors to obtain the vector [1,2,3,4], which is used as the output result and input to the classification layer.
[0076] The classification layer is used to determine the distribution probability that the next word adjacent to the specific word is a named entity based on the concatenated vector output by the input layer using a fully connected neural network.
[0077] In this embodiment, the purpose of the second model is to determine the leading word in the text, where the leading word refers to a word or phrase before the named entity class segmentation. For example, in the sentence "I go to Beijing", "Beijing" is a place name, and the corresponding "go" is the leading word. The second model identifies the probability that the input word is the leading word, and gives the corresponding distribution probability that the next word of the word is a named entity class word. In other words, the greater the probability that the word recognized by the second model is the leading word, the greater the probability that the next word of the word is a named entity. And by identifying the leading word, it is equivalent to increasing the dimension of named entity recognition, thereby avoiding the phenomenon of model overfitting caused by different applications of a word with multiple meanings in different scenarios.
[0078] After the above structural description of the first model and the second model, the natural language processing model also needs to determine the loss function for it. In an embodiment of the present invention, the loss function is composed of the loss function of the first model and the loss function of the second model. Specifically, the loss function of the first model of the first model adopts the CRF likelihood function, that is, in the case of a given label, see which label's probability distribution is output, and determine one of the optimal given labels. For the loss function of the second model, the cross entropy loss function is adopted, and the cross entropy loss function is the cross entropy of the predicted distribution probability and the actual distribution probability of the second model for the named entity. Cross entropy is generally used to find the gap between the target and the predicted value in deep learning. According to the description of the above embodiment, the predicted distribution probability refers to the probability of whether the word is a leading word, and the actual distribution probability refers to whether the word in the sample is a leading word. In actual applications, the training sample only marks the segmentation or word corresponding to the named entity, and does not mark the leading word. For this, it is necessary to see whether the next word or word of the word is a marked named entity when processing the segmentation. If so, the word is considered to be a leading word, otherwise, it is determined that the word is not a leading word.
[0079] Furthermore, the loss function of the natural language processing model is obtained by weighting and summing the loss function of the first model and the loss function of the second model according to preset weights, that is, the greater the weight of the loss function of the second model, the greater the influence of the leading word on the named entity segmentation recognition. To this end, it is necessary to set weights for the loss function of the first model and the loss function of the second model in advance.
[0080] According to the above description of the construction of the natural language processing model, the specific structure of the natural language processing model is as follows Figure 2As shown, in the first model, the input layer, BiLSTM layer and CRF layer are shown respectively, and the fully connected layer is not shown. The second model includes the input layer and the classification layer, and the input data of the input layer comes from the input layer and the BiLSTM layer of the first model. The loss functions of the first model and the second model together constitute the loss function of the natural language processing model.
[0081] Finally, the above-mentioned natural language processing model is trained using the labeled training samples, that is, the optimal solution of its loss function is obtained through the training samples. In the present invention, since there are multiple independent models in the natural language processing model, and different models process different tasks, during training, multi-task learning can be used to train to simultaneously optimize the loss functions of multiple models, so that the loss function of the natural language processing model can obtain the optimal solution. Specifically in this embodiment, the stochastic gradient descent method is used to optimize the loss function. The stochastic gradient descent method is a common method for solving the loss function in machine learning. Therefore, the principle and process of the stochastic gradient descent method for finding the optimal solution of the loss function in this embodiment are no longer explained.
[0082] The above is an explanation of the application of the natural language processing model proposed in the embodiment of the present invention in the named entity recognition scenario. It can be seen from the natural language processing model constructed above that it can help the first model to more accurately recognize the named entities in the sentence by effectively recognizing the leading word in the sentence through the second model, avoiding the overfitting phenomenon of polysemous words, and improving the recognition accuracy of the model. Among them, the named entity recognition scenario can specifically be a scenario for entity recognition such as names of people, places, and organization names, which can be specifically applied to real-time scenarios such as dialogues and barrage texts, and can also be applied to ordinary text recognition scenarios.
[0083] Furthermore, as a response to the above Figure 1 The embodiment of the present invention provides a model training device based on natural language processing, the main purpose of which is to improve the accuracy of natural language processing model recognition. For ease of reading, this device embodiment will no longer repeat the details of the above method embodiments, but it should be clear that the device in this embodiment can correspond to all the contents of the above method embodiments. Figure 3 As shown, specifically including:
[0084] A setting unit 21 is used to construct a natural language processing model, wherein the natural language processing model includes a first model and a second model, wherein the first model is used to identify named entities, and the second model is used to identify leading words of named entities;
[0085] A training unit 22 is used to train the natural language processing model constructed by the setting unit 21 by adopting a multi-task learning method to obtain a loss function of the natural language processing model;
[0086] The optimization unit 23 is used to optimize the loss function of the natural language processing model obtained by the training unit 22 to obtain an optimized natural language processing model.
[0087] Furthermore, the first model set by the setting unit 21 adopts a BiLSTM-CRF model, and its structure includes:
[0088] An input layer, used to convert an input sentence into a vector sequence, wherein the input sentence includes a plurality of words;
[0089] A BiLSTM layer, used to combine the vector sequence converted by the input layer with context information to generate a corresponding feature vector;
[0090] A fully connected layer is used to receive the feature vector generated by the BiLSTM layer and calculate the distribution probability of the output label corresponding to each word on all labels;
[0091] The CRF layer is used to determine the output sequence of all labels according to the distribution probability output by the fully connected layer according to a preset rule.
[0092] Furthermore, the input layer is specifically used to: split the input sentence into a word sequence; obtain a vector representation of each word to obtain a vector sequence of the input sentence.
[0093] Furthermore, the BiLSTM layer includes multiple LSTM units, each LSTM unit is used to output a feature vector with a fixed length corresponding to a word vector in the vector sequence, wherein the multiple LSTM units have an association relationship corresponding to the arrangement order of the word vectors in the vector sequence.
[0094] Further, the structure of the second model set by the setting unit 21 includes: an input layer and a classification layer;
[0095] An input layer, used to obtain a vector representation of each word in an input sentence and a feature vector generated by the BiLSTM layer, and concatenate the vector representation of a specific word and the corresponding vector of the specific word in the feature vector to obtain a concatenated vector;
[0096] The classification layer is used to determine the distribution probability that the next word adjacent to the specific word is a named entity according to the concatenated vector by using a fully connected neural network.
[0097] Further, such as Figure 4As shown, the loss function of the natural language processing model obtained by the training unit 22 is obtained in the following manner, which specifically includes:
[0098] A weight setting module 221, used to set weights for the loss function of the first model and the loss function of the second model respectively;
[0099] The determination module 222 is used to calculate the loss function of the natural language processing model according to the loss function of the first model and the loss function of the second model after the weights are set by the weight setting module 221.
[0100] Furthermore, the loss function of the first model adopts a CRF likelihood function; and / or,
[0101] The loss function of the second model adopts a cross entropy loss function, and the cross entropy loss function is the cross entropy of the predicted distribution probability and the actual distribution probability of the named entity by the second model.
[0102] In addition, an embodiment of the present invention further provides a processor, which is used to run a program, wherein when the program is running, the model training method based on natural language processing provided by any one of the above embodiments is executed.
[0103] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0104] It is understandable that the related features in the above methods and devices can be referenced to each other. In addition, the "first", "second" and the like in the above embodiments are used to distinguish the embodiments, but do not represent the advantages and disadvantages of the embodiments.
[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0106] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the present invention is not directed to any specific programming language either. It should be understood that various programming languages can be utilized to realize the content of the present invention described herein, and the description of the above specific languages is for disclosing the preferred embodiment of the present invention.
[0107] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0108] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0109] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0110] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0112] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0113] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0114] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0115] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0116] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0117] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A model training method based on natural language processing, the method comprising: Constructing a natural language processing model, wherein the natural language processing model includes a first model and a second model, wherein the first model is used to identify named entities, and the second model is used to identify leading words of named entities; The first model includes: an input layer, used to convert an input sentence into a vector sequence, wherein the input sentence includes a plurality of words, a BiLSTM layer, used to combine the vector sequence converted by the input layer with context information to generate a corresponding feature vector, and a fully connected layer, used to receive the feature vector generated by the BiLSTM layer, and calculate the distribution probability of the output label corresponding to each word on all labels; The structure of the second model includes: an input layer and a classification layer, wherein the input layer is used to obtain the vector representation of each word in the input sentence and the feature vector generated by the BiLSTM layer, concatenate the vector representation of the specific word and the corresponding vector of the specific word in the feature vector to obtain a concatenated vector, and the classification layer is used to determine the distribution probability that the next word adjacent to the specific word is a named entity according to the concatenated vector using a fully connected neural network; The natural language processing model is trained by adopting a multi-task learning method to obtain a loss function of the natural language processing model; Optimization is performed according to the loss function of the natural language processing model to obtain an optimized natural language processing model.
2. The method according to claim 1, characterized in that The first model adopts the BiLSTM-CRF model, and its structure also includes: The CRF layer is used to determine the output sequence of all labels according to the distribution probability output by the fully connected layer according to a preset rule.
3. The method according to claim 2, characterized in that The step of converting the input sentence into a vector sequence includes: Splitting the input sentence into word sequences; The vector representation of each word is obtained to obtain the vector sequence of the input sentence.
4. The method according to claim 2, characterized in that: The BiLSTM layer includes multiple LSTM units, each LSTM unit is used to output a feature vector with a fixed length corresponding to a word vector in the vector sequence, wherein the multiple LSTM units have an association relationship corresponding to the arrangement order of the word vectors in the vector sequence.
5. The method according to claim 1, characterized in that The loss function of the natural language processing model is obtained as follows: Set weights for the loss function of the first model and the loss function of the second model respectively; The loss function of the natural language processing model is calculated according to the loss function of the first model and the loss function of the second model after setting the weights.
6. The method according to claim 5, characterized in that The loss function of the first model adopts the CRF likelihood function; and / or, The loss function of the second model adopts a cross entropy loss function, and the cross entropy loss function is the cross entropy of the predicted distribution probability and the actual distribution probability of the second model for the named entity.
7. A model training device based on natural language processing, the device comprising: A setting unit is used to construct a natural language processing model, wherein the natural language processing model includes a first model and a second model, wherein the first model is used to identify a named entity, and the second model is used to identify a leading word of a named entity, and the first model includes: an input layer, used to convert an input sentence into a vector sequence, wherein the input sentence includes a plurality of words, a BiLSTM layer, used to combine the vector sequence converted by the input layer with context information to generate a corresponding feature vector, and a fully connected layer, used to receive the feature vector generated by the BiLSTM layer, and calculate the distribution probability of the output label corresponding to each word on all labels; the structure of the second model includes: an input layer and a classification layer, wherein the input layer is used to obtain the vector representation of each word in the input sentence and the feature vector generated by the BiLSTM layer, and concatenate the vector representation of a specific word and the corresponding vector of the specific word in the feature vector to obtain a concatenated vector, and the classification layer is used to determine the distribution probability of the next word adjacent to the specific word being a named entity according to the concatenated vector using a fully connected neural network; a training unit, used to train the natural language processing model constructed by the setting unit in a multi-task learning manner to obtain a loss function of the natural language processing model; The optimization unit is used to optimize the loss function of the natural language processing model obtained by the training unit to obtain an optimized natural language processing model.
8. The device according to claim 7, characterized in that The first model adopts the BiLSTM-CRF model, and its structure also includes: The CRF layer is used to determine the output sequence of all labels according to the distribution probability output by the fully connected layer according to a preset rule.
9. The device according to claim 8, characterized in that The structure of the second model includes: an input layer and a classification layer; An input layer, used to obtain a vector representation of each word in an input sentence and a feature vector generated by the BiLSTM layer, and concatenate the vector representation of a specific word and the corresponding vector of the specific word in the feature vector to obtain a concatenated vector; The classification layer is used to determine the distribution probability that the next word adjacent to the specific word is a named entity according to the concatenated vector by using a fully connected neural network.
10. The device according to claim 9, characterized in that The loss function of the natural language processing model is obtained as follows: Set weights for the loss function of the first model and the loss function of the second model respectively; The loss function of the natural language processing model is calculated according to the loss function of the first model and the loss function of the second model after setting the weights.
11. The device according to claim 10, characterized in that The loss function of the first model adopts the CRF likelihood function; and / or, The loss function of the second model adopts a cross entropy loss function, and the cross entropy loss function is the cross entropy of the predicted distribution probability and the actual distribution probability of the second model for the named entity.
12. A processor, characterized in that: The processor is used to run a program, wherein the program, when running, executes the model training method based on natural language processing described in any one of claims 1 to 6.
Citation Information
Patent Citations
Neural network training method and device, and named entity recognition method and device
CN109062901A