Named entity recognition method and device

Through multi-task learning combining named entity recognition and sentence classification tasks, sharing layer and task-specific layer, the problem of NER model training difficulties in low-resource language environments is solved, and efficient named entity recognition effect is achieved.

CN113536791BActive Publication Date: 2025-05-13ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010314468.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-20
Publication Date
2025-05-13
Estimated Expiration
2040-04-20

AI Technical Summary

Technical Problem

In low-resource locale environments, the lack of manual annotation corpus results in difficulty in training named entity recognition (NER) models, affecting the effectiveness of the model.

Method used

The multi-task learning method is adopted, combining named entity recognition tasks and sentence classification tasks, sharing layers and task-specific layers, and the generalization ability and performance of the model are improved by pre-training sentence classification models and jointly training the NER model.

Benefits of technology

It realizes effective training of NER models in low-resource language environments, improves the accuracy and efficiency of naming entity recognition, and reduces the dependence on manual annotation corpus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113536791B_ABST
    Figure CN113536791B_ABST
Patent Text Reader

Abstract

A method and device for named entity recognition are disclosed. With the named entity recognition task as the main task and the sentence classification task as the auxiliary task, a recognition model is trained in a multi-task learning manner. The recognition model has a shared layer shared by the named entity recognition task and the sentence classification task and task-specific layers used for the named entity recognition task and the sentence classification task respectively. Text is input into the trained recognition model to obtain the corresponding named entity recognition result. Thus, model training can be conveniently and effectively performed for low-resource language NER, thereby achieving better NER effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a named entity recognition scheme, and in particular to an entity recognition scheme for setting category labels for word segmentation of product titles. Background Art

[0002] In the field of e-commerce, for example, sellers will set promotional statements such as product titles for the products they sell. The product title is a brief description of the product on sale by the seller, such as "…flexible silicone shell…", which contains a lot of information related to the product, such as "flexible" for its style / attributes, "silicone" for its material, and "shell" for what product it is. The Named Entity Recognition (NER) system can be used to extract information from the product title to set a corresponding category label for the product.

[0003] The goal of the NER system is to identify spans of tokens in the input text and classify them into predefined categories. The input text can be a sentence. The predefined categories are typically "person name", "place name", "organization name", etc. For specific fields such as e-commerce, the input text can be a product title, and the predefined categories can be "product name", "brand", "consumer group", etc. This depends on the design of the target named entity type.

[0004] Currently, manually annotated data is used to train NER models in NER systems.

[0005] For example, given a text segment "...Flexible Silicone Case..." taken from a product title, its annotated version would be "...[Flexible]_Style[Silicone]_Material[Case]_Product..." In this example, the tag "Silicone" is annotated (tagged) as "Material" and the tag "Case" is annotated as "Product".

[0006] Token-level NER annotations like the example above require a lot of human effort.

[0007] Therefore, in low-resource language NER for languages ​​lacking manually annotated language resources or corpora, such as category label identification for product titles, there is still a need for a convenient and effective NER model training method to achieve better NER effects. Summary of the invention

[0008] A technical problem to be solved by the present disclosure is to provide a NER solution for conveniently and effectively training NER models for low-resource languages.

[0009] According to the first aspect of the present disclosure, a method for named entity recognition is provided, comprising: taking a named entity recognition task as a main task and a sentence classification task as an auxiliary task, training a recognition model in a multi-task learning manner, wherein the recognition model has a shared layer shared by the named entity recognition task and the sentence classification task and task-specific layers respectively used for the named entity recognition task and the sentence classification task; and inputting text into the trained recognition model to obtain a corresponding named entity recognition result.

[0010] Optionally, the step of training the recognition model includes: pre-training the sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, the sentence classification model including a shared layer and a task-specific layer for a sentence classification task; and training the recognition model using training samples with named entity recognition labels and sentence classification labels.

[0011] Optionally, to train a sentence classification model, the following multi-class cross entropy loss function L C Minimize:

[0012]

[0013] Where i represents the sentence index, N is the number of training samples, K is the number of target categories, and s k is the standardized prediction score of the kth target category after applying the softmax function, and t is the true label of the one-hot encoding. In order to train the named entity recognition model, the negative log-likelihood function L of the correct label sequence relative to the training set is NER Minimize:

[0014]

[0015] Among them, y represents the label sequence, p(y (i) |H' (i) ) is based on the final hidden representation H' corresponding to the i-th sentence (i) The probability of the obtained label sequence y, the named entity recognition model includes a shared layer and a task-specific layer for the named entity recognition task, combined with L C and L NER , we get the joint loss function L JOINT :

[0016]

[0017] Among them, λ is a balancing parameter that minimizes the joint loss function during the training of the recognition model.

[0018] Optionally, the shared layer includes at least one of the following: a word embedding layer, a projection layer, a BiLSTM layer, and an attention layer, which outputs a final hidden representation; and / or the task-specific layer for the named entity recognition task includes a conditional random field layer, which obtains a named entity recognition result based on the final hidden representation; and / or the task-specific layer for the sentence classification task includes a pooling layer and a linear layer, the pooling layer performs pooling processing on the final hidden representation to obtain an input to the linear layer, and the linear layer outputs a sentence classification result.

[0019] Optionally, the word embedding layer is a pre-trained word embedding layer; and / or the input of the word embedding layer is a word segmentation sequence obtained after word segmentation processing of the training sample or text; and / or the word embedding layer represents the words in the input word segmentation sequence as corresponding word embedding vectors; and / or the projection layer projects the word embedding vector to obtain the input vector corresponding to the word segmentation of the BiLSTM layer; and / or the BiLSTM layer outputs a hidden representation corresponding to the word segmentation; and / or the attention layer applies an attention mechanism to the hidden representation to obtain a final hidden representation corresponding to the word segmentation; and / or the pooling layer performs maximum pooling on the final hidden representation to create a fixed-size global vector as the input of the linear layer; and / or the linear layer obtains the prediction score for each classification based on the fixed-size global vector output by the pooling layer.

[0020] Optionally, the attention layer obtains the final hidden representation by the following formula:

[0021] H′=concat(head1,...,head n )W O +H,

[0022] head j =attention(Q j , K j , V j ),

[0023]

[0024] Where H is the hidden representation, H′ is the final hidden representation, n is the number of heads in the self-attention mechanism, j is the sequence number of the corresponding head, 1≤j≤n, W O is the weight matrix, concat() is the connection function, and attention() is the attention function.

[0025] Optionally, the attention function is calculated by the following formula:

[0026]

[0027]

[0028] Where w is the weight vector, d h is the dimension of the hidden representation.

[0029] Optionally, the step of inputting the text into the trained recognition model includes: performing word segmentation processing on the text to obtain a word segmentation sequence; and inputting the word segmentation sequence into a word embedding layer in the shared layer.

[0030] Optionally, the method is applied to an e-commerce scenario; and / or the text is a product title; and / or the named entity recognition result is a category corresponding to each word in the product title.

[0031] According to a second aspect of the present disclosure, a method for named entity recognition is provided, comprising: providing a named entity recognition model, the named entity recognition model is obtained by training the named entity recognition model and the sentence classification model in a multi-task learning manner with a named entity recognition task as a main task and a sentence classification task as an auxiliary task, wherein the named entity recognition model and the sentence classification model have a common shared layer and respective task-specific layers; and inputting text into the named entity recognition model to obtain a corresponding named entity recognition result.

[0032] According to a third aspect of the present disclosure, a method for setting category labels for word segmentations of a product title is provided, comprising: providing a category recognition model, the category recognition model is obtained by taking a category recognition task as a main task and a sentence classification task as an auxiliary task, and training the category recognition model and the sentence classification model in a multi-task learning manner, wherein the category recognition model and the sentence classification model have a common shared layer and respective task-specific layers; performing word segmentation processing on the product title to obtain a word segmentation sequence; and inputting the word segmentation sequence into the category recognition model to obtain a category label corresponding to each word in the word segmentation sequence.

[0033] Optionally, the category label is a category label in a predetermined category label set.

[0034] According to a fourth aspect of the present disclosure, a method for training a named entity recognition model is provided, comprising: obtaining a sentence with a sentence classification label as a first training sample; obtaining a sentence with a named entity recognition label and a sentence classification label as a second training sample; and using the first training sample and the second training sample, with the named entity recognition task as the main task and the sentence classification task as the auxiliary task, to train the recognition model in a multi-task learning manner, wherein the recognition model has a shared layer shared by the named entity recognition task and the sentence classification task, and task-specific layers used for the named entity recognition task and the sentence classification task, respectively.

[0035] Optionally, the step of training the recognition model includes: pre-training the sentence classification model using the first training sample to obtain pre-trained shared layer parameters, the sentence classification model including a shared layer and a task-specific layer for the sentence classification task; and training the recognition model using the second training sample.

[0036] Optionally, to train a sentence classification model, the following multi-class cross entropy loss function L C Minimize:

[0037]

[0038] Where i represents the sentence index, N is the number of training samples, K is the number of target categories, and s k is the standardized prediction score of the kth target classification after applying the softmax function, and t is the true label of the one-hot encoding. In order to train the NER model, the negative log-likelihood function L of the correct label sequence relative to the training set is NER Minimize:

[0039]

[0040] Among them, y represents the label sequence, p(y (i) |H' (i) ) is based on the final hidden representation H' corresponding to the i-th sentence (i) The probability of the obtained label sequence y, the named entity recognition model includes a shared layer and a task-specific layer for the named entity recognition task, combined with L C and L NER , we get the joint loss function L JOINT :

[0041]

[0042] Among them, λ is a balancing parameter that minimizes the joint loss function during the training of the named entity recognition model and the sentence classification model.

[0043] According to a fifth aspect of the present disclosure, a named entity recognition device is provided, comprising: a model training device for training a recognition model in a multi-task learning manner using a named entity recognition task as a main task and a sentence classification task as an auxiliary task, wherein the recognition model has a shared layer shared by the named entity recognition task and the sentence classification task and task-specific layers respectively used for the named entity recognition task and the sentence classification task; and a recognition device for inputting text into a trained named entity recognition model to obtain a corresponding named entity recognition result.

[0044] Optionally, the model training device includes: a first training device, which pre-trains a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, the sentence classification model including a shared layer and a task-specific layer for a sentence classification task; and a second training device, which trains a recognition model using training samples with named entity recognition labels and sentence classification labels.

[0045] According to a sixth aspect of the present disclosure, a named entity recognition device is provided, comprising: a preparation device for providing a named entity recognition model, wherein the named entity recognition model is obtained by training the named entity recognition model and the sentence classification model in a multi-task learning manner with a named entity recognition task as a main task and a sentence classification task as an auxiliary task, wherein the named entity recognition model and the sentence classification model have a common shared layer and respective task-specific layers; and a recognition device for inputting text into the named entity recognition model to obtain a corresponding named entity recognition result.

[0046] According to a seventh aspect of the present disclosure, there is provided a device for setting category labels for word segmentations of a product title, comprising: a preparation device for providing a category recognition model, the category recognition model being obtained by training the category recognition model and the sentence classification model in a multi-task learning manner with a category recognition task as a main task and a sentence classification task as an auxiliary task, wherein the category recognition model and the sentence classification model have a common shared layer and respective task-specific layers; a word segmentation device for performing word segmentation processing on a product title to obtain a word segmentation sequence; and a recognition device for inputting the word segmentation sequence into a named entity recognition model to obtain a category label corresponding to each word in the word segmentation sequence.

[0047] According to an eighth aspect of the present disclosure, there is provided a device for training a named entity recognition model, comprising: a first acquisition device for acquiring sentences with sentence classification labels as first training samples; a second acquisition device for acquiring sentences with named entity recognition labels and sentence classification labels as second training samples; and a model training device for using the first training samples and the second training samples, with the named entity recognition task as the main task and the sentence classification task as the auxiliary task, to train the recognition model in a multi-task learning manner, wherein the recognition model has a shared layer shared by the named entity recognition task and the sentence classification task, and task-specific layers respectively used for the named entity recognition task and the sentence classification task.

[0048] According to a ninth aspect of the present disclosure, a computing device is provided, comprising: a processor; and a memory on which executable codes are stored, and when the executable codes are executed by the processor, the processor executes the methods described in the first to fourth aspects above.

[0049] According to the tenth aspect of the present disclosure, a non-temporary machine-readable storage medium is provided, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor executes the method described in the first to fourth aspects above.

[0050] As a result, model training can be conveniently and effectively performed for low-resource language NER, thereby achieving better NER effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.

[0052] Figure 1 Schematic diagram of the system architecture of the recognition model according to the present disclosure.

[0053] Figure 2 It is a schematic flowchart of the model training method according to the present disclosure.

[0054] Figure 3 is a schematic block diagram of a NER model training device that can be used to implement the model training method according to the present disclosure.

[0055] Figure 4 It is a schematic flow chart of the detailed process of the training steps.

[0056] Figure 5 yes Figure 3 Schematic block diagram of the training device in .

[0057] Figure 6 is a schematic flowchart of a named entity recognition method according to the present disclosure.

[0058] Figure 7 is a schematic block diagram of a NER device that can be used to implement the named entity recognition method according to the present disclosure.

[0059] Figure 8 It is a schematic flowchart of a method for named entity recognition according to an embodiment of the present disclosure.

[0060] Fig. 9 is a schematic block diagram of a NER device that can be used to implement a named entity recognition method according to an embodiment of the present disclosure.

[0061] Fig.10 The present invention is a schematic flowchart of a method for setting category labels for word segmentation of product titles.

[0062] Fig.11It is a schematic block diagram of a device that can be used to implement the category label setting method according to the present disclosure.

[0063] Fig.12 A schematic diagram of the structure of a computing device that can be used to implement the above method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0065] The inventors of the present disclosure have noticed that while tag-level label annotation is expensive, sentence-level labels are usually easier to obtain. For example, in the field of e-commerce, the product titles in the above example belong to the category of "electronic devices" specified by the seller. We can use the product categories as noisy sentence-level labels and perform standard text classification.

[0066] In this way, we can consider transferring useful information from sentence-level classification labels to improve token-level NER.

[0067] The inventor proposes that the idea of ​​multi-task learning can be used to implement the training of the NER model.

[0068] Multi-task learning is a technique that enables a model to generalize better to the target task by sharing representations or layers between related tasks.

[0069] In this way, the NER task can be used as the main task, and the sentence classification task can be used as the auxiliary task. The NER model and the sentence classification model can be jointly trained in a multi-task learning manner.

[0070]

Identification model

[0071] Below, reference Figure 1 The system architecture of the recognition model of the present disclosure is described.

[0072] Figure 1 The system architecture of the recognition model according to the present disclosure is schematically shown.

[0073] like Figure 1 As shown, the recognition model of the present disclosure can be regarded as a combination of a NER model and a sentence classification model.

[0074] like Figure 1The recognition model of the present disclosure may include two main components: a shared layer and a task-specific layer.

[0075] The NER model and the sentence classification model share the same layers. At the same time, the NER model and the sentence classification model each have their own task-specific layers.

[0076] In other words, the NER task and sentence classification task share the shared layers, and their respective task-specific layers are used for the NER task and sentence classification task respectively.

[0077] The shared layers may include, for example, a word embedding layer, a projection layer, a bidirectional long short-term memory (BiLSTM) layer, and an attention layer, and output the final hidden representation H'.

[0078] The task-specific layers for the NER task may include a conditional random field layer (CRF), which obtains the NER result based on the final hidden representation.

[0079] The task-specific layers for the sentence classification task may include a pooling layer and a linear layer. The pooling layer performs pooling on the final hidden representation to obtain the input of the linear layer, and the linear layer outputs the sentence classification result.

[0080] The following describes in detail the shared layer in the recognition model of the present disclosure.

[0081] The word embedding layer represents the words in the input word segmentation sequence as corresponding word embedding vectors. Its input is the word segmentation sequence obtained after word segmentation of the training sample or the recognition object text (i.e., "sentence"). Here, each word in the word segmentation sequence is used as a token.

[0082] Let w1, w2, w3, ..., w T is the input token / word sequence, where T is the sentence length, or T is the number of tokens / words contained in the sentence. Figure 1 Medium T=5.

[0083] The word embedding layer can use pre-trained word embedding vectors e t To represent each w t , 1≤t≤T.

[0084] The word embedding layer can be a pre-trained word embedding layer. For example, the publicly available fastText pre-trained word embedding layer can be used.

[0085] You can not t is fine-tuned, and a projection layer is used to project it to the new space x t .

[0086] The projection layer embeds the word vector e t Projection is performed to obtain the input vector X=[x1, x2, x3, …, x T ].

[0087] X = [x1, x2, x3, ..., x T ] is fed to the BiLSTM (bidirectional long short-term memory) network to obtain the hidden representation H corresponding to the word segmentation = [h1, h2, h3, ..., h T ].

[0088]

[0089] d h It is a hidden representation h t Dimension.

[0090] The attention layer applies an attention mechanism to H to help the recognition model focus on specific tags. This process can be used to obtain the final hidden representation H′ corresponding to the word segmentation through the following formula:

[0091] H′=concat(head1,...,head n )W O +H,

[0092] head j =attention(Q j , K j , V j ),

[0093]

[0094] Where H is the hidden representation, H′ is the final hidden representation, H′ = [h1′, h2′, h3′, …, h T ′]. n is the number of heads in the self-attention mechanism, j is the serial number of the corresponding head, 1≤j≤n, W O is a trainable weight matrix, concat() is the connection function, and attention() is the attention function.

[0095] Here, the self-attention mechanism can be used. The self-attention mechanism is a technique that allows the model to pay attention to (i.e. assign more weight to) different tokens of the input sentence.

[0096] The attention function can be calculated by the following formula:

[0097]

[0098]

[0099] Among them, w is the trainable weight vector, d h is the dimension of the hidden representation.

[0100] Q, K, V are the above Q j , K j 、V j The matrix obtained by integration. T represents the matrix transpose.

[0101] The softmax function can be regarded as a generalization of the Sigmoid function, which can be used for multi-classification. The function expression of the softmax function of the i-th element in the array Z can be expressed as:

[0102]

[0103] The mathematical expression of the softplus function can be expressed as:

[0104] softplus(x)=log(1+e x )

[0105] Since the output values ​​generated by the softplus function range from (0 to ∞), the range of the tth element of δ is

[0106] The recognition model of the present disclosure uses the scale factor δ obtained through learning. In this way, the recognition model can dynamically adjust the scale factor δ without increasing a large amount of computational cost.

[0107] As described above, the recognition model of the present disclosure may include a new variation of the multi-head self-attention mechanism, which uses a learned scaling factor δ to enable the model to control the distribution of attention to multiple tokens in a sentence. The learned scaling factor can be simply calculated using a linear transformation without adding a lot of computational cost.

[0108] The task specific layers are described below.

[0109] The task-specific layer for the NER task may include a conditional random field layer (CRF). The conditional random field layer obtains the probability of the label sequence y based on the final hidden representation H′, y1, y3, y3, ..., y t , thus obtaining the result of named entity recognition.

[0110] For example, "B-PRODUCT" and "E-PRODUCT" mean "product" (for example, they can respectively indicate the beginning and end of the "product" label); "S-MATERIAL" means "material"; "S-PATTERN" means "style": "o" means no label.

[0111] Task-specific layers for sentence classification tasks may include pooling layers and linear layers.

[0112] The pooling layer performs max pooling on the final hidden representation H′ to create a fixed-size global vector as input to the linear layer. This allows the model to capture the most useful local features encoded in the hidden layer states.

[0113] The pooling layer feeds this fixed-size global vector to the linear layer. The linear layer s obtains the unstandardized prediction score for each category based on this vector, thereby outputting the sentence classification result. For example, "ELECTRONICS" represents the category "electronic equipment".

[0114]

Model training

[0115] As described above, the present disclosure uses the NER task as the main task and the sentence classification task as the auxiliary task, and trains the recognition model in a multi-task learning manner.

[0116] The recognition model can be viewed as a combination of a NER model and a sentence classification model, with shared layers common to both NER and sentence classification tasks and task-specific layers for both NER and sentence classification tasks, respectively.

[0117] Figure 2 It is a schematic flowchart of the model training method according to the present disclosure.

[0118] Figure 3 is a schematic block diagram of a model training device 300 that can be used to implement the model training method according to the present disclosure.

[0119] like Figure 2 As shown, in step S210, for example, a sentence with a sentence classification label can be obtained by the first obtaining device 310 as a first training sample.

[0120] In the example of an e-commerce scenario where product titles are sentences, sellers often add classification tags to product titles. For example, a product titled "wireless headphones" is tagged "electronic equipment". In this way, titles with sentence classification tags (sentence-level tags) are readily available in large quantities.

[0121] In step S220, for example, by the second acquisition device 320, sentences with NER labels and sentence classification labels are acquired as second training samples.

[0122] Sentences with NER labels (tags / word-level labels) tend to be fewer.

[0123] The number of the first training samples is much larger than that of the second training samples, because the first training samples (eg, product titles with their categories) are readily available, while the second training samples need to be manually labeled.

[0124] This is also the starting point for the present disclosure to train the recognition model in a multi-task learning manner.

[0125] In step S230, for example, by using the first training sample and the second training sample through the training device 330, the recognition model is trained in a multi-task learning manner with the NER task as the main task and the sentence classification task as the auxiliary task.

[0126] Figure 4 It is a schematic flowchart of the detailed process of the above training step S230.

[0127] Figure 5 is a schematic block diagram of the training device 330 .

[0128] In step S230, the first training sample may be used to pre-train the sentence classification model, and then the second training sample may be used to train the entire recognition model.

[0129] like Figure 4 As shown, in step S232, for example, by the first training device 332, the sentence classification model is pre-trained using the first training sample to obtain pre-trained shared layer parameters, and the sentence classification model includes a shared layer and a task-specific layer for the sentence classification task.

[0130] Through the pre-training here, all the model parameters of the shared layer that have been pre-trained can be generated.

[0131] These pre-trained model parameters are then used to initialize shared layers including projection layers and BiLSTM layers.

[0132] In step S234, the recognition model is further trained using the second training samples, for example, by the second training device 334.

[0133] The training operation of step S234 may be repeated until convergence. Convergence may be determined by observing the performance of NER on the sample data. In each cycle, the joint loss function may be calculated And stochastic optimization methods can also be used to update model parameters.

[0134] To train the sentence classification model, the following multi-class cross-entropy loss function is used: Minimize:

[0135]

[0136] Where i represents the sentence index, N is the number of training samples, K is the number of target categories, and s k is the normalized prediction score of the kth target class after applying the softmax function, and t is the one-hotencoded true label.

[0137] For NER, H’ is fed into a CRF layer to obtain the probability p of the label sequence y.

[0138] To train the NER model, the negative log-likelihood function of the correct label sequence relative to the training set is Minimize:

[0139]

[0140] Among them, y represents the label sequence, p(y (i) |H' (i) ) is based on the final hidden representation H' corresponding to the i-th sentence (i) The probability of the obtained label sequence y.

[0141] Combination and Get the joint loss function

[0142]

[0143] Wherein, λ is a balancing parameter. In one embodiment, λ can be simply set to 1.

[0144] here, It can be used as a regularization term (regularization is a technique to deal with overfitting that occurs during model training), which helps reduce overfitting in NER tasks.

[0145] During the training of the NER model and the sentence classification model, the joint loss function is minimized.

[0146] As mentioned above, the present disclosure is not only Figure 1As shown in the figure, the sentence classification model and the NER model are jointly trained, and a large number of training samples with only sentence labels are used to pre-train the sentence classification model. The pre-trained hidden representation will help the recognition model generalize better in the NER task.

[0147] Therefore, the recognition model obtained by combining the NER model and the sentence classification model is conveniently and effectively trained.

[0148] The following describes a named entity recognition scheme according to the present disclosure.

[0149] Named Entity Recognition

[0150] Figure 6 is a schematic flowchart of a named entity recognition method according to the present disclosure.

[0151] Figure 7 is a schematic block diagram of a NER device 700 that can be used to implement the named entity recognition method according to the present disclosure.

[0152] like Figure 6 As shown, in step S610, for example, through the model training device 300, the recognition model is trained in a multi-task learning manner with the NER task as the main task and the sentence classification task as the auxiliary task.

[0153] As described above, the recognition model has shared layers common to the NER task and the sentence classification task and task-specific layers for the NER task and the sentence classification task, respectively.

[0154] The model training device 300 here can be Figure 3 The model training device 300 shown in FIG. 1 , correspondingly, the training step of step S610 can refer to Figure 2 and Figure 4 The steps described are performed.

[0155] Then, in step S620, for example, through the recognition device 400, the text can be input into the trained recognition model to obtain the corresponding NER result.

[0156] Here, the text can be segmented to obtain a segmented word sequence, and then the segmented word sequence is input into the word embedding layer in the shared layer.

[0157] The NER solution disclosed in the present invention can be applied to e-commerce scenarios, for example. The product title set by the seller for the product is used as the input text (i.e., a sentence). Through named entity recognition (NER), the category label corresponding to each token / segmentation in the product title is obtained.

[0158] In this way, a list of relevant information can be generated for each product.

[0159] For example, if there is a text segment "…flexible silicone shell…" in the product title, the following list of information can be obtained:

[0160] Style: Flexible;

[0161] Material: Silicone;

[0162] Product: Shell.

[0163] In addition, as a NER solution, after the model training is completed, the auxiliary task, namely the sentence classification task, can no longer be performed. In this way, only the part of the recognition model used for the NER task, namely the NER model, can be retained.

[0164] Figure 8 It is a schematic flowchart of a method for named entity recognition according to an embodiment of the present disclosure.

[0165] Fig. 9 is a schematic block diagram of a NER device 900 that can be used to implement a named entity recognition method according to an embodiment of the present disclosure.

[0166] like Figure 8 As shown, in step S810, a NER model is provided, for example, by preparing device 910. The NER model here is obtained by training the NER model and the sentence classification model in a multi-task learning manner, using the NER task as the main task and the sentence classification task as the auxiliary task, as described above. The NER model and the sentence classification model have a common shared layer and respective task-specific layers.

[0167] In this way, a trained recognition model combining the NER model and the sentence classification model can be obtained. In the subsequent NER recognition process, only the NER model can be used.

[0168] Therefore, in step S820, for example, through the recognition device 400, the text can be input into the NER model to obtain the corresponding NER result.

[0169] As described above, the NER solution of the present disclosure can be used to set category labels for the word segmentation of product titles.

[0170] Fig.10 The present invention is a schematic flowchart of a method for setting category labels for word segmentation of product titles.

[0171] Fig.11 1 is a schematic block diagram of a category label setting device 1100 that can be used to implement the category label setting method according to the present disclosure.

[0172] like Fig.10 As shown, in step S1010, a category recognition model is provided, for example, by preparing device 1110.

[0173] The category recognition model can be a NER model that takes product titles as input text or training data, i.e., sentences. The NER model sets category labels (e.g., "materials") for tokens / segmented words (e.g., "silicone") in the product title.

[0174] As described above, the category recognition task can be used as the main task, the sentence classification task can be used as the auxiliary task, and the category recognition model and the sentence classification model can be trained in a multi-task learning manner.

[0175] Likewise, the category recognition model and sentence classification model have common shared layers and their own task-specific layers.

[0176] In step S1020, for example, the word segmentation device 1120 is used to segment the product title to obtain a word segmentation sequence.

[0177] For example, the text segment "...flexible silicone shell..." of the product title can be divided into the segment words "flexible", "silicone", and "shell".

[0178] Then in step S1030, for example, through the recognition device 1130, the word segmentation sequence, such as "flexible, silicone, shell", is input into the category recognition model to obtain the category label corresponding to each word in the word segmentation sequence.

[0179] For example, the category label for "flexible" is "style", the category label for "silicone" is "material", and the category label for "shell" is "product".

[0180] In this way, a list of relevant information can be generated for each product.

[0181] For example, if there is a text segment "…flexible silicone shell…" in the product title, the following list of information can be obtained:

[0182] Style: Flexible;

[0183] Material: Silicone;

[0184] Product: Shell.

[0185] Here, the category label may be a category label in a predetermined category label set. For example, the probability that each tag / segmentation corresponds to each tag in the predetermined label set may be calculated to determine which category label each tag / segmentation corresponds to.

[0186] So far, the system architecture, training scheme and corresponding recognition scheme of the recognition model according to the present disclosure have been described in detail.

[0187] In this way, an information classification list can be generated quickly and easily.

[0188] As described above, the present disclosure does not rely on artificially generated features to produce output.

[0189] In addition, the present disclosure adopts multi-task learning (MTL) to fully utilize the training signals of the auxiliary task (sentence classification).

[0190] Furthermore, the present disclosure supports multi-class classification which can be applied to NER.

[0191] In addition, in the multi-head self-attention mechanism, the present disclosure uses a scaling factor obtained through learning.

[0192] The disclosed system includes a novel neural network architecture and a training algorithm thereof for low-resource NER. In particular, the neural network architecture and the training algorithm thereof are based on a combination of (1) multi-task learning and (2) pre-training.

[0193] According to the first aspect, the neural network architecture has shared layers and task-specific layers. Sharing these layers between the two tasks (NER and sentence classification) helps the neural network generalize better and prevents overfitting on low-resource NER tasks.

[0194] According to the second aspect, the training algorithm uses the pre-trained model parameters obtained from the auxiliary sentence classification task to initialize the shared layers such as the shared word projection layer and the BiLSTM layer. Compared with the scheme using random initialization, the use of pre-trained model parameters provides a better starting point for the training process.

[0195] The training algorithm can also fine-tune the model parameters by optimizing a joint loss function.

[0196] Fig.12 A schematic diagram of the structure of a computing device that can be used to implement the above method according to an embodiment of the present invention is shown.

[0197] See also Fig.12 , the computing device 1200 includes a memory 1210 and a processor 1220 .

[0198] The processor 1220 may be a multi-core processor or may include multiple processors. In some embodiments, the processor 1220 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), etc. In some embodiments, the processor 1220 may be implemented using a customized circuit, such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0199] The memory 1210 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, ROM can store static data or instructions required by the processor 1220 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even after the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as a permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as a dynamic random access memory. The system memory may store some or all instructions and data required by the processor at run time. In addition, the memory 1210 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the memory 1210 may include a readable and / or writable removable storage device, such as a laser disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and transient electronic signals transmitted wirelessly or wired.

[0200] The memory 1210 stores executable codes, and when the executable codes are processed by the processor 1220 , the processor 1220 can execute the identification, setting and training methods mentioned above.

[0201] The identification, setting and training according to the present invention have been described above in detail with reference to the accompanying drawings.

[0202] In addition, the method according to the present invention may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.

[0203] Alternatively, the present invention may also be implemented as a non-temporary machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor executes the various steps of the above-mentioned method according to the present invention.

[0204] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or a combination of both.

[0205] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system and method according to multiple embodiments of the present invention. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0206] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for named entity recognition, comprising: Taking the named entity recognition task as the main task and the sentence classification task as the auxiliary task, the recognition model is trained in a multi-task learning manner, wherein the recognition model has a shared layer shared by the named entity recognition task and the sentence classification task and task-specific layers respectively used for the named entity recognition task and the sentence classification task; as well as Input the text into the trained recognition model to obtain the corresponding named entity recognition results. The step of training the recognition model includes: Pre-training a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and The recognition model is trained using training samples with named entity recognition labels and sentence classification labels. The recognition model includes the sentence classification model and the named entity recognition model. The named entity recognition model includes the shared layer and a task-specific layer for the named entity recognition task. The task-specific layer for the named entity recognition task receives the output of the shared layer and obtains a result of named entity recognition.

2. The method for named entity recognition according to claim 1, wherein: To train a sentence classification model, the following multi-class cross entropy loss function is used Minimize: Where i represents the sentence index, N is the number of training samples, K is the number of target categories, and s k is the normalized prediction score of the kth target class after applying the softmax function, and t is the one-hot encoded true label, To train a named entity recognition model, the negative log-likelihood function of the correct label sequence relative to the training set is Minimize: Among them, y represents the label sequence, p(y (i) |H' (i) ) is based on the final hidden representation H' corresponding to the i-th sentence (i) The probability of the obtained label sequence y, Combination and Get the joint loss function Where λ is the equilibrium parameter, In the process of training the recognition model, the joint loss function is minimized.

3. The method for named entity recognition according to claim 1, wherein: The shared layers include at least one of the following: a word embedding layer, a projection layer, a BiLSTM layer, an attention layer, and output a final hidden representation; and / or The task-specific layer for the named entity recognition task includes a conditional random field layer, which obtains the named entity recognition result based on the final hidden representation; and / or The task-specific layers for the sentence classification task include a pooling layer and a linear layer. The pooling layer pools the final hidden representation to obtain the input of the linear layer, and the linear layer outputs the sentence classification result.

4. The method for named entity recognition according to claim 3, wherein: The word embedding layer is a pre-trained word embedding layer; and / or The input of the word embedding layer is a word segmentation sequence obtained after word segmentation processing of the training sample or the text; and / or The word embedding layer represents the words in the input word sequence as corresponding word embedding vectors; and / or The projection layer projects the word embedding vector to obtain the input vector corresponding to the word segmentation of the BiLSTM layer; and / or The BiLSTM layer outputs hidden representations corresponding to the tokens; and / or The attention layer applies attention to the hidden representation to obtain the final hidden representation corresponding to the word segmentation; and / or The pooling layer performs max pooling on the final hidden representation to create a fixed-size global vector as input to the linear layer; and / or The linear layer obtains the prediction score for each category based on the fixed-size global vector output by the pooling layer.

5. The method for named entity recognition according to claim 4, wherein: The attention layer obtains the final hidden representation through the following formula: H′=concat(head1,..,head n )W O +H, head j =attention(Q j ,K j ,V j ), Where H is the hidden representation, H' is the final hidden representation, n is the number of heads in the self-attention mechanism, j is the sequence number of the corresponding head, 1≤j≤n, W j Q , W j K , W j V , W O is the weight matrix, concat() is the connection function, and attention() is the attention function.

6. The method for named entity recognition according to claim 5, wherein: The attention function is calculated by the following formula: Where w is the weight vector, d h is the dimension of the hidden representation.

7. The method for named entity recognition according to claim 1, wherein: The step of inputting text into the trained recognition model includes: Performing word segmentation processing on the text to obtain a word segmentation sequence; The word segmentation sequence is input into the word embedding layer in the shared layer.

8. The method for named entity recognition according to claim 1, wherein: The method is applied in an e-commerce context; and / or The text is a product title; and / or The named entity recognition result is the category corresponding to each word in the product title.

9. A method for named entity recognition, comprising: A named entity recognition model is provided, wherein the named entity recognition model is obtained by training the named entity recognition model and the sentence classification model in a multi-task learning manner using a named entity recognition task as a main task and a sentence classification task as an auxiliary task, wherein the named entity recognition model and the sentence classification model have a common shared layer and respective task-specific layers; as well as Inputting text into the named entity recognition model to obtain corresponding named entity recognition results, wherein the step of training the named entity recognition model and the sentence classification model includes: Pre-training a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and The sentence classification model and the named entity recognition model are trained using training samples with named entity recognition labels and sentence classification labels. The named entity recognition model includes the shared layer and a task-specific layer for the named entity recognition task. The task-specific layer for the named entity recognition task receives the output of the shared layer and obtains a result of named entity recognition.

10. A method for setting a category label for a word segmentation of a product title, comprising: Providing a category recognition model, wherein the category recognition model is obtained by training the category recognition model and the sentence classification model in a multi-task learning manner using the category recognition task as a main task and the sentence classification task as an auxiliary task, wherein the category recognition model and the sentence classification model have a common shared layer and respective task-specific layers; Perform word segmentation processing on the product title to obtain a word segmentation sequence; as well as Input the word segmentation sequence into the category recognition model to obtain the category label corresponding to each word in the word segmentation sequence, The step of training the category recognition model and the sentence classification model includes: Pre-training a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and The sentence classification model and the named entity recognition model are trained using training samples with category identification labels and sentence classification labels. The named entity recognition model includes the shared layer and a task-specific layer for the named entity recognition task. The task-specific layer for the named entity recognition task receives the output of the shared layer and obtains a result of named entity recognition.

11. The method according to claim 10, wherein: The category label is a category label in a predetermined category label set.

12. A method for training a named entity recognition model, comprising: Obtain a sentence with a sentence classification label as the first training sample; Obtaining sentences with named entity recognition labels and sentence classification labels as second training samples; as well as Using the first training sample and the second training sample, the recognition model is trained in a multi-task learning manner with the named entity recognition task as the main task and the sentence classification task as the auxiliary task. The recognition model has a shared layer for the named entity recognition task and the sentence classification task and task-specific layers for the named entity recognition task and the sentence classification task respectively. The step of training the recognition model includes: Pre-training a sentence classification model using the first training sample to obtain pre-trained shared layer parameters, the sentence classification model comprising the shared layer and a task-specific layer for a sentence classification task, the task-specific layer for the sentence classification task receiving an output of the shared layer and outputting a sentence classification result; and The recognition model is trained using a second training sample, wherein the recognition model includes the sentence classification model and the named entity recognition model, wherein the named entity recognition model includes the shared layer and a task-specific layer for a named entity recognition task, wherein the task-specific layer for the named entity recognition task receives an output of the shared layer and obtains a result of named entity recognition.

13. The method according to claim 12, wherein: To train a sentence classification model, the following multi-class cross entropy loss function is used Minimize: Where i represents the sentence index, N is the number of training samples, K is the number of target categories, and s k is the normalized prediction score of the kth target class after applying the softmax function, and t is the one-hot encoded true label, To train the NER model, the negative log-likelihood function of the correct label sequence relative to the training set is Minimize: Among them, y represents the label sequence, p(y (i) |H' (i) ) is based on the final hidden representation H' corresponding to the i-th sentence (i) The probability of the obtained label sequence y, Combination and Get the joint loss function Where λ is the equilibrium parameter, During the training of the named entity recognition model and the sentence classification model, the joint loss function is minimized.

14. A named entity recognition device, comprising: A model training device, used to train a recognition model using a named entity recognition task as a main task and a sentence classification task as an auxiliary task in a multi-task learning manner, wherein the recognition model has a shared layer shared by the named entity recognition task and the sentence classification task and task-specific layers respectively used for the named entity recognition task and the sentence classification task; as well as The recognition device is used to input the text into the trained named entity recognition model to obtain the corresponding named entity recognition result. Wherein, the model training device comprises: A first training device pre-trains a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and The second training device uses training samples with named entity recognition labels and sentence classification labels to train the recognition model, wherein the recognition model includes the sentence classification model and the named entity recognition model, and the named entity recognition model includes the shared layer and a task-specific layer for the named entity recognition task, and the task-specific layer for the named entity recognition task receives the output of the shared layer and obtains a result of named entity recognition.

15. A named entity recognition device, comprising: Preparing a device for providing a named entity recognition model, wherein the named entity recognition model is obtained by training the named entity recognition model and the sentence classification model in a multi-task learning manner using a named entity recognition task as a main task and a sentence classification task as an auxiliary task, wherein the named entity recognition model and the sentence classification model have a common shared layer and respective task-specific layers; as well as The recognition device is used to input the text into the named entity recognition model to obtain the corresponding named entity recognition result, The training of the named entity recognition model and the sentence classification model includes: Pre-training a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and The sentence classification model and the named entity recognition model are trained using training samples with named entity recognition labels and sentence classification labels. The named entity recognition model includes the shared layer and a task-specific layer for the named entity recognition task. The task-specific layer for the named entity recognition task receives the output of the shared layer and obtains a result of named entity recognition.

16. A device for setting category labels for word segmentation of a product title, comprising: Preparing a device for providing a category recognition model, wherein the category recognition model is obtained by training the category recognition model and the sentence classification model in a multi-task learning manner using the category recognition task as a main task and the sentence classification task as an auxiliary task, wherein the category recognition model and the sentence classification model have a common shared layer and respective task-specific layers; A word segmentation device, used for performing word segmentation processing on the product title to obtain a word segmentation sequence; as well as The recognition device is used to input the word segmentation sequence into the category recognition model to obtain the category label corresponding to each word segmentation in the word segmentation sequence. The training of the category recognition model and the sentence classification model includes: Pre-training a sentence classification model using training samples with sentence classification labels to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and The sentence classification model and the category recognition model are trained using training samples with category recognition labels and sentence classification labels. The category recognition model includes the shared layer and a task-specific layer for the category recognition task. The task-specific layer for the category recognition task receives the output of the shared layer and obtains a result of category recognition.

17. A device for training a named entity recognition model, comprising: A first acquisition device is used to acquire a sentence with a sentence classification label as a first training sample; A second acquisition device is used to acquire sentences with named entity recognition labels and sentence classification labels as second training samples; as well as The model training device is used to use the first training sample and the second training sample, take the named entity recognition task as the main task, take the sentence classification task as the auxiliary task, and train the recognition model in a multi-task learning manner, The recognition model has a shared layer for the named entity recognition task and the sentence classification task and task-specific layers for the named entity recognition task and the sentence classification task respectively. Wherein, the model training device comprises: a first training device, which pre-trains a sentence classification model using a first training sample to obtain pre-trained shared layer parameters, wherein the sentence classification model includes the shared layer and a task-specific layer for a sentence classification task, wherein the task-specific layer for the sentence classification task receives an output of the shared layer and outputs a sentence classification result; and A second training device uses a second training sample to train the recognition model, wherein the recognition model includes the sentence classification model and the named entity recognition model, and the named entity recognition model includes the shared layer and a task-specific layer for the named entity recognition task, and the task-specific layer for the named entity recognition task receives the output of the shared layer and obtains a result of named entity recognition.

18. A computing device comprising: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 13.

19. A non-transitory machine-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • A sentence backbone analysis method and system based on multi-task depth neural network of character segmentation and named entity recognition

    CN109255119A

  • Medical named entity identification method based on medical dictionary

    CN110569506A