Named entity recognition device and program

The named entity recognition device simplifies calculations and enhances entity classification by using a model processing unit and attention mechanism to improve the accuracy of identifying named entities, addressing the limitations of existing methods.

JP2025110032APending Publication Date: 2025-07-28NIPPON HOSO KYOKAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024003720
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-28

AI Technical Summary

Technical Problem

Existing named entity recognition methods, such as sequential labeling and span-based models, face challenges in accurately classifying nested structured entities and require complex calculations, leading to suboptimal performance.

Method used

A named entity recognition device utilizing a model processing unit, embedding unit, attention mechanism unit, and feed-forward network to convert token sequences into vector sequences, calculate similarities, and determine entity types with weighted sums, allowing for improved classification without complex calculations.

Benefits of technology

The device achieves high-performance named entity extraction by simplifying calculations while accurately identifying and categorizing named entities, including person names, place names, and organization names, outperforming existing methods in certain datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110032000001_ABST
    Figure 2025110032000001_ABST
Patent Text Reader

Abstract

To provide a named entity recognition device and a program capable of improving the performance of named entity recognition without requiring complicated calculations.SOLUTION: A named entity recognition device is provided with: a model processing unit for converting a token string corresponding to a partial word string into a vector string; an embedding unit for outputting a task-type value; an attention mechanism unit for, with the token string inputted as a value (V) and a key (K), and the task-type value inputted as a query (Q), calculating a similarity between the key and the query, calculating a weighted sum of the values using the similarity as a weight, and outputting the weighted sum as estimation information about the partial word string; and a feed-forward network for calculating a score of a label indicating whether or not the partial word string is a named entity based on the estimation information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a named entity recognition device and a program.

Background Art

[0002] There is a technique for extracting named entities from natural language text. This technique is called named entity extraction and is one of the basic techniques in natural language processing. A named entity is a general term for proper nouns such as personal names, organization names, and place names, as well as date expressions and time expressions.

[0003] As methods for named entity recognition according to the prior art, a method called sequential labelling and a method using a span-based model can be mentioned.

[0004] In the sequential labelling method, words are vectorized, and a column of those vectors is input into an LSTM (Long Short Term Memory) or the like, and a label is assigned to each word. When a named entity is a sequence of multiple words, labels representing the start, middle, and end of the named entity are assigned. When a named entity is a single word, a label indicating that it is a single-word named entity is assigned.

[0005] In the span-based model method, based on a column of vectors output corresponding to a sequence of words, by collecting subsequences of that column into one vector, it is estimated whether the subsequence is a named entity or not.

[0006] Non-Patent Document 1 describes a method for performing named entity extraction using a span-based model. In the technique described in Non-Patent Document 1, mainly MaxPooling is used for the part of collecting the outputs from the pre-trained model.

[0007] Non-Patent Document 2 describes a technique for labeling words using a sequential labeling method. The text targeted in Non-Patent Document 2 is the text included in medical documents.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0009] However, the sequential labeling method has a problem that it cannot finely classify nested structured named entities. For example, when the expression "NHK Broadcasting Technology Research Institute" is given, the sequential labeling method can only extract named entities of either "NHK" or "NHK Broadcasting Technology Research Institute". Also, when using LSTM, good performance could not be achieved.

[0010] In addition, the method using a span-based model is said to have good recognition performance, but there is a problem that the calculation tends to be complicated. In other words, there is room for improvement in the summarization processing part of the method using a span-based model.

[0011] The present invention has been made in consideration of the above circumstances, and aims to provide a named entity recognition device and a program that do not require complicated calculations as in the method using a span-based model and can improve the performance of named entity extraction.

Means for Solving the Problem

[0012] [1] To solve the above problems, a named entity recognition device according to an aspect of the present invention includes a model processing unit that converts a token sequence corresponding to a partial word sequence, which is a part of a text, into a vector sequence, an embedding unit that outputs a task type value corresponding to the type of task, an attention mechanism unit that inputs the token sequence output from the model processing unit as a value (V) and a key (K), inputs the task type value output from the embedding unit as a query (Q), calculates the similarity between the key and the query, calculates a weighted sum of the values with the similarity as a weight, and outputs the weighted sum as estimation information about the partial word sequence, and a feed-forward network that calculates a score of a label indicating whether the partial word sequence is a named entity based on the estimation information.

[0013] [2] Also, one aspect of the present invention is configured such that, in the named entity recognition device of [1] above, the internal parameters of the model of the embedding unit, the attention mechanism unit, and the feed-forward network can be adjusted based on a set of pairs of an input token sequence and a correct label corresponding to the token sequence, which is learning data.

[0014] [3] Also, one aspect of the present invention is that, in the named entity recognition device of [1] or [2] above, the label represents whether the partial word sequence is a named entity, and further, when the partial word sequence is a named entity, it represents which type of named entity the partial word sequence is.

[0015] [4] Also, one aspect of the present invention is that, in the named entity recognition device of [3] above, the type of named entity includes at least a person name, a place name, and an organization name.

[0016] [5] Also, one aspect of the present invention is that, in the named entity recognition device of any one of [1] to [4] above, it further includes a post-processing unit that determines a label to be assigned to the partial word sequence based on the scores of each of the plurality of labels.

[0017] [6] Also, one aspect of the present invention is a model processing unit that converts a token sequence corresponding to a partial word sequence, which is a text part, into a vector sequence, an embedding unit that outputs a task type value corresponding to the type of task, and inputs the token sequence output from the model processing unit as a value (V) and a key (K), inputs the task type value output from the embedding unit as a query (Q), calculates the similarity between the key and the query, calculates the weighted sum of the values with the similarity as the weight, and outputs the weighted sum as estimation information about the partial word sequence. An attention mechanism unit, and a feed-forward network that calculates a score of a label indicating whether the partial word sequence is a named entity based on the estimation information. A program for causing a computer to function as a named entity recognition device including

Advantages of the Invention

[0018] According to the present invention, the named entity recognition device can calculate a score of a label indicating whether a partial word sequence is a named entity based on a token sequence corresponding to the partial word sequence by relatively simple calculations.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0020] Next, an embodiment of the present invention will be described with reference to the drawings. The named entity recognition device 1 according to this embodiment performs high-performance named entity recognition (NER, Named Entity Recognition, which may also be called "named entity extraction") from natural language text.

[0021] The natural language text input to the named entity recognition device 1 is a sequence of words. Each word is converted into a token, and further converted into a vector for each token by the pre-trained model processing unit 22. Thereafter, the named entity recognition device 1 aggregates the vectors of named entity candidates composed of a plurality of words in the input text by the attention mechanism unit 26, and further determines whether it is a named entity by a neural network (feed-forward network 28). The characteristic configuration of this embodiment is the configuration of the part that aggregates the vectors of named entity candidates by the attention mechanism described above. That is, this embodiment is a device that devises the aggregation part of the span-based model.

[0022] In this embodiment, in the above-mentioned "summarization" process, weighted sum calculation using an attention mechanism is performed. As a result, the performance is improved compared to the summarization in the prior art.

[0023] Max Pooling, which was used in the summarization in the prior art, is a method of obtaining only the maximum value of vectors in a word sequence, and the importance of words was not considered in that method.

[0024] Also, although the importance of words should be considered in the LSTM of the prior art, it is considered that the performance of named entity extraction did not improve so much probably because the influence of word order is too strong.

[0025] Figure 1 is a block diagram showing the schematic functional configuration of the named entity recognition device according to this embodiment. As shown in the figure, the named entity recognition device 1 includes a pre-trained model processing unit 22, an embedding unit 24, an attention mechanism unit 26, a feed-forward network 28, and a post-processing unit 30. Each of these functional units can be realized by, for example, a computer and a program. Also, each functional unit has a storage means as necessary. The storage means is, for example, a variable in a program or a memory allocated by the execution of a program. Also, if necessary, non-volatile storage means such as a magnetic hard disk device or a solid state drive (SSD) may be used. Also, at least a part of the functions of each functional unit may be realized as a dedicated electronic circuit instead of a program.

[0026] The pre-trained model processing unit 22 inputs a token sequence and outputs a vector sequence corresponding to the token sequence. The pre-trained model processing unit 22 can be realized in a machine-learnable form using, for example, a neural network. In the present embodiment, the pre-trained model processing unit 22 is pre-trained. That is, in the learning process of the named entity recognition device 1 described later, the model held by the pre-trained model processing unit 22 only performs adjustment at a low learning rate called fine-tuning or is not updated. Note that the pre-trained model processing unit 22 may be simply referred to as the "model processing unit".

[0027] The token sequence input to the pre-trained model processing unit 22 corresponds to the input sentence text. That is, when the input sentence text is a word sequence "···, w1, w2, ···", the tokens corresponding to the words w1 and w2 are as shown in (1) and (2) below, respectively.

[0028]

Number

[0029]

Number

[0030] Here, |t w1 | and |t w2 | are the lengths of the token sequences corresponding to the words w1 and w2, respectively. Each token can be represented as a vector, for example.

[0031] Note that the named entity recognition device 1 may have a function of converting a word into a token (which may be called a conversion unit). Alternatively, a function of converting a word into a token may be provided inside the pre-trained model processing unit 22.

[0032] The vector sequence output by the pre-trained model processing unit 22 corresponds to the above token sequence. The vector sequences corresponding to the above token sequences (1) and (2) are represented as (3) and (4) below, respectively.

[0033]

Number

[0034]

Number

[0035] That is, the pre-trained model processing unit 22 (model processing unit) converts the input token sequence into a vector sequence. The input token sequence corresponds to a sub-word sequence that is a part of the input text (the text to be subject to named entity extraction).

[0036] The embedding unit 24 inputs a task ID and outputs a predetermined vector (which may be called a "task type value"). The task ID input to the embedding unit 24 is information for identifying a task. In this embodiment, there is only one type of named entity recognition task. The task ID representing named entity recognition may be any value. For example, it may be a numerical value such as 0 or 1, or other values. The vector calculated and output by the embedding unit 24 based on the above task ID serves as a Q (query) input to the attention mechanism unit 26 described later.

[0037] The embedding unit 24 is configured to be machine-learnable using a neural network or the like. In the learning process described later, the internal parameters of the embedding unit 24 are also updated. That is, through learning, the values of the internal parameters of the embedding unit 24 are adjusted so as to output a value of Q suitable for the processing of named entity recognition. That is, the embedding unit 24 outputs a task type value corresponding to the type of task.

[0038] The attention mechanism unit 26 calculates and outputs a vector based on the input values of Q, V, and K. As described above, the Q input to the attention mechanism unit 26 is a query. Also, V is a value. And K is a key. A vector sequence is input to each of V and K in the attention mechanism unit 26. The vector sequences input to V and K are the sequences obtained by concatenating the vector sequence of (3) and the vector sequence of (4). The output from the upper embedding unit 24 is input to Q in the attention mechanism unit 26.

[0039] The vector output by the attention mechanism unit 26 is an estimated value regarding the named entity corresponding to the word sequence w t and is represented by (5) below. That is, the vector of (5) below indicates whether the word sequence w t is a named entity, and also has information indicating what type of named entity (for example, the type such as a person's name, a place name, an organization name, etc.) the word sequence w t is when it is a named entity. In the example shown in FIG. 1, the word sequence w t is a sequence of [word w1 - word w2]. It corresponds to an arbitrary word sequence.

[0040]

Equation

[0041] The attention mechanism unit 26 performs processing by the attention mechanism. The attention mechanism itself is based on the prior art. Specifically, the attention mechanism unit 26 calculates the similarity between Q (query) and K (key), and obtains and outputs the weighted sum of V (value) with that similarity as the weight. In other words, the attention mechanism unit 26 extracts information useful for solving Q (a specific task. In this embodiment, the task of named entity recognition) from V (here, the vector sequence corresponding to the input word sequence. For example, the information of the vector sequence obtained by concatenating the vector sequence of (3) and the vector sequence of (4) above).

[0042] The processing by the attention mechanism is represented by, for example, the following formula (6).

[0043]

Equation

[0044] In formula (6), Q, K, and V are inputs (vectors) to the attention mechanism. QK T is the inner product of Q and K, representing the similarity between the two. The function Softmax() performs calculations to make the sum of the similarities equal to 1 (normalization). The √D in the denominator of the argument of the function Softmax() is given to prevent the vanishing gradient. Here, D is the dimensionality of the vectors Q, K, and V. That is, the Output (output of the attention mechanism) in formula (6) is the weighted sum of V by the similarity.

[0045] That is, the attention mechanism unit 26 inputs the token sequence output from the pre-trained model processing unit 22 as values (V) and keys (K), and also inputs the task type value output from the embedding unit 24 as a query (Q). Then, the attention mechanism unit 26 calculates the similarity between the key and the query, calculates the weighted sum of the values with the similarity as the weight, and outputs the weighted sum as the estimation information about the original sub-word sequence.

[0046] The feed-forward network 28 inputs the vector in (5) above and outputs information on the label scores. The feed-forward network 28 is realized by, for example, a three-layer neural network. The values of the internal parameters of the feed-forward network 28 are adjusted during learning.

[0047] The feed-forward network 28 outputs numerical values of scores for each label. For example, when the labels are four types: "PLACE" (place name), "PERSON" (person name), "ORGANIZATION" (organization name), and "0" (not a named entity), the feed-forward network 28 outputs four numerical values corresponding to each of these labels. The numerical values may be normalized as appropriate. The numerical value corresponding to a label represents the likelihood that the original input word sequence corresponds to that label.

[0048] That is, based on the estimation information output from the attention mechanism unit 26, the feed-forward network 28 calculates a score of a label indicating whether the input partial word sequence is a named entity. Note that the label in the present embodiment indicates whether the partial word sequence is a named entity, and further, when the partial word sequence is a named entity, it indicates which type of named entity the partial word sequence is. The types of named entities may include at least person names, place names, and organization names. Also, the label may represent types of named entities other than these.

[0049] The post-processing unit 30 performs post-processing based on the numerical information output by the feed-forward network 28. Specifically, the post-processing unit 30 determines the label to be assigned to the words included in the input text. That is, the post-processing unit 30 determines the label to be assigned to the input partial word sequence based on the scores of each of the plurality of labels. An example of the specific processing procedure of the post-processing unit 30 will be described later with reference to the flowchart.

[0050] That is, as described above, the named entity recognition device 1 extracts the named entities included in the input text (word sequence) based on the input text.

[0051] Next, the machine learning of the named entity recognition device 1 will be described. The named entity recognition device 1 performs learning using learning data that is a set of pairs of an input token sequence and a correct label corresponding to the token sequence. That is, the named entity recognition device 1 is configured to be able to adjust the internal parameters of the models respectively possessed by the embedding unit 24, the attention mechanism unit 26, and the feed-forward network 28 based on the above learning data. Note that the general machine learning method itself belongs to the prior art.

[0052] FIG. 2 is a schematic diagram showing an example of input data used when performing learning of the named entity recognition device 1. As shown in the figure, the input data is the text of a sentence. The input sentence is a sequence of words. The input data may originally be given as a sequence of words, or a sequence of words may be obtained by performing morphological analysis processing or the like on the input sentence. In the example shown in the figure, the input data is a word sequence of "I" - "am" - "Saitama" - [from] - [is].

[0053] FIG. 3 is a schematic diagram showing learning data created based on the input sentence shown in FIG. 2. As shown in the figure, the learning data in this embodiment is given as a set of pairs of a word sequence and a correct judgment result label. The 15 types of word sequences shown in FIG. 3 are all subsequences obtained based on the word sequence of "I" - "am" - "Saitama" - [from] - [is] shown in FIG. 2. Each row in the table shown in FIG. 3 is conveniently assigned a number from 1 to 15. For example, the subsequence in the first row is "I". Also, the subsequence in the second row is "I" - "am". The same applies to the third row and subsequent rows. Corresponding to each of these subsequences, a correct judgment result label is given. The judgment result label corresponding to a named entity may be, for example, PLACE (place name), PERSON (personal name), ORGANIZATION (organization name), etc. Also, for the word sequence that is not a named entity (negative example in the learning of the named entity recognition device 1) in the tenth row, the judgment result label in this example is set to "0".

[0054] Note that if there are too many negative examples in learning, the efficiency of the learning process will decrease. Therefore, the negative example data may be appropriately reduced. The data shown in FIG. 3 lists all the patterns of the partial word sequences of the input sentences shown in FIG. 2. If the data shown in FIG. 3 is used for learning as it is, the efficiency of the learning process will decrease. Therefore, in this embodiment, for example, about 80% to 90% of the negative example data is randomly selected and deleted.

[0055] FIG. 4 is a schematic diagram showing an example of the learning data obtained as a result of randomly reducing the negative example data. The data shown in FIG. 4 is the learning data obtained as a result of deleting the first row, the third row to the ninth row, and the eleventh row to the fourteenth row from the data shown in FIG. 3.

[0056] The named entity recognition device 1 performs learning using the learning data as described above and optimizes the parameters of each model. However, as described above, the pre-trained model processing unit 22 may be excluded from the learning target. During learning, based on the input word sequence, the output value from the feed-forward network 28 is calculated. The output value from the feed-forward network 28 is the score for each label. The output value from the feed-forward network 28 can be represented as a vector having the dimension of the number of label types (including the label "0"). On the other hand, the learning data has the correct answer of the label. That is, a vector (which may be called a correct answer vector) can be created such that the value corresponding to the label that hits the correct answer is 1 and the values corresponding to the other labels are 0. The loss (error) for learning is calculated as the loss between the vector output by the feed-forward network 28 and the above correct answer vector. As the loss, for example, the cross-entropy of the two score spaces may be used. The named entity recognition device 1 updates the parameter values of the internal model (excluding the pre-trained model processing unit 22) by the error backpropagation method based on the calculated loss.

[0057] FIG. 5 is a schematic diagram showing an example of input data when extracting named entities using the learned named entity recognition device 1. This input data is a word sequence based on the sentence "I am from Saitama".

[0058] FIG. 6 is a schematic diagram showing an example of a list of determination result labels and their scores calculated by the learned named entity recognition device 1 based on a word sequence. Row numbers are attached to the table in this figure for convenience. The 15 word sequences shown in FIG. 6 are all patterns of partial word sequences of the word sequence "I" - "am" - "from" - "Saitama" - "." shown in FIG. 5. The determination result label in each row of the table in this figure is determined as the label with the highest score based on the scores of each label calculated based on the input word sequence.

[0059] That is, the pre-trained model processing unit 22 acquires token sequences corresponding to the partial word sequences in each row, and outputs vectors corresponding to them. The attention mechanism unit 26 outputs v(hat) based on the vectors output from the pre-trained model processing unit 22 and the embedding values output from the embedding unit 24. wt Based on the above v(hat), wt the feed-forward network 28 calculates the scores of each label. Then, the post-processing unit 30 determines the label with the score having the largest value.

[0060] In the example of FIG. 6, the determination result labels for "Saitama" in the 10th row and "Saitama" - "from" in the 11th row are "PLACE" (place name). Also, the determination result labels for the word sequences in the other rows (from the 1st row to the 9th row and from the 12th row to the 15th row) are "0" (not a named entity). Note that the scores corresponding to the labels are values normalized to be 0.0 or more and 1.0 or less. The score of the label "PLACE" corresponding to "Saitama" in the 10th row is 0.997, and the score of the label "PLACE" for "Saitama" - "from" in the 11th row is 0.744. Scores are also calculated for the determination result label "0" in the other rows.

[0061] The post-processing unit 30 extracts candidates for named entities from the results shown in FIG. 6. In other words, the post-processing unit 30 extracts only the rows in the results shown in FIG. 6 where the determination result label is not "0".

[0062] FIG. 7 is a schematic diagram showing the data of the named entity candidates extracted by the post-processing unit 30 from the data shown in FIG. 6. That is, the data shown in FIG. 7 is the result of extracting only the cases where the determination result label is not "0". That is, here, only the data of the 10th row and the 11th row in the data shown in FIG. 6 are extracted.

[0063] The post-processing unit 30 assigns labels to the words based on the data shown in FIG. 7. At that time, the post-processing unit 30 assigns labels in order from the candidate with the highest score. However, when assigning labels in that order, candidates that include words already determined to be named entities are rejected. An example of the processing procedure for this label assignment will be described below with reference to FIG. 8.

[0064] FIG. 8 is a schematic diagram for explaining an example of the procedure in which the post-processing unit 30 assigns labels in order from the candidate with the highest score. As a premise of this processing procedure, the data of the 10th row shown in FIG. 7 (conveniently called candidate 1) and the data of the 11th row (conveniently called candidate 2) are each candidates. Hereinafter, an example of the processing procedure will be described. The input sentence (word sequence) to be processed is "I" - "am" - "from" - "Saitama" - ".".

[0065] In step 0 of FIG. 8, the post-processing unit 30 sets the initial values of the label candidates corresponding to each word. That is, the initial value of all the words included in "I" - "am" - "from" - "Saitama" - "." is set to the label "0".

[0066] Next, in step 1, the post-processing unit 30 attempts to assign candidate 1 in descending order of score. The word sequence of candidate 1 is "Saitama", and its score is 0.997. At that time, the label of the word "Saitama" is "0", and no label indicating that it is a proper name is assigned. Therefore, the post-processing unit 30 assigns the label "PLACE" of candidate 1 to the word "Saitama". That is, in step 1 of FIG. 8, the label of the word "Saitama" is replaced with "PLACE".

[0067] Next, in step 2, the post-processing unit 30 attempts to assign candidate 2, which has the next highest score. The word sequence of candidate 2 is "Saitama" - "origin", and its score is 0.744. At that time, the label of the word "Saitama" is "PLACE", and the label of the word "origin" is "0". That is, regarding candidate 2, a label (a label other than "0") indicating that the word "Saitama" in the word sequence "Saitama" - "origin" is already a proper name has already been assigned. Therefore, the post-processing unit 30 rejects candidate 2. That is, in step 2 of FIG. 8, the labels of the words "Saitama" and "origin" are not updated. That is, the label of the word "Saitama" is "PLACE" (candidate 1), and the label of the word "origin" remains "0".

[0068] Since candidate 1 and candidate 2 are all the candidates, the post-processing unit 30 thus ends the process of label assignment.

[0069] FIG. 9 is a flowchart showing the procedure of the process for determining the label to be assigned to a word based on the data of candidates for proper names (FIG. 7). Hereinafter, the processing procedure will be described according to this flowchart.

[0070] In step S11, the post-processing unit 30 initializes the labels for all the words included in the word sequence corresponding to the input sentence. Note that the initial value is set to "0" (indicating that it is not a proper name). This process corresponds to giving the initial value "0" as the label for each word included in the word sequence "I" - "am" - "Saitama" - "origin" - "am" in the example shown in FIG. 8.

[0071] In step S12, the post-processing unit 30 determines whether there is data of unprocessed specific expression candidates. If there is data of unprocessed specific expression candidates (step S12: YES), the process proceeds to the next step S13. If there is no data of unprocessed specific expression candidates (step S12: NO), the processing of the entire flowchart ends.

[0072] When proceeding to step S13, the post-processing unit 30 sets the candidate with the highest score among the unprocessed candidates as the processing target. The post-processing unit 30 performs the processing from the following step S14 to S16 on this candidate as the processing target.

[0073] In step S14, the post-processing unit 30 determines whether all the words included in the word sequence of the processing target are unassigned labels. If all the words included in the word sequence of the processing target are unassigned labels (step S14: YES), the process proceeds to step S15. If at least a part of the words included in the word sequence of the processing target has a label (a label indicating a specific expression) assigned (step S14: NO), the process proceeds to step S16.

[0074] When proceeding to step S15, the post-processing unit 30 assigns the label that the current candidate as the processing target has to each word included in the word sequence of the processing target. Thereby, the specific expression recognition device 1 extracts (recognizes) the current word sequence of the processing target as a specific expression. After the processing of this step ends, the process returns to step S12.

[0075] When proceeding to step S16, the post-processing unit 30 rejects the current candidate as the processing target. After the processing of this step ends, the process returns to step S12. This search method is generally called a greedy algorithm. Other search methods, such as beam search, may be used.

[0076] According to the example of the post-processing unit 30 described above, the post-processing unit 30 can assign labels to a word sequence based on the output from the feed-forward network 28. Specifically, the post-processing unit 30 can assign a label indicating that the word sequence is a named entity (and a label indicating the type of named entity (such as a person name, a place name, an organization name, etc.)) to the word sequence determined to be a named entity.

[0077] FIG. 10 is a block diagram showing an example of the internal configuration of the named entity recognition device 1. The named entity recognition device 1 can be implemented using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be realized using existing technologies. The central processing unit 901 executes instructions included in a program read from the RAM 902 or the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic operations and logical operations according to each instruction. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "Random Access Memory". The input / output port 903 is a port for the central processing unit 901 to exchange data with external input / output devices and the like. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads and writes data in the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port 903 via the bus 906.

[0078] Note that at least some functions of the specific expression recognition device 1 in the embodiment can be realized by a computer and a program. In that case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed. Here, the "computer system" is assumed to include hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, a DVD-ROM, a USB memory, or a storage device such as a hard disk built into a computer system. That is, the "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Furthermore, the "computer-readable recording medium" also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, something that temporarily and dynamically holds a program, and something that holds a program for a certain period of time, like a volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and may further be something that can be realized in combination with a program already recorded in the computer system for realizing the aforementioned functions.

[0079] The embodiments have been described above, but the following modification examples can also be implemented. As long as combinations are possible, the features of multiple modification examples may be combined and implemented.

[0080] [Modification Example 1] In the above-described embodiment, after the post-processing unit 30 assigned the label "PLACE", which indicates that it is a specific expression, to the subsequence "Saitama" with a higher score, it did not assign a label indicating that it is a specific expression to other subsequences "Saitama" - "origin" containing this "Saitama". That is, it was rejected as a candidate (step S16 in the flowchart of FIG. 9). In other words, when the post-processing unit 30 assigned a label indicating that the partial character string A is a specific expression, it rejected another partial character string B that completely included the partial character string A as a candidate. Here, for the partial character string B to completely include the partial character string A means that all the words included in the partial character string A exist in the partial character string B in that order.

[0081] As a modification, even when the post-processing unit 30 assigns a label indicating that the partial character string A is a specific expression, another partial character string B that completely includes the partial character string A may be a candidate for the specific expression. Further, in that case, the post-processing unit 30 can assign a label indicating that the partial character string B is a specific expression.

[0082] In the case of this modification, when the specific expression recognition device 1 includes an expression such as "NHK Broadcasting Technology Research Institute" in the input text, it can assign a label indicating that it is a specific expression to both the word sequence "NHK" and the word sequence "NHK Broadcasting Technology Research Institute". In other words, the specific expression recognition device 1 can assign the label "ORGANIZATION" (organization name) to both the word sequence "NHK" and the word sequence "NHK Broadcasting Technology Research Institute". That is, the specific expression recognition device 1 can recognize (extract) both the word sequence "NHK" and the word sequence "NHK Broadcasting Technology Research Institute" as specific expressions.

[0083] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.

[0084] [Modification 2] In the above-described embodiment, based on the learning data, machine learning of the embedding unit 24, the attention mechanism unit 26, and the feed-forward network 28 of the named-entity recognition device 1 was enabled. On the other hand, as a modification 2, the named-entity recognition device 1 may be configured using the already learned embedding unit 24, the attention mechanism unit 26, and the feed-forward network 28. In this case, the named-entity recognition device 1 can perform an estimation for an unknown input (sub-word sequence) based on the internal parameter values that have already been optimized by learning. That is, the named-entity recognition device 1 can calculate the score of each label corresponding to the unknown sub-word sequence. Further, the named-entity recognition device 1 can assign a label to the sub-word sequence based on the score.

[0085] [Modification 3] In the above-described embodiment, the label handled by the named-entity recognition device 1 indicates whether the sub-word sequence is a named entity, and further, when the sub-word sequence is a named entity, it indicates which type of named entity the sub-word sequence is. Also, in the above-described embodiment, the types of named entities include at least types such as person name (PERSON), place name (PLACE), and organization name (ORGANIZATION).

[0086] In contrast, in Modification 3, the label may indicate whether the sub-word sequence is a named entity, but may not indicate which type of named entity it is. In this case, the label has two types: a label indicating that it is a named entity (for example, "NE") and a label indicating that it is not a named entity (for example, "0").

[0087] [Experimental Results] An experiment was conducted to evaluate the performance of the embodiments described above, and the results are described below. In the experiment, three types of datasets, WNUT-16, WNUT-17, and CoNLL-03, were used as datasets for named entity extraction. The number of training data, development data, and evaluation data in each of these datasets is as shown in Table 1 below.

[0088]

Table 1

[0089] In this experiment, the methods used for named entity extraction were: 1) the method of the named entity recognition device according to the above embodiment, 2) SoTA (State-of-the-Art), 3) LSTM (Long Short Term Memory), 4) Max Pooling, and 5) Avg Pooling. Note that the methods from 2) to 5) are methods according to the prior art. Note that SoTA is the method with the best performance among the methods using the same data as of December 2023.

[0090] The experimental results (evaluation values) when using WNUT-16 and WNUT-17 are shown in Table 2, and the experimental results (evaluation values) when using CoNLL-03 are shown in Table 3. Note that for the methods other than 2) SoTA, three trials were conducted respectively, and the average value and the maximum value are described in their respective tables.

[0091]

Table 2

[0092]

Table 3

[0093] As shown in Table 2, for WNUT-16, 1) the maximum value (60.33) when using the method of this embodiment was the result with the highest performance. Also, for WNUT-16, the cases where the performance was better than SoTA were 1) the maximum value when using the method of this embodiment, 3) the average value and the maximum value when using LSTM, 4) the average value and the maximum value when using Max Pooling, and 5) the maximum value when using Avg Pooling.

[0094] Also, for WNUT-17, 1) the average value (62.21) and the maximum value (62.66) when using the method of this embodiment were the results with the highest performance. Also, for WNUT-17, the cases where the performance was better than SoTA were 1) the average value and the maximum value when using this embodiment, 3) the maximum value when using LSTM, and 4) the average value and the maximum value when using Max Pooling.

[0095] Also, as shown in Table 3, for CoNLL-03, 2) SoTA was the result with the highest performance. That is, for CoNLL-03, there was no method that obtained better performance than 2) SoTA. That is, for CoNLL-03, 2) SoTA was the result with the highest performance.

[0096] As described above, from the experimental results, it was confirmed that the method for extracting named entities described in this embodiment shows better performance than other methods for the datasets WNUT-16 and WNUT-17.

Industrial Applicability

[0097] The present invention can be used, for example, in processing for natural language. However, the scope of use of the present invention is not limited to those exemplified here.

Explanation of Signs

[0098] 1 Named Entity Recognition Device 22 Pre-trained Model Processing Unit (Model Processing Unit) 24 Embedding Unit 26 Attention Mechanism Unit 28 Feed-Forward Network 30 Post-Processing Unit 901 Central Processing Unit 902 RAM 903 Input / Output Port 904, 905 Input / Output Devices 906 Bus

Claims

1. A model processing unit that converts a token sequence corresponding to a partial word sequence that is a part of text into a vector sequence, An embedding unit that outputs a task type value corresponding to the type of task, The token sequence output from the model processing unit is input as value (V) and key (K), the task type value output from the embedding unit is input as query (Q), the similarity between the key and the query is calculated, and the weighted sum of the value with the similarity as the weight is calculated, and the weighted sum is output as estimation information about the partial word sequence. An attention mechanism unit, A feed-forward network that calculates a score of a label indicating whether the partial word sequence is a named entity based on the estimation information, A named entity recognition device comprising:

2. Based on a set of learning data that is a pair of an input token sequence and a correct label corresponding to the token sequence, the internal parameters of the models of the embedding unit, the attention mechanism unit, and the feed-forward network can be adjusted. Configured as such, The named entity recognition device according to claim 1.

3. The label represents whether the partial word sequence is a named entity, and further, when the partial word sequence is a named entity, it represents which type of named entity the partial word sequence is, The named entity recognition device according to claim 1.

4. The types of named entities include at least personal names, geographical names, and organization names, The named entity recognition device according to claim 3.

5. A post-processing unit that determines a label to be assigned to the partial word sequence based on the scores of each of the plurality of labels, The named entity recognition device according to claim 1, further comprising:

6. A model processing unit that converts a token sequence corresponding to a partial word sequence that is a part of text into a vector sequence, An embedding unit that outputs a task type value corresponding to the type of task, The token sequence output from the model processing unit is input as value (V) and key (K), the task type value output from the embedding unit is input as query (Q), the similarity between the key and the query is calculated, and the weighted sum of the value with the similarity as the weight is calculated, and the weighted sum is output as estimation information about the partial word sequence. An attention mechanism unit, A feed-forward neural network that calculates a score of a label indicating whether or not the partial word sequence is a named entity based on the estimated information; A program for causing a computer to function as a named entity recognition device including the same.