Predictor Interactive Learning System, Predictor Interactive Learning Method, and Program
The predictor interactive learning system addresses inefficiencies in conventional learning by using interest score calculations and dialogue-based data selection to enhance learning efficiency and accuracy with reduced data requirements.
Patent Information
- Application Number
- JP2021570044
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-06
- Filing Date
- 2021-01-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-01-05
AI Technical Summary
Conventional predictor learning systems require large amounts of teacher data and labels to improve prediction accuracy, but the effectiveness of these data and labels is unclear, leading to inefficient learning and prolonged time to achieve desired accuracy.
A predictor interactive learning system that calculates interest scores based on the appearance rate, predicted values, variance, and average values of words in a corpus, extracts teacher data using a dialogue learning frame, and performs machine learning with question-and-answer processes to select effective teacher data and labels.
Enables efficient predictor learning in less time with fewer teacher data and labels by selecting relevant teacher data based on interest scores, improving learning efficiency and accuracy.
Smart Images

Figure 0007708672000001 
Figure 0007708672000002 
Figure 0007708672000003
Abstract
Description
Technical Field
[0001] The present invention relates to a predictor interactive learning system, a predictor interactive learning method, and a program for learning a predictor (prediction model) that is a language analysis model. This application claims priority based on Japanese Patent Application No. 2020-443 filed in Japan on January 6, 2020, and incorporates its content herein by reference.
Background Art
[0002] Conventionally, a predictor as a language analysis model that performs named entity extraction on an input having a sequence structure such as words (characters, character strings, and symbols) in text, documents, and other text data has been known. When this predictor is a machine learning model, the predictor first performs machine learning using each combination of teacher data and teacher labels. Next, the predictor inputs the text data to be analyzed and predicts the named entity expression of each word in the text data, that is, predicts the possibility that the word is the label learned by the predictor.
[0003] Here, for example, the predictor outputs a value in the range from "0" to "1" for each word in the text data as the predicted value of the named entity expression. The closer this predicted value is to "1", the higher the possibility that it is that label. In addition, when training the above predictor, in order to improve the accuracy of the prediction of the label, it is necessary to train using a large amount of teacher data and teacher labels. For this reason, a learning device that trains the predictor using teacher data and teacher labels is used (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the learning of the above-described predictor, in order to improve the accuracy of the prediction of the eigenvalue representation, it is necessary to perform learning using a large amount of teacher data and teacher labels. However, it is unclear whether each of the teacher data and teacher labels used for learning is effective for the learning of the predictor. That is, it is not known whether each of the teacher data and teacher labels used for learning is data that improves the prediction accuracy of the predictor. For this reason, there is a drawback that even if a large amount of teacher data and teacher labels are used, the prediction accuracy of the predictor cannot be improved, or it takes a long time to learn until a predetermined prediction accuracy is obtained. That is, there is a drawback that efficient learning of the predictor cannot be performed.
[0006] The present invention has been made in view of such circumstances. An object of the present invention is to provide a predictor interactive learning system, a predictor interactive learning method, and a program that can perform learning of a predictor in a shorter time by using less teacher data and teacher labels as compared with the prior art.
Means for Solving the Problems
[0007] The predictor interactive learning system of the present invention performs machine learning on a predictor that outputs a prediction value indicating the probability that an input word is a predetermined specific expression using predetermined teacher data and teacher labels, and for each word in the corpus used for the machine learning, the prediction value output by the predictor for the word and statistical data regarding the word in the corpus fromAn interest score calculation unit that calculates an interest score, a dialogue learning frame unit that extracts, from the corpus, the word that will be the teacher data in the next learning for the predictor based on the interest score, and a question-and-answer unit that outputs a question as to whether the extracted teacher data is the specific expression for which the predictor predicts a probability, and obtains a teacher label corresponding to the teacher data as a response to the question. The machine learning unit performs machine learning on the predictor using the teacher data extracted by the dialogue learning frame unit and the teacher label obtained by the question-and-answer unit for the teacher data. i. The interest score calculation unit calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction values 。 Further, the predictor interactive learning system of the present invention includes a machine learning unit that performs machine learning on a predictor that outputs a prediction value indicating the accuracy that an input word is a predetermined unique expression using predetermined teacher data and teacher labels, and for each word in the corpus used for the machine learning, an interest score calculation unit that obtains an interest score from the prediction value output by the predictor for the word and statistical data related to the word in the corpus, a dialogic learning frame unit that extracts the word that will be the teacher data in the next learning for the predictor from the corpus based on the interest score, a question and answer unit that outputs a question as to whether the extracted teacher data is the unique expression for which the predictor predicts the accuracy and obtains a teacher label corresponding to the teacher data as a response to the question, and a first interest score calculation unit that calculates a first interest score, which is an interest score for selecting teacher data for the first question from each of the words in the corpus, corresponding to a predetermined calculation rule using the combination of the teacher data and the teacher label input as a seed when starting the machine learning of the predictor. The machine learning unit performs machine learning of the predictor using the teacher data extracted by the dialogic learning frame unit and the teacher label obtained by the question and answer unit for the teacher data, and the calculation rule is a rule for obtaining the first interest score of each of the words corresponding to the degree of coincidence between other words arranged adjacent to the front and rear of the word of the teacher data input as the seed in the corpus and other words arranged adjacent to the front and rear of each of the words in the corpus
[0010] In the predictor dialogue learning system of the present invention, the dialogue learning frame unit may extract the teacher data corresponding to the initial interest score from each of the words in the corpus, and the question-and-answer unit may obtain a response indicating whether each of the teacher data corresponding to the initial interest score is a teacher label corresponding to the teacher data input with each of the teacher data as a seed.
[0011] The predictor dialogue learning method of the present invention includes a machine learning process in which a machine learning unit performs machine learning on a predictor to be learned using teacher information including predetermined teacher data and teacher labels, and an interest score calculation unit calculates, for each word in the corpus that is the teacher information, the predicted value output by the predictor for the word and statistical data regarding the word in the corpus from an interest score calculation process for obtaining an interest score, a dialogue learning process in which a dialogue learning frame unit extracts teacher words that will be teacher data for the predictor from the corpus based on the interest score, and a question-and-answer process in which a question-and-answer unit outputs a question as to whether the teacher word is a specific value expression as a label of the predictor, and obtains a teacher label corresponding to the teacher word as a response to the question. The machine learning unit performs machine learning on the predictor using the teacher data extracted by the dialogue learning frame unit and the label obtained by the question-and-answer unit for the teacher data as the teacher label. (i) The interest score calculation unit calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction value. 。 Further, in the predictor interactive learning method of the present invention, a machine learning process in which a machine learning unit performs machine learning on a predictor that outputs a prediction value indicating the probability that a word input using predetermined teacher data and teacher labels is a predetermined unique expression, an interest score calculation process in which an interest score calculation unit obtains an interest score for each word in the corpus used for the machine learning from the prediction value output by the predictor for the word and the statistical data regarding the word in the corpus, a dialogic learning frame process in which a dialogic learning frame unit extracts the word that will be the teacher data in the next learning for the predictor from the corpus based on the interest score, a question and answer process in which a question and answer unit outputs a question as to whether the extracted teacher data is the unique expression for which the predictor predicts the probability and obtains a teacher label corresponding to the teacher data as a response to the question, and a first-time interest score calculation process in which a first-time interest score calculation unit calculates, in accordance with a predetermined calculation rule, a first-time interest score, which is an interest score for selecting teacher data for a first-time question from each of the words in the corpus using a combination of the teacher data and the teacher label input as a seed when starting the machine learning of the predictor. The machine learning unit performs machine learning of the predictor using the teacher data extracted by the dialogic learning frame unit and the teacher label obtained by the question and answer unit for the teacher data, and the calculation rule is a rule for obtaining the first-time interest score for each of the words corresponding to the degree of coincidence between other words arranged adjacent to the front and rear of the word of the teacher data input as the seed in the corpus and other words arranged adjacent to the front and rear of each of the words in the corpus.
[0012] The program of the present invention is a machine learning means for performing machine learning on a computer using predetermined teacher data and teacher labels to train a predictor that outputs a predicted value indicating the probability that an input word is a predetermined named entity, and for each word in the corpus used for the machine learning, the predicted value output by the predictor for that word and statistical data regarding the word in the corpus from interest score calculation means for calculating an interest score, dialogue learning framework means for extracting, from the corpus, the word that will be the teacher data in the next learning for the predictor based on the interest score, question-and-answer means for outputting a question as to whether the extracted teacher data is the named entity for which the predictor predicts a probability and obtaining the teacher label corresponding to the teacher data as a response to the question, and causing the machine learning means to perform machine learning on the predictor using the teacher data extracted by the dialogue learning framework means and the teacher label obtained by the question-and-answer means for that teacher data (i) The interest score calculating means calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction value. 。 Further, the program of the present invention causes a computer to perform machine learning means for performing machine learning on a predictor that outputs a prediction value indicating the probability that an input word is a predetermined unique expression using predetermined teacher data and teacher labels, and for each word in the corpus used for the machine learning, an interest score calculating means for obtaining an interest score from the prediction value output by the predictor for the word and the statistical data regarding the word in the corpus, a dialog learning frame means for extracting the word that will be the teacher data in the next learning for the predictor from the corpus based on the interest score, a question and answer means for outputting a question as to whether the extracted teacher data is the unique expression for which the predictor predicts the probability, and obtaining a teacher label corresponding to the teacher data as a response to the question, and when starting the machine learning of the predictor, an initial interest score calculating means for calculating an initial interest score, which is an interest score for selecting teacher data for the first question from each of the words in the corpus, corresponding to a predetermined calculation rule using the combination of the teacher data and the teacher label input as a seed. The machine learning means performs machine learning of the predictor using the teacher data extracted by the dialog learning frame means and the teacher label obtained by the question and answer means for the teacher data, and the interest score calculating means calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction value.
Advantages of the Invention
[0013] According to this invention, it is possible to provide a predictor dialogue-type learning system, a predictor dialogue-type learning method, and a program that can train a predictor in less time with less teacher data and teacher labels compared to the prior art.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0015] Hereinafter, with reference to the drawings, a predictor interactive learning system according to an embodiment of the present invention will be described. In this embodiment, the machine learning model that the predictor interactive learning system causes to learn is a predictor that predicts the unique expression of each word in a natural language sentence (sentence data). FIG. 1 is a diagram showing a configuration example of a predictor interactive learning system according to an embodiment of the present invention. In FIG. 1, the predictor interactive learning system 100 includes a data input / output unit 101, a first interest score calculation unit 102, a machine learning unit 103, an interest score calculation unit 104, a dialog-based learning frame unit 105, a question and answer unit 106, a corpus data storage unit 107, a rule-based question data storage unit 108, a teacher data storage unit 109, and a predictor weight coefficient storage unit 110.
[0016] The data input / output unit 101 writes and stores the corpus data (corpus data) input from an external device into the corpus data storage unit 107. FIG. 2A shows the character array of the learning sentence data as an example of the corpus. This learning sentence data structures a natural language sentence and is decomposed into words (including symbols) one by one by morphological analysis or the like.
[0017] FIG. 3 is a sequence diagram showing the concept of each step of the learning process of the predictor interactive learning system 100 in this embodiment. In the predictor interactive learning system 100 of this embodiment, the interest score calculation step, the question-answer step, and the learning step shown in FIG. 3 are repeated. When sentence data is input, the predictor interactive learning system 100 performs learning of a predictor that outputs a predicted value of the unique expression of each word. This learning is the learning of the weight coefficients of each function in the predictor, and the details will be described later.
[0018] Here, in the interest score calculation step, when each word in the corpus is input to the interest score calculation unit 104, the interest score is obtained for each word in the corpus from the predicted value output by the predictor being trained and the statistical value for each word. This interest score indicates the degree of learning efficiency for the predictor and is used to extract words to be used as teacher data in the next learning.
[0019] Also, in the question - answer step, the question - answering unit 106 asks the user whether the words extracted as teacher data from the interest score correspond to the named entities predicted by the predictor. Next, in this question - answer step, the user is asked to obtain a teacher label for the word as an answer to the question, and a combination of teacher data and teacher label is generated.
[0020] In the learning step, the predictor is trained using the teacher data and teacher label obtained in the question - answer step. The interactive learning system 100 obtains the predicted value for each word in the corpus of this predictor after learning and then proceeds to the interest score calculation step.
[0021] The above - mentioned interest score calculation step, question - answer step, and learning step are repeated, for example, a predetermined number of times to obtain a predictor that outputs a predicted value for the named entity in the text data. Each of the predictors is preferably trained by selecting a corpus (including words as similar technical terms) for each specialty in the field for which extraction is desired when used for extracting the named entities of technical terms in a document.
[0022] Returning to FIG. 1, the data input / output unit 101 writes the learning text data of the corpus used for training the predictor, which is supplied from an external device, to the corpus data storage unit 107 for storage. Also, the data input / output unit 101 writes the seed (data of the initial teacher data and teacher label) input from input means such as a keyboard or a touch panel to the rule - based question data storage unit 108 for storage.
[0023] The first-time interest score calculation unit 102 calculates the first-time interest scores of all words in the corpus using the first-time teacher data supplied from the outside and the seed as the teacher label. In the present embodiment, learning of a predictor to be learned is performed using teacher data and a teacher label, the corpus is input to this predictor to obtain predicted values of each word, and an interest score is calculated corresponding to this predicted value. Based on this interest score, teacher data and a teacher label to be used for the next learning of the predictor are acquired. For this reason, at the time of starting learning, since the predicted value of the predictor cannot be obtained, the interest score cannot be obtained, and thus the first-time interest score based on a rule is calculated. The calculation of this first-time interest score will be described in detail later.
[0024] The machine learning unit 103 uses the teacher data and the teacher label to generate a predictor that is a machine learning model, that is, to perform learning of the weight coefficients of each function in the predictor (learning step in FIG. 3). The teacher data and the teacher label are recorded in the teacher data table shown in FIG. 12 created from the question-and-answer list in FIG. 3 described later. The learning of the weight coefficients of each function in the predictor is, for example, if it is a neural network such as an RNN (Recurrent Neural Network), the learning of the parameters of the weight coefficients of the input of the function in each layer. Hereinafter, this learning will be simply referred to as "learning of the predictor". Next, the machine learning unit 103 writes the weight coefficients of each function in the predictor obtained by learning to the predictor weight coefficient storage unit 110 for storage and updates them to the new weight coefficients obtained each time learning is performed.
[0025] When the number of times the machine learning unit 103 has performed learning of the predictor reaches a predetermined set end number of times, the machine learning unit 103 ends the learning of the predictor and generates and outputs a predictor corresponding to the weight coefficients at that time. The machine learning unit 103 may be configured to end the learning of the predictor when the predicted value output by the predictor during learning becomes equal to or greater than a preset end threshold value.
[0026] The interest score calculation unit 104 acquires the predicted value of each word used when extracting the specific expression in the corpus from the predictor being learned, and calculates the interest score based on the word including this predicted value and the statistical value corresponding to the predicted value (the interest score calculation step in FIG. 3). The predicted value indicates a numerical value in the range from "0" to "1" output by the predictor (here, the predictor being learned) as the label of a predetermined specific expression (for example, "animal", etc.) of each word. The closer the predicted value is to "1", the higher the accuracy of the predicted predetermined specific expression. In this embodiment, the above-mentioned interest score is obtained, for example, as the average value of four interest scores: the first interest score, the second interest score, the third interest score, and the fourth interest score.
[0027] The first interest score is obtained corresponding to the normalized rank (range of 0 to 1) of the appearance rate obtained by dividing the number of the same words in the corpus by the total number of words in the corpus. For example, when the corpus is formed by 15 words of A, A, A, A, B, B, B, C, C, C, C, C, D, D, E where each of the words A, B, C, D, and E is used, the appearance rate of word A is 4 / 15, the appearance rate of word B is 3 / 15, the appearance rate of word C is 5 / 15, the appearance rate of word D is 2 / 15, and the appearance rate of word E is 1 / 15. The rank of the appearance rate is that word A is in the 2nd place, word B is in the 3rd place, word C is in the 1st place, word D is in the 4th place, and word E is in the 5th place. The normalized rank of the appearance rate is 0.75 for A, 0.5 for B, 1.0 for C, 0.25 for D, and 0.0 for word E.
[0028] The second interest score is obtained corresponding to the predicted value (in the range of 0 to 1) output by the predictor for a word. The third interest score is obtained corresponding to the average value (in the range of 0 to 1) obtained by averaging the predicted values of the same word. The fourth interest score is obtained corresponding to the variance (in the range of 0 to 1) of the predicted values of the same word. As will be described later, each of the first interest score, the second interest score, the third interest score, and the fourth interest score is normalized by its respective maximum value. Here, in the present embodiment, each graph of the first interest score to the fourth interest score is created while changing the coordinate values (two-dimensional coordinates consisting of the vertical axis and the horizontal axis) used in the spline interpolation (cubic spline interpolation) described later so that the learning efficiency is better in the experiment of training the predictor using a plurality of corpora.
[0029] FIG. 4 is a graph showing the correspondence between the first interest score and the appearance rate. In the graph of FIG. 4, the vertical axis represents the numerical value of the first interest score, and the horizontal axis represents the appearance rate. The interest score curve L1 shows the correspondence between the first interest score and the appearance rate. This interest score curve L1 is obtained by cubic spline interpolation of five coordinate values (0.1, 1.0), (0.25, 2.0), (0.5, 2.0), (0.75, 5.0), (1.0, 7.0), where the appearance rate on the horizontal axis is x = (0.1, 0.25, 0.5, 0.75, 1.0) and the first interest score on the vertical axis is y = (1.0, 2.0, 2.0, 5.0, 7.0). The interest score curve L1 is such that the higher the appearance rate, the higher the first interest score, and it is considered that the position where the same word exists in the corpus is more, and the shape is such that information about the same word regarding a wider corpus can be obtained.
[0030] FIG. 5 is a graph showing the correspondence between the second interest score and the predicted value. In the graph of FIG. 5, the vertical axis represents the numerical value of the second interest score, and the horizontal axis represents the predicted value. The interest score curve L2 shows the correspondence between the second interest score and the predicted value. This interest score curve L2 is obtained by cubic spline interpolation of five coordinate values (0.1, 1.0), (0.25, 2.0), (0.5, 3.0), (0.75, 3.0), (1.0, 6.0), where the prediction rate on the horizontal axis is x = (0.1, 0.25, 0.5, 0.75, 1.0) and the second interest score on the vertical axis is y = (1.0, 2.0, 3.0, 3.0, 6.0). In the combination of teacher data and teacher labels in the specific expression, there are many data with false (= 0) in the interest score curve L2. The higher the predicted value, the higher the probability that the predicted value of the teacher data to be questioned is true (= 1). Therefore, the teacher data with true (= 1) teacher labels and the teacher data with false (= 0) teacher labels are well balanced and can be used for predictive learning. Therefore, it is considered to be in a shape that can improve the accuracy of the predicted value of the predictor and reduce the number of learning times of the predictor.
[0031] Figure 6 is a graph showing the correspondence between the third interest score and the average value of the predicted values. In the graph of Figure 6, the vertical axis shows the numerical value of the third interest score, and the horizontal axis shows the average value of the predicted values. The interest score curve L3 shows the correspondence between the third interest score and the average value of the predicted values. This interest score curve L3 is obtained by cubic spline interpolation of five coordinate values (0.1, 1.0), (0.25, 2.0), (0.5, 2.0), (0.75, 5.0), (1.0, 7.0), where the average value of the predicted values on the horizontal axis is x = (0.1, 0.25, 0.5, 0.75, 1.0) and the third interest score on the vertical axis is y = (1.0, 2.0, 2.0, 5.0, 7.0). Similar to the interest score curve L2, in the combination of teacher data and teacher labels in the specific expression, there are many data with false (= 0) in the interest score curve L3. However, the higher the average value of the predicted value, the higher the probability that the predicted value of the teacher data to be questioned is true (= 1). Therefore, the teacher data with true (= 1) teacher labels and the teacher data with false (= 0) teacher labels are well balanced and can be used for predictive learning. Therefore, it is considered to be in a shape that can improve the accuracy of the predicted value of the predictor and reduce the number of learning times of the predictor.
[0032] FIG. 7 is a graph showing the correspondence between the fourth interest score and the variance of the predicted values. In the graph of FIG. 7, the vertical axis represents the numerical value of the fourth interest score, and the horizontal axis represents the variance of the predicted values. The interest score curve L4 shows the correspondence between the fourth interest score and the variance of the predicted values. This interest score curve L4 is obtained by cubic spline interpolation of five coordinate values (0.1, 1.0), (0.25, 2.0), (0.5, 1.0), (0.75, 3.0), and (1.0, 2.0), where the variance of the predicted values on the horizontal axis is x = (0.1, 0.25, 0.5, 0.75, 1.0) and the fourth interest score on the vertical axis is y = (1.0, 2.0, 1.0, 3.0, 2.0). The interest score curve L4 shows that the higher the variance of the predicted values, the higher the fourth interest score. At the positions where the same word exists in the corpus, the magnitudes of the predicted values are different. That is, since each of these same words appears in various different contexts, it is considered that the shape has a higher learning efficiency compared to words that appear in the same context.
[0033] FIG. 8 is a conceptual diagram for explaining the appearance rate, predicted value, average value of the predicted values, and variance of the predicted values of each word in the corpus. In FIG. 8, the word "dog" is used as an example of a word in the corpus for explanation. The interest score calculation unit 104 refers to the corpus and obtains the number of words, which is the number of the word "dog" in this corpus. Next, the interest score calculation unit 104 divides the number of the word "dog" by the total number of words, which is the total number of all words in the corpus, to obtain the appearance rate of the word "dog".
[0034] That is, the interest score calculation unit 104 obtains the appearance rate for obtaining the first interest score of each word as the ratio of the same word in the corpus. This appearance rate is normalized in the range from 0 to 1 by the maximum value of the appearance rate of each word in the corpus. For example, in the corpus 200 shown in FIG. 8, when the total number of words is 100, since the word "dog" exists at three positions 201, 202, and 203, the appearance rate is "0.03". Also, in the case where the word "dog" has the highest appearance rate, by normalization, the appearance rate of the word "dog" becomes "1".
[0035] The interest score calculation unit 104 normalizes each predicted value of the words output by the predictor during learning, and uses it as the predicted value for obtaining the second interest score for each word. Also, the predicted value indicates a numerical value in the range from "0" to "1" output by the predictor as the label of a predetermined unique expression (for example, an animal, etc.) of each word. The closer the predicted value is to "1", the higher the probability that it is the predicted predetermined unique expression. In FIG. 8, the predicted value of the word "dog" at position 201 in the corpus 200 is 0.64, the predicted value of the word "dog" at position 202 is 0.79, and the predicted value of the word "dog" at position 203 is 0.56.
[0036] Next, the interest score calculation unit 104 obtains the average value of the predicted values (either the unnormalized numerical value or the normalized numerical value can be used) for each identical word, and uses it as the average value of the predicted values for obtaining the third interest score for each identical word. In FIG. 8, since the predicted values of the word "dog" at each of positions 201, 202, and 203 are 0.64, 0.79, and 0.56, the average value is 0.66.
[0037] The interest score calculation unit 104 obtains the variance of the predicted values for each identical word, and uses it as the variance of the predicted values for obtaining the fourth interest score for each identical word. In FIG. 8, since the predicted values of the word "dog" at each of positions 201, 202, and 203 are 0.64, 0.79, and 0.56, and the average value of the predicted values is 0.66, the variance of the predicted values is 0.1.
[0038] The interest score calculation unit 104 obtains the average value of the first interest score, the second interest score, the third interest score, and the fourth interest score of each word in the corpus, and sets this average value as the interest score of each word. For example, in the corpus 200 of FIG. 8, the first interest score of the word "dog" at position 201 is obtained as 8 according to FIG. 4 because the appearance rate is 1. Similarly, the second interest score of the word "dog" is obtained as 2 according to FIG. 5 because the predicted value is 0.64. The third interest score of the word "dog" is obtained as 1.5 according to FIG. 6 because the average value of the predicted values is 0.66.
[0039] Also, the fourth interest score of the word "dog" is obtained as 1.8 according to FIG. 7 because the variance of the predicted values is 0.10. Therefore, the interest score of the word "dog" at position 201 is obtained as (8 + 2 + 1.5 + 1.8) / 4 = 5.88. Next, the interest score calculation unit 104 calculates the interest scores of the words "dog" at positions 202 and 203, obtains the average value of the interest scores from positions 201 to 203, and sets this average value as the interest score of the word "dog" in the corpus 200.
[0040] The dialog learning frame unit 105 refers to the interest scores of each word (the same word) in the corpus 200, and distributes the probability that each word is selected (extracted) as teacher data corresponding to the magnitude of each interest score. The dialog learning frame unit 105 extracts a predetermined number of words as teacher data according to the probability. For example, if there are only two words, word A and word B, the interest score of word A is 1, and the interest score of word B is 2, the probability that word A is selected as teacher data is set as 1 / 3, and the probability that word B is selected as teacher data is set as 2 / 3 and distributed to each word. As a result, the probability of being selected as teacher data is twice that of word B compared to word A, and the ease of selection of word B compared to word A is twice.
[0041] At this time, the interactive learning frame unit 105 refers to the teacher data table (the table in FIG. 12 described later) in the teacher data storage unit 109, and excludes the words that have already been extracted as teacher data from the target of extracting teacher data. Next, the interactive learning frame unit 105 performs a process of randomly extracting words to be used as teacher data from the words in the corpus 200 according to the above-described selected probabilities of the respective words. The interactive learning frame unit 105 outputs the extracted teacher data as a question list, assigns respective interest scores thereto, and outputs the result to the question and answer unit 106.
[0042] Also, the interactive learning frame unit 105 is supplied with a question and answer list for this question list from the question and answer unit 106. Then, the interactive learning frame unit 105 outputs a question and answer table in which teacher labels, which are answers to the questions, are respectively assigned to each of the teacher data that are questions in the question list, to the machine learning unit 103. FIG. 9 is a diagram showing a configuration example of a question list generated by the interactive learning frame unit 105. For each record, teacher data and an interest score are respectively associated.
[0043] The question and answer unit 106 presents questions corresponding to the question list supplied from the interactive learning frame unit 105 to the user as, for example, a question table, and acquires teacher labels corresponding to the teacher data that are answers to the questions from the user (question-answer step in FIG. 3). Then, the question and answer unit 106 acquires the response from the user to the question for the teacher data, and outputs a question-answer list in which this teacher data and the response are associated with each other, to the interactive learning frame unit 105.
[0044] FIG. 10 is a diagram showing a configuration example of a question table presented by the question and answer unit 106 to the user. In FIG. 10, for each record, teacher data and a question regarding this teacher data are shown. When the predictor to be learned is a machine learning model that outputs a predicted value indicating the probability that "animal" is a unique expression of each word in the text data, the question regarding the teacher data is "Is it an animal?". In the present embodiment, although it is described as a question table as an example, other formats may be used as the output format.
[0045] FIG. 11 is a diagram showing a configuration example of a question-and-answer table output by the question-and-answer unit 106. In FIG. 11, for each record, teacher data and a response (answer) regarding this teacher data are shown. The response shows either yes (true) or no (false) as the answer from the user to "Is it an animal?" for each of the teacher data. Here, yes (true) indicates "1" as the teacher label, and no (false) indicates "0" as the teacher label. In the present embodiment, although it is described as a question-and-answer table as an example, other formats may be used as the output format.
[0046] FIG. 12 is a diagram showing a configuration example of a teacher data table written in the teacher data storage unit 109. In FIG. 12, for each record, teacher data and "0" or "1" as the teacher label corresponding to this teacher data are shown. In FIG. 12, "yes (true)" in FIG. 11 is set as the teacher label "1", and "no (false)" is set as the teacher label "0". The interactive learning frame unit 105 adds new teacher data and teacher labels as new records to the teacher data table in the teacher data storage unit 109 based on the question-and-answer list supplied from the question-and-answer unit 106.
[0047] Next, the calculation of the initial interest score by the initial interest score calculation unit 102 will be described. In this embodiment, the initial interest score is calculated based on the seed of the teacher data. When training the predictor using the corpora of FIGS. 2A and 2B, for example, if the named entity to be predicted is an animal, the seeds are words such as "dog" and "cat". The user inputs the seed words one by one into the input field on the initial question selection screen (not shown). As a result, the data input / output unit 101 writes and stores, in the rule-based question data storage unit 108, a seed table showing seeds such as "dogs" and "cats" and the named entity "animal". In this embodiment, a seed table is described as an example, but other formats may be used as the output format.
[0048] FIG. 13 is a diagram showing a configuration example of the seed table in the rule-based question data storage unit 108. In the seed table of FIG. 13, for each record, a seed word (teacher data) and a label of the seed word (data indicating 1 or 0, teacher label) corresponding to the seed word (teacher data) and the named entity are shown.
[0049] Thereby, the initial interest score calculation unit 102 refers to the seed table in the rule-based question data storage unit 108 and extracts adjacent words (adjacent words) to the input seed words, "dog" and "cat", from the corpus shown in FIG. 2A. In this embodiment, as shown in FIG. 2B, four words, two in the front stage and two in the rear stage, of the target seed word are extracted as an example. The number of adjacent words can be arbitrarily set. For example, in FIG. 2B, the adjacent words of the seed word "dogs" at position 301 are four words, "example" and "," in the front stage and "," and "cats" in the rear stage.
[0050] The adjacent words of the seed word "cats" at position 302 are four words, "dogs" and "," in the front stage and "," and "and" in the rear stage. The adjacent words of the seed word "dogs" at position 303 are four words, namely the two words "mammals" and "," in the previous segment and the two words "," and "bears" in the subsequent segment. The adjacent words are characters (including symbols such as ", (comma)" and ": (colon)") or character strings (including single characters).
[0051] In the process of the first interest score calculation unit 102 for the first question selection, according to the following rules, it calculates the interest score used to extract the teacher data of the predictor from the words other than the seed. First, the first interest score calculation unit 102 extracts the adjacent words of each word in the entire Co-Pass text by the same process as in the case of the above-mentioned seed word (hereinafter referred to as the seed word). For example, as shown in FIG. 2B, from the corpus in FIG. 2A, the first interest score calculation unit 102 extracts, as the adjacent words of the word "bears" at position 304, which is a word other than the seed, the words "dogs" and "," in the previous segment and the words "," and "and" in the subsequent segment. Similarly, the first interest score calculation unit 102 extracts, as the adjacent words of the word "meat" at position 305, the words "dogs" and "eat" in the previous segment and the words "and" and "cats" in the subsequent segment. The first interest score calculation unit 102 compares each of the adjacent words of the seed word with the adjacent words of all the words in the entire Co-Pass text, and outputs the matching state (the number of matching words) of the adjacent words of the seed word and each word as the comparison result, that is, the interest score.
[0052] When comparing the adjacent words of the word "bears" and the seed word "cats" (at position 302), since the words "dogs" and "," in the previous segment and the words "," and "and" in the subsequent segment all match, that is, all four words in the previous and subsequent segments match, the first interest score of the word "bears" is 4. Also, when comparing the adjacent words of the word "meat" with the seed word "cats" (at position 302), in terms of the comparison results between the words before and after "meat", namely "dogs" and "eat", and "and" and "cats", and between the words before and after "cats", namely "dogs" and ",", and ",", and "and", "dogs" and "and" match in both "meat" and "cats".
[0053] However, in the rules of this embodiment, even if there are words that match in adjacent words, if they do not match the first adjacent words adjacent to both sides of the seed word, the interest score is 0. Therefore, in the comparison of the adjacent words of the above-mentioned word "meat" and the seed word "cats" (at position 302), since the matching "dogs" and "and" are not the first adjacent words but the second adjacent words located across the first adjacent words, the interest score of the word "meat" is set to 0.
[0054] And when the same word exists at multiple locations in the corpus, for the target word, the maximum value of the interest scores obtained at multiple locations is obtained as the initial interest score of that same word. Here, the initial interest score calculation unit 102 outputs the initial interest score of each word in the corpus to the interactive learning frame unit 105. When the initial interest score supplied from the initial interest score calculation unit 102 to the interactive learning frame unit 105 reaches a preset threshold, for example, when the threshold is 3, the interactive learning frame unit 105 extracts each word with an initial interest score of 3 or more.
[0055] Also, the interactive learning frame unit 105 generates a question list using the extracted words as teacher data and outputs the generated question list to the question and answer unit 106. Thereby, as already described, the question and answer unit 106 performs the process of obtaining from the user the teacher label of each piece of teacher data in the question list (that is, whether it is an animal (=1) or not an animal (=0)).
[0056] Then, for the initial learning of the predictor, the seed words and each of the words extracted by the interactive learning frame unit 105 based on the initial interest score are used as teacher data, and the predictor is learned so that the teacher label corresponding to this teacher data becomes the predicted value. As described above, the interest score for extracting the teacher data used for learning after the second time is obtained by the interest score calculation unit 104 for each word in the corpus by inputting each word and corresponding to the predicted value output by the predictor.
[0057] Figure 14 is a flowchart showing an operation example of the learning process of the predictor performed by the predictor interactive learning system 100 according to the present embodiment. Step S1: The user inputs seed words to the predictor interactive learning system 100. For example, when the predictor outputs a predicted value indicating the angle at which the word represents an animal as a unique expression of the word, "dogs" and "cats", which are words representing animals in the corpora of FIGS. 2A and 2B, are input to the predictor interactive learning system 100 as seed words.
[0058] Step S2: The data input / output unit 101 outputs each of the input seed words to the initial interest score calculation unit 102. The initial interest score calculation unit 102 compares the adjacent words of the seed words with the adjacent words of all the words in the corpus, and calculates the initial interest score of each word in the corpus. Next, the initial interest score calculation unit 102 outputs the calculated initial interest score to the interactive learning frame unit 105.
[0059] Step S3: As a result, the interactive learning frame unit 105 extracts words corresponding to the initial interest scores equal to or higher than a preset threshold from the initial interest scores supplied from the initial interest score calculation unit 102. The interactive learning frame unit 105 generates a question list using the words corresponding to the initial interest scores as teacher data. Next, the interactive learning frame unit 105 provides the generated question list to the question / answer unit 106. In addition, the dialog learning frame unit 105 generates a question list for the second and subsequent times by randomly extracting words from the corpus according to the probability of being selected as teacher data corresponding to the interest score generated by the interest score calculation unit 104, and generates a question list using the extracted words as teacher data.
[0060] Step S4: The dialog learning frame unit 105 outputs the generated question list (Figure 9) to the question-answer unit 106. Based on the supplied question list, the question-answer unit 106 displays a question table (Figure 10) on a display screen (not shown), asks whether the label for the teacher data is correct, and prompts the user for a response.
[0061] Step S5: The question-answer unit 106 generates a question-answer table (Figure 11) based on the response (either label = true (= 1) or false (= 0)) for each piece of teacher data in the question table. Then, the question-answer unit 106 outputs the generated question-answer table to the dialog learning frame unit 105. The dialog learning frame unit 105 outputs the question-answer table supplied from the question-answer unit 106 to the machine learning unit 103.
[0062] Step S6: The machine learning unit 103 extracts new teacher data and teacher labels (\"1\" if the word is a named entity, or \"0\" if it is not) from the question-answer table, adds them to the teacher data table (Figure 12) in the teacher data storage unit 109, and stores them by writing. Next, the machine learning unit 103 performs a learning process of adjusting the weight coefficients in the function of the predictor based on the teacher data and teacher labels shown in the teacher data table. Thereby, the machine learning unit 103 writes the adjusted weight coefficients to the predictor weight coefficient storage unit 110 and stores them.
[0063] Step S7: The machine learning unit 103 determines whether the number of learning times for learning the predictor has reached a preset end number of times. At this time, if the number of learning times is not equal to or more than a preset end number of times, that is, if it is less than the end number of times, the machine learning unit 103 proceeds with the process to step S8. On the other hand, if the number of learning times reaches the preset end number of times or more, the machine learning unit 103 proceeds with the process to step S10.
[0064] Step S8: The machine learning unit 103 outputs, to the interest score calculation unit 104, the predicted value output by the predictor for each word in the corpus at the current time. The interest score calculation unit 104 acquires, from the machine learning unit 103, the predicted value output by the predictor for each word in the corpus.
[0065] Step S9: The interest score calculation unit 104 calculates the interest score for each word based on the predicted value of each acquired word. Then, the interest score calculation unit 104 outputs the calculated interest score to the dialog learning frame unit 105.
[0066] Step S10: The machine learning unit 103 outputs a predictor having a weight coefficient in the function of the predictor, which is stored in the predictor weight coefficient storage unit 110.
[0067] Accordingly, according to the predictor dialog learning system of the present embodiment, it is possible to acquire teacher data and teacher labels used in the learning of the predictor based on the interest score indicating the degree of learning efficiency of the predictor, and to perform efficient learning of the predictor with a smaller number of teacher data and teacher labels and a smaller number of learning times compared to the conventional case.
[0068] The predictor to be learned by the predictor dialog learning system of the present embodiment outputs, for each word in the text data, a probability (a numerical value in the range from 1 indicating true to 0 indicating false) that the word indicates a predetermined specific expression (for example, "animal"). When constructing a named-entity extractor that extracts words of a predetermined named entity from any text data using this predictor, words with a predicted value equal to or higher than a preset predicted value are extracted as predetermined named entities corresponding to the predictor, and each of the extracted words and their respective appearance positions are output as extraction data.
[0069] When this named-entity extractor includes a first predictor that outputs a predicted value indicating the probability that the named entity is "animal" and a second predictor that extracts a predictor that indicates the probability that the named entity is "plant", for example, when the input text data is "Most rabbits eat cabbage,and cat eat Cat Grass.", a data sequence of [{"animal": [[2, 1], [7, 1],...]}, "plant": [[4, 1], [9, 2],...]} is output as the output.
[0070] In the above data sequence, "animal": [[2, 1], [7, 1],...] means that the words with the named entity "animal" are expressed as [2 (the second word), 1 (consisting of 1 word (rabbits))], [7 (the seventh word), 1 (1 word (rabbits))], etc. Also, "plant": [[2, 1], [7, 1],...] means that the words with the named entity "plant" are expressed as [4 (the fourth word), 1 (consisting of 1 word (cabbage))], [9 (the ninth word), 2 (2 words (Cat Grass))], etc.
[0071] In this embodiment, although it is described as a single word as the training data, the words whose probabilities are estimated by the predictors learned by the predictor interactive learning system 100 are not limited to words composed of only one word, and multiple words in which two or more words are arranged continuously are also target words for extraction (that is, words as units for estimating probabilities). Therefore, as a word segmentation (definition of words) process in the preprocessing of the corpus, not only single words but also multiple words, for example, multiple words in which two words are continuous, are registered as words.
[0072] Here, the maximum number of consecutive words to be defined as one word is preset at this preprocessing stage. For example, when setting that a maximum of two consecutive words are to be regarded as a word, in the corpora of FIGS. 2A and 2B, words such as "rabbits", "dogs", "cabbage", "meat", "rabbits eat", "Cat Grass", "Most rabbits", etc., are registered as words, which are units for estimating named entities, including single words and multi-words consisting of two consecutive words respectively.
[0073] Also, in the combination of teacher data and teacher labels as the initial seed, as teacher data, phrases consisting of each of one and multiple words are set and used for the initial learning of the predictor. For example, in the corpora of FIGS. 2A and 2B, taking "dogs" and "Cat Grass" as teacher data, and taking "is a plant (true)" and "is not a plant (false)" as teacher labels in the estimation of named entities, the first interest scores of all words (words consisting of single words and words consisting of two consecutive words) in the corpus are calculated, and a question list is generated.
[0074] Regarding the generation of the question list after the second time, as already described, the predictor is made to perform learning corresponding to the first question list to obtain the estimation result of the predictor. Thereby, the interest score calculation unit 104 corresponds to the interest score generated from the estimation result of the predictor to be learned, and randomly extracts words from the corpus according to the probability of being selected as teacher data, and generates a question list for obtaining the teacher data and teacher labels used for the next learning.
[0075] Through learning using the above-described question list, the predictor to be learned estimates the probability that an entity expression, such as "plant", is present for a single word like "cabbage" or "dogs", and for two consecutive words like "Cat Grass" or "various animals". As a result, as an entity expression extractor, words with the entity expression "plant" can be output as extraction data in the form of [4 (the 4th word), 1 (composed of 1 word (cabbage))] and [9 (the 9th word), 2 (2 words (Cat Grass))] as described above.
[0076] In addition, a program for realizing a function of performing a learning process of an interactive predictor by question / answer using the predictor interactive learning system shown in FIG. 1 is recorded on a computer-readable recording medium, and the program recorded on this recording medium is read into a computer system and executed, thereby performing a process of performing a learning process of an interactive predictor by question / answer. Here, the "computer system" is assumed to include hardware such as an OS and peripheral devices.
[0077] In addition, the "computer system" is assumed to include a homepage providing environment (or display environment) if the WWW system is used. In addition, the "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer system. Further, the "computer-readable recording medium" also includes things that hold a program dynamically for a short time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and things that hold a program for a certain period of time, such as volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the above-described functions, and may also be for realizing the above-described functions in combination with a program already recorded in the computer system.
[0078] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.
Description of Reference Numerals
[0079] 100 Predictor Interactive Learning System 101 Data Input / Output Unit 102 Initial Interest Score Calculation Unit 103 Machine Learning Unit 104 Interest Score Calculation Unit 105 Interactive Learning Frame Unit 106 Question and Answer Unit 107 Corpus Data Storage Unit 108 Rule-Based Question Data Storage Unit 109 Teacher Data Storage Unit 110 Predictor Weight Coefficient Storage Unit
Claims
1. A machine learning unit that performs machine learning on a predictor that outputs a prediction value indicating the probability that an input word is a predetermined named entity using predetermined teacher data and teacher labels; An interest score calculation unit that calculates an interest score from the prediction value output by the predictor for each word in the corpus used for the machine learning and the statistical data related to the word in the corpus; A dialog learning frame unit that extracts, from the corpus, the word that will be the teacher data in the next learning for the predictor based on the interest score; A question-and-answer unit that outputs a question as to whether the extracted teacher data is the named entity for which the predictor predicts the probability, and obtains the teacher label corresponding to the teacher data as a response to the question; Comprising: The machine learning unit: Performs machine learning on the predictor using the teacher data extracted by the dialog learning frame unit and the teacher label obtained by the question-and-answer unit for the teacher data; The interest score calculation unit calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction value; A predictor dialog-style learning system.
2. A machine learning unit that performs machine learning on a predictor that outputs a prediction value indicating the probability that an input word is a predetermined named entity using predetermined teacher data and teacher labels; An interest score calculation unit that calculates an interest score from the prediction value output by the predictor for each word in the corpus used for the machine learning and the statistical data related to the word in the corpus; A dialog learning frame unit that extracts, from the corpus, the word that will be the teacher data in the next learning for the predictor based on the interest score; A question-and-answer unit that outputs a question as to whether the extracted teacher data is the named entity for which the predictor predicts the probability, and obtains the teacher label corresponding to the teacher data as a response to the question; When starting the machine learning of the predictor, a first interest score, which is an interest score for selecting teacher data for the first question from each of the words in the corpus using the combination of the teacher data and the teacher labels input as seeds, is calculated corresponding to a predetermined calculation rule by a first interest score calculation unit; comprising the machine learning unit performs machine learning of the predictor using the teacher data extracted by the dialogic learning frame unit and the teacher label acquired by the question-and-answer unit for the teacher data; the calculation rule is a rule for obtaining the first interest score of each of the words corresponding to the degree of coincidence between other words arranged adjacent to the front and rear stages of the word of the teacher data input as the seed in the corpus and other words arranged adjacent to the front and rear stages of each of the words in the corpus; a predictor dialogic learning system.
3. the dialogic learning frame unit extracts the teacher data corresponding to the first interest score from each of the words in the corpus; the question-and-answer unit acquires a response indicating whether each of the teacher data corresponding to the first interest score is a teacher label corresponding to the teacher data input as a seed. The predictor dialogic learning system according to claim 2.
4. A machine learning process in which a machine learning unit performs machine learning on a predictor that outputs a prediction value indicating the probability that a word input using predetermined teacher data and teacher labels is a predetermined specific expression; An interest score calculation process in which an interest score calculation unit obtains an interest score for each word in the corpus used for the machine learning from the prediction value output by the predictor for the word and the statistical data regarding the word in the corpus; A dialogic learning frame process in which a dialogic learning frame unit extracts the word that will be the teacher data in the next learning for the predictor from the corpus based on the interest score; A question-and-answer process in which a question-and-answer unit outputs a question as to whether the extracted teacher data is the specific expression for which the predictor predicts the probability and acquires the teacher label corresponding to the teacher data as a response to the question including the machine learning unit performs machine learning of the predictor using the teacher data extracted by the dialogic learning frame unit and the teacher label acquired by the question-and-answer unit for the teacher data; The interest score calculation unit calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction values. Predictor interactive learning method.
5. A machine learning process in which a machine learning unit performs machine learning on a predictor that outputs a prediction value indicating the probability that a word input using predetermined teacher data and teacher labels is a predetermined unique expression. An interest score calculation process in which an interest score calculation unit obtains an interest score for each word in the corpus of the machine learning from the prediction value output by the predictor for the word and the statistical data regarding the word in the corpus. A dialogic learning frame process in which a dialogic learning frame unit extracts the word that will be the teacher data in the next learning for the predictor from the corpus based on the interest score. A question and answer process in which a question and answer unit outputs a question as to whether the extracted teacher data is the unique expression for which the predictor predicts the probability, and obtains a teacher label corresponding to the teacher data as a response to the question. A first interest score calculation process in which a first interest score calculation unit calculates, in accordance with a predetermined calculation rule, a first interest score, which is an interest score for selecting teacher data for a first question from each of the words in the corpus using the combination of the teacher data and the teacher label input as a seed when starting the machine learning of the predictor. including the machine learning unit performs machine learning on the predictor using the teacher data extracted by the dialogic learning frame unit and the teacher label obtained by the question and answer unit for the teacher data. The calculation rule is a rule for obtaining the first interest score of each of the words corresponding to the degree of coincidence between other words arranged adjacent to the front and rear of the word of the teacher data input as the seed in the corpus and other words arranged adjacent to the front and rear of each of the words in the corpus. Predictor interactive learning method.
6. A computer Machine learning means for performing machine learning on a predictor that outputs a prediction value indicating the probability that an input word is a predetermined named entity using predetermined teacher data and teacher labels. Interest score calculation means for obtaining an interest score from the prediction value output by the predictor for each word in the corpus used for the machine learning and the statistical data related to the word in the corpus. Dialogue learning framework means for extracting the word to be the teacher data in the next learning for the predictor from the corpus based on the interest score. Question and answer means for outputting a question as to whether the extracted teacher data is the named entity for which the predictor predicts the probability, and obtaining the teacher label corresponding to the teacher data as a response to the question. Function as The machine learning means Performs machine learning on the predictor using the teacher data extracted by the dialogue learning framework means and the teacher label obtained by the question and answer means for the teacher data. The interest score calculation means calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction values. Program.
7. A computer Machine learning means for performing machine learning on a predictor that outputs a prediction value indicating the probability that an input word is a predetermined named entity using predetermined teacher data and teacher labels. Interest score calculation means for obtaining an interest score from the prediction value output by the predictor for each word in the corpus used for the machine learning and the statistical data related to the word in the corpus. Dialogue learning framework means for extracting the word to be the teacher data in the next learning for the predictor from the corpus based on the interest score. Question and answer means for outputting a question as to whether the extracted teacher data is the named entity for which the predictor predicts the probability, and obtaining the teacher label corresponding to the teacher data as a response to the question. When starting the machine learning of the predictor, the first interest score, which is an interest score for selecting teacher data for the first question from each of the words in the corpus using the combination of the teacher data and the teacher label input as a seed, is calculated according to a predetermined calculation rule by first interest score calculation means. Function as The machine learning means Performs machine learning of the predictor using the teacher data extracted by the dialogical learning frame means and the teacher label acquired by the question-and-answer means for the teacher data. The interest score calculation means calculates the interest score by obtaining the average value of each of a first interest score corresponding to the normalized rank of the appearance rate of the word in the corpus, a second interest score corresponding to the prediction value output by the predictor when each of the words in the corpus is input, a third interest score corresponding to the variance of the prediction value, and a fourth interest score corresponding to the average value of the prediction value. Program.
Citation Information
Patent Citations
Language analysis model learning device, language analysis model learning method, language analysis model learning program, and recording medium with the same
JP2008225907A