Learning device, label estimation device and program
The learning device and label estimation device use a trained model to efficiently assign multiple labels to sentences by considering label appropriateness and co-occurrence, reducing the labor and computational effort in label estimation tasks.
Patent Information
- Application Number
- JP2021192928
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-02
- Filing Date
- 2021-11-29
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Existing label estimation methods require significant labor when multiple labels need to be assigned to a sentence, as each additional label increases the complexity and effort required.
A learning device and label estimation device that utilize a label estimation model trained with machine learning, incorporating label appropriateness information and co-occurrence data to estimate appropriate labels for sentences, reducing the computational and manual effort.
The method effectively suppresses the increase in labor required for label estimation by accurately assigning multiple labels with high precision, using a trained model that considers label appropriateness and co-occurrence probabilities.
Smart Images

Figure 0007763643000009 
Figure 0007763643000010 
Figure 0007763643000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device, a label estimation device, and a program. [Background technology]
[0002] It may be desirable to assign labels related to news articles to facilitate searches, etc. For example, if a news article is about a company whose stock price has fluctuated due to an infectious disease, terms such as infectious disease, stock price, and business are assigned as labels. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-53730 [Non-patent literature]
[0004] [Non-Patent Document 1] Grigorios, Tsoumakas, Ioannis Katakis, “Multi-Label Classification: An Overview” [Non-patent document 2] Ankit Pal, Muru Selvakumar and Malaikannan Sankarasubbu, “MAGNET: Multi-Label Text Classification using Attention-based Graph Neural Network” arXiv:2003.11644v1 [Non-patent document 3] Ashutosh Adhikari, Achyudh Ram, Raphael Tang, and Jimmy Lin, “Rethinking Complex Neural Network Architectures for Document Classification” Proceedings of NAACL-HLT 2019, pages 4046-4051 Summary of the Invention [Problem to be solved by the invention]
[0005] Although it may be sufficient to assign one label to a sentence, since sentences often consist of multiple words, there are cases where one label is insufficient. In other words, as shown in the example above, it may be desirable to assign multiple labels to a single sentence. However, the more labels there are to be assigned to a sentence, the greater the effort required for label estimation.
[0006] In view of the above circumstances, an object of the present invention is to provide a technique for suppressing an increase in the labor required for label estimation work. [Means for solving the problem]
[0007] One aspect of the present invention is a learning device that includes a model learning unit that updates a label estimation model, which is a mathematical model that estimates a label to be assigned to a sentence indicated by input sentence information, by a machine learning method using model training data that includes sentence information indicating a sentence and label appropriateness information that indicates the degree to which a plurality of labels, which are predetermined as candidate labels to be assigned to the sentence, are appropriate as labels for the sentence, and the label appropriateness information is information obtained based on true / false information that indicates a label that satisfies a predetermined condition regarding the probability that the label will be assigned to the sentence indicated by the model training data, and label co-occurrence information that is information that indicates the probability of co-occurrence between any two of a plurality of labels, which are predetermined as candidate labels to be assigned to the sentence.
[0008] One aspect of the present invention is a method for processing a document, comprising: an object acquisition unit that acquires object information that indicates a sentence to be processed; The label estimation device includes: a model learning unit that updates a label estimation model, which is a mathematical model that estimates a label to be assigned to a sentence indicated by input sentence information, by a machine learning method using model learning data including sentence information that indicates a sentence and label appropriateness information that indicates the degree to which a plurality of labels, which are predetermined as candidate labels to be assigned to the sentence, are appropriate as labels for the sentence; and an estimation unit that estimates a label to be assigned to a sentence indicated by target information acquired by the sentence acquisition unit, using the trained label estimation model obtained by a learning device, which is information obtained based on true / false information that indicates a label that satisfies a predetermined condition regarding the probability that the label will be assigned to the sentence indicated by the model learning data, and label co-occurrence information that is information that indicates the probability of co-occurrence between any two of a plurality of labels, which are predetermined as candidate labels to be assigned to the sentence.
[0009] One aspect of the present invention is a program for causing a computer to function as the learning device described above.
[0010] One aspect of the present invention is a program for causing a computer to function as the label estimation device described above. [Effects of the Invention]
[0011] The present invention makes it possible to provide a technique for suppressing an increase in the labor required for label estimation work. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is an explanatory diagram illustrating a label estimation system according to an embodiment. [Figure 2] FIG. 4 is a diagram showing an example of label co-occurrence information according to the embodiment. [Figure 3] 5A to 5C are diagrams illustrating an example of label suitability information generation processing according to an embodiment. [Figure 4]FIG. 1 is a diagram showing an example of the hardware configuration of a learning device 1 according to an embodiment. [Figure 5] FIG. 2 is a diagram showing an example of the configuration of a control unit 11 in the embodiment. [Figure 6] 4 is a flowchart showing an example of the flow of processing executed by the learning device 1 in the embodiment. [Figure 7] FIG. 2 is a diagram showing an example of the hardware configuration of a label estimation device 2 according to the embodiment. [Figure 8] FIG. 2 is a diagram showing an example of the configuration of a control unit 21 in the embodiment. [Figure 9] 4 is a flowchart showing an example of the flow of processing executed by the label estimation device 2 in the embodiment. [Figure 10] FIG. 1 is a first diagram showing an example of experimental results using the label estimation system according to the embodiment. [Figure 11] FIG. 2 is a second diagram showing an example of experimental results using the label estimation system according to the embodiment. [Figure 12] FIG. 3 is a third diagram showing an example of experimental results using the label estimation system according to the embodiment. [Figure 13] FIG. 10 is a diagram showing an example of the configuration of a control unit in a modified example. DETAILED DESCRIPTION OF THE INVENTION
[0013] (Embodiment) FIG. 1 is an explanatory diagram illustrating a label estimation system 100 according to an embodiment. The label estimation system 100 is a system that estimates a label to be assigned to a sentence. The label estimation system 100 obtains a mathematical model for estimating a label to be assigned to a sentence by a machine learning method. The label estimation system 100 uses the obtained mathematical model to estimate a label to be assigned to an input sentence.
[0014] More specifically, when text information is input, the label estimation system 100, having acquired the mathematical model, uses the acquired mathematical model to estimate a label to be assigned to a sentence indicated by the text information, based on the text information. The text information is information indicating a sentence. The label estimation system 100 will be described in more detail. The label estimation system 100 includes a learning device 1 and a label estimation device 2.
[0015] The learning device 1 obtains a trained label estimation model by updating the label estimation model using a machine learning method. The label estimation model is a mathematical model that estimates a label to be assigned to a sentence (hereinafter referred to as a "target sentence") indicated by input sentence information based on input sentence information, and is a mathematical model before a predetermined termination condition for learning is satisfied.
[0016] More specifically, the result estimated by the label estimation model is label appropriateness information. The label appropriateness information is information indicating the label appropriateness for each label candidate. The label appropriateness is the degree to which a label is appropriate as a label for the target sentence. The label candidates are each of a plurality of labels that are predetermined as candidates for the label to be assigned to the target sentence. The label candidates are terms that can be associated with the target sentence, such as "infectious disease," "business," "sports," and "stock prices."
[0017] A trained label estimation model is a label estimation model at the point in time when a predetermined termination condition for learning (hereinafter referred to as a "learning termination condition") is satisfied. The learning termination condition is, for example, a condition that a change in the label estimation model due to learning is smaller than a predetermined change. The learning termination condition may also be, for example, a condition that the number of times learning has been performed reaches a predetermined number.
[0018] Hereinafter, the process by which the learning device 1 obtains a trained label estimation model is referred to as a model learning process. The machine learning method may be any method that can obtain a trained label estimation model. The machine learning method may be, for example, a method that uses CNN (Convolutional Neural Networks), a method that uses LSTM (Long short-term memory), or a method that uses BERT (Bidirectional Encoder Representations from Transformers).
[0019] In a machine learning method for obtaining a trained label estimation model, data having text information as an explanatory variable is used. A response variable corresponding to the explanatory variable indicates label appropriateness information. Hereinafter, data having text information as an explanatory variable and label appropriateness information as a response variable is referred to as model training data. Model training data is data used to obtain a trained label estimation model. In other words, model training data is data used to train a label estimation model. Hereinafter, the label appropriateness information contained in the model training data is referred to as training data.
[0020] <Examples of label appropriateness information> A specific example of representation of label appropriateness information will be described. When there are N label candidates (N is a natural number), the label appropriateness information is represented, for example, by an N-dimensional vector. Each element of the N-dimensional vector corresponds to one of the N label candidates, and elements with different indexes n correspond to different label candidates. n is a natural number between 1 and N. Note that index n is an index that distinguishes between label candidates and also indicates the n-th element of the N-dimensional vector. For ease of explanation, the label estimation system 100 will be described below taking as an example a case where there are N label candidates.
[0021] Each element of the N-dimensional vector representing label appropriateness information indicates the label appropriateness of each corresponding label candidate. The label appropriateness is expressed, for example, as a value between 0 and 1. In such a case, the closer the value of each element of the N-dimensional vector representing label appropriateness information is to 0, for example, the more inappropriate the corresponding label candidate is as a label for the sentence indicated by the sentence information. On the other hand, the closer the value of each element of the N-dimensional vector representing label appropriateness information is to 1, for example, the more appropriate the corresponding label candidate is as a label for the sentence indicated by the sentence information.
[0022] <Model learning process and loss function> The model learning process will be described in more detail. As described above, the model learning process is a process of updating the label estimation model by a machine learning method using model learning data until the learning termination condition is met. In the model learning process, the label estimation model is updated so as to reduce the loss calculated using a loss function. Note that the loss calculated using the loss function is the value of the loss function, and is, for example, a value that represents the degree of mismatch between the output of the label estimation model and the training data.
[0023] The loss function is an index expressed using the degree of agreement and the degree of disagreement between the training data and the estimation results of the label estimation model. The loss function may be, for example, the binary cross-entropy defined by the following equation (1).
[0024]
number
[0025] y n is the label appropriateness indicated by the training data and indicates the label appropriateness of the label candidate with index n. n is the label appropriateness estimated by the label estimation model, and indicates the label appropriateness of the label candidate with index n. Note that A{^} indicates a symbol in which a circumflex is added to the symbol A. Therefore, y{^} nmeans the symbol y with a circumflex and a subscript n. More specifically, y{^} n means the symbol in the following formula (2).
[0026]
number
[0027] The term expressed by the following equation (3) in equation (1) indicates the quantitative difference between the estimation result of the label estimation model and the training data when the estimation result of the label estimation model and the training data qualitatively match.
[0028]
number
[0029] The term expressed by the following equation (4) in equation (1) indicates the quantitative difference between the estimation result of the label estimation model and the training data when there is a qualitative mismatch between the estimation result of the label estimation model and the training data.
[0030]
number
[0031] A specific example of updating to reduce the loss function is a process of updating the label estimation model using the loss function shown in Equation (1) so as not to increase the probability that the label estimation model will estimate incorrect label information. Incorrect label information is label appropriateness information that qualitatively does not match the training data.
[0032] The label estimation device 2 estimates a label to be assigned to a target sentence indicated by the input text information, using the trained label estimation model acquired by the learning device 1. More specifically, the label estimation device 2 estimates the label appropriateness of each label candidate for the target sentence indicated by the input text information, using the trained label estimation model acquired by the learning device 1.
[0033] <Generating label appropriateness information included in model training data> An example of a method for generating label appropriateness information included in model learning data will be described below. The label appropriateness information is generated, for example, manually or by a device, based on correct / incorrect information and label co-occurrence information.
[0034] The correct / incorrect information is information indicating a label that satisfies a predetermined condition regarding the likelihood of being assigned to a sentence indicated by sentence information included in the model training data. In other words, the correct / incorrect information is information indicating a label that satisfies a predetermined condition regarding the likelihood of label appropriateness for a sentence indicated by the model training data. The predetermined condition regarding the likelihood of label appropriateness (hereinafter referred to as the "label appropriateness condition") is, for example, a condition that the label appropriateness is the highest. If there are multiple labels that satisfy the label appropriateness condition, the correct / incorrect information may indicate multiple labels. The correct / incorrect information is expressed, for example, as an N-dimensional vector in which only the element corresponding to the label with the highest likelihood of being assigned has a value of 1 and the other elements have values of 0. For ease of explanation, the label estimation system 100 will be described below using an example in which the label appropriateness condition is a condition that the label appropriateness is the highest.
[0035] Label co-occurrence information is information indicating the probability of co-occurrence between any two label candidates out of N label candidates. Specifically, the probability of co-occurrence is the probability that one label candidate appears in a sentence when the other label candidate appears in the sentence. Note that label co-occurrence information may also indicate the probability of co-occurrence between identical label candidates. Since the probability of co-occurrence between identical label candidates refers to autocorrelation, the probability of co-occurrence between identical label candidates is 1. Note that label co-occurrence information does not necessarily need to indicate the probability of co-occurrence between identical label candidates, and in such cases, the probability of co-occurrence between identical label candidates indicated by the label co-occurrence information is, for example, 0.
[0036] 2 is a diagram showing an example of label co-occurrence information in an embodiment. The label co-occurrence information is expressed, for example, as a positive definite matrix whose element values are between 0 and 1. In the example of FIG. 2, the columns and rows each represent label candidates, and the diagonal components represent autocorrelation.
[0037] More specifically, the label co-occurrence information in Fig. 2 is a matrix indicating the Positive Pointwise Mutual Information (PPMI) scores between label candidates. The PPMI score is defined by the following equation (5).
[0038]
number
[0039] In equation (5), l n denotes the label candidate for index n, and l m indicates a label candidate for index m. Note that m is an integer between 1 and N. m may be the same value as n or may be different. C(l n ) indicates the number of occurrences of the label candidate of index n in a set of predetermined sentences prepared in advance (hereinafter referred to as the "pre-sentence set"). m ) indicates the number of occurrences of the label candidate with index m in the pre-set sentences. n , lm) indicates the number of co-occurrences between the label candidate of index n and the label candidate of index m in the pre-sentence set.
[0040] Hereinafter, the process of generating label appropriateness information based on correct / incorrect information and label co-occurrence information is referred to as the label appropriateness information generation process. In the label appropriateness information generation process, for example, for a label candidate whose element value of a vector representing correct / incorrect information is 1, a process is executed in which the probability of co-occurrence with other label candidates is obtained using the label co-occurrence information. When there are multiple label candidates whose element value of a vector representing correct / incorrect information is 1, for example, the probability of co-occurrence with other label candidates is obtained for the multiple label candidates whose element value is 1, and the sum of the co-occurrence probabilities for each of the other label candidates is calculated. Next, in the label appropriateness information generation process, a process is executed in which the probability of co-occurrence with other label candidates is converted to a value between 0 and 1 using a function such as a sigmoid function that limits the value of an independent variable to a predetermined value between 0 and 1.
[0041] In the label appropriateness information generation process, a process is executed in which the element values of the vector representing the correct / incorrect information, which have a value of 0, are replaced with converted values (hereinafter referred to as "replacement process"). The correct / incorrect information in which the element values have been changed by the replacement process is the label appropriateness information.
[0042] Fig. 3 is a diagram illustrating an example of label appropriateness information generation processing in an embodiment. More specifically, Fig. 3 is an explanatory diagram illustrating an example of label appropriateness information generation processing in a case where label co-occurrence information is a matrix indicating PPMI scores of label candidates (hereinafter referred to as "PPMI matrix").
[0043] FIG. 3 shows images G1 to G5. Image G1 shows an example of correct / incorrect information. The correct / incorrect information for image G1 shows five label candidates: "sports," "business," "health," "vaccine," and "infectious disease." Image G1 shows that "business" and "infectious disease" have the highest label appropriateness. In FIG. 3, "business" and "infectious disease" are label candidates that satisfy the label appropriateness conditions.
[0044] Image G2 shows an example of a PPMI matrix. Image G3 shows a sigmoid function. Image G4 shows the process of obtaining the vector sum of rows of the PPMI matrix that satisfy the label appropriateness condition. Specifically, it shows the process of obtaining the vector sum of the row indicating the probability that a label candidate co-occurs with the label candidate "business" and the row indicating the probability that a label candidate co-occurs with the label candidate "infectious disease." Image G5 shows an example of label appropriateness information.
[0045] In the label appropriateness information generation process, as shown in image G4, a process of adding up rows in a PPMI matrix that indicate the probability of co-occurrence between a correct label and a label candidate (hereinafter referred to as a "main co-occurrence row") is executed. Hereinafter, the process of adding up main co-occurrence rows in a PPMI matrix is referred to as the addition process. In the example of FIG. 3, "business" and "infectious disease" are correct labels, and the process of adding up the row of "business" and the row of "infectious disease" is the addition process. A correct label is a label candidate that corresponds to an element with a value of 1 among the label candidates that correspond to the elements of an N-dimensional vector that indicates correctness information. In other words, a correct label is a label candidate that satisfies the label appropriateness condition.
[0046] By executing the summation process, the PPMI scores of incorrect labels that tend to co-occur with the correct label are added together for each incorrect label. The information obtained as a result of the summation is expressed, for example, as an N-dimensional vector. In the example of Figure 3, the summation process adds together the PPMI scores of incorrect labels that tend to co-occur with "business" and the PPMI scores of incorrect labels that tend to co-occur with "infectious disease."
[0047] An incorrect label is a label candidate that corresponds to an element with a value of 0 among the label candidates corresponding to the elements of the N-dimensional vector that indicates correctness information. In other words, an incorrect label is a label candidate that is not a correct label among the label candidates. In the example of Figure 3, the incorrect labels are "sports", "health", and "vaccine".
[0048] In the label appropriateness information generation process, next, the N-dimensional vector obtained by executing the addition process (hereinafter referred to as the "addition result vector") is input to the sigmoid function shown in image G3, whereby the value of each element of the addition result vector is normalized to a value between 0 and 1. Hereinafter, the process of normalizing the value of each element of the addition result vector to a value between 0 and 1 by inputting the addition result vector to the sigmoid function is referred to as the first normalization process. The first normalization process is defined by the following equation (6).
[0049]
number
[0050] P nm denotes the PPMI matrix. σ(·) represents the sigmoid function.
[0051] In the label appropriateness information generation process, after the first normalization process is performed, the obtained score S n and vector y n In the label appropriateness information generation process, after the first normalization process is performed, the normalization process (hereinafter referred to as the second normalization process) shown in the following equations (7) and (8) is also performed.
[0052]
number
[0053]
number
[0054] p´ nm means the element in the nth row and mth column of the PPMI matrix P'. m is a vector whose row index indicates the probability of co-occurrence between the correct label and the candidate label of m. nis a vector. α is a hyperparameter (a coefficient between 0 and 1) that indicates the strength of smoothing. α is, for example, a coefficient between 0 and 1 that is called a scaling rate. In this way, the replacement process includes the addition process, the first normalization process, the smoothing process, and the second normalization process. The y obtained in this way is n is an example of label suitability information.
[0055] 4 is a diagram showing an example of the hardware configuration of a learning device 1 according to an embodiment. The learning device 1 includes a control unit 11 having a processor 91, such as a CPU (Central Processing Unit), and a memory 92 connected via a bus, and executes a program. By executing the program, the learning device 1 functions as a device including the control unit 11, an input unit 12, a communication unit 13, a memory unit 14, and an output unit 15.
[0056] More specifically, the processor 91 reads out a program stored in the storage unit 14 and stores the read out program in the memory 92. When the processor 91 executes the program stored in the memory 92, the learning device 1 functions as a device including a control unit 11, an input unit 12, a communication unit 13, a storage unit 14, and an output unit 15.
[0057] The control unit 11 controls the operation of various functional units included in the learning device 1. The control unit 11 executes, for example, a model learning process. The control unit 11 may also execute, for example, a label appropriateness information generation process. As described above, the label appropriateness information generation process may be performed manually, or may be executed by a device. Below, the label estimation system 100 will be described using an example in which the learning device 1 executes the label appropriateness information generation process.
[0058] The control unit 11 controls, for example, the operation of the output unit 15. The control unit 11 records, for example, various pieces of information generated by the execution of the model learning process in the storage unit 14. The control unit 11 records, for example, the obtained label appropriateness information in the storage unit 14.
[0059] Input unit 12 includes input devices such as a mouse, keyboard, and touch panel. Input unit 12 may be configured as an interface that connects these input devices to learning device 1. Input unit 12 accepts input of various types of information to learning device 1.
[0060] The communication unit 13 includes a communication interface for connecting the learning device 1 to an external device. The communication unit 13 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits true / false information. The external device is, for example, a device that transmits label co-occurrence information. The external device is, for example, a device that transmits model training data. The external device is, for example, the label estimation device 2. Note that the true / false information, label co-occurrence information, and model training data do not necessarily need to be input via the communication unit 13, and may be input to the input unit 12.
[0061] The storage unit 14 is configured using a computer-readable storage medium device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 14 stores various information related to the learning device 1. The storage unit 14 stores information input via the input unit 12 or the communication unit 13, for example. The storage unit 14 stores various information generated by executing a model learning process, for example. The storage unit 14 stores label appropriateness information, for example. The storage unit 14 stores a label estimation model in advance. The storage unit 14 may also store an obtained trained label estimation model.
[0062] The output unit 15 outputs various types of information. The output unit 15 includes a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display. The output unit 15 may be configured as an interface that connects these display devices to the learning device 1. The output unit 15 outputs, for example, information input to the input unit 12. The output unit 15 may display, for example, the execution results of the model learning process. The output unit 15 may display, for example, label suitability information.
[0063] 5 is a diagram showing an example of the configuration of the control unit 11 in the embodiment. The control unit 11 includes a label appropriateness information acquisition unit 110, a model learning unit 120, a storage control unit 130, a communication control unit 140, and an output control unit 150.
[0064] The label appropriateness information acquisition unit 110 acquires label appropriateness information. The label appropriateness information acquisition unit 110 acquires label appropriateness information by executing a label appropriateness information generation process based on the correct / incorrect information and label co-occurrence information input to the input unit 12 or the communication unit 13.
[0065] The model learning unit 120 updates the label estimation model until the learning termination condition is satisfied, using the label appropriateness information and the model learning data input to the input unit 12 or the communication unit 13. That is, the model learning unit 120 obtains a learned label estimation model by performing a model learning process using the label appropriateness information and the model learning data input to the input unit 12 or the communication unit 13.
[0066] The storage control unit 130 records various information in the storage unit 14. The communication control unit 140 controls the operation of the communication unit 13. The output control unit 150 controls the operation of the output unit 15.
[0067] 6 is a flowchart showing an example of the flow of processing executed by the learning device 1 in the embodiment. The label appropriateness information acquisition unit 110 acquires label appropriateness information (step S101). Next, model training data is input to the input unit or communication unit 13 (step S102). Next, the model training unit 120 estimates label appropriateness information by inputting sentence information indicated by the model training data into a label estimation model (step S103). Next, the model training unit 120 updates the label estimation model based on the label appropriateness information included in the model training data and the estimation result of step S103 (step S104). Next, the model training unit 120 determines whether a learning termination condition is satisfied (step S105). If the learning termination condition is satisfied (step S105: YES), the processing ends. On the other hand, if the learning termination condition is not satisfied (step S105: NO), the processing returns to step S102.
[0068] The processes from step S102 to step S105, which are repeated until the learning end condition is satisfied, are an example of a model learning process.
[0069] 7 is a diagram showing an example of the hardware configuration of a label estimation device 2 in an embodiment. The label estimation device 2 includes a control unit 21 having a processor 93 such as a CPU and a memory 94 connected by a bus, and executes a program. By executing the program, the label estimation device 2 functions as a device including the control unit 21, an input unit 22, a communication unit 23, a storage unit 24, and an output unit 25.
[0070] More specifically, the processor 93 reads out a program stored in the storage unit 24 and stores the read program in the memory 94. When the processor 93 executes the program stored in the memory 94, the label estimation device 2 functions as a device including a control unit 21, an input unit 22, a communication unit 23, a storage unit 24, and an output unit 25.
[0071] The control unit 21 controls the operation of various functional units included in the label estimation device 2. The control unit 21 executes, for example, a trained label estimation model. The control unit 21 controls, for example, the operation of the output unit 25. The control unit 21 records, for example, various pieces of information generated by the execution of the trained label estimation model in the storage unit 24.
[0072] The input unit 22 includes input devices such as a mouse, a keyboard, and a touch panel. The input unit 22 may be configured as an interface that connects these input devices to the label estimation device 2. The input unit 22 accepts input of various information to the label estimation device 2.
[0073] The communication unit 23 includes a communication interface for connecting the label estimation device 2 to an external device. The communication unit 23 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits text information. The external device is, for example, the learning device 1. The communication unit 23 acquires a trained label estimation model by communicating with the learning device 1. Note that the text information does not necessarily need to be input to the communication unit 23, and may be input to the input unit 22.
[0074] The storage unit 24 is configured using a computer-readable storage medium device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 24 stores various information related to the label estimation device 2. The storage unit 24 stores information input via, for example, the input unit 22 or the communication unit 23. The storage unit 24 stores various information generated by executing, for example, a trained label estimation model. The storage unit 24 stores the trained label estimation model.
[0075] The output unit 25 outputs various types of information. The output unit 25 includes a display device such as a CRT display, a liquid crystal display, or an organic EL display. The output unit 25 may be configured as an interface that connects these display devices to the label estimation device 2. The output unit 25 outputs, for example, information input to the input unit 22. The output unit 25 may also display, for example, the execution result of the trained label estimation model.
[0076] 8 is a diagram showing an example of the configuration of the control unit 21 in the embodiment. The control unit 21 includes an object acquisition unit 210, an estimation unit 220, a memory control unit 230, a communication control unit 240, and an output control unit 250. The object acquisition unit 210 acquires text information input to the input unit 22 or the communication unit 23.
[0077] The estimation unit 220 executes the trained label estimation model on the text information acquired by the object acquisition unit 210. The estimation unit 220 estimates label appropriateness information for the text information acquired by the object acquisition unit 210 by executing the trained label estimation model.
[0078] The storage control unit 230 records various information in the storage unit 24. The communication control unit 240 controls the operation of the communication unit 23. The output control unit 250 controls the operation of the output unit 25.
[0079] 9 is a flowchart showing an example of the flow of processing executed by the label estimation device 2 in the embodiment. The object acquisition unit 210 acquires text information input to the input unit 22 or the communication unit 23 (step S201). Next, the estimation unit 220 executes the trained label estimation model to estimate label appropriateness information for the text information acquired by the object acquisition unit 210 (step S202). Next, the output control unit 250 controls the operation of the output unit 25 to cause the output unit 25 to output the acquired label appropriateness information (step S203).
[0080] (Experimental results) Here, the results of an experiment using the label estimation system 100 will be described. In the experiment, benchmarks used in multi-label classification were used as datasets. Specifically, three datasets were used: Reuters-21578, Arxiv Academic Paper Dataset (AAPD), and 20Newsgroups. In the experiment, machine learning models used in natural language processing were used as machine learning models. Specifically, BERT (Bidirectional Encoder Representations from Transformers), Bi-LSTM (Long Short Term Memory), and CNN (Convolution Neural Network) were used. In the experiment, Micro-f1 and Macro-f1 were used as evaluation metrics.
[0081] Fig. 10 is a first diagram showing an example of experimental results using the label estimation system 100 of the embodiment. In Fig. 10, the rows with "BERT w / ALS," "LSTM w / ALS," and "CNN w / ALS" in the "Method" column indicate results using the label estimation system 100. The rows with "BERT only," "LSTM only," and "CNN only" in the "Method" column indicate estimation results using a trained label estimation model obtained using correct / incorrect information without using label appropriateness information.
[0082] Note that "BERT" in "BERT w / ALS" and "BERT only" indicates that the machine learning model used in the experiment was BERT. Note that "LSTM" in "LSTM w / ALS" and "LSTM only" indicates that the machine learning model used in the experiment was Bi-LSTM. Note that "CNN" in "CNN w / ALS" and "CNN only" indicates that the machine learning model used in the experiment was CNN.
[0083] "Macro-f1" in "Rueters-21578" indicates the value of Macor-f1 when the dataset used is Reuters-21578. "Micro-f1" in "Rueters-21578" indicates the value of Micor-f1 when the dataset used is Reuters-21578. "Macro-f1" in "AAPD" indicates the value of Macor-f1 when the dataset used is AAPD. "Micro-f1" in "AAPD" indicates the value of Micor-f1 when the dataset used is AAPD. "Macro-f1" in "20Newsgroups" indicates the value of Macor-f1 when the dataset used is 20Newsgroups. "Micro-f1" in "20Newsgroups" indicates the value of Micor-f1 when the dataset used is 20Newsgroups.
[0084] The results in Fig. 10 show the results of five experiments conducted with different random seeds. The numbers in parentheses in Fig. 10 indicate standard deviations. The results in Fig. 10 show that the label estimation system 100 can estimate labels with high accuracy regardless of a specific machine learning model such as CNN or Bi-LSTM.
[0085] FIG. 11 is a second diagram showing an example of experimental results using the label estimation system 100 of the embodiment. More specifically, FIG. 11 shows the results of an experimental evaluation of the accuracy of estimation of low-frequency label candidates. Note that FIG. 10 shows the results of an experimental evaluation of the accuracy of estimation of both low-frequency label candidates and non-low-frequency label candidates. Note that a low-frequency label candidate refers to a label candidate that, among multiple label candidates, is ranked below the median in terms of the number of times it appears in a dataset.
[0086] 11, the rows with "BERT w / ALS," "LSTM w / ALS," and "CNN w / ALS" in the "Method" column indicate the results using the label estimation system 100. The rows with "BERT only," "LSTM only," and "CNN only" in the "Method" column indicate the results of estimation using a trained label estimation model obtained using correct / incorrect information without using label appropriateness information.
[0087] Note that "BERT" in "BERT w / ALS" and "BERT only" indicates that the machine learning model used in the experiment was BERT. Note that "LSTM" in "LSTM w / ALS" and "LSTM only" indicates that the machine learning model used in the experiment was Bi-LSTM. Note that "CNN" in "CNN w / ALS" and "CNN only" indicates that the machine learning model used in the experiment was CNN.
[0088] "Macro-f1" in "Rueters-21578" indicates the value of Macor-f1 when the dataset used is Reuters-21578. "Micro-f1" in "Rueters-21578" indicates the value of Micor-f1 when the dataset used is Reuters-21578. "Macro-f1" in "AAPD" indicates the value of Macor-f1 when the dataset used is AAPD. "Micro-f1" in "AAPD" indicates the value of Micor-f1 when the dataset used is AAPD. "Macro-f1" in "20Newsgroups" indicates the value of Macor-f1 when the dataset used is 20Newsgroups. "Micro-f1" in "20Newsgroups" indicates the value of Micor-f1 when the dataset used is 20Newsgroups.
[0089] The results in Fig. 11 show the results of five experiments conducted with different random seeds. The numbers in parentheses in Fig. 11 indicate standard deviations. The results in Fig. 11 show that the label estimation system 100 can estimate low-frequency label candidates with high accuracy, regardless of a specific machine learning model such as CNN or Bi-LSTM.
[0090] FIG. 12 is a third diagram showing an example of experimental results using the label estimation system 100 of the embodiment. The horizontal axis of FIG. 12 indicates the number of training rounds. The vertical axis of FIG. 12 indicates the Micro-f1 value. "CNN only (train)" indicates the result of estimation of training data using a trained label estimation model obtained using correct / incorrect information without using label appropriateness information. "CNN only (valid)" indicates the result of estimation of development data using a trained label estimation model obtained using correct / incorrect information without using label appropriateness information. "CNN with ALS (train)" indicates the result of estimation of training data using the label estimation system 100. "CNN with ALS (valid)" indicates the result of estimation of development data using the label estimation system 100. Note that the development data is training data used in learning in an experiment to measure the estimation accuracy of the label estimation system 100 for each learning session.
[0091] The results in FIG. 12 show that learning by the label estimation system 100 can suppress overlearning in the early stages of learning compared to learning using correct / incorrect information without using label appropriateness information.
[0092] The learning device 1 in this embodiment configured as described above obtains a trained label estimation model using label appropriateness information. Therefore, it is possible to obtain a mathematical model that estimates more labels to be assigned with high accuracy than a device that obtains a trained label estimation model based only on correct / incorrect information rather than label appropriateness information. As a result, the learning device 1 can suppress an increase in the labor required for the task of estimating labels.
[0093] Furthermore, the label estimation device 2 in this embodiment configured as above estimates labels to be assigned to sentences using a trained label estimation model obtained using label appropriateness information. Therefore, compared to a device that obtains a trained label estimation model based only on correct / incorrect information rather than label appropriateness information, it can estimate more labels to be assigned with high accuracy. As a result, the label estimation device 2 can suppress an increase in the labor required for the task of estimating labels.
[0094] (Variation) As described above, the label suitability information may be generated manually. In such a case, the label suitability information is input to the input unit 12 or the communication unit 13 instead of the correct / incorrect information and the label co-occurrence information. In such a case, the label suitability information acquisition unit 110 acquires the label suitability information input to the input unit 12 or the communication unit 13, instead of acquiring the label suitability information based on the correct / incorrect information and the label co-occurrence information.
[0095] The output control unit 250 may output a predetermined label from the label appropriateness information of the estimation result of the estimation unit 220 to the output unit 25 as a genre.
[0096] As described above, the label appropriateness indicated by the label appropriateness information is, for example, expressed as a value between 0 and 1. However, the label appropriateness does not necessarily have to be expressed as a value between 0 and 1. Therefore, the value of each element of the N-dimensional vector representing the label appropriateness information may include a negative value. An example of the label appropriateness information generation process was described above, and the process used therein was an example of a process using a function, such as a sigmoid function, that limits the value of an independent variable to a predetermined value between 0 and 1. This is an example of a process using an example where the label appropriateness is expressed as a value between 0 and 1. Therefore, if the label appropriateness does not need to be a value between 0 and 1, there is no need to perform a process using a function, such as a sigmoid function, that limits the value of an independent variable to a predetermined value between 0 and 1.
[0097] The control unit 21 may further include either or both of a sentence similarity estimation unit 260 and a key phrase extraction unit 270. Hereinafter, the control unit 21 including the sentence similarity estimation unit 260 and the key phrase extraction unit 270 will be referred to as control unit 21a. FIG. 13 is a diagram showing an example of the configuration of the control unit 21a in a modified example. The control unit 21a differs from the control unit 21 in that it includes the sentence similarity estimation unit 260 and the key phrase extraction unit 270.
[0098] The text similarity estimation unit 260 estimates the degree of similarity between two pieces of text information (hereinafter referred to as "text similarity"). At least one of the two pieces of text information is text information acquired by the target acquisition unit 210. Therefore, both pieces of text information may be text information acquired by the target acquisition unit 210, or one may be text information acquired by the target acquisition unit 210 and the other may be text information previously stored in the storage unit 24. The text similarity estimation unit 260 inputs the two pieces of text information to the estimation unit 220 and causes the estimation unit 220 to estimate label appropriateness information for both pieces of text information. Based on the two pieces of label appropriateness information estimated by the estimation unit 220, the text similarity estimation unit 260 acquires the degree of agreement of the label appropriateness information as the text similarity between the two pieces of text information. For example, the text similarity estimation unit 260 acquires the value of the inner product of each vector corresponding to each of the two pieces of label appropriateness information as the text similarity.
[0099] The text similarity estimation unit 260 may determine that the two pieces of text information are similar when the acquired text similarity is equal to or greater than a predetermined level. In such a case, the output control unit 250 may cause the output unit 25 to output one or both of the two pieces of text information determined by the text similarity estimation unit 260 to be similar.
[0100] The key phrase extraction unit 270 acquires key phrases in the text indicated by the text information using GiNZA, an open source library for Japanese natural language processing. The output control unit 250 may cause the output unit 25 to output the key phrases acquired by the key phrase extraction unit 270.
[0101] The learning device 1 may be implemented using multiple information processing devices connected to each other via a network, in which case the functional units of the learning device 1 may be distributed and implemented across the multiple information processing devices.
[0102] The label estimation device 2 may be implemented using a plurality of information processing devices communicably connected via a network. In this case, the functional units of the label estimation device 2 may be distributed and implemented in the plurality of information processing devices.
[0103] Note that all or part of the functions of the learning device 1 and the label estimation device 2 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may be transmitted via a telecommunications line.
[0104] The text information acquired by the object acquiring unit 210 is an example of object information. The text indicated by the text information acquired by the object acquiring unit 210 is an example of a processing object.
[0105] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]
[0106] 100...label estimation system, 1...learning device, 2...label estimation device, 11...control unit, 12...input unit, 13...communication unit, 14...memory unit, 15...output unit, 110...label appropriateness information acquisition unit, 120...model learning unit, 130...memory control unit, 140...communication control unit, 150...output control unit, 21...control unit, 22...input unit, 23...communication unit, 24...memory unit, 25...output unit, 210...object acquisition unit, 220...estimation unit, 230...memory control unit, 240...communication control unit, 250...output control unit, 91...processor, 92...memory, 93...processor, 94...memory, 21a...control unit, 260...sentence similarity estimation unit, 270...key phrase extraction unit
Claims
1. a model learning unit that updates a label estimation model, which is a mathematical model that estimates a label to be assigned to a sentence indicated by input text information, by a machine learning method using model learning data that includes text information indicating a sentence and label appropriateness information indicating the degree to which a plurality of labels predetermined as candidates for labels to be assigned to the sentence are appropriate as labels for the sentence; Equipped with The label appropriateness information is information obtained based on true / false information indicating a label that satisfies a predetermined condition regarding the probability of being assigned to a sentence indicated by the model training data, and label co-occurrence information that is information indicating the probability of co-occurrence between any two of a plurality of labels that are predetermined as candidates for labels to be assigned to the sentence. Learning device.
2. The predetermined condition is a condition that the probability of being assigned to a sentence represented by the model training data is the highest. The learning device according to claim 1 .
3. an object acquisition unit that acquires object information that indicates a sentence to be processed; a model learning unit that updates a label estimation model, which is a mathematical model that estimates a label to be assigned to a sentence indicated by input sentence information, by a machine learning method using model learning data including sentence information indicating a sentence and label appropriateness information indicating the degree to which a plurality of labels predetermined as candidate labels to be assigned to the sentence are appropriate as labels for the sentence, wherein the label appropriateness information is information obtained based on true / false information indicating a label that satisfies a predetermined condition regarding the likelihood of being assigned to the sentence indicated by the model learning data, and label co-occurrence information, which is information indicating the probability of co-occurrence between any two of a plurality of labels predetermined as candidate labels to be assigned to the sentence; and an estimation unit that estimates a label to be assigned to the sentence indicated by the target information acquired by the target acquisition unit, using the trained label estimation model obtained by a learning device, the label appropriateness information being information obtained based on true / false information indicating a label that satisfies a predetermined condition regarding the likelihood of being assigned to the sentence indicated by the model learning data, and label co-occurrence information, which is information indicating the probability of co-occurrence between any two of a plurality of labels predetermined as candidate labels to be assigned to the sentence. A label estimation device comprising:
4. A program for causing a computer to function as the learning device according to claim 1 or 2.
5. A program for causing a computer to function as the label estimation device according to claim 3.
Citation Information
Patent Citations
Article theme component decomposition method and device, equipment and storage medium
CN109918641A
Learning device, method for learning, program parameter, and learning program
JP2018028872A
Information processing apparatus and information processing program
JP2018185601A
Deep-learning learning method for category classification of documents and system for the same
JP2019053730A
Multi-label classification using a learned combination of base classifiers
US20110302111A1