Learning device, estimation device, learning method, and program
The learning device efficiently applies knowledge distillation to complex classification problems by using a teacher model with hierarchical functions and targeted loss updates, enhancing the student model's accuracy in capturing short-term and long-term contexts.
Patent Information
- Application Number
- JP2023541156
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2041-08-10
AI Technical Summary
Existing knowledge distillation techniques are not effectively applied to complex classification problems involving complex contexts, particularly in resource-constrained environments like mobile devices, where high classification accuracy is needed.
A learning device and method that utilizes a teacher model with hierarchical Q functions to estimate labels, incorporating a student model that mimics the teacher model's hierarchical structure, using hard and soft target losses to update parameters, ensuring the student model accurately captures both short-term and long-term context.
Enables the learning of a student model with high classification accuracy by efficiently distilling knowledge from a teacher model, even in resource-constrained environments, improving labeling accuracy in complex contexts.
Smart Images

Figure 0007709642000005 
Figure 0007709642000006 
Figure 0007709642000007
Abstract
Description
Technical Field
[0001] The present invention relates to a speech sequence labeling technique that takes a text sequence as input and outputs a label corresponding to the text sequence.
Background Art
[0002] In recent years, for the purpose of understanding conversations and dialogues, a technique of speech sequence labeling has been proposed that takes a speech sequence as input and estimates a label corresponding to the response scene of a conversation or dialogue for each speech.
[0003] For example, in Non-Patent Document 1, a model using a deep neural network (hereinafter also referred to as a "labeling model") is provided to realize speech sequence labeling that takes as input the text of the speech recognition results of an operator and a customer in a contact center and estimates a label for any one of the response scenes of opening, matter grasping, identity confirmation, response, and closing for each speech. According to Non-Patent Document 1, the labeling model is configured as shown in the schematic diagram of FIG. 1, stacking a network that understands short-term context at the word unit (hereinafter also referred to as a "short-term context understanding network") and a network that understands long-term context at the sentence unit (hereinafter also referred to as a "long-term context understanding network"), and inputting the obtained intermediate features into a network that predicts a label (hereinafter also referred to as a "label prediction network") to estimate the label of the response scene.
[0004] In order to achieve high classification accuracy in a labeling model such as Non-Patent Document 1 for utterance sequence labeling, it is necessary to increase the number of trainable parameters for each of the short-term context understanding network and the long-term context understanding network. Inference using such a labeling model requires a rich computing environment. However, it is particularly difficult to prepare a rich computing environment, especially in a mobile environment or an environment where multiple inferences are executed simultaneously in parallel. Here, a knowledge distillation technique has been proposed in which a model with a small number of trainable parameters and lightweight (hereinafter also referred to as a "student model") is efficiently learned using the knowledge acquired by a model with a large number of trainable parameters and high classification accuracy (hereinafter also referred to as a "teacher model").
[0005] For example, according to Non-Patent Document 2, as schematically shown in FIG. 2, in order to train a student model, in addition to using a loss (hereinafter also referred to as a "hard target loss") for making the probability distribution output by the student model approach the probability distribution of the correct label, a loss (hereinafter also referred to as a "soft target loss") for making the probability distribution output by the student model approach the probability distribution output by the teacher model is used. As a result, the student model can be trained to imitate the teacher model, and knowledge distillation for distilling the knowledge possessed by the teacher model into the student model can be realized.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
[0007] However, the method of Non-Patent Document 2 applies the knowledge distillation technique to a simple classification problem, and a configuration in which the knowledge distillation technique is applied to a complex classification problem considering a complex context such as Non-Patent Document 1 has not been considered.
[0008] An object of the present invention is to provide a learning device, an estimation method, a learning method, and a program that apply the knowledge distillation technique to a complex classification problem considering a complex context. [Means for Solving the Problems]
[0009] To solve the above problems, according to one aspect of the present invention, a learning device uses a teacher model, which is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more, to estimate labels for texts included in a learning dataset. The learning device includes a teacher model label estimation unit, a student model label estimation unit, a hard target loss evaluation unit, a soft target loss evaluation unit, and a parameter update unit. The student model label estimation unit uses a student model, which is a model hierarchically including the same Q functions as the teacher model, to estimate labels for texts included in the learning dataset. The hard target loss evaluation unit obtains a hard target loss using the correct labels for texts included in the learning dataset and the estimation results of the student model label estimation unit. The soft target loss evaluation unit obtains a soft target loss using the estimation results of the teacher model label estimation unit and the estimation results of the student model label estimation unit. The parameter update unit updates the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss.
Advantages of the Invention
[0010] According to the present invention, knowledge distillation can be realized for complex classification problems considering complex contexts, and an effect is achieved in that a student model with high classification accuracy can be learned.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described. In the drawings used in the following description, components having the same function and steps performing the same process are denoted by the same reference numerals, and redundant description is omitted. In the following description, symbols "~", "]]" etc. used in the text should be described directly above the immediately following character, but due to text notation limitations, they are described immediately before the character. In the formula, these symbols are described in their original positions. Also, the processing performed for each element unit of a vector or matrix is applied to all elements of the vector or matrix unless otherwise specified. - 」 etc. should be described directly above the immediately following character, but due to text notation limitations, they are described immediately before the character. In the formula, these symbols are described in their original positions. Also, the processing performed for each element unit of a vector or matrix is applied to all elements of the vector or matrix unless otherwise specified.
[0013] <Highlights of the First Embodiment> The highlight of this embodiment is the application of the knowledge distillation technique to the utterance sequence labeling problem. Conventionally, many studies have been conducted on the knowledge distillation technique for the purpose of model compression of machine translation models and BERT (Bidirectional Encoder Representations from Transformers) models, but this embodiment is the first to apply it to the problem of utterance sequence labeling. In this embodiment, the teacher model has a configuration that performs multi-stage processing, and the student model also performs knowledge distillation while maintaining the multi-stage processing configuration. In this embodiment, by implementing model compression through knowledge distillation, it is possible to achieve high classification accuracy labeling even in situations where it is particularly difficult to prepare a rich computing environment.
[0014] <First Embodiment> Hereinafter, taking knowledge distillation for model lightweighting in a neural network for utterance sequence labeling that takes an utterance text sequence in a contact center as input and outputs a label corresponding to a response scene, such as Non-Patent Document 1, as an example, it will be described. However, this embodiment is not limited to the utterance text sequence in the contact center or the utterance sequence labeling of the response scene. That is, it can be applied to any sequence labeling problem that requires consideration of context. It can be applied to a problem of assigning a label for each sentence or a specific unit when a text sequence is given. For example, it can be applied to a neural network as follows.
[0015] · The input layer is configured to receive text (or something having equivalent information such as its vector representation).
[0016] · The output layer corresponds to the estimated result of the label.
[0017] · The intermediate layer performs multi-stage processing or uses something that can handle context, such as a Transformer encoder (see Reference 1).
[0018] (Reference 1) Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, "Attention is All you need", 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017 Furthermore, this embodiment is not limited to the situation where the number of parameters that can be learned by the student model is smaller than that of the teacher model, and it may also be the case where the number of parameters that can be learned by the student model is larger than or equal to that of the teacher model. Note that, in the situation where the number of parameters is larger than or equal to (the sizes are equivalent), and the problem of the prior art that "a rich computing environment is required" can be solved, configurations such as "when there are multiple teacher models and one student model learns from them" are assumed.
[0019] <Estimation system> FIG. 3 is a diagram showing a configuration example of the estimation system according to the first embodiment.
[0020] The estimation system includes a learning device 100 and an estimation device 200.
[0021] The learning device 100 takes as inputs a learning dataset D = (X, - P) and a teacher model TM, and by means of the knowledge distillation technique, the student model SM learns to imitate the teacher model TM, and outputs the learned student model SM. The learning dataset D = (X, - P) is composed of a series of utterance texts X n =(x n,1 , x n,2 , …, x n,T_n ) and a series n of correct labels n,t p - corresponding to each utterance text x n,t of the series of utterance texts X - P n =( - p n,1 , - p n,2 , …, - p n,T_n ). It is a dataset constructed by collecting a large amount (N conversations' worth) of such pairs as the conversation data for one conversation, and X = (X1, X2, …, X N ), - P = ( - P1, - P2, …, - P N) where n is the index of the call data, and n = 1, 2, …, N. Also, the speech text x n,t means the t-th speech data included in the call data n, and the subscript A_B means A B means, and T n is the number of speech texts included in the call data n, and t = 1, 2, …, T n is.
[0022] The estimation device 200 receives a pre-trained student model SM, takes as input call data X test including one or more text sequences to be estimated, estimates the corresponding label sequence, and outputs the estimated label sequence P test is.
[0023] The learning device 100 and the estimation device 200 are special devices configured by loading a special program into a known or dedicated computer having, for example, a central processing unit (CPU: Central Processing Unit), a main memory device (RAM: Random Access Memory), etc. The learning device 100 and the estimation device 200 execute each process under the control of, for example, a central processing unit. The data input to the learning device 100 and the estimation device 200 and the data obtained by each process are stored, for example, in a main memory device, and the data stored in the main memory device is read out to the central processing unit as needed and used for other processes. At least a part of each processing unit of the learning device 100 and the estimation device 200 may be configured by hardware such as an integrated circuit. Each storage unit included in the learning device 100 and the estimation device 200 can be configured by, for example, a main memory device such as a RAM (Random Access Memory), or middleware such as a relational database or a key-value store. However, each storage unit does not necessarily have to be provided inside the learning device 100 and the estimation device 200, and may be configured by an auxiliary storage device configured by a semiconductor memory element such as a hard disk, an optical disk, or a flash memory (Flash Memory), and may be provided outside the learning device 100 and the estimation device 200.
[0024] First, the learning device 100 will be described.
[0025] <Overview of the processing of the learning device 100> In the learning process, in order to efficiently learn a student model with few learnable parameters and light weight, the knowledge acquired by a teacher model with many learnable parameters and high classification accuracy is used.
[0026] Here, the number of learnable parameters is defined by, for example, the number of layers, the intermediate output dimension, etc. when the short-term context understanding network or the long-term context understanding network is configured by an LSTM (long short-term memory) or a fully connected neural network.
[0027] Also, the number of learnable parameters is defined by, for example, the number of blocks, the intermediate output dimension in the fully connected neural network of each block, the number of heads of the multi-head attention, and the output dimension, etc. when the short-term context understanding network or the long-term context understanding network is configured by a Transformer encoder block.
[0028] Furthermore, the number of learnable parameters is defined by, for example, the number of layers, the intermediate output dimension, etc. when the label prediction network is configured by a fully connected neural network.
[0029] In short, the teacher model and the student model share the "hierarchical structure" by the short-term context understanding network, long-term context understanding network, and label prediction network as shown in FIG. 1, but it is assumed that the model sizes are different because the number of parameters of each network is different. The "hierarchical structure" mentioned here does not mean a simple neural network, but a structure including a plurality of functions for performing processing in a predetermined unit. The function for performing processing in a predetermined unit has an intention in the unit of processing. For example, the long-term context understanding network performs document unit processing with the intention of understanding the long-term context between sentences, and the short-term context understanding network performs sentence unit processing with the intention of understanding the short-term context within a sentence.
[0030] In the learning process, the student model is learned using the loss defined by the combination of the two losses schematically shown in FIG. 4. Here, the teacher model is not the learning target, and the parameters are fixed.
[0031] The hard target loss is a loss for making the probability distribution output by the student model approach the probability distribution of the correct label.
[0032] The soft target loss is a loss for making the probability distribution output by the student model approach the probability distribution output by the teacher model.
[0033] In the learning process, the learning dataset may be used to perform learning by the error backpropagation method or the like so as to optimize a loss function in which the hard target loss and the soft target loss are linearly combined at a certain ratio, for example.
[0034] Next, a configuration example of a learning device for implementing the above processing will be described.
[0035] <Learning Device 100> FIG. 5 shows a functional block diagram of the learning device 100 according to the first embodiment, and FIG. 6 shows its processing flow. FIG. 7 is a diagram for explaining the processing outline of the learning device 100.
[0036] The learning device 100 includes a teacher model label estimator 110, a student model label estimator 120, a hard target loss evaluator 130, a soft target loss evaluator 140, and a parameter updater 150.
[0037] <Teacher model label estimator 110> The teacher model label estimator 110 receives a teacher model TM in advance. The teacher model TM is a hierarchical model by a neural network and includes a short-term context understanding network, a long-term context understanding network, and a label prediction network in the present embodiment.
[0038] The teacher model label estimator 110 uses the teacher model TM, which is a hierarchical model, to process the N utterance text sequences X n (n = 1, 2, …, N) included in the N conversations of the learning dataset D, receives the utterance text sequences X n and estimates the labels for the utterance texts x n,t included in X (S110), and outputs the probability distribution ~z n,t (n = 1, 2, …, N, t = 1, 2, …, T n where T n is the number of utterance texts included in the utterance text sequence X n as the estimation result. For example, the processing is performed as follows.
[0039] The teacher model label estimator 110 uses the short-term context understanding network to obtain the intermediate feature ~s n for the utterance text x n,t included in the utterance text sequence X n,t (~L) (S110A). Here, ~L indicates the number of layers of the short-term context understanding network of the teacher model. The short-term context understanding network is a neural network that understands the short-term context at the word level and captures what kind of utterance was made within the sentence. The intermediate feature ~s n,t (~L) includes features for understanding the short-term context at the word level.
[0040] Next, the teacher model label estimation unit 110 uses the long-term context understanding network to obtain the intermediate feature quantity ~s n,t (~L) for the intermediate feature quantity ~u n,t (~M) (S110B). However, ~M indicates the number of layers of the long-term context understanding network of the teacher model. Note that the long-term context understanding network is a neural network that understands the long-term context at the sentence level and follows the flow of the topic by capturing the time series of the utterances. The intermediate feature quantity ~u n,t (~M) contains features for understanding the long-term context at the sentence level.
[0041] Furthermore, the teacher model label estimation unit 110 predicts the label for the intermediate feature quantity ~u n,t (~M) using the label prediction network (S110C) and outputs the probability distribution ~z n,t of the prediction. The label prediction network is a neural network that predicts labels. In this embodiment, the output layer of the label prediction network includes a temperature-softmax function, and the teacher model label estimation unit 110 outputs the probability distribution ~z n,t which is the output of the temperature-softmax function. Note that ~v in FIG. 7 t is the output of the fully connected layer one before the output layer of the label prediction network.
[0042] <Student model label estimation unit 120> The student model label estimation unit 120 initializes the student model SM in advance. As an initialization method of the neural network, existing techniques can be used. The student model SM is a hierarchical model by a neural network, similar to the teacher model TM, and includes a short-term context understanding network, a long-term context understanding network, and a label prediction network in this embodiment.
[0043] The student model label estimation unit 120 uses the student model SM, which is a hierarchical model, to process the N utterance text series X nReceives (n = 1, 2, …, N) and the utterance text series X n estimates the label for the utterance text x n,t contained in (S120), and the probability distribution p n,t , z n,t (n = 1, 2, …, N, t n = 1, 2, …, T n ) is output. For example, the processing is performed as follows.
[0044] The student model label estimation unit 120 uses the short-term context understanding network to obtain the intermediate feature quantity s n for the utterance text x n,t contained in the utterance text series X (S120A). However, L indicates the number of layers of the short-term context understanding network of the student model. For example, let L ≦ ~L. n,t (L) Next, the student model label estimation unit 120 uses the long-term context understanding network to obtain the intermediate feature quantity u
[0045] for the intermediate feature quantity s (S120B). However, M indicates the number of layers of the long-term context understanding network of the student model. For example, let M ≦ ~M. n,t (L) Furthermore, the student model label estimation unit 120 uses the label prediction network to predict the label for the intermediate feature quantity u n,t (M) (S120C), and outputs the probability distribution p
[0046] of the prediction, p n,t (M) , z n,t , z n,t is output. The output layer of the label prediction network of the student model label estimation unit 120 includes a softmax function and a temperature-softmax function, and the student model label estimation unit 120 outputs the probability distribution p n,t which is the output of the softmax function and the probability distribution z n,t which is the output of the temperature-softmax function. Note that v t in FIG. 7 is the output of the fully connected layer one layer before the output layer of the label prediction network.
[0047] <Hard target loss evaluation unit 130> The hard target loss evaluation unit 130 receives the series of correct labels - p n,t and the probability distribution of the prediction by the student model - P n =( - p n,1 , - p n,2 ,…, - p n,T_n ) and - p n,1 ,p n,2 ,…,p n,T_n (n = 1, 2, …, N), obtains the hard target loss L HT and outputs it. The distance between the probability distribution obtained from the correct label and the probability distribution of the prediction may be evaluated using any loss function such as the cross-entropy loss. For example, the hard target loss L HT is obtained by the following formula.
Equation
[0048] <Soft target loss evaluation unit 140> The soft target loss evaluation unit 140 receives the probability distributions ~z n,1 , ~z n,2 , …, ~z n,T_n predicted by the teacher model and the probability distributions z n,1 , z n,2 , …, z n,T_n (n = 1, 2, …, N), obtains the soft target loss L ST , and outputs it. The distance between the two probability distributions may be evaluated using any loss function such as cross-entropy loss or mean squared error. For example, the soft target loss L ST is obtained by the following equation. [Equation] Note that τ is a parameter of the temperature-softmax function.
[0049] [Parameter update unit 150] The parameter update unit 150 receives the hard target loss L HT and the soft target loss L ST , and updates the parameters of the student model so as to optimize the loss L obtained from the hard target loss L HT and the soft target loss L ST (S150). For example, the hard target loss L HT and the soft target loss L ST are combined linearly at a certain ratio to obtain a loss function L.
[0050] L = L HT + λL ST However, λ is a parameter indicating the combination ratio of the hard target loss and the soft target loss. The parameter update unit 150 updates the parameters of the student model so as to optimize the loss function L. For example, the learning device 100 may be trained using the learning dataset D by the error backpropagation method or the like. The ratio λ may be defined in advance in the learning schedule and changed according to the number of learning steps based on it. For example, at the beginning of learning, the soft target loss LST Learn only using this, and gradually learn to give the hard target loss L HT It may be learned to give.
[0051] The parameter update unit 150 outputs the updated parameters to the student model label estimator 120 until a predetermined condition is satisfied, and repeats S120, S130, S140, and S150 (NO in S150-2). The predetermined condition is, for example, that the number of repetitions exceeds a predetermined number, or the difference between the parameters before and after the update is equal to or less than a predetermined threshold value. In short, it is a condition for determining whether or not the parameter update has converged.
[0052] Next, the estimation device 200 will be described.
[0053] <Estimation device 200> FIG. 8 shows a functional block diagram of the estimation device 200 according to the first embodiment, and FIG. 9 shows its processing flow.
[0054] The estimation device 200 includes an estimation unit 210.
[0055] The estimation unit 210 receives a pre-trained student model SM.
[0056] The estimation unit 210 uses the call data X test including one or more text sequences to be estimated as input, and uses the student model SM to estimate the labels corresponding to the respective utterance texts of the call data X test in order (S210), and outputs the estimated label sequence P test to output.
[0057] <Effect> With the above configuration, knowledge distillation is realized for a complex classification problem considering a complex context, and an effect that a student model with high classification accuracy can be learned is obtained.
[0058] <Modification example> In this embodiment, the spoken text is the processing target, but it is not necessarily limited to the text based on speech. For example, it is applicable to a text sequence including exchanges using texts without speech used in chats, emails, various SNSs, etc.
[0059] <Highlights of the Second Embodiment> The highlights of this embodiment are the following two points.
[0060] 1. The long-term context understanding network of the student model learns to imitate the long-term context understanding network of the teacher model.
[0061] 2. The short-term context understanding network of the student model learns to imitate the short-term context understanding network of the teacher model.
[0062] According to the above 1., the intermediate features of the long-term context output by the long-term context understanding network of the student model are learned to approach the intermediate features of the long-term context output by the long-term context understanding network of the teacher model. As a result, since the long-term context understanding network of the student model can learn to imitate the long-term context understanding network of the teacher model, the student model can imitate the teacher model more precisely than when only imitating the probability distribution of each label, leading to an improvement in the classification accuracy of the student model.
[0063] According to the above 2., the intermediate features of the short-term context output by the short-term context understanding network of the student model are learned to approach the intermediate features of the short-term context output by the short-term context understanding network of the teacher model. As a result, since the short-term context understanding network of the student model can learn to imitate the short-term context understanding network of the teacher model, the robustness of the short-term context understanding network with respect to the content of the spoken text is improved, leading to an improvement in the classification accuracy of the student model.
[0064] In the first embodiment, by introducing the knowledge distillation technique realized by Non-Patent Document 2 into the labeling model for the utterance sequence labeling problem shown in Non-Patent Document 1, a lightweight student model can be efficiently learned using the knowledge acquired by the teacher model.
[0065] However, the method of Non-Patent Document 2 is applied to a simple classification problem, and there may be cases where the knowledge of intermediate features cannot be efficiently distilled.
[0066] Therefore, in this embodiment, for the utterance sequence labeling problem considering complex contexts, a student model with high classification accuracy is learned by efficiently distilling the knowledge of intermediate features from the teacher model.
[0067] In this embodiment, in order to introduce knowledge distillation such as that of Non-Patent Document 2 into the labeling network for utterance sequence labeling considering complex contexts, the intermediate features of the long-term context and the intermediate features of the short-term context output by the student model are learned to mimic those of the teacher model, thereby efficiently distilling the knowledge of intermediate features from the teacher model.
[0068] <Second Embodiment> The description will focus on the parts different from the first embodiment.
[0069] <Estimation System> FIG. 3 is a diagram showing a configuration example of an estimation system according to the second embodiment.
[0070] The estimation system includes a learning device 300 and an estimation device 200.
[0071] The second embodiment is different from the first embodiment in the content of the learning process.
[0072] <Outline of Processing of Learning Device 300> In the learning process of the second embodiment, the student model is learned using a loss defined by the combination of the four losses schematically shown in FIG. 10. The hard target loss and the soft target loss are common to the first embodiment.
[0073] The long-term context loss is a loss function for learning so that the intermediate features of the long-term context output by the long-term context understanding network of the student model mimic the intermediate features of the long-term context output by the long-term context understanding network of the teacher model. When the output dimensionality of the long-term context understanding network is different between the student model and the teacher model, as schematically shown in FIG. 11, for example, a fully connected layer for aligning the dimensionality is branched and provided to the output of the long-term context understanding network of the teacher model, and the student model and the fully connected layer for aligning the dimensionality may be learned so that the features output by the fully connected layer for aligning the dimensionality are close to the intermediate features of the long-term context output by the long-term context understanding network of the student model.
[0074] The short-term context loss is a loss function for learning so that the intermediate features of the short-term context output by the short-term context understanding network of the student model mimic the intermediate features of the short-term context output by the short-term context understanding network of the teacher model. When the output dimensionality of the short-term context understanding network is different between the student model and the teacher model, as schematically shown in FIG. 11, for example, a fully connected layer for aligning the dimensionality is branched and provided to the output of the short-term context understanding network of the teacher model, and the student model and the fully connected layer for aligning the dimensionality may be learned so that the features output by the fully connected layer for aligning the dimensionality are close to the intermediate features of the short-term context output by the short-term context understanding network of the student model.
[0075] In the learning process of the second embodiment, the learning dataset may be used to perform learning by the error backpropagation method or the like so as to optimize a loss function in which the hard target loss, the soft target loss, the long-term context loss, and the short-term context loss are linearly combined at a certain ratio, for example.
[0076] Next, a configuration example of a learning device for performing the above processing will be described.
[0077] <Learning device 300> FIG. 5 is a functional block diagram of the learning device 300 according to the second embodiment, and FIG. 6 shows its processing flow. FIG. 7 is a diagram for explaining the processing outline of the learning device 300.
[0078] The learning device 300 includes a teacher model label estimation unit 110, a student model label estimation unit 120, a hard target loss evaluation unit 130, a soft target loss evaluation unit 140, a short-term context loss evaluation unit 360, a long-term context loss evaluation unit 370, and a parameter update unit 350.
[0079] <Short-term context loss evaluation unit 360> The short-term context loss evaluation unit 360 receives intermediate feature vectors ~s n,t (~L) , s n,t (L) (n = 1, 2,..., N, t = 1, 2,..., T n ), obtains the short-term context loss L UC (S360), and outputs it. The distance between the two intermediate features may be evaluated using any loss function such as the mean squared error. For example, the short-term context loss L UC is obtained by the following equation.
Equation
[0080] <Long-term context loss evaluation unit 370> The long-term context loss evaluation unit 370 receives intermediate feature vectors ~u n,t (~M) , u n,t (M) (n = 1, 2,..., N, t = 1, 2,..., T n ), and the long-term context loss L DCObtain (S370) and output. The distance between the two intermediate features may be evaluated using any loss function such as the mean squared error. For example, the long-term context loss L is obtained by the following formula. DC is obtained.
Number
[0081] <Parameter update unit 350> The parameter update unit 350 receives the hard target loss L HT and the soft target loss L ST and the short-term context loss L UC and the long-term context loss L DC and obtains a loss function L obtained by linearly combining the hard target loss L HT and the soft target loss L ST and the short-term context loss L UC and the long-term context loss L DC at a certain ratio.
[0082] L = L HT + λL ST + αL UC + βL DC However, λ, α, and β are parameters indicating the combination ratio of the hard target loss, the soft target loss, the short-term context loss, and the long-term context loss. The parameter update unit 350 updates the parameters of the student model so as to optimize the loss function L (S350). For example, the learning device 300 may perform learning by the error backpropagation method or the like using the learning dataset D. The ratios λ, α, and β may be defined in advance for the learning schedule and the learning may be performed while changing according to the number of learning steps based on this. For example, in the initial stage of learning, learning may be performed using only the short-term context loss, and gradually the long-term context loss, the soft target loss, and the hard target loss may be given in this order for learning.
[0083] As described above, when a fully connected layer for aligning the number of dimensions is provided as shown in FIG. 11, the parameters of the fully connected layer are also updated accordingly.
[0084] The parameter update unit 350 outputs the updated parameters to the student model label estimator 120 until a predetermined condition is satisfied, and repeats S120 - S140, S360, S370, and S350 (NO in S150-2).
[0085] <Effect> With such a configuration, in the utterance sequence labeling problem considering complex contexts, the student model can more precisely imitate a teacher model with a high ability to capture the characteristics of short-term context and long-term context. As a result, since it becomes possible to efficiently distill the knowledge acquired by the teacher model by the student model, the labeling accuracy of the student model can be improved. The "characteristics of short-term context" are expressed by one vector for each sentence, and by learning so that the information becomes close between the teacher model and the student model, the method of expressing the characteristics of the sentence can be imitated as it is. Also, the "characteristics of long-term context" are expressed by one vector for the flow of the topic, and by learning so that the information becomes close between the teacher model and the student model, the method of expressing the flow of the topic can be imitated as it is.
[0086] <Verification experiment results> A verification experiment was conducted on the task of utterance sequence labeling that takes the utterance text sequence in the contact center as input and outputs labels corresponding to the response scenarios. Using pseudo-response data in a Japanese contact center, the number of data in the learning dataset was 327 calls, and the number of data in the test dataset was 37 calls. The classification targets were five labels: opening, matter grasping, identity confirmation, response, and closing. The number of parameters of the teacher model was 13.11M, and the number of parameters of the student model was 3.65M. The classification accuracy of each response scenario was compared between a baseline where the student model was simply learned from scratch and a learning method using the learning methods according to the first embodiment and the second embodiment.
[0087] In addition, in the second embodiment, the case where neither the short-term context loss nor the long-term context loss is used was also compared. For the evaluation, the accuracy rate based on perfect match was used. The results of the verification experiment are shown in FIG. 12. From FIG. 12, it can be seen that by applying the second embodiment in particular, even a lightweight student model can obtain a classification accuracy close to that of the teacher model.
[0088] <Modification Example> When it is hierarchical like the long-term context understanding network and the short-term context understanding network as in this embodiment, the intermediate features of each layer may be compared. In this embodiment, there are two layers, the long-term context understanding network and the short-term context understanding network, but this embodiment can also be applied to three or more layers. When the networks constituting the teacher model and the student model are divided into Q hierarchical (block) for each function, the output of any layer constituting the hierarchical (block) is compared with Q or more, Q or more losses are calculated, the Q or more losses are combined to obtain the final loss, and the parameters of the student model may be updated so as to optimize the final loss. When the number of losses to be calculated exceeds Q, the comparison target for the Q + 1-th and subsequent ones may be from any block. In other words, when the number of losses to be calculated exceeds Q, two or more comparison targets (intermediate features) may be extracted from one block, and a total of Q + 1 or more comparison targets may be extracted from Q blocks.
[0089] Also, as described in the verification experiment, it may be configured to have only one of the short-term context loss evaluation unit 360 and the long-term context loss evaluation unit 370. When divided into Q hierarchical (block) for each task, the output of any layer constituting the hierarchical (block) is compared with one or more, one or more losses are calculated, the one or more losses are combined to obtain the final loss, and the parameters of the student model may be updated so as to optimize the final loss.
[0090] <Other Modification Examples> The present invention is not limited to the above-described embodiments and modifications. For example, the various processes described above may be executed not only in time series according to the description, but also in parallel or individually according to the processing capabilities of the device executing the processes or as required. In addition, appropriate changes can be made without departing from the spirit of the present invention.
[0091] <Program and Recording Medium> The various processes described above can be implemented by causing the storage unit 2020 of the computer shown in FIG. 13 to read a program for executing each step of the above method and causing it to operate on the control unit 2010, the input unit 2030, the output unit 2040, etc.
[0092] The program describing this processing content can be recorded on a computer-readable recording medium. As the computer-readable recording medium, for example, any of a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, etc. may be used.
[0093] In addition, the distribution of this program can be carried out, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or a CD-ROM on which the program is recorded. Further, the program may be stored in the storage device of a server computer and transferred from the server computer to other computers via a network to distribute the program.
[0094] A computer that executes such a program first stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer in its own storage device. Then, when executing the process, the computer reads the program stored in its own recording medium and executes the process according to the read program. As another execution form of this program, the computer may directly read the program from the portable recording medium and execute the process according to the program. Further, each time a program is transferred from the server computer to this computer, the computer may sequentially execute the process according to the received program. Also, the above-described process may be executed in a configuration of a so-called ASP (Application Service Provider) type service that does not transfer the program from the server computer to this computer but realizes the processing function only by the execution instruction and result acquisition. Note that the program in this embodiment includes information used for processing by an electronic computer and similar to the program (data having a property of defining the processing of the computer but not being a direct instruction to the computer).
[0095] In this embodiment, the apparatus is configured by causing a computer to execute a predetermined program, but at least a part of these processing contents may be realized hardware-wise.
[0096] <Modification Example> In the above embodiment, the program executed by the CPU after reading software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacture, such as a GPU (Graphics Processing Unit) and an FPGA (Field-Programmable Gate Array), and a dedicated electric circuit which is a processor having a circuit configuration designed specifically for executing specific processing, such as an ASIC (Application Specific Integrated Circuit). Also, the program may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a plurality of FPGAs, a combination of a CPU and an FPGA, etc.). Further, the hardware structure of these various processors is, more specifically, an electric circuit combining circuit elements such as semiconductor elements.
[0097] Also, in the above embodiment, it has been described that the program is stored (installed) in the storage in advance, but it is not limited to this. The program may be provided in a form stored in a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), and a USB (Universal Serial Bus) memory. Also, the program may be in a form downloaded from an external device via a network.
[0098] Regarding the above embodiments, the following additional remarks are further disclosed.
[0099] (Additional Clause 1) A memory, At least one processor connected to the memory, comprising, wherein the processor Using a teacher model, which is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more, execute teacher model label estimation processing for estimating a label for text included in a learning dataset. Using a student model, which is a model hierarchically including the same Q functions as the teacher model, execute student model label estimation processing for estimating a label for text included in a learning dataset. Using the correct label for the text included in the learning dataset and the estimation result of the student model label estimation processing, obtain a hard target loss. Using the estimation result of the teacher model label estimation processing and the estimation result of the student model label estimation processing, obtain a soft target loss. Update the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss. Learning device.
[0100] (Appended item 2) A non-transitory storage medium storing a program executable by a computer to execute learning processing, The learning processing includes: Using a teacher model, which is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more, execute teacher model label estimation processing for estimating a label for text included in a learning dataset. Using a student model, which is a model hierarchically including the same Q functions as the teacher model, execute student model label estimation processing for estimating a label for text included in a learning dataset. Using the correct label for the text included in the learning dataset and the estimation result of the student model label estimation processing, obtain a hard target loss. Using the estimation result of the teacher model label estimation processing and the estimation result of the student model label estimation processing, obtain a soft target loss. Update the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss. Non-transitory storage medium.
Claims
A teacher model label estimation unit that estimates a label for text included in a learning dataset using a teacher model that is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more; A student model label estimation unit that estimates a label for text included in a learning dataset using a student model that is a model hierarchically including the same Q functions as the teacher model; A hard target loss evaluation unit that obtains a hard target loss using the correct label for the text included in the learning dataset and the estimation result of the student model label estimation unit; A soft target loss evaluation unit that obtains a soft target loss using the estimation result of the teacher model label estimation unit and the estimation result of the student model label estimation unit; A parameter update unit that updates the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss; Including a q-th loss evaluation unit that obtains a q-th loss from a q-th intermediate feature amount obtained from the q-th layer of the teacher model and a q-th intermediate feature amount obtained from the q-th layer of the student model, where q is any integer from 1 to Q; The parameter update unit updates the parameters of the student model so as to optimize the loss obtained from the hard target loss, the soft target loss, and the q-th loss. A learning device. A teacher model label estimation unit that estimates a label for text included in a learning dataset using a teacher model that is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more; A student model label estimation unit that estimates a label for text included in a learning dataset using a student model that is a model hierarchically including the same Q functions as the teacher model; A hard target loss evaluation unit that obtains a hard target loss using the correct label for the text included in the learning dataset and the estimation result of the student model label estimation unit; A soft target loss evaluation unit that obtains a soft target loss using the estimation result of the teacher model label estimation unit and the estimation result of the student model label estimation unit; Including a parameter update unit that updates the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss. The teacher model and the student model include a short-term context understanding network as the first layer, a long-term context understanding network as the second layer, and a label prediction network as the third layer. A short-term context loss evaluation unit that obtains a short-term context loss from a first intermediate feature amount obtained from the short-term context understanding network of the teacher model and a second intermediate feature amount obtained from the short-term context understanding network of the student model. It includes a long-term context loss evaluation unit that obtains a long-term context loss from a third intermediate feature amount obtained from the long-term context understanding network of the teacher model and a fourth intermediate feature amount obtained from the long-term context understanding network of the student model. The parameter update unit updates the parameters of the student model so as to optimize the loss obtained from the hard target loss, the soft target loss, the short-term context loss, and the long-term context loss. Learning device.
3. An estimation device that uses a student model learned by the learning device according to any one of Claims 1 to 2, including an estimation unit that estimates a label corresponding to a text to be estimated using the learned student model. Estimation device.
4. A learning method using a learning device, wherein the learning device uses a teacher model that is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more, to estimate a label for a text included in a learning dataset. A teacher model label estimation step; a student model label estimation step in which the learning device uses a student model that is a model hierarchically including the same Q functions as the teacher model to estimate a label for a text included in a learning dataset; a hard target loss evaluation step in which the learning device obtains a hard target loss using the correct label for the text included in the learning dataset and the estimation result of the student model label estimation step; a soft target loss evaluation step in which the learning device obtains a soft target loss using the estimation result of the teacher model label estimation step and the estimation result of the student model label estimation step; a parameter update step in which the learning device updates the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss. The learning device includes a q-th loss evaluation step of obtaining a q-th loss from a q-th intermediate feature amount obtained from the q-th layer of the teacher model and a q-th intermediate feature amount obtained from the q-th layer of the student model, where q is any integer from 1 to Q, The parameter update step updates the parameters of the student model so as to optimize the loss obtained from the hard target loss, the soft target loss, and the q-th loss. Learning method. **Claim 5** A learning method using a learning device, The learning device uses a teacher model, which is a model hierarchically including Q functions that perform processing in a predetermined unit, where Q is any integer of 2 or more, to estimate a label for text included in a learning dataset in a teacher model label estimation step. The learning device uses a student model, which is a model hierarchically including the same Q functions as the teacher model, to estimate a label for text included in a learning dataset in a student model label estimation step. The learning device includes a hard target loss evaluation step of obtaining a hard target loss by using a correct label for text included in the learning dataset and an estimation result of the student model label estimation step. The learning device includes a soft target loss evaluation step of obtaining a soft target loss by using an estimation result of the teacher model label estimation step and an estimation result of the student model label estimation step. The learning device includes a parameter update step of updating the parameters of the student model so as to optimize the loss obtained from the hard target loss and the soft target loss. The teacher model and the student model each include a short-term context understanding network as the first layer, a long-term context understanding network as the second layer, and a label prediction network as the third layer. The learning device includes a short-term context loss evaluation step of obtaining a short-term context loss from a first intermediate feature amount obtained from the short-term context understanding network of the teacher model and a second intermediate feature amount obtained from the short-term context understanding network of the student model. The learning device includes a long-term context loss evaluation step of obtaining a long-term context loss from a third intermediate feature amount obtained from the long-term context understanding network of the teacher model and a fourth intermediate feature amount obtained from the long-term context understanding network of the student model. The parameter update step updates the parameters of the student model so as to optimize the loss obtained from the hard target loss, the soft target loss, the short-term context loss, and the long-term context loss. Learning method. **Claim 6** A program for causing a computer to function as the learning device according to any one of Claims 1 to 2. **Claim 7** A program for causing a computer to function as the estimation device according to Claim 3.
Citation Information
Patent Citations
Entity name recognition model training method and device, equipment and storage medium
CN112613312A
Model learning device, method therefor, and program
WO2018051841A1