Natural language processing device and program
The natural language processing system addresses the inefficiency of training models across multiple domains by using an influence analysis unit to identify domain-sensitive components, allowing for targeted learning and maintaining performance across different domains.
Patent Information
- Application Number
- JP2021131559
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-08-12
AI Technical Summary
Existing natural language processing technologies using machine learning techniques face inefficiencies in training models for multiple domains, as they require different processing models and domain-specific training data, leading to potential performance degradation when domains differ.
A natural language processing system that includes an influence analysis unit to analyze the impact of domain differences on specific components of word-only embedded expressions, allowing for targeted learning and re-learning across domains by identifying components susceptible to domain influences.
This approach enables efficient re-learning across different domains, preventing performance degradation and improving learning efficiency by focusing on domain-specific components.
Smart Images

Figure 0007682053000014 
Figure 0007682053000015 
Figure 0007682053000016
Abstract
Description
[Technical field]
[0001] The present invention relates to a natural language processing device and a program. [Background technology]
[0002] Techniques for processing natural language using machine learning mechanisms such as neural networks are being researched and used. Natural language processing here refers to tasks such as translating sentences written in natural language, answering questions written in natural language, inputting sentences written in natural language and outputting classification labels (classification information) for the sentences, etc.
[0003] Non-Patent Document 1 describes BERT, a model that can be used to perform various analysis tasks of natural language. BERT stands for "Bidirectional Encoder Representations from Transformers." [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv:1810.04805 [cs.CL], 2018. Summary of the Invention [Problem to be solved by the invention]
[0005] In natural language processing technology using machine learning techniques, in order to obtain high-quality learning results, it is necessary to perform learning appropriate for each domain. Specifically, it is required to use different processing models for each domain and to train these models using learning data suitable for the domain. Here, a domain is an area to which texts with similar properties belong. Specifically, a domain may be, for example, the field of content expressed by text. For example, a set of texts in the journalism field and a set of texts in the entertainment field have different properties. These sets of texts in the journalism field and the entertainment field belong to the journalism domain and the entertainment domain, respectively.
[0006] Accommodating different domains is called domain adaptation or cross-domain adaptation.
[0007] It is possible to train the BERT model described in Non-Patent Document 1 using different training data for each domain. However, training all domains from scratch may be inefficient. This is because a trained model may have a part specific to a domain and a common part that is independent of the domain.
[0008] When training models for multiple domains, it is highly desirable to be able to efficiently obtain high-quality training results.
[0009] The present invention has been made based on the above-mentioned problem recognition, and aims to provide a natural language processing device and a program that can efficiently learn models for performing natural language processing for multiple domains. [Means for solving the problem]
[0010] [1] In order to solve the above problem, a natural language processing system according to one embodiment of the present invention includes an influence analysis unit that, for a machine-learning capable model that processes at least a word-single embedded expression corresponding to a natural language expression, analyzes whether or not the specific component is easily influenced by a difference between the first domain and the second domain based on the magnitude of the absolute value of the difference between the likelihood of a loss function calculated from a result of removing a specific component of a word-single embedded expression based on first training data belonging to a first domain and the likelihood of a loss function calculated from a result of removing the specific component of a word-single embedded expression based on second training data belonging to a second domain, and outputs information indicating whether or not the specific component is easily influenced by a difference between the first domain and the second domain.
[0011] [2] In another aspect of the present invention, in the natural language processing device described above, the specific component is information of each dimension of the word-only embedding representation represented as a vector.
[0012] [3] In addition, in one aspect of the present invention, in the above-mentioned natural language processing device, the influence analysis unit classifies each dimension of the vector of the word alone embedded expression into either a side that is easily affected by the difference between the first domain and the second domain or a side that is not easily affected by the difference between the first domain and the second domain based on the magnitude of the absolute value of the difference between the likelihood of a loss function calculated when information of a specific dimension of the word alone embedded expression based on the first training data is forcibly changed to a predetermined value and the likelihood of a loss function calculated when information of the specific dimension of the word alone embedded expression based on the second training data is forcibly changed to the predetermined value, and outputs information indicating which side each dimension belongs to as information indicating for each dimension whether it is a component that is easily affected by the difference between the first domain and the second domain.
[0013] [4] In one aspect of the present invention, in the above-mentioned natural language processing device, the natural language processing model further includes a learning unit that learns only certain components of the natural language processing model using the second learning data based on a natural language processing model trained with the first learning data and information indicating whether each specific component is susceptible to the influence of the difference between the first domain and the second domain output by the impact analysis unit.
[0014] [5] Furthermore, one aspect of the present invention is a program for causing a computer to function as a natural language processing device including an influence analysis unit that, for a machine-learning capable model that processes at least a word-single embedded expression corresponding to a natural language expression, analyzes whether or not a specific component is susceptible to the influence of a difference between a first domain and a second domain based on the magnitude of the absolute value of the difference between the likelihood of a loss function calculated from a result of removing a specific component of a word-single embedded expression based on first training data belonging to a first domain and the likelihood of a loss function calculated from a result of removing the specific component of a word-single embedded expression based on second training data belonging to a second domain, and outputs information indicating whether or not the specific component is susceptible to the influence of a difference between the first domain and the second domain. Effect of the Invention
[0015] According to the present invention, the influence analysis unit performs an analysis process based on the first training data and the second training data, and outputs information indicating whether a specific component of a word-only embedded expression corresponding to a natural language expression is a component that is easily affected by the difference between the first domain and the second domain. By outputting this information, for example, for a model trained with the first training data, it is possible to know which component of a word-only embedded expression corresponding to the second training data is effective to use for training.
[0016] That is, in technology for processing natural language expressions (sentences, etc.), it is possible to efficiently perform re-learning when the domain is different from the time of model learning. In other words, it becomes relatively easy to prevent performance degradation when the domain is different from the time of model learning. [Brief description of the drawings]
[0017] [Figure 1] 1 is a block diagram showing a schematic functional configuration of a learning control device (natural language processing device) according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a block diagram showing a schematic functional configuration of a model used for natural language processing in the embodiment. [Diagram 3] 2 is a schematic diagram showing the relationship between input symbols (words, etc.) for a model for natural language processing in the embodiment and various feature amounts (embeddings) associated with the input symbols. FIG. [Figure 4] FIG. 2 is a schematic diagram showing an example of 1-of-K coding used to represent words in the embodiment. [Diagram 5] 13 is a schematic diagram for explaining a dimension removal (component removal) operation used by the influence analysis unit according to the embodiment for analyzing the influence due to differences in domains. FIG. [Figure 6] 2 is a block diagram showing an example of an internal configuration of a learning control device (natural language processing device) according to the embodiment. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] Next, an embodiment of the present invention will be described with reference to the drawings. The learning control device 1 of this embodiment is a device that controls the learning of a model capable of machine learning for processing natural language. The learning control device 1 is also called a "learning support device" or a "natural language processing device". As described below, the learning control device 1 performs processing for discriminating between a part independent of a domain (a subspace of the algebraic space) and a part specific to a domain (similarly, a subspace) in the space (algebraic space) in which the model operates. In addition, as a result of the learning control device 1 making such a discrimination, when a model that has been learned in a certain domain is re-learned in another domain, it becomes possible to perform re-learning specialized in the part specific to the domain. In other words, the learning control device 1 makes the learning of the model more efficient.
[0019] FIG. 1 is a block diagram showing a schematic functional configuration of a learning control device according to this embodiment. As shown in the figure, the learning control device 1 includes a learning data storage unit 21, a learning data storage unit 22, an influence analysis unit 23, a learning unit 31, and a model 32. Each of these functional units can be realized, for example, by a computer and a program. Each functional unit also has a storage means as necessary. The storage means is, for example, a variable in a program or a memory allocated by the execution of a program. Also, non-volatile storage means such as a magnetic hard disk device or a solid state drive (SSD) may be used as necessary. Also, at least a part of the functions of each functional unit may be realized as a dedicated electronic circuit rather than a program. The functions of each unit are as follows.
[0020] The learning data storage unit 21 stores data for learning a model for natural language processing. Specifically, the learning data storage unit 21 stores learning data for a first domain (hereinafter, referred to as T 1 (called T 1 may be referred to as the first learning data.
[0021] The learning data storage unit 22 stores data for learning a model for natural language processing. Specifically, the learning data storage unit 21 stores learning data for a second domain different from the first domain (hereinafter, referred to as T 2 (called T 2 may be referred to as second learning data.
[0022] The influence analysis unit 23 performs an analysis process on a machine-learnable model that processes at least a word-only embedded expression corresponding to a natural language expression. Specifically, the influence analysis unit 23 calculates the likelihood of a loss function calculated from a result of removing a specific component of a word-only embedded expression based on the first learning data belonging to the first domain. The influence analysis unit 23 also calculates the likelihood of a loss function calculated from a result of removing the specific component of a word-only embedded expression based on the second learning data belonging to the second domain. The influence analysis unit 23 calculates the difference (absolute value) between these two likelihoods. The influence analysis unit 23 analyzes whether or not the specific component is a component that is easily influenced by the difference between the first domain and the second domain based on the magnitude of the calculated difference (absolute value). As a result of this analysis process, the influence analysis unit 23 outputs information (for convenience, referred to as dimension information. This "dimension" is one aspect of a component) indicating whether or not the specific component is a component that is easily influenced by the difference between the first domain and the second domain.
[0023] As shown in the figure, the influence analysis unit 23 performs the above analysis process by using model D i M 1 (where 1≦i≦d) i M 1 , and how it is used by impact analysis unit 23 will be described later.
[0024] The specific component may be information on each dimension of the word alone embedded expression represented as a vector. In this case, the influence analysis unit 23 may be as follows. That is, the influence analysis unit 23 calculates the likelihood of a loss function calculated when the information on a specific dimension of the word alone embedded expression based on the first learning data is forcibly changed to a predetermined value (for example, zero). Also, the influence analysis unit 23 calculates the likelihood of a loss function calculated when the information on the specific dimension of the word alone embedded expression based on the second learning data is forcibly changed to the predetermined value (a value similar to the above predetermined value). Based on the magnitude of the difference (absolute value of) between them, the influence analysis unit 23 classifies each dimension of the vector of the word alone embedded expression into either a side that is easily influenced by the difference between the first domain and the second domain or a side that is not easily influenced by the difference between the first domain and the second domain. The influence analysis unit 23 outputs information (dimension information) indicating which side each dimension belongs to as information indicating whether or not each dimension is a component that is easily influenced by the difference between the first domain and the second domain. The influence analysis unit 23 can pass the obtained dimensional information to the learning unit 31 .
[0025] The details of the processing by the impact analysis unit 23 will be explained further later.
[0026] The learning unit 31 learns the model 32 with the second learning data based on information indicating whether or not a specific component (specific dimension) is susceptible to the influence of the difference between the first domain and the second domain output by the influence analysis unit 23. That is, when re-learning the model 32 (natural language processing model) that has been trained in the first domain, the learning unit 31 controls so that learning is performed only on predetermined components.
[0027] The model 32 receives information on a word string (sentence) in a natural language and outputs a processing result of the word string. The model 32 is configured to be machine-learned using, for example, a neural network. That is, the model 32 has updatable internal parameters, and is configured to be able to adjust the values of the internal parameters by a machine learning technique using learning data. An example of the internal configuration of the model 32 will be described later with reference to FIG. 2.
[0028] The model 32 can be trained to perform various processes based on an input natural language sentence using the learning data. The model 32 can be trained to classify the input sentence and output a classification result (class, classification label), for example. In this case, the model 32 may be a model that outputs one classification for one input sentence. The model 32 may also be a model that outputs the same number of classifications as the number of words in an input sentence (word string). The model 32 may also be a model that outputs any other number of classifications for one input sentence. The model 32 can also be trained to translate an input source language sentence and output a sentence in another language (target language). The model 32 can also be trained to output an answer to a question expressed by the input sentence. The model 32 can also be trained to output a sentence equivalent to a summary sentence of the input sentence.
[0029] The model 32 may be a natural language processing model that has been trained in advance with the first training data. When the model 32 has been trained in advance with the first training data, the model 32 can train only a predetermined component among the components of the word-only embedding expression based on the dimension information that the influence analysis unit 23 passes to the training unit 31. The training unit 31 can control which dimension of the model 32 to train based on the dimension information.
[0030] 2 is a block diagram showing a schematic functional configuration of a model (such as the above model 32) used for natural language processing in this embodiment. As shown in the figure, the model 100 is configured to include a fully connected neural network layer 101, a word-only embedded representation 111, an embedded representation 112, a synthesis unit 113, an attention layer 114, and a softmax layer 115. The configuration of this model 100 is the same as that of BERT. A word string (sentence) and information associated therewith are input to this model 100. The model 100 outputs a processing result (for example, information on a classification result (class) corresponding to the input word string).
[0031] The fully connected neural network layer 101 is a layer that inputs (a sequence of) words in 1-of-K coding representation and outputs (a sequence of) word-only embedding representations (token embeddings) in multidimensional continuous values.
[0032] The word independent embedding representation 111 is an embedding representation (token embedding) for each individual word corresponding to the word string (sentence) that is the input to the model. The word independent embedding representation 111 is realized, for example, using a semiconductor memory device. Hereinafter, the word independent embedding representation 111 may be represented as x (bold).
[0033] Other embedded expressions 112 are embedded expressions corresponding to information representing word positions and information representing special symbols. A word position is the position of each word included in a word string. Word position information is, for example, automatically generated based on a word string. Special symbols are symbols representing the beginning of a sentence, a sentence boundary, etc. The beginning of a sentence is represented, for example, as [CLS]. A sentence boundary is represented, for example, as [SEP]. The embedded expressions 112 are realized, for example, by using a semiconductor memory device.
[0034] The synthesis unit 113 synthesizes the information of the word-only embedded expression 111 and the information of the embedded expression 112. How the synthesis unit 113 synthesizes the information will be described later with reference to another figure.
[0035] The attention layer 114 is a layer (neural network) that outputs a signal based on the output from the synthesis unit 113.
[0036] The softmax layer 115 outputs the result of natural language processing by applying a softmax function to the output from the attention layer 114. Note that an example of the output from the softmax layer 115 shown in Fig. 2 is a class or a class sequence. The output from the softmax layer 115 is not limited to class information, and may be other types of information that are the result of natural language analysis processing.
[0037] Learning of a model using training data is performed as follows. That is, the cross-entropy of the output from the model based on the input data to the model that the training data has, relative to the correct answer that the training data has, is used as the loss function. The loss function is expressed as L(M,T). Here, M is the model, and T is the training data. Learning of model M is a process of optimizing the parameters contained in model M so as to minimize the loss L(M,T). Note that backpropagation can be used as a parameter optimization technique in neural networks. Cross-entropy minimization is equivalent to negative log-likelihood maximization, and is expressed by the following formula (1).
[0038]
number
[0039] In formula (1), t is one piece of data (one sentence or one text) included in the set of learning data T. Also, p M (t) is the likelihood that model M outputs when the above data t is applied to model M.
[0040] The first domain (the training data is T 1 ) is trained as M 1 Then, L(M,T 1 ) that minimizes M 1is expressed by the following equation (2).
[0041]
number
[0042] In the above equation (2), argmin represents the model M obtained by optimizing the parameter values that the model M can take over the entire space of parameter values so as to minimize the argument value.
[0043] Similar to the first domain, the second domain (whose training data is T 2 ) is trained as M 2 Then, L(M,T 2 ) that minimizes M 2 is expressed by the following equation (3).
[0044]
number
[0045] In the BERT model used in this embodiment, in addition to the embedding representation of a single word (token embedding), an embedding representation representing a word position (position embedding) and a word segment embedding representation (segment embedding) are used.
[0046] FIG. 3 is a schematic diagram showing the relationship between input symbols (words, etc.) to the model 100 and various features (embeddings, embedded representations) associated with the input symbols. In the illustrated example, sentences (English sentences) input to the model 100 are "my dog is cute" and "he likes play ##ing". The input sentence is divided into words, and each word is converted into a 1-of-K coding representation. The 1-of-K representation is a representation in which the elements of the vocabulary V of the input word are either 0 or 1 in the |V| dimension. Note that |V| is the number of words in the vocabulary V (vocabulary size). In the 1-of-K representation, only the element of the dimension corresponding to the word is 1, and the elements of all other dimensions are 0.
[0047] The calculation method for token embedding (word-only embedding layer) is as follows. As described above, one word in the input sentence (input word sequence) is represented by 1-of-K coding. If the representation in this 1-of-K coding is Z, the output X to the word-only embedding layer is expressed by the following formula (4).
[0048]
number
[0049] In equation (4), f is a sigmoid function. A is a matrix and B is a vector. The sigmoid function f operates on each dimension. Since Z is a 1-of-K coding, it is a |V|-dimensional binary value. Of the |V| elements of vector Z, only one element (for one dimension) has a value of 1, and all other elements have values of 0. The number of dimensions of the word-only embedding representation (token embedding) is d. In other words, A above is a matrix with a size of d×|V|. B above is a d-dimensional vector. Each of the elements of matrix A and vector B can take continuous values. These elements are the parameters of the model.
[0050] A special symbol called [CLS] is placed at the beginning of an input sentence (string of words). [CLS] is a sentence class identifier. In addition, a special symbol called [SEP] is placed at the end and boundary of the sentence. [SEP] is an identifier inserted at the division point of the sentence. A unit separated by the special symbol [SEP] may be called a segment.
[0051] As shown in Fig. 3, each input symbol is associated with a token embedding, a segment embedding, and a position embedding. Each of the token embedding, segment embedding, and position embedding is a feature represented by a multidimensional vector.
[0052] In the example shown, the token embedding corresponding to the leading special symbol [CLS] is E [CLS] The token embedding for the next word "my" is E my The token embedding is as follows: dog , E is , ,E [SEP] The token embedding is a feature output by the fully connected neural network layer 101, which receives the vector of the 1-of-K coding representation as input. The token embedding is a multidimensional vector whose elements are continuous values.
[0053] A segment embedding is a feature corresponding to a segment. In this example, the segment embedding corresponding to each symbol included in the input symbol sequence [CLS], my, dog, is, cute, [SEP] (for convenience, called segment A) is E A The segment embedding corresponding to each symbol in the input symbol sequence he, likes, play, ##ing, [SEP] (called segment B for convenience) following segment A is expressed as E B It is expressed as follows.
[0054] Position embedding is a feature that depends on the position of a symbol in an input symbol string. In this example, the embedding is 0 , E 1 , E 2 , E 3 , E 4 , E 5 , E 6 , E 7 , E 8 , E 9 , E 10 The position embedding is associated with
[0055] In this embodiment, the input embedding corresponding to each input symbol is the sum of the above token embedding, segment embedding, and position embedding.
[0056] In this embodiment, the influence analysis unit 23 analyzes which dimensional information of the token embedding vector corresponding to the input word shown in FIG. 3 has a large influence when learning across different domains. Specifically, the influence analysis unit 23 forcibly changes the value of a specific dimension of the token embedding corresponding to each of the words my, dog, is, cute, he, likes, play, and ##ing, and analyzes the influence of that dimension on model learning. In other words, the influence analysis unit 23 analyzes which dimensional information affects learning across multiple domains by removing information of a specific dimension of the token embedding. Note that token embeddings corresponding to special symbols [CLS] and [SEP] are not subject to dimension removal.
[0057] Next, a detailed description will be given of the analysis process performed by the impact analysis unit 23. As described above, the impact analysis unit 23 analyzes the impact of information on each dimension when a model that has been trained using training data of a certain domain is trained using training data of another domain.
[0058] FIG. 4 is a schematic diagram showing an example of the aforementioned 1-of-K coding for each word. The illustrated example shows vectors by 1-of-K coding corresponding to each of the words my, dog, is, and cute. These vectors are |V|-dimensional vectors, and the element values are either 0 or 1. In the vector corresponding to the word "my", only the value of the dimension corresponding to the word "my" is "1", and all other dimensions are "0". Similarly, in the vectors corresponding to the words "dog", "is", and "cute", only the value of the dimension corresponding to the words "dog", "is", and "cute", respectively, is "1", and all other dimensions are "0".
[0059] In this embodiment, the embedding representation of a single word is subjected to a dimension removal operation, which will be described next.
[0060] FIG. 5 is a schematic diagram for explaining the operation of dimension removal used by the impact analysis unit 23 for analysis. In the figure, x (in bold) is a token embedding (output of the word-only embedding layer) corresponding to one input word. The token embedding is represented as a d-dimensional vector (d is a natural number. As an example, the value of d is 768). That is, x (in bold) is expressed as (x 1 ,x 2 ,···,x d ) is a vector. The impact analysis unit 23 calculates the vector D i This operation can be performed. i The operation D replaces the i-th dimension element of a vector with a specified value (for example, 0). That is, vector x (bold type) is expressed by the following formula (5). In this case, if operation D is applied to vector x (bold type), i D obtained by applying i x (in bold) is as shown in equation (6) below, where 1≦i≦d.
[0061]
number
[0062]
number
[0063] The impact analysis unit 23 performs the above operation D i The analysis is performed as follows. Here, two domains, the first domain and the second domain, are analyzed. The set of learning data for the first domain is T 1 Let the set of learning data for the second domain be T 2 The influence analysis unit 23 uses the learning data T 1 A model M trained on 1 Here, model D i M 1 (1≦i≦d) is the machine-learned model M 1 In the token embedding corresponding to the input word, operation D i This is a model that has been modified to include the above. i M 1 In the model D, the i-th dimension of the token embedding corresponding to the input word is forced to 0. In other words, i M 1 forces the removal of information in the i-th dimension of the token embedding corresponding to the input word.
[0064] The influence analysis unit 23 calculates L expressed by the following formula (7) for each i in the range of 1≦i≦d. i Request.
[0065]
number
[0066] L calculated by the above formula (7) i is the training data T i The model M trained with i For the training data T 2This is the change in the loss function when applying the L in Equation (7) (when there is no information on the i-th dimension). i is calculated as the difference in log-likelihood. i is the logarithm of the likelihood ratio. Also, L i is calculated as a value specific to the i-th dimension of the word-only embedding layer. That is, L i represents the degree of adverse effect of each dimension of the word-only embedding layer.
[0067] Here, L(M,T) is the loss when learning model M with learning data T. That is, L expressed by Equation (3) i is a model D that removes the i-th dimension of the token embedding. i M 1 (M 1 is the data T 1 (the model trained in 1 Loss from data T in the case 2 This is the value obtained by subtracting the loss in the case
[0068] The impact analysis unit 23 performs the above L for each i (1≦i≦d). i Calculate L for the first and second domains. i The larger the value of L, the more susceptible the i-th dimension is to domain differences. i The smaller the value of is, the less the i-th dimension is affected by the difference in domains. In other words, the training data for the first domain T 1 A model M trained on 1 Based on this, the training data for the second domain T 2 When learning using L i By performing learning only on dimensions where the value of L is relatively large, the learning efficiency can be improved. i Even if learning is omitted (parameter updates are not performed) for dimensions with a relatively large value, the impact of omission is relatively small.
[0069] The impact analysis unit 23 calculates the L i The values of are sorted in descending order, and the dimensions corresponding to the top N values (N is a positive integer) are taken as dimensions that are susceptible to domain differences. As a result, the word embedding layer is divided into a part that is sensitive to domain differences and a part that is not. The word embedding layer is a d-dimensional vector (whose elements are continuous values). In other words, the part of the information in this word embedding layer that is sensitive to domain differences is an N-dimensional continuous value, R N The part that is not affected by the difference in domain is a continuous value of (dN) dimension, R d-N In other words, the space is expressed as in the following formula (8).
[0070]
number
[0071] The representation x of the embedding layer for a single word is expressed as the following equation (9) by rearranging the dimensions that are susceptible to domain influences to the front in the dimensional notation. If this x is expressed as the sum of the part that is susceptible to domain differences and the part that is not, as shown in equation (8) above, it is expressed as the following equation (10).
[0072]
number
[0073]
number
[0074] In the above example, the impact analysis unit 23 separates the dimension that depends on the difference in domain from the dimension that does not depend on the difference in domain based on a predetermined value N. The impact analysis unit 23 may separate the dimension of the vector into the dimension that depends on the difference in domain and the dimension that does not depend on the difference in domain based on the value of N determined by other methods. For example, the impact analysis unit 23 may divide the L calculated by the above formula (7) intoi The dimension whose value is equal to or greater than a predetermined threshold may be regarded as a dimension susceptible to the influence of the domain, and the other dimensions may be regarded as dimensions not susceptible to the influence of the domain. i By using a clustering technique on the values of i The clusters with large values of L i In other words, even in these cases, the influence analysis unit 23 divides the d-dimensional word-only embedding layer representation into N dimensions (dimensions that are easily affected by domain differences) and (dN) dimensions (dimensions that are not easily affected by domain differences).
[0075] Based on the results of the above analysis, the impact analysis unit 23 outputs information (dimension information) indicating which dimensions are likely to be affected by the domain and which dimensions are not likely to be affected by the domain. Specifically, the impact analysis unit 23 passes this dimension information to the learning unit 31. In other words, the dimension information output by the impact analysis unit 23 enables the learning unit 31 to learn the model 32 for only specific dimensions. When the learning unit 31 receives this dimension information, it becomes possible for the learning unit 31 to efficiently learn about a new domain.
[0076] After the influence analysis unit 23 divides the dimensions in the above manner, the training unit 31 can perform training in different domains. The training unit 31 retrains only the internal parameters of the model 32 that correspond to the dimensions that are susceptible to domain influence. In other words, the training unit 31 updates the parameters of the fully connected network of the word embedding layer of the model 32 only for the dimensions that are identified by the influence analysis unit 23 as being susceptible to domain influence in the word embedding layer alone.
[0077] The output X to the word-only embedding layer is calculated using equation (11) below.
[0078]
number
[0079] The word-only embedding representation obtained by X above is passed to the subsequent processing stages of the model. When the model is BERT, X is added to other embeddings and passed to a model called Attention (attention layer 114 in FIG. 2). A softmax function is applied in the final stage of the model (softmax layer 115 in FIG. 2), and the natural language analysis result (in this example, a single classification label or a classification label sequence) is output.
[0080] In normal learning, all the parameters of A and B and the parameters of the other parts are updated so as to minimize the cross entropy (loss) of the value before the softmax function in the final stage. However, the learning unit 31 in this embodiment updates only some of the parameters of A and B when re-learning the model 32 using a different domain. In other words, the learning unit 31 only learns about the dimensions (the N dimensions) that are easily influenced by the domain. That is, the learning unit 31 does not further learn about the dimensions (the (dN) dimensions) that are not easily influenced by the domain. This is because the dimensions that are not easily influenced by the domain have already been learned and the results of that learning may be used. This makes it possible to efficiently proceed with learning even if the data of the domain to be newly learned is small.
[0081] The i-th component of AZ+B (argument to the sigmoid function) shown in the above equation (11) is expressed by the following equation (12).
[0082]
number
[0083] The learning unit 31 of this embodiment selects A such that, among the internal parameters of the model 32, i corresponds to a dimension that depends on the domain. ij and b iPerform learning such that only [parameters] are the targets of parameter update.
[0084] FIG. 6 is a block diagram showing an example of the internal configuration of the learning control device 1 in the above embodiment. The learning control device 1 can be realized using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be realized using existing technologies. The central processing unit 901 executes instructions included in a program read from the RAM 902 or the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic operations and logical operations according to each instruction. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "random access memory". The input / output port 903 is a port for the central processing unit 901 to exchange data with external input / output devices and the like. The input / output devices 904 and 905 are input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads and writes data in the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906.
[0085] At least some of the functions of the learning control device 1 can be realized by a computer and a program. In that case, the program for realizing the functions may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into a computer system and executed to realize the functions. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, CD-ROMs, DVD-ROMs, and USB memories, and storage devices such as hard disks built into a computer system. In other words, the term "computer-readable recording medium" may be a non-transitory computer-readable recording medium. The term "computer-readable recording medium" may also include a medium that temporarily and dynamically holds a program, such as a communication line when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain period of time, such as a volatile memory inside a computer system that is a server or client in that case. The above-mentioned program may be a program for realizing some of the functions described above, or may be a program that can be realized in combination with a program already recorded in the computer system.
[0086] Although the embodiment has been described above, the present invention can also be embodied in the following modified examples. A plurality of modified examples may be combined to the extent that they can be combined.
[0087] [Modification 1: Example in which the device does not have the learning unit 31 and the model 32] In the above embodiment, the learning control device 1 includes the learning unit 31 and the model 32. Based on the dimension information output by the influence analysis unit 23, the learning unit 31 performs learning such that only the internal parameters related to a specific dimension in the feature space among the internal parameters of the model 32 are updated. In this modified example 1, the learning control device 1 does not have the learning unit 31 or the model 32. In this modified example 1, the influence analysis unit 23 outputs information on the dimension (dimension of the subspace) determined. In other words, the influence analysis unit 23 outputs information indicating which dimension is susceptible to the influence of the difference in domains. The learning control device 1 (which may be called a "learning support device") outputs this dimension information to the outside. That is, in this modified example, the learning control device 1 outputs information indicating whether the dimension in the feature space is susceptible or not susceptible to the influence of the difference in domains for two given domains. In this case, the external device that receives the information can perform learning only on the specific dimension of the model based on the received information. In other words, when an external device wants to learn a model that has already been learned in one domain in another domain, the external device can perform learning that updates only parameters related to a specific dimension based on the information output by the learning control device 1 of Modification 1. This improves the efficiency of learning.
[0088] [Variation 2: When there are three or more domains] In the above embodiment, the process for learning a model for two different domains has been described, but as a modification, the number of domains may be three or more.
[0089] First, we will explain the case where the number of domains is three. Let the learning data corresponding to the three domains be T 1 ,T 2 ,T 3 Let R be the domain-specific subspace of the token embedding. What this modification aims to achieve is to 1 ,T 2 ,T 3The goal is to find the subspace R in the token embedding feature space when
[0090] As a first step, we select two of the three domains. The training data sets corresponding to the two selected domains are (T 1 ,T 2 ), (T 2 ,T 3 ), (T 3 ,T 1 ) There are three ways.
[0091] As a second step, for each of the three sets, the domain-specific part (feature subspace) R N Here again, N is a positive integer that is predetermined or determined by other suitable methods. The domain-specific parts corresponding to each of the above pairs are respectively denoted as R 12 , R 23 , R 31 The value of N, which represents the number of dimensions of the subspace, is expressed as R 12 , R 23 , R 31 may be different for each of the
[0092] As a third step, the above determined R 12 , R 23 , R 31 Based on the above, a subspace R for the original three domains is determined. Here, R may be, for example, any of the following. For example, R can be R 12 , R 23 , R 31 The Japanese space (R 12 ∪R 23 ∪R 31 ) or R may be R 12 , R 23 , R 31 The product space of (R 12 ∩R 23 ∩R 31By at least one of these methods, a subspace R of the feature space can be determined for the three given domains. For example, by conducting an experiment, the above sum space (R 12 ∪R 23 ∪R 31 ) or product space (R 12 ∩R 23 ∩R 31 ) may be determined as the subspace R.
[0093] In the above, we have explained how to obtain the subspace R when there are three domains. However, when there are four or more domains, the subspace R of the features for those domains can be obtained by using a similar idea. That is, a set of two domains selected from four or more domains (the learning data of those domains is called T U , T V For each of the subspaces R UV and calculate all subspaces R UV For example, the sum or product space of these domains may be a subspace R for the entire domain. If the number of domains is Z, the method for selecting two domains from these Z domains is as follows: Z C 2 ) as stated above.
[0094] [Variation 3: Example of using a model other than BERT] In the above embodiment, a model based on BERT is used as a model for performing natural language processing. In the third modification, a deep learning model other than BERT is used. In this case, the method described in the above embodiment is applied to the output of the embedding layer for words alone. In this case, it is also possible to identify only the dimensions in the feature space that are susceptible to influence across domains. In other words, it is possible to perform learning only on the dimensions that are susceptible to influence across domains. In other words, the efficiency of learning the model for natural language processing is improved.
[0095] [Variation 4: No segment embedding or position embedding, or combined with other embeddings] In the above embodiment, as described with reference to FIG. 3, the sum of token embedding, segment embedding, and position embedding is used as the input embedding corresponding to each input symbol. In this modification, either one or both of the segment embedding and the position embedding may be omitted. Alternatively, the input embedding may include other types of embedding. Other types of embedding are embeddings based on other information, such as embeddings of meta information such as modification relationships (dependencies) between input tokens and the time of sending a sentence. In any of these cases, the focus is on decomposing the algebraic space using word-only embeddings.
[0096] [Variation 5: Operation of the Impact Analysis Unit] In the above embodiment, as described with reference to FIG. 5, the influence analysis unit 23 performs a process of replacing the numerical value of the i-th dimension of the word alone embedding layer with 0 in order to examine the influence of the i-th dimension of the word alone embedding layer. Instead, in this modified example, the model parameters A and B in equation (4), which is the output function to the word alone embedding layer, are determined so that the i-th dimension in the sigmoid is always a fixed value (e.g., 0). In other words, when equation (4) is written in components, the i-th component is as shown in the following equation (13).
[0097]
number
[0098] That is, the element in the i-th row and j-th column of A, ij and B, the i-th element of B i For example, all of X and V are set to 0. Here, j is the dimension of the vocabulary and has a value ranging from 1 to |V|. In this modification, X iSince the argument of the sigmoid in equation (13) for calculating is always 0, the sigmoid function value is always 0.5. In both the above embodiment and this modification, the numerical value of the i-th dimension of the word independent embedding layer can be set to a fixed value, and the influence analysis unit 23 can analyze the degree of influence of each dimension by analysis using this.
[0099] The above describes in detail the embodiments and modified examples of the present invention with reference to the drawings, but the specific configurations are not limited to these embodiments, and also include designs that do not deviate from the gist of the present invention. [Industrial Applicability]
[0100] The present invention can be used in the field of natural language processing using trainable models to adapt the models to multiple domains. As an example, the present invention can be used in the process of classifying natural language expressions (sentences, word strings), but the application is not limited thereto. [Explanation of symbols]
[0101] 1 Learning control device (learning support device, natural language processing device) 21 Learning data storage unit (T 1 ) 22 Learning data storage unit (T 2 ) 23 Impact Analysis Department 31 Learning Department 32 Models 100 Models 101 fully connected neural network layers 111 Word-only Embeddings 112 Embedded Expressions 113 Synthesis Section 114 Attention Layer 115 Softmax Layer 901 Central Processing Unit 902 RAM 903 Input / Output Ports 904,905 Input / Output Devices 906 Bus
Claims
1. an influence analysis unit that, for a machine-learned model that processes at least word-only embedded expressions corresponding to natural language expressions, analyzes whether or not a specific component is susceptible to the influence of a difference between a first domain and a second domain based on the magnitude of the absolute value of the difference between the likelihood of a loss function calculated from a result of removing a specific component of a word-only embedded expression based on first training data belonging to a first domain and the likelihood of a loss function calculated from a result of removing the specific component of a word-only embedded expression based on second training data belonging to a second domain, and outputs information indicating whether or not the specific component is susceptible to the influence of a difference between the first domain and the second domain; A natural language processing device comprising:
2. The specific component is information of each dimension of the word-only embedding representation represented as a vector. The natural language processing device according to claim 1 .
3. the influence analysis unit classifies each dimension of the vector of the word alone embedded representation into either a side that is easily influenced by the difference between the first domain and the second domain or a side that is not easily influenced by the difference between the first domain and the second domain based on the magnitude of the absolute value of the difference between the likelihood of a loss function calculated when information of a specific dimension of the word alone embedded representation based on the first training data is forcibly changed to a predetermined value and the likelihood of a loss function calculated when information of the specific dimension of the word alone embedded representation based on the second training data is forcibly changed to the predetermined value, and outputs information indicating which side each dimension belongs to as information indicating for each dimension whether it is a component that is easily influenced by the difference between the first domain and the second domain. The natural language processing device according to claim 2 .
4. A natural language processing model trained on the first training data; a learning unit that performs learning using the second learning data for only a predetermined component of the natural language processing model based on information indicating whether or not the specific component is susceptible to an influence of a difference between the first domain and the second domain output by the influence analysis unit; The natural language processing apparatus according to claim 1 , further comprising:
5. Computer, an influence analysis unit that, for a machine-learned model that processes at least word-only embedded expressions corresponding to natural language expressions, analyzes whether or not a specific component is susceptible to the influence of a difference between a first domain and a second domain based on the magnitude of the absolute value of the difference between the likelihood of a loss function calculated from a result of removing a specific component of a word-only embedded expression based on first training data belonging to a first domain and the likelihood of a loss function calculated from a result of removing the specific component of a word-only embedded expression based on second training data belonging to a second domain, and outputs information indicating whether or not the specific component is susceptible to the influence of a difference between the first domain and the second domain; A natural language processing device comprising: A program to function as a
Citation Information
Patent Citations
Document data processing device, document data processing method and program
JP2021096552A