Learning device, method for learning, and learned model
The learning device enhances the performance of base models by aligning noisy features with original features through a transportation cost-based update, addressing limitations in existing model performance and improving robustness for downstream tasks.
Patent Information
- Application Number
- JP2023192264
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-22
AI Technical Summary
The performance of a base model in tasks like document classification and information extraction is often limited by its pre-training on a large public corpus, and it is difficult to confirm latent model characteristics that affect downstream target model performance.
A learning device that includes a data acquisition unit, a generation unit, a feature calculation unit, a cost calculation unit, and an update unit, which acquires input data, adds noise to it, calculates features, determines a transportation cost to align noisy features with original features, and updates the model based on this cost.
This approach improves the performance of the base model by reflecting the overall structure in the loss function and focusing on contextual information, resulting in a robust model that is better suited for downstream tasks.
Smart Images

Figure 2025079533000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a learning device, a learning method, and a trained model. [Background technology]
[0002] When performing tasks such as document classification and information extraction using a trained model, it is necessary to construct a target model specialized for the purpose in order to obtain high performance.When constructing a target model, a two-stage approach is often adopted, in which a model pre-trained on a huge public corpus is used as a base model, and the base model is further trained on a corpus corresponding to the purpose.
[0003] Since the target model inherits the characteristics of the base model, the final performance of the target model may depend on the performance of the base model. Furthermore, the performance of the base model does not fully contribute to the performance of the downstream target model, and latent model characteristics that do not appear as performance values may affect the performance of the target model. However, it is difficult to confirm such latent model characteristics during the model training stage, and it is also difficult to obtain feedback from users. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2022-2080 Summary of the Invention [Problem to be solved by the invention]
[0005] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a learning device, a method, and a trained model that improve the performance of a base model. [Means for solving the problem]
[0006] The learning device according to this embodiment includes a data acquisition unit, a generation unit, a feature calculation unit, a cost calculation unit, and an update unit. The data acquisition unit acquires a first token sequence obtained by dividing input data into token units. The generation unit generates a second token sequence by adding noise to the first token sequence. The feature calculation unit calculates a first feature from the first token sequence using a model that extracts features, and calculates a second feature from the second token sequence. The cost calculation unit calculates a transportation cost when the second feature is brought closer to the first feature. The update unit updates the model based on the transportation cost. [Brief description of the drawings]
[0007] [Figure 1] FIG. 1 is a block diagram showing a learning device according to a first embodiment. [Diagram 2] 4 is a flowchart showing an example of the operation of the learning device according to the first embodiment. [Diagram 3] FIG. 13 is a diagram showing an example of a transport matrix. [Figure 4] FIG. 13 is a diagram showing an example of a relationship diagram corresponding to a transport matrix. [Diagram 5] FIG. 11 is a block diagram showing a learning device according to a second embodiment. [Figure 6] 10 is a flowchart showing an example of the operation of the learning device according to the second embodiment. [Figure 7] FIG. 11 is a view showing a display example of a user interface according to the second embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a hardware configuration of a learning device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Hereinafter, a learning device, a learning method, and a trained model according to the present embodiment will be described in detail with reference to the drawings. In the following embodiments, parts with the same reference numerals perform the same operations, and duplicated descriptions will be omitted as appropriate.
[0009] (First embodiment) The learning device according to the first embodiment will be described with reference to the block diagram of FIG. The learning device 10 according to the first embodiment includes a storage unit 101, a data acquisition unit 102, a division unit 103, a generation unit 104, a feature calculation unit 105, a cost calculation unit 106, and an update unit 107.
[0010] The storage unit 101 stores a machine learning model, input data used for training the machine learning model, a trained model, and the like. The machine learning model is a model capable of extracting features in natural language processing, and is assumed to be, for example, a large-scale language model (LLM) such as BERT (Bidirectional Encoder Representations from Transformers) or GPT-3, GPT-3.5, GPT-4 of the GPT (Generative Pre-trained Transformer) series. Note that the present invention is not limited to this, and any machine learning model may be used as long as it can be a base model for pre-training before downstream tasks. The input data is assumed to be text data such as sentences. The trained model includes a network layer that processes the input data and infers output data.
[0011] The data acquiring unit 102 acquires input data or a first token sequence obtained by dividing the input data into token units. The division unit 103 divides the input data into token units.
[0012] The generating unit 104 generates a second token sequence by adding noise to the first token sequence. The feature calculation unit 105 uses a model for extracting features to calculate a first feature from the first token sequence, and calculates a second feature from the second token sequence.
[0013] The cost calculation unit 106 calculates the transportation cost required to bring the second feature closer to the first feature. The transportation cost is the cost required to restore the second token sequence to the first token sequence. The update unit 107 updates the model based on the transportation costs and generates a trained model.
[0014] Next, an example of the operation of the learning device 10 according to the first embodiment will be described with reference to the flowchart of FIG. In step SA1, the data acquiring unit 102 acquires input data, for example, from the storage unit 101. The input data is assumed to be, for example, text that can be acquired from a corpus. Note that the data acquiring unit 102 is not limited to acquiring the input data from the storage unit 101, and may acquire the input data from an external server that stores a large-scale corpus.
[0015] In step SA2, the division unit 103 divides the input data into token units to generate a first token sequence. The division of the input data into tokens may be performed using a general method such as dividing the input data into morpheme units or word units using a tokenizer. Tokens that are difficult to handle, such as symbols and numbers, may be normalized or removed. The division unit 103 may divide the input data into more granular token units using a tokenizer capable of finer division.
[0016] In step SA3, the generation unit 104 copies the first token sequence, and adds at least one of masking (filling) and rearrangement to the copied token sequence as noise. The masking may be performed by masking "input length of token sequence × p_{mask}" tokens according to a probability value p_{mask}. Specifically, if the probability value P is "0.1", 10% of the tokens in the token sequence will be masked. Thus, if the token sequence is formed of 10 tokens, one token will be masked.
[0017] The shuffling process can be done by changing the order of "token string input length × p_{shuffle}" tokens according to the probability value p_{shuffle}. Specifically, if the probability value is "0.5", the order of 50% of the tokens in the token string will be changed, so if the order of 10 tokens is "1, 2, 3, 4, 5, 6, 7, 8, 9, 10", the tokens will be shuffled as "1, 2, 3, 6, 4, 5, 7, 10, 9, 8". In this way, at least one of the masking and rearrangement processes is applied, i.e., noise is added, to generate a second token sequence.
[0018] In step SA4, the feature calculation unit 105 uses a model for extracting features to extract a first feature from the first token sequence and a second feature from the second token sequence. The model may be any model capable of extracting a feature expression, and may use, for example, BERT to extract a feature expression vector for each token as the first feature for the first token sequence, and generate a vector sequence for the first token sequence. The second feature may be generated in the same manner as the first feature.
[0019] In step SA5, the cost calculation unit 106 calculates the transportation cost when the second token sequence, which is a token sequence to which noise has been added, is returned to the first token sequence, which is a token sequence that does not contain noise. Since the second token sequence is the same token sequence as the first token sequence before noise is added, the correspondence between them can be grasped by the learning device 10. The cost calculation unit 106 may calculate the transportation cost by solving the process of returning the second token sequence to the first token sequence as an optimal transportation problem. Note that a general method may be used for the optimal transportation problem, and therefore a detailed description thereof will be omitted here.
[0020] In step SA6, the update unit 107 calculates a loss value based on a loss function including a transportation cost. For example, if the model is BERT, the loss function L MLMand the loss function L for NSP (Next Sentence Prediction) NSP and the loss function L for transportation costs OP For example, the loss function L in equation (1) can be used as the loss function for the entire model. L=L MLM +L NSP +L OP (1) In addition, the loss function L for the entire model is the loss function L for transportation costs. OP Alternatively, the loss value may be calculated using only each loss function L MLM ,L NSP and L OP The weighted sum of these may be used as the loss function L for the entire model.
[0021] In step SA7, the update unit 107 judges whether the training of the model is completed. For example, it judges whether the loss value is equal to or less than a threshold and has converged. If the loss value is equal to or less than a threshold and has converged, it is determined that the training of the model is completed, and the process proceeds to step SA9. On the other hand, if the loss value is greater than the threshold or has not converged, the process proceeds to step SA8. Note that, although an example of judging the end of training based on the loss value is shown here, the present invention is not limited to this, and a general training end judgment method for training a model using a loss function L, such as whether training has been repeated for a predetermined number of epochs, may be used.
[0022] In step SA8, the update unit 107 updates the model by updating parameters such as the weights and biases of the model, and then the process returns to step SA6 to repeat the same processing. In step SA9, the storage unit 101 stores the generated trained model.
[0023] Next, an example of the transportation cost will be described with reference to FIG. 3 and FIG. FIG. 3 shows a transport matrix 30 (also called an alignment matrix) in which the second token sequence 32 is the vertical column and the first token sequence 31 is the horizontal column, and the combinations of the corresponding relationships are illustrated in a matrix format.
[0024] In FIG. 3, the first token sequence 31 is assumed to be "There is a problem with the temperature control of the power supply and the temperature characteristics of the resistor," and the second token sequence 32 is assumed to be "There is a problem with the temperature control of the resistor and the power supply [MASK]." The first token sequence 31 and the second token sequence 32 are each assumed to be converted into a vector sequence of features. In addition, the underscore "_" in each of the above token sequences is a special character that indicates the separator between tokens. As the output result of the model, the grid of the combination with the highest degree of certainty of correspondence between each token in the first token sequence 31 and a token in the second token sequence 32 is indicated by diagonal lines. For example, as a result of feature calculation by the model, it can be seen that "[MASK]" in the second token sequence 32 is estimated to correspond to the token "temperature characteristic" in the first token sequence 31. Note that, for convenience of explanation, FIG. 3 shows the transport matrix 30 as a grid, and a value of certainty is associated with each grid.
[0025] Figure 4 is a relationship diagram 40 that displays the correspondence relationship on a token-by-token basis to facilitate visual understanding of the transport matrix 30 shown in Figure 3. Specifically, in Figure 4, the second token string 32 is arranged in the left column, and the first token string 31 is arranged in the right column, with corresponding tokens connected by lines.
[0026] Feature value x of token in the first token sequence 31 i and the token feature quantity y j The distance between the feature representations is the transport amount, and the similarity between the feature representations is the cosine similarity (cos sim (x i , y j )) transportation volume = 1-cos sim (x i ,y j ), where i and j are natural numbers greater than or equal to 1, and x i and y j is a vector.
[0027] In other words, a similar feature expression has a small transportation amount, and a dissimilar feature expression has a large transportation amount. In this embodiment, since the second token sequence 32 is generated based on the first token sequence 31, the learning device 10 can grasp the transportation route for returning from the second token sequence 32, which is a token sequence with noise added, to the first token sequence 31, which is a token sequence with no noise added. Therefore, the cost calculation unit 106 can calculate the transportation cost and the transportation matrix by solving the optimal transportation problem based on the transportation amount between each feature expression of the first token sequence 31 and the second token sequence 32 and calculating the optimal transportation route.
[0028] In the example of Fig. 3, as described above, [MASK] in the first token sequence 31 and "temperature characteristics" in the second token sequence 32 are different on the surface (character strings themselves), but a transport matrix is generated as corresponding to them. A model that can generate such a correct correspondence can be said to be a highly accurate model that can correctly read the context.
[0029] According to the first embodiment described above, a model is applied to a second token sequence generated by adding at least one of masking and rearrangement noise to a first token sequence that is input. As an output result of the model, a correspondence relationship between token features (feature vectors) is output. Based on the correspondence relationship, an optimal transportation problem for returning the second token sequence to the first token sequence is solved, and a transportation cost is calculated. The model is updated using a loss function including the calculated transportation cost, and the model is trained. This allows us to train the model by reflecting the overall structure in the loss function and focusing on the entire token while taking into account contextual information. Therefore, we can quantitatively evaluate the quality of the feature representation from the loss value, and also qualitatively evaluate it from the transport matrix obtained as a by-product of the optimal transport problem, resulting in the construction of a base model that is robust to noise. In other words, we can improve the performance of the base model for downstream target models, such as similar fault document search and faulty part extraction.
[0030] Second embodiment The second embodiment differs from the first embodiment in that a transportation matrix obtained by solving an optimal transportation problem is presented to a user, feedback information is obtained from the user, and the feedback information is utilized for training a model.
[0031] A learning device 10 according to the second embodiment will be described with reference to the block diagram of FIG. The learning device 10 of the second embodiment includes a storage unit 101, a data acquisition unit 102, a division unit 103, a generation unit 104, a feature calculation unit 105, a cost calculation unit 106, an update unit 107, a display control unit 201, and a feedback acquisition unit 202.
[0032] The display control unit 201 controls a user interface (UI) so as to display the transportation queue, transportation costs, and the like to the user via a display device (not shown).
[0033] The feedback acquisition unit 202 acquires, as feedback information, input values and operations input by the user via the UI for model training.
[0034] Next, an example of the operation of the learning device 10 according to the second embodiment will be described with reference to the flowchart of FIG. Incidentally, steps SA1 to SA4 and steps SA7 to SA9 are similar to those in FIG.
[0035] In step SB1, the cost calculation unit 106 calculates a transport matrix. For example, a transport matrix obtained when solving an optimal transport problem may be used. In step SB2, the display control unit 201 displays the transport matrix.
[0036] In step SB3, the feedback acquisition unit 202 acquires at least one of an input of a value and an operation on the transport matrix by the user via the UI as feedback information. In step SB4, the cost calculation unit 106 calculates the transportation cost based on the transportation matrix reflecting the feedback information. The cost calculation unit 106 calculates the loss value based on a loss function including the transportation cost. Then, in step SA7, it is determined whether the loss value is equal to or smaller than the threshold value. If the loss value is greater than the threshold value, the model is updated in step SA8, and then the process returns to step SB1 to repeat the same process.
[0037] Next, a display example of a UI by the display control unit 201 according to the second embodiment will be described with reference to FIG. 7 is a token string with added noise, and here it is assumed that the second token string 32 is displayed. The user drags and drops each token in the second token string 32 on the UI with a cursor 71 to rearrange them in the order of a token string that does not contain noise as expected by the user, thereby obtaining user-specified tokens 72.
[0038] Specifically, the second token string 32, “Problem_of_temperature_control_of_resistor_and_of_power_[MASK]”, is rearranged according to the token order of the first token string 31 while still including the "[MASK]" token, to generate the user-specified token 72. That is, it is assumed that the token is corrected to “Problem_of_power_[MASK]_and_of_temperature_control_of_resistor”. This correction makes it possible to measure the difference between the token correspondence indicated by the learning device 10 and the token correspondence expected by the user. Furthermore, since it is possible to distinguish the difference in the recognition of the token correspondence between the learning device 10 and the user, it can be used to identify the parts that need to be corrected in the target model generated downstream.
[0039] The center and right side of Fig. 7 show the correspondence between a user-specified token 72 and a first token string 31. Here, the correspondence estimated by the learning device 10 is shown by a dashed path 73. The user can correct the path connecting the tokens on the UI by using a cursor 71. The correspondence corrected by the user is shown by a solid path 74.
[0040] In the example of Fig. 7, the learning device 10 estimates that the token "temperature control" in the user-specified tokens 72 corresponds to the token "temperature characteristic" in the first token string 31, and connects them by a path 73. The user recognizes that there is an error in the correspondence between the tokens, and in order to correct the correspondence, the user connects the token "temperature control" in the user-specified tokens 72 to the token "temperature control" in the first token string 31, thereby creating a path 74. The feedback acquisition unit 202 may acquire the processing result related to the correction as feedback information.
[0041] Also, assume that the learning device 10 estimates that the token "[MASK]" in the user-specified tokens 72 corresponds to the token "temperature control" in the first token string 31, and that they are connected by a path 73. If the user recognizes that there is an error in the correspondence between the tokens, but that this is an acceptable error, the user can set multiple correspondences. Here, the user connects the token "[MASK]" and the token "temperature control" by a path 74, and sets the accuracy of the correspondence to "0.2". Furthermore, the token "[MASK]" and the token "temperature characteristics" are connected by a path 74, and set the accuracy of the correspondence to "0.8".
[0042] In this way, for example, even if the correspondence between tokens presented by the learning device 10 is incorrect, if the error is acceptable in consideration of the future use of the target model, it may be adopted as the correct answer, or multiple correspondences may be set. Similarly, even if the correspondence between tokens is correct, if it is unacceptable in consideration of the future use of the target model, it may be corrected as an error.
[0043] According to the second embodiment described above, the transport matrix is presented to the user, and feedback information is obtained from the user. By referring to the transport matrix displayed on the UI, the user can provide appropriate feedback while checking the training status of the model, such as whether the training is progressing as expected. This allows the learning device to perform optimal training while checking the potential model characteristics at the training stage through qualitative evaluation based on the feedback. As a result, the performance of the base model can be improved.
[0044] It should be noted that the user does not need to input information for all items presented in the UI, and the feedback acquisition unit 202 can accept and process information presented by the user to the extent that the user can input information.
[0045] Next, an example of the hardware configuration of the learning device 10 according to the above embodiment is shown in a block diagram in FIG. The learning device 10 includes a CPU (Central Processing Unit) 81, a RAM (Random Access Memory) 82, a ROM (Read Only Memory) 83, a storage 84, a display device 85, an input device 86, and a communication device 87, each of which is connected by a bus.
[0046] The CPU 81 is a processor that executes arithmetic processing and control processing according to a program. The CPU 81 uses a predetermined area of the RAM 82 as a working area and executes the processing of each part of the learning device 10 described above in cooperation with programs stored in the ROM 83 and the storage 84. Note that each process of the learning device 10 may be executed by one processor, or may be executed in a distributed manner by multiple processors.
[0047] The RAM 82 is a memory such as a Synchronous Dynamic Random Access Memory (SDRAM). The RAM 82 functions as a work area for the CPU 81. The ROM 83 is a memory that stores programs and various information in a non-rewritable manner.
[0048] The storage 84 is a device that writes and reads data to a magnetic recording medium such as a hard disk drive (HDD), a semiconductor storage medium such as a flash memory, a magnetically recordable storage medium such as a HDD, an optically recordable storage medium, etc. The storage 84 writes and reads data to the storage medium in response to control from the CPU 81.
[0049] The display device 85 is a display device such as an LCD (Liquid Crystal Display), etc. The display device 85 displays various information based on a display signal from the CPU 81. The input device 86 is an input device such as a mouse, a keyboard, etc. The input device 86 receives information input by a user as an instruction signal, and outputs the instruction signal to the CPU 81. The communication device 87 communicates with external devices via a network under the control of the CPU 81 .
[0050] The instructions shown in the processing procedure shown in the above-mentioned embodiment can be executed based on a program, which is software. A general-purpose computer system can store this program in advance and obtain the same effect as the control operation of the learning device described above by reading this program. The instructions described in the above-mentioned embodiment are recorded as a program that can be executed by a computer on a magnetic disk (flexible disk, hard disk, etc.), an optical disk (CD-ROM, CD-R, CD-RW, DVD-ROM, DVD±R, DVD±RW, Blu-ray (registered trademark) Disc, etc.), a semiconductor memory, or a recording medium similar thereto. The recording medium may be in any storage format as long as it is readable by a computer or an embedded system. The computer can realize the same operation as the control of the learning device in the above-mentioned embodiment by reading the program from this recording medium and executing the instructions described in the program based on this program with a CPU. Of course, when the computer acquires or reads the program, it may acquire or read it through a network. In addition, an OS (operating system), database management software, network, or other MW (middleware) running on a computer may execute some of the processes required to realize this embodiment based on instructions from a program installed on the computer or embedded system from a recording medium. Furthermore, the recording medium in this embodiment is not limited to a medium independent of a computer or an embedded system, but also includes a recording medium that stores or temporarily stores a program downloaded via a LAN, the Internet, or the like. Furthermore, the number of recording media is not limited to one, and cases in which the processing in this embodiment is executed from multiple media are also included in the recording media in this embodiment, and the media may have any configuration.
[0051] The computer or embedded system in this embodiment is for executing each process in this embodiment based on a program stored in a recording medium, and may be configured as any one of a device such as a personal computer or a microcomputer, or a system in which multiple devices are connected to a network. In addition, the computer in this embodiment is not limited to a personal computer but also includes an arithmetic processing device, a microcomputer, etc. included in information processing equipment, and is a general term for equipment or devices that can realize the functions in this embodiment by a program.
[0052] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included in the scope and spirit of the invention, and are included in the scope of the invention and its equivalents described in the claims. [Explanation of symbols]
[0053] 10...Learning device, 30...Transport matrix, 31...First token sequence, 32...Second token sequence, 40...Relationship diagram, 71...Cursor, 72...User-specified token, 73, 74, 75...Route, 84...Storage, 85...Display device, 86...Input device, 87...Communication device, 101...Storage unit, 102...Data acquisition unit, 103...Splitting unit, 104...Generation unit, 105...Feature calculation unit, 106...Cost calculation unit, 107...Update unit, 201...Display control unit, 202...Feedback acquisition unit
Claims
1. a data acquisition unit that acquires a first token sequence obtained by dividing input data into tokens; a generation unit that generates a second token sequence by adding noise to the first token sequence; a feature calculation unit that calculates a first feature from the first token sequence and calculates a second feature from the second token sequence using a model that extracts features; a cost calculation unit that calculates a transportation cost when the second feature amount is brought closer to the first feature amount; an update unit that updates the model based on the transportation costs; A learning device comprising:
2. The learning device according to claim 1 , wherein the cost calculation unit calculates the transportation cost based on a transportation matrix in an optimal transportation problem.
3. The learning device according to claim 1 , wherein the generation unit executes at least one of a token rearrangement process and a token mask process as the noise addition process.
4. The learning device according to claim 1 , wherein the update unit determines whether training of the model is terminated based on a loss function including the transportation cost.
5. The learning device according to claim 1 , further comprising a display control unit that displays a transportation matrix related to the calculation of the transportation cost to a user on a display device.
6. A feedback acquisition unit that acquires feedback information from the user regarding training of the model based on the transport matrix, The learning device according to claim 5 , wherein the cost calculation unit calculates a new transportation cost based on the feedback information.
7. A first token sequence is obtained by dividing the input data into tokens; generating a second token sequence by adding noise to the first token sequence; calculating a first feature from the first token sequence and a second feature from the second token sequence using a model for extracting a feature; calculating a transportation cost in a case where the second characteristic amount is brought closer to the first characteristic amount; updating the model based on the transportation costs; How to learn.
8. A trained model comprising a network layer that processes input data and infers output data, a generating step of generating a second token sequence by adding noise to the first token sequence; a feature calculation step of calculating a first feature from the first token sequence and calculating a second feature from the second token sequence using a model for extracting a feature; a cost calculation step of calculating a transportation cost when the second feature amount is brought closer to the first feature amount; updating the model based on the transportation costs; Trained by inputting the input data into the network layer with the updated parameters assigned thereto to infer the output data; A trained model for making a computer work.
Citation Information
Patent Citations
Artificial intelligence system for classifying data based on contrastive learning
JP2022002080A