Information processing device, information processing method, program
The described system addresses the challenge of selecting training data for machine learning models by generating and analyzing evaluation distributions to enhance data suitability for alignment, thereby improving model reliability and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-31
AI Technical Summary
The challenge of selecting appropriate training data for aligning machine learning models, particularly large language models, is difficult due to variations in evaluator backgrounds and changing preferences over time.
An information processing device and method that acquires datasets, generates distributions of evaluation values, calculates indices like diversity and universality, and outputs suitable datasets for alignment, considering evaluator characteristics and time-based changes.
Facilitates easy selection of training data for machine learning models, improving reliability and accuracy by accounting for evaluator biases and temporal variations.
Smart Images

Figure 2026055256000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In recent years, machine learning models that provide output in response to user requests have been utilized in various situations. For example, as described in Patent Document 1, large language models (LLMs) trained in language processing are being used, allowing users to input questions in language and obtain linguistic responses. After such large language models are trained using a vast amount of training data, they undergo a process called alignment to fine-tune them to produce outputs that are preferable to humans. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Patent No. 7404596 [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] However, as mentioned above, there is a problem in selecting the training data to be used when performing alignment after machine learning a large-scale language model. Furthermore, this problem is not limited to large-scale language models, but applies to various other machine learning models as well.
[0005] Therefore, one of the purposes of this disclosure is to solve the aforementioned problem of difficulty in selecting training data to be used when aligning machine learning models. [Means for solving the problem]
[0006] An information processing device, which is one form of this disclosure, An acquisition unit that acquires a set of data sets including input data input to a machine learning model, output data output from the machine learning model according to the input data, and an evaluation value representing an evaluation of the output data with respect to the input data; A generation unit that generates a distribution of the evaluation values in the set of data sets; A calculation unit that calculates a value of a preset index in the data set based on the distribution; Comprising It has the following configuration. Also, an information processing method according to an aspect of the present disclosure An information processing apparatus acquires a set of data sets including input data input to a machine learning model, output data output from the machine learning model according to the input data, and an evaluation value representing an evaluation of the output data with respect to the input data, generates a distribution of the evaluation values in the set of data sets, calculates a value of a preset index in the data set based on the distribution, It has the following configuration. Also, a program according to an aspect of the present disclosure Causes an information processing apparatus to acquire a set of data sets including input data input to a machine learning model, output data output from the machine learning model according to the input data, and an evaluation value representing an evaluation of the output data with respect to the input data, generate a distribution of the evaluation values in the set of data sets, calculate a value of a preset index in the data set based on the distribution, and execute the process, It has the following configuration.
Advantages of the Invention
[0007] With the present disclosure configured as described above, it is possible to easily select learning data used when aligning a machine learning model.
Brief Description of the Drawings
[0008] [Figure 1] This block diagram shows an example of the configuration of the evaluation device relating to this disclosure. [Figure 2] This block diagram shows an example of the configuration of the evaluation device relating to this disclosure. [Figure 3] This figure shows an example of the processing process of the evaluation device related to this disclosure. [Figure 4] This figure shows an example of the processing process of the evaluation device related to this disclosure. [Figure 5] This figure shows an example of the processing process of the evaluation device related to this disclosure. [Figure 6] This flowchart shows an example of the processing operation of the evaluation device related to this disclosure. [Figure 7] This block diagram shows an example of the hardware configuration of the information processing device related to this disclosure. [Figure 8] This is a block diagram showing an example of the configuration of the information processing device related to this disclosure. [Modes for carrying out the invention]
[0009] <First Embodiment> A first embodiment of this disclosure will be described with reference to the drawings. The drawings may be relevant to any embodiment.
[0010] The evaluation device 10 of this disclosure is used, for example, to assist in the selection of training data to be used when performing alignment, which is an additional training process that fine-tunes a machine learning model trained using training data to produce an output that is preferable to humans. In this embodiment, as an example, the machine learning model to be aligned is described as a large language model. However, the machine learning model to be aligned is not limited to large language models (LLMs), and may be any machine learning model that produces any output for any input. Accordingly, the training data used for machine learning may be any data.
[0011] The evaluation device 10 is composed of one or more information processing devices having an arithmetic unit and a storage unit. As shown in FIG. 1, the evaluation device 10 includes an acquisition unit 11, a generation unit 12, a calculation unit 13, and an output unit 14. Each function of the acquisition unit 11, the generation unit 12, the calculation unit 13, and the output unit 14 can be realized by executing a program for realizing each function stored in the storage unit by the arithmetic unit. Further, the evaluation device 10 includes a dataset storage unit 16 configured by a storage unit.
[0012] The acquisition unit 11 acquires a set of datasets of the query q
[0012] , i , i which is input data input to the LLM 20, i the output x i which is output data output from the LLM 20 in response to the query q, i the evaluation y i of the output x i for the query q, i and stores them in the dataset storage unit 16 (step S1 in FIG. 6). At this time, the dataset consisting of the query q i and the output x i and the query q i and the evaluation y may be an existing dataset, or may be a dataset acquired from the operating LLM 20 as shown in FIG. 2.
[0013] Here, a specific example of the dataset is shown in FIG. 3. The query q i of the dataset is a language-based question for the LLM 20, and as an example, "What is the capital of Japan?" can be cited. Further, the output x i of the dataset is a language-based answer from the LLM 20 to the above question, and as an example, "Tokyo." can be cited. And the evaluation y i of the dataset is an evaluation by a predetermined evaluator 21 for the answer to the above question, and as an example, there is "good". Note that the evaluation y iFor each answer to a question, if evaluator 21 evaluates it positively as reasonable or agreeable, it is represented as "good," and if they evaluate it negatively as not reasonable or disagreeable, it is represented as "bad." Therefore, evaluation y i Even for the same question, the evaluation score can vary depending on the evaluator's background knowledge and the passage of time (changes in the times). i This can be represented by discrete values, such as the binary values "good" and "bad" mentioned above, or by a wider range of stepped values, or by continuous values in the range of "0.0" to "1.0", where a more positive evaluation results in a higher value.
[0014] As shown in Figure 3, another example of a dataset is query q. i "Who is Japan's best soccer player?", Output x i "This is player XX, who holds the record for the most goals scored in history.", evaluation y i "bad" is one example. Another example of a dataset is query q i "Translate 'I love you' into Japanese," output x i "The moon is beautiful, isn't it?", rating y i "0.5" is one example. Another example of a dataset is query q i "Translate 'I love you' into Japanese," output x i "A literal translation would be 'I love you,' but there's a famous translation that Natsume Soseki taught people would understand if they translated it as 'The moon is beautiful, isn't it?'" (evaluation y) i "0.9" is one example. Thus, the dataset contains identical or similar queries q i Output x with different content i This also includes items that are as described above.
[0015] The acquisition unit 11 may also acquire option information associated with the dataset. The option information for the dataset may include query q. i There is input characteristic information that represents the characteristics of the subject, and evaluator characteristic information that represents the characteristics of the evaluator. Input characteristic information is, for example, query qi It is a type of query, and as an example, query q i The types of questions based on their content include questions, translations, summaries, etc. i Examples of types of expression include text and audio. In addition, evaluator characteristic information includes, for example, the evaluator's attributes, such as gender, age, place of residence, nationality, and identification information.
[0016] The generation unit 12 evaluates the set of datasets described above. i The distribution is generated (step S2 in Figure 6). At this time, the generation unit 12 further processes the set of datasets using the same or similar query q i And such identical or similar queries q i The same or similar output x i The data is classified into a similarity set (Q,X), which is a set of pairs of and , and a distribution of evaluations (evaluation Y) corresponding to each dataset in the similarity set (Q,X) is generated. In this case, the similarity set (Q,X) is a set of queries q with high similarity to sentences according to a predetermined criterion. i The set of and the query q i Output x for i Output x with high similarity to texts based on pre-set criteria. i It is composed of a set of and . The similar set (Q,X) may be classified, for example, by a large-scale language model, by another information processing device, or by human error. In addition, the dataset collected by the acquisition unit 11 may already form a similar set (Q,X).
[0017] The generation unit 12 then generates a distribution of evaluation Y in the similar set (Q,X) that represents the degree of variability and the degree of change over time of evaluation Y. For example, the generation unit 12 generates a distribution of discrete evaluation y i For this, a distribution of the degree of variation is generated as shown in Figure 4(4-1), and the evaluation of the continuous value y i For this, a distribution of the degree of variation is generated as shown in Figure 4(4-2). In addition, the generation unit 12 evaluates y iThe distribution of the degree of variation may be generated for each time interval, and the degree of change over time may be generated. Furthermore, the generation unit 12 evaluates the discrete value y i Regarding this, we generate odds ratios and cumulative probabilities of the binomial distribution, and evaluate continuous values y i For this, kurtosis, skewness, etc., may be generated. The generation unit 12 may generate any distribution as the distribution of evaluation Y in the similar set (Q,X).
[0018] Furthermore, the generation unit 12 generates query q as described above. i and output x i The similarity sets (Q,X) classified by the similarity of the queries may be further classified by the similarity of the optional information in the dataset. For example, query q i The queries may be classified into similar sets (Q,X) based on their type (question, translation, summary, text, audio, etc.) and the attributes of the evaluators (gender, age, place of residence, nationality, identification information, etc.). Then, the generation unit 12, as described above, classifies the queries q i A distribution of evaluation Y in similar sets (Q,X) classified by type and evaluator attributes may be generated.
[0019] The calculation unit 13 calculates the value of a pre-set index in the dataset based on the distribution of evaluation Y in the similar set (Q,X) generated as described above. Specifically, the calculation unit 13 calculates the index of evaluation Y that can be read from the distribution of evaluation Y in the similar set (Q,X). As an example, the calculation unit 13 calculates the diversity and universality of evaluation Y as an index of evaluation Y in the similar set (Q,X) (step S3 in Figure 6).
[0020] Specifically, the calculation unit 13 can calculate a diversity value corresponding to the degree of variability from the distribution of the degree of variability of evaluation Y as shown in Figure 4. For example, it can calculate a value such that the greater the variability, the greater the diversity. Furthermore, the calculation unit 13 can calculate a universality value corresponding to the degree of agreement of the distribution over time from the distribution of the degree of change of evaluation Y. For example, it can calculate a value such that the greater the degree of agreement of the distribution over time, the greater the universality. This allows for the calculation of queries q in a dataset as shown in Figure 3. i "What is the capital of Japan?", Output x i For the similar set (Q,X) corresponding to "Tokyo," the diversity value may be low and the universality value high because the answer to the question is factual and universal. On the other hand, for queries in a dataset like the one shown in Figure 3, q i "Who is Japan's best soccer player?", Output x i For the similar set (Q,X) corresponding to "This is player XX who has recorded the most goals scored in history," the diversity value may be high and the universality value low because the answers to the question depend on the respondent and change over time. Note that the calculation unit 13 may calculate any index from any distribution as described above, not limited to the diversity and universality mentioned above.
[0021] Furthermore, the calculation unit 13 calculates the anomaly score of a predetermined dataset included in the similar set (Q,X) as an index of that dataset, based on the distribution of evaluation Y of the similar set (Q,X) generated as described above (step S4 in Figure 6). Specifically, the calculation unit 13 calculates the evaluation y of a predetermined dataset within the similar set (Q,X) i The probability of occurrence is calculated, and the anomaly score is calculated according to this probability. In other words, the evaluation of a given dataset y relative to a standard evaluation within the similar set (Q,X) is calculated. i The degree of deviation is calculated as the anomaly score for a given dataset. For example, it can be calculated in a range of "0.0" to "1.0", where a lower probability of occurrence corresponds to a higher anomaly score.
[0022] The output unit 14 outputs the dataset based on the indicators such as diversity, universality, and anomaly calculated as described above (step S5 in Figure 6). For example, the output unit 14 outputs the dataset to the operator who operates the machine learning model, instructing them to display the dataset on a display device according to the indicators. At this time, as shown in Figure 5, the output unit 14 outputs the indicators such as diversity and universality calculated as described above, associated with the dataset. In addition, if the output unit 14 has calculated the anomaly of the dataset, it may also output this anomaly associated with the dataset. Furthermore, the output unit 14 outputs the query q, which is optional information associated with the dataset. i The type of evaluation and the attributes of the evaluators may be output in association with the dataset.
[0023] Furthermore, the output unit 14 may output datasets considering their suitability for use in alignment, based on the calculated indicators such as diversity, universality, and anomaly. For example, the output unit 14 may output datasets that meet criteria such as low diversity, high universality, and low anomaly as datasets suitable for use in alignment. Alternatively, the output unit 14 may output datasets that are unsuitable for alignment, depending on the indicators.
[0024] As described above, this disclosure calculates indicators such as diversity, universality, and anomaly score from the distribution of evaluations of the dataset. By referring to these indicators, machine learning model operators can easily select training data to be used when aligning their machine learning models. For example, machine learning model operators can select training data that prioritizes datasets with low diversity and high universality, or select training data that excludes datasets with high anomaly scores. Furthermore, by outputting optional information associated with the dataset, it becomes possible to select training data that takes into account evaluation biases due to the type of input to the dataset and the attributes of the evaluators. As mentioned above, by selecting training data in this way, machine learning models can be created at low cost, and the reliability and accuracy of the machine learning model's output can be improved.
[0025] <Second Embodiment> Next, a second embodiment of the present disclosure will be described with reference to the drawings. This embodiment shows an outline of the evaluation apparatus, etc., described in the above-described embodiment. Note that the drawings may be relevant to any embodiment.
[0026] First, the hardware configuration of the information processing device 100 in this disclosure will be described. The information processing device 100 is composed of a general information processing device, and as an example, it is equipped with the following hardware configuration as shown in Figure 7. ·CPU(Central Processing Unit)101(Arithmetic unit) ROM (Read Only Memory) 102 (Storage Device) • RAM (Random Access Memory) 103 (Storage Device) • Program group 104 loaded into RAM 103 • Storage device 105 for storing the program group 104 • Drive device 106 for reading and writing to external storage medium 110 of the information processing device. • Communication interface 107 connecting to a communication network 111 outside the information processing device. • Input / output interface 108 for data input and output. • Bus 109 connecting each component
[0027] Figure 7 shows an example of the hardware configuration of the information processing device 100, and the hardware configuration of the information processing device is not limited to the case described above. For example, the information processing device may consist of only a part of the configuration described above, such as not having a drive device 106. In addition, the information processing device may use a GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating point number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof instead of the CPU described above.
[0028] The information processing device 100 can be equipped with the acquisition unit 121, generation unit 122, and calculation unit 123 shown in Figure 8 by having the CPU 101 acquire the program group 104 and execute it. The program group 104 is, for example, stored in advance in a storage device 105 or ROM 102, and the CPU 101 loads it into RAM 103 and executes it as needed. The program group 104 may also be supplied to the CPU 101 via a communication network 111, or it may be stored in advance in a storage medium 110, and the drive device 106 reads the program and supplies it to the CPU 101. However, the acquisition unit 121, generation unit 122, and calculation unit 123 described above may be constructed with dedicated electronic circuits to realize such means.
[0029] The acquisition unit 121 acquires a set of datasets comprising input data input to a machine learning model, output data output from the machine learning model in accordance with the input data, and evaluation values representing the evaluation of the output data relative to the input data. The generation unit 122 generates a distribution of the evaluation values in the set of datasets. The calculation unit 123 calculates the values of pre-set indicators in the datasets based on the distribution.
[0030] As described above, this disclosure makes it easy to select a dataset as training data to be used when aligning a machine learning model by calculating an index from the distribution of evaluation values in the dataset and referring to such an index.
[0031] Furthermore, at least one of the functions of the acquisition unit 121, generation unit 122, and calculation unit 123 described above may be performed on an information processing device installed and connected at any location on the network, that is, it may be performed using so-called cloud computing.
[0032] Furthermore, the aforementioned programs can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). Programs may also be supplied to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can be supplied to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.
[0033] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure are possible, as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each of the embodiments described above can be combined with other embodiments as appropriate.
[0034] <Note> Some or all of the above embodiments may also be described as follows. The general configuration of the information processing apparatus, information processing method, and program in this disclosure is described below. However, this disclosure is not limited to the configurations described below. Furthermore, some or all of the configurations and functions described in Appendices 2 to 8, which are dependent on Appendice 1 below, may also be dependent on Appendices 9 and 10 in the same way as Appendices 2 to 8. Moreover, not limited to Appendices 1, 9, and 10, some or all of the configurations and functions described as appendices may also be dependent on similar hardware, software, various recording means for recording software, or systems, without departing from the embodiments described above. (Note 1) An acquisition unit acquires a set of datasets comprising input data input to a machine learning model, output data output from the machine learning model in accordance with the input data, and evaluation values representing the evaluation of the output data with respect to the input data. A generation unit that generates the distribution of the evaluation values in the set of datasets, A calculation unit that calculates the value of a pre-set indicator in the dataset based on the distribution, Equipped with an information processing device. (Note 2) The information processing device described in Appendix 1, The system includes an output unit that outputs information identifying the dataset based on the calculation results of the aforementioned indicators. Information processing device. (Note 3) The information processing device described in Appendix 2, The output unit outputs the calculation result of the indicator in association with information that identifies the dataset. Information processing device. (Note 4) The information processing device described in Appendix 1, The calculation unit calculates the diversity of the evaluation values in the dataset as an index based on the distribution. Information processing device. (Note 5) The information processing device described in Appendix 1, The calculation unit calculates the universality of the evaluation value in the dataset as an index based on the distribution. Information processing device. (Note 6) The information processing device described in Appendix 1, The calculation unit calculates the degree of abnormality of the evaluation value in the dataset as an index based on the distribution. Information processing device. (Note 7) The information processing device described in Appendix 2, The acquisition unit acquires the dataset, which includes input characteristic information representing the characteristics of the input data. The output unit outputs the input characteristic information in association with information that identifies the dataset. Information processing device. (Note 8) The information processing device described in Appendix 2, The acquisition unit acquires the dataset which includes evaluator characteristic information representing the characteristics of the evaluators of the evaluation values. The output unit outputs the evaluator characteristic information in association with the information that identifies the dataset. Information processing device. (Note 9) Information processing device, Obtain a set of datasets comprising input data to be input to a machine learning model, output data output from the machine learning model in response to the input data, and evaluation values representing the evaluation of the output data with respect to the input data. The distribution of the evaluation values in the set of the aforementioned datasets is generated. Based on the distribution, calculate the value of a pre-set indicator in the dataset. Information processing methods. (Note 9.1) The information processing method described in Appendix 9, Based on the calculation results of the aforementioned indicators, information identifying the dataset is output. Information processing methods. (Appendix 9.2) The information processing method described in Appendix 9.1, The calculation results of the aforementioned indicators are output in association with information that identifies the aforementioned dataset. Information processing methods. (Appendix 9.3) The information processing method described in Appendix 9.1, A dataset is obtained that includes input characteristic information representing the characteristics of the input data. The input characteristic information is output in association with information that identifies the dataset. Information processing methods. (Appendix 9.4) The information processing method described in Appendix 9.1, The dataset containing evaluator characteristic information representing the characteristics of the evaluators of the evaluation values is obtained, The evaluator characteristics information is output in association with the information that identifies the dataset. Information processing methods. (Note 10) In an information processing device, Obtain a set of datasets comprising input data to be input to a machine learning model, output data output from the machine learning model in response to the input data, and evaluation values representing the evaluation of the output data with respect to the input data. The distribution of the evaluation values in the set of the aforementioned datasets is generated. Based on the distribution, calculate the value of a pre-set indicator in the dataset. A program that executes a process. [Explanation of Symbols]
[0035] 10 Evaluation device 11 Acquisition Department 12 Generation part 13 Calculation Section 14 Output section 16 Dataset Storage Unit 20 LLM 21 evaluators 100 Information Processing Devices 101 CPU 102 ROM 103 RAM 104 Program Groups 105 Storage device 106 Drive unit 107 Communication Interface 108 Input / Output Interfaces 109 Bus 110 Storage medium 111 Communication Network 121 Acquisition Department 122 Generation part 123 Calculation Section
Claims
1. An acquisition unit acquires a set of datasets comprising input data input to a machine learning model, output data output from the machine learning model in accordance with the input data, and evaluation values representing the evaluation of the output data with respect to the input data. A generation unit that generates the distribution of the evaluation values in the set of datasets, A calculation unit that calculates the value of a pre-set indicator in the dataset based on the distribution, Equipped with an information processing device.
2. An information processing apparatus according to claim 1, The system includes an output unit that outputs information identifying the dataset based on the calculation results of the aforementioned indicators. Information processing device.
3. An information processing apparatus according to claim 2, The output unit outputs the calculation result of the indicator in association with information that identifies the dataset. Information processing device.
4. An information processing apparatus according to claim 1, The calculation unit calculates the diversity of the evaluation values in the dataset as an index based on the distribution. Information processing device.
5. An information processing apparatus according to claim 1, The calculation unit calculates the universality of the evaluation value in the dataset as an index based on the distribution. Information processing device.
6. An information processing apparatus according to claim 1, The calculation unit calculates the degree of abnormality of the evaluation value in the dataset as an index based on the distribution. Information processing device.
7. An information processing apparatus according to claim 2, The acquisition unit acquires the dataset, which includes input characteristic information representing the characteristics of the input data. The output unit outputs the input characteristic information in association with information that identifies the dataset. Information processing device.
8. An information processing apparatus according to claim 2, The acquisition unit acquires the dataset which includes evaluator characteristic information representing the characteristics of the evaluators of the evaluation values. The output unit outputs the evaluator characteristic information in association with the information that identifies the dataset. Information processing device.
9. Information processing device, Obtain a set of datasets comprising input data to be input to a machine learning model, output data output from the machine learning model in response to the input data, and evaluation values representing the evaluation of the output data with respect to the input data. The distribution of the evaluation values in the set of the aforementioned datasets is generated. Based on the distribution, calculate the value of a pre-set indicator in the dataset. Information processing methods.
10. In an information processing device, Obtain a set of datasets comprising input data to be input to a machine learning model, output data output from the machine learning model in response to the input data, and evaluation values representing the evaluation of the output data with respect to the input data. The distribution of the evaluation values in the set of the aforementioned datasets is generated. Based on the distribution, calculate the value of a pre-set indicator in the dataset. A program that executes a process.
Citation Information
Patent Citations
Information processing method, program, and information processing system
JP7404596B1