Data generation program, data generation method, and information processing device

WO2026203247A1PCT designated stage Publication Date: 2026-10-01FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012629
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-10-01

Smart Images

  • Figure JP2025012629_01102026_PF_FP_ABST
    Figure JP2025012629_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device detects a first agent that generates data and a second agent that evaluates the generated data, generates first data indicating a summary of content and a collection of questions and answers by executing the first agent with the content being used as input, generates an evaluation result of the first data by executing the second agent with the first data being used as input, and outputs RAG data by executing the first agent with the evaluation result of the first data being used as input.
Need to check novelty before this filing date? Find Prior Art

Description

Data generation program, data generation method and information processing apparatus

[0001] The present invention relates to a data generation program, a data generation method and an information processing apparatus.

[0002] In generative AI (Artificial Intelligence) such as LLM (Large Language Model), a technology is known that uses RAG (Retrieval Augmented Generation) to search a database for information related to a query and improve answer quality by utilizing the search results. Improving answer quality requires improving the quality of the search results themselves, and thus improving the quality of RAG data, which is data used for RAG, is demanded. In recent years, a technology that generates a QA (Question & Answer) collection from an original document and uses it as RAG data has been known.

[0003] Laurent Mombaerts, Terry Ding, Adi Banerjee, Florian Felice, Jonathan Taws, Tarik Borogovac, “Meta Knowledge for Retrieval Augmented Large Language Models”, arXiv: 2408.09017v1, 16, Aug, 2024

[0004] However, with the above technology, when the quality of the original document that is the source of RAG data is poor, the quality of the generated RAG data also decreases. For example, when the original document is a Web page, it contains unnecessary information such as links and advertisements, so incorrect QA may be generated, which reduces the quality of RAG data.

[0005] In one aspect, an object of the present invention is to provide a data generation program, a data generation method and an information processing apparatus that can improve the quality of RAG data.

[0006] In the first proposal, the data generation program instructs the computer to detect a first agent that generates data and a second agent that evaluates the generated data, to run the first agent with content as input to generate first data showing a summary of the content and a collection of questions and answers, to run the second agent with the first data as input to generate the evaluation result of the first data, and to run the first agent with the evaluation result of the first data as input to output RAG (Retrieval Augmented Generation) data.

[0007] In one respect, it can improve the quality of RAG data.

[0008] Figure 1 is a diagram illustrating the information processing device according to Embodiment 1. Figure 2 is a functional block diagram showing the functional configuration of the information processing device according to Embodiment 1. Figure 3 is a diagram illustrating the generation of RAG data. Figure 4 is a diagram illustrating the evaluation of RAG data. Figure 5 is a diagram illustrating the evaluation results of RAG data. Figure 6 is a diagram illustrating the regeneration of RAG data. Figure 7 is a flowchart showing the flow of the RAG data generation process according to Embodiment 1. Figure 8 is a diagram illustrating an example of hardware configuration.

[0009] The following describes in detail, with reference to the drawings, embodiments of the data generation program, data generation method, and information processing device according to the present invention. However, the present invention is not limited by these embodiments. Each embodiment can be combined as appropriate within a non-consistent range.

[0010] (Description of Information Processing Device) Figure 1 is a diagram illustrating the information processing device 10 according to Embodiment 1. The information processing device 10 shown in Figure 1 is an example of a computer that improves the quality of RAG data by accurately extracting the content of documents and other content, rather than using the content as RAG data as is. Furthermore, the generation of high-quality RAG data improves the accuracy of RAG (RAG search) performed by a language model including a large-scale language model, and also improves the accuracy of the output results of the language model. Thus, the processing performed by the information processing device 10 can be applied to applications that generate RAG data used by an LLM, or applications in which an LLM performs RAG and generates answers to questions. In this embodiment, an LLM is used as an example of a language model for explanation.

[0011] Specifically, the information processing device 10 detects a first agent that generates data and a second agent that evaluates the generated data. The information processing device 10 generates first data, which shows a summary of the content and a collection of questions and answers, by executing the first agent with the content as input. The information processing device 10 generates the evaluation result of the first data by executing the second agent with the first data as input. The information processing device 10 outputs RAG data by executing the first agent with the evaluation result of the first data as input.

[0012] Here, we will explain in detail using the original document D1 shown in Figure 1 as an example. For example, as shown in Figure 1, the information processing device 10 executes a generation agent 50, which is an example of a first agent, and an evaluation agent 60, which is an example of a second agent, and has each agent perform the subsequent processing. If each agent is already being executed, the information processing device 10 detects each agent that is being executed from the execution process, etc. Here, the generation agent 50 and the evaluation agent 60 are examples of AI agents that, when given a goal, have a language model generate tasks to achieve the goal, collect information to have the language model execute the generated tasks, and execute the tasks by inputting the collected information into the language model.

[0013] Specifically, the generation agent 50 inputs the original document D1, which is the source content for the RAG data, into the LLM and generates first data structured into a summary of the original document D1, tags representing the characteristics of the summary, and a collection of questions and answers (hereinafter sometimes referred to as QA) based on the original document D1, along with tags representing the characteristics of the QA.

[0014] Next, the evaluation agent 60 inputs the first data and the original document D1 into the LLM, scores the quality of the summary and tags, and the quality of the QA and tags, and generates an evaluation result in which improvements to each quality are added to the first data.

[0015] The generation agent 50 then inputs the original document D1 and the evaluation results into the LLM, regenerates the first data with the improvements included in the evaluation results, and outputs the regenerated first data as RAG data.

[0016] In this way, the information processing device 10 repeatedly performs the generation and evaluation of RAG data. Therefore, if the source document is a web page, the information processing device 10 can gradually remove links and advertisements from the RAG data generated on the first attempt, thereby suppressing the generation of erroneous QA and improving the quality of the RAG data.

[0017] Subsequently, if an LLM (Look-and-Make Machine) is used to generate answers to user questions, the LLM can use the RAG data generated in the above process in the RAG when it receives input such as questions from the user, thus improving the accuracy of the answers to the questions.

[0018] (Functional Configuration) Figure 2 is a functional block diagram showing the functional configuration of the information processing device 10 according to Embodiment 1. As shown in Figure 2, the information processing device 10 has a communication unit 11, an output unit 12, a storage unit 13, and a control unit 20.

[0019] The communication unit 11 is a processing unit that performs the transmission and reception of various types of data with other devices, and is implemented, for example, by a communication interface. For example, the communication unit 11 receives content that will become the source data for RAG data from a user terminal or the like.

[0020] The output unit 12 is a processing unit that outputs various types of data, and is implemented by, for example, a display or a touch panel. For example, the output unit 12 outputs RAG data in the process of generation, final RAG data, and evaluation results, which will be described later.

[0021] The storage unit 13 is a processing unit that stores various data and programs executed by the control unit 20, and is implemented by, for example, memory or a hard disk. This storage unit 13 stores the original document 14, RAG data 15, model DB (DataBase) 16, etc. In addition, the storage unit 13 also stores various data that are being generated in the process by the control unit 20, which will be described later.

[0022] Original document 14 is an example of content and is the data from which RAG data is generated. For example, original document 14 is data related to a document and can be in any format, such as text data, a web page, or video data containing a document.

[0023] RAG data 15 is RAG data generated from the original document 14 by the control unit 20, which will be described later. For example, RAG data 15 is data structured into summaries and tags, and QA and tags, and can be in any format, such as text data or web data. This RAG data 15 is used for the RAG of LLM.

[0024] Model DB16 is a database that stores various machine learning models. For example, Model DB16 stores the above-mentioned LLM. Note that Model DB16 may store one LLM, or it may store multiple LLMs prepared for each of the above-mentioned AI agents.

[0025] The control unit 20 is a processing unit that oversees the entire information processing device 10, and is realized, for example, by a processor or a combination of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). This control unit 20 includes a RAG data generation unit 21, a RAG data evaluation unit 22, and a RAG data output unit 23. The RAG data generation unit 21, the RAG data evaluation unit 22, and the RAG data output unit 23 are realized by electronic circuits of the processor or processes executed by the processor.

[0026] The RAG data generation unit 21 is a processing unit that generates first data showing a summary of the original document 14 and a collection of questions and answers by executing the generation agent 50 with the original document 14, which is an example of content, as input. Specifically, the RAG data generation unit 21 is a processing unit that executes the generation agent 50 and causes the generation agent 50 to perform the processing described later.

[0027] For example, the generation agent 50 inputs the source document 14 into the LLM. The generation agent 50 then uses the LLM to generate first data structured into a summary, tags representing the characteristics of the summary, and QA based on the content of the source document 14, along with tags representing the characteristics of the QA.

[0028] Furthermore, the generation agent 50 also updates (regenerates) the first data using the evaluation results described later. For example, the generation agent 50 inputs the original document 14 and the evaluation results into the LLM. Then, using the LLM, the generation agent 50 generates first data with improved summary and summary feature tags, and QA and QA feature tags, based on improvements to the quality of the summary and the tags representing the characteristics of the summary, and improvements to the quality of the QA and the tags representing the characteristics of the QA.

[0029] The RAG data evaluation unit 22 is a processing unit that generates an evaluation result for the first data by executing the evaluation agent 60 with the first data generated by the RAG data generation unit 21 as input. Specifically, the RAG data evaluation unit 22 is a processing unit that executes the evaluation agent 60 and causes the evaluation agent 60 to perform the processing described later.

[0030] For example, the evaluation agent 60 inputs the original document 14 and the first data into the LLM. The evaluation agent 60 then uses the LLM to score the first quality of the summary and the tags representing the characteristics of the summary, and the second quality of the QA and the tags representing the characteristics of the QA, and generates an evaluation result in which improvement points for each quality are added to the first data.

[0031] The RAG data output unit 23 is a processing unit that outputs RAG data by executing the generation agent 50 with the evaluation result of the first data as input. Specifically, the RAG data output unit 23 is a processing unit that executes the generation agent 50 and causes the generation agent 50 to perform the processing described later.

[0032] For example, the generation agent 50 updates the first data using the evaluation generated by the evaluation agent 60 and outputs the updated generated data as RAG data. For example, the generation agent 50 stores the updated generated data as RAG data 15 in the storage unit 13, outputs the updated generated data to a specified storage area, or transmits the updated generated data to a specified device.

[0033] (Specific Example) Next, we will explain an example of RAG data generation using Figures 3 to 6. Figure 3 is a diagram illustrating the generation of RAG data, Figure 4 is a diagram illustrating the evaluation of RAG data, Figure 5 is a diagram illustrating the evaluation results of RAG data, and Figure 6 is a diagram illustrating the regeneration of RAG data. Here, as an example, we will explain an example of generating RAG data from a web page containing the New Year's greeting of a company president.

[0034] First, using Figure 3, we will explain an example of generating the first data (RAG data) from the source document 14. As shown in Figure 3, the generation agent 50 inputs the source document 14 into the LLM and also provides (inputs) a prompt 51 to the LLM. Here, the prompt 51 is set to "generate a summary of the input data", "generate as much QA as possible from the input data", and "output structured RAG data using Markdown". As a result, the generation agent 50 generates RAG data A1 from the source document 14.

[0035] For example, as shown in Figure 3, RAG data A1 includes the "Title" of the RAG data, "Data" which is the date and time the RAG data was created, and "Contents" which is a structured representation of the original document's content.

[0036] The "Contents" section contains the "summary and tags" specified in prompt 51. Specifically, the "Contents" section includes "Summary," which is a summary of the original document 14, and "Tags," which represent the characteristics of the summary, such as "Mr. / Ms. XX, △△ Company, management, digital transformation (DX), ~".

[0037] Furthermore, the "Contents" section contains the "QA and tags derived from the original document 14" specified in prompt 51. Specifically, the "Contents" section includes, as the first "Question," "What specific role does '△△ CO.' play?", as the "Answer" to that "Question," "'△△ CO.' aims to provide concrete solutions for △△ Company to contribute to social issues through its business activities. ~~~", and as "Tags" representing the characteristics of the QA, "About △△ Company, Management, ~~~".

[0038] Furthermore, the "Contents" section includes a second "Question" which asks, "How do you view 'ethics' in the use of technology?", and the "Answer" to that "Question" which reads, "Mr. / Ms. XX emphasizes the importance of 'ethics' in the use of technology. ~~~". The "Tags" section, which represents the characteristics of the Q&A, includes "~~~, AI, technology, sustainability (SX), social issues".

[0039] The number of QAs included in "Contents" is dependent on the LLM's control, but since prompt 51 instructs "Generate as many QAs as possible," it is controlled so that one or more are generated from the original document 14.

[0040] Next, we will explain an example of generating evaluation results for the first data (RAG data A1) using Figures 4 and 5. As shown in Figure 4, the evaluation agent 60 inputs the original document 14 and the RAG data A1 generated in Figure 3 into the LLM, and also gives prompt 61 to the LLM. Here, prompt 61 is set to "refer to the original data, evaluate whether the summary and tags of the RAG data are appropriate, assign a score, and generate improvement points" and "add the evaluation score and improvement points to the RAG data".

[0041] As a result, the evaluation agent 60 generates RAG data A' by adding evaluation results to RAG data A1. Specifically, as shown in Figure 5, RAG data A' is data in which evaluation scores and areas for improvement are described in the "Contents" of RAG data A. For example, in "Summary," an evaluation score of "95" and the area for improvement "The examples and achievement goals lack specificity; adding concise details would make it more persuasive" are added. Also, in the first "Question," an evaluation score of "93" and the area for improvement "Adding specific success stories and numerical data would make it more persuasive" are added. Also, in the second "Question," an evaluation score of "92" and the area for improvement "Including specific measures and examples of AI ethics would increase clarity" are added.

[0042] Finally, an example of regenerating (updating) the first data (RAG data A1) using evaluation results will be described with reference to FIG. 6. As shown in FIG. 6, the generation agent 50 inputs the original document 14 and RAG data A1' including the evaluation result generated in FIG. 5 to the LLM, and provides a prompt 51 to the LLM. Here, in addition to "generate a summary of input data", "generate as much QA as possible from the input data", and "output structured RAG data in Markdown" included in the prompt 51, "regenerate RAG data based on the evaluation result" is set in the prompt 52. As a result, the generation agent 50 generates RAG data A2 from the original document 14.

[0043] As described above, the information processing apparatus 10 can generate high-quality RAG data by repeating the above-described processing until an end condition is satisfied and repeating the regeneration of RAG data. Note that the end condition can be arbitrarily set, for example, such that the processing is performed a predetermined number of times, the average of evaluation scores in RAG data is equal to or higher than a threshold, the lowest evaluation score in RAG data is equal to or higher than a threshold, the number of QAs in RAG data is equal to or higher than a threshold, or the like.

[0044] (Processing Flow) FIG. 7 is a flowchart showing the flow of RAG data generation processing according to the first embodiment. As shown in FIG. 7, the information processing apparatus 10 reads the original document 14 (S101), and executes the RAG data generation described in FIG. 3 (S102).

[0045] Subsequently, the information processing apparatus 10 reads the generated RAG data (S103), executes the evaluation of RAG data described in FIG. 4 and FIG. 5 (S104), and executes regeneration of RAG data using the evaluation result described in FIG. 6 (S105).

[0046] Then, when the processing from S103 to S106 has not been executed a predetermined number of times (S106: No), the information processing apparatus 10 executes the processing from S103 onward using the regenerated RAG data. On the other hand, when the processing from S103 to S106 has been executed a predetermined number of times (S106: Yes), the information processing apparatus 10 outputs the final RAG data (S107).

[0047] (Effect) As described above, the information processing apparatus 10 can generate structured RAG data including "summary and tags" and "QA and tags" from the original document 14, so that the extraction accuracy of relevant information for queries when an LLM uses RAG can be improved. Furthermore, the information processing apparatus 10 can achieve speedup of processing of applications using an LLM or the like. In addition, since the information processing apparatus 10 can generate high-quality RAG data, it eliminates the need for the LLM to execute RAG multiple times or prepare a plurality of pieces of RAG data, and can also reduce the processing load on a processor or the like and reduce memory usage.

[0048] Further, since the information processing apparatus 10 can generate and output a score indicating the quality of content and points for improvement, it can generate information that can be visually confirmed by a user. Therefore, when the score is low, the user can consider measures such as generating one piece of RAG data from a plurality of original documents. As described above, the information processing apparatus 10 can output information effective for development of applications using RAG and the like.

[0049] Further, since the information processing apparatus 10 repeatedly executes the generation, evaluation, and regeneration of RAG data, the summaries, tags, QAs, and tags of the RAG data are refined, and the quality of the finally generated RAG data can be improved.

[0050] Hitherto, embodiments of the present invention have been described, but the present invention may be implemented in various different forms other than the above-described embodiments.

[0051] (Numerical Values, Application Examples, etc.) The parameters, specific examples, numerical values, application examples and the like used in the above embodiments are merely examples and can be arbitrarily changed. The processing flow described in each flowchart can also be appropriately changed within a consistent range.

[0052] (Language Models) Language models are pre-trained models such as multimodal, large-scale, and small-scale language models. For example, a language model is a transformer-based model trained using a first token set generated from a token set in which some tokens are masked from multiple tokens.

[0053] Specifically, the language model is a transformer-based model trained using a token set generated from a token set in which some tokens are masked from a set of multiple tokens. For example, the information processing device 10 trains the language model using a token set (an unsupervised learning dataset). The language model consists of a deep neural network. For example, the language model is a machine learning model that incorporates an architecture called a transformer. In other words, the language model is a transformer-based language model. One such transformer is known to be BERT (Bidirectional Encoder Representation from Transformers). The information processing device 10 trains the language model by masking some of the tokens included in the token set and estimating the masked tokens.

[0054] (Application) An example of an application that utilizes RAG data is described below. For example, RAG data is applied to a conversation application between a user and an avatar that simulates a specific person. First, the information processing device 10 obtains information from a question posed to an avatar that simulates a specific person and generates information according to the input information. Next, the information processing device 10 searches the RAG data using the obtained question information and generates an answer to the question by inputting the retrieved RAG data into a language model. After that, the information processing device 10 outputs the generated answer to the display screen as the avatar's answer. Note that a specific person is a person with a specific role, for example, an expert in a specific field. Also, an avatar that simulates a person with a specific role is, for example, an avatar that simulates an expert in academics, business, sports, arts, culinary research, etc.

[0055] More specifically, the information processing device 10 displays an avatar representing the president of a certain company on the display screen. Next, the information processing device 10 receives questions from the user who is interacting with the avatar representing the president. Then, the information processing device 10 searches the RAG data using the received questions and generates answers to the questions by inputting prompts incorporating the search results into the language model. After that, the information processing device 10 has the avatar representing the president explain the generated answers.

[0056] Here, we will explain the data that forms the basis for generating RAG data. The content includes, for example, text data, web pages, and video data including documents that describe information about the person that the avatar is modeling. For example, the content may be articles that introduce the person's past statements, achievements, attributes, and affiliations.

[0057] More specifically, when the information processing device 10 sets the president of a certain company as the person the avatar will represent, it searches for web pages containing information about the president of that company and collects the content. The information processing device 10 generates RAG data from the content of the web page containing the New Year's greeting from the president of that company.

[0058] Furthermore, when the information processing device 10 sets the president of a certain company as the specific person that the avatar will represent, it displays the avatar of the president of that company on the display screen. Next, the information processing device 10 reads content related to the president of the company from among multiple contents stored in the memory unit and generates RAG data from the read content. Then, the information processing device 10 uses the RAG data to perform speech processing for the avatar. For example, the avatar makes a statement. This makes it possible to make the statements of the avatar that represents a specific person closer to the actual statements of that person.

[0059] (System) Unless otherwise specified, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings may be changed at will.

[0060] Furthermore, the specific forms of distribution and integration of the components of each device are not limited to those shown in the figures. For example, the RAG data generation unit 21, the RAG data evaluation unit 22, and the RAG data output unit 23 may be integrated. In other words, all or part of the components may be functionally or physically distributed or integrated in any unit depending on various loads and usage conditions. Moreover, all or any part of the processing functions of each device may be realized by a CPU and a program that is analyzed and executed by the CPU, or by hardware using wired logic.

[0061] Furthermore, each processing function performed by each device may be implemented, in whole or in part, by a CPU and a program executed by that CPU, or by wired logic hardware. Note that a portion of the program may be executed by a GPU (Graphics Processing Unit).

[0062] (Hardware) Figure 8 is a diagram illustrating an example of hardware configuration. As shown in Figure 8, the information processing device 10 includes a communication device 10a, an HDD (Hard Disk Drive) 10b, memory 10c, and a processor 10d. Furthermore, each of the parts shown in Figure 8 is interconnected by a bus or the like.

[0063] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and databases that operate the functions shown in Figure 2.

[0064] The processor 10d operates a process that performs the functions described in Figure 2 by reading a program that performs the same processing as each processing unit shown in Figure 2 from the HDD 10b or the like and loading it into memory 10c. This processor 10d may consist of one or more processors, or it may be a combination of a CPU and a GPU. For example, this process performs the same functions as each processing unit of the information processing device 10. Specifically, the processor 10d reads a program that has the same functions as the RAG data generation unit 21, the RAG data evaluation unit 22, the RAG data output unit 23, etc. from the HDD 10b or the like. Then, the processor 10d executes a process that performs the same processing as the RAG data generation unit 21, the RAG data evaluation unit 22, the RAG data output unit 23, etc.

[0065] Thus, the information processing device 10 operates as an information processing device that executes an information processing method by reading and executing a program. Furthermore, the information processing device 10 can also achieve the same functionality as in the above-described embodiment by reading the program from a recording medium using a media reader and executing the read program. Note that the program referred to in this other embodiment is not limited to being executed by the information processing device 10. For example, the above embodiment may also be applied similarly when another computer or server executes the program, or when they collaborate to execute the program.

[0066] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, flexible disk (FD), CD-ROM, MO (Magneto-Optical disk), or DVD (Digital Versatile Disc), and executed by being read from the recording medium by a computer.

[0067] 10 Information Processing Device 11 Communication Unit 12 Output Unit 13 Storage Unit 14 Original Document 15 RAG Data 16 Model DB 20 Control Unit 21 RAG Data Generation Unit 22 RAG Data Evaluation Unit 23 RAG Data Output Unit

Claims

1. A data generation program that causes a computer to perform the following processes: detect a first agent that generates data and a second agent that evaluates the generated data; generate first data showing a summary of the content and a collection of questions and answers by executing the first agent with content as input; generate evaluation results of the first data by executing the second agent with the first data as input; and output RAG (Retrieval Augmented Generation) data by executing the first agent with the evaluation results of the first data as input.

2. The data generation program according to claim 1, wherein the process for generating the first data includes inputting the content into a language model, and using the language model to generate the first data structured into a summary, tags representing the characteristics of the summary, and a set of questions and answers based on the content, and tags representing the characteristics of the set of questions and answers.

3. The data generation program according to claim 2, wherein the process for generating the evaluation result inputs the content and the first data into a language model, uses the language model to score the first quality of the summary and the tags representing the characteristics of the summary, and the second quality of the question and answer set and the tags representing the characteristics of the question and answer set, and generates the evaluation result by adding points for improvement to the first data.

4. The data generation program according to claim 3, wherein the process for generating the first data includes inputting the content and the evaluation results into a language model, and using the language model to generate the first data in which the summary and tags representing the characteristics of the summary, and the question and answer set and tags representing the characteristics of the question and answer set have been improved based on the improvements.

5. A data generation program according to claim 1, wherein the program provides an avatar that simulates a person in a specific role, and which generates information in response to input information. The program obtains information from questions posed to the avatar, uses the obtained information to search the RAG data, inputs a prompt incorporating the retrieved RAG data into a language model to generate an answer to the question, and causes the computer to output the generated answer as the avatar's answer, wherein the content is collected data from any of the following: text data describing information about a person in a specific role simulated by the avatar, a web page describing information about the person, or video data including a document describing information about the person.

6. The data generation program according to claim 5, wherein, when a first person is set as a person who plays a specific role that the avatar simulates, the program displays the avatar of the first person on the display screen, loads content containing information about the first person, inputs the loaded content to the first agent to generate the RAG data from the content, and uses the generated RAG data to cause the computer to perform speech processing for the avatar of the first person.

7. The data generation program according to claim 5, wherein the first agent and the second agent are AI agents that, when given a goal, cause a language model to generate tasks to achieve the goal and execute the generated tasks.

8. A data generation method that performs the following processes: a computer detects a first agent that generates data and a second agent that evaluates the generated data; it runs the first agent with content as input to generate first data showing a summary of the content and a collection of questions and answers; it runs the second agent with the first data as input to generate an evaluation result of the first data; and it runs the first agent with the evaluation result of the first data as input to output RAG (Retrieval Augmented Generation) data.

9. An information processing device having a control unit that detects a first agent that generates data and a second agent that evaluates the generated data; generates first data showing a summary of the content and a collection of questions and answers by executing the first agent with content as input; generates an evaluation result of the first data by executing the second agent with the first data as input; and outputs RAG (Retrieval Augmented Generation) data by executing the first agent with the evaluation result of the first data as input.