Information processing device, information processing method, and information processing program
Patent Information
- Application Number
- PCT/JP2025/013026
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025013026_01102026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and information processing program
[0001] This invention relates to an information processing device, an information processing method, and an information processing program.
[0002] Traditionally, multi-agent debate (MAD) systems have been known. According to MAD, multiple agents can discuss a query and generate an answer. For example, agents are large language models (LLMs) that are assigned roles.
[0003] Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. Encouraging divergent thinking in large language models through multi-agent debate. In Empirical Methods in Natural Language Processing, 2024.
[0004] However, conventional techniques have a problem in that it is difficult to find an appropriate set of agents for multi-agent debate.
[0005] For example, Non-Patent Document 1 describes how two agents are selected by a human. In multi-agent debate, accuracy improves with a larger number of agents. On the other hand, selecting many agents while considering computational resources is difficult for humans because it complicates the problem.
[0006] To solve the aforementioned problems, the information processing device of the present invention is characterized by comprising: an acquisition unit that acquires a plurality of queries; and a selection unit that selects a second set from a first set of agents that generate answers to queries such that the combination of answers generated by each of the included agents for each of the plurality of queries is optimized.
[0007] According to the present invention, it is possible to determine an appropriate set of agents for a multi-agent debate.
[0008] Figure 1 shows an example configuration of an information processing device according to the first embodiment. Figure 2 is a diagram illustrating the selection unit of the information processing device. Figure 3 is a diagram illustrating variables and sets. Figure 4 is a diagram illustrating an example algorithm for selecting an agent. Figure 5 is a flowchart showing the processing flow of the information processing device. Figure 6 is a flowchart showing the processing flow for selecting an agent. Figure 7 shows an example configuration of a computer that executes an information processing program.
[0009] The embodiments for carrying out the present invention will be described below with reference to the drawings. The present invention is not limited to these embodiments.
[0010] [First Embodiment] [Configuration of the First Embodiment] Figure 1 is a diagram showing an example configuration of an information processing device according to the first embodiment. As shown in Figure 1, the information processing device 10 receives a query input and outputs a response to the query.
[0011] The information processing device 10 includes an acquisition unit 111, a selection unit 112, and a debate unit 113. The acquisition unit 111, the selection unit 112, and the debate unit 113 are realized by the control unit executing a program. For example, the control unit is an electronic circuit such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), or GPU (Graphics Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).
[0012] The information processing device 10 stores the QADB 121 and the agent DB 122. The information processing device 10 stores information using storage devices such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or optical disc.
[0013] QADB121 stores query-answer pairs (combinations). Both queries and answers may be natural language text. Queries and answers may also be pre-prepared manually.
[0014] Agent DB122 contains information for building agents. For example, Agent DB122 may be a prompt for assigning roles to the LLM for each agent ID.
[0015] For example, roles may represent occupation, social standing, and areas of knowledge, such as economic analyst, university student, or company manager. In a debate, LLMs engage in discussion while playing their assigned roles. Each agent may also be given prior knowledge. Furthermore, agents can generate answers to input queries.
[0016] Furthermore, Agent DB122 stores the parameters of LLM. By combining the parameters stored in Agent DB122 with prompts, the agent is constructed.
[0017] As shown in Figure 1, the acquisition unit 111 passes the set of query-answer pairs acquired from the QADB 121 to the selection unit 112. The acquisition unit 111 may also extract a set of pairs from the QADB 121 that includes queries similar to the queries input to the information processing device 10, and pass the extracted set of pairs to the selection unit 112.
[0018] Queries in QADB 121 that are similar to the input query may be specified in advance by a human. Alternatively, the acquisition unit 111 may calculate the similarity between the input query and each query in QADB 121, and extract a set of pairs based on the calculated similarity.
[0019] The selection unit 112 uses the set of pairs received from the acquisition unit 111 to select a set of agents from the set of agents stored in the agent DB 122. The selection unit 112 selects a set of agents with a predetermined number of agents or less as elements that improves the accuracy of the answers to queries obtained through debate.
[0020] The debate unit 113 has the agents in the set selected by the selection unit 112 conduct a debate and obtains answers. For example, the debate unit 113 can have multiple agents conduct a debate using the method described in Non-Patent Literature 1.
[0021] The processing of the selection unit 112 will be explained in detail using Figure 2. Figure 2 is a diagram illustrating the selection unit of the information processing device.
[0022] As shown in Figure 2, the selection unit 112 includes a generation unit 112a, a similarity calculation unit 112b, an objective function calculation unit 112c, and an optimization unit 112d.
[0023] Here, using Figure 3, we will explain the variables and sets related to the process by which the selection unit 112 selects an agent. Figure 3 is a diagram illustrating the variables and sets.
[0024] The selection unit 112 is the set of agents M = {m 1 ,m 2 ,…,m N A set of k agents X is selected from}. The number of agents selected by the selection unit 112 may be less than or equal to k, but shall not exceed k. k may be determined manually in advance. Set M is an example of the first set. Set X is an example of the second set.
[0025] The selection unit 112 preferably selects an agent that generates the correct answer to the specified query q, i.e., the query input to the information processing device 10. Therefore, the acquisition unit 111 selects the dataset D = {(q)} of QADB 121. i , a i )} iFrom this, query-answer pairs (q', a') similar to q are obtained. Note that the acquisition of pairs by the acquisition unit 111 may be performed manually. The dataset D of the acquired pairs q This is a subset of dataset D.
[0026] Queries in QADB 121 that are similar to the input query may be specified in advance by a human. Alternatively, the acquisition unit 111 may calculate the similarity between the input query and each query in QADB 121, and extract a set of pairs based on the calculated similarity.
[0027] As shown in Figure 3, l(m, t) is the output text when agent m receives the instruction t. In other words, l(m, t) is the text of the response generated by agent m based on query t.
[0028] The processing of the selection unit 112 is defined as a combinatorial optimization problem in which a set of k agents X that generate the correct answer for query q is selected from a set of agents M, as shown in equation (1).
[0029]
[0030] d(・,・) is the distance between two sentences (texts). -d(・,・) is the similarity between the sentences. For example, the distance between sentences can be calculated using Sentence-BERT (see Reference 1).
[0031] Reference 1: Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 2019.
[0032] The selection unit 112 can optimize equation (1) using an existing combinatorial optimization method. Alternatively, the selection unit 112 may optimize equation (1) using a solver.
[0033] As described above, the selection unit 112 selects, from the agent set M, a set X such that the combination of answers generated by each included agent for each of a plurality of queries is optimized.
[0034] Specifically, as shown in equation (1), the selection unit 112 calculates the sum of similarity degrees between answers generated by each included agent for each of a plurality of queries and answers corresponding to each of the plurality of queries. Then, the selection unit 112 selects the set X that maximizes the similarity degree.
[0035] For example, the number of query and answer combinations for which the similarity degree is calculated in equation (1) is the number of queries (D q size) multiplied by the number of agents (size of X).
[0036] FIG. 4 shows an algorithm when the selection unit 112 optimizes equation (1) using a greedy method. FIG. 4 is a diagram illustrating an example of an algorithm for selecting agents.
[0037] As shown in FIG. 4, the selection unit 112 first initializes the set X 0 as an empty set (line 1). The selection unit 112 repeatedly executes the processes of the third line and the fourth line until k agents are selected, that is, for i = 1...k (lines 2 and 5).
[0038] The processing in the i-th iteration will be described. In the third line, the generation unit 112a inputs each query included in the data set D to each agent not included in the set X among the agents included in the set M, and obtains an answer. Let Y i-1 to each query included in the dataset D q , and obtain answers. Among the agents included in the set M, agents not included in the set X i-1 form a set, which is denoted as Y i .
[0039] The objective function calculation unit 112c calculates the set Y i For each of the agents m included in the set, the objective function S q (X i-1 ∪{m i})-S q (X i-1 The optimization unit 112d calculates the following: The optimization unit 112d finds m such that the objective function is optimized (maximized in the example in Figure 4), and i as X i-1 Add this to the fourth line.
[0040] In this way, the selection unit 112 can add a predetermined number (for example, k) of agents from set M to set X, in order of the amount of increase in the degree of optimization of the answer combinations when added to set X. The objective function in the third row of Figure 4 represents the amount of increase in the degree of optimization.
[0041] [Processing of the First Embodiment] The processing flow of the information processing device 10 will be explained using Figure 5. Figure 5 is a flowchart showing the processing flow of the information processing device.
[0042] First, as shown in Figure 5, the acquisition unit 111 accepts the specification of a query (step S101). The acquisition unit 111 then retrieves a set of pairs of queries that are similar to the specified query from a dataset whose elements are query-answer pairs (step S102). For example, the acquisition unit 111 retrieves a set of pairs of queries that have a similarity to the specified query that is above a threshold.
[0043] Next, the selection unit 112 uses the acquired set of pairs to select a set of agents X from the set of agents M (step S103). The detailed processing flow of step S103 will be explained later.
[0044] Next, the debate unit 113 has the agents of set X conduct a debate and obtains answers (step S104). Then, the debate unit 113 outputs the answers (step S105).
[0045] Figure 6 will be used to explain the details of step S103 in Figure 5. Figure 6 is a flowchart showing the process flow for selecting an agent.
[0046] As shown in Figure 6, first, the selection unit 112 determines whether or not there are any unselected pairs among the acquired pairs (step S201).
[0047] If there are no unselected pairs (step S201; No), the selection unit 112 outputs set X (step S207).
[0048] If there are unselected pairs (step S201; Yes), the selection unit 112 selects an unselected pair (step S202).
[0049] Here, the selection unit 112 inputs the queries of the selected pair to the agent and generates an answer (step S203). The selection unit 112 calculates the similarity between the answers of the selected pair and the generated answers (step S204).
[0050] The selection unit 112 calculates an objective function based on the similarity (step S205). Then, the selection unit 112 adds agents to the set X so that the objective function is optimized (step S206).
[0051] After that, the selection unit 112 returns to step S201 and repeats the process until there are no more unselected pairs.
[0052] [Effects of the First Embodiment] As described above, the acquisition unit 111 acquires multiple queries. The selection unit 112 selects a second set from a first set of agents that generate answers to queries such that the combination of answers generated by each of the included agents for each of the multiple queries is optimized.
[0053] In this way, the information processing device 10 can automatically select a set of responding agents according to each agent's response to the query. As a result, the information processing device 10 can find an appropriate set of agents for a multi-agent debate.
[0054] The debate unit 113 has the agents of the second set conduct a debate and obtain answers to the specified queries. As a result, the information processing device 10 can obtain accurate answers to the multi-agent debate.
[0055] The selection unit 112 selects a second set such that the sum of the similarities between the answers each included agent generates for each of the multiple queries and the answers associated with each of the multiple queries is maximized. For example, the selection unit 112 adds a predetermined number of agents from the first set to the second set, in order of the increase in the degree of optimization of the answer combinations when added to the second set.
[0056] As a result, the information processing device 10 can use an algorithm to solve the selection of an appropriate set of agents for a multi-agent debate as an optimization problem.
[0057] [System Configuration, etc.] Furthermore, the components of each part shown in the diagram are functional concepts and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown in the diagram, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. In addition, all or any part of the processing functions performed by each device can be realized by a CPU and the program executed on that CPU, or by hardware using wired logic.
[0058] Furthermore, among the processes described in the embodiments described above, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified.
[0059] [Program] The information processing device 10 described above can be implemented by installing a program (information processing program) as packaged software or online software on a desired computer. For example, by having the information processing device execute the above program, the information processing device can be made to function as the information processing device 10. The information processing device referred to here includes mobile communication terminals such as smartphones, mobile phones and PHS (Personal Handyphone System), and terminals such as PDA (Personal Digital Assistant).
[0060] Figure 7 shows an example configuration of a computer that executes an information processing program. Computer 1000 has, for example, memory 1010 and CPU 1020. Computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0061] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0062] The hard disk drive 1090 stores, for example, the OS 1091, application programs 1092, program modules 1093, and program data 1094. That is, the programs that define each process executed by the information processing device 10 are implemented as program modules 1093 in which executable code for a computer is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to the functional configuration of the information processing device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0063] Furthermore, the data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.
[0064] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.
[0065] The following additional information is disclosed regarding the embodiments described above.
[0066] (Note 1) An information processing device comprising: memory and at least one processor connected to the memory, wherein the processor selects a second set from a first set of agents that acquire a plurality of queries and generate answers to the queries, such that the combination of answers generated by each of the constituent agents for each of the plurality of queries is optimized. (Note 2) An information processing device according to Note 1, wherein the processor acquires a plurality of queries whose similarity to a specified query is greater than or equal to a threshold. (Note 3) An information processing device according to Note 2, wherein the processor causes the agents of the second set to conduct a debate and obtain answers to the specified queries. (Note 4) An information processing device according to Note 1, wherein the processor selects a second set such that the sum of the similarities between the answers generated by each of the constituent agents for each of the plurality of queries and the answers associated with each of the plurality of queries is maximized. (Appendix 5) An information processing device as described in Appendix 1, wherein the processor adds a predetermined number of agents from the first set to the second set in order of the magnitude of the increase in the degree of optimization of the answer combination when added to the second set. (Appendix 6) A non-temporary storage medium storing a program executable by a computer, wherein the program causes the computer to execute a process of selecting a second set from a first set of agents that acquire a plurality of queries and generate answers to the queries, such that the combination of answers generated by each of the constituent agents for each of the plurality of queries is optimized.
[0067] 10 Information Processing Device 111 Acquisition Unit 112 Selection Unit 112a Generation Unit 112b Similarity Calculation Unit 112c Objective Function Calculation Unit 112d Optimization Unit 113 Debate Unit 121 QADB 122 Agent DB
Claims
1. An information processing device comprising: an acquisition unit that acquires multiple queries; and a selection unit that selects a second set from a first set of agents that generate answers to queries such that the combination of answers generated by each of the included agents for each of the multiple queries is optimized.
2. The information processing apparatus according to claim 1, characterized in that the acquisition unit acquires a plurality of queries whose similarity to a specified query is equal to or greater than a threshold.
3. The information processing apparatus according to claim 2, further comprising a debate unit that causes agents of the second set to conduct a debate and obtains answers to the specified queries.
4. The information processing apparatus according to claim 1, characterized in that the selection unit selects the second set such that the sum of the similarities between the answers generated by each of the included agents for each of the plurality of queries and the answers associated with each of the plurality of queries is maximized.
5. The information processing apparatus according to claim 1, characterized in that the selection unit adds a predetermined number of agents from the first set to the second set in order of the amount of increase in the degree of optimization of the combination of answers when added to the second set.
6. An information processing method performed by a computer, comprising: an acquisition step of acquiring a plurality of queries; and a selection step of selecting a second set from a first set of agents that generate answers to queries such that the combination of answers generated by each of the included agents for each of the plurality of queries is optimized.
7. An information processing program characterized by causing a computer to perform the following steps: an acquisition step of acquiring multiple queries; and a selection step of selecting a second set from a first set of agents that generate answers to queries such that the combination of answers generated by each of the included agents for each of the multiple queries is optimized.