Prompt adjustment program, prompt adjustment method, and information processing device

The prompt adjustment program addresses the computational expense of conventional prompt generation by determining adjustment items for language models based on combinations of configuration information, resulting in optimized prompts with reduced calculation costs.

WO2025126517A1PCT designated stage expired Publication Date: 2025-06-19FUJITSU LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/021509
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-06-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Conventional prompt generation methods for language models are computationally expensive due to the high calculation cost associated with large vocabularies, making it challenging to optimize prompts efficiently.

Method used

A prompt adjustment program that generates multiple sample prompts based on various combinations of configuration information representing the prompt as a combination of blocks and their possible values, and determines adjustment items by identifying the items included in the combination with the highest inference accuracy.

Benefits of technology

Enables the optimization of prompts with a significantly reduced calculation cost, allowing for efficient generation of high-precision prompts for language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024021509_19062025_PF_FP_ABST
    Figure JP2024021509_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention generates a plurality of sample prompts that correspond to a plurality of types of combinations on the basis of structural information that expresses the structure of a prompt as a combination of a plurality of items that correspond to a plurality of blocks that are to constitute the prompt and values that can be assumed by the items and determines the items included in the combination used to generate the sample prompt that has the highest inference accuracy of results obtained by inputting the plurality of sample prompts into a language model to be adjustment items for adjusting the prompt. The present invention can thereby optimize the prompt at a low calculation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Prompt adjustment program, prompt adjustment method, and information processing device

[0001] The present invention relates to a prompt adjustment program, a prompt adjustment method, and an information processing device.

[0002] In recent years, language generation AI (Artificial Intelligence) has been attracting attention, and language generation AI is being used to perform tasks such as QA (Question and Answer) and information recommendation. In order to improve the accuracy of individual tasks in language generation AI, fine-tuning of pre-trained models has conventionally been performed, but fine-tuning requires enormous costs (computation time and computational resources).

[0003] Furthermore, changing the sentences (prompts) input to the pre-trained model can improve the accuracy of the output of the pre-trained model. For this reason, for example, a technique (prompting technique) for generating sentences to be input as prompts to large language models (LLMs) has been attracting attention (see, for example, Patent Literature 1). Conventionally, a technique for optimizing short prompts of a few words has also been known (see, for example, Non-Patent Literature 1).

[0004] Patent No. 7325152 US Patent Application Publication No. 2023 / 0316001 JP 2023-73220 A US Patent Application Publication No. 2023 / 0316006

[0005] Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P. Xing, Zhiting Hu, “RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning”, [online], May 25, 2022, CVPR2021, [Retrieved November 29, 2023], Internet <URL: / arxiv.org / abs / 2205.12548>

[0006] However, in such conventional prompt generation methods, the computational cost is calculated as the length of the prompt raised to the power of the number of vocabulary elements, so the computational cost becomes enormous when attempting to generate prompts with a large (long) vocabulary element.

[0007] In one aspect, the present invention aims to enable prompt optimization to be achieved with low computational cost.

[0008] Therefore, this prompt adjustment program causes a computer to execute a process of generating a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that make up the prompt and the possible values ​​that each item can take, and determining, as an adjustment item for adjusting the prompt, an item included in the combination used to generate the sample prompt that has the highest inference accuracy in the results obtained by inputting the plurality of sample prompts into a language model.

[0009] According to one embodiment, prompt optimization can be achieved at low computational cost.

[0010] 10 is a block diagram illustrating an example of the hardware (HW) configuration of a computer that realizes the functions of a prompt adjustment device as an example of the first embodiment. FIG. 11 is a diagram illustrating an example of the functional configuration of a prompt adjustment device as an example of the first embodiment. FIG. 12 is a diagram illustrating an example of a prompt configuration. FIG. 13 is a diagram illustrating an example of formulation information in a prompt adjustment device as an example of the first embodiment. FIG. 14 is a diagram illustrating an example of formulation information and an example of all combinations of possible attribute values ​​in a prompt adjustment device as an example of the first embodiment. FIG. 15 is a diagram illustrating an example of management information created by an optimization unit of a prompt adjustment device according to the first embodiment. FIG. 16 is a diagram illustrating an example of prompt design information generated by an optimization unit of a prompt adjustment device according to an example of the first embodiment. FIG. 17 is a flowchart illustrating the processing of the optimization unit of a prompt adjustment device according to the first embodiment. FIG. 18 is a flowchart illustrating the details of the processing of step S4 in the flowchart shown in FIG. 8. FIG. 19 is a diagram illustrating an example of the functional configuration of a prompt adjustment device as an example of the second embodiment. FIG. 19 is a diagram illustrating an example of data mining input information in a prompt adjustment device according to the second embodiment. FIG. 19 is a diagram illustrating an example of interactions generated by an interaction extraction unit of a prompt adjustment device according to the second embodiment. FIG. 19 is a diagram illustrating an example of a limited search range. 18 is a flowchart for explaining the processing of an optimization unit of the prompt adjustment device according to the second embodiment. FIG. 19 is a flowchart for explaining the processing of a pre-processing unit of the prompt adjustment device according to the second embodiment. FIG. 20 is a flowchart for explaining the processing of a post-processing unit of the prompt adjustment device according to the second embodiment. FIG. 21 is a flowchart for explaining the processing of an optimization processing unit of the prompt adjustment device according to the second embodiment. FIG. 22 is a flowchart for explaining a modification of the flowchart shown in FIG. 17. FIG. 23 is a flowchart for explaining a modification of the flowchart shown in FIG. 17. FIG. 24 is a diagram showing a modification of formulation information. FIG. 25 is a diagram illustrating the functional configuration of a prompt adjustment device as an example of a third embodiment. FIG. 26 is a diagram for explaining a method of creating training data by a history data mining processing unit of a prompt adjustment device as an example of the third embodiment. FIG. 27 is a diagram illustrating training data in a prompt adjustment device as an example of the third embodiment.10 is a diagram illustrating output information in a prompt adjustment device as an example of a third embodiment. FIG. 11 is a flowchart for explaining processing in a training phase of a history data mining processing unit in a prompt adjustment device according to a third embodiment. FIG. 12 is a flowchart for explaining processing in a prediction phase of a prompt adjustment device according to a third embodiment.

[0011] Hereinafter, embodiments of the prompt adjustment program, prompt adjustment method, and information processing device will be described with reference to the drawings. However, the embodiments shown below are merely examples, and are not intended to exclude various modifications or applications of techniques not explicitly stated in the embodiments. In other words, each embodiment can be implemented with various modifications (e.g., combinations of the embodiments and modifications) within the scope of its intent. Furthermore, each figure does not intend to include only the components shown in the figure, but may include other functions, etc.

[0012] (I) Description of the First Embodiment (A) Configuration (A-1) Hardware Configuration Fig. 1 is a block diagram showing an example of the hardware (HW) configuration of a computer 10 that realizes the functions of a prompt adjustment device 1a as an example of the first embodiment. When multiple computers are used as HW resources that realize the functions of the prompt adjustment device 1a, each computer may have the HW configuration exemplified in Fig. 1.

[0013] The prompt adjustment device 1a is an information processing device that adjusts prompts to be input to a language model so that the output from the language model is highly accurate. The language model may be a large-scale language model (LLM).

[0014] As shown in FIG. 1, the computer 10 may include, as its HW configuration, a processor 10a, a graphics processing unit 10b, a memory 10c, a storage unit 10d, an IF (Interface) unit 10e, an IO (Input / Output) unit 10f, and a reading unit 10g, for example.

[0015] The processor 10a is an arithmetic processing unit that performs various controls and calculations, and is an example of a control unit. The processor 10a may be connected to each block in the computer 10 via a bus 10j so that they can communicate with each other. The processor 10a may be a multiprocessor including multiple processors, a multi-core processor having multiple processor cores, or a configuration having multiple multi-core processors.

[0016] Examples of the processor 10a include integrated circuits (ICs) such as a CPU, MPU, APU, DSP, ASIC, and FPGA. Note that the processor 10a may be a combination of two or more of these integrated circuits. CPU is an abbreviation for Central Processing Unit, MPU is an abbreviation for Micro Processing Unit, APU is an abbreviation for Accelerated Processing Unit, DSP is an abbreviation for Digital Signal Processor, ASIC is an abbreviation for Application Specific IC, and FPGA is an abbreviation for Field-Programmable Gate Array.

[0017] The graphics processing device 10b controls screen display for an output device such as a monitor in the IO unit 10f. The graphics processing device 10b may also be configured as an accelerator that executes machine learning processing and inference processing using a machine learning model. The graphics processing device 10b may be any of a variety of arithmetic processing devices, such as a graphics processing unit (GPU), an APU, a DSP, an ASIC, an FPGA, or other integrated circuits (ICs).

[0018] The memory 10c is an example of HW that stores various types of data, programs, and other information. The memory 10c may be, for example, a volatile memory such as a dynamic random access memory (DRAM) or a non-volatile memory such as a persistent memory (PM), or both.

[0019] The storage unit 10d is an example of HW that stores various types of data, programs, and other information. Examples of the storage unit 10d include various storage devices such as a magnetic disk device such as a hard disk drive (HDD), a semiconductor drive device such as a solid state drive (SSD), and a nonvolatile memory. Examples of nonvolatile memory include a flash memory, a storage class memory (SCM), and a read-only memory (ROM).

[0020] The storage unit 10d may store a program 10h (prompt adjustment program) that realizes all or part of the various functions of the computer 10.

[0021] For example, the processor 10a of the prompt adjustment device 1a can implement a prompt adjustment function, which will be described later, by loading a program 10h stored in the storage unit 10d into the memory 10c and executing it.

[0022] The IF unit 10e is an example of a communication IF that controls the connection and communication between the computer 10 and other computers. For example, the IF unit 10e may include an adapter that complies with a LAN (Local Area Network) such as Ethernet (registered trademark) or optical communication such as FC (Fibre Channel). The adapter may support either or both wireless and wired communication methods.

[0023] For example, the prompt adjustment device 1a may be connected to other information processing devices, databases, etc. (not shown) via the IF unit 10e and a network so that they can communicate with each other. The program 10h may be downloaded to the computer 10 from the network via the communication IF and stored in the storage unit 10d.

[0024] The IO unit 10f may include one or both of an input device and an output device. Examples of input devices include a keyboard, a mouse, and a touch panel. Examples of output devices include a monitor, a projector, and a printer. The IO unit 10f may also include a touch panel that combines an input device and a display device. The output device may be connected to the graphics processing device 10b.

[0025] The reading unit 10g is an example of a reader that reads data and program information recorded on the recording medium 10i. The reading unit 10g may include a connection terminal or device to which the recording medium 10i can be connected or inserted. Examples of the reading unit 10g include an adapter compliant with USB (Universal Serial Bus) or the like, a drive device that accesses a recording disk, and a card reader that accesses a flash memory such as an SD card. Note that the recording medium 10i may store the program 10h, and the reading unit 10g may read the program 10h from the recording medium 10i and store it in the memory unit 10d.

[0026] Examples of the recording medium 10i include non-transitory computer-readable recording media such as magnetic / optical disks and flash memories. Examples of magnetic / optical disks include flexible disks, CDs (Compact Discs), DVDs (Digital Versatile Discs), Blu-ray Discs, and HVDs (Holographic Versatile Discs). Examples of flash memories include semiconductor memories such as USB memories and SD cards.

[0027] The above-described HW configuration of the computer 10 is an example. Therefore, the HW in the computer 10 may be increased or decreased (for example, adding or deleting any block), divided, integrated in any combination, or the HW may be added or deleted as needed.

[0028] (A-2) Example of Functional Configuration FIG. 2 is a diagram illustrating an example of the functional configuration of a prompt adjustment device 1a as an example of the first embodiment.

[0029] 2, the prompt adjustment device 1a may illustratively include functions as an input processing unit 11, an optimization unit 12a, and an output processing unit 17. The functions of the input processing unit 11, the optimization unit 12a, and the output processing unit 17 may be realized, for example, by the above-mentioned processor 10a executing a program 10h (prompt adjustment program).

[0030] The input processing unit 11 receives input of multiple (multiple types of) attributes included in the prompt and possible values ​​for each attribute by, for example, a user via a keyboard, mouse, network, or the like (not shown).

[0031] Input data is also input to the input processing unit 11. The input data includes a question (question sentence) and its correct answer. The input data including the question and the correct answer may be called sample data with correct answer. The input data may also be input by a user or the like via a keyboard, mouse, network, or the like.

[0032] The input data is changed as appropriate depending on the task. Here, for example, a movie recommendation task is taken as an example, in which a user inputs a list of movies they have watched and a list of movies they would like to recommend next, and a movie that the user is likely to watch next is selected from the list of movies and output.

[0033] The "questions" included in the input data correspond to the "list of movies the user has watched" and the "list of multiple movies the user would like to recommend next," which are input as described above. The "correct answers" correspond to the "correct movie candidates." Below are specific examples of questions and their correct answers.

[0034] Example of a question: History list "AAA, BBB, CCC" (AAA etc. are names of movies), Recommended candidate list "DDD, EEE, FFF, GGG, HHH" (DDD etc. are names of movies) Example of a correct answer: GGG (GGG is the name of a movie) The input processing unit 11 passes each piece of input information to the optimization unit 12a. The input processing unit 11 may, for example, store each piece of input information in a predetermined storage area of ​​the memory 10c or the storage unit 10d. The optimization unit 12a may obtain the information by reading out the information stored in these storage areas.

[0035] The optimization unit 12a generates information for adjusting (creating) a prompt based on the input data. Information for adjusting or creating a prompt can be referred to as prompt design information. By inputting a prompt created based on this prompt design information into the LLM, a highly accurate answer can be obtained that improves task performance using the LLM (e.g., question answering or information recommendation). In other words, the optimization unit 12a generates prompt design information for creating an optimized prompt.

[0036] As shown in FIG. 2, the optimization unit 12a functions as a formulation processing unit 13.

[0037] The formulation processor 13 formulates the configuration of the prompt. The formulation processor 13 formulates the prompt by combining multiple attributes included in the prompt and the values ​​that each attribute can take. Formulation can also be called parameterization or formatting.

[0038] The formulation processor 13 divides the prompt into a plurality of blocks. Each block is assigned a function in the prompt. A block is an element that defines the prompt. A block can also be called an attribute.

[0039] The formulation processing unit 13 may, for example, treat each of the multiple sentences included in the prompt as a block, or may determine the block based on specific words included in the sentence, and can be implemented with appropriate modifications.

[0040] FIG. 3 is a diagram illustrating the configuration of a prompt.

[0041] In FIG. 3, symbol A indicates an example of a prompt, and symbol B indicates a block configuration corresponding to the prompt indicated by symbol A.

[0042] FIG. 3 illustrates an example prompt for an information recommendation task that allows the LLM to make movie recommendations.

[0043] A prompt is composed of a combination of multiple blocks. The prompt shown in Figure 3 has a task block, an input format block, an output format block, an output note block, an example block, and a question block. The task block, input format block, output format block, output note block, example block, and question block are each examples of blocked items.

[0044] These task blocks, input format blocks, output format blocks, output note blocks, example blocks, and question blocks correspond to prompts P1, P2, P3, P4, P5, and P6, respectively, indicated by the symbol A in Figure 3. The input contents for P5 and P6 correspond to the "questions" in the input data described above. In particular, the input content for P6 is a question to the LLM, and the LLM generates an answer to this question.

[0045] As mentioned above, each block is assigned a function in the prompt: for example, the task block indicates the task, and the question block poses the actual question to the LLM.

[0046] In the prompt indicated by symbol A in Figure 3, the task block is at the beginning and the question block is at the end, but the accuracy of the LLM's response is likely to change if the task block is placed second, third, etc., at the end, or if the question block is placed first, second, third, etc. In particular, the attribute related to the order of "which block is placed in what order" in the prompt is considered to be important in terms of performance.

[0047] Hereinafter, the low / high accuracy of the results obtained by inputting a prompt into the LLM may be simply referred to as low / high accuracy.

[0048] The formulation processor 13 formulates a prompt by treating the elements that define the prompt as attributes and setting possible values ​​for each of the multiple attributes. This makes it possible to treat the prompt as a problem in which the optimal combination is automatically determined. Hereinafter, the information that associates each attribute with the possible values ​​for each attribute, created by the formulation processor 13 when formulating the prompt, may be referred to as formulation information.

[0049] The formulation information represents attribute settings associated with formulating a prompt. The formulation processing unit 13 may create the formulation information based on the input sample prompt. The sample prompt may be a prompt generated by applying a combination of parameters to each piece of sample data (for example, a history list including specific movie names and a recommendation candidate list) and outputting it.

[0050] FIG. 4 is a diagram showing an example of formulation information in the prompt adjustment device 1a as an example of the first embodiment.

[0051] In the formulation information illustrated in FIG. 4, for example, the order from the beginning of the prompt (1st to 6th) is used as an attribute, and possible values ​​include task block, input format block, output format block, output note block, example block, and question block. The order from the beginning of the prompt is an example of the position of a block in the prompt. Therefore, the attribute of the order from the beginning of the prompt is an example of an item that represents the position of a block in the prompt.

[0052] Thus, each of the first to sixth attributes is set to one of a task block, an input format block, an output format block, an output note block, an example block, and a question block.

[0053] In the example shown in Figure 4, in addition to the first to sixth attributes (in order from the beginning), the language used for instruction statements and the language used for output statements are also set as attributes. The possible values ​​for the language used for instruction statements and the language used for output statements are English and Japanese. The "language used for instruction statements" corresponds to the language of the task block, input / output format block, and output note block, and the "language used for output statements" corresponds to the language of the example block and question block.

[0054] Furthermore, the above-mentioned possible values ​​of the language used in the instruction sentence and the language used in the output sentence, namely, English and Japanese, are examples of variations in the expression of the block.

[0055] The optimization unit 12a may perform the following settings (prompt design), for example, by selecting and setting a value selected from the possible values ​​for each attribute.

[0056] 1st: Task block 2nd: Input format block 3rd: Output format block 4th: Output notes block 5th: Example block 6th: Question block Language used for instructions: English Language used for output statements: English The optimization unit 12a creates multiple (multiple types) of prompt design information by switching the value selected from the possible values ​​for each attribute. The values ​​selected from the possible values ​​set for each attribute can be called parameters, input parameters, or hyperparameters.

[0057] The optimization unit 12a generates all combinations of parameters for all attributes by switching between possible values ​​(parameters) for each attribute.

[0058] FIG. 5 is a diagram showing an example of formulation information and examples of all combinations of possible attribute values ​​in the prompt adjustment device 1a as an example of the first embodiment.

[0059] 5, symbol A indicates another example of formulation information, and symbol B indicates all possible combinations of attribute values ​​corresponding to the formulation information indicated by symbol A. However, in this example, the question block is formulated as having only short.

[0060] In the formulation information illustrated by symbol A in FIG. 5, the order from the beginning of the prompt (first to third) is used as an attribute, and the possible values ​​are an instruction block, a question block, and a format block.

[0061] 5, in addition to the first to third blocks (in order from the beginning), the length of the description of the first block (1stlen), the length of the description of the second block (2ndlen), and the length of the description of the third block (3rdlen) are also set as attributes. The possible values ​​for the length of the description of the first block (1stlen), the length of the description of the second block (2ndlen), and the length of the description of the third block (3rdlen) are short and long. However, if the first block (1st) {or the second block (2nd), or the third block (3rd)}} is a question, the only possible value for 1stlen (or 2ndlen, 3rdlen) is short.

[0062] The aforementioned possible values ​​of "short" and "long" for the length of the explanation for the first block (1stlen), the length of the explanation for the second block (2ndlen), and the length of the explanation for the third block (3rdlen) are examples of variations in block representation. For example, the accuracy of prompts can change significantly depending on whether the prompt for a task block is short (brief and concise) or long (long and detailed).

[0063] In this example of attribute settings indicated by symbol A in Figure 5, the total number of possible combinations of values ​​for each attribute is 48 (= 6 x 2 x 2 x 2) ÷ 2 (because question can only be short) = 24, as indicated by symbol B. The number of possible combinations of values ​​for each attribute is represented as K. In the example indicated by symbol B in Figure 3, K = 24.

[0064] The formulation processing unit 13 stores the generated formulation information and all combinations of values ​​that each attribute can take in a predetermined storage area of ​​the memory 10c.

[0065] The formulation processing unit 13 creates formulation information (configuration information) that represents the configuration of a prompt as a combination of multiple (attributes: order, language; see FIG. 4) or multiple (attributes: order, length; see FIG. 5) corresponding to the multiple blocks that make up the prompt, and the values ​​that each item can take.

[0066] The optimization unit 12a generates a plurality (K) of prompts (sample prompts) by performing processing such as rearranging the sample data with correct answers to match each of all (K) combinations of possible values ​​of each attribute generated by the formulation processing unit 13. The optimization unit 12a generates K×N sample prompts by using a plurality (N) of sample data with correct answers to generate a plurality (K) of prompts (sample prompts).

[0067] The optimization processing unit 18 generates a plurality of sample prompts corresponding to a plurality of combinations based on the formulation information and a plurality of sample data with correct answers.

[0068] The optimization unit 12a inputs each of the generated K × N sample prompts into the LLM, compares the output of the LLM with the correct answer, and calculates the accuracy for each combination. Here, the accuracy for each combination may be calculated by, for example, calculating the average accuracy of the N sample data with correct answers. Note that the accuracy calculation method is not limited to this and can be modified as appropriate.

[0069] The optimization unit 12 a may create management information for managing the accuracy of each combination. The management information includes combinations of multiple sample data with correct answers and combinations of values ​​that can be taken by multiple attributes in the formulation information.

[0070] FIG. 6 is a diagram illustrating management information created by the optimization unit 12a of the prompt adjustment device 1a according to the first embodiment.

[0071] FIG. 6 shows management information in the case where four (N=4) sample data with correct answers are used based on the formulation information exemplified in FIG.

[0072] In the management information exemplified in FIG. 6, a combination id, a sample_id, an attribute combination pattern, and whether or not the answer is correct are associated with each other.

[0073] The attribute combination pattern may be referred to as a parameter set. The combination id is identification information that identifies the attribute combination pattern. The sample_id is identification information that identifies the sample data with correct answers. Whether or not the answer is correct may indicate whether or not the result obtained by inputting a prompt generated based on the corresponding sample data with correct answers and the attribute combination pattern into the LLM is correct, or may be calculated using a commonly used accuracy index. Here, an example of simply whether or not the answer is correct is shown.

[0074] FIG. 7 is a diagram illustrating prompt design information generated by the optimization unit 12a of the prompt adjustment device 1a as an example of the first embodiment.

[0075] In the example shown in FIG. 7, symbol A indicates an example in which prompt design information is represented in a table format, and symbol B indicates an example in which prompt design information is represented in JSON (JavaScript Object Notation) format.

[0076] The optimization unit 12a passes the generated prompt design information to the output processing unit 17. For example, the optimization unit 12a may store the prompt design information in a predetermined storage area of ​​the memory 10c or the storage unit 10d. The output processing unit 17 may obtain the information by reading out the information stored in these storage areas.

[0077] The optimization unit 12a extracts (generates) the most accurate combination as prompt design information. The attributes and possible values ​​included in the combination included in this prompt design information are used as adjustment items for adjusting the prompt.

[0078] That is, the optimization unit 12a determines, as the adjustment item for adjusting the prompt, the item included in the combination used to generate the sample prompt that has the highest inference accuracy when multiple sample prompts are input into an LLM (language model).

[0079] The output processing unit 17 performs processing for outputting the prompt design information generated by the optimization unit 12a. For example, the output processing unit 17 may output the prompt design information to another system.

[0080] (B) Operation The processing of the optimization unit 12a of the prompt adjustment device 1a according to the first embodiment configured as described above will be described with reference to the flowchart (steps S1 to S6) shown in FIG.

[0081] In step S1, the optimization unit 12a extracts, from the input data stored by the input processing unit 11, a plurality of (N) sample data sets with correct answers to be used for optimizing the prompts.

[0082] In step S2, the formulation processing unit 13 parameterizes (formulates) the configuration of one or more sample prompts to create formulation information.

[0083] In step S3, the optimization unit 12a generates all (K) combinations of parameters for all attributes by switching the values ​​(parameters) that can be taken for each attribute in the formulation information. The optimization unit 12a also generates K × N sample prompts by using N pieces of sample data with correct answers to generate K sample prompts.

[0084] In step S4, the optimization unit 12a inputs each of the generated K × N sample prompts into the LLM and calculates the accuracy of each output LLM output. Here, for each sample prompt, it calculates whether each output LLM output is correct (0: incorrect, 1: correct).

[0085] In step S5, the optimization unit 12a calculates the accuracy for each combination by calculating the average accuracy of the N sample prompts, and in step S6, the optimization unit 12a extracts (generates) the combination with the highest accuracy as prompt design information, after which the process ends.

[0086] Next, the details of the process of step S4 in the flowchart shown in FIG. 8 will be described with reference to the flowchart (steps S41 to S48) shown in FIG.

[0087] In step S41, the optimization unit 12a creates a table having N x K rows, each of which has a column with items such as combination ID, sample_id, attribute combination pattern, and whether or not the answer is correct. Note that the "correct or not" column is left blank. In this table, rows are created in ascending order of combination ID, and for the same combination ID, rows are created in ascending order of sample_id. The attribute combination pattern is the value of each attribute corresponding to the combination ID.

[0088] In step S42, a loop process is started in which the control up to step S47 is repeatedly performed for all nodes present in the N samples. The optimization unit 12a extracts one sample from the N samples. This extracted sample is called sample n.

[0089] In step S43, a loop process is started in which the control up to step S46 is repeatedly performed for all K combinations. The optimization unit 12a extracts one combination from the K combinations. This extracted combination is referred to as combination k.

[0090] In step S44, the optimization unit 12a generates a prompt for the combination k based on the sample n.

[0091] In step S45, the optimization unit 12a inputs the generated prompt into the LLM and obtains the result (output).

[0092] In step S46, the optimization unit 12a registers the result of the LLM in the (N×k+n)th row of the management information created in step S41. For example, if the prompt is a classification question, the optimization unit 12a registers information indicating whether the result is the correct answer, and if the prompt is a ranking question, the optimization unit 12a registers information on the rank and calculation results such as a general accuracy index based on the rank (rank) of the correct answer.

[0093] In the management information illustrated in FIG. 6, when the first row is row 0, the combination id=k, sample_id=n is the N×k+nth row.

[0094] In step S47, loop end processing corresponding to step S43 is performed. When processing for all combinations is completed, the process proceeds to step S48.

[0095] In step S48, a loop end process corresponding to step S42 is performed. When the process for all samples is completed, the process ends.

[0096] (C) Effect As described above, according to the prompt adjustment device 1a of the first embodiment, the formulation processing unit 13 (optimization unit 12a) formulates the prompt by dividing it into multiple blocks, and creates formulation information that associates multiple attributes with the values ​​that each attribute can take.

[0097] Furthermore, the optimization unit 12a processes the sample data with correct answers in accordance with all possible combinations of values ​​for each attribute, thereby generating a plurality of sample prompts.

[0098] Then, the optimization unit 12a inputs each of the generated sample prompts into the LLM, compares the output of the LLM with the correct answer, calculates the accuracy for each combination, and extracts (generates) the combination with the highest accuracy as prompt design information.

[0099] This makes it easy to generate prompts that can get highly accurate results from the LLM.

[0100] (II) Description of the Second Embodiment (A) Configuration In the prompt adjustment device 1a of the first embodiment described above, the optimization unit 12a inputs each of the generated sample prompts into the LLM, compares the output of the LLM with the correct answer, and calculates the accuracy for each combination. Therefore, for example, if the formulation information has a large number of attributes, the number of possible combinations of values ​​for each attribute becomes enormous, which may increase the calculation cost.

[0101] Therefore, the prompt adjustment device 1b as an example of the second embodiment aims to easily generate prompts that can obtain highly accurate results from an LLM and to reduce calculation costs. The prompt adjustment device 1b is an information processing device that adjusts prompts to be input to a language model so that the output from the language model is highly accurate.

[0102] FIG. 10 is a diagram illustrating a functional configuration of a prompt adjustment device 1b as an example of the second embodiment.

[0103] 10, the prompt adjustment device 1b of the second embodiment includes an optimization unit 12b instead of the optimization unit 12a of the first embodiment, and other parts are configured in the same manner as the prompt adjustment device 1a of the first embodiment. Also, the prompt adjustment device 1b of the second embodiment has the same hardware configuration as the prompt adjustment device 1a of the first embodiment.

[0104] The functions of the input processing unit 11, the optimization unit 12b, and the output processing unit 17 may be realized, for example, by the processor 10a described above executing the program 10h (prompt adjustment program).

[0105] The optimization unit 12b includes an interaction extraction unit 14, a pre-processing unit 15, a post-processing unit 16, and an optimization unit 18 in addition to the formulation processing unit 13 similar to that of the first embodiment.

[0106] In the drawings, the same reference numerals as those already mentioned indicate the same parts, and therefore the description thereof will be omitted.

[0107] The interaction extraction unit 14 uses a known data mining technique to detect interactions between multiple attributes in the input data. Data mining that detects interactions between multiple attributes in the input data is an example of the first data mining. An interaction is a synergistic effect that appears when multiple attributes are combined. The interaction is an example of a result of the first data mining.

[0108] The interaction extraction unit 14 receives, for example, CSV (Comma Separated Values) data in which the combination of hyperparameters is the attribute and whether the output is correct or not is the label, and extracts interactions having characteristics of correct answers (incorrect answers).

[0109] The interaction extraction unit 14 may receive data mining input information in CSV format, such as the example shown in FIG.

[0110] FIG. 11 is a diagram illustrating data mining input information in the prompt adjustment device 1b according to the second embodiment.

[0111] The optimization unit 12b may create data mining input information such as the example shown in Fig. 11. The data mining input information is information generated based on formulation information and sample data with correct answers, similar to the management information in the first embodiment, and is formed by associating each combination of attributes with information indicating whether the output from the LLM is correct or not.

[0112] The interaction extraction unit 14 performs data mining on the input data mining input information to calculate (extract, detect) interactions between multiple attributes. An interaction is a synergistic effect that appears when multiple attributes are combined.

[0113] The interaction extraction unit 14 extracts interactions between attributes and combinations of attributes through data mining. The interactions extracted by the interaction extraction unit 14 are input to calculation terms for Bayesian optimization by the optimization processing unit 18, which will be described later. Wide Learning (registered trademark), for example, may be used as a data mining method.

[0114] The interaction extraction unit 14 performs data mining using as input data mining input information that associates the results (whether correct or not) obtained by inputting multiple sample prompts into an LLM (language model) with the combinations used to generate the sample prompts.

[0115] 12 and 13 are diagrams illustrating interactions generated by the interaction extraction unit 14 of the prompt adjustment device 1b according to the second embodiment. Note that Fig. 12 illustrates an interaction with high accuracy, and Fig. 13 illustrates an interaction with low accuracy.

[0116] 12 and 13, labels correspond to chunks. Label=1 indicates good accuracy, and label=0 indicates poor accuracy. Chunks indicate interactions.

[0117] Also, "^" means "and" and corresponds to AND in a logical expression. =0 / =1 indicates that the value of the attribute is not applicable / is not relevant.

[0118] For example, the first line of Figure 12, "1stlen_short=0 ∧ 2nd_instruction=1 ∧ 2ndlen_short=0", indicates that the attribute "1stlen" is not "short" (=0), and the attribute "2nd" is "instruction" (=1), and the attribute "2ndlen" is not "short" (=0). In other words, a prompt where a long (non-short) instruction comes second and the first is long (non-short) indicates good precision.

[0119] 13, "1stlen_short=1 ∧ 2nd_instruction=1 ∧ 2ndlen_short=1" indicates that the attribute "1stlen" is "short" (=1), the attribute "2nd" is "instruction" (=1), and the attribute "2ndlen" is "short" (=1). In other words, a short instruction coming second and a short prompt coming first indicate poor accuracy.

[0120] The interaction extraction unit 14 passes the generated interactions to the post-processing unit 16 and the optimization processing unit 18. The interaction extraction unit 14 may, for example, store the generated interactions in a predetermined storage area of ​​the memory 10c or the storage unit 10d. The post-processing unit 16 and the optimization processing unit 18 may obtain the information stored in these storage areas by reading it out.

[0121] The preprocessing unit 15 reduces the data to be processed (data mining) by the interaction extraction unit 14 .

[0122] The preprocessing unit 15 lists in advance, for example, single attributes or combinations of attributes that are qualitatively known to result in low accuracy (poor performance). This list may be called an exclusion list or a skip list. The exclusion list is a list of pairs of attributes and their values ​​determined by formulation. The exclusion list is stored, for example, in a predetermined storage area of ​​the storage unit 10d.

[0123] The exclusion list may store, for example, the condition that the "question block" has the characteristic of completion (the sentence ends midway and the remaining part is left to be answered) and that the "question block" is placed at the beginning of the prompt. The conditions registered in the exclusion list may be called exclusion conditions.

[0124] In addition, the exclusion list may store a condition that it contains expressions that do not match the learning data (general sentences) of the LLM.

[0125] Furthermore, the exclusion list may store conditions where the language changes frequently during the course of a conversation.

[0126] In addition, the pre-processing unit 15 uses, for example, sample data with correct answers to create a sample prompt that matches a single attribute pattern or a combination of attributes that corresponds to the conditions of the exclusion list, and checks in advance whether the results obtained by inputting this sample prompt into the LLM are actually of poor accuracy.

[0127] For example, the preprocessing unit 15 uses sample data with correct answers to set the "question block" as the m-th block from the beginning, and while randomly changing the other blocks, creates prompts, inputs them into the LLM, and checks the accuracy of the output results. This trial is repeated while switching m=1, 2, 3, ..., and if the accuracy is significantly low (for example, the accuracy is below a threshold), the value of the "question block" is not entered in the m-th block.

[0128] That is, the preprocessing unit 15 generates sample prompts (second sample prompts) using sample data with correct answers under conditions (exclusion conditions) that are qualitatively known to result in low accuracy (poor performance) for single-attribute patterns or combinations of attributes, and then inputs these sample prompts into the LLM to verify in advance the accuracy of the results obtained. If the verification results show that the accuracy is significantly low, such as below a preset threshold, the preprocessing unit 15 stores the combination patterns that fall under these exclusion conditions as a skip list, and skips the processing (input to the LLM and processing by the interaction extraction unit 14 (data mining)) for combinations whose conditions are included in the list.

[0129] In this way, the preprocessing unit 15 reduces the variations in the combinations of attributes to be included in the data mining.

[0130] If the experiment verifies that the accuracy is low, the preprocessing unit 15 excludes the combination patterns that meet the conditions in the exclusion list from the processing (data mining) of the interaction extraction unit 14 .

[0131] The preprocessing unit 15 compares the combinations of multiple attributes and their possible values ​​proposed by Bayesian optimization by the optimization processing unit 18 based on the formulation information created by the formulation processing unit 13 with the above-mentioned skip list, and if the proposed attribute pattern or attribute combination pattern is included in the exclusion conditions registered in the skip list, skips subsequent processing (input to the LLM and its accuracy evaluation, and input to the interaction extraction unit 14). Skipping subsequent processing (input to the LLM and its accuracy evaluation, and input to the interaction extraction unit 14) for attribute patterns or attribute combination patterns that meet the skip list conditions can be referred to as pruning.

[0132] The pre-processing unit 15 does not generate combination patterns in advance and then exclude them, but when a combination that matches the skip list is presented from Bayesian optimization, it proceeds to the next proposal loop without processing.

[0133] For example, if "1st: question" is in the skip list, and "1st: question" is included in the Bayesian optimization proposal, the process will be skipped and the next proposal will be advanced. As a result, the skipped combination will not be added to the input for data mining. In other words, the skipped combination will not be submitted to the LLM to calculate accuracy, and the result will not be added to the input for data mining.

[0134] The preprocessing unit 15 excludes configuration information that includes a combination that satisfies the exclusion condition from the target of data mining.

[0135] In addition, before excluding configuration information containing combinations that satisfy the exclusion conditions from the target of data mining, the preprocessing unit 15 generates a second sample prompt using sample data for the combination that satisfies the exclusion conditions, and if the accuracy of the result obtained by inputting this second sample prompt into the language model is below a threshold, excludes the combination that satisfies the exclusion conditions from the target of data mining.

[0136] The post-processing unit 16 uses statistical tests to reduce the interactions output from the interaction extraction unit 14 based on p-values, where p-values ​​represent statistical probabilities.

[0137] The post-processing unit 16 tests the combinations that appear in the interactions output from the interaction extraction unit 14 and calculates p-values. A known statistical testing method, such as the Breslow-Day test, may be used for the test. The post-processing unit 16 excludes interactions with p-values ​​equal to or greater than a threshold value (e.g., 0.05) from the targets of processing by the optimization processing unit 18, which will be described later. The exclusion of interactions from the targets of processing by the optimization processing unit 18 may be referred to as pruning.

[0138] The post-processing unit 16 may rank (sort) and cut off the interactions obtained as a result of data mining by the interaction extraction unit 14 by p-values ​​using a statistical test to determine whether the interactions are significant or not.

[0139] The post-processing unit 16 tests whether interactions between parameter combinations affect inference performance. The post-processing unit 16 excludes interactions whose p-values ​​are equal to or greater than a threshold value (e.g., 0.05), i.e., interactions that have no effect, from the processing targets (search candidates) of the optimization unit 18.

[0140] The post-processing unit 16 excludes interactions (data mining results) whose p-value (a value representing the impact of data mining results on inference performance) exceeds a threshold (in this second embodiment, is equal to or greater than the threshold) from the optimization target.

[0141] The post-processing unit 16 performs modeling for cutoff. The post-processing unit 16 performs logistic modeling using the effects of individual factors and interactions, and performs testing. Specifically, the post-processing unit 16 performs modeling using interactions whose coefficients calculated by logistic regression are not 0. The post-processing unit 16 calculates the effects of each factor and its interactions using logistic regression.

[0142] The optimization processing unit 18 uses the results of data mining by the interaction extraction unit 14 to reduce the number of attribute combinations by using a known optimization method, and determines combinations for generating prompts. As a known optimization method, for example, Bayesian optimization may be used.

[0143] The optimization processor 18 explicitly narrows down the search space of the Bayesian optimization acquisition function and selects parameters therefrom for generating prompt design information.

[0144] The optimization processing unit 18 searches for adjustment items for adjusting prompts through optimization based on the results (interactions) of the data mining.

[0145] The optimization processing unit 18 receives the remaining interactions generated by the interaction extraction unit 14 after the pruning by the post-processing unit 16 .

[0146] The optimization processing unit 18 processes only the interactions resulting from combinations of attributes, which are output from the interaction extraction unit 14 and are cut off by the post-processing unit 16, as calculation terms for Bayesian optimization, thereby reducing the time required for optimization.

[0147] Bayesian optimization is a method for arriving at an optimal solution with a small number of trials by sampling points with a high probability of being the optimal solution. Note that when performing Bayesian optimization, variables must be independent.

[0148] Here, the process of narrowing down the search range using interactions in Bayesian optimization in the optimization processing unit 18 will be described.

[0149] For example, in the case of all possible combinations of formulation information and attribute values ​​shown in Fig. 5, the original search range has three possibilities for "1st", "2nd", and "3rd", and two possibilities for "1stlen", "2ndlen", and "3rdlen". However, if 1st (2nd, 3rd) is "question", there is only one possibility for 1stlen (2ndlen, 3rdlen) (only "short").

[0150] Here, for example, it is assumed that the following two interactions are extracted from the multiple interactions illustrated in FIG.

[0151] Condition for good accuracy: 1stlen_short=0 ∧ 2nd_instruction=1 ∧ 2ndlen_short=0 Condition for good accuracy: 1st_instruction=0 ∧ 1stlen_short=0 ∧ 2ndlen_short=0 The optimization processing unit 18 calculates the AND condition of the above interactions. As a result, 1st_instruction=0 ∧ 1stlen_short=0 ∧ 2nd_instruction=1 ∧ 2ndlen_short=0 is obtained.

[0152] In this AND condition, the 1st search range is instruction=0, meaning anything other than instruction. 1stlen is short=0, meaning anything other than short. 2nd is instruction=1, meaning instruction. 2ndlen is short=0, meaning anything other than short.

[0153] Therefore, the search range is limited as shown in the limited range of FIG.

[0154] Fig. 14 is a diagram illustrating an example of a limited search range, in which the limited search range for Bayesian optimization is associated with each attribute.

[0155] For example, as described above, the search range for 1st is other than "instruction," so in Fig. 14, "format" and "question" are set as a limited range in association with the attribute 1st (first block). Also, the search range for 2nd is "instruction," so "instruction" is set in association with the attribute 2nd (second block). The search range for 3rd is other than 1st and 2nd, so it is inevitably determined once 1st is determined. A parameter set searched for with a limited range is expected to be highly accurate.

[0156] In the prompt adjustment device 1b of the second embodiment, the interaction extraction unit 14 calculates the interactions between attributes, and then the optimization processing unit 18 inputs the interaction terms into Bayesian optimization. The optimization processing unit 18 narrows down and determines the optimal combination through Bayesian optimization. That is, the results of data mining are used to explicitly narrow down the search space of the acquisition function for Bayesian optimization, and parameters are selected from there.

[0157] The optimization processing unit 18 determines the most accurate combination among the combinations proposed by Bayesian optimization as the adjustment item. For example, a trial number indicating how many times the proposal will be made may be set in advance by Bayesian optimization, and the proposal may be repeated until the trial number is reached.

[0158] This allows the optimization unit 12b to perform Bayesian optimization and optimize the prompt without trying all possible combinations of data.

[0159] Immediately after the start of processing by the optimization unit 12b, the number of data is small, making it difficult to extract useful interactions. However, as the processing progresses, the amount of data increases, making it easier to ensure statistical significance, and the optimization processing unit 18 can accelerate the narrowing down process.

[0160] (B) Operation The processing of the optimization unit 12b of the prompt adjustment device 1b according to the second embodiment configured as described above will be described with reference to the flowchart (steps S11 to S21) shown in FIG.

[0161] In step S11, the optimization unit 12b extracts, from the input data stored by the input processing unit 11, a plurality of (N) sample data sets with correct answers to be used for optimizing the prompts.

[0162] In step S12, the formulation processing unit 13 parameterizes (formulates) the configuration of one or more sample prompts to create formulation information.

[0163] In step S13, the preprocessing unit 15 creates a skip list from the qualitative exclusion list and stores the skip list in a predetermined storage area of ​​the memory 10c or the storage unit 10d. Details of the process of step S13 will be described later with reference to FIG. 16.

[0164] In step S14, the optimization unit 12a sets the number of loops to N (number of loops=N).

[0165] In step S15, the optimization processing unit 18 proposes parameters by Bayesian optimization based on the accuracy of the parameters from the first loop to the previous loop. The proposed parameters may be called proposed parameters.

[0166] In step S16, the preprocessing unit 15 checks whether the proposed parameter is included in the skip list. If the check result shows that the proposed parameter is included in the skip list (see the YES route in step S16), the process returns to step S15.

[0167] If the proposed parameter is not included in the skip list (see the NO route from step S16), the process proceeds to step S17.

[0168] In step S17, the optimization unit 12b generates a sample prompt using the proposed parameters, calculates whether the LLM is correct for this created sample prompt, and stores the result in the corresponding item indicating whether it is correct or not in the data mining input information.

[0169] In step S18, the interaction extraction unit 14 performs data mining using as input the parameters from the first to last loop in the data mining input information and a table indicating whether the answer is correct or not, to calculate interactions.

[0170] In step S19, the post-processing unit 16 performs testing, calculates the p-value of each interaction output from the interaction extraction unit 14, and performs pruning by excluding interactions whose p-values ​​are equal to or greater than a threshold value and leaving only interactions whose p-values ​​are less than the threshold value.

[0171] In step S20, the optimization processing unit 18 adds the remaining interactions to the Bayesian optimization term and selects the next parameter set.

[0172] In step S21, the optimization unit 12b checks whether a termination condition is satisfied. The termination condition may be, for example, detecting that the accuracy of the LLM output no longer improves, or reaching a predetermined number of loops, and may be implemented in an appropriate modified form.

[0173] If the termination condition is not satisfied (see the NO route in step S21), the process returns to step S15, whereas if the termination condition is satisfied (YES in step S21), the process ends.

[0174] Next, the processing of the preprocessing unit 15 of the prompt adjustment device 1b according to the second embodiment will be described with reference to the flowchart (steps S201 to S208) shown in Fig. 16. The flowchart shown in Fig. 16 shows details of the processing of step S13 in the flowchart shown in Fig. 15.

[0175] In step S201, the preprocessing unit 15 sets the skip list to empty. In step S202, the preprocessing unit 15 automatically compares the exclusion list with the attributes and possible values ​​formulated for the target data, thereby determining the items to be verified in the exclusion list. This allows for correspondence to be achieved even if the attribute names / values ​​written in the exclusion list do not completely match the attribute names / values ​​set in the formulation, provided that they substantially match. For example, the exclusion list may contain "QUESTION:1," but the attribute / possible value set in the formulation may be "1st: question." This corresponds to a case where the attribute and possible values ​​are swapped or the case is different, such as when the attribute and possible values ​​are swapped or the case is different.

[0176] In step S203, a loop process is started in which the control up to step S206 is repeatedly performed for all items to be verified in the exclusion list.

[0177] In step S204, the pre-processing unit 15 performs a pre-check using the sample data with correct answers. The pre-processing unit 15 verifies the accuracy of the LLM results using the sample prompts created using the sample data with correct answers.

[0178] In step S205, the preprocessing unit 15 checks whether the accuracy of the LLM result is low by, for example, comparing the accuracy of the LLM result with a preset threshold. If the accuracy of the LLM result is less than the threshold, i.e., if the accuracy is low (see the YES route in step S205), the process proceeds to step S206.

[0179] In step S206, the preprocessing unit 15 adds the combination patterns that meet the corresponding conditions in the exclusion list to the skip list, and then proceeds to step S207.

[0180] Also, if the result of checking in step S205 shows that the accuracy of the LLM result is equal to or greater than the threshold, that is, if the accuracy is high (see the NO route in step S205), the process proceeds to step S207.

[0181] In step S207, loop end processing corresponding to step S203 is performed. Here, when processing for all items to be verified in the exclusion list is completed, the skip list is returned in step S208. Thereafter, processing ends. Note that in the optimization processing unit 18, if the parameter set proposed by Bayesian optimization matches the skip list, subsequent processing is not performed and the optimization processing unit 18 proceeds to the next Bayesian optimization proposal.

[0182] Next, the processing of the post-processing unit 16 of the prompt adjustment device 1b according to the second embodiment will be described with reference to the flowchart (steps S31 to S35) shown in FIG.

[0183] In step S31, a loop process is started in which the control of step S32 is repeatedly performed for all interactions generated by the interaction extraction unit 14.

[0184] In step S32, the post-processing unit 16 calculates a p-value using statistical verification.

[0185] In step S33, a loop end process corresponding to step S31 is performed. When the process for all interactions is completed, the process proceeds to step S34.

[0186] In step S34, the post-processing unit 16 sorts all of the interactions based on the p-values. For example, the post-processing unit 16 sorts all of the interactions in ascending order based on the p-values.

[0187] In step S35, the post-processing unit 16 performs pruning on the sorted interaction column. For example, the post-processing unit 16 may select and retain only interactions whose p-values ​​are less than a preset threshold (p-value<threshold). Alternatively, the post-processing unit 16 may select (extract) and retain only a predetermined number (m) of interactions whose p-values ​​are less than the threshold and are ranked from the top of the p-values. Then, the processing ends.

[0188] Next, the Bayesian optimization process by the optimization processing unit 18 of the prompt adjustment device 1b according to the second embodiment will be described with reference to the flowchart (steps S71 to S75) shown in Fig. 18. Note that Fig. 18 illustrates a process for maximizing accuracy.

[0189] In step S71, the optimization processing unit 18 sets the loop count to N (loop count=N). N corresponds to the number of trials that represent how many times a proposal is made.

[0190] In step S72, the optimization processing unit 18 starts a loop process in which steps S73 to S74 are repeatedly performed for all loops represented by the total number of loops. The number of parameter proposal loops is represented by i. i is a natural number equal to or less than N, and its initial value is 1 (i=1).

[0191] In step S73, the optimization processing unit 18 proposes parameters (proposed parameters) by Bayesian optimization based on the accuracy of the parameters from loop count 1 to (i-1).

[0192] In step S74, the optimization processing unit 18 calculates the accuracy using the proposed parameters. In step S75, loop end processing corresponding to step S72 is performed. When the number of loops reaches N, this flow ends.

[0193] In this way, in the Bayesian optimization by the optimization processing unit 18, the next parameters are proposed using the results obtained up to that point.

[0194] (C) Effect As described above, according to the prompt adjustment device 1b of the second embodiment, the interaction extraction unit 14 detects interactions of multiple attributes using a data mining technique based on data mining input information generated based on formulation information and sample data with correct answers.

[0195] The optimization processor 18 can then improve the efficiency of the optimization technique by narrowing down the combinations of attributes using the results of data mining by the interaction extractor 14, and can determine combinations (prompt design information) for generating accurate prompts from a huge number of combinations with a limited number of Bayesian optimization loops. In other words, the optimization processor 18 explicitly narrows down the search space of the acquisition function for Bayesian optimization using the interactions obtained by data mining, and selects parameters from there.

[0196] This allows prompts that can provide highly accurate results from LLM to be generated in a short time with a small amount of calculation, which means that the calculation time and cost required to generate the prompts can be reduced.

[0197] The preprocessing unit 15 checks attribute combinations in advance based on an exclusion list that registers conditions that are qualitatively known to result in low accuracy (poor performance), generates a skip list, and performs pruning. In other words, if an attribute combination proposed by Bayesian optimization is included in the skip list, the process of generating a prompt using that combination and submitting it to the LLM to calculate accuracy, and the process of using the result in data mining, are skipped. This reduces the number of variations in attribute combinations to be included in data mining in the interaction extraction unit 14, thereby also reducing calculation time and calculation costs.

[0198] Furthermore, the preprocessing unit 15 uses sample data with correct answers to create sample prompts in accordance with combination patterns corresponding to the conditions in the exclusion list, and inputs these sample prompts into the LLM to check in advance whether the results obtained are truly poor in accuracy. This improves the reliability of pruning by the preprocessing unit 15.

[0199] The post-processing unit 16 tests the combinations that appear in the interactions output from the interaction extraction unit 14 and calculates p-values. The post-processing unit 16 then performs a cutoff to exclude interactions with p-values ​​equal to or greater than a threshold from the targets of processing by the optimization unit 18, which will be described later. This reduces the calculation time and cost required for Bayesian optimization by the optimization unit 18.

[0200] In addition, the post-processing unit 16 uses statistical testing to reduce the interactions output from the interaction extraction unit 14 based on p-values. This makes it possible to improve calculation efficiency by not considering interactions that are not statistically significant (that are unlikely to directly affect accuracy). Furthermore, although interactions of ordinal variables are not independent, performing testing also guarantees independence between attributes, making it possible to use Bayesian optimization.

[0201] The interaction extraction unit 14 calculates interactions between statistically significant attributes, thereby ensuring the independence between terms to be included in Bayesian optimization in the optimization processing unit 18. Furthermore, the post-processing unit 16 performs testing to calculate p-values, and by using these p-values ​​to limit the number of interaction terms and run Bayesian optimization, it is possible to optimize prompts while reducing calculation time and cost without trying every possible combination of data.

[0202] (III) Description of the third embodiment As described above, in the prompt adjustment device 1a of the first embodiment, for example, if the number of attributes in the formulation information is large, the number of possible combinations of values ​​for each attribute becomes enormous, which may increase the calculation cost.

[0203] In the prompt adjustment device 1b of the second embodiment, the results of data mining are used to narrow down the combinations of attributes, but even after such narrowing down, the number of combinations of attributes may still be enormous. For example, even if 700 million candidates are narrowed down to 1 / 1000, the number of candidates is still 700,000.

[0204] Furthermore, even if the combination of attributes can be successfully narrowed down, it is difficult to determine which combination of attributes will result in highly accurate output from the language model. It is difficult to predict the optimal combination based only on the information that "this combination will produce this level of accuracy."

[0205] In the prompt adjustment device 1c of this third embodiment, the degree to which accuracy will improve for a combination of attributes is predicted, and attribute combinations that will improve accuracy are output as recommended combinations, thereby achieving more efficient prompt optimization.

[0206] The prompt adjustment device 1c is an information processing device that adjusts prompts to be input to a language model so that the output from the language model is highly accurate.

[0207] (A) Configuration FIG. 22 is a diagram illustrating the functional configuration of a prompt adjustment device 1c as an example of the third embodiment.

[0208] The prompt adjustment device 1c of the third embodiment illustrated in Figure 22 is equipped with an optimization unit 12c instead of the optimization unit 12b of the second embodiment, and other parts are configured in the same way as the prompt adjustment device 1b of the second embodiment.

[0209] Moreover, the optimization unit 12c includes a history management unit 19, a history data mining processing unit 20, and a trial parameter determination unit 21 instead of the preprocessing unit 15 of the optimization unit 12b of the second embodiment. Note that other parts are configured in the same way as the prompt adjustment device 1b of the second embodiment. Moreover, the prompt adjustment device 1c of the third embodiment has the same hardware configuration as the prompt adjustment device 1a of the first embodiment and the prompt adjustment device 1b of the second embodiment.

[0210] The functions of the input processing unit 11, the optimization unit 12c, and the output processing unit 17 may be realized, for example, by the processor 10a described above executing the program 10h (prompt adjustment program).

[0211] In the drawings, the same reference numerals as those already mentioned indicate the same parts, and therefore the description thereof will be omitted.

[0212] The history management unit 19 manages the association between the combination of parameters (item combination) determined as an adjustment item by the optimization processing unit 18 and the accuracy of the result obtained by inputting prompts (sample prompt, first sample prompt, second sample prompt) created based on the combination into the LLM. Inputting the prompts created based on the item combination into the LLM may be referred to as a trial. Furthermore, the accuracy of the result obtained by inputting the prompts created based on the combination into the LLM may simply be referred to as the accuracy of the combination. The accuracy of the result obtained by inputting the prompts into the LLM is an example of the first result and the second result.

[0213] The history management unit 19 associates each combination determined by the optimization processing unit 18 with its accuracy and stores it as past history information 22 in a specified storage area such as the storage unit 10d or the memory 10c.

[0214] The history data mining processing unit 20 performs data mining on the past history information 22. Therefore, after multiple trials are performed while changing parameters, the history data mining processing unit 20 performs processing based on the past history information 22 in which the results of these trials are stored.

[0215] The history data mining processing unit 20 performs second data mining using the mining model 23 .

[0216] The historical data mining processing unit 20 creates (trains) a mining model 23 in a training phase (learning phase), and in a prediction phase, evaluates item combinations (second data mining) using the mining model 23. The mining model 23 is a machine learning model that receives input of past parameters (third combination) and candidate parameters (fourth combination candidate) and outputs the prediction accuracy for the candidate parameters.

[0217] In the training phase, the history data mining processing unit 20 first creates training data 24 to be input to the mining model 23 .

[0218] FIG. 23 is a diagram for explaining a method for creating training data 24 by the history data mining processing unit 20 of the prompt adjustment device 1c as an example of the third embodiment.

[0219] 23, symbol A indicates an example of training data 24, symbol B indicates an example of past history information 22 used to create the training data 24. Symbol C indicates possible values ​​of each parameter in the past history information 22.

[0220] The past history information 22 shown in FIG. 23 associates, for a plurality of trials, the combination of attributes (parameters) in each trial with information indicating the accuracy thereof, and includes trial, parameter, and accuracy.

[0221] "Trial" is identification information for identifying a trial, and an integer is used in the example shown in Fig. 23. The past history information 22 shown in Fig. 23 indicates a state in which the third trial (trial=3) has been completed, and three trials performed in the order trial=1, trial=2, and trial=3 are registered.

[0222] A parameter has multiple attributes. In the example shown in Fig. 23, the parameter has four attributes: Inst, 1Q, 1O, and 1C. In addition, in the example shown in Fig. 23, the possible values ​​of these attributes Inst, 1Q, 1O, and 1C are m, v1, and v2, as shown by the symbol C.

[0223] The accuracy is the accuracy of the result obtained by inputting a prompt created based on the parameters (combination) into the LLM, that is, the accuracy of the combination. It is assumed that accuracy calculation is performed for each parameter in each trial. The accuracy calculation for the parameters in each trial may be performed, for example, by the optimization processing unit 18 using a known method.

[0224] All trials registered in the past history information 22 have been input into the LLM and have undergone accuracy calculation, and therefore can be said to be past trials.

[0225] The history data mining processing unit 20 creates training data 24 using past history information 22 .

[0226] The training data 24 illustrated in FIG. 23 includes a trial set, parameter #1, parameter #2, and a label accuracy difference.

[0227] A trial set in the training data 24 is a combination of two trials selected from a plurality of past trials included in the past history information 22. The two trials constituting the trial set may be referred to as a first trial and a second trial. The second trial may be one for which accuracy calculation was performed after the first trial. In particular, the second trial may be referred to as a current trial, and the first trial may be referred to as a past trial.

[0228] Parameter #1 in the training data 24 is a parameter in a past trial (first trial), and parameter #2 is a parameter in a current trial (second trial).

[0229] The label accuracy difference in the training data 24 represents the comparison result between the accuracy of parameter #1 and the accuracy of parameter #2. For example, 1 is set when the accuracy of parameter #2 is higher than the accuracy of parameter #1, and 0 is set when the accuracy of parameter #2 is equal to or lower than the accuracy of parameter #1. In other words, the label accuracy difference represents which parameter has better accuracy for the two trials that make up the trial set. The label accuracy difference is an example of comparison information that represents the comparison result between the first result (the accuracy of parameter #1) and the second result (the accuracy of parameter #2).

[0230] In the training data 24, parameter #1 corresponds to a first combination of possible values ​​for the items used to generate a first sample prompt of the sample prompts, and parameter #2 corresponds to a second combination of possible values ​​for the items used to generate a second sample prompt of the sample prompts.

[0231] The accuracy of a first sample prompt in the past history information 22 corresponds to the first result obtained by inputting the first sample prompt into the LLM. The accuracy of a second sample prompt corresponds to the second result obtained by inputting the second sample prompt into the LLM. The label accuracy difference can be said to include both the first and second results.

[0232] The combination of parameter #1 and the accuracy (first result) obtained by inputting a first sample prompt created using this parameter #1 into the LLM may be referred to as first information. This first information may also be referred to as past trial information. The combination of parameter #2 and the accuracy (second result) obtained by inputting a second sample prompt created using this parameter #2 into the LLM may be referred to as second information. This second information may also be referred to as current trial information.

[0233] It is desirable that all combinations (trial sets) of multiple trials be registered in the training data 24. For example, if three trials, trial=1, trial=2, and trial=3, are registered in the past history information 22, it is desirable that three trial sets, (1, 2), (1, 3), and (2, 3), be registered in the training data 24.

[0234] The history data mining processing unit 20 uses the created training data 24 to create a mining model 23. The history data mining processing unit 20 uses a combination of parameter #1 and parameter #2 in the training data 24 as one line of data (input data) and the label accuracy difference as a label (correct data, teacher data) to create the mining model 23. The history data mining processing unit 20 combines past trials to generate training data 24 and performs data mining from it.

[0235] An example of first information is information associating the accuracy (first result) obtained by inputting a first sample prompt (first prompt) into the LLM with parameter #1 used to generate the first prompt. An example of second information is information associating the accuracy (second result) obtained by newly inputting a second sample prompt (second prompt) into the LLM with parameter #2 used to generate the second prompt. In the training phase, a second data mining is performed using the first information and the second information as input.

[0236] The history data mining processing unit 20 can predict "whether accuracy will improve" depending on the parameters, for example, by adding the results of data mining to the Bayesian optimization term.

[0237] Furthermore, in the prediction phase, the history data mining processing unit 20 creates prediction data 25 to be input to the mining model 23 .

[0238] FIG. 24 is a diagram illustrating the prediction data 25 in the prompt adjustment device 1c as an example of the third embodiment.

[0239] The prediction data 25 shown in FIG. 24 includes a trial set, parameter #3, parameter #4, a predicted value, and an average predicted value.

[0240] A trial set in the prediction data 25 is a combination of one trial (past trial) selected from multiple past trials included in the past history information 22 and a trial (candidate trial) based on a candidate parameter (candidate parameter) proposed by the prompt adjustment device 1c.

[0241] In the example shown in FIG. 24, candidate trials are identified using an identifier that combines the string "candidate" with an integer.

[0242] It is desirable for the history data mining processing unit 20 to prepare multiple types of candidate parameters. The candidate parameters have parameters that are different from those of past trials registered in the past history information 22. The candidate parameters may be created using various known methods, and detailed descriptions of the methods for creating candidate parameters will be omitted. As an example, the candidate parameters may be created by narrowing down the candidate parameters using interactions. In the example shown in FIG. 24, N types of candidate parameters (candidate 1 to candidate N) are shown.

[0243] Of the two trials that make up the trial set in the prediction data 25, the past trial may be referred to as the third trial, and the candidate trial may be referred to as the fourth trial. The third trial may be referred to as the past trial, and the fourth trial may be referred to as the current trial.

[0244] Parameter #3 in the prediction data 25 is a parameter in a past trial (third trial). Parameter #3 in the prediction data 25 may be referred to as a past parameter. Parameter #3 is an example of a third combination of possible values ​​for multiple items used to generate a previously created sample prompt (third sample prompt).

[0245] Parameter #4 in the prediction data 25 is a parameter in the candidate trial (fourth trial). Parameter #4 in the prediction data 25 may be called a candidate parameter. Parameter #4 is an example of a fourth candidate combination of possible values ​​for multiple items to be used in generating a new prompt (fourth sample prompt).

[0246] 24, for convenience, the prediction data 25 is sorted based on the identifiers that identify the candidate trials, so that candidate trials with the same identifiers are consecutive.

[0247] The predicted value in the prediction data 25 is a predicted value of the accuracy of the candidate trial (fourth trial), and is undetermined (??) because it is a value predicted by the mining model 23. The average predicted value is the average of the predicted values, so this average predicted value is also undetermined (??).

[0248] In the prediction data 25, it is desirable that all trials in the past history information 22 are set in parameter #3. It is also desirable that a plurality of candidate trials are set in parameter #4.

[0249] As shown in Figure 24, for example, if three trials, trial = 1, trial = 2, and trial = 3, are registered in the past history information 22 and N candidate trials are set, it is desirable that all combinations of the three past trials and the N candidate trials be set as trial sets in the training data 24.

[0250] The history data mining processing unit 20 inputs the created prediction data 25 into the mining model 23 to obtain an output.

[0251] The mining model 23 outputs a predicted value of the accuracy of the candidate trial (fourth trial) for each trial set of the prediction data 25. In the prediction phase, the history data mining processing unit 20 inputs parameter #3 (past parameter) and parameter #4 (candidate parameter) into the mining model 23 (machine learning model) to obtain the prediction accuracy for the candidate parameter.

[0252] The history data mining processing unit 20 may create output information 26 using the predicted value for each trial set obtained from the mining model 23 .

[0253] FIG. 25 is a diagram illustrating output information 26 in a prompt adjustment device 1c as an example of the third embodiment.

[0254] The output information 26 shown in FIG. 25 is configured by registering the predicted values ​​obtained from the mining model 23 in the prediction data 25 shown in FIG.

[0255] In the output information 26 illustrated in Fig. 25, for example, for a trial set (1, candidate 1) consisting of a past trial represented by "trial = 1" and a candidate trial represented by "candidate 1", "0.3" is set as the predicted value of the accuracy of the candidate trial represented by "candidate 1". Also, for example, for a trial set (2, candidate 1) consisting of a past trial represented by "trial = 2" and a candidate trial represented by "candidate 1", "0.3" is set as the predicted value of the accuracy of the candidate trial represented by "candidate 1".

[0256] Furthermore, the history data mining processing unit 20 may calculate an average of the predicted values ​​for each candidate trial based on each predicted value calculated by the mining model 23 and register the average in the output information 26 .

[0257] 25, for example, the predicted value of the trial set (1, candidate 1) is "0.3," the predicted value of the trial set (2, candidate 1) is "0.3," and the predicted value of the trial set (3, candidate 1) is "0.6." Therefore, the average predicted value of the candidate trial represented by "candidate 1" is calculated as 0.4 {=(0.3+0.3+0.6) / 3}.

[0258] That is, the output information 26 illustrated in Figure 25 is created by registering the predicted values ​​of each trial set obtained by the mining model 23 for the prediction data 25 shown in Figure 24 and the average predicted values ​​for each candidate trial calculated by the historical data mining processing unit 20.

[0259] The history data mining processing unit 20 stores the created output information 26 in a predetermined storage area such as the memory 10c or the storage unit 10d.

[0260] The trial parameter determination unit 21 determines the optimal combination (parameters) based on the output information 26 created by the history data mining processing unit 20 .

[0261] The trial parameter determination unit 21 determines the candidate trial parameter (candidate parameter) with the highest average predicted value, i.e., the most accurate, from the output information 26 as the parameter for the next trial.

[0262] For example, in the example shown in FIG. 25, the average predicted value of the candidate trial represented by "candidate 2" is 0.4333, which is the highest.

[0263] In such a case, the trial parameter determination unit 21 determines (adopts) the parameters (v1, v1, v1, v1) of the candidate trial represented by "candidate 2" as the parameters to be used in the next trial.

[0264] Furthermore, when determining multiple parameters, the trial parameter determination unit 21 may select the required number of candidate trials from the multiple candidate trials in the output information 26 in descending order of the predicted value average.

[0265] (B) Operation The processing in the training phase of the historical data mining processing unit 20 in the prompt adjustment device 1c of the third embodiment configured as described above will be explained according to the flowchart (steps S101 to S106) shown in Figure 26.

[0266] In step S101, the history data mining processing unit 20 creates a trial set by combining two trials (a past trial and a current trial) selected from the multiple past trials included in the past history information 22. It is desirable that the history data mining processing unit 20 creates all combinations of the multiple trials in the past history information 22 as trial sets.

[0267] In step S102, a loop process is started in which the control up to step S104 is repeatedly performed for all trial sets.

[0268] In step S103, the history data mining processing unit 20 compares the accuracy of parameter #1 for the past trial with the accuracy of parameter #2 for the current trial to set a label accuracy difference.

[0269] For example, the historical data mining processing unit 20 sets the label accuracy difference to 1 if the accuracy of parameter #2 is higher than the accuracy of parameter #1, and to 0 if the accuracy of parameter #2 is equal to or lower than the accuracy of parameter #1.

[0270] In step S104, the history data mining processing unit 20 outputs one row of data combining the parameter #1, the parameter #2, and the label accuracy difference as one element of the training data 24.

[0271] In step S105, the loop end process corresponding to step S102 is performed. When the process for all trial sets is completed, the process proceeds to step S106.

[0272] In step S106, the history data mining processing unit 20 uses the created training data 24 to create a mining model 23. Thereafter, the process ends.

[0273] Next, the process in the prediction phase of the prompt adjustment device 1c according to the third embodiment will be described with reference to the flowchart (steps S81 to S91) shown in FIG.

[0274] First, in the processes of steps S81 to S86 below, the prediction data 25 is created. In step S81, the history data mining processing unit 20 generates multiple parameter candidates for a candidate trial using a known method. The history data mining processing unit 20 may generate multiple candidate parameters by narrowing down the parameter candidates based on interactions, for example.

[0275] In step S82, a loop process is started in which the control of step S84 is repeatedly performed for all candidate parameters.

[0276] In step S83, a loop process is started in which the control up to step S84 is repeatedly performed for all trial pairs (pairs of past trials and current trials).

[0277] In step S84, the history data mining processing unit 20 creates elements of the prediction data 25 by combining the parameters of the past trial (parameter #3) and the parameters of the current trial (parameter #4).

[0278] In step S85, a loop end process corresponding to step S83 is performed. When the process for all trial sets is completed, the process proceeds to step S86.

[0279] In step S86, loop end processing corresponding to step S82 is performed. When processing for all candidate parameters is completed, the process proceeds to step S88.

[0280] In step S87, the history data mining processing unit 20 inputs the prediction data 25 into the mining model 23 and obtains a predicted value of the accuracy of the candidate trial (fourth trial) for each trial set (each element) of the prediction data 25. The history data mining processing unit 20 may register the predicted value output from the mining model 23 in the output information 26.

[0281] In step S88, a loop process is started in which the control of step S89 is repeatedly performed for all candidate parameters.

[0282] In step S89, the history data mining processing unit 20 calculates the average of the predicted values ​​for each candidate trial based on the predicted values ​​calculated by the mining model 23. The history data mining processing unit 20 may register the calculated average of the predicted values ​​in the output information 26.

[0283] In step S90, a loop end process corresponding to step S88 is performed. When the process for all candidate parameters is completed, the process proceeds to step S91.

[0284] In step S91, the trial parameter determination unit 21 determines the parameter (candidate parameter) of the candidate trial having the highest average predicted value, i.e., the highest accuracy, as the parameter of the next trial.

[0285] (C) Effects According to the prompt adjustment device 1c of the third embodiment, the same effects as those of the second embodiment described above can be obtained.

[0286] Furthermore, the history data mining processing unit 20 trains the mining model 23 using training data 24 that associates the parameters of two trials (first trial, second trial) selected from a plurality of past trials included in the past history information 22 with the label accuracy difference between these trials. Then, prediction data 25 that combines past parameters of the past trials with candidate parameters of a candidate trial is input to this mining model 23, thereby obtaining a predicted value for the accuracy of the candidate trial.

[0287] This allows the candidate trial with the highest predicted accuracy to be selected from among multiple candidate trials, and allows prompts that can obtain highly accurate results from the LLM to be generated in a short time, thereby reducing the computation time and cost required to generate the prompts.

[0288] Furthermore, since parameters (combinations) that improve accuracy can be output, prompt optimization can be efficiently achieved.

[0289] Furthermore, the past history information 22 manages each parameter and each accuracy of the two trials that were actually performed. Then, based on this past history information 22, the history data mining processing unit 20 creates training data 24 that associates each combination of parameters from the two past trials with the difference in accuracy between them. By inputting prediction data 25 that combines past parameters and candidate parameters into a mining model 23 created using such training data 24, it is possible to predict how much the accuracy of the candidate parameters will improve compared to the past parameters. Therefore, it is possible to easily identify candidate parameters that are expected to improve accuracy.

[0290] (IV) Others The disclosed technology is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of each embodiment. The configurations and processes of each embodiment can be selected or combined as needed.

[0291] For example, in the prompt adjustment device 1b of the second embodiment described above, at least one of the pre-processing unit 15 and the post-processing unit 16 may be excluded from the configuration.

[0292] Furthermore, in the second embodiment described above, the post-processing unit 16 performs a statistical test to narrow down the interactions, but the present invention is not limited to this.

[0293] For example, the post-processing unit 16 may extract multiple (n) interactions using a specific method and include only the n interactions in the Bayesian optimization term.

[0294] As a specific method for extracting the n interactions, the post-processing unit 16 may, for example, randomly extract the n interactions.

[0295] Fig. 19 shows a modification of the flowchart shown in Fig. 17. The post-processing unit 16 may execute the process of the flowchart shown in Fig. 19 instead of the process shown in the flowchart of Fig. 17.

[0296] In the flowchart shown in FIG. 19, the post-processing unit 16 randomly extracts n interactions in step S51, and then ends the process.

[0297] The post-processing unit 16 may also calculate the weight of each interaction by logistic regression, extract m interactions in order of absolute values ​​of the coefficients, and randomly extract n interactions from these m interactions. Note that the m interactions extracted in order of absolute values ​​of the coefficients are important interactions for accuracy prediction.

[0298] Fig. 20 shows another modification of the flowchart shown in Fig. 17. The post-processing unit 16 may execute the process of the flowchart shown in Fig. 20 instead of the process shown in the flowchart of Fig. 17.

[0299] In the flowchart shown in FIG. 20, in step S61, the post-processing unit 16 calculates the weight of each interaction by logistic regression, and extracts m interactions in order of absolute values ​​of the coefficients.

[0300] Then, in step S62, the post-processing unit 16 randomly extracts n interactions from the m interactions, and then ends the process.

[0301] In each of the above-described embodiments, the formulation processing unit 13 may use a Lehmer code to represent the order of prompts from the beginning in the formulation information.

[0302] Fig. 21 is a diagram showing a modified example of the formulation information, in which the attribute of order is expressed using Lehmer code.

[0303] In the formulation information shown in FIG. 21, the attribute is an instruction block, a question block, and a format block, and the possible values ​​are shown in correspondence with the values ​​that can be taken as Lehmer codes.

[0304] This allows the possible values ​​for each attribute (order attribute) to be expressed in Lehmer code (see symbol P1). By expressing the order of attributes in Lehmer code, for example, even if the Instruction block is 0, the order of other blocks can also be 0, and the variables become independent. For example, Bayesian optimization requires that variables be independent, but applying Lehmer code ensures independence, making it applicable to optimization.

[0305] 21, in addition to the above-mentioned Instruction block, Format block, and Question block, the length of the explanation for the Instruction block, the length of the explanation for the Format block, and the length of the explanation for the Question block are also set as attributes. The possible values ​​for the explanation length for the Instruction block, the length of the explanation for the Format block, and the length of the explanation for the Question block are short and long.

[0306] In the second embodiment described above, the preprocessing unit 15 lists single attributes or combinations of attributes with low accuracy (poor performance), but the present invention is not limited to this. The preprocessing unit 15 may also list single attributes or combinations of attributes with high accuracy (good performance).

[0307] The configuration of the prompt is not limited to the example shown in Fig. 3. Furthermore, the formulation information is not limited to the examples shown in Figs.

[0308] In the prompt adjustment device 1c of the third embodiment described above, the past history information 22 is stored by correlating each combination with its accuracy for multiple combinations determined by the optimization processing unit 18 described in detail in the second embodiment, and the history data mining processing unit 20 performs data mining on this past history information 22, but this is not limited to this.

[0309] As the past history information 22, for example, a plurality of combinations determined based on the prompt design information extracted (generated) by the optimization unit 12a described in detail in the first embodiment may be stored in association with each combination and its accuracy, and the history data mining processing unit 20 may perform data mining on this past history information 22. In other words, the prompt adjustment device 1a of the first embodiment and the prompt adjustment device 1c of the third embodiment may be combined and implemented.

[0310] Furthermore, in the third embodiment described above, the history data mining processing unit 20 creates a combination of multiple trials (trial sets), as shown in FIG. 26, and then creates training data 24 by setting the label accuracy difference through loop processing, and inputs the training data 24 created in this way into the mining model 23, but this is not limited to this.

[0311] Each time a trial is performed, the historical data mining processing unit 20 may create training data 24 by combining parameter #1 of the first trial (past trial), parameter #2 of the third trial (current (latest) trial), and the accuracy difference between them, and input this data into the mining model 23.

[0312] Furthermore, in the third embodiment described above, the history data mining processing unit 20 calculates the average of the predicted values ​​for each candidate trial based on the predicted values ​​calculated by the mining model 23 and registers the average in the output information 26, but this is not limitative. For example, instead of the average of the predicted values, the "average value of the probability that accuracy will improve for parameter #3" may be used.

[0313] Furthermore, the prompt adjustment device 1c of the third embodiment described above includes an optimization unit 12c that includes a history management unit 19, a history data mining processing unit 20, and a trial parameter determination unit 21 instead of the preprocessing unit 15 of the optimization unit 12b of the second embodiment, but is not limited to this. The optimization unit 12c of the third embodiment may also include the preprocessing unit 15, and various modifications can be made.

[0314] Furthermore, the above disclosure enables those skilled in the art to implement and manufacture each embodiment.

[0315] 1a, 1b, 1c Prompt adjustment device 10 Computer 10a Processor 10b Graphics processing device 10c Memory 10d Storage unit 10e IF unit 10f IO unit 10g Reading unit 10h Program 10i Recording medium 10j Bus 11 Input processing unit 12a, 12b, 12c Optimization unit 13 Formulation processing unit 14 Interaction extraction unit 15 Preprocessing unit 16 Postprocessing unit 17 Output processing unit 18 Optimization processing unit 19 History management unit 20 History data mining processing unit 21 Trial parameter determination unit 22 Past history information 23 Mining model 24 Training data 25 Prediction data 26 Output information

Claims

1. A prompt adjustment program that causes a computer to execute the following process: generate a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that make up the prompt and the possible values ​​of each item; and determine, as an adjustment item for adjusting the prompt, an item included in the combination used to generate the sample prompt with the highest inference accuracy as a result obtained by inputting the plurality of sample prompts into a language model.

2. A prompt adjustment program that causes a computer to execute the following processes: generate a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that make up the prompt and the possible values ​​of each item; perform a first data mining process using as input information that associates the results obtained by inputting the plurality of sample prompts into a language model with the combinations used to generate the sample prompts; and search for adjustment items for adjusting the prompts by optimization based on the results of the first data mining.

3. The prompt adjustment program of claim 2, characterized in that the computer is caused to execute a process of skipping the first data mining for configuration information in which the combination contained in the configuration information satisfies an exclusion condition.

4. The prompt adjustment program of claim 3, characterized in that the computer is caused to execute a process of generating a second sample prompt using sample data for a combination that satisfies the exclusion condition before skipping the first data mining for configuration information including a combination that satisfies an exclusion condition, and skipping the first data mining for the combination that satisfies the exclusion condition if the accuracy of the result obtained by inputting the second sample prompt into the language model is equal to or lower than a threshold value.

5. A prompt adjustment program as described in any one of claims 2 to 4, characterized in that the computer is caused to execute a process of excluding from the optimization the results of the first data mining for which a value representing the impact that the results of the first data mining have on inference performance exceeds a threshold value.

6. The prompt adjustment program according to claim 2, characterized in that the process of searching for the adjustment items includes a process of explicitly narrowing down a search space of the optimization acquisition function using the results of the first data mining and selecting the adjustment items from within the narrowed down range.

7. The prompt adjustment program according to claim 1 or 2, characterized in that the item represents the position of the block in the prompt.

8. The prompt adjustment program according to claim 1 or 2, characterized in that the items represent expression variations of the blocks in the prompt.

9. A prompt adjustment method, characterized in that a computer executes the process of: generating a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that constitute the prompt and the possible values ​​of each item; and determining, as an adjustment item for adjusting the prompt, an item included in the combination used to generate the sample prompt that has the highest inference accuracy as a result obtained by inputting the plurality of sample prompts into a language model.

10. A method for adjusting a prompt, characterized in that a computer executes the following processes: generating a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that make up the prompt and the possible values ​​of each item; performing a first data mining process using as input information that associates the results obtained by inputting the plurality of sample prompts into a language model with the combinations used to generate the sample prompts; and searching for adjustment items for adjusting the prompts by optimization based on the results of the first data mining.

11. The prompt adjustment method according to claim 10, characterized in that the computer executes a process of skipping the first data mining for configuration information that includes a combination that satisfies an exclusion condition.

12. The prompt adjustment method of claim 11, characterized in that the computer executes a process of generating a second sample prompt using sample data for a combination that satisfies the exclusion condition before skipping the first data mining for configuration information including a combination that satisfies an exclusion condition, and skipping the first data mining for the combination that satisfies the exclusion condition if the accuracy of the result obtained by inputting the second sample prompt into the language model is equal to or lower than a threshold value.

13. A prompt adjustment method as described in any one of claims 10 to 12, characterized in that the computer executes a process of excluding from the optimization any result of the first data mining for which a value representing the impact of the result of the first data mining on inference performance exceeds a threshold value.

14. The prompt adjustment method of claim 10, wherein the process of searching for adjustment items includes a process of explicitly narrowing down a search space of the optimization acquisition function using the results of the first data mining, and selecting the adjustment items from within the narrowed down range.

15. The method of adjusting a prompt according to claim 9 or 10, characterized in that the item represents the position of the block in the prompt.

16. The method of adjusting a prompt according to claim 9 or 10, characterized in that the items represent expression variations of the blocks in the prompt.

17. An information processing device comprising: a control unit that executes a process of generating a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that constitute the prompt and the possible values ​​of each item; and determining, as an adjustment item for adjusting the prompt, an item included in the combination used to generate the sample prompt with the highest inference accuracy as a result obtained by inputting the plurality of sample prompts into a language model.

18. An information processing device comprising: a control unit that executes the following processes: generating a plurality of sample prompts corresponding to a plurality of combinations based on configuration information that represents the configuration of a prompt as a combination of a plurality of items corresponding to a plurality of blocks that constitute the prompt and the possible values ​​of each item; performing a first data mining process using as input information that associates the results obtained by inputting the plurality of sample prompts into a language model with the combinations used to generate the sample prompts; and searching for adjustment items for adjusting the prompts by optimization based on the results of the first data mining.

19. The information processing device according to claim 18, characterized in that the control unit executes a process of skipping the first data mining for configuration information that includes a combination that satisfies an exclusion condition.

20. The information processing device according to claim 19, characterized in that the control unit executes a process in which, before skipping the first data mining for configuration information including a combination that satisfies an exclusion condition, a second sample prompt is generated using sample data for the combination that satisfies the exclusion condition, and if the accuracy of the result obtained by inputting the second sample prompt into the language model is equal to or lower than a threshold, the control unit executes a process in which the control unit executes the process of skipping the first data mining for the combination that satisfies the exclusion condition.

21. An information processing device as described in any one of claims 18 to 20, characterized in that the control unit executes a process of excluding from the optimization a result of the first data mining for which a value representing the impact of the result of the first data mining on inference performance exceeds a threshold value.

22. The information processing device according to claim 18, characterized in that, in the process of searching for the adjustment items, the control unit explicitly narrows down the search space of the optimization acquisition function using the results of the first data mining, and selects the adjustment items from within the narrowed down range.

23. An information processing apparatus according to claim 17 or 18, characterized in that the item represents the position of the block in the prompt.

24. An information processing device according to claim 17 or 18, characterized in that the items represent expression variations of the blocks in the prompt.

25. A prompt adjustment program that causes a computer to execute a process of performing a second data mining using as input a first result obtained by inputting a first sample prompt into a language model and first information that associates a first combination of possible values ​​for each of the multiple items used to generate the first sample prompt, and second result obtained by inputting a second sample prompt into a language model and second information that associates a second combination of possible values ​​for each of the multiple items used to generate the second sample prompt.

26. The prompt adjustment program described in claim 25, characterized in that the process of performing the second data mining involves training a machine learning model using the first combination and the second combination as input data and comparison information representing a comparison result between the first result and the second result as correct answer data.

27. The prompt adjustment program according to any one of claims 1 to 4, characterized in that the process of generating the plurality of sample prompts generates the plurality of sample prompts including a first sample prompt and a second sample prompt according to the plurality of types of combinations based on the configuration information, and further performs a second data mining process using as inputs a first result obtained by inputting the first sample prompt into a language model and first information that associates a first combination of possible values ​​for the plurality of items used in generating the first sample prompt, and a second result obtained by newly inputting the second sample prompt into a language model and second information that associates a second combination of possible values ​​for the plurality of items used in generating the second sample prompt.

28. The prompt adjustment program described in claim 27, characterized in that the process of performing the second data mining involves training a machine learning model using the first combination and the second combination as input data and comparison information representing a comparison result between the first result and the second result as correct answer data.

29. A prompt adjustment program that causes a computer to execute a process of obtaining a prediction accuracy for the fourth combination candidate by inputting into a machine learning model a third combination of possible values ​​for the multiple items used to generate a third sample prompt, and a fourth combination candidate of possible values ​​for the multiple items to be used to generate a fourth sample prompt.

30. The prompt adjustment program of any one of claims 1 to 4, characterized in that the process of generating the plurality of sample prompts generates the plurality of sample prompts including a third sample prompt corresponding to the plurality of types of combinations based on the configuration information, and further causes the computer to execute a process of obtaining a prediction accuracy for the fourth combination candidate by inputting into a machine learning model a third combination of possible values ​​for the plurality of items used in generating the third sample prompt, and a fourth combination candidate of possible values ​​for the plurality of items to be used in generating a fourth sample prompt.

31. A method for prompt adjustment, characterized in that a computer executes a process of performing a second data mining using as input a first result obtained by inputting a first sample prompt into a language model and first information that associates a first combination of possible values ​​for a plurality of items used to generate the first sample prompt, and second result obtained by inputting a second sample prompt into a language model and second information that associates a second combination of possible values ​​for a plurality of items used to generate the second sample prompt.

32. The prompt adjustment method described in claim 31, characterized in that the second data mining process trains a machine learning model using the first combination and the second combination as input data and comparison information representing a comparison result between the first result and the second result as correct answer data.

33. The prompt adjustment method according to any one of claims 9 to 12, characterized in that the computer executes the process of generating the plurality of sample prompts by generating the plurality of sample prompts, including a first sample prompt and a second sample prompt, according to the plurality of types of combinations, based on the configuration information, and further performing a second data mining process using as inputs a first result obtained by inputting the first sample prompt into a language model and first information that associates a first combination of possible values ​​for the plurality of items used in generating the first sample prompt, and a second result obtained by newly inputting the second sample prompt into a language model and second information that associates a second combination of possible values ​​for the plurality of items used in generating the second sample prompt.

34. The prompt adjustment method described in claim 33, characterized in that the second data mining process trains a machine learning model using the first combination and the second combination as input data and comparison information representing a comparison result between the first result and the second result as correct answer data.

35. A prompt adjustment method, comprising: a computer executing a process of inputting a third combination of possible values ​​for a plurality of items used to generate a third sample prompt; and a fourth combination of possible values ​​for a plurality of items to be used to generate a fourth sample prompt into a machine learning model, thereby obtaining a prediction accuracy for the fourth combination candidate.

36. The prompt adjustment method according to any one of claims 9 to 12, characterized in that the process of generating the plurality of sample prompts is performed by the computer to generate the plurality of sample prompts including a third sample prompt corresponding to the plurality of types of combinations based on the configuration information, and further to obtain a prediction accuracy for the fourth combination candidate by inputting into a machine learning model a third combination of possible values ​​for the plurality of items used in generating the third sample prompt, and a fourth combination candidate of possible values ​​for the plurality of items to be used in generating a fourth sample prompt.

37. An information processing device comprising a control unit that executes a process of performing a second data mining process using as input a first result obtained by inputting a first sample prompt into a language model and first information that associates a first combination of possible values ​​with a plurality of items used to generate the first sample prompt, and second information that associates a second result obtained by inputting a second sample prompt into a language model and a second combination of possible values ​​with a plurality of items used to generate the second sample prompt.

38. The information processing device according to claim 37, characterized in that, in the process of performing the second data mining, the control unit trains a machine learning model using the first combination and the second combination as input data and comparison information representing a comparison result between the first result and the second result as correct answer data.

39. The information processing device according to any one of claims 17 to 20, characterized in that, in the process of generating the plurality of sample prompts, the control unit generates the plurality of sample prompts including a first sample prompt and a second sample prompt corresponding to the plurality of types of combinations based on the configuration information, and performs a second data mining process using as inputs a first result obtained by inputting the first sample prompt into a language model and first information that associates a first combination of possible values ​​for the plurality of items used in generating the first sample prompt, and a second result obtained by newly inputting the second sample prompt into a language model and second information that associates a second combination of possible values ​​for the plurality of items used in generating the second sample prompt.

40. The information processing device of claim 39, wherein the control unit, in the process of performing the second data mining, trains a machine learning model using the first combination and the second combination as input data and comparison information representing a comparison result between the first result and the second result as correct answer data.

41. An information processing device comprising: a control unit that executes a process of obtaining a prediction accuracy for the fourth combination candidate by inputting into a machine learning model: a third combination of possible values ​​for a plurality of items used to generate a third sample prompt; and a fourth combination candidate of possible values ​​for a plurality of items to be used to generate a fourth sample prompt.

42. The information processing device according to any one of claims 17 to 20, characterized in that, in the process of generating the plurality of sample prompts, the control unit generates the plurality of sample prompts including a third sample prompt corresponding to the plurality of types of combinations based on the configuration information, and obtains a prediction accuracy for the fourth combination candidate by inputting into a machine learning model a third combination of possible values ​​for the plurality of items used in generating the third sample prompt, and a fourth combination candidate of possible values ​​for the plurality of items to be used in generating a fourth sample prompt.

Citation Information

Patent Citations

  • Method of generating response using utterance and apparatus therefor

    JP2023073220A

  • Text generation device and text generation method

    JP7325152B1

  • System and method with entity type clarification for fine-grained factual knowledge retrieval

    US20230316001A1

  • Computer implemented methods for the automated analysis or use of data, including use of a large language model

    US20230316006A1

  • Automatic prompt generation and optimization method for Chinese large-scale language model

    CN116522926A