Methods, computer programs, computer systems (determining improved parameter values ​​for running large-scale language models using machine learning)

The method optimizes LLM parameter settings using a machine learning module to improve computational efficiency and response quality in question and answer tasks, addressing inefficiencies in existing LLM operations.

JP2026089693APending Publication Date: 2026-06-01INTERNATIONAL BUSINESS MACHINE CORPORATION

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2025-11-20
Publication Date
2026-06-01

AI Technical Summary

Technical Problem

The use of Large Language Models (LLMs) for question and answer tasks is inefficient due to the time and computational resources required to find optimal parameter settings for additional input content and operating modes, especially with retrieval-augmented generation methods.

Method used

A method and system for determining improved parameter values for LLMs by loading questions and answers, generating document chunks and input content, and using a machine learning module to optimize parameter settings based on score comparisons and training datasets.

Benefits of technology

This approach accelerates the search for optimal parameter values, improving LLM performance by reducing computational steps and enhancing the quality of responses, particularly in text-based and multimedia question and answer tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026089693000001_ABST
    Figure 2026089693000001_ABST
Patent Text Reader

Abstract

A method for determining improvement parameter values ​​for running a Large-Scale Language Model (LLM). [Solution] The method includes generating chunks of a document and input content dependent on those chunks according to a plurality of different modes, where each mode is specified by a set of values ​​for a parameter; generating prompts for an LLM depending on the input content and each question; providing the prompts as input to the LLM and receiving each tentative answer as a response from the LLM; and performing a comparison between the tentative answers and the target answers to obtain a score for each question as a result; training a machine learning module using the set of parameter values ​​and the scores for the questions as training data; and using the trained ML module to search for improved values ​​for the parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital computer systems, and more specifically, to a method for determining an improved set of values for parameters for operating a Large Language Model (LLM).

Summary of the Invention

Problems to be Solved by the Invention

[0002] The use of LLM for question and answer tasks is becoming more common. In addition to questions of interest, additional input content can be provided to the prompts for LLM. For example, there are many possible variations in how additional input content can be obtained using a retrieval-augmented generation method (RAG), and how the LLM can be instructed. Therefore, it may take time and require an excessive amount of computational resources to find the optimal setting of parameter values for defining how additional input content can be obtained and / or for defining the operating mode of the LLM.

Means for Solving the Problems

[0003] As described by the subject matter of the independent claims, various embodiments provide a method for determining an improved set of values for parameters for operating an LLM, a computer program product, and a computer system. Advantageous embodiments are described in the dependent claims. Embodiments of the present invention can be freely combined with each other if they are not mutually exclusive.

[0004] In one embodiment, the present invention relates to a method for determining an improved set of values ​​for parameters to operate a large-scale language model (LLM) depending on a set of documents. The method includes the step of loading a set of questions and answers. Each of the answers corresponds to one of the questions. The questions relate to content provided by the set of documents. The method further includes the step of performing an iteration. Performing the iteration includes using each mode. Furthermore, performing the iteration includes generating a set of chunks of the documents depending on the set of documents, and generating input content depending on the set of chunks of the documents. The generation of the chunks and the input content is performed according to each mode. Each mode is different for each iteration and is specified by each set of values ​​for the parameters. The values ​​for each parameter specify how the generation of the chunks or the input content is performed.

[0005] Furthermore, the execution of the iteration includes generating a prompt for the LLM depending on the input content and one of the questions.

[0006] Furthermore, the execution of the iteration includes providing the prompt as input to the LLM and receiving each provisional response as a response from the LLM.

[0007] Furthermore, the execution of the iteration includes performing a comparison between the respective answers and the respective provisional answers corresponding to a certain question that the prompt has generated on its behalf, and as a result obtaining a score for the question.

[0008] The method further includes the steps of generating a total score for each of the modes depending on the resulting score for the question, and generating a training dataset for each of the modes, including the set of values ​​for the parameters that specify each of the modes, and the total score for each of the modes.

[0009] The method further includes a step of training a machine learning module (ML module) using the set of values ​​of the parameters of the training dataset as input, and using the overall score of the training dataset as the target output of the ML module.

[0010] The method further includes the step of using the trained ML module to search for the set of improved values ​​for the parameter, wherein the set of improved values ​​for the parameter is input to the trained ML module, resulting in an improved score greater than the maximum overall score of the training dataset.

[0011] In one embodiment, the present invention relates to a computer program product comprising a computer-readable storage medium in which computer-readable program code is embodied, the computer-readable program code is configured to implement a procedure for determining an improved set of values ​​for parameters to operate a large-scale language model (LLM) depending on a set of documents, the procedure including a procedure for loading a set of questions and answers, where each answer corresponds to one of the questions, and the questions relate to content provided by the set of documents; a procedure for performing iterations, where the iterations use each mode; generating a set of each chunk of the documents depending on the set of documents, and generating input content depending on the set of chunks of the documents, where The generation of the set of chunks and the input content is performed according to the respective modes, each of which differs for each iteration and is specified by a set of values ​​for each parameter, the value of which specifies how the generation of the set of chunks or the input content is performed; generating a prompt for the LLM depending on the input content and a question among the questions; providing the prompt as input to the LLM and receiving each tentative answer as a response from the LLM; and performing a comparison between the respective answers and each tentative answer corresponding to a question generated depending on the prompt, thereby obtaining a score for the question. The procedure further includes: generating a total score for each of the modes depending on the resulting score for the question; generating a training dataset for each of the modes, including the set of values ​​for the parameters that specify each of the modes, and the total score for each of the modes; training a machine learning module (ML module) using the set of values ​​for the parameters of the training dataset as input and the total score of the training dataset as the target output of the ML module; and using the trained ML module to perform a search for the improved set of values ​​for the parameters, wherein the improved set of values ​​for the parameters is input to the trained ML module, resulting in an improved score greater than the maximum total score of the training dataset.

[0012] In one embodiment, the present invention relates to a computer system for determining an improved set of parameter values ​​for generating improved prompts for a Large-Scale Language Model (LLM) depending on a set of documents, wherein the computer system includes a set of processors; a computer-readable storage medium; and program instructions stored on the computer-readable storage medium, the program instructions causing the set of processors to perform a procedure. The computer system is configured to load a set of questions and answers, where each answer corresponds to one of the questions, and the questions relate to content provided by the set of documents. Furthermore, the computer system is configured to perform iterations, the iterations including using each mode.

[0013] The execution of an iteration includes using each mode. Furthermore, the execution of an iteration includes generating a set of chunks of each document depending on the set of documents, and generating input content depending on the set of chunks of the documents. The generation of the set of chunks and the input content is performed according to the mode. The mode is different for each iteration and is specified by a set of values ​​for each of the parameters. The values ​​for each of the parameters specify how the generation of the set of chunks or the input content is performed.

[0014] Furthermore, the execution of the iteration includes generating a prompt for the LLM depending on the input content and one of the questions.

[0015] Furthermore, the execution of the iteration includes providing the prompt as input to the LLM and receiving each provisional response as a response from the LLM.

[0016] Furthermore, the execution of the iteration includes performing a comparison between the answer and the provisional answer corresponding to a certain question that the prompt has generated on its behalf, and as a result obtaining a score for the question.

[0017] Furthermore, the computer system is configured to generate an overall score for each of the modes, depending on the score obtained as a result of the question.

[0018] Furthermore, the computer system is configured to generate a training dataset for each mode, which includes the set of values ​​for each of the parameters that specify the mode, and the respective overall scores.

[0019] Furthermore, the computer system is configured to train a machine learning module (ML module) using, as an input, the set of values of the parameters of the training data set and using, as a target output of the ML module, the overall score of the training data set.

[0020] Furthermore, the computer system is configured to use the trained ML module to perform a search for the set of improved values of the parameters, where inputting the set of improved values of the parameters into the trained ML module results in an improved score greater than the maximum overall score of the training data set.

Brief Description of the Drawings

[0021] Hereinafter, embodiments of the present invention will be described in more detail by way of example only with reference to the following drawings.

[0022] [Figure 1] A flowchart of a method for determining a set of improved values of parameters for operating a large language model (LLM) using a set of questions and answers according to an example of the present subject matter.

[0023] [Figure 2] Illustrates the generation of chunks dependent on a set of documents.

[0024] [Figure 3] Depicts a first set of chunks of the set of chunks shown in FIG. 2 and embedding vectors representing the chunks of the first set of chunks.

[0025] [Figure 4] Shows questions of a set of questions and question embedding vectors representing the questions.

[0026] [Figure 5]This diagram illustrates the use of LLM to generate a provisional answer based on prompts.

[0027] [Figure 6] This diagram depicts a table showing the parameter values ​​used to control the generation of the prompts shown in Figure 5 and the operation of the LLM to generate the provisional answers shown in Figure 5.

[0028] [Figure 7] This shows the training dataset for training the machine learning module (machine-learning (ML) module).

[0029] [Figure 8] Figure 7 illustrates a trial set for performing a search for the improved set of parameter values ​​shown in Figure 6 using the trained ML module shown.

[0030] [Figure 9] This is a flowchart illustrating how to use a trained ML module to obtain new answers to new questions related to the document shown in Figure 2.

[0031] [Figure 10] This describes a computing environment using an example related to this subject. [Modes for carrying out the invention]

[0032] The descriptions of various embodiments of the present invention are presented for illustrative purposes only and are not intended to be exhaustive or limitful to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best describe the principles, practical applications, or technological improvements over the technologies available on the market of the embodiments, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

[0033] A set of improved parameter values ​​(hereinafter also referred to as improved values) may enable the generation of improved prompts for the LLM depending on the improved values, and the reception of improved responses from the LLM in response to inputting the improved prompts into the LLM (hereinafter referred to as improved use). Improved use may include generating a new set of chunks of documents depending on a set of documents, and generating new input content depending on the new set of chunks. The new set of chunks and the new input content may be generated according to a new mode, which may be specified by the set of improved parameter values. Each value in the set of improved parameter values ​​may specify how the new set of chunks or the new input content is generated.

[0034] According to one modified version of the improved use, a new prompt for the LLM may be generated depending on new input content and a selected question from a set of questions. For example, the new prompt may include new input content and a selected question. The new prompt may be provided as new input to the LLM, and a new tentative answer may be received as a response from the LLM. In most cases, the quality of the new tentative answer, e.g., accuracy, may be higher than any tentative answer generated in the iteration, because the improved score is greater than the maximum overall score of the training dataset. Therefore, according to this modification, the new tentative answer may be considered an improved answer compared to the tentative answer obtained in the iteration. Thus, the performance of the LLM may be improved.

[0035] According to another variation of the improved use, a new prompt for the LLM may be generated depending on new input content and on a new question that may be relevant to the document instead of using a selected question. For example, the new prompt may include new input content and a new question. Similar to the variation above, the new prompt may be provided as new input to the LLM, and a new answer may be received as a response from the LLM. In most cases, the improved score may be greater than the maximum overall score of the training dataset, and therefore the quality of the new answer, e.g., accuracy, may be higher than when using the LLM to obtain an underlying new answer, from which the new input would be generated depending on one of the chunks of document produced in one of the iterations. Thus, the performance of the LLM may be enhanced when the LLM is used to generate new answers to new questions.

[0036] By using a trained ML module when performing improved value searches, the improved value search can be accelerated compared to using an LLM for the improved value search. Using an LLM for the improved value search would imply performing the above iterations to re-evaluate each trial of the set of parameter values. However, generating a set of document-dependent chunks, and generating input content (commonly known as the information acquisition process), can require a number of computational steps that may be about 10 or 100 or even 1000 times greater than the number of computational steps required to re-evaluate each trial of the set of parameter values ​​using a trained ML module. Each trial of the set of parameter values ​​may involve a new combination of parameter values.

[0037] Generally, the fewer computational steps required, the shorter the time needed to evaluate each trial of the parameter value set. Therefore, the search for improved values ​​can be accelerated by using a trained ML module. In addition, if the LLM runs on a first computing device, e.g., a central server or a set of first processors, and the generation of scores for iterations, training of the ML module, and / or the execution of the search are performed on a second computing device, e.g., client devices or a set of second processors, or a single second processor, then using a trained ML module can reduce data traffic between these two computing devices. In this case, only a limited amount of data exchange may be performed to generate the training dataset set, e.g., sending prompts to the first computing device and receiving provisional answers for each training dataset from the first computing device.

[0038] The mode of iteration can differ from iteration to iteration. Considering two modes that can be specified by a set of two parameter values, the difference between these two modes suggests that at least one parameter value in that set of two values ​​is different. Since there are at least two parameters for specifying the mode, iteration can be performed in the form of nested loops. Within each loop of the nested loop, the value of each of these parameters may change each time the loop is iterated.

[0039] In one example, the execution of the innermost loop may include further iterations, i.e., the execution of further loops. Each time a further loop is iterated, a different question (hereinafter referred to as each question) may be selected. Therefore, performing further iterations multiple times, i.e., performing further loops multiple times, may involve the repetition of: generating input content according to the mode; generating prompts for the LLM that depend on the input content and each question; providing the prompts as input to the LLM and receiving each tentative answer as a response from the LLM; and performing a comparison between the answers and tentative answers corresponding to each question that the prompts depended on, resulting in obtaining a score for each question. Performing further iterations may include keeping the mode unchanged. The scores for the questions may be stored to calculate the respective overall score for each mode.

[0040] In one example, questions and answers may be in the form of a question data file and an answer data file, respectively. Questions and answers may contain words or phrases. In one example, a question and / or answer may contain one or more images and / or image sequences and / or one or more audio files. The content provided by a collection of documents may be in the form of words, phrases and sentences of the documents. In one example, a document may contain one or more images and / or image sequences and / or one or more audio files. In this application, LLM is most often used for text-based question and answer (Q&A) tasks. However, since Q&A tasks may involve increased use of images in the future, questions, answers and / or documents may be in the form of media files containing text, images, image sequences and / or audio files. In one example, one or more of the questions, answers and documents may contain only images and / or image sequences and / or audio files.

[0041] The questions may be related to the content of the document, and as a result, at least one word, image, or audio file in each question may have the same meaning as at least one word, image, or audio file in the document.

[0042] A prompt for the LLM may contain a command instructing the LLM to use the input content to generate each tentative answer. For example, the command might be a phrase such as, "Use the input content contained in this prompt to answer the following questions." In most applications, each prompt generated in each iteration may contain the respective question, the input content generated in each iteration, and the command.

[0043] An LLM (Language Learning Module) can be an artificial intelligence (AI) module designed to understand and generate human-like text based on incoming input. An LLM may be configured to generate a provisional response in response to each prompt it receives. An LLM can be trained with vast amounts of text and / or image data from books, articles, and other sources, thereby enabling it to learn language patterns, grammar, facts, and even some logical reasoning abilities. Images for training an LLM can be converted into image embedding vectors, if the LLM is trained with image data. LLM can be designed, for example, in the form of one of OpenAI's generative pre-trained transformer (GPT) models (e.g., GPT model, GPT-2 model, GPT-3®, GPT-4®), or in the form of a BERT model, T5 model, RoBERTa model, XLNet model, GPT-Neo model, GPT-J model, LLAMA® model, BLOOM model, Claude® model, Cohere® model, OPT model, or ERNIE model.

[0044] As described above, the generation of a set of chunks dependent on a document, and the generation of input content dependent on that set of chunks for each iteration, can be called an information retrieval process. In most cases, an information retrieval process may involve the generation of chunks. However, there may be applications in which information retrieval generates input content without generating chunks dependent on a document in some parts of the iteration. In this case, the first part of the iteration may involve the generation of a set of chunks dependent on a document, and the second part of the iteration may use a predefined set of chunks. The predefined set of chunks may be generated in one of the previous iterations, or separately from the iteration.

[0045] The generation of a set of document-dependent chunks as part of the information retrieval process may involve dividing a set of documents into chunks. For example, a chunk may be a proper subset of a document. Alternatively, chunks may have intersections. The information retrieval process may further involve using an embedding model to compute chunk-dependent embedding vectors. Each embedding vector may represent the semantic meaning of each chunk in the set of chunks generated in each iteration.

[0046] The generation of input content for each iteration may involve generating input content dependent on each question. For example, the generation of input content for each iteration may involve comparing an embedding vector with one or more further embedding vectors representing each question. In one example, one further embedding vector may represent each question, which will also be referred to below as the question embedding vector. In this case, the information retrieval process may involve using an embedding model to compute question embedding vectors dependent on each question.

[0047] In one example, it may be possible to determine a set of contextualized embedding vectors for each question based on each question embedding vector. A contextualized embedding vector can be an example of further embedding vectors. In this case, generating input content for each iteration may involve comparing the embedding vectors with one or more vectors from the set of contextualized embedding vectors for each question. In one example, the information retrieval process may involve calculating a concatenated embedding vector, depending on the contextualized embedding vectors and linear factors. Each linear factor may weight each contextualized embedding vector to obtain the concatenated embedding vector. The concatenated embedding vector can be an example of further embedding vectors. In this case, generating input content for each iteration may involve comparing the embedding vectors with the concatenated embedding vector.

[0048] Comparing an embedding vector to each question embedding vector or a concatenated embedding vector may involve calculating a similarity score for each embedding vector. In one example, the similarity score could be cosine similarity. The information retrieval process may further involve ranking the embedding vectors according to their similarity scores and selecting a predefined number of chunks from a set of chunks for each iteration. In one example, the selected chunks might be the chunks that are highest ranked according to their similarity scores. Prompt generation may involve generating input content depending on the selected chunks. In one example, the input content might consist of the selected chunks.

[0049] The score for each question can be an indicator of how similar the provisional answer is to the answer corresponding to each question from which the prompt was generated (hereinafter referred to as the target answer). Hereinafter, a higher score is assumed to indicate better performance of the LLM in producing a provisional answer as similar as possible to the target answer. The indicator could be a performance metric such as an F1 score or accuracy.

[0050] The F1 score is a common metric used to assess the performance of question-answering large-scale language models (LLMs), particularly when partial correctness of responses is acceptable. It measures the overlap between words in the tentative response and words in the target response. The F1 score is the harmonic mean of Precision and Recall, providing a balanced measure that takes both false positives and false negatives into account. Precision refers to the number of correctly predicted words in the tentative response divided by the total number of words in the tentative response. It essentially measures the accuracy of the words provided by the LLM. Recall, on the other hand, is the number of correctly predicted words in the tentative response divided by the total number of words in the target response, assessing the number of relevant words that the model successfully retrieved. The F1 score is calculated using the formula F1 score = 2 × (Precision × Recall) / (Precision + Recall). A high F1 score indicates a large overlap between the tentative and target responses, which may suggest that the LLM is performing well in terms of obtaining relevant and accurate information. This metric can generally be valuable in Q&A tasks where perfect agreement is not required and partial correctness can still be useful.

[0051] Other examples show that metrics can be Exact Match (EM), Mean Reciprocal Rank (MRR), Human Rating, Precision, Recall, or passage-level metrics such as Precision at k.

[0052] As used herein, the term "module" refers to any known or future-developed hardware, executable program or other software, artificial intelligence, fuzzy logic, or combination thereof, used to perform a function associated with a "module," or as a result of performing a function associated with a "module."

[0053] For example, an ML module may include one or more artificial neural networks and / or one or more decision trees. Training an ML module may involve learning from the training dataset by tuning the ML module as a whole to best represent the patterns, relationships, or structures in the training dataset.

[0054] For example, an ML module may contain a mathematical function with parameters. The functions of an ML module (hereinafter also referred to as ML functions) can be linked to one another. For example, the output of one or more ML functions may be linked to the input of one or more other ML functions. The entire set of ML functions can be considered a compound function. The parameter values ​​of the ML functions may be adaptable to the training dataset. An ML module may be configured such that a cost or loss function can be minimized by adapting the parameter values ​​of the ML functions.

[0055] The cost or loss function may be a function of the deviation of the trained output values ​​of the composite function from the overall score of the training dataset. Training an ML module may involve calculating each trained output value of the composite function depending on the set of parameter values ​​of each training dataset. Training may further involve calculating the cost or loss function based on the deviations, where each deviation may represent the deviation of each trained output value from the overall score of each training dataset. Training may further involve adapting the parameter values ​​of the ML function depending on the cost or loss function, for example, depending on the partial derivatives of the cost or loss function. Training may involve calculating each partial derivative as the derivative of the cost or loss function with respect to each parameter of the ML function. Training may involve performing various iterations of adapting the parameter values ​​of the ML function until the cost or loss function falls below a predefined threshold.

[0056] If the ML module includes a neural network, the parameters of the ML function may be weights indicating the strength of connections between neurons in that neural network. Training may involve performing one or more learning algorithms such as gradient descent, stochastic gradient descent, minibatch gradient descent, momentum, Nesterov-accelerated gradient, Adagrad, RMSProp, Adam, backpropagation, step decay, exponential decay, adaptive learning rate, L1 and L2 regularization, dropout, and early termination, and / or performing initialization methods such as He initialization or Xavier initialization.

[0057] Performing an improved value search involves generating a trial set of parameter values ​​(hereinafter also referred to as the trial set). Each trial set may contain trail values ​​for each parameter. Each trial set may have the same dimensions as the set of values ​​for each parameter in the training dataset. All possible mathematical combinations of parameter values ​​can exist in the vector space (hereinafter also referred to as the search space). In one example, the trial set may contain random values ​​for each parameter. In this case, the search can be a random search.

[0058] In another example, the boundaries of an acceptable region within the search space can be defined. The acceptable region can specify all acceptable combinations of parameter values. Furthermore, the acceptable region can be divided into subregions. The size of these subregions may depend on a predefined resolution value. For example, the center of mass of each subregion can specify each combination of parameter values, thereby specifying each trial set within the trial set.

[0059] The search for improved values ​​may involve using trial sets one by one to generate trial output values ​​for the ML module. The ML module may use an adapted ML function to calculate each trial output value depending on the values ​​in each trial set. The trial set that the ML module uses to calculate the maximum trial output value may be specified as a set of improved values ​​for the parameter.

[0060] Performing iterations to generate scores for questions and overall scores for each mode may involve setting values ​​for each set of parameters (hereinafter also referred to as initialization). Initialization may involve performing a DOE to set values ​​for each set of parameters, for example, according to a full factor design or a partial factor design. In one example, the values ​​for each set of parameters may be set randomly. Alternatively, initialization may involve dividing the allowed region into further subregions and selecting a portion of those subregions. In this case, the center of mass of each selected further subregion may represent a set of parameters. The number of subregions defined to generate the trial set may be greater than the number of further subregions, for example, 10 times, 100 times, or up to 1000 times greater.

[0061] According to one example, the method may further include receiving the aforementioned new questions relating to the set of documents. According to this example, the method may further include generating a new set of chunks of documents depending on the set of documents, and generating new input content depending on the new set of chunks. The generation of the new set of chunks and the new input content may be performed according to the aforementioned new modes. The new modes may be specified by an improved set of parameter values. Each value in the improved set of parameter values ​​may specify how the generation of the new set of chunks or the new input content is performed. According to this example, the method may further include generating a new prompt for the LLM depending on the new input content and a new question. The method may further include providing the new prompt as a new input to the LLM and receiving a new answer from the LLM as a response. The method may further include providing a new answer as a response to a new question. This example illustrates one of the aforementioned variations of the improved use of the LLM.

[0062] For example, an ML module may include decision trees. Training an ML module may include training decision trees. Decision trees can offer several advantages over artificial neural networks (ANNs), primarily due to their interpretability and simplicity. Decision trees are often easy to understand and visualize, provide a structured and clear decision-making process, and differ from the "black box" nature of ANNs. In addition, training decision trees may be computationally intensive, require less data preprocessing, and may perform better with smaller datasets compared to ANNs. Furthermore, decision trees can handle missing values ​​more effectively and can naturally grasp complex relationships between parameters without assuming any linear relationships between parameters and the overall score. Additionally, decision trees can provide insights into the importance of each parameter. Decision trees may essentially ignore irrelevant parameters. By using decision trees, ML modules can be more efficient, robust, and interpretable compared to using ANNs.

[0063] For example, an ML module may include a collection of decision trees. Training an ML module may include generating a proper subset of the training dataset, where the proper subset is a random sample of the training dataset. Training an ML module may further include training each decision tree using each subset of the training dataset. Performing the training of an ML module may further include performing a random forest or gradient boosting tree method.

[0064] Using a set of decision trees instead of a single decision tree can offer several advantages. First, the combined action of multiple trees can average out errors and provide more reliable results, potentially improving the predictive accuracy of the ML module. A single decision tree can overfit to the training dataset. However, a set of decision trees can reduce this overfitting of the ML module to the training dataset, thereby enhancing its generalization ability. Additionally, using a set of decision trees can increase the robustness of the ML module to noise and data variability within the training dataset, and reduce its sensitivity to outliers. A set of decision trees may model complex nonlinear relationships between parameters more effectively than a single decision tree, capturing intricate interactions between parameters. The variance of the ML module is also reduced, resulting in greater stability and less susceptibility to small changes in the training dataset. Furthermore, using a set of trees can provide a more reliable estimate of the importance of each parameter, offering valuable insights into which parameters are likely to have the greatest impact on the overall score.

[0065] For example, performing a search may include performing a grid search, a random search, or a search based on a genetic algorithm. Performing a grid search or a random search may involve dividing the search space into smaller regions to specify a set of trials. Performing a grid search allows all specified combinations of parameter values ​​to be tested. Such a test may involve inputting each set of trials into an ML module and receiving the trial score from the ML module as a response. The set of trials on which the ML module depends to calculate the highest trial score may be specified as a set of improved parameter values. Performing a grid search may increase the likelihood of finding improved values ​​as the optimal set of parameter values. Grid searches may be easy to implement and their results can be reproducible because they are deterministic. Additionally, grid searches can be performed in parallel and, as a result, can be efficient and scalable when using multiple processors or distributed systems.

[0066] Performing searches based on genetic algorithms can increase search efficiency and scalability compared to performing grid searches. Genetic algorithms can efficiently explore large and complex search spaces without evaluating all possible combinations of parameter values, saving time and computational resources. Genetic algorithms can balance exploration and utilization by exploring new combinations of parameter values ​​while focusing on promising areas of the search space using evolutionary techniques. Using genetic algorithms for searches can be more effective in finding the optimal or near-optimal set of parameter values.

[0067] For example, the parameters of a set of parameters may represent the input features of an ML module. Training the ML module may involve determining a feature importance score for each input feature. The feature importance score may indicate the relative impact that each input feature has on the ML module's prediction. In this context, the ML module's prediction may take the form of an overall score generated for each mode. For example, training the ML module may involve running a random forest method, such as a random forest regression method, and as a result obtaining feature importance scores for the parameters.

[0068] Feature importance scores can make it possible to select parameters whose values ​​may change more significantly than those of other parameters when performing a search. For example, performing a search may involve subdividing the acceptable region into smaller regions, depending on the feature importance score of the parameters. In this case, the acceptable region may be subdivided more finely along the dimensions of the search space corresponding to parameters with higher feature importance scores. The higher the feature importance score of a parameter, the higher the resolution of the subregions, and consequently the resolution of the trial set, in the dimension corresponding to each parameter.

[0069] For example, the method may further include selecting the chunk size of the chunks in the set of chunks for each iteration. The values ​​of the parameters in the set of parameters may specify the chunk size for each iteration (hereinafter referred to as the first parameter). Using the first parameter as one of the parameters may have the following advantages: Typically, smaller chunks may improve the retrieval of highly relevant details in a document, while larger chunks may preserve a broader context. By performing a search for improved values ​​and thereby searching for improved values ​​of the first parameter, the opportunity to optimize the chunk size in terms of the balance between information retrieval accuracy and richness of context may be increased.

[0070] For example, the method may further include selecting chunks from a set of chunks and, for each iteration, generating input content depending on the selected chunks. The value of a parameter in a set of parameters may specify the number of chunks selected (hereinafter referred to as the second parameter). As described above, the information acquisition process may include ranking the embedding vectors according to their similarity score and selecting a predefined number of chunks from the set of chunks for each iteration. The predefined number of chunks may be the number of chunks selected, i.e., the value of the second parameter.

[0071] A larger number of selected chunks may allow for the presentation of more document content as input to the LLM. Therefore, using more selected chunks may increase the chances of improving the overall score. However, using more selected chunks may increase computational costs. Therefore, a smaller number of selected chunks may allow for more iterations and the acquisition of a larger training dataset. This may result in a more accurate ML module. For example, a smaller number of selected chunks may allow the LLM to access highly specific and relevant fragments of the document, thus improving information retrieval accuracy. This may reduce the inclusion of irrelevant content in the document. Performing improved value searches, and thereby searching for improved values ​​for the second parameter, may increase the chances of generating the optimal set of training datasets for producing the most accurate ML module possible.

[0072] For example, the method may further include selecting the size of the intersection of the set of chunks for each iteration. The method may further include generating the set of chunks such that, for each iteration, each chunk in the set of chunks contains at least one intersection that is contained in another chunk in the set of chunks. The value of a parameter in the set of parameters may specify the size of the intersection (hereinafter referred to as the third parameter). The intersection, or overlap, makes it possible to consider document information that spans chunk boundaries. The intersection can prevent the loss of significant context that may be necessary to fully understand or interpret a given fragment of document content. However, given that there is a maximum number of tokens for prompts, a smaller intersection size means that more document information can be included in the prompt for each iteration. Therefore, by performing an improved value search and thereby searching for an improved value of the third parameter, there is a greater chance of finding the optimal balance between preserving significant context contained in one or more chunks and including as much information as possible in a new prompt.

[0073] For example, the method may further include, for each iteration, selecting the type of embedding model to generate each embedding vector depending on each chunk in the set of chunks. Each embedding vector may represent each chunk, as described above. The values ​​of the parameters in the set of parameters may specify the type of embedding model (hereinafter referred to as the fourth parameter). The selection of the type of embedding model for each iteration may include selecting the embedding model as one of the models in the list of embedding models. The list of embedding models may include the Word2Vec model, GloVe model, FastText model, ELMo model, BERT model and / or Sentence-BERT model, and transformer-based models such as the GPT model, RoBERTa model and / or DistilBERT model. The value of the fourth parameter may be a natural number.

[0074] Word2Vec and GloVe models are efficient and can capture word relationships well. The FastText model can improve the handling of extralexical words by using subword information. The ELMo model can use contextualized embeddings, which can enable the handling of ambiguity. The BERT model can provide very rich contextualized embeddings suitable for various natural language processing (NLP) tasks. The Sentence-BERT model can be optimized for sentence-level similarity tasks. In general, various embedding models may perform differently depending on the content of the document set. Because the content of the documents may depend on the application of the computer system, one embedding model in the list may perform better than another when applied to a certain industrial field such as automotive, and vice versa when applied to a different industrial field such as mining or medicine. Therefore, by performing a search for improved values ​​and thereby searching for improved values ​​for the fourth parameter, the chances of finding the best embedding model from the models in the list for the information retrieval process using the documents can be increased.

[0075] For example, the method may further include selecting the size of each embedding vector for each iteration. Each value of the parameter in the parameter set may specify the size of each embedding vector (hereinafter referred to as the fifth parameter). Smaller embedding vectors may capture a more general and broader semantic meaning of the document's content, thereby increasing the generalization ability of the computer system. This may be beneficial when the questions used in the iterations may contain significant variability in their context. Larger embedding vectors are better suited to capturing more subtle and specific features, thereby improving the accuracy of tentative answers. This increases the capabilities of the computer system, allowing for the provision of specific tentative answers. By performing a search for improved values ​​and thereby searching for improved values ​​of the fifth parameter, a balance between generalization and specificity can be achieved in light of the document's content. The value of the fifth parameter may change depending on changes in the document's content, i.e., depending on the application domain of the computer system.

[0076] For example, the method may further include selecting the size of the provisional answer for each iteration. The value of a parameter in the set of parameters may specify the size of the provisional answer (hereinafter referred to as the sixth parameter). Selecting the size of the provisional answer may mean restricting the size of the provisional answer to the selected size. The prompt for each iteration may include a command that describes the provisional answer so that it has a maximum size of the selected size.

[0077] The smaller the size of the provisional answer, the more precise the resulting provisional answer may be. The larger the size of the provisional answer, the more detailed and contextually rich the resulting provisional answer may be. By varying the limited size of the provisional answer in iterations, a nearly optimal size for the new answer can be found, based on the search for trained ML modules and improved values, taking into account the set of documents and, consequently, the application domain of the computer system. A balance between conciseness and comprehensiveness can be found by searching for improved values ​​of the sixth parameter with respect to the application of the computer system.

[0078] For example, the method may further include selecting a value for a seventh parameter in the set of parameters for each iteration. The seventh parameter may indicate the degree of repetition of similar contexts in the tentative response. The seventh parameter may also be known as the generation repetition penalty parameter. The higher the value of the seventh parameter, the less likely the LLM is to produce repetitive or redundant content in each tentative response. The LLM may be configured to assign a cost to repetitive words or phrases in the tentative response, depending on the value of the seventh parameter. The higher the value of the seventh parameter, the greater the variation in the content of the tentative response.

[0079] Setting the value of the seventh parameter too high can cause the LLM to avoid necessary repetition, resulting in unnatural or incoherent text. In natural language, some repetition can be important for readability and coherence. If the value of the seventh parameter is too high, the model may be forced to paraphrase or omit important information. This can result in unnatural, confusing, or incomplete content in the provisional and new answers. Therefore, by performing a search for improved values ​​and thereby searching for improved values ​​of the seventh parameter, the LLM can generate a good balance between undesirable redundancy and clarity of content in the provisional and, consequently, the content of the new answers.

[0080] For example, the method may further include selecting a type of LLM for each iteration. The values ​​of the parameters in the set of parameters may specify the selected type of LLM (hereinafter referred to as the eighth parameter). The computer system may be configured to select an LLM as one LLM from a list of LLMs. The list may include GPT models such as the OpenAI GPT-3 model or GPT-4 model, or further LLMs such as the BERT model, LLAMA model, BLOOM model, Claude model, Cohere model, OPT model, or ERNIE model. Depending on the content of the documents, one of the LLMs in the list may work better than the others. Therefore, by performing a search for improved values ​​and thereby searching for improved values ​​of the eighth parameter, the best type of LLM for the set of documents and a given question and answer can be found.

[0081] For example, the method may further include selecting a value for a ninth parameter in the set of parameters for each iteration, where the value of the ninth parameter may indicate the degree of modification of the probability distribution of the LLM. Depending on how the probability distribution used by the LLM to generate the tentative answers is modified, the randomness of the tentative answers may be altered. The ninth parameter may be known as the generation temperature parameter. The ninth parameter may allow adjustment of how deterministically or creatively the LLM can generate tentative and new answers. Therefore, by performing a search for improved values ​​and thereby searching for improved values ​​of the ninth parameter, it may be possible to find a balance between the creativity and reliability of the content of tentative and new answers.

[0082] For example, the method may further include selecting a type of information retrieval method for each iteration. Generating input content may include obtaining input content by performing the selected information retrieval method. The values ​​of the parameters in the set of parameters may specify the type of information retrieval method selected (hereinafter referred to as the 10th parameter). The type of information retrieval method selected may be one of the information retrieval methods in the list, including, for example, the BM25 method, the TF-IDF method, dense vector retrieval methods (e.g., FAISS and ScaNN), semantic search methods, elastic search methods, approximate nearest neighbor (ANN) search methods, hybrid retrieval methods, maximum inner product search (MIPS) methods, graph-based retrieval methods, and memory-augmented methods. BM25 and TF-IDF methods can be considered keyword-based methods that rely on statistical measures of term frequency and document importance. Dense vector information acquisition and semantic retrieval methods use neural network-based embedding models to grasp the semantic meaning of chunks, potentially enabling more context-aware matching. ElasticSearch methods can combine both keyword-based and dense vector retrieval capabilities. Nearest Neighbor (ANN) search methods can find approximate matches between question embedding vectors and embedding vectors. This can be advantageous when the document set is very large and the set of chunks has a large number of chunks. Graph-based information acquisition can organize chunk information within a graph structure to grasp the relationships between chunks. This can enable faster searches based on connectivity or similarity.

[0083] These different information retrieval methods may differ in their approach to matching the questions for each iteration with the document and / or chunks. Depending on which information retrieval method is selected, the input content may be generated more quickly. Furthermore, depending on which of these information retrieval methods is selected, the semantic understanding of the document and / or chunks, and / or the accuracy of the provisional answers, may differ. Therefore, by performing improved value searches and thereby searching for improved values ​​for the tenth parameter, it may be possible to adapt the semantic understanding capabilities of the computer system for generating input content to the content of the document.

[0084] For example, the method may further include, for each iteration, selecting a chunking method from a set of chunking methods for generating chunks depending on the document. The method may further include generating a set of chunks according to the selected chunking method. The values ​​of the parameters in the set of parameters may specify the selected chunking method (hereinafter referred to as the 11th parameter). The chunking methods may differ in how the document can be divided into chunks. For example, these different chunking methods may be designed to divide the document by sentences, paragraphs, total number of fixed words, or based on semantic units. The set of chunking methods may include, for example, fixed-length chunking, sentence-based chunking, paragraph-based chunking, semantic chunking, and / or sliding window chunking.

[0085] Implementing a fixed-length chunking method may involve dividing a document into chunks of a set number of words or characters (e.g., 200 words per chunk) without considering sentence or paragraph boundaries. Implementing a sentence-based chunking method may involve dividing the document so that each chunk contains a complete sentence, often grouping three to five sentences together. Implementing a paragraph-based chunking method may involve dividing the document so that chunks are created based on paragraph divisions, the size of which may vary depending on the required level of detail. Implementing a semantic chunking method may involve dividing the document so that chunks can be in a contextually meaningful form, thereby organizing the document's content by topic or subheadings. Implementing a sliding window chunking method may generate overlaps between chunks by moving a window across the text. Implementing the sliding window method may result in the aforementioned overlaps. Implementing a topic-based chunking method may involve running algorithms such as topic modeling to separate the document into distinctly different themes. In this case, each chunk can correspond to its own cohesive object. Therefore, by performing an improved value search and thereby searching for an improved value for the 11th parameter, it may be possible to find the best chunking method for a set of documents that may change with respect to the application of the computer system.

[0086] For example, the method may further include generating an embedding vector for each iteration, where each embedding vector represents each chunk in the set of chunks. The method may further include generating the aforementioned additional embedding vectors representing the questions. Alternatively, the method may further include generating a set of additional embedding vectors representing the questions, such as contextualized embedding vectors. The method may further include selecting a type of similarity metric. The values ​​of the parameters in the set of parameters may specify the type of similarity metric selected (hereinafter referred to as the 12th parameter). The method may further include performing a comparison between the embedding vectors and the additional embedding vectors using the selected similarity metric. Alternatively, the method may further include performing a comparison between the embedding vectors and the set of additional embedding vectors using the selected similarity metric. The method may further include selecting a subset of chunks depending on the result of the comparison. Performing the comparison may produce a similarity metric value for each chunk. The method may further include creating a ranking list. The ranking list may enumerate chunks according to their values ​​of the similarity metric. The selection of a subset of chunks may involve selecting a predefined number of chunks that contain the highest similarity metric. This predefined number of chunks may be the aforementioned number of selected chunks. The method may further include generating input content depending on the selected subset of chunks.

[0087] Selecting a type of similarity metric may involve choosing a similarity metric from a list of similarity metrics. The list of similarity metrics may include, for example, cosine similarity, Euclidean distance, Manhattan distance, and dot product. By performing an improved value search and thereby searching for improved values ​​for the twelfth parameter, it may be possible to find the most appropriate similarity metric for the set of documents and the question. For example, selecting a type of information retrieval method may involve selecting a type of similarity metric.

[0088] For example, the method may further involve selecting values ​​for parameters of the embedding model to generate each embedding vector depending on each chunk in the set of chunks. The values ​​of the parameters in the set of parameters may specify the values ​​of the parameters of the embedding model (hereinafter referred to as the 13th parameter). For example, the parameters of the embedding model may include which type of query, value, or key vector to use, if the embedding model may contain a set of different queries, values, or key vectors to select. For example, the parameters of the embedding model may be parameters of functions in the embedding model, such as a softmax function for computing contextualized embedding vectors within the embedding model. By performing an improved value search and thereby searching for improved values ​​for the 13th parameter, it may be possible to tune the embedding model to adapt it to the document content and / or questions and / or answers.

[0089] Figure 1 is a flowchart illustrating a method for determining an improved set of parameter values ​​for operating a large-scale language model (LLM) in a manner dependent on a set of documents, as is the case in this subject. For illustrative purposes, the method described in Figure 1 may be implemented in the system illustrated in Figure 10, but is not limited to this implementation of the operational steps performed by the processor set 810 of computer 801.

[0090] The method may include the following steps. The processor set 810 of computer 801 shown in Figure 10 may be configured to execute the steps in the form of operational steps. In step 101, a set of questions and answers may be loaded. The operating system 822 of computer 801 may be configured to load the set of questions and answers into the cache 821 of computer 801. Furthermore, the operating system 822 may be configured to load code for block 900 into the cache 821 for determining an improved set of parameter values ​​for running a large language model according to one of the above modifications. The code for block 900 may include instructions. The instructions may cause the processor set 810 to execute the steps of the method when executing the instructions. Computer 801 containing the code for block 900 may be considered configured to execute the steps of the method. More precisely, computer 801 having a cache 821 storing the code for block 900 may be considered configured to execute the steps of the method.

[0091] Each answer may correspond to one of the questions. The questions may relate to the content provided by the collection of documents. The code in block 900 may contain one or more commands that cause the processor set 810 to execute iteratives when the code in block 900 is executed. Each iteration may contain steps 102, 103, 104, and 105 and may be executed according to their respective modes. The modes may differ for each iteration.

[0092] In step 102, each set of chunks of a document may be generated depending on the set of documents. In addition, in step 102, input content may be generated depending on the set of chunks of documents. The generation of chunk sets and input content may be performed according to a mode. The mode may be specified by a set of values ​​for each parameter. The values ​​of each parameter may specify how the generation of chunk sets or input content is performed.

[0093] In step 103, a prompt for the LLM may be generated depending on the input content and a question among the questions. In step 104, the prompt may be provided as input to the LLM, and each tentative answer may be received as a response from the LLM. In step 105, a comparison may be performed between the answer and the tentative answer corresponding to the question on which the prompt was generated. The comparison between the answer and the tentative answer may result in a score for the question on which the prompt was generated.

[0094] Steps 102, 103, 104, and 105 for each mode are repeated for each question in the set of questions, while keeping each mode in an unmodified state, and as a result, a score for each question (hereinafter referred to as the question score) may be obtained.

[0095] In step 106, the overall score for each mode may be generated depending on the resulting score obtained for each question. For example, the overall score could be the average of the question scores.

[0096] Step 107 of the method may generate training datasets. Each training dataset may contain a set of values ​​for each parameter that specifies the respective mode used in each iteration, and the respective overall score.

[0097] In step 107, a machine learning module (ML module) can be trained using the set of parameter values ​​from the training dataset as input, and the overall score of the training dataset as the target output of the ML module.

[0098] In step 108, a search for a set of improved parameter values ​​may be performed using a trained ML module. The set of improved parameter values ​​(also called improved values) may be determined such that, by inputting the set of improved parameter values ​​into the trained ML module, an improved score greater than the maximum overall score of the training dataset can be obtained.

[0099] Figure 2 illustrates an example of a document set, which includes a first document 11, a second document 12, and a third document 13. In another application, the document set may include 100 or more documents, for example, 1000 documents. The processor set 810 may be configured to perform document partitioning and input content generation according to one of the modifications described above and according to the mode of each iteration. Each mode may be defined by the selection of values ​​for each of these parameters.

[0100] The parameters may include the first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, twelfth, and / or thirteenth parameters described above. The first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, eleventh, and / or twelfth parameters may be labeled as the first parameter 201, the second parameter 202, the third parameter 203, the fourth parameter 204, the fifth parameter 205, the sixth parameter 206, the seventh parameter 207, the eighth parameter 208, the ninth parameter 209, the tenth parameter 210, the eleventh parameter 211, and the twelfth parameter 212, as shown in Figure 6.

[0101] Figure 6 depicts Table 200, which shows the selection of parameter values ​​for six exemplary iterations. Each row in Table 200 may represent the selected parameter value for each iteration. For example, the first row may represent the selection of parameter values ​​for the first iteration. The selected parameter value for the first iteration is described in detail below. The first row may show that the value of the first parameter 201 is set to equal 512, which may suggest that the chunk size of the chunk is equal to 512 tokens.

[0102] Tokens, such as 512 tokens, can be the smallest string of characters that have semantic meaning. For example, in natural language processing (NLP), a token can refer to a word, syllable, or punctuation mark. For instance, in the sentence "Hello, world!", the tokens could be "Hello", ",", "world", and "!". A token can contain one or more characters.

[0103] Figure 2 may illustrate a document divided into nine chunks, each of which is shown in Figure 2 as an unlabeled rectangle. Thus, chunking the first document 1 may result in a set of first chunks 10. Similarly, chunking the second document 2 may result in a set of second chunks 20. By analogy, chunking the third document 3 may result in a set of third chunks 30. The combined set of the first chunks 10, the second chunks 20, and the third chunks 30 can be considered an example of the respective chunk sets of the document generated in step 102 for each iteration. In the example shown in Table 200, the value of the 11th parameter 211 in the first row is set to equal to 1, which may suggest that a first chunking method may be used to create a set of chunks. For example, the first chunking method may be sentence-based chunking. The second chunking method, indicated by a value of 2 in Table 200, may be a paragraph-based chunking method. The third chunking method, indicated by a value of 3 in Table 200, may be a semantic chunking method.

[0104] As shown in the example in Table 200, the value of the third parameter 203 is set to 128, which may suggest that the size of the intersection in the first chunk set 10 may be equal to 128 tokens. The value of the fourth parameter 204 for the first iteration is set to "1", which may indicate that the selected embedding model for generating the embedding vector for the chunk may be a first type of embedding model contained in the code of block 900. The first type of embedding model may be a Word2Vec model. The second type of embedding model may be a BERT model, which may be specified by the value "2" in Table 200. The third type of embedding model may be a FastText model, which may be specified by the value "3" in Table 200.

[0105] Computer 801 may be configured to generate an embedding vector for each chunk using a selected embedding model, as shown in Figure 3. For example, a set of processors 810 may compute a first embedding vector 100.1 for the first chunk 10.1 of a first set of chunks 10, based on a selected embedding model. In one example, processor 810 may load one or more parameter values ​​of one or more neural networks of the selected embedding model into cache 821 to compute the embedding vector. Processor 810 may further provide the first chunk 10.1 as input to the selected embedding model and receive the first embedding vector 100.1 as a response from the selected embedding model. For simplicity, entries in the first embedding vector 100.1 are shown in the form of unlabeled rectangles. Entries may be real numbers. Similarly, a second embedding vector 100.2 may be generated for the second chunk 10.2 of the first set of chunks 10, a third embedding vector 100.3 may be generated for the third chunk 10.3, and a ninth embedding vector 100.9 may be generated for the ninth chunk 10.9 of the first set of chunks 10. These can combine to form the first set of embedding vectors 100.

[0106] Figure 4 illustrates question 40, which is shown as an example of each question used for each iteration. Each question may be selected from a set of questions. The processor 810 may provide question 40 as input to the selected embedding model and receive question embedding vector 40.1 as a response from the selected embedding model.

[0107] Typically, the execution of an iteration can include a subset of multiple iterations. An iteration can be considered as multiple nested loops. Within each of the nested loops, the value of one of these parameters may be changed each time the loop is traversed. Thus, each parameter can be associated with one loop. The order of the loops does not matter. One of the loops may be associated with a variation of a question for generating a prompt (hereinafter referred to as the question loop). The question loop can be an example of the further loops described above. It may be practical to select the question loop as the innermost loop of these loops. This allows the overall score for each combination of parameter values ​​to be generated without having to remember the scores generated for more than one combination of parameter values.

[0108] It is possible to perform the iterations such that only the values ​​of a subset of the 13 parameters mentioned above are modified after all iterations are complete. Depending on the application of this method, not all of the 13 parameters mentioned above may be of interest. For example, in some cases it may be practical to keep the values ​​of the first parameter 201, the third parameter 203, and the eleventh parameter 211 constant for all iterations. As a result, a set of document chunks may be generated in the first loop pass of the iteration, i.e., the first execution, and remain unchanged for the remaining executions of the iteration. In that context, step 102 again states that the value of each of these parameters may specify how the set of chunks or the generation of input content is performed. The value of each parameter that is changed when each of the nested loops is executed may define how the chunks or input content are created.

[0109] Referring again to Figure 6, the first row of Table 200 shows that the value of the fifth parameter 205 is set to equal 384, which could mean that the size of the embedding vectors 100 and 40.1 could be equal to 384. The same could be true for the size of the embedding vectors representing the chunks of the second set of chunks 20 and the third set of chunks 30 (called the embedding vectors of other chunks). The embedding vectors of other chunks can be calculated in the same way as embedding vector 100, and are not shown in the figure for clarity.

[0110] The set of processors 810 can individually compare the question embedding vector 40.1 with the embedding vector 100 and the embedding vectors of other chunks using a similarity metric. In one example, the similarity metric can be selected. The first row depicts that for the first iteration, the value of the 12th parameter 12 is set to equal "1". As shown in the example in Figure 6, the value "1" may indicate a first type of similarity metric. The first type could be, for example, cosine similarity. The value "0" may indicate a second type of similarity metric. The second type could be, for example, Euclidean distance.

[0111] The set of processors 810 may be configured to calculate similarity scores for each of the embedding vectors, 100 and the other chunks, using a selected similarity metric. Furthermore, the set of processors 810 may be configured to rank these vectors according to their similarity scores and to select a predefined number of chunks corresponding to the embedding vector with the highest similarity score.

[0112] The predefined number of chunks to be selected may be the number of chunks selected as described above, i.e., the value of the second parameter 202. In the first iteration, the value of the second parameter 202 may be equal to 5, as shown in the first row of Table 200.

[0113] Furthermore, the set of processors 810 may be configured to generate input content using a selected information retrieval method. For a first iteration, a first type of information retrieval method may be selected, indicated by the value of the tenth parameter 210 in the first row of Table 200 being equal to "1". In one example, the first type of information retrieval method may be dense vector information retrieval. A second type of information retrieval method may be an ElasticSearch method, indicated by the value of the tenth parameter 210 being "2". According to the example shown in the figure, the input content may include the second chunk 10.2 and the third chunk 10.3 of the first set of chunks 10.

[0114] Furthermore, the set of processors 810 may be configured to generate prompts 500 for LLM 50, as shown in Figure 5. Prompts 500 may include input content, a question 40, and a command 51. Command 51 may compel LLM 50 to use the input content to generate a provisional answer 52. For example, processor 810 may generate a prompt for each question in the set of questions. Each prompt may be created in a similar manner to prompts 500 in each iteration of the question loop, with question 40 being the respective question. This may produce a provisional answer for each question.

[0115] In step 105, the set of processors 810 may compare each tentative answer with the corresponding answer to that question (hereinafter referred to as the respective target answer) that each prompt generates in each iteration. The set of processors 810 may calculate the respective score for each question depending on the result of the comparison between each tentative answer and the respective target answer. To perform the comparison between each tentative answer and the respective target answer, the set of processors 810 may generate the embedding vector for each target answer and the embedding vector for each tentative answer using the same embedding model used to generate the embedding vectors 100 and 40.1. Performing the comparison between each tentative answer and the respective target answer may involve calculating the respective further similarity score depending on the embedding vector for each tentative answer and the embedding vector for each target answer, using a similarity metric selected for the iteration or another similarity metric. The score for each question may be the respective further similarity score.

[0116] The set of processors 810 can operate LLM50 according to the setting of the values ​​of the sixth parameter 206, the seventh parameter 207, the eighth parameter 208, and / or the ninth parameter 209. In the example shown in Figure 6, for the first iteration shown by the first row, the value of the sixth parameter 206 may be set to equal 1000 (number of tokens), the value of the seventh parameter 207 may be set to equal 1.3, the value of the eighth parameter 208 may be set to equal 1, indicating a first type of LLM to be operated as LLM50, and / or the value of the ninth parameter 209 may be set to equal 0.9. The first type of LLM may be the GPT-4 model. The second type of LLM may be the BERT model, shown with a value of 2 in Table 200. The third type of LLM may be the LLAMA model, shown with a value of 3 in Table 200.

[0117] Therefore, in one example, the values ​​of each parameter may define how chunks can be created, or how input content can be created based on the selected chunks, or how LLM50 is operated to generate each provisional answer.

[0118] The set of processors 810 can further calculate an overall score for each iteration based on the scores for the questions, for example, on an additional similarity score. In practice, the overall score for each iteration may be equal to the average of the scores for the questions, for example, the average of the additional similarity scores. As shown in the example in Table 200, the overall score for the first iteration may be equal to 0.40. The overall score for an iteration is labeled with reference numeral 220.

[0119] Table 200 illustrates the parameter values ​​and the overall scores obtained for each of the six iterations. Each iteration may include a loop through all questions in the set of questions. Therefore, the iterations can be considered as patterns, and each pattern represents a different parameter setting. If the value of at least one of these parameters is different, it can be considered a different parameter setting. The settings may be indicated by the numbers shown in the first column of Table 200. In one example, the values ​​in Table 200 may be set manually. In another example, the values ​​in Table 200 may be obtained by performing the DOE method as described above. Table 200 may be stored in storage 824 or volatile memory 812.

[0120] For example, prompt 500, in particular command 51, may include instructions for setting the values ​​of the sixth parameter 206, the seventh parameter 207, the eighth parameter 208, and / or the ninth parameter 209, according to the respective settings for operating LLM 50 in each iteration.

[0121] The set of processors 810 can generate each training dataset depending on the parameter values ​​of each pattern and the overall score obtained for each setting. Therefore, each row in Table 200 can represent each training dataset. In most applications, it can be seen that overall scores can be determined for more than 6 settings, for example, 100 or up to 1000 settings.

[0122] As an example, the first training dataset 1 for the first configuration is shown in Figure 7. Similarly, the second training dataset 2 for the second configuration (second row in Table 200), the third training dataset 3 for the third configuration (third row), the fourth training dataset 4 for the fourth configuration (fourth row), the fifth training dataset 5 for the fifth configuration (fifth row), and the sixth training dataset 6 for the sixth configuration (sixth row) are depicted in Figure 7. ML module 70 can be considered an example of the above ML module. For example, ML module 70 may contain multiple decision trees. The set of processors 810 may train ML module 70 depending on training datasets 1, 2, 3, 4, 5, and 6 using one of the above training methods, for example, a random forest or a gradient boosting tree.

[0123] Training may involve using the combined score of each training dataset as the target output value for the ML module 70. The set of processors 810 may use the ML module 70 to compute each training output value 71 for each parameter setting. To this end, the set of processors 810 may provide each set of parameter values ​​for each setting as input to the ML module 70 and receive each training output value 71 as a response.

[0124] Furthermore, the set of processors 810 may calculate the value of the cost or loss function described above, depending on the deviation of the training output value of the ML module 70 from the overall score of the training dataset. Training may further include adapting the parameter values ​​of the ML module 70 depending on the cost or loss function, for example, depending on the partial derivative of the cost or loss function with respect to the parameters of the ML module 70. If the value of the cost or loss function reaches a predefined threshold, training may be interrupted, and the ML module 70 may be in a trained state.

[0125] Performing step 109 may involve generating "n" trial sets 80. Figure 8 shows each trial set 80i, more specifically the first trial set 801, the second trial set 802, and the nth trial set 80. n This provides a schematic description. Each trial set 80i may contain values ​​for a parameter whose value is modified in an iteration. Trial set 80 may have the same dimension as the set of values ​​for each parameter whose value is modified in an iteration. Performing a search for improved values ​​may involve providing each trial set 80i as input to the trained ML module 70 and receiving each trail score 81i as a response from the ML module 70. The result of performing the search may be to identify the improved value as the value of the trial set 80 that yields the best trial score (hereinafter referred to as the best trial set). The search may be performed according to one of the modifications described above, for example, as a grid search, or using one or more genetic algorithms.

[0126] Figure 9 is a flowchart of the application method for applying the trained ML module 70. In step 301, a new question related to a set of documents, for example, documents 11, 12, and 13, may be received. In step 302, a new set of chunks of documents may be generated depending on the set of documents. In addition, in step 302, new input content may be generated depending on the new set of chunks of documents. The generation of the new chunks and the new input content is performed according to a new mode. The new mode is specified by a set of improved parameter values, i.e., by the parameter values ​​given by the best trial set. Each value in the set of improved parameter values ​​may specify how the generation of the new chunks or the new input content is performed.

[0127] In step 303, a new prompt for LLM50 may be generated depending on new input content and a new question. The new prompt may be generated similarly to prompt 500, but a new question may be used instead of question 40, and new input content may be used instead of input content.

[0128] In step 304, a new prompt may be provided as a new input to the LLM50, and a new response may be received from the LLM50 as a response.

[0129] In step 305, a new answer may be provided as a response to a new question.

[0130] According to one application, the calculation of LLM50 may be performed on an additional computing device, for example, on a remote server 804 or public cloud 805 as shown in Figure 10. In this case, prompt 500 may be sent from computer 801 to the additional computing device, for example, the remote server 804 or public cloud 805. According to this application, each provisional answer to each question in the set of questions may be sent from the additional computing device, for example, from the remote server 804 or public cloud 805, to computer 801. By using a trained ML module 70 to search for improved values, data traffic between computer 801 and additional computing devices such as the remote server 804 or public cloud 805 can be reduced. Without the trained ML module 70, LLM50 would have to be used for each trial set. As mentioned above, the number of trial sets 80 may be greater than the number of training datasets. Training of the ML module may be performed on computer 801. Since training the ML module 70 can be considered computationally intensive, the computer 801 may have less computing power than additional computing devices. Therefore, by using the ML module 70 to find improved values, it may be possible to use computing devices with relatively little power.

[0131] The computing environment 800 includes an example of an environment for executing at least a portion of the computer code involved in carrying out the method of the invention, such as the code in block 900 for determining an improved set of parameter values ​​for operating a large-scale language model according to one of the modifications described above. The computing environment 800 may be an example of the computer system described above. In addition to block 900, the computing environment 800 includes, for example, a computer 801, a wide area network (WAN) 802, an end user device (EUD) 803, a remote server 804, a public cloud 805, and a private cloud 806. In this embodiment, the computer 801 includes a processor set 810 (including a processing circuit configuration 820 and a cache 821), a communication fabric 811, volatile memory 812, persistent storage 813 (including an operating system 822 and identified blocks 900), a peripheral device set 814 (including a user interface (UI), a device set 823, storage 824, and an Internet of Things (IoT) sensor set 825), and a network module 815. The remote server 804 includes a remote database 830. The public cloud 805 includes a gateway 840, a cloud orchestration module 841, a host physical machine set 842, a virtual machine set 843, and a container set 844.

[0132] Computer 801 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device currently known or to be developed in the future that is capable of running programs, accessing networks, or querying databases such as remote database 830. As is well understood in the field of computer technology, and depending on the technology, the execution of the computer implementation method may be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of the computing environment 800, in order to keep the presentation as simple as possible, the detailed discussion focuses on a single computer, specifically computer 801. Although computer 801 is not shown in the cloud in Figure 10, it may be located in the cloud. On the other hand, computer 801 does not need to be in the cloud, except to any extent that can be definitively shown.

[0133] The processor set 810 includes one or more computer processors of any type currently known or to be developed in the future. The processing circuit configuration 820 may be distributed across multiple packages, for example, multiple interconnected integrated circuit chips. The processing circuit configuration 820 may implement multiple processor threads and / or multiple processor cores. The cache 821 is memory located within the processor chip package and is typically used for data or code that should be available for high-speed access by threads or cores running on the processor set 810. The cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuit configuration. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, the processor set 810 may operate using qubits and be designed to perform quantum computing.

[0134] Computer-readable program instructions are typically loaded onto computer 801 to cause the processor set 810 of computer 801 to execute a series of operational steps, thereby realizing a computer implementation method. As a result, the instructions thus executed instantiate methods specified in the flowcharts and / or descriptions of computer implementation methods contained herein (collectively referred to as the methods of the invention), for example, a method for determining an improved set of parameter values ​​for operating a large-scale language model according to one of the modifications described above. These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 821 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 810 to control and direct the execution of the methods of the invention. In the computing environment 800, at least some of the instructions for executing the methods of the invention may be stored in block 900 in persistent storage 813.

[0135] The communication fabric 811 is a signal conduction path that enables various components of the computer 801 to communicate with one another. Typically, this fabric is made up of switches and conductive paths, such as buses, bridges, physical input / output ports, and similar components. Other types of signal communication paths, such as fiber optic communication paths and / or wireless communication paths, may be used.

[0136] Volatile memory 812 is any type of volatile memory currently known or to be developed in the future. Examples include dynamic random-access memory (RAM) or static RAM. Typically, volatile memory 812 is characterized by random access, but this is not required unless explicitly stated. In computer 801, volatile memory 812 is located in a single package and resides inside computer 801, but alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally to computer 801.

[0137] Persistent storage 813 is any form of non-volatile storage for a computer, currently known or to be developed in the future. Non-volatility of this storage means that the stored data is maintained regardless of whether power is supplied to computer 801 and / or directly to persistent storage 813. Persistent storage 813 may be read-only memory (ROM), but typically at least a portion of the persistent storage allows for writing, deleting, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 822 can take multiple forms, such as various known proprietary operating systems or open-source portable operating system interface operating systems employing a kernel. The code contained in block 900 typically includes at least a portion of computer code involved in performing the method of the present invention.

[0138] The peripheral device set 814 includes a collection of peripheral devices for computer 801. Data communication connections between computer 801's peripheral devices and other components can be implemented in various ways, including Bluetooth® connections, Near-Field Communication (NFC) connections, connections made by cables (e.g., Universal Serial Bus (USB) type cables), insertable connections (e.g., Secure Digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, the UI device set 823 may include components such as display screens, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 824 is external storage such as an external hard drive, or insertable storage such as an SD card. Storage 824 may be persistent and / or volatile. In some embodiments, storage 824 may take the form of a quantum computing memory device for storing data in the form of qubits. In embodiments where computer 801 is required to have a large amount of storage (for example, computer 801 locally stores and manages a large database), this storage may be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 825 consists of sensors that can be used in an Internet of Things application. For example, one sensor may be a thermometer and another may be a motion detector.

[0139] The network module 815 is a collection of computer software, hardware, and firmware that enables computer 801 to communicate with other computers via the WAN 802. The network module 815 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of the network module 815 are performed on the same physical hardware device. In other embodiments (e.g., embodiments using Software-Defined Networking (SDN)), the control and forwarding functions of the network module 815 are performed on physically separate devices, such that the control function manages multiple different network hardware devices. Computer-readable program instructions for performing the method of the present invention can typically be downloaded to computer 801 from an external computer or external storage device via a network adapter card or network interface included in the network module 815.

[0140] WAN802 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances by any currently known or future-developed technology for transmitting computer data. In some embodiments, WAN802 may be replaced and / or supplemented by a local area network (LAN), such as a Wi-Fi® network, designed to transmit data between devices located in a local area. WANs and / or LANs typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0141] The end-user device (EUD) 803 is any computer system used and controlled by an end-user (e.g., a customer of the company operating computer 801) and can take any of the forms discussed above in relation to computer 801. EUD 803 typically receives useful and valuable data from the operation of computer 801. For example, in a hypothetical case where computer 801 is designed to provide recommendations to the end-user, these recommendations would typically be transmitted from the network module 815 of computer 801 to EUD 803 via WAN 802. Thus, EUD 803 can display or otherwise present recommendations to the end-user. In some embodiments, EUD 803 may be a client device such as a thin client, heavy client, mainframe computer, and desktop computer.

[0142] The remote server 804 is any computer system that provides at least some data and / or functionality to computer 801. The remote server 804 may be controlled and used by the same entity that operates computer 801. The remote server 804 represents a machine that collects and stores useful and valuable data for use by other computers, such as computer 801. For example, in a hypothetical case where computer 801 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 801 from the remote database 830 of the remote server 804.

[0143] Public Cloud 805 is any computer system available for use by multiple entities, providing on-demand availability of computer system resources and / or the processing power of other computers, particularly data storage (cloud storage) and computing power, without direct, active management by the user. Cloud computing typically leverages resource sharing to achieve consistency and economies of scale. Direct and active management of the computing resources of Public Cloud 805 is performed by the computer hardware and / or software of the Cloud Orchestration Module 841. The computing resources provided by Public Cloud 805 are typically implemented by virtual computing environments running on various computers that make up the host physical machine set 842, which is a universe of physical computers located within and / or available to Public Cloud 805. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 843 and / or containers from the container set 844. These VCEs can be stored as images and transferred either as images or after VCE instantiation, among and between various physical machine hosts. The cloud orchestration module 841 manages the transfer and storage of images, deploys new VCE instantiations, and manages the active instantiation of VCE deployments. The gateway 840 is a collection of computer software, hardware, and firmware that enables the public cloud 805 to communicate over the WAN 802.

[0144] Here, some further explanation of virtualized computing environments (VCEs) is provided. A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature in the kernel that allows for the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave as actual computers from the perspective of the programs running within them. Computer programs running on a normal operating system can use all of that computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to the container; this feature is known as containerization.

[0145] Private Cloud 806 is similar to Public Cloud 805, except that its computing resources are available for use by a single enterprise only. While Private Cloud 806 is described as being in communication with WAN 802, in other embodiments, a private cloud may be completely isolated from the internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a distinct and separate entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple configured clouds. In this embodiment, both Public Cloud 805 and Private Cloud 806 are part of a larger hybrid cloud.

[0146] Cloud computing services and / or microservices (not shown separately in Figure 10): Private and public clouds are programmed and configured to provide cloud computing services and / or microservices (unless otherwise indicated, the term “microservices” is to be interpreted as encompassing larger “services,” regardless of size). Cloud services are typically infrastructure, platforms, or software hosted by a third-party provider and made available to users over the internet. Cloud services facilitate the flow of user data from front-end clients (e.g., user-side servers, tablets, desktops, laptops) to the provider’s systems and back over the internet. In some embodiments, cloud services may be configured and orchestrated according to the “as a service” technology paradigm, where something is presented to internal or external customers in the form of cloud computing services. The as a service offering typically provides endpoints that various customers interface with. These endpoints are typically based on a set of application programming interfaces (APIs). One category of as-a-service offerings is Platform as a Service (PaaS), where a service provider provisions, instantiates, runs, and manages modular bundles of code that customers can use to instantiate a computing platform and one or more applications without the complexity of building and maintaining the infrastructure typically associated with them. Another category is Software as a Service (SaaS), where software is centrally hosted and allocated on a subscription basis.SaaS is also known as on-demand software, web-based software, or web-hosted software. The four technical subfields involved in cloud services are deployment, integration, on-demand, and virtual private networks.

[0147] Various aspects of this disclosure are described by explanatory text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of computer program products (CPPs). With respect to any flowchart, operations may be performed in a different order than those shown in a given flowchart, depending on the technology involved. For example, again, depending on the technology involved, two operations shown in consecutive blocks of a flowchart may be performed in reverse order, as a single integrated step, simultaneously, or with at least partial time overlap.

[0148] Embodiments of a computer program product ("CPP Embodiment" or "CPP") are terms used in this disclosure to describe any set of one or more storage media (also called "mediums") that collectively comprise a set of one or more storage devices that collectively comprise machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. "Storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. Computer-readable storage media may be, but are not limited to, electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any preferred combination thereof. Some known types of storage devices, including these media, include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as pits / lands formed on the main surface of a punch card or disk), or any suitable combination of the foregoing. When the term "computer-readable storage medium" is used in this disclosure, it shall not be construed as storage in the form of transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses passing through optical fiber cables, electrical signals communicated through wires, and / or other transmission media.As will be understood by those skilled in the art, data is typically moved at several intermittent points during the normal operation of a storage device, such as during access, defragmentation, or garbage collection; however, data is not transient while it is stored, and this does not mean that the storage device is transient.

[0149] This subject may include the following clauses.

[0150] (Clause 1) A method for determining an improved set of values ​​for parameters to operate a large-scale language model (LLM) depending on a set of documents, the method comprising the steps of loading a set of questions and answers, wherein each answer corresponds to one of the questions, and the questions relate to the content provided by the set of documents; The method comprises a step of performing an iteration, the iteration using a respective mode; generating a set of chunks of the documents depending on the set of documents, and generating input content depending on the set of chunks of the documents, wherein the generation of the chunks and the input content is performed according to a mode, the mode being different for each iteration and specified by a set of values ​​of the parameters, the values ​​of the parameters specifying how the generation of the chunks or the input content is performed; generating a prompt for the LLM depending on the input content and a question among the questions; providing the prompt as input to the LLM and receiving a provisional answer from the LLM as a response; and performing a comparison between the answer and the provisional answer corresponding to a question generated depending on the prompt, thereby obtaining a score for the question, the method further includes A method comprising: generating an overall score for each of the modes depending on the resulting score for the aforementioned question; generating a training dataset for each of the modes, including the set of values ​​for the parameters that specify the mode, and the overall score for each of the modes; training a machine learning (ML) module (ML module) using the set of values ​​for the parameters of the training dataset as input and the overall score of the training dataset as the target output of the ML module; and using the trained ML module to perform a search for the improved set of values ​​for the parameters, wherein the improved set of values ​​for the parameters is input to the trained ML module, resulting in an improved score greater than the maximum overall score of the training dataset.

[0151] (Article 2) The method according to Clause 1, further comprising: receiving a new question relating to the set of documents; generating a new set of chunks of documents depending on the set of documents, and generating new input content depending on the new set of chunks of documents, wherein the generation of the new set of chunks and the new input content is performed according to a new mode, the new mode being specified by the improved set of values ​​of the parameters, each of which value of the improved set of values ​​of the parameters specifies how the generation of the new set of chunks or the new input content is performed; generating a new prompt for the LLM depending on the new input content and the new question; providing the new prompt as a new input to the LLM and receiving a new answer from the LLM as a response; and providing the new answer as a response to the new question.

[0152] (Article 3) The method according to clause 1 or 2, wherein the ML module includes a decision tree, and the training of the ML module includes training the decision tree.

[0153] (Article 4) The method according to Clause 3, wherein the ML module includes a set of decision trees, the ML module's training includes generating a proper subset of the training dataset, where the proper subset is a random sample of the training dataset, and training each of the decision trees using each of the respective subsets of the training dataset.

[0154] (Article 5) The execution of a search includes performing a grid search, a random search, or a search based on a genetic algorithm, as described in any of the preceding clauses 1 to 4.

[0155] (Article 6) The method according to any one of the preceding clauses 1 to 5, wherein the parameters of the set of parameters represent the input features of the ML module, and the training of the ML module comprises determining a feature importance score for each of the input features, where the feature importance score indicates the relative impact that each of the input features has on the predictions of the ML module.

[0156] (Article 7) The method according to any one of the preceding clauses 1 to 6, further comprising the step of selecting a chunk size for the chunks in the set of chunks for each of the iterations, wherein the value of the parameter in the set of parameters specifies the chunk size.

[0157] (Clause 8) The method according to any one of the preceding clauses 1 to 7, further comprising the steps of selecting chunks from the set of chunks for each iteration and generating the input content depending on the selected chunks, wherein the value of a parameter in the set of parameters specifies the number of selected chunks.

[0158] (Article 9) The method according to any one of the preceding clauses 1 to 8, further comprising: for each iteration, the step of selecting the size of the intersection of the set of chunks; and the step of generating the set of chunks such that each of the chunks in the set of chunks contains at least one intersection that is included in another chunk of the set of chunks, wherein the value of the parameter of the set of parameters specifies the size of the intersection.

[0159] (Clause 10) The method according to any one of the preceding clauses 1 to 9, further comprising the step of selecting, for each iteration, a type of embedding model for generating each embedding vector depending on each of the chunks in the set of chunks, wherein each embedding vector represents each of the chunks and the parameter values ​​of the set of parameters specify the type of embedding model.

[0160] (Article 11) The method according to any one of the preceding clauses 1 to 10, further comprising the step of selecting the size of each embedding vector for each iteration, wherein each embedding vector represents each chunk of the set of chunks, and each value of the parameter of the set of parameters specifies the size of each embedding vector.

[0161] (Article 12) The method according to any of the preceding clauses 1 to 11, further comprising the step of selecting the size of the provisional answer for each of the iterations, wherein the value of the parameter of the set of parameters specifies the size of the provisional answer.

[0162] (Article 13) The method according to any of the preceding clauses 1 to 12, further comprising the step of selecting a parameter value from the set of parameters for each of the preceding iterations, wherein the parameter value indicates the degree of repetition of similar contexts in the provisional answer.

[0163] (Article 14) The method according to any one of the preceding clauses 1 to 13, further comprising the step of selecting the type of the LLM for each of the iterations, wherein the parameter values ​​of the set of parameters specify the selected type of the LLM.

[0164] (Article 15) The method according to any one of the preceding clauses 1 to 14, further comprising the step of selecting a parameter value from the set of parameters for each of the preceding iterations, wherein the parameter value indicates the degree of modification of the probability distribution of the LLM.

[0165] (Article 16) The method according to any one of the preceding clauses 1 to 15, further comprising, for each of the preceding iterations, a step of selecting a type of information acquisition method, and the generation of the input content comprising performing the selected information acquisition method to obtain the input content, wherein the value of a parameter in the set of parameters specifies the type of the selected information acquisition method.

[0166] (Article 17) The method according to any one of the preceding clauses 1 to 16, further comprising the steps of: for each iteration, selecting a chunking method from a set of chunking methods for generating chunks in a document-dependent manner, and generating the set of chunks according to the selected chunking method, wherein the values ​​of the parameters in the set of parameters specify the selected chunking method.

[0167] (Article 18) The method further comprises, for each iteration, the steps of: generating an embedding vector, wherein each embedding vector represents each of the chunks in the set of chunks; generating a further embedding vector representing the question, or a set of further embedding vectors representing the question; selecting a type of similarity metric, wherein the values ​​of the parameters in the set of parameters specify the type of the selected similarity metric; performing a comparison between the embedding vector and the further embedding vector using the selected similarity metric, or performing a comparison between the embedding vector and the set of further embedding vectors using the selected similarity metric; selecting a subset of the chunks depending on the result of the comparison; and generating the input content depending on the selected subset of the chunks.

[0168] (Article 19) The method according to any one of the preceding clauses 1 to 18, further comprising the step for each iteration of performing a selection of parameter values ​​of an embedding model for generating an embedding vector depending on each of the chunks in the set of chunks, wherein each embedding vector represents each of the chunks, and the parameter values ​​of the set of parameters specify the values ​​of the parameters of the embedding model.

[0169] (Article 20) A computer program product comprising a computer-readable storage medium in which computer-readable program code is embodied, wherein the computer-readable program code is configured to implement any of the methods described in the preceding clauses 1 to 19.

[0170] (Article 21) A computer system for determining an improved set of parameter values ​​for generating improved prompts for a Large Language Model (LLM) depending on a set of documents, wherein the computer system is configured to load a set of questions and answers, where each answer corresponds to one of the questions, and each question relates to content provided by the set of documents; to perform iterations, where the iterations use the respective modes; to generate a set of chunks of the documents depending on the set of documents, and to generate input content depending on the set of chunks of the documents, where the set of chunks and the input content The generation is performed according to the mode, which differs for each of the iterations and is specified by a set of values ​​for each of the parameters, the values ​​of each of the parameters specifying how the generation of the set of chunks or the input content is performed; generating a prompt for the LLM depending on the input content and a question among the questions; providing the prompt as input to the LLM and receiving each tentative answer as a response from the LLM; and performing a comparison between the answer and the tentative answer corresponding to a question generated depending on the prompt, thereby obtaining a score for the question. The computer system is further configured to: generate a total score for each of the modes depending on the resulting score for the question; generate a training dataset for each mode, including the set of values ​​for the parameters that specify the mode, and the total score for each mode; train a machine learning module (ML module) using the set of values ​​for the parameters of the training dataset as input and the total score of the training dataset as the target output of the ML module; and use the trained ML module to perform a search for the improved set of values ​​for the parameters, wherein the improved set of values ​​for the parameters is input to the trained ML module, resulting in an improved score greater than the maximum total score of the training dataset.

[0171] (Article 22) The computer system according to Clause 21 is configured to: receive a new question relating to the set of documents; generate a new set of chunks of the documents depending on the set of documents, and generate new input content depending on the new set of chunks of the documents, wherein the generation of the new set of chunks and the new input content is performed according to a new mode, the new mode being specified by the improved set of values ​​of the parameters, each of which value of the improved set of values ​​of the parameters specifies how the generation of the new set of chunks or the new input content is performed; generate a new prompt for the LLM depending on the new input content and the new question; provide the new prompt as a new input to the LLM and receive a new answer from the LLM as a response; and provide the new answer as a response to the new question.

[0172] (Article 23) The computer system according to Clause 21 or 22, wherein the ML module includes a set of decision trees, and the training of the ML module includes generating a proper subset of the training dataset, wherein the proper subset is a random sample of the training dataset, and training each of the decision trees using each of the respective subsets of the training dataset.

Claims

1. A method for determining an improved set of values ​​for parameters to operate a large-scale language model (LLM) depending on a set of documents, the method being: A stage in which a set of questions and answers is loaded, where each answer corresponds to one of the questions, and the questions relate to the content provided by the set of documents; The stage of performing iterations Equipped with, The aforementioned iteration is Use the appropriate mode; The process involves generating a set of chunks of the documents depending on the set of documents, and generating input content depending on the set of chunks of the documents, wherein the generation of the chunks and the input content is performed according to the respective modes, each of which differs for each iteration and is specified by a set of values ​​for the respective parameters, each of which specifies how the generation of the chunks or the input content is performed; To generate a prompt for the LLM depending on the input content and one of the questions; Providing the aforementioned prompts as input to the LLM, and receiving the respective provisional answers as responses from the LLM; and The prompt performs a comparison between the respective answers and the respective provisional answers corresponding to a certain question generated in response to that prompt, and as a result obtains a score for the question. Includes, The aforementioned method further, A step of generating a total score for each of the modes, depending on the score obtained as a result of the aforementioned question; A step of generating a training dataset for each of the modes, including the set of values ​​for the parameters that specify each of the modes, and the overall score for each mode; A step of training a machine learning module (ML module) using the set of values ​​of the parameters of the training dataset as input, and using the overall score of the training dataset as the target output of the ML module; and The step involves using the trained ML module to search for the set of improved values ​​for the parameters, where the set of improved values ​​for the parameters is input into the trained ML module, resulting in an improved score greater than the maximum overall score of the training dataset. Equipped with, method.

2. The aforementioned method further, The stage of receiving new questions related to the aforementioned collection of documents; A step of generating a new set of chunks of documents depending on the set of documents, and generating new input content depending on the new set of chunks of documents, wherein the generation of the new set of chunks and the new input content is performed according to a new mode, the new mode being specified by the improved set of values ​​of the parameters, each of which value of the improved set of values ​​of the parameters specifies how the generation of the new set of chunks or the new input content is performed; A step of generating a new prompt for the LLM depending on the new input content and the new question; The steps include providing the new prompt as a new input to the LLM and receiving a new response from the LLM; and The step of providing the aforementioned new answer as a response to the aforementioned new question. The method according to claim 1, comprising:

3. The method according to claim 1, wherein the ML module includes a decision tree, and the training of the ML module includes training the decision tree.

4. The method according to claim 3, wherein the ML module includes a set of decision trees including the decision tree, the training of the ML module includes generating a proper subset of the training dataset, wherein the proper subset is a random sample of the training dataset, and training each of the decision trees using each of the respective subsets of the training dataset.

5. The method according to claim 1, wherein the execution of the search includes performing a grid search, a random search, or performing the search based on a genetic algorithm.

6. The method according to claim 1, wherein the parameters of the set of parameters represent the input features of the ML module, and the training of the ML module includes determining a feature importance score for each input feature, where the feature importance score indicates the relative impact that each input feature has on the prediction of the ML module.

7. The method further, for each of the iterations, In the step of selecting the chunk size of the chunks in the set of chunks, the values ​​of the parameters in the set of parameters specify the chunk size. The method according to claim 1, comprising:

8. The method further, for each of the iterations, A step of selecting chunks from the set of chunks and generating the input content depending on the selected chunks, wherein the value of the parameter in the set of parameters specifies the number of selected chunks. The method according to claim 1, comprising:

9. The method further, for each of the iterations, The step of selecting the size of the intersection of the set of chunks; and The step of generating the set of chunks such that each of the chunks in the set of chunks includes at least one intersection of another chunk in the set of chunks, wherein the parameter values ​​of the set of parameters specify the size of the at least one intersection. The method according to claim 1, comprising:

10. The method further, for each of the iterations, A step of selecting the type of embedding model for generating each embedding vector depending on each of the chunks in the set of chunks, where each embedding vector represents each of the chunks, and the parameter values ​​of the set of parameters specify the type of embedding model. The method according to claim 1, comprising:

11. The method further, for each of the iterations, In the step of selecting the size of each embedding vector, where each embedding vector represents the respective chunk of the set of chunks, and each value of the parameter in the set of parameters specifies the respective size of the embedding vector. The method according to claim 1, comprising:

12. The method further, for each of the iterations, In the step of selecting the size of the provisional answer, the values ​​of the parameters in the set of parameters specify the size of the provisional answer. The method according to claim 1, comprising:

13. The method further, for each of the iterations, The step of selecting the parameter values ​​from the set of parameters, where the parameter values ​​indicate the degree of repetition of similar contexts in the provisional answer. The method according to claim 1, comprising:

14. The method further, for each of the iterations, In the step of selecting the type of LLM, the parameter values ​​of the set of parameters specify the selected type of LLM. The method according to claim 1, comprising:

15. The method further, for each of the iterations, The step of selecting the parameter values ​​from the set of parameters, where the parameter values ​​indicate the degree of modification of the probability distribution of the LLM. The method according to claim 1, comprising:

16. The method further, for each of the iterations, The step of selecting the type of information acquisition method, the generation of the input content includes performing the information acquisition method to obtain the input content, wherein the value of the parameter of the set of parameters specifies the type of the selected information acquisition method. The method according to claim 1, comprising:

17. The method further, for each of the iterations, A step of selecting a chunking method from a set of chunking methods for generating chunks depending on the document, and generating the set of chunks according to the selected chunking method, wherein the values ​​of the parameters in the set of parameters specify the selected chunking method. The method according to claim 1, comprising:

18. The method further, for each of the iterations, The step of generating embedding vectors, where each of the embedding vectors represents each of the chunks in the set of chunks; A step of generating a further embedding vector representing the aforementioned question, or a set of further embedding vectors representing the aforementioned question; In the step of selecting the type of similarity index, the values ​​of the parameters in the set of parameters specify the type of similarity index selected; A step of performing a comparison between the embedding vector and the further embedding vector using the similarity index, or a step of performing a comparison between the embedding vector and the set of further embedding vectors using the similarity index; A step of selecting a subset of the chunks depending on the result of the comparison; and The step of generating the input content depending on the selected subset of the chunk. The method according to claim 1, comprising:

19. The method further, for each of the iterations, A step of selecting parameter values ​​for an embedding model to generate an embedding vector depending on each of the chunks in the set of chunks, where each embedding vector represents each of the chunks, and the parameter values ​​of the set of parameters specify the values ​​of the parameters in the embedding model. The method according to any one of claims 1 to 18, comprising:

20. A computer program comprising computer-readable program code, wherein the computer-readable program code is configured to cause a set of processors to execute a procedure for determining an improved set of values ​​for parameters to operate a large-scale language model (LLM) depending on a set of documents, and the procedure is: A procedure for loading a set of questions and answers, where each answer corresponds to one of the questions, and the questions relate to the content provided by the set of documents; Steps to perform the iteration Includes, The aforementioned iteration, Use the appropriate mode; The process involves generating a set of chunks of the documents depending on the set of documents, and generating input content depending on the set of chunks of the documents, wherein the generation of the chunks and the input content is performed according to the respective modes, each of which differs for each iteration and is specified by a set of values ​​for the respective parameters, each of which specifies how the generation of the chunks or the input content is performed; To generate a prompt for the LLM depending on the input content and one of the questions; Providing the aforementioned prompts as input to the LLM, and receiving the respective provisional answers as responses from the LLM; and The prompt performs a comparison between the respective answers and the respective provisional answers corresponding to a certain question generated in response to that prompt, and as a result obtains a score for the question. Includes, The above procedure further, A procedure for generating a total score for each of the modes, depending on the score obtained as a result of the aforementioned questions; A procedure for generating a training dataset for each of the modes, including the set of values ​​for the parameters that specify each of the modes, and the overall score for each mode; A procedure for training a machine learning module (ML module) using the set of values ​​of the parameters of the training dataset as input, and using the overall score of the training dataset as the target output of the ML module; and A procedure to perform a search for the set of improved values ​​for the parameter using the trained ML module, wherein by inputting the set of improved values ​​for the parameter into the trained ML module, an improved score greater than the maximum overall score of the training dataset is obtained. including, Computer program.

21. A computer system for determining an improved set of parameter values ​​for generating improved prompts for a Large-Scale Language Model (LLM) depending on a set of documents, wherein the computer system is Processor set; Computer-readable storage media; and Program instructions stored on the computer-readable storage medium Equipped with, The program instruction is given to the processor set, A procedure for loading a set of questions and answers, where each answer corresponds to one of the questions, and the questions relate to the content provided by the set of documents; Steps to perform the iteration Perform the steps that include, The aforementioned iteration, Use the appropriate mode; The process involves generating a set of chunks of the documents depending on the set of documents, and generating input content depending on the set of chunks of the documents, wherein the generation of the chunks and the input content is performed according to the respective modes, each of which differs for each iteration and is specified by a set of values ​​for the respective parameters, each of which specifies how the generation of the chunks or the input content is performed; To generate a prompt for the LLM depending on the input content and one of the questions; Providing the aforementioned prompts as input to the LLM, and receiving the respective provisional answers as responses from the LLM; and The prompt performs a comparison between the respective answers and the respective provisional answers corresponding to a certain question generated in response to that prompt, and as a result obtains a score for the question. Includes, The above procedure further, A procedure for generating a total score for each of the modes, depending on the score obtained as a result of the aforementioned questions; A procedure for generating a training dataset for each of the modes, including the set of values ​​for the parameters that specify each of the modes, and the overall score for each mode; A procedure for training a machine learning module (ML module) using the set of values ​​of the parameters of the training dataset as input, and using the overall score of the training dataset as the target output of the ML module; and A procedure to perform a search for the set of improved values ​​for the parameter using the trained ML module, wherein by inputting the set of improved values ​​for the parameter into the trained ML module, an improved score greater than the maximum overall score of the training dataset is obtained. including, Computer system.