Method and electronic device for generating training data for preference optimization learning

WO2026160842A1PCT designated stage Publication Date: 2026-07-30SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-21
Publication Date
2026-07-30

Smart Images

  • Figure KR2026001242_30072026_PF_FP_ABST
    Figure KR2026001242_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for generating training data for preference optimization learning by: inferring a predefined number of times by using a first generative model by inputting the same input sequence into the first generative model the predefined number of times; evaluating preference for a result of the inference by the first generative model by using a second generative model; selecting, from the result of the inference, a plurality of response pairs each comprising a non-preferred response and a preferred response, on the basis of a result of the evaluation; and generating training data by processing the plurality of response pairs.
Need to check novelty before this filing date? Find Prior Art

Description

Method for generating training data for preference optimization learning and electronic device

[0001] This relates to a method for generating training data for preference optimization learning and an electronic device.

[0002] Artificial intelligence-based technologies are being utilized in various fields across industries. A diverse range of AI models have been developed and are being applied in each sector, and the application of AI-based solutions is rapidly increasing not only in manufacturing but also in diverse industries such as robotics, transportation and logistics, healthcare, education, and pharmaceuticals and biotechnology. The adoption of such AI-based technologies is leading to the strengthening of competitiveness for companies and nations.

[0003] Generative Artificial Intelligence (GAI) technology is widely used in various fields, such as text summarization, answering questions, translation, and image generation. In the case of GAI-based services, when a user inputs a prompt—a command containing a request or question—into a generative model, the model can generate a response corresponding to the prompt by performing operations between the matrix corresponding to the input prompt and the matrices included in the model's layers.

[0004] According to one embodiment of the present disclosure, a method for generating training data for preference optimization learning is provided. The method for generating training data for preference optimization learning includes the step of inputting the same input sequence to a first generative model a predetermined number of times and inferring a predetermined number of times using the first generative model. Additionally, the method for generating training data for preference optimization learning includes the step of evaluating the preference for the inference result of the first generative model using a second generative model. Additionally, the method for generating training data for preference optimization learning includes the step of selecting a plurality of response pairs, each comprising a non-preferred response and a preferred response, from the inference result based on the evaluation result of the evaluation. Additionally, the method for generating training data for preference optimization learning includes the step of generating training data by processing the plurality of response pairs.

[0005] According to one embodiment of the present disclosure, a computer-readable recording medium having a program for executing the above-described method recorded thereon may be provided.

[0006] According to one embodiment of the present disclosure, an electronic device for generating training data for preference optimization learning is provided. The electronic device for generating training data for preference optimization learning includes a memory comprising one or more storage media for storing instructions, and at least one processor comprising a processing circuit. By executing one or more instructions by at least one processor, the electronic device inputs the same input sequence to a first generative model a predetermined number of times and infers a predetermined number of times using the first generative model. Additionally, by executing one or more instructions by at least one processor, the electronic device evaluates the preference for the inference result of the first generative model using a second generative model. Additionally, by executing one or more instructions by at least one processor, the electronic device selects a plurality of response pairs, each including a non-preferred response and a preferred response, from the inference result based on the evaluation result. Additionally, by executing one or more instructions by at least one processor, the electronic device generates training data by processing the plurality of response pairs.

[0007] FIG. 1 is a diagram illustrating the process of executing a generation model according to an input sequence including a prompt in an electronic device according to one embodiment of the present disclosure.

[0008] FIG. 2 is a diagram illustrating a preference optimization learning method according to one embodiment of the present disclosure.

[0009] FIG. 3 is a flowchart illustrating a method for generating training data for preference optimization learning according to one embodiment of the present disclosure.

[0010] FIG. 4 is a flowchart for explaining the process of inferring using a first generative model according to one embodiment of the present disclosure.

[0011] FIG. 5 is a diagram showing an example of a prompt for evaluating a preference for an inference result using a second generative model according to one embodiment of the present disclosure.

[0012] FIG. 6 is a diagram showing an example of a result of evaluating a preference for an inference result using a second generative model according to one embodiment of the present disclosure.

[0013] FIG. 7 is a diagram illustrating an example of a plurality of response pairs that can be selected from an inference result based on an evaluation result according to one embodiment of the present disclosure.

[0014] FIG. 8 is a flowchart illustrating the process of generating preference learning data by processing a plurality of response pairs according to one embodiment of the present disclosure.

[0015] FIG. 9 is a diagram showing an example of a prompt and a verification result for verifying a part that is deducted from the evaluation in order to process a non-preferred response according to one embodiment of the present disclosure.

[0016] FIG. 10 is a drawing showing an example of the result of processing a non-preferred response according to one embodiment of the present disclosure.

[0017] FIG. 11 is a diagram showing an example of a prompt and a verification result for verifying a part of a preferred response that is deducted from the evaluation in order to process a preferred response according to one embodiment of the present disclosure.

[0018] FIG. 12 is a diagram showing an example of the result of processing a preference response according to one embodiment of the present disclosure.

[0019] FIG. 13 is a diagram showing an example of a prompt for evaluating a preference for a preference response processed using a second generative model according to one embodiment of the present disclosure and a result of evaluating the preference.

[0020] FIG. 14 is a block diagram illustrating the configuration and operation of an electronic device for generating training data for preference optimization learning according to one embodiment of the present disclosure.

[0021] FIG. 15 is a block diagram for explaining in detail the operation of a learning data generation module in an electronic device for generating learning data for preference optimization learning according to one embodiment of the present disclosure.

[0022] The terms used in this specification will be briefly explained, and the present disclosure will be described in detail. In the present disclosure, the expression "at least one of a, b, or c" may refer to "a," "b," "c," "a and b," "a and c," "b and c," or "all of a, b, and c."

[0023] The terms used in this disclosure have been selected to be as widely used and general as possible, taking into account their functions within this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant explanatory sections. Therefore, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.

[0024] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art as described in this specification. Additionally, terms including ordinal numbers, such as "first" or "second," used in this specification may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another.

[0025] When a part of a specification is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "part" or "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0026] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0027] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0028] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.

[0029] 'Generative AI (generative artificial intelligence)' can refer to artificial intelligence technology configured to generate new text, images, etc., in response to prompts and input data (e.g., text, images, etc.).

[0030] "Generative model" may refer to a neural network model that implements generative AI technology. A generative model can generate text or images, etc., according to the intent contained in a prompt. Additionally, by learning the patterns and structures of training data, a generative model can generate new data having characteristics similar to the input data or new data corresponding to the input data. For example, if the prompt is text containing a question, the generative model can generate and output an answer to the question. For example, if the prompt is text containing a request, the generative model can output text or an image generated according to the request. A transformer executed in an electronic device according to one embodiment of the present disclosure corresponds to a generative model. Instead of, or in addition to, the term "generative model," terms such as "generative artificial intelligence model," "language model," "neural network model," or "model" may be used.

[0031] "Prompt" refers to a sentence or keyword for interaction between a user and a model, and may be text used by the user to convey questions or commands to the model. A prompt may refer to text or other forms of input (e.g., audio or visual images) that guide the model on what kind of output to generate. In this disclosure, "execute a prompt" or "execute a generative model according to a prompt" may refer to an action in which the generative model performs a task in accordance with the request of the prompt, that is, an action in which the generative model performs an operation to generate a result corresponding to the prompt as the prompt is input to the generative model. The prompt may be generated in real-time reflecting the result of the operation and input to the generative model, or a pre-prepared prompt may be input by a processor according to conditions.

[0032] 'Input data' refers to the actual data that a model needs to process or analyze. Input data can take various forms, such as text, images, and audio. For example, if a user requests a translation by entering a prompt like "Translate the following sentence into Korean" into the model, the text to be translated can be considered the input data. Alternatively, for instance, if a user requests an image edit by entering a prompt like "Remove the clouds in the sky" into the model, the image to be edited can be considered the input data. Terms such as 'source data' or 'input values' may also be used instead of 'input data'.

[0033] An 'input sequence' refers to the input actually applied to a model, which can mean the entire input that the model must process. In other words, an input sequence can refer to all data delivered to the model's input layer, and may include not only text prompts but also other forms of input data such as images and audio. That is, an input sequence can be a combination of prompts and input data. Instead of, or in addition to, the terms 'complete input,' 'input stream,' or 'input series' may also be used.

[0034] To explain with a specific example, if a user inputs the sentence "He always inspires me" as input data along with the prompt "Translate the following sentence into Korean," the input sequence could be "Translate the following sentence into Korean. He always inspires me." Alternatively, if a user inputs an image as input data along with the prompt "Remove the clouds in the sky," the input sequence could be a combined form of "Remove the clouds in the sky of the photo" and the image.

[0035] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily implement the present invention. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0036] The present disclosure will be described in detail below with reference to the attached drawings.

[0037] FIG. 1 is a diagram illustrating the process of executing a generation model according to an input sequence including a prompt in an electronic device (100) according to one embodiment of the present disclosure. Referring to FIG. 1, the process of the electronic device (100) providing an output corresponding to the input sequence using a generation model is shown.

[0038] An input sequence including prompts and input data may be input through the user interface of the electronic device (100). The user may request processing of the input sequence through the user interface of the electronic device (100). The prompts or input data included in the input sequence may be prepared in advance in the electronic device (100), automatically generated by the system of the electronic device (100), or input by the user. While performing a predetermined task in the electronic device (100), a plurality of input sequences may be generated or input step by step and processed sequentially. According to one embodiment, the prompts and input data may be provided to the generation model simultaneously. Also, according to one embodiment, the prompts and input data may be provided to the generation model at different times.

[0039] A prompt includes the user's intention to achieve through the generative model and conditions specifying what result the generative model should output. A prompt and input data in the form of text may be input to the generative model, and other forms of input data, such as images or audio, may also be input. For example, a prompt requesting the translation of an input sentence into a specific language and input data of the sentence to be translated may be input to a first generative model. Alternatively, a prompt requesting the performance of a specific image processing on an input image and input data of the image to be edited may be input to a second generative model. According to one embodiment, the prompt may be a natural language query containing text.

[0040] The preprocessor can process the prompt into small data units called tokens to understand the intent and conditions of the prompt and to clarify the actions to be performed by the generative model. The preprocessor may be implemented as an integral part of the generative model or as a separate entity detached from the generative model, as illustrated in FIG. 1. The preprocessor may include a rule-based system or a neural network model that performs tokenization. The preprocessor can improve the efficiency and accuracy of the generative model by removing unnecessary tokens from the prompt through token pruning or vector quantization. For example, the preprocessor may divide the prompt into multiple tokens, evaluate the importance of each token to distinguish between core tokens and auxiliary tokens, retain the core tokens, and remove the auxiliary tokens. According to one embodiment, at least one auxiliary token may be retained.

[0041] A generative model can be a text generation model that performs text generation, translation, summarization, or chatbot functions. A generative model can be an image generation model that generates images based on text descriptions. A generative model can be a music generation model that creates new music by generating melodies, harmonies, and rhythms. A generative model can be a code generation model that generates code based on natural language descriptions. A generative model can be a multimodal generative model that generates new data by analyzing and processing various types of data, such as text, images, audio, and video.

[0042] Users can view new text, images, audio, etc., generated by the generative model in response to prompts and input data through the user interface. For example, if the prompt and input data are text containing a question, the generative model can generate an answer to the question and provide it to the user through the user interface.

[0043] FIG. 2 is a diagram illustrating a preference optimization learning method according to one embodiment of the present disclosure.

[0044] Even if large language models are trained using vast training datasets, they may provide outputs that do not align with human values. For example, a trained large language model may provide biased outputs based on the training dataset, or it may produce harmful or misleading outputs. To address these issues, trained large language models can improve their outputs through Reinforcement Learning from Human Feedback (RLHF). However, Reinforcement Learning from Human Feedback requires significant computational power and data processing using reward models.

[0045] According to one embodiment of the present disclosure, the preference optimization learning method may be Direct Preference Optimization learning. Unlike reinforcement learning through human feedback, Direct Preference Optimization learning for a large language model performs a preference optimization process using preference learning data without a reward model. While reinforcement learning through human feedback guides the policy of a large language model through an iterative feedback loop using a reward model, Direct Preference Optimization learning can optimize the output of a large language model to match responses preferred by humans using preference learning data. Since Direct Preference Optimization learning does not require a separate reward model, the burden on computing resources can preferably be reduced.

[0046] While training data for direct preference optimization is available as open datasets, such datasets are general-purpose datasets usable by various models for specific purposes and cannot be considered suitable for a particular model. Below, we describe a method for generating training data for preference optimization suitable for a specific model.

[0047] Referring to FIG. 2, the process of an electronic device (100) generating preference learning data for a first generative model using a first generative model and a second generative model is illustrated for the preference optimization learning of a first generative model. The electronic device (100) can generate preference learning data configured for preference optimization for a first generative model by generating preference learning data using the inference results of the first generative model for the preference optimization learning of the first generative model.

[0048] In the electronic device (100), the first generative model may receive as input an input sequence including a prompt and input data. The first generative model may perform inference based on the input sequence and output an inference result. The first generative model may be a large language model on which Supervised Fine Tuning (SFT) has been performed.

[0049] In the electronic device (100), the second generative model may receive the output result of the first generative model as input. The second generative model may evaluate the preference for the inference result and output the evaluation result. The second generative model may score the inference result of the first generative model according to a predetermined standard, giving a high score to a response preferred by a human (user) (hereinafter, preferred response) and a low score to a response not preferred by a human (user) (hereinafter, non-preferred response). For example, the evaluation score for a preferred response may be greater than the evaluation score for a non-preferred response. The second generative model may be a preference evaluation model and may be a large language model that performs preference evaluation according to an input sequence providing the inference result and evaluation data of the first generative model.

[0050] The electronic device (100) can generate preference learning data from the inference results of the first generative model based on the evaluation results of the second generative model. The electronic device (100) can select response pairs consisting of a non-preferred response that received a first score among the inference results of the first generative model and a preferred response that received a second score among the inference results of the first generative model. The electronic device (100) can generate multiple response pairs. The electronic device (100) can generate preference learning data by processing at least one of the non-preferred response and the preferred response constituting the response pair. According to one embodiment, the evaluation score can be compared with a response threshold, and a response in which the evaluation score exceeds the response threshold is identified as a preferred response, and a response in which the evaluation score is less than the response threshold is identified as a non-preferred response.

[0051] The electronic device (100) can perform preference optimization learning for the first generative model using the generated preference learning data. For example, if the electronic device (100) has a learning module that performs preference optimization learning, the electronic device (100) can perform direct preference optimization learning for the first generative model using the preference learning data. Alternatively, the electronic device (100) can transmit the generated preference learning data to another external device. For example, the preference learning data generated by the electronic device (100) can be transmitted to another electronic device or server having the same first generative model. According to one embodiment, the preference learning data can be transmitted to one or more servers operating as a distributed cloud that performs the functions of the generative models disclosed herein.

[0052] FIG. 3 is a flowchart illustrating a method for generating training data for preference optimization learning according to one embodiment of the present disclosure.

[0053] Referring to FIG. 3, in step S310, the electronic device (100) can infer the same input sequence a predetermined number of times using a first generative model. The first generative model is a large language model on which SFT learning has been performed, and may be a model on which SFT learning has been performed in the electronic device (100) or received from an external source. The input sequence may include a prompt and input data. The input sequence may be pre-prepared, or part of the input sequence may be input received directly from the user. For example, both the prompt and the input data may be predetermined, or at least one of the prompt and the input data may be input received from the user.

[0054] The electronic device (100) can infer the same input sequence a predetermined number of times using a first generative model. To generate various inference results for the same input sequence, a setting value regarding the degrees of freedom of the output of the first generative model can be set to a predetermined value within the range from '0' to '1'. For example, if the setting value regarding the degrees of freedom of the output of the first generative model is set to '0', only the same inference result can be generated for the same input sequence. For example, if the setting value regarding the degrees of freedom of the output of the first generative model is set to '0.7', various inference results can be generated for the same input sequence. After adjusting the setting regarding the degrees of freedom of the output of the first generative model to generate various inference results for the same input sequence, the electronic device (100) can infer the same input sequence a predetermined number of times using the first generative model. According to one embodiment, if the predetermined number of times is N, the input sequence can be input to the first generative model N times to obtain N results.

[0055] FIG. 4 is a flowchart for explaining the process of inferring using a first generative model according to one embodiment of the present disclosure.

[0056] In step S410, the electronic device (100) may obtain prompts and input data. The electronic device (100) may obtain prompts and input data from an input sequence. The prompts may include inference elements or inference guide information used for inference of the first generative model, but are not necessarily required. The input data may be the object to be inferred or the question itself.

[0057] In step S420, the electronic device (100) can perform inference by inputting a prompt and input data to a first generating model. The first generating model can output an inference result corresponding to the input data according to the information or intent included in the prompt.

[0058] In step S430, the electronic device (100) can determine whether the number of inferences using the first generative model corresponds to a predefined number. For example, it can be determined whether the electronic device (100) has performed a predefined number of inferences for the same input sequence. The predefined number can be adjusted by one or more settings, such as user-defined settings. For example, if the predefined number is set to '100 times', the electronic device (100) can perform inference using the first generative model until the number of inferences using the first generative model reaches '100 times'. When the number of inferences using the first generative model reaches '100 times', the electronic device (100) can terminate the inference using the first generative model.

[0059] Referring again to FIG. 3, at step S320, the electronic device (100) can evaluate a preference for the inference result of the first generative model using a second generative model. The second generative model may be a large language model that evaluates a preference for the inference result of the first generative model according to an input sequence that provides the inference result and evaluation data of the first generative model.

[0060] The electronic device (100) may acquire an input sequence that provides the inference result of a first generative model and evaluation data. The input sequence may include the inference result of the first generative model as input data in a prompt containing evaluation data. The evaluation data may include one or more of evaluation criteria, evaluation elements, evaluation methods, and evaluation utterances.

[0061] The electronic device (100) can input the inference results and evaluation data of the first generative model provided in the input sequence into the second generative model to evaluate the preference for the inference results of the first generative model. According to one embodiment, the preference for the inference results may be designated as either “preferred” or “non-preferred.” The second generative model may score the inference results of the first generative model according to the evaluation data, output a high score or a high-ranking evaluation for the preferred response, and output a low score or a low-ranking evaluation for the non-preferred response.

[0062] For example, the second generative model can obtain the similarity between the sentence corresponding to the inference result of the first generative model and the sentence generated based on evaluation data, and evaluate sentences with higher similarity as having higher preference. Each sentence can be converted into a vector through embedding, and semantic similarity between sentences can be obtained by calculating the cosine similarity between the vectors. As another example, the second generative model can determine whether the sentence corresponding to the inference result of the first generative model contains all the words, phrases, and sentences generated based on evaluation data, and evaluate sentences with higher completeness as having higher preference.

[0063] FIG. 5 is a diagram showing an example of a prompt for evaluating a preference for an inference result using a second generative model according to one embodiment of the present disclosure.

[0064] Referring to Fig. 5, the prompt requests that the user evaluate the quality of the bot's response to a user's question. The prompt includes evaluation criteria, evaluation factors, and information necessary for the response. In the example prompt in Fig. 5, information necessary for the response is provided separately as a document item, but it is not mandatory to provide it. The prompt indicates that basic response evaluation factors such as accuracy, friendliness, usefulness, and stability should be considered. The prompt includes factors that result in deductions in the evaluation, one or more examples of deducting evaluation points, and a method for displaying the evaluation results. Fig. 5 indicates that after providing an explanation regarding the evaluation, the preference for the response should be evaluated on a scale from 1 to 10.

[0065] Looking at the example prompt in Fig. 5, in response to the evaluative utterance question "Who is the President of Korea?", the inference result "The current President of the Republic of Korea is Yoon Seok-yeol. Yoon Seok-yeol took office as the 20th President on May 10, 2022" can be provided as an answer. The prompt may indicate that, as evaluation criteria and evaluation elements for the answer, a total of four pieces of information—the President's inauguration date, year of birth, educational background, and career—must be included. Based on these evaluation criteria and evaluation elements, the second generative model can evaluate the answer of the inference result of the first generative model for the evaluative utterance question and output an evaluation result.

[0066] FIG. 6 is a diagram showing an example of a result of evaluating a preference for an inference result using a second generative model according to one embodiment of the present disclosure.

[0067] Referring to Fig. 6, the results of an evaluation based on evaluation criteria and evaluation factors are shown for the inference result answer, "The current President of the Republic of Korea is Yoon Suk-yeol. Yoon Suk-yeol took office as the 20th President on May 10, 2022," in response to the evaluation utterance question "Who is the President of Korea?" in Fig. 5. As shown in Fig. 6, in accordance with the request written in the prompt of Fig. 5, the answer of the inference result was evaluated according to basic response evaluation factors such as accuracy, familiarity, usefulness, and stability, and corresponds to "rating [[6]]". It explains that the answer of the inference result in Fig. 5 is an answer that does not meet some evaluation criteria, has problems in terms of accuracy and usefulness, and that all essential information mentioned as evaluation factors must be included in the answer.

[0068] Referring again to FIG. 3, in step S330, the electronic device (100) may select a plurality of response pairs consisting of a non-preferred response and a preferred response from the inference results of the first generative model based on the evaluation results. The electronic device (100) may perform inference a predetermined number of times using the first generative model and obtain a result of evaluating the preference using the second generative model for each inference result. For example, if there were a total of 100 different inference results using the first generative model, 100 evaluation results corresponding to the 100 inference results may be obtained using the second generative model. Based on the 100 evaluation results, the electronic device (100) may select a plurality of response pairs consisting of an inference evaluated as a non-preferred response and an inference evaluated as a preferred response among the 100 inference results.

[0069] The electronic device (100) can determine the inference result that received the best evaluation score from the evaluation results of the second generation model as the preferred response. The electronic device (100) can obtain at least one sentence corresponding to the inference result that received the best evaluation score. If there are multiple inference results that received the best evaluation score, some of the inference results can be randomly selected. If there is only one inference result that received the best evaluation score, the sentence corresponding to that inference result can be determined as the preferred response for each of the multiple response pairs.

[0070] The electronic device (100) can determine the inference result with the minimum evaluation score and the inference result with the mode evaluation score from the evaluation results of the second generative model as non-preferred responses. The electronic device (100) can select the sentence corresponding to the inference result with the minimum evaluation score as the first non-preferred response. The reason for selecting the sentence corresponding to the inference result with the minimum evaluation score as the first non-preferred response is that it is the worst case among the inference results of the first generative model, so it is to be advantageously used to train the generative model to output a better result than this. The electronic device (100) can select the sentence corresponding to the inference result with the mode evaluation score as the second non-preferred response. The reason for selecting the sentence corresponding to the inference result with the mode evaluation score as the second non-preferred response is to use it to train the generative model to output a better result than the sentence that the first generative model can generate with the highest probability.

[0071] The electronic device (100) selects multiple pairs of responses from the inference results in which preferences are evaluated, and by including two non-preferred responses, it can minimize cases of side effects (e.g., inaccurate results) that may occur during preference optimization learning. That is, based on the two non-preferred responses included in the multiple pairs of responses, sentences similar to these among the sentences that the first generative model can generate can be induced or designated to be non-preferred.

[0072] FIG. 7 is a diagram illustrating an example of a plurality of response pairs that can be selected from an inference result based on an evaluation result according to one embodiment of the present disclosure.

[0073] Looking at Fig. 7, it can be seen that the first preferred response of the first response pair and the second preferred response of the second response pair are both inference results that received the highest evaluation results. The first non-preferred response of the first response pair is the inference result that received the lowest evaluation result, and the second non-preferred response of the second response pair may be the inference result corresponding to the most evaluation results. For example, according to the examples in Figs. 5 and 6 above, the inference result with an evaluation result of 'rating [

[0010] ]' may correspond to the first preferred response of the first response pair and the second preferred response of the second response pair. The inference result with an evaluation result of 'rating [[1]]' may be the first non-preferred response of the first response pair. If the inference result with an evaluation result of 'rating [[6]]' was the most frequent, the inference result with an evaluation result of 'rating [[6]]' may be the second non-preferred response of the second response pair.

[0074] Referring again to FIG. 3, in step S340, the electronic device (100) can generate preference learning data by processing a plurality of response pairs. The electronic device (100) can process at least one of the non-preferred response and the preferred response included in the plurality of response pairs. The reason the electronic device (100) processes the selected response pairs is that the quality of the non-preferred response and the preferred response constituting the response pair affects the accuracy and directionality of the model. For example, the non-preferred response included in the response pair may contain factual data or meaningful data. The preference learning data of the unprocessed non-preferred response can be used to train the generative model to reject even the parts corresponding to factual or meaningful data. Furthermore, the preferred response included in the response pair alone may not be sufficient to train the generative model to provide a satisfactory level of output. The preference learning data of the unprocessed preferred response can be used to train the generative model to output results similar to that level, even though it is not at a satisfactory level. Therefore, it is necessary to generate training data that minimizes parts that may have an adverse effect on the preference optimization learning of the generative model. The electronic device (100) can perform a process of processing the non-preferred response for the non-preferred response among the multiple response pairs, and a process of processing the preferred response for the preferred response among the multiple response pairs in parallel.

[0075] FIG. 8 is a flowchart illustrating the process of generating preference learning data by processing a plurality of response pairs according to one embodiment of the present disclosure.

[0076] In step S810, the electronic device (100) can distinguish between a non-preferred response and a preferred response in each of the multiple response pairs. For the non-preferred response, the process of step S820 can be performed, and for the preferred response, the process of step S840 can be performed.

[0077] In step S820, the electronic device (100) can identify the part of the non-preferred response that is deducted from the evaluation using the second generation model. For example, the electronic device (100) can identify the part of the sentence of the non-preferred response that is initially deducted. The electronic device (100) can identify the sentence containing the part that is deducted from the evaluation using the second generation model by including the sentence corresponding to the non-preferred response as input data in the prompt.

[0078] In step S830, the electronic device (100) can filter out the non-preferred response so that the part that is deducted remains. For example, the electronic device (100) can keep the non-preferred response up to the first part that is deducted in the sentence of the non-preferred response. The electronic device (100) can keep the non-preferred response up to the sentence containing the first part that is deducted among the sentences corresponding to the non-preferred response.

[0079] FIG. 9 is a diagram showing an example of a prompt and a verification result for verifying a part that is deducted from the evaluation in order to process a non-preferred response according to one embodiment of the present disclosure.

[0080] Referring to FIG. 9, the prompt requests to find a sentence containing the first deductible part based on the question, the answer, and the evaluation of the answer. The prompt includes evaluation criteria and evaluation elements, and an evaluation of the answer based on basic response evaluation elements such as accuracy, kindness, usefulness, and stability, as well as perfect responses and elements that deduct points from the evaluation. FIG. 9 indicates that after providing a description of the sentence containing the first deductible part, the identification number of the sentence must be indicated according to a specific format.

[0081] Looking at the example prompt in Fig. 9, in response to the evaluative utterance question "Who is the President of Korea?", the sentences "1. The current President of the Republic of Korea is Moon Jae-in. 2. Moon Jae-in took office as the 19th President on May 10, 2017" are entered as an example of a non-preferred response among the response pairs selected from the inference result. As an evaluation of these non-preferred response sentences, the prompt indicates that in terms of accuracy, it incorrectly specified that the current President of the Republic of Korea is Moon Jae-in, and in terms of usefulness, it provides incorrect information about the current President of the Republic of Korea and omits detailed information such as the President's year of birth, education, and career.

[0082] The electronic device (100) can output a sentence containing the first part that is deducted among the sentences of the unpreferred response, along with an explanation, as an identification number of the sentence according to a specific format by executing a second generation model according to a prompt. Referring to the example of the verification result in FIG. 9, it can be seen that the sentence containing the first part that is deducted is the first sentence, provides an explanation that the current president was incorrectly identified, and is displayed in the format 'Sentence: [[1]]'.

[0083] FIG. 10 is a drawing showing an example of the result of processing a non-preferred response according to one embodiment of the present disclosure.

[0084] The electronic device (100) can process the non-preferred responses that constitute a response pair. The electronic device (100) can filter the non-preferred responses so that the sentence containing the first part that is deducted remains. Referring to FIG. 10, it can be seen that the non-preferred responses are filtered so that the first sentence containing the first part that is deducted remains among the sentences of the non-preferred responses of FIG. 9. Such processing of the non-preferred responses causes the second sentence corresponding to the fact among the sentences of the non-preferred responses to be excluded from the non-preferred responses. Processing of the non-preferred responses generates preference learning data for the non-preferred responses so as not to train the generative model to negate even the data corresponding to the fact.

[0085] Referring again to FIG. 8, at step S840, the electronic device (100) can identify the part of the preferred response that is deducted from the evaluation using a second generative model. For example, the electronic device (100) can identify the part of the sentence of the preferred response that is deducted for the first time. The electronic device (100) can identify the sentence containing the part that is deducted from the evaluation using a second generative model by including the sentence corresponding to the preferred response as input data in the prompt. However, if the preferred response has already received the highest evaluation, the preferred response can be included in the preference learning data.

[0086] In step S850, the electronic device (100) can filter the preferred response so that the deductible part is removed. For example, the electronic device (100) can keep the preferred response up to the first deductible part in the sentence of the preferred response, and exclude the part after the first deductible part from the preferred response. The electronic device (100) can keep the sentences corresponding to the preferred response up to the sentence containing the first deductible part as the preferred response.

[0087] In step S860, the electronic device (100) can supplement the part that adds points to the evaluation of the filtered preference response. For example, the electronic device (100) can add a sentence that adds points after the sentence corresponding to the filtered preference response. The electronic device (100) can generate a supplemented preference response using a second generation model by including the sentence corresponding to the filtered preference response as input data in the prompt.

[0088] In step S870, the electronic device (100) can evaluate the preference for the supplemented preference response using a second generative model. For example, the electronic device (100) can evaluate the preference for the sentence corresponding to the supplemented preference response using a second generative model by including the sentence corresponding to the supplemented preference response in the prompt as input data.

[0089] In step S880, the electronic device (100) may determine whether to apply a process of processing the preference response to the supplemented preference response based on the evaluation result of the supplemented preference response. The electronic device (100) may determine whether the evaluation result of the preference for the supplemented preference response satisfies a predefined criterion. For example, the electronic device (100) may determine whether the sentence corresponding to the supplemented preference response received the highest score evaluation. The electronic device (100) may execute a process of processing the supplemented preference response until the evaluation result of the preference for the supplemented preference response satisfies a predetermined criterion.

[0090] FIG. 11 is a diagram showing an example of a prompt and a verification result for checking a part of a preferred response that is deducted from the evaluation in order to process a preferred response according to one embodiment of the present disclosure.

[0091] Referring to FIG. 11, the prompt requests to find a sentence containing the first deductible part based on the question, the answer, and the evaluation of the answer. The prompt includes evaluation criteria and evaluation factors, and an evaluation of the answer based on basic response evaluation factors such as accuracy, kindness, usefulness, and stability, as well as perfect responses and factors that deduct points from the evaluation. FIG. 11 indicates that after providing a description of the sentence containing the first deductible part, the identification number of the sentence must be indicated according to a specific format.

[0092] Looking at the example prompt in Fig. 11, in response to the evaluative utterance question "Who is the President of Korea?", the sentences "1. The current President of the Republic of Korea is Yoon Suk-yeol. 2. Yoon Suk-yeol took office as the 20th President on May 10, 2022" are input as an example of a preferred response among the response pairs selected from the inference result. As an evaluation of these preferred response sentences, the prompt indicates that in terms of accuracy, the current President of the Republic of Korea was correctly identified, but detailed information such as the President's year of birth, education, and career was not provided, and in terms of usefulness, it is somewhat useful but does not fully satisfy the evaluation criteria.

[0093] The electronic device (100) can output a sentence containing the first deductible part among the sentences of the preferred response, along with an explanation, as an identification number of the sentence according to a specific format by executing a second generation model according to a prompt. Referring to the example of the verification result in FIG. 11, it can be seen that the sentence containing the first deductible part is the second sentence, and it provides an explanation that additional essential information such as the president's year of birth, education, and career was not provided, and is displayed in the format 'Sentence: [[2]]'.

[0094] FIG. 12 is a diagram showing an example of the result of processing a preference response according to one embodiment of the present disclosure.

[0095] The electronic device (100) can process the preferred responses constituting the response pairs. The electronic device (100) can filter the preferred responses so that the sentence containing the first deductible part is removed. Referring to FIG. 12, it can be seen that the preferred responses are filtered so that the second sentence containing the first deductible part among the sentences of the preferred responses of FIG. 11 is removed. Such processing of the preferred responses ensures that the second sentence, which does not properly reflect the evaluation criteria and evaluation elements among the sentences of the preferred responses, is excluded from the preferred responses. The electronic device (100) can supplement the parts that add points to the evaluation in the filtered preferred responses. Referring to FIG. 12, it can be seen that the parts that add points to the filtered preferred responses are supplemented by adding sentences containing information on the President's inauguration date, year of birth, educational background, and career to the preferred response in which only the first sentence of the preferred response of the selected response pair remains. Consequently, the processing of the preferred responses generates preference learning data for the preferred responses so as not to train the generative model to output results of an appropriate level rather than a satisfactory level.

[0096] FIG. 13 is a diagram showing an example of a prompt for evaluating a preference for a preference response processed using a second generative model according to one embodiment of the present disclosure and a result of evaluating the preference.

[0097] Referring to Fig. 13, as described earlier in Fig. 5, the prompt requests that the quality of the bot's response to the user's question be evaluated. The prompt indicates that the response must be evaluated according to evaluation criteria, and that basic response evaluation elements such as accuracy, friendliness, usefulness, and stability must be considered. The prompt includes elements that result in deductions in the evaluation, deductions from the evaluation score, and methods for displaying the evaluation results.

[0098] Looking at the example of the prompt in FIG. 13, in response to the question of the evaluation utterance "Who is the President of Korea?", sentences of processed preferred responses are input as answers by supplementing the parts that add points to the filtered preferred responses described in FIG. 12. The electronic device (100) can execute a second generation model according to the prompt in FIG. 13 and can output an evaluation result that evaluates the answers of the processed preferred responses to the question of the evaluation utterance.

[0099] Referring to FIG. 13, an example of the result of evaluating preference according to evaluation criteria and evaluation factors is shown for the processed preference response. According to the request written on the prompt in FIG. 13, the processed preference response was evaluated according to basic response evaluation factors such as accuracy, friendliness, usefulness, and stability, and it can be seen that it corresponds to 'rating [

[0010] ]'. That is, the processed preference response in FIG. 13 is a response that satisfies the evaluation criteria and contains all necessary information, and is the result of the highest score evaluation. In this case, the electronic device (100) can terminate the process of processing the preference response because the sentence corresponding to the processed preference response received the highest score evaluation.

[0100] As described above, the electronic device (100) can generate preference learning data by processing a plurality of response pairs consisting of non-preferred responses and preferred responses. Since the preference learning data generated in this way is based on the inference results of the first generative model and the preference evaluation results for the inference results of the first generative model, it may be preference learning data exclusive to the first generative model.

[0101] FIG. 14 is a block diagram illustrating the configuration and operation of an electronic device (100) for generating training data for preference optimization learning according to one embodiment of the present disclosure.

[0102] Referring to FIG. 14, an electronic device (100) for generating training data according to one embodiment may include at least one memory (110), a processor (120), a communication interface (130), and an input / output interface (140). The electronic device (100) is a processing device configured to perform operations capable of executing a generation model, and may be a user device such as a smartphone, smart glasses, a wearable device, a tablet computer, a laptop computer, a desktop computer, or a server.

[0103] A memory (110) according to one embodiment of the present disclosure may store a program for processing and controlling a processor (120), and may store instructions, data structures, and program code that the processor (120) can read. In one embodiment of the present disclosure, operations performed by the processor (120) may be implemented by executing the instructions or code of the program stored in the memory (110).

[0104] A memory (110) according to one embodiment of the present disclosure may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a non-volatile memory including at least one of ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, a magnetic disk, and an optical disk, and a volatile memory such as RAM (Random Access Memory) or SRAM (Static Random Access Memory).

[0105] A memory (110) according to one embodiment of the present disclosure may store one or more instructions and / or programs that control an electronic device (100) generating training data to train an artificial intelligence model or to use an artificial intelligence model. For example, the memory (110) may store an inference module, an evaluation module, and a training data generation module. The electronic device (100) generating training data may further be equipped with a training module for training an artificial intelligence model.

[0106] A processor (120) according to one embodiment of the present disclosure can control operations or functions performed by an electronic device (100) that generates learning data by executing instructions or programmed software modules stored in memory (110). The processor (120) may be composed of hardware components that perform arithmetic, logic, and input / output operations and signal processing. The processor (120) can control the overall operations of the electronic device (100) that generates learning data by executing one or more instructions stored in memory (110).

[0107] A processor (120) according to one embodiment of the present disclosure may include various processing circuits and / or a plurality of processors. For example, the term “processor” as used herein, including in the claims, may include at least one processor and various processing circuits. “At least one processor” may be configured to perform the various functions described herein individually and / or collectively. As used herein, “processor,” “at least one processor,” and “one or more processors” may be configured to perform various functions. However, these terms cover, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor can perform all functions. Additionally, “at least one processor” may include a combination of processors performing various functions of the disclosed functions in a distributed manner. “At least one processor” may execute program instructions to achieve or perform various functions.

[0108] A processor (120) according to one embodiment of the present disclosure may be composed of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), an Application Processor, a Neural Processing Unit, or at least one of an AI-dedicated processor designed with a hardware structure specialized for processing AI models, but is not limited thereto. Each processor constituting the processor (120) may be a dedicated processor for performing a predetermined function.

[0109] An artificial intelligence (AI) processor according to one embodiment of the present disclosure can perform computation and control for processing a task set to be performed by an electronic device (100) that generates training data using an artificial intelligence (AI) model. The AI ​​processor may be manufactured in the form of a dedicated hardware chip for artificial intelligence (AI), or it may be manufactured as part of a general-purpose processor (e.g., CPU or application processor) or a graphics-dedicated processor (e.g., GPU) and mounted on the electronic device (100).

[0110] A communication interface (130) according to one embodiment of the present disclosure can perform wired or wireless communication with other devices or networks. The communication interface (130) may include a communication circuit or a communication module that supports at least one of various wired or wireless communication methods. For example, the communication interface (130) can perform data communication between an electronic device (100) that generates training data and other devices by using at least one of a data communication method including wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, WFD (Wi-Fi Direct), infrared communication (IrDA, infrared Data Association), BLE (Bluetooth Low Energy), NFC (Near Field Communication), Wibro (Wireless Broadband Internet), WiMAX (World Interoperability for Microwave Access), SWAP (Shared Wireless Access Protocol), WiGig (Wireless Gigabit Alliances), or RF communication.

[0111] A communication interface (130) according to one embodiment of the present disclosure may receive an artificial intelligence model or data used to generate training data from an external device. For example, the communication interface (130) may receive an artificial intelligence model or database trained on a server or another device from a server. The communication interface (130) may transmit an artificial intelligence model trained on an electronic device (100) that generates training data or generated training data to an external device or server.

[0112] An input / output interface (140) according to one embodiment of the present disclosure may include an output unit that provides information and may further include an input unit that receives input. The input / output interface (140) may be in a form where the output unit and the input unit are separated, or in a single form where they are integrated, such as a touchscreen. The input / output interface (140) may receive input information from a user and provide output information to the user. The output unit may output an audio signal or a video signal. The output unit may include a display unit that displays information processed by an electronic device (100) that generates learning data and displays a user interface for receiving user operations. The output unit may include a speaker or buzzer that outputs an audio signal. The input unit may obtain user input for controlling the electronic device (100) that generates learning data. For example, the input unit may be a key pad, a touch pad, a touch panel (contact-type capacitive method, pressure-type resistive method, infrared detection method, surface-type ultrasonic conduction method, integral-type tension measurement method, piezo-effect method, etc.), a microphone, etc. In addition, the input unit may include an eye-tracking sensor, a jog wheel, etc., but is not limited thereto. User input may be in the form of text, voice, or gestures, but is not limited thereto. The user may input prompts, sentence-type search commands, search terms, etc. through the user interface or microphone of the electronic device (100) that generates learning data.

[0113] An input / output interface (140) according to one embodiment of the present disclosure can obtain setting values ​​or inputs for generating training data. The input / output interface (140) can receive user input regarding information or control commands required during the process of generating training data.

[0114] According to one embodiment of the present disclosure, an electronic device (100) for generating training data for preference optimization learning comprises a memory (110) comprising one or more storage media for storing instructions and at least one processor (120) comprising a processing circuit. The at least one processor (120) can execute one or more instructions to load and execute instructions or code for at least one module for generating training data. Prompts and input data may be generated or input step-by-step according to a process while performing a predetermined task in the electronic device (100) and may be executed by a generation model of the electronic device (100).

[0115] According to one embodiment of the present disclosure, one or more instructions are executed by at least one processor (120) so that the electronic device (100) may infer the same input sequence a predetermined number of times using a first generative model. The first generative model may be a large language model on which Supervised Fine Tuning (SFT) learning has been performed. The input sequence may be prepared in advance within the electronic device (100) or may be automatically generated by the system of the electronic device (100), and the input data of the input sequence may be input received from a user.

[0116] The inference module (121) of the processor (120) can adjust the settings regarding the degrees of freedom of the output of the first generating model so that various inference results are generated for the same input sequence. The inference module (121) can obtain an input sequence including a prompt and input data. The inference module (121) can perform inference a predetermined number of times for the same input sequence using the first generating model and output an inference result containing a predetermined number of inferences.

[0117] According to one embodiment of the present disclosure, one or more instructions are executed by at least one processor (120) so that an electronic device (100) may evaluate a preference for the inference result of a first generative model using a second generative model. The second generative model may be a large language model that performs a preference evaluation according to an input sequence providing the inference result of the first generative model and evaluation data.

[0118] The evaluation module (123) of the processor (120) may obtain an input sequence that provides the inference results and evaluation data of the first generative model. The input sequence may include the inference results of the first generative model as input data in a prompt containing evaluation data. The evaluation module (123) may input the inference results and evaluation data of the first generative model into the second generative model to evaluate the preference for the inference results of the first generative model. The evaluation module (123) may output a high score or a high-ranking evaluation for inferences corresponding to preferred responses among the inference results of the first generative model, and output a low score or a low-ranking evaluation for inferences corresponding to non-preferred responses.

[0119] According to one embodiment of the present disclosure, by executing one or more instructions by at least one processor (120), the electronic device (100) may be enabled to select a plurality of response pairs consisting of a non-preferred response and a preferred response from an inference result based on an evaluation result. The training data generation module (125) of the processor (120) may select a plurality of response pairs consisting of an inference evaluated as a non-preferred response and an inference evaluated as a preferred response from among inference results including a predetermined number of inferences.

[0120] According to one embodiment of the present disclosure, one or more instructions are executed by at least one processor (120), thereby enabling an electronic device (100) to generate preference learning data by processing a plurality of response pairs. A learning data generation module (125) of the processor (120) can process at least one of a non-preferred response and a preferred response included in the plurality of response pairs.

[0121] For example, the training data generation module (125) of the processor (120) can process the non-preferred responses constituting multiple response pairs by using the second generation model to identify the parts of the non-preferred responses that result in deductions in the evaluation and filtering the non-preferred responses so that the deducted parts remain. The training data generation module (125) of the processor (120) can process the preferred responses by using the second generation model to identify the parts of the preferred responses that result in deductions in the evaluation and filtering the preferred responses so that the deducted parts are removed, and supplementing the filtered preferred responses with the parts that result in additions in the evaluation. The training data generation module (125) of the processor (120) can evaluate the preference of the supplemented preferred responses using the second generation model and, based on the evaluation results of the supplemented preferred responses, determine whether to further apply a process of processing the preferred responses to the supplemented preferred responses.

[0122] FIG. 15 is a block diagram for explaining in detail the operation of a learning data generation module (125) in an electronic device (100) that generates learning data for preference optimization learning according to one embodiment of the present disclosure.

[0123] The learning data generation module (125) may include a response pair selection module (125-1), a deduction factor verification module (125-2), a filtering module (125-3), a bonus factor supplementation module (125-4), and a learning data storage module (125-5), but is not limited to the names of each module, and unlike what is shown in FIG. 15, some modules may be integrated modules or separated into more subdivided modules.

[0124] The response pair selection module (125-1) can determine the inference result of the first generative model that received the best evaluation score from the evaluation results of the second generative model as the preferred response. The response pair selection module (125-1) can determine the inference result that received the minimum evaluation score and the inference result that received the mode evaluation score from the evaluation results of the second generative model as the non-preferred response. The response pair selection module (125-1) can select a first response pair consisting of a non-preferred response having the minimum value of the evaluation result and a preferred response having the maximum value of the evaluation result, and a second response pair consisting of a non-preferred response having the mode of the evaluation result and a preferred response having the maximum value of the evaluation result.

[0125] The deduction factor verification module (125-2) can identify the part that is deducted in the evaluation of the non-preferred response included in a plurality of response pairs using a second generative model. For example, the deduction factor verification module (125-2) can identify the part that is deducted first in the sentence of the non-preferred response by including the sentence of the non-preferred response as input data in the prompt and using the second generative model.

[0126] The deduction factor verification module (125-2) can identify the part of the preferred response included in a plurality of response pairs that is subject to a deduction in evaluation by using a second generation model. For example, the deduction factor verification module (125-2) can include the sentence of the preferred response as input data in the prompt and identify the part of the sentence of the preferred response that is initially subject to a deduction by using a second generation model.

[0127] The filtering module (125-3) can filter out unfavorable responses so that the parts subject to deduction remain. For example, the filtering module (125-3) can keep the unfavorable response up to the first part subject to deduction in the sentence of the unfavorable response. The filtering module (125-3) can keep the unfavorable response up to the sentence containing the first part subject to deduction among the sentences corresponding to the unfavorable response. Finally, the processed unfavorable response can be stored in the training data storage module (125-5) as preference training data for preference optimization training.

[0128] The filtering module (125-3) can filter the preferred response so that the part that causes deduction is removed. For example, the filtering module (125-3) can keep the preferred response up to the part that causes the first deduction in the sentence of the preferred response, and exclude the part that causes the first deduction from the preferred response. The filtering module (125-3) can keep the preferred response up to the sentence containing the part that causes the first deduction among the sentences corresponding to the preferred response.

[0129] The bonus element supplement module (125-4) can supplement the parts that add to the evaluation of the filtered preference response. The bonus element supplement module (125-4) can generate a supplemented preference response by including the sentence of the filtered preference response as input data in the prompt and adding a sentence that adds to the end of the sentence of the filtered preference response using a second generative model. The bonus element supplement module (125-4) can evaluate the preference of the supplemented preference response using a second generative model. For example, the bonus element supplement module (125-4) can evaluate the preference of the sentence of the supplemented preference response using a second generative model by including the sentence of the supplemented preference response as input data in the prompt. Based on the evaluation result of the supplemented preference response, the bonus element supplement module (125-4) can determine whether to further apply a process of processing the preference response to the supplemented preference response. The bonus element supplement module (125-4) can terminate the process of processing the preference response if the evaluation result of the preference for the supplemented preference response satisfies a predefined criterion. Finally, the processed preference response can be stored in the training data storage module (125-5) as preference training data for preference optimization training.

[0130] As illustrated in FIG. 14, when a learning module (127) is implemented in a processor (120), the learning module (127) can perform preference optimization learning for a first generative model using preference learning data generated by a learning data generation module (125). Since the preference learning data generated by the learning data generation module (125) is generated based on the inference results of the first generative model and the preference evaluation results for the inference results of the first generative model, it can be most effective for preference optimization learning of the first generative model.

[0131] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include computer storage media and communication media. Computer storage media include both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include other data of modulated data signals, such as computer-readable instructions, data structures, or program modules.

[0132] Additionally, computer-readable storage media may be provided in the form of non-transitory storage media. Here, 'non-transitory storage media' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, 'non-transitory storage media' may include a buffer in which data is stored temporarily.

[0133] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0134] According to one embodiment of the present disclosure, a method for generating training data for preference optimization learning is provided. The method for generating training data for preference optimization learning may include the step (S310) of inputting the same input sequence to a first generative model a predetermined number of times and inferring a predetermined number of times using the first generative model. Additionally, the method for generating training data for preference optimization learning may include the step (S320) of evaluating the preference for the inference result of the first generative model using a second generative model. Additionally, the method for generating training data for preference optimization learning may include the step (S330) of selecting a plurality of response pairs, each containing a non-preferred response and a preferred response, from the inference result based on the preference evaluation result for the inference result of the first generative model. Additionally, the method for generating training data for preference optimization learning may include the step (S340) of generating training data by processing the selected plurality of response pairs.

[0135] Additionally, according to one embodiment of the present disclosure, the step of selecting a plurality of response pairs (S330) may select a first response pair consisting of a non-preferred response having a minimum value of the preference evaluation result and a preferred response having a maximum value of the preference evaluation result, and a second response pair consisting of a non-preferred response having a mode of the preference evaluation result and a preferred response having a maximum value of the preference evaluation result.

[0136] Additionally, according to one embodiment of the present disclosure, the step of generating training data (S340) may include a step of processing non-preferred responses included in each of a plurality of response pairs. The step of processing non-preferred responses may include a step (S820) of identifying parts of the non-preferred responses that result in deductions in the evaluation using a second generative model. Additionally, the step of processing non-preferred responses may include a step (S830) of filtering the non-preferred responses such that the parts of the non-preferred responses that do not cause deductions are removed, and the parts of the non-preferred responses that result in deductions remain.

[0137] Additionally, according to one embodiment of the present disclosure, the step of generating training data (S340) may include a step of processing preferred responses included in each of a plurality of response pairs. The step of processing preferred responses may include a step (S840) of identifying parts of the preferred responses that result in deductions in the evaluation using a second generative model. Additionally, the step of processing preferred responses may include a step (S850) of filtering preferred responses so that parts of the preferred responses that result in deductions are removed while parts of the preferred responses that result in deductions are retained. Additionally, the step of processing preferred responses may include a step (S860) of supplementing parts of the filtered preferred responses that result in additions in the evaluation.

[0138] Additionally, according to one embodiment of the present disclosure, the step of processing the preference response may include a step (S870) of evaluating the preference for the supplemented preference response using a second generative model. Additionally, the step of processing the preference response may include a step (S880) of determining whether to apply the step of processing the preference response to the supplemented preference response based on the evaluation result of the supplemented preference response.

[0139] Additionally, according to one embodiment of the present disclosure, the inference step (S310) may include a step of adjusting one or more settings regarding the degrees of freedom of the output of the first generating model so that different inference results are generated for the same input sequence.

[0140] Additionally, according to one embodiment of the present disclosure, the first generative model may be a large language model on which Supervised Fine Tuning (SFT) learning has been performed. Additionally, the second generative model may be a large language model that performs preference evaluation according to an input sequence providing inference results and evaluation data of the first generative model.

[0141] According to one embodiment of the present disclosure, a computer-readable recording medium is provided on which a program for executing a method for generating training data for the above preference optimization learning is recorded.

[0142] According to one embodiment of the present disclosure, an electronic device (100) for generating training data for preference optimization learning is provided. The electronic device (100) may include a memory (110) comprising one or more storage media for storing instructions and at least one processor (120) comprising a processing circuit. By executing one or more instructions by the at least one processor (120), the electronic device (100) may input the same input sequence to a first generative model a predetermined number of times and infer a predetermined number of times using the first generative model. Additionally, by executing one or more instructions by the at least one processor (120), the electronic device (100) may evaluate the preference for the inference result of the first generative model using a second generative model. Additionally, by executing one or more instructions by at least one processor (120), the electronic device (100) can select a plurality of response pairs, each including a non-preferred response and a preferred response, from the inference result based on the preference evaluation result of the inference result of the first generative model. Additionally, by executing one or more instructions by at least one processor (120), the electronic device (100) can generate preference learning data by processing the selected plurality of response pairs.

[0143] Additionally, according to one embodiment of the present disclosure, by executing one or more instructions by at least one processor (120), the electronic device (100) can select a first pair of responses consisting of a non-preferred response having a minimum value of the preference evaluation result and a preferred response having a maximum value of the preference evaluation result, and a second pair of responses consisting of a non-preferred response having a mode of the preference evaluation result and a preferred response having a maximum value of the preference evaluation result.

[0144] Additionally, according to one embodiment of the present disclosure, by executing one or more instructions by at least one processor (120), the electronic device (100) can process a plurality of response pairs of unfavorable responses by using a second generation model to identify a part of the unfavorable response that is subject to a deduction in evaluation, removing the part of the unfavorable response that does not cause the deduction, and filtering the unfavorable response so that the part of the unfavorable response that is subject to the deduction remains.

[0145] Additionally, according to one embodiment of the present disclosure, by executing one or more instructions by at least one processor (120), the electronic device (100) can process a plurality of response pairs by using a second generation model to identify a part of the preferred response that is deducted from the evaluation, filtering the preferred response so that the part of the preferred response that is deducted is removed while the part of the preferred response that does not cause the deduction is retained, and supplementing the filtered preferred response with a part that is added to the evaluation.

[0146] Additionally, according to one embodiment of the present disclosure, by executing one or more instructions by at least one processor (120), the electronic device (100) can evaluate the preference for the supplemented preference response using a second generation model and, based on the evaluation result for the supplemented preference response, determine whether to apply a process for processing the preference response to the supplemented preference response.

[0147] Additionally, according to one embodiment of the present disclosure, by executing one or more instructions by at least one processor (120), the electronic device (100) can adjust one or more settings regarding the degrees of freedom of the output of the first generating model so that different inference results are generated for the same input sequence.

[0148] Additionally, according to one embodiment of the present disclosure, the first generative model may be a large language model on which Supervised Fine Tuning (SFT) learning has been performed. Additionally, the second generative model may be a large language model that performs preference evaluation according to an input sequence providing inference results and evaluation data of the first generative model.

[0149] The foregoing description of the present disclosure is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0150] The scope of the present disclosure is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present disclosure.

Claims

1. A step (S310) of inputting the same input sequence to a first generative model a predetermined number of times and inferring a predetermined number of times using the first generative model; A step (S320) of evaluating the preference for the inference result of the first generative model using the second generative model; Based on the evaluation results of the above evaluation, a step (S330) of selecting a plurality of response pairs, each including a non-preferred response and a preferred response, from the inference results; and A step (S340) of generating training data by processing the above plurality of response pairs; A method for generating training data for preference optimization learning that includes 2. In Paragraph 1, The above-mentioned selection step (S330) is, A method for selecting a first response pair consisting of a non-preferred response having the minimum value of the evaluation result and a preferred response having the maximum value of the evaluation result, and a second response pair consisting of a non-preferred response having the mode of the evaluation result and a preferred response having the maximum value of the evaluation result.

3. In Paragraph 1 or 2, The step (S340) of generating the above training data is, The method includes a step of processing the non-preferred response included in each of the plurality of response pairs, and the step of processing the non-preferred response is A step (S820) of identifying the part of the non-preferred response that is deducted from the evaluation using the second generation model above; and A method comprising the step (S830) of filtering the unpreferred response so that the portion of the unpreferred response that does not cause the aforementioned deduction is removed and the portion of the unpreferred response that causes the aforementioned deduction remains.

4. In any one of paragraphs 1 to 3, The step (S340) of generating the above training data is, The step of processing the preferred response included in each of the plurality of response pairs is included, and the step of processing the preferred response is A step (S840) of identifying the parts of the preference response that result in a deduction in evaluation using the second generation model above; A step of filtering the preferred response (S850) so that the part of the preferred response that does not cause the aforementioned deduction is retained and the part of the preferred response that causes the aforementioned deduction is removed; and A method comprising the step (S860) of supplementing the part that adds points to the evaluation in the filtered preference response.

5. In any one of paragraphs 1 to 4, The step of processing the above preference response is, A step (S870) of evaluating the preference for the supplemented preference response using the second generation model; and A method further comprising a step (S880) of determining whether to apply a step of processing the preferred response to the supplemented preferred response based on the evaluation result of the supplemented preferred response.

6. In any one of paragraphs 1 to 5, The above inference step (S310) is, A method further comprising the step of adjusting one or more settings regarding the degrees of freedom of the output of the first generating model so as to generate different inference results for the same input sequence.

7. In any one of paragraphs 1 through 6, The above-mentioned first generative model is a large language model on which Supervised Fine Tuning (SFT) has been performed, and A method in which the second generative model is a large language model that performs preference evaluation according to an input sequence providing inference results and evaluation data of the first generative model.

8. A computer-readable recording medium having a program recorded thereon for executing the method of any one of paragraphs 1 through 7.

9. Memory (110) comprising one or more storage media for storing instructions; and It includes at least one processor (120) including a processing circuit, and By executing one or more instructions by the above-mentioned at least one processor (120), an electronic device (100) that generates training data for preference optimization learning, The same input sequence is input to a first generative model a predetermined number of times, and inference is performed a predetermined number of times using the first generative model, and Evaluate the preference for the inference result of the first generative model using the second generative model, and Based on the evaluation results of the above evaluation, a plurality of response pairs, each including a non-preferred response and a preferred response, are selected from the inference results, and An electronic device (100) that generates learning data by processing the above multiple response pairs.

10. In Paragraph 9, By executing one or more instructions by the above at least one processor (120), the electronic device (100) is made to, An electronic device (100) for selecting a first response pair consisting of a non-preferred response having the minimum value of the evaluation result and a preferred response having the maximum value of the evaluation result, and a second response pair consisting of a non-preferred response having the mode of the evaluation result and a preferred response having the maximum value of the evaluation result.

11. In Paragraph 9 or 10, By executing one or more instructions by the above at least one processor (120), the electronic device (100) is made to, An electronic device (100) that processes the unfavorable responses constituting the plurality of response pairs by using the second generation model to identify the parts of the unfavorable responses that result in a deduction in evaluation, removing the parts of the unfavorable responses that do not cause the deduction, and filtering the unfavorable responses so that the parts of the unfavorable responses that result in the deduction remain.

12. In any one of paragraphs 9 through 11, By executing one or more instructions by the above at least one processor (120), the electronic device (100) is made to, An electronic device (100) that processes the preferred responses constituting the plurality of response pairs by using the second generation model to identify the parts of the preferred response that cause a deduction in the evaluation, filtering the preferred response so that the parts of the preferred response that cause a deduction are removed while the parts of the preferred response that do not cause a deduction are retained, and supplementing the parts of the filtered preferred response that cause an addition in the evaluation.

13. In any one of paragraphs 9 through 12, By executing one or more instructions by the above at least one processor (120), the electronic device (100) is made to, An electronic device (100) that evaluates the preference for the supplemented preference response using the second generation model and determines whether to apply a process for processing the preference response to the supplemented preference response based on the evaluation result of the supplemented preference response.

14. In any one of paragraphs 9 through 12, By executing one or more instructions by the above at least one processor (120), the electronic device (100) is made to, An electronic device (100) that evaluates the preference for the supplemented preference response using the second generation model and determines whether to apply a process for processing the preference response to the supplemented preference response based on the evaluation result of the supplemented preference response.

15. In any one of paragraphs 9 through 14, The above-mentioned first generative model is a large language model on which Supervised Fine Tuning (SFT) has been performed, and The above second generative model is an electronic device (100) which is a large language model that performs preference evaluation according to an input sequence providing inference results and evaluation data of the above first generative model.