Method for removing hallucination from inference results of neural network model and electronic device for performing same

The method and device address neural network hallucinations by context-based detection and correction, enhancing response reliability and model performance through feedback.

WO2025226076A1PCT designated stage Publication Date: 2025-10-30SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/005618
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Neural network models, particularly large-scale language models, suffer from hallucinations due to insufficient or incorrect training data, which current methods struggle to completely prevent, necessitating effective detection and removal techniques.

Method used

A method and electronic device that evaluate neural network responses based on context to detect and modify hallucinations, utilizing a response modification module with components like context extraction, evaluation item generation, and hallucination removal units to correct responses, and further enhance model reliability through feedback and reinforcement learning.

Benefits of technology

Effectively detects and removes hallucinations, improving the reliability and accuracy of neural network responses, and enhances model performance through feedback mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025005618_30102025_PF_FP_ABST
    Figure KR2025005618_30102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for removing hallucination from inference results of a neural network model may comprise the steps of: obtaining a response from a neural network model on the basis of a prompt provided to the neural network model; determining, on the basis of the context included in the prompt, whether hallucination has occurred in the response; and if it is determined that the hallucination has occurred, modifying and outputting the response.
Need to check novelty before this filing date? Find Prior Art

Description

Method for removing hallucination from inference results of neural network models and electronic device for performing the same

[0001] The present disclosure relates to a method for removing hallucinations from the inference results of a neural network model and an electronic device therefor. Specifically, the present disclosure relates to a method for detecting hallucinations by evaluating responses based on the context of a prompt and modifying the responses to remove hallucinations.

[0002] Neural network models, such as large-scale language models, are rapidly improving in performance and are being utilized in various fields. However, hallucinations can occur due to insufficient or incorrect training data or input data used during neural network model training. Methods to prevent hallucinations include improving the quality of training data or strengthening input data validation. However, despite these efforts, complete prevention of hallucinations remains challenging, requiring methods to detect and remove them.

[0003] According to one aspect of the present disclosure, a method for removing hallucination from an inference result of a neural network model may include a step of obtaining a response of the neural network model based on a prompt provided to the neural network model, a step of determining whether hallucination has occurred in the response based on a context included in the prompt, and a step of modifying and outputting the response if it is determined that hallucination has occurred.

[0004] According to one aspect of the present disclosure, an electronic device for removing hallucination from an inference result of a neural network model may be provided. The electronic device may include at least one processor, which includes an input / output interface for receiving a prompt for a neural network model and outputting a response of the neural network model to the prompt, a memory storing one or more instructions for detecting hallucination in the response, and a processing circuit. The at least one processor may execute the one or more instructions alone or in cooperation, thereby obtaining a response of the neural network model based on the prompt provided to the neural network model, determining whether hallucination has occurred in the response based on a context included in the prompt, and if it is determined that hallucination has occurred, modifying and outputting the response.

[0005] According to one aspect of the present disclosure, a non-transitory computer-readable recording medium storing one or more commands may be provided. When the one or more commands are executed by the one or more processors, the one or more processors may perform the neural network model.

[0006] According to one aspect of the present disclosure, a computer program may be stored in a non-transitory storage medium for performing at least one of the embodiments of the disclosed method on a computer.

[0007] FIG. 1 is a diagram illustrating configurations for performing a method for removing hallucination from an inference result of a neural network model according to one embodiment of the present disclosure.

[0008] FIG. 2 is a diagram illustrating the configuration of a response modification module according to one embodiment of the present disclosure.

[0009] FIG. 3 is a diagram illustrating a configuration of an electronic device according to one embodiment of the present disclosure.

[0010] FIG. 4 is a diagram illustrating a configuration for performing reinforcement learning on a neural network model based on a modified response according to one embodiment of the present disclosure.

[0011] FIG. 5 is a diagram illustrating an example in which a neural network model outputs a response corresponding to a prompt according to one embodiment of the present disclosure.

[0012] FIG. 6 is a diagram illustrating an example of a context extraction unit extracting a key context or key token from a prompt according to one embodiment of the present disclosure.

[0013] FIG. 7 is a diagram illustrating an example in which an evaluation item generation unit generates an evaluation item based on a key context or a key token according to one embodiment of the present disclosure.

[0014] FIG. 8 is a diagram illustrating an example of detecting hallucination by having an evaluation performing unit perform an evaluation on a response for each evaluation item according to one embodiment of the present disclosure.

[0015] FIG. 9 is a diagram illustrating an example of a hallucination removal unit removing hallucination from a response according to one embodiment of the present disclosure.

[0016] FIG. 10 is a diagram illustrating an example in which a neural network model outputs a response corresponding to a prompt according to one embodiment of the present disclosure.

[0017] FIG. 11 is a diagram illustrating an example of a context extraction unit extracting a key context from a prompt according to one embodiment of the present disclosure.

[0018] FIG. 12 is a diagram illustrating an example in which an evaluation item generation unit generates an evaluation item based on a key context according to one embodiment of the present disclosure.

[0019] FIG. 13 is a diagram illustrating an example of detecting hallucination by having an evaluation performing unit perform an evaluation on a response for each evaluation item according to one embodiment of the present disclosure.

[0020] FIG. 14 is a diagram illustrating an example of a hallucination removal unit removing hallucination from a response according to one embodiment of the present disclosure.

[0021] FIGS. 15 to 19 are flowcharts illustrating a method for removing hallucination from inference results of a neural network model according to embodiments of the present disclosure.

[0022] In this disclosure, the expression “at least one of a, b or c” may refer to “a”, “b”, “c”, “a and b”, “a and c”, “b and c”, “all of a, b and c”, or variations thereof.

[0023] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are to be understood to include plural referents. Thus, for example, reference to "a surface of a composition" may also include reference to one or more of such surfaces.

[0024] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. This is to avoid obscuring the gist of this disclosure by omitting unnecessary explanations and to convey it more clearly. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification as a whole.

[0025] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.

[0026] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.

[0027] It should be understood that the blocks and combinations of flowcharts in each of the flowcharts in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.

[0028] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.

[0029] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.

[0030] All functions or operations described in the present disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors may be a circuitry that performs processing, and may include circuitry such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), an Integrated Chip (IC), etc.

[0031] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.

[0032] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0033] According to one embodiment of the present disclosure, when a prompt is input to a neural network model and a response is output, an electronic device can evaluate the response based on the prompt to determine whether hallucination has occurred in the response, and if hallucination has occurred in the response, the electronic device can modify and output the response. Specifically, the electronic device can evaluate the response according to evaluation items determined based on a context included in the prompt to determine whether the response includes an error, and if the response includes an error, the electronic device can remove the error included in the response and output it.

[0034] Additionally, an electronic device according to one embodiment can improve the reliability of a neural network model by providing feedback to the neural network model using the modified response. Specifically, the electronic device can retrain the neural network model using training data including the prompt and the modified response.

[0035] First, the process of detecting and removing hallucinations in a response is described in its entirety with reference to FIGS. 1 and 2, and the configuration of an electronic device that performs the operation of detecting and removing hallucinations in a response is described with reference to FIG. 3. Then, the process of retraining a neural network model based on a modified response is described with reference to FIG. 4.

[0036] The following describes in detail the process of detecting and removing hallucinations, with specific examples of prompts and responses (Figures 5 to 14).

[0037] 1. Detect and remove hallucinations in responses based on the context contained in the prompt.

[0038] FIG. 1 is a diagram illustrating configurations for performing a method for removing hallucinations from inference results of a neural network model according to one embodiment of the present disclosure. Referring to FIG. 1, a process for detecting and removing hallucinations from a response (20) output in response to a prompt (10) will be described.

[0039] The neural network model (100), response modification module (200), and detailed components (210, 211, 212, 213, 220) included in the response modification module (200) illustrated in FIG. 1 may be components classified based on function or role. The components (100, 200, 210, 211, 212, 213, 220) illustrated in FIG. 1 may be software components implemented by the processor (320) of the electronic device (300) of FIG. 3, which will be described later, executing a program stored in the memory (330), and may also be virtual components in which no matching hardware device actually exists. In other words, the operations performed by the processor (320) of the electronic device (300) by executing the program stored in the memory (330) may be classified into a plurality of groups by function or purpose, and the entities performing the operations included in each classified group may be expressed as the components (100, 200, 210, 211, 212, 213, 220) of FIG. 1. Accordingly, the operations described as being performed by the components (100, 200, 210, 211, 212, 213, 220) of FIG. 1 may be viewed as actually being performed by the processor (320) of the electronic device (300) by executing the program stored in the memory (330).

[0040] At least one of the components, elements, modules and units (collectively referred to as “components” in this paragraph) illustrated in block form in FIGS. 1, 2, 8, 9, 13 and 14 may utilize direct circuit structures such as memories, processors, logic circuits, look-up tables, etc., and may perform corresponding functions under the control of one or more microprocessors or other control devices. In addition, at least one of these components may be embodied as a part of a module, program or code including one or more executable instructions for performing specific logical functions, and may be executed by one or more microprocessors or other control devices. In addition, at least one of these components may include or be implemented by a processor such as a central processing unit (CPU), a microprocessor, etc., that performs the corresponding functions.

[0041] Referring to FIG. 1, when a prompt (10) is input to a neural network model (100) according to one embodiment of the present disclosure, the neural network model (100) may perform inference and output a response (20). According to one embodiment of the present disclosure, the neural network model (100) may be configured to include a Large Language Model (LLM) and may answer a question, perform an action according to a request (e.g., writing an email), or perform a translation or summary, etc. Of course, the neural network model (100) is not limited thereto, and may be a model that has various forms of input and output and is trained to perform actions for various purposes. The prompt (10) or the response (20) may be in the form of text, an image, or various other forms of input and output.

[0042] Hallucination may occur in the response (20) due to insufficient or incorrect training data or input data (prompt (10)) used in learning the neural network model (100), or due to various other causes. For example, the response (20) may contain incorrect information. Or, for example, the information requested in the prompt (10) may not be reflected in the response (20), or conversely, information not requested in the prompt (10) may be included in the response (20). In addition, hallucination may occur in various forms, such as the response (20) being generated to contain information that does not fit the context of the prompt (10) or is false.

[0043] To prevent hallucination, methods such as improving the quality of training data or strengthening the validation of input data can be used. However, despite these efforts, it is difficult to completely prevent hallucination. Therefore, in embodiments of the present disclosure, the electronic device can detect and remove hallucination in the response (20).

[0044] The response modification module (200) can detect hallucination in the response (20) and output a modified response (30) by removing the detected hallucination.

[0045] The response modification module (200) may include a hallucination detection unit (210) and a hallucination removal unit (220), and the hallucination detection unit (210) may include a context extraction unit (211), an evaluation item generation unit (212), and an evaluation execution unit (213).

[0046] The input and output of the neural network model (100), i.e., the prompt (10) and the response (20), can be input to the response modification module (200). According to one embodiment of the present disclosure, the response modification module (200) can determine whether hallucination has occurred in the response (20) based on the context included in the prompt (10), and if hallucination has occurred, can remove the hallucination.

[0047] An example of a method for detecting hallucination in a response (20) by a hallucination detection unit (210) is described as follows.

[0048] The context extraction unit (211) can extract context from the prompt (10). According to one embodiment of the present disclosure, the context extraction unit (211) can extract at least one key context or key token from the prompt (10).

[0049] Key context may refer to a major context among the contexts included in a prompt (10). Key context may be a context that directly influences the generation of a response (20). Therefore, key context may be used to determine whether hallucination has occurred in a response (20).

[0050] A key token may refer to a key token among the tokens included in a prompt (10) and may be extracted from a key context. A key token may also be a token that directly influences the generation of a response (20). Therefore, a key token may be used to determine whether hallucination has occurred in a response (20).

[0051] An embodiment of extracting a key context or key token from a prompt (10) is described in detail below with reference to FIG. 6.

[0052] The assessment item generation unit (212) can generate assessment items based on the extracted context. For example, the assessment item generation unit (212) can generate at least one assessment item based on at least one key context or key token extracted from the prompt (10).

[0053] The evaluation item generation unit (212) can generate an assessment list containing multiple evaluation items and can also assign priorities to the multiple evaluation items. If priorities are assigned to multiple evaluation items, the evaluation execution unit (213) can determine whether hallucination has occurred by performing an evaluation by applying a weight corresponding to the priority to each evaluation item.

[0054] Evaluation items may include content that must be verified to determine whether hallucination has occurred. For example, evaluation items may include content that verifies whether a response (20) contains specific information, and if not, determines that hallucination has occurred in the response (20). Additionally, evaluation items may include content that determines whether hallucination has occurred based on various criteria or rules. One embodiment of a method for generating evaluation items based on context extracted from a prompt (10) is as follows.

[0055] According to one embodiment of the present disclosure, the evaluation item generation unit (212) can determine information to be included in the response (20) based on a key context or a key token, and can generate an evaluation item with content that determines that hallucination has occurred if the determined information is not included in the response (20). For example, if the neural network model (100) provides an email writing service, the evaluation item generation unit (212) can determine information to be included in the email (response (20)) based on a key context or a key token extracted from the prompt (10). Accordingly, the evaluation item generation unit (212) can generate an evaluation item with content that determines that hallucination has occurred if the determined information is not included in the email (response (20)).

[0056] According to one embodiment of the present disclosure, the evaluation item generation unit (212) may generate an evaluation item that determines that hallucination has occurred if the response (20) does not include content corresponding to the key context or key token. For example, if the neural network model (100) provides a service that answers questions, the evaluation item generation unit (212) may generate an evaluation item that determines that hallucination has occurred if the response (20) does not include an answer to the question (key context or key token) included in the prompt (10).

[0057] According to one embodiment of the present disclosure, the evaluation item generation unit (212) may generate an evaluation item for determining that hallucination has occurred if the response (20) includes content that does not match the key context or key token. For example, if the neural network model (100) provides a service for translating or summarizing text, the evaluation item generation unit (212) may generate an evaluation item for determining that hallucination has occurred if the response (20) includes content that does not match the content (key context or key token) of the text to be translated or summarized.

[0058] According to one embodiment of the present disclosure, the evaluation item generation unit (212) can check the reference document (source document) used when generating a response based on a key context or a key token, and generate an evaluation item with content that determines that hallucination has occurred if the response (20) includes content that does not match the reference document. For example, if the neural network model (100) provides a service that answers questions, the evaluation item generation unit (212) can check the document that the neural network model (100) referenced when generating an answer to a question included in a prompt (10), and generate an evaluation item with content that determines that hallucination has occurred if the response (20) includes content that does not match the reference document.

[0059] The evaluation performing unit (213) can perform an evaluation on the response (20) for each evaluation item and determine whether hallucination has occurred in the response (20) based on the evaluation result. According to one embodiment of the present disclosure, the evaluation performing unit (213) can determine that the evaluation result is 'failure' if the response (20) does not satisfy the condition required by the evaluation item. Conversely, the evaluation performing unit (213) can determine that the evaluation result is 'success' if the response (20) satisfies the condition required by the evaluation item.

[0060] If there is only one evaluation item output from the evaluation item generation unit (212), the evaluation execution unit (213) can determine whether hallucination has occurred by considering only the evaluation result of the corresponding evaluation item. According to one embodiment of the present disclosure, the evaluation execution unit (213) can determine that hallucination has not occurred in the response (20) if the evaluation result of the evaluation item is 'success', and conversely, if the evaluation result of the evaluation item is 'failure', it can determine that hallucination has occurred in the response (20). For example, the evaluation execution unit (213) can check whether specific information is included in the response (20) depending on the evaluation item, and if the specific information is not included, it can determine that hallucination has occurred in the response (20).

[0061] As previously explained, the number of evaluation items generated by the evaluation item generation unit (212) may be multiple. If the number of evaluation items output from the evaluation item generation unit (212) is multiple, the evaluation execution unit (213) may perform an evaluation for each evaluation item and then comprehensively consider the evaluation results of the evaluation items to determine whether hallucination has occurred. To this end, rules or criteria for determining whether hallucination has occurred may be prepared based on the evaluation results for multiple evaluation items.

[0062] According to one embodiment of the present disclosure, when there are multiple evaluation items, the evaluation performing unit (213) can determine that hallucination has occurred if the evaluation result of any one of the multiple evaluation items is 'failure'.

[0063] Or, according to one embodiment of the present disclosure, when there are multiple evaluation items, the evaluation performing unit (213) obtains the results (e.g., scores) of the evaluation performed for each evaluation item, and compares the sum of the evaluation results (evaluation scores) with a preset threshold value to determine whether hallucination has occurred in the response (20).

[0064] For example, if the condition required by the evaluation item (e.g., "Does the response (20) contain specific information?") is satisfied, the evaluation result (evaluation score) can be assumed to be 0, and if the condition required by the evaluation item is not satisfied, the evaluation result (evaluation score) can be assumed to be 1. The evaluation performing unit (213) can perform an evaluation on the response (20) for multiple evaluation items, and then add up all the evaluation results (0 or 1) obtained for each evaluation item and compare the result with a preset threshold value (e.g., 2). If the result of adding up all the evaluation results is greater than the preset threshold value, the evaluation performing unit (213) can determine that hallucination has occurred in the response (20).

[0065] As previously explained, priorities can be assigned to multiple evaluation items generated by the evaluation item generation unit (212), and the evaluation execution unit (213) can perform evaluation by applying a weight corresponding to the priority to each evaluation item.

[0066] For example, the evaluation performing unit (213) can determine whether hallucination has occurred by multiplying the evaluation result (evaluation score) for each evaluation item by a weight corresponding to the evaluation item, and then adding up all of the multiplied weights for the evaluation results (evaluation scores) and comparing the result with a preset threshold value.

[0067] If the hallucination detection unit (210) determines that hallucination has occurred in the response (20), the hallucination removal unit (220) can remove the hallucination by modifying the response (20). According to one embodiment of the present disclosure, the hallucination removal unit (220) can remove the hallucination by modifying the response (20) based on at least one of the context (key context or key token) or the evaluation items. An exemplary method of the hallucination removal unit (220) modifying the response (20) will be described in detail below with reference to the examples illustrated in FIGS. 9 and 14.

[0068] The response modification module (200) can output a modified response (30) with hallucination removed.

[0069] 2. Apply the tone determined based on the prompt to the response.

[0070] According to one embodiment of the present disclosure, the response modification module (200) may further include configurations for changing the tone of the response (20). The configuration of the response modification module (200) according to one embodiment of the present disclosure is illustrated in FIG. 2.

[0071] Referring to FIG. 2, the response modification module (200) may additionally include a tone determination unit (230) and a tone application unit (240) compared to the embodiment illustrated in FIG. 1.

[0072] The tone determination unit (230) can determine the tone to be applied to the response (20) based on the context included in the prompt (10). The tone determination unit (230) can determine the tone to be applied to the response (20) based on one or more factors that can be judged or predicted based on the context of the prompt (10).

[0073] According to one embodiment of the present disclosure, the tone determination unit (230) can determine the tone to be applied to the response (20) according to the target to be provided with the response (20).

[0074] For example, the tone determination unit (230) may infer a relationship between a user (requester) who inputs the prompt (10) and a recipient (recipient) who will receive the response (20) based on the context included in the prompt (10), and determine a tone according to the inferred relationship. The tone determination unit (230) may determine a tone to be applied to the response (10) according to which of the requesters is older than the recipient. Alternatively, the tone determination unit (230) may determine a tone to be applied to the response (10) according to the relationship between the recipient and the requester (e.g., superior-subordinate relationship in a company, teacher-student relationship, family relationship, etc.).

[0075] Or, for example, the tone determination unit (230) may predict the age group of the subject to be provided with the response (20) based on the context included in the prompt (10) and determine the tone according to the predicted age group.

[0076] According to one embodiment of the present disclosure, the tone determination unit (230) may determine a tone to be applied to the response (20) based on the tone of the prompt (10). For example, the tone determination unit (230) may determine that the same tone as the tone of the prompt (10) is to be applied to the response (20). Alternatively, for example, the tone determination unit (230) may determine that a tone that is more polite and formal than the tone of the prompt (10) is to be applied to the response (20).

[0077] According to one embodiment of the present disclosure, the tone determination unit (230) may determine a tone to be applied to a response (20) based on the request of the prompt (10). For example, the tone determination unit (230) may determine the request of the prompt (10) based on the context included in the prompt (10) (e.g., whether it requests an answer regarding specialized knowledge or requests the provision of information in a light conversational format, etc.), and determine a tone based on the determined request.

[0078] When the tone determination unit (230) determines a tone to be applied to the response (20), the tone application unit (240) can modify the response (20) according to the determined tone and output a modified response (30). The tone application unit (240) can receive the response (20) from which hallucination has been removed from the hallucination removal unit (220) and modify the expression of the response (20) according to the determined tone. For example, if the tone determined by the tone determination unit (230) is a 'friendly tone', the tone application unit (240) can modify the sentences included in the response (20) into friendly expressions and output a modified response (30).

[0079] 3. Configuration of electronic devices

[0080] FIG. 3 is a diagram illustrating a configuration of an electronic device according to one embodiment of the present disclosure.

[0081] An electronic device (300) according to one embodiment of the present disclosure may be a computing device with computational processing capabilities. For example, the electronic device (300) may be a server managed by a service operator providing a service utilizing a neural network model, or may be a computing device such as a desktop or laptop. Of course, the electronic device (300) is not limited thereto, and may be various types of devices, such as a smartphone or tablet.

[0082] As previously described, the configurations (100, 200, 210, 211, 212, 213, 220, 230, 240) illustrated in FIGS. 1 and 2 may be software configurations implemented by the processor (320) of the electronic device (300) of FIG. 3 executing a program stored in the memory (330).

[0083] Although only one electronic device (300) is illustrated in FIG. 3, according to one embodiment of the present disclosure, the configurations (100, 200, 210, 211, 212, 213, 220, 230, 240) illustrated in FIGS. 1 and 2 may be software configurations implemented by a combination of two or more electronic devices.

[0084] Hereinafter, the configuration of an electronic device (300) according to one embodiment of the present disclosure will be described with reference to FIG. 3. Referring to FIG. 3, the electronic device (300) may include an input / output interface (310), a processor (320), and a memory (330).

[0085] The input / output interface (310) may include an input interface (e.g., communication port, touch screen, keyboard, hard button, microphone, etc.) for receiving control commands, execution requests, data, etc. from a user or an external device, and an output interface (e.g., communication port, display panel, speaker, etc.) for outputting the result of an operation according to a request or the status of the electronic device (300).

[0086] According to one embodiment of the present disclosure, the electronic device (300) can receive a prompt from a user or an external device through an input / output interface (310), and can output a response generated according to the prompt or a modified response through the input / output interface (310).

[0087] Although not shown in FIG. 3, an electronic device (300) according to one embodiment of the present disclosure may further include a communication interface for performing wired or wireless communication with an external device or network.

[0088] The processor (320) controls a series of processes to operate the electronic device (300) according to the embodiments described below, and may be composed of one or more processors. The one or more processors included in the processor (320) may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. The one or more processors included in the processor (320) may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), a Digital Signal Processor (DSP), a graphics-only processor such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), an artificial intelligence-only processor such as a Neural Processing Unit (NPU), or a communication-only processor such as a Communication Processor (CP). When the one or more processors included in the processor (320) are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0089] The processor (320) can write data to the memory (330) or read data stored in the memory (330), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program or at least one instruction stored in the memory (330). Accordingly, the processor (320) can perform operations described in the following embodiments, and operations described as being performed by the electronic device (300) or software components implemented by the electronic device (300) (100, 200, 210, 211, 212, 213, 220, 230, 240 of FIGS. 1 and 2) in the following embodiments can be regarded as being performed by the processor (320) unless otherwise specified.

[0090] The memory (330) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (330) may not exist separately and may be configured to be included in the processor (320). The memory (330) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program or at least one instruction for performing operations according to embodiments described below may be stored in the memory (330). The memory (330) may also provide stored data to the processor (320) at the request of the processor (320).

[0091] 4. Additional training of the neural network model (feedback based on modified responses)

[0092] According to one embodiment of the present disclosure, the neural network model (100) of FIG. 1 may be additionally trained based on the correction response (30). The electronic device (300) may perform fine tuning or reinforcement learning on the neural network model (100) based on the correction response (30).

[0093] The electronic device (300) can perform fine tuning by retraining the neural network model (100) using training data including prompts (10) and correction responses (30).

[0094] The electronic device (300) may obtain a reward based on the modified response (30) and perform reinforcement learning on the neural network model (100) based on the obtained reward. The electronic device (300) may obtain a score measuring the quality of the response (20) or the modified response (30) as a reward and perform reinforcement learning on the neural network model (100) in a direction that increases the reward. A method by which the electronic device (300) performs reinforcement learning on the neural network model (100) will be described in detail below with reference to FIG. 4.

[0095] FIG. 4 is a diagram illustrating a configuration for performing reinforcement learning on a neural network model (100) based on a modified response (30) according to one embodiment of the present disclosure. The neural network model (100), the response modification module (200), and the quantitative assessment module (400) may be a software configuration implemented by a processor (320) of an electronic device (300) executing a program or instruction stored in a memory (330).

[0096] The electronic device (300) can obtain a reward (40) based on the modified response (30) output from the response modification module (200) and the evaluation result output from the quantitative evaluation module (400). The quantitative evaluation module (400) can receive the modified response (30) from the response modification module (200) and perform a quantitative evaluation on the modified response (30) to thereby calculate a score for factual consistency.

[0097] The electronic device (300) can obtain a reward (40) by inputting the score for factual consistency produced by the quantitative evaluation module (400) and the corrected response (30) into a reward function. At this time, the reward function used may be a function that quantifies the reliability, harmlessness, and usefulness of the corrected response (30) and produces a score. The electronic device (300) can obtain a score that quantifies the quality of the response (20) or the corrected response (30) by using the reward function.

[0098] The electronic device (300) can perform reinforcement learning on the neural network model (100) using the acquired reward (40). That is, the electronic device (300) can update the parameters included in the neural network model (100) in a direction that increases the reward (40).

[0099] Below, the process of detecting and removing hallucinations by an electronic device (300) is described in detail through specific examples of services utilizing a neural network model (100). Examples of a neural network model (100) writing an email according to a prompt are illustrated in FIGS. 5 to 9, and examples of a neural network model (100) retrieving information and responding to a request from a prompt are illustrated in FIGS. 10 to 14.

[0100] 5. Examples of detecting and removing hallucinations (Figs. 5 to 9)

[0101] A neural network model (100) according to one embodiment of the present disclosure can perform various actions in response to a prompt request. For example, the neural network model (100) may compose an email, translate, or summarize an email in response to a request contained in a prompt.

[0102] FIG. 5 is a diagram illustrating an example of a neural network model (100) outputting a response corresponding to a prompt according to one embodiment of the present disclosure. In FIG. 5, a prompt (510) input to the neural network model (100) includes content requesting the writing of an email. The prompt (510) may be input to the electronic device (300) through the input / output interface (310) of the electronic device (300) or through the interface of an external device connected to the electronic device (300) and transmitted to the neural network model (100).

[0103] The neural network model (100) can output an email written in response to a request from a prompt (510) as a response (520). As illustrated in FIG. 5, a response (520) in the format of an email, including the content requested by the prompt (510), can be output from the neural network model (100).

[0104] According to one embodiment of the present disclosure, the electronic device (300) may not provide the user with the response (520) initially output from the neural network model (100), but may only provide the user with the response after detection and removal of hallucinations are completed. Of course, the electronic device (300) may also output the response (520) through the input / output interface (310) or an external device, thereby allowing the user to confirm the process of modifying the response (520).

[0105] FIG. 6 is a diagram illustrating an example of a context extraction unit (211) extracting a key context or key token from a prompt according to one embodiment of the present disclosure. In the embodiment illustrated in FIG. 6, the context extraction unit (211) extracts four key contexts from the prompt (510). The key contexts are major contexts among the contexts included in the prompt (510) and may be contexts that directly influence the generation of a response (520).

[0106] The context extraction unit (211) can analyze the contents of the prompt (510) and extract the contents to be included in the email as key context. Referring to Fig. 6, the context extraction unit (211) can extract the main contents that the user (the person entering the prompt) wants to convey to the recipient of the email as key context from the prompt (510).

[0107] The context extraction unit (211) may have criteria or rules set in advance for extracting key context from the prompt (510), and the context extraction unit (211) may be implemented as a neural network model.

[0108] The context extraction unit (211) can also extract key tokens from key contexts. Referring to FIG. 6, the context extraction unit (211) can extract key tokens by dividing each key context into multiple tokens.

[0109] Key tokens are major tokens among the tokens included in the prompt (510) and may be tokens that directly affect the generation of the response (520). The context extraction unit (211) may extract key tokens directly from the prompt (510). Criteria or rules for the context extraction unit (211) to extract key tokens from the prompt (510) may be preset, and the context extraction unit (211) may be implemented as a neural network model.

[0110] FIG. 7 is a diagram illustrating an example in which an evaluation item generation unit generates an evaluation item based on a key context or a key token according to one embodiment of the present disclosure.

[0111] Referring to FIG. 7, the evaluation item generation unit (212) can determine information to be included in an email (response (520)) based on a key context or key token, and can generate evaluation items that confirm whether the determined information is included in the response (520). To this end, the evaluation item generation unit (212) can generate evaluation items by combining a plurality of key contexts or key tokens.

[0112] The evaluation item generation unit (212) may generate evaluation items corresponding to each key context, generate evaluation items corresponding to two or more key contexts, or generate two or more evaluation items corresponding to one key context. In Fig. 7, the evaluation item generation unit (212) generates one evaluation item (evaluation item 1) corresponding to one key context (key context 1), and also generates one evaluation item (evaluation item 3) corresponding to two key contexts (key context 2, key context 3).

[0113] According to one embodiment of the present disclosure, the evaluation item generation unit (212) may additionally generate evaluation items for matters requiring verification depending on the service provided by the neural network model (100). Referring to FIG. 7, since the neural network model (100) provides an email writing service, the evaluation item generation unit (212) may generate evaluation items (evaluation item 2, evaluation item 5) for verifying the recipient and sender.

[0114] According to one embodiment of the present disclosure, the evaluation item generation unit (212) may assign priorities to multiple evaluation items (evaluation items 1 to 5). The degree to which each evaluation item contributes to or influences the detection of hallucinations may be determined based on the assigned priority.

[0115] FIG. 8 is a diagram illustrating an example of detecting hallucination by having an evaluation performing unit perform an evaluation on a response for each evaluation item according to one embodiment of the present disclosure.

[0116] Referring to FIG. 8, the evaluation performing unit (213) can perform an evaluation for each of multiple evaluation items for the response (520).

[0117] The evaluation performing unit (213) can check the contents described in the first area (810) of the response (520) and determine that the condition required for evaluation item 1 is satisfied. That is, the evaluation performing unit (213) can determine that the evaluation result for evaluation item 1 is 'success'. In addition, the evaluation performing unit (213) can check the contents described in the fourth area (840) of the response (520) and determine that the condition required for evaluation item 4 is satisfied. That is, the evaluation performing unit (213) can also determine that the evaluation result for evaluation item 4 is 'success'.

[0118] Since the second area (820) of the response (520) lacks information about the recipient of the email, the evaluation performing unit (213) may determine that the condition required for evaluation item 2 is not satisfied. In other words, the evaluation performing unit (213) may determine that the evaluation result for evaluation item 2 is 'failure'.

[0119] The evaluation execution unit (213) can check the contents described in the third area (830) of the response (520) and determine that the conditions required for evaluation item 3 are not satisfied. According to evaluation item 3, the time (4 PM) and location (minutes) of the meeting should be included in the response (520), but the third area (830) only states that the sender will attend the meeting to be held 'in the afternoon'. Therefore, the evaluation execution unit (213) can determine that the evaluation result for evaluation item 3 is 'failure'.

[0120] Since the information about the sender of the email is missing in the fifth area (850) of the response (520), the evaluation performing unit (213) may determine that the condition required for evaluation item 5 is not satisfied. In other words, the evaluation performing unit (213) may determine that the evaluation result for evaluation item 5 is 'failure'.

[0121] Since the evaluation results for evaluation items 2, 3, and 5 are 'failure', the evaluation performing unit (213) can determine that hallucination has occurred in the response (520). As described above, the evaluation performing unit (213) can determine that hallucination has occurred in the response (520) if the evaluation result for any one of the multiple evaluation items (evaluation items 1 to 5) is 'failure', or can obtain the evaluation result for each evaluation item and determine whether hallucination has occurred by comparing the sum of the evaluation results with a predetermined threshold value. In addition, the evaluation performing unit (213) can determine whether hallucination has occurred by applying a weight corresponding to the priority given to each evaluation item.

[0122] FIG. 9 is a diagram illustrating an example of a hallucination removal unit removing hallucination from a response according to one embodiment of the present disclosure.

[0123] The hallucination removal unit (220) can remove hallucinations by modifying the response (520). The hallucination removal unit (220) can modify the response (520) based on at least one of the context (key context or key token) or the evaluation items. For example, the hallucination removal unit (220) can modify the response (520) so that the response (520) includes the key context or key token. Alternatively, for example, the hallucination removal unit (220) can modify the response (520) so that it satisfies a condition required by the evaluation items.

[0124] Referring to FIG. 9, the hallucination removal unit (220) can add information about the recipient of the email to the second area (920) to satisfy the conditions required by evaluation item 2. In addition, the hallucination removal unit (220) can add information about the meeting time (4:00 PM) and location (minutes), which are the contents of the key context, to the third area (930a, 930b) to satisfy the conditions required by evaluation item 3. In addition, the hallucination removal unit (220) can add information about the sender of the email to the fifth area (950) to satisfy the conditions required by evaluation item 5.

[0125] The hallucination removal unit (220) can output a modified response (530) after removing hallucination by modifying the response (520). The electronic device (300) can output the modified response (530) through an input / output interface (310) or an external device.

[0126] 6. Examples of detecting and removing hallucinations (Figs. 10 to 14)

[0127] As previously described, the neural network model (100) may answer questions contained in a prompt or provide information requested by the prompt.

[0128] FIG. 10 is a diagram illustrating an example of a neural network model outputting a response corresponding to a prompt, according to one embodiment of the present disclosure. In FIG. 10, a prompt (1010) input to a neural network model (100) includes content requesting information. The prompt (1010) may be input through an input / output interface (310) of an electronic device (300) or through an interface of an external device connected to the electronic device (300) and transmitted to the neural network model (100).

[0129] The neural network model (100) can search for information according to a request from a prompt (1010) and output the search result as a response (1020). As illustrated in FIG. 10, a response (1020) including information on the 'childcare leave policy' requested in the prompt (1010) can be output from the neural network model (100).

[0130] According to one embodiment of the present disclosure, the electronic device (300) may not provide the user with the response (1020) initially output from the neural network model (100), but may only provide the user with the response after detection and removal of hallucinations are completed. Of course, the electronic device (300) may also allow the user to confirm the process of modifying the response (1020) by outputting the response (1020) through the input / output interface (310) or an external device.

[0131] FIG. 11 is a diagram illustrating an example of a context extraction unit extracting a key context from a prompt according to one embodiment of the present disclosure. In the embodiment illustrated in FIG. 11, the context extraction unit (211) extracts one key context from a prompt (1010). The extracted key context can indicate what information the prompt (1010) includes in its inquiry.

[0132] FIG. 12 is a diagram illustrating an example in which an evaluation item generation unit generates an evaluation item based on a key context according to one embodiment of the present disclosure.

[0133] Referring to FIG. 12, the evaluation item generation unit (212) can generate an evaluation item that verifies that the reference document used when generating a response (1020) based on the key context is a company rule and verifies whether the response (1020) includes content that does not match the company rules. The evaluation item generation unit (212) can check the reference document used by the neural network model (100) when generating the response (1020) from a database connected to the Internet or an electronic device (300).

[0134] FIG. 13 is a diagram illustrating an example of detecting hallucination by having an evaluation performing unit perform an evaluation on a response for each evaluation item according to one embodiment of the present disclosure.

[0135] Referring to FIG. 13, the evaluation performing unit (213) can perform an evaluation on evaluation item 1 for the response (1020). The evaluation performing unit (213) can confirm that the content described in the first area (1310) of the response (1020) is not included in the reference document, and can determine that the evaluation result for evaluation item 1 is 'failure'. Accordingly, the evaluation performing unit (213) can determine that hallucination has occurred in the response (1020).

[0136] FIG. 14 is a diagram illustrating an example of a hallucination removal unit removing hallucination from a response according to one embodiment of the present disclosure.

[0137] The hallucination removal unit (220) can remove hallucination by modifying the response (1020). The hallucination removal unit (220) can modify the response (1020) to satisfy the condition required by evaluation item 1 (that the response does not include content that is not in the reference document). Referring to FIG. 9, the hallucination removal unit (220) can remove content that is not in the standard from the response (1020) and output a modified response (1030). By comparing the first area (1410) of the modified response (1030) with the response (1020), it can be confirmed that content that is not in the standard has been deleted.

[0138] 7. Flowchart of exemplary methods and processes for detecting and removing hallucinations based on the context contained in the prompt.

[0139] Hereinafter, with reference to FIGS. 15 to 19, a method for removing hallucinations from the inference results of a neural network model according to embodiments of the present disclosure will be described. The steps included in the flowcharts of FIGS. 15 to 19 are performed by the configurations illustrated in FIGS. 1 to 4, and therefore, even if omitted below, the contents previously described with reference to FIGS. 1 to 4 can be equally applied to FIGS. 15 to 19.

[0140] Referring to FIG. 15, at step 1501, an electronic device can input a prompt to a neural network model and obtain a response from the neural network model. The neural network model can answer a question, perform an action based on the request (e.g., writing an email), or perform a translation or summary.

[0141] At step 1502, the electronic device can determine whether hallucination occurred in the response based on the context contained in the prompt. The detailed steps included in step 1502 are illustrated in FIG. 16.

[0142] Referring to FIG. 16, in step 1601, the electronic device can extract at least one key context or key token from the prompt. The key context may refer to a primary context among the contexts included in the prompt. The key context may be a context that directly affects the generation of a response. Therefore, the key context may be used to determine whether hallucination has occurred in the response. The key token may refer to a primary token among the tokens included in the prompt and may be extracted from the key context. The key token may also be a token that directly affects the generation of a response. Therefore, the key token may be used to determine whether hallucination has occurred in the response. Criteria or rules for extracting the key context or key token from the prompt may be preset.

[0143] In step 1602, the electronic device can generate at least one evaluation item based on at least one extracted key context or key token. Detailed steps included in step 1602 are illustrated in FIG. 17.

[0144] Referring to FIG. 17, in step 1701, the electronic device can determine information to be included in the response based on at least one extracted key context or key token. In step 1702, the electronic device can generate an evaluation item that determines that hallucination has occurred in the response if the determined information is not included in the response.

[0145] In addition, the electronic device may generate an evaluation item for determining that hallucination has occurred if the response does not contain content corresponding to at least one extracted key context or key token. Alternatively, the electronic device may generate an evaluation item for determining that hallucination has occurred if the response contains content that does not match at least one extracted key context or key token. Alternatively, the electronic device may check a reference document used in generating the response (20) based on at least one extracted key context or key token, and generate an evaluation item for determining that hallucination has occurred if the response contains content that does not match the reference document.

[0146] Returning to FIG. 16, in step 1603, the electronic device may evaluate the response for at least one evaluation item. In step 1604, the electronic device may determine whether hallucination has occurred in the response based on the evaluation results.

[0147] Returning to Figure 15, at step 1503, if hallucination occurs in the response, the electronic device can modify and output the response. The detailed steps included in step 1503 are illustrated in Figure 18.

[0148] Referring to FIG. 18, at step 1801, the electronic device may modify the response based on at least one evaluation item. For example, the electronic device may modify the response to include a key context or a key token. Alternatively, the electronic device may modify the response to satisfy a condition required by the evaluation item.

[0149] At step 1802, the electronic device can determine the tone based on the context contained in the prompt. The electronic device can determine the tone to be applied to the response based on one or more factors that can be judged or predicted based on the context of the prompt. For example, the electronic device can determine the tone based on the intended recipient of the response. Alternatively, the electronic device can determine the tone to be applied to the response based on the tone of the prompt. Alternatively, the electronic device can determine the tone to be applied to the response based on the request of the prompt.

[0150] At step 1803, the electronic device can apply the determined tone to the response.

[0151] Returning to Figure 15, at step 1504, the electronic device can further train a neural network model based on the modified response. The detailed steps involved in step 1504 are illustrated in Figure 19.

[0152] Referring to FIG. 19, at step 1901, the electronic device can perform a quantitative evaluation on the modified response to produce a score for factual consistency.

[0153] At step 1902, the electronic device can obtain a reward by applying the generated score and the modified response to a reward function. The reward function used here may be a function that quantifies the reliability, harmlessness, and usefulness of the modified response and produces a score. By utilizing the reward function, the electronic device can obtain a score that quantifies the quality of the response or modified response.

[0154] At step 1903, the electronic device can perform reinforcement learning on the neural network model using the acquired reward. The electronic device can update the parameters included in the neural network model in a way that increases the reward.

[0155] According to the embodiments described above, the electronic device determines whether hallucination has occurred in a response based on the context contained in the prompt, enabling rapid and accurate detection of hallucination. Furthermore, the electronic device modifies the response to remove the detected hallucination and outputs it, thereby improving the quality of services provided through the neural network model.

[0156] A method for removing hallucination from an inference result of a neural network model according to one embodiment of the present disclosure may include a step of inputting a prompt to a neural network model to obtain a response of the neural network model, a step of determining whether hallucination has occurred in the response based on a context included in the prompt, and a step of modifying and outputting the response if hallucination has occurred in the response.

[0157] According to one embodiment, the judging step may include extracting at least one key context or key token from the prompt, generating at least one assessment item based on the extracted at least one key context or key token, performing an evaluation on the response for each of the at least one assessment item, and judging whether hallucination has occurred in the response based on the evaluation result.

[0158] According to one embodiment, the step of generating the at least one evaluation item may include the step of determining information to be included in the response based on the extracted at least one key context or key token, and the step of generating the evaluation item, wherein if the determined information is not included in the response, it is determined that hallucination has occurred in the response.

[0159] According to one embodiment, the step of generating at least one evaluation item may generate an evaluation item that determines that hallucination has occurred in the response if the response does not contain content corresponding to at least one extracted key context or key token.

[0160] According to one embodiment, the step of generating at least one evaluation item may generate an evaluation item that determines that hallucination has occurred in the response if the response includes content that does not match at least one extracted key context or key token.

[0161] According to one embodiment, the step of generating the at least one evaluation item may include the step of checking a reference document used in generating the response based on the at least one extracted key context or key token, and the step of generating the evaluation item, which determines that hallucination has occurred if the response includes content that does not match the reference document.

[0162] According to one embodiment, the method may further include a step of additionally training the neural network model based on the modified response.

[0163] In one embodiment, the step of additionally training the neural network model may include retraining the neural network model using training data including the prompt and the modified response.

[0164] According to one embodiment, the step of additionally training the neural network model may include a step of performing a quantitative assessment on the modified response to calculate a score for factual consistency, a step of obtaining a reward by inputting the calculated score and the modified response into a reward function, and a step of performing reinforcement learning on the neural network model using the obtained reward.

[0165] In one embodiment, the step of modifying and outputting the response may include the step of modifying the response based on the at least one evaluation item, the step of determining a tone based on a context included in the prompt, and the step of applying the determined tone to the response.

[0166] An electronic device according to one embodiment of the present disclosure includes an input / output interface for receiving a prompt for a neural network model and outputting a response of the neural network model to the prompt, a memory storing a program or instruction for detecting hallucination in the response, and at least one processor, wherein the at least one processor executes the program or instruction stored in the memory, thereby allowing the electronic device to input the prompt to the neural network model to obtain a response of the neural network model, determine whether hallucination has occurred in the response based on a context included in the prompt, and then modify and output the response if hallucination has occurred in the response.

[0167] According to one embodiment, the electronic device may extract at least one key context or key token from the prompt to determine whether the hallucination has occurred, generate at least one assessment item based on the extracted at least one key context or key token, perform an evaluation on the response for each of the at least one assessment item, and then determine whether hallucination has occurred in the response based on the evaluation result.

[0168] According to one embodiment, the electronic device may generate an evaluation item by determining information to be included in the response based on the extracted at least one key context or key token when generating the at least one evaluation item, and then determining that hallucination has occurred in the response if the determined information is not included in the response.

[0169] According to one embodiment, the electronic device may generate an evaluation item that determines that hallucination has occurred in the response if, when generating the at least one evaluation item, the response does not include content corresponding to the at least one extracted key context or key token.

[0170] According to one embodiment, the electronic device may generate an evaluation item in which, when generating the at least one evaluation item, if the response includes content that does not match the at least one extracted key context or key token, it is determined that hallucination has occurred in the response.

[0171] According to one embodiment, the electronic device can generate an evaluation item by, when generating the at least one evaluation item, checking a reference document used in generating the response based on the at least one extracted key context or key token, and determining that hallucination has occurred if the response includes content that does not match the reference document.

[0172] According to one embodiment, the electronic device can additionally train the neural network model based on the modified response.

[0173] In one embodiment, the electronic device can retrain the neural network model using training data including the prompt and the modified response to further train the neural network model.

[0174] According to one embodiment, the electronic device may perform a quantitative assessment on the modified response to additionally train the neural network model, thereby calculating a score for factual consistency, and obtain a reward by applying the calculated score and the modified response to a reward function, and then perform reinforcement learning on the neural network model using the obtained reward.

[0175] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.

[0176] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.

[0177] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0178] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.

[0179] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. A method executed singly or cooperatively by at least one processor including a processing circuit, A step of obtaining a response of the neural network model based on a prompt provided to the neural network model; A step of determining whether hallucination has occurred in the response based on the context included in the prompt; and A method comprising a step of modifying and outputting the response when it is determined that the above hallucination has occurred.

2. In paragraph 1, The step of determining whether the above hallucination has occurred is: A step of extracting at least one key context or at least one key token from the above prompt; A step of generating at least one assessment item based on the at least one key context or the at least one key token; A step of performing an evaluation on the response for each of the at least one evaluation item; and A method characterized by comprising a step of determining whether hallucination has occurred in the response based on the above evaluation.

3. In either of paragraphs 1 and 2, The step of generating at least one evaluation item comprises: determining information to be included in the response based on at least one key context or at least one key token; and A method characterized by comprising a step of generating at least one evaluation item for determining that hallucination has occurred in the response if the determined information is not included in the response.

4. In any one of paragraphs 1 to 3, The step of generating at least one evaluation item comprises: A method characterized in that at least one evaluation item is generated, determining that the hallucination has occurred if the response does not contain content corresponding to at least one key context or at least one key token.

5. In any one of paragraphs 1 to 4, The step of generating at least one evaluation item comprises: A method characterized in that at least one evaluation item is generated, determining that hallucination has occurred in the response if the response contains content that does not match at least one key context or at least one key token.

6. In any one of paragraphs 1 to 5, The step of generating at least one evaluation item comprises: A step of verifying a reference document used in generating the response based on the at least one key context or the at least one key token; and A method comprising the step of generating at least one evaluation item, wherein hallucination is determined to have occurred if the response contains content that does not match the above reference document.

7. In any one of paragraphs 1 to 6, A method characterized by further comprising a step of training the neural network model based on the modified response.

8. In any one of paragraphs 1 to 7, The step of training the above neural network model is: A method characterized in that the neural network model is retrained using training data including the above prompt and the above modified response.

9. In any one of paragraphs 1 to 8, The step of training the above neural network model is: A step of calculating a score related to factual consistency based on a quantitative assessment of the above modified response; A step of obtaining a reward based on the above score, the modified response, and a reward function; and A method characterized by comprising a step of performing reinforcement learning on the neural network model based on the above reward.

10. In any one of paragraphs 1 to 9, The step of modifying and outputting the above response is: A step of modifying the response based on at least one evaluation item; a step of determining a tone based on the context included in the above prompt; and A method characterized by comprising a step of applying the tone to the response.

11. In the electronic device (300), An input / output interface (310) for receiving a prompt for a neural network model (100) and outputting a response of the neural network model (100) to the prompt; A memory (330) storing one or more commands for detecting hallucination in the above response; and At least one processor (320) comprising a processing circuit, The electronic device (300) executes the one or more instructions by the at least one processor (320) alone or in cooperation. Obtaining a response of the neural network model (100) based on the prompt provided to the neural network model (100), Based on the context included in the above prompt, determine whether hallucination occurred in the response. An electronic device that modifies and outputs the response when it is determined that the above hallucination has occurred.

12. In paragraph 11, The electronic device (300) determines whether the hallucination has occurred or not. Extract at least one key context or at least one key token from the above prompt, Generating at least one assessment item based on at least one key context or at least one key token, After evaluating the response for at least one of the above evaluation items, An electronic device characterized in that it determines whether hallucination has occurred in the response based on the above evaluation.

13. In any one of paragraphs 11 and 12, The electronic device (300) generates at least one evaluation item, After determining the information to be included in the response based on at least one key context or at least one key token, An electronic device characterized in that it generates at least one evaluation item for determining that hallucination has occurred in the response if the determined information is not included in the response.

14. In any one of paragraphs 11 to 13, The electronic device (300) generates at least one evaluation item, An electronic device characterized in that it generates at least one evaluation item for determining that the hallucination has occurred if the response does not include content corresponding to the at least one key context or the at least one key token.

15. In any one of paragraphs 11 to 14, The electronic device (300) generates at least one evaluation item, An electronic device characterized in that it generates at least one evaluation item for determining that hallucination has occurred in the response if the response contains content that does not match at least one key context or at least one key token.

Citation Information

Patent Citations

  • Illusion determination method and device and electronic equipment

    CN117764180A

  • Biomarker for colon cancer diagnosis and pharmaceutical composition for preventing or treating colon cancer

    KR102791823B1

  • Computer implemented methods for the automated analysis or use of data, including use of a large language model

    US20240095468A1