Electronic device for performing prompt tuning and control method thereof

The electronic device addresses user satisfaction issues with giant language models by disambiguating user inputs and refining prompts, resulting in improved response accuracy and relevance.

WO2025143597A1PCT designated stage expired Publication Date: 2025-07-03SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019407
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

User satisfaction with giant language models is low due to varying user expressions and incorrect questioning, leading to ambiguous inputs.

Method used

An electronic device identifies user commands from spoken voice, disambiguates words, and tunes prompts for input into large language models, using a processor to refine inputs and update satisfaction scores based on response feedback.

Benefits of technology

Enhances user satisfaction by improving the accuracy and relevance of responses from large language models through prompt tuning, ensuring clearer and more appropriate outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019407_03072025_PF_FP_ABST
    Figure KR2024019407_03072025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The present electronic device may include: a memory; and one or more processors connected to the memory and controlling the electronic device, wherein the processors: identify a user command from a user utterance voice stored in the memory; acquire, from the user utterance voice, a first word representing the location of at least one word which is an object of the user command; acquire a second word describing a preconfigured word when the preconfigured word is identified from the user utterance voice; and acquire a prompt to be input to a large language model (LLM) on the basis of the user utterance voice, the first word, and the second word.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for performing prompt tuning and method for controlling the same

[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device performing prompt tuning and a method for controlling the same.

[0002] Advances in electronic technology have led to the development of electronic devices offering a variety of functions. In particular, recent developments have seen the development of electronic devices utilizing large language models (LLMs), enhancing user convenience.

[0003] A giant language model is a language model composed of an artificial neural network with numerous parameters, and can be a model capable of understanding and generating natural language.

[0004] However, there is a problem that user satisfaction is low because each user of the giant language model may express questions in a different way or ask questions to the giant language model incorrectly.

[0005] According to one embodiment of the present disclosure for achieving the above object, an electronic device includes a memory and one or more processors connected to the memory and controlling the electronic device, wherein the processor identifies a user command from a user's spoken voice stored in the memory, obtains a first word indicating a position of at least one word that is a target of the user command from the user's spoken voice, and when a preset word is identified from the user's spoken voice, obtains a second word describing the preset word, and obtains a prompt for inputting into a large language model (LLM) based on the user's spoken voice, the first word, and the second word.

[0006] Additionally, the memory further stores the giant language model, and the processor can input the prompt into the giant language model to obtain response information corresponding to the prompt.

[0007] And, the processor can identify a word having an ambiguous meaning in the user's spoken voice as the preset word, and obtain the second word that the preset word means in the user's spoken voice among the ambiguous meanings.

[0008] In addition, the display further includes a processor that controls the display to display a UI for selecting one of the ambiguous meanings, and when one of the ambiguous meanings is selected, the second word can be obtained based on the selection.

[0009] In addition, the memory further stores a score indicating the user's satisfaction with each of a plurality of words and a second word indicating the meaning of a word below a preset score, and the processor can identify a word below the preset score in the user's spoken voice as the preset word, and obtain the second word based on the word below the preset score.

[0010] Additionally, the memory can further store the large language model, and the processor can input the prompt into the large language model to obtain response information corresponding to the prompt, receive a user's satisfaction with the response information, and update the score based on the user's satisfaction.

[0011] And, the memory further stores the large language model, and the score can be obtained based on first sample response information obtained by inputting each of a plurality of sample user speech voices into the large language model and second sample response information obtained by inputting each of the plurality of sample prompts corresponding to each of the plurality of sample user speech voices into the large language model.

[0012] Additionally, the processor can obtain the prompt by changing a word corresponding to the user command into a type of command.

[0013] And, the processor can identify a position where the second word is to be added in the user's spoken voice based on the type of the preset word and the first word.

[0014] In addition, the processor further includes a communication interface, and the processor controls the communication interface to transmit the prompt to an external server, and can receive response information corresponding to the prompt from the external server through the communication interface.

[0015] Meanwhile, according to one embodiment of the present disclosure, a control method of an electronic device may include a step of identifying a user command from a user's spoken voice, a step of obtaining a first word indicating a position of at least one word that is a target of the user command from the user's spoken voice, a step of obtaining a second word describing the preset word when a preset word is identified from the user's spoken voice, and a step of obtaining a prompt for inputting into a large language model (LLM) based on the user's spoken voice, the first word, and the second word.

[0016] Additionally, the method may further include a step of inputting the prompt into a large language model to obtain response information corresponding to the prompt.

[0017] And, the step of obtaining the second word may identify a word having an ambiguous meaning in the user's spoken voice as the preset word, and obtain the second word that the preset word means in the user's spoken voice among the ambiguous meanings.

[0018] In addition, the step of obtaining the second word may display a UI for selecting one of the ambiguous meanings, and when one of the ambiguous meanings is selected, the second word may be obtained based on the selection.

[0019] And, the electronic device stores a score indicating the user's satisfaction with each of a plurality of words and a second word indicating the meaning of a word below a preset score, and the step of obtaining the second word may identify a word below the preset score in the user's spoken voice as the preset word, and obtain the second word based on the word below the preset score.

[0020] In addition, the method may further include a step of inputting the prompt into a large language model to obtain response information corresponding to the prompt, a step of receiving a user's satisfaction with the response information, and a step of updating the score based on the user's satisfaction.

[0021] In addition, the score may be obtained based on first sample response information obtained by inputting each of a plurality of sample user speech voices into a large language model, and second sample response information obtained by inputting each of a plurality of sample prompts corresponding to each of the plurality of sample user speech voices into the large language model.

[0022] Additionally, the step of obtaining the above prompt can obtain the prompt by changing the word corresponding to the user command into the type of the command.

[0023] And, the step of obtaining the prompt can identify a position where the second word is to be added in the user's spoken voice based on the type of the preset word and the first word.

[0024] Additionally, the method may further include a step of transmitting the prompt to an external server and a step of receiving response information corresponding to the prompt from the external server.

[0025] FIG. 1 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure.

[0026] FIG. 2 is a block diagram showing a detailed configuration of an electronic device according to an embodiment of the present disclosure.

[0027] FIG. 3 is a diagram for explaining a prompt processing process according to one embodiment of the present disclosure.

[0028] FIG. 4 and FIG. 5 are drawings for explaining a first word according to one embodiment of the present disclosure.

[0029] FIGS. 6 to 8 are drawings for explaining a method for processing words with ambiguous meanings according to one embodiment of the present disclosure.

[0030] FIGS. 9 to 11 are drawings for explaining a method for processing words below a preset score according to one embodiment of the present disclosure.

[0031] FIG. 12 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.

[0032] An object of the present disclosure is to provide an electronic device and a control method thereof for performing prompt tuning before a user's spoken voice is input into a large language model.

[0033] It should be understood that the various embodiments and terms used in this document are not intended to limit the technical features described in this document to specific embodiments, but rather to include various modifications, equivalents, or substitutes of the embodiments.

[0034] In connection with the description of the drawings, similar reference numerals may be used for similar or related components.

[0035] The singular form of a noun corresponding to an item may include one or more items, unless the context clearly indicates otherwise.

[0036] In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in that phrase, or all possible combinations thereof.

[0037] Terms such as "first," "second," or "first" or "second" may be used simply to distinguish one component from another and do not qualify the components in any other respect (e.g., importance or order).

[0038] When a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component), with or without the terms “functionally” or “communicatively,” it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0039] The terms “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in this document, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0040] When a component is said to be “connected,” “coupled,” “supported,” or “in contact with” another component, this includes not only cases where the components are directly connected, coupled, supported, or in contact, but also cases where the components are indirectly connected, coupled, supported, or in contact through a third component.

[0041] When we say that a component is “on” another component, this includes not only cases where the component is in contact with the other component, but also cases where there is another component between the two components.

[0042] The term “and / or” includes any combination of a plurality of related described elements or any one of a plurality of related described elements.

[0043] The operating principle and embodiments of the present invention will be described with reference to the attached drawings below.

[0044] FIG. 1 is a block diagram showing the configuration of an electronic device (100) according to one embodiment of the present disclosure.

[0045] The electronic device (100) is a device that obtains a prompt and can be implemented as a TV, a desktop PC, a laptop, a video wall, a large format display (LFD), a digital signage, a digital information display (DID), a projector display, a smartphone, a tablet PC, etc. For example, the electronic device (100) may be a device that obtains a prompt from a user's speech or text input, etc., and tunes the prompt. Alternatively, the electronic device (100) may be a device that receives a prompt from an external device and tunes the prompt. Here, the prompt is a type of instruction message and may be information input into a large language model (LLM).

[0046] However, it is not limited thereto, and the electronic device (100) may be any device that obtains a prompt.

[0047] According to FIG. 1, the electronic device (100) includes a memory (110) and a processor (120). However, the present invention is not limited thereto, and the electronic device (100) may be implemented in a form in which some components are excluded.

[0048] Memory (110) may refer to hardware that stores information such as data in an electrical or magnetic form so that a processor (120) or the like can access it. To this end, memory (110) may be implemented as at least one piece of hardware from among non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), solid state drive (SSD), RAM, ROM, etc.

[0049] The memory (110) may store at least one instruction required for the operation of the electronic device (100) or the processor (120). Here, the instruction is a code unit that instructs the operation of the electronic device (100) or the processor (120), and may be written in machine language, which is a language that a computer can understand. Alternatively, the memory (110) may store EDID and DPCD for the display (120).

[0050] The memory (110) may store data in bit or byte units that can represent characters, numbers, images, etc. For example, the memory (110) may store user speech, a large language model, a score indicating the user's satisfaction with each of a plurality of words, and information on the meaning of words below a preset score.

[0051] The memory (110) is accessed by the processor (120), and reading / writing / modifying / deleting / updating instructions, instruction sets, or data can be performed by the processor (120).

[0052] The processor (120) controls the overall operation of the electronic device (100). Specifically, the processor (120) is connected to each component of the electronic device (100) and can control the overall operation of the electronic device (100). For example, the processor (120) is connected to components such as a memory (110), a display (not shown), etc. and can control the operation of the electronic device (100).

[0053] The processor (120) may be implemented with one or more processors. In this case, the one or more processors may include one or more of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit), a MIC (Many Integrated Core), a DSP (Digital Signal Processor), an NPU (Neural Processing Unit), a hardware accelerator, or a machine learning accelerator. The one or more processors may control one or any combination of other components of the electronic device (100) and perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in the memory (110). For example, the one or more processors may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in the memory (110).

[0054] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an AI-specific processor). For example, a process of quantizing a neural network model according to an embodiment of the present disclosure may be performed by a general-purpose processor, and a process of learning or inferring the quantized neural network model may be performed by an AI-specific processor.

[0055] One or more processors may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.

[0056] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.

[0057] In embodiments of the present disclosure, one or more processors may refer to a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but the embodiments of the present disclosure are not limited thereto. However, for convenience of explanation, the operation of the electronic device (100) is described below using the expression processor (120).

[0058] The processor (120) can identify a user command from a user's spoken voice stored in the memory (110). For example, the processor (120) can identify the user's intention from the user's spoken voice. Here, the user's spoken voice may be information stored in the memory (110) according to the user's speech. Alternatively, the user's spoken voice may be information received from an external device.

[0059] The processor (120) may obtain a first word indicating the position of at least one word that is the target of a user command in the user's spoken voice. For example, the processor (120) may identify at least one word that is the target of a user command in the user's spoken voice and obtain a first word indicating the position of the word, such as "in the text below."

[0060] When a preset word is identified in the user's spoken voice, the processor (120) can obtain a second word that describes the preset word.

[0061] For example, the processor (120) can identify a word with ambiguous meaning in a user's spoken voice as a preset word, and obtain a second word that the preset word in the user's spoken voice means among the ambiguous meanings. For example, the electronic device (100) further includes a display, and the processor (120) can control the display to display a UI for selecting one of the ambiguous meanings, and when one of the ambiguous meanings is selected, the processor (120) can obtain a second word based on the selection.

[0062] The processor (120) can obtain a prompt for input into the large-scale language model based on the user's speech, a first word, and a second word. For example, the processor (120) can obtain a prompt by adding the first word and the second word to the user's speech. Furthermore, the processor (120) can also obtain a prompt by changing a word corresponding to a user command to a command type. This operation can be referred to as prompt tuning, and even if each user has a different way of expressing themselves or an incorrect way of asking questions, a prompt that can improve the performance of the large-scale language model can be obtained through prompt tuning.

[0063] The memory (110) further stores a large language model, and the processor (120) inputs a prompt into the large language model to obtain response information corresponding to the prompt. However, this is not limited to this, and the large language model may be stored in an external server. In this case, the processor (120) may transmit the prompt to the external server and receive response information obtained by processing the prompt by the large language model from the external server.

[0064] In the above, the preset word is an example of a word having an ambiguous meaning, but it is not limited thereto. For example, the memory (110) further stores a score indicating the user's satisfaction with each of a plurality of words and a second word indicating the meaning of a word below a preset score, and the processor (120) may identify a word below a preset score in the user's spoken voice as a preset word and obtain the second word based on the word below the preset score. Here, the score may be information obtained based on first sample response information obtained by inputting each of a plurality of sample user spoken voices into a large language model and second sample response information obtained by inputting each of a plurality of sample prompts corresponding to each of the plurality of sample user spoken voices into the large language model.

[0065] The processor (120) may input a prompt into a large language model to obtain response information corresponding to the prompt, receive a user's satisfaction rating for the response information, and update a score based on the user's satisfaction rating. Through these operations, a prompt that is adaptive to the user can be obtained.

[0066] The processor (120) can identify a position in the user's spoken speech where the second word is to be added based on the preset word type and the first word. For example, if the preset word type is a word with a preset score in the user's spoken speech and the first word is a word such as "in the text below," the processor (120) can add the second word to the beginning of the user's spoken speech. This operation can prevent the second word from being mixed with at least one word that is the target of the user command.

[0067] Meanwhile, while the processor (120) has been described as acquiring the first and second words, this is not a limitation. For example, the processor (120) may acquire at least one of the first or second words, and obtain a prompt for input into the large language model based on the acquired word and the user's spoken voice.

[0068] Additionally, the processor (120) may first obtain the second word and then obtain the first word.

[0069] The functions related to artificial intelligence according to the present disclosure can be operated through a processor (120) and a memory (110).

[0070] The processor (120) may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, AP, DSP, etc., a graphics-only processor such as a GPU or VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU.

[0071] One or more processors are controlled to process input data according to predefined operating rules or artificial intelligence models stored in the memory (110). Alternatively, if one or more processors are dedicated artificial intelligence processors, the dedicated artificial intelligence processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model. The predefined operating rules or artificial intelligence models are characterized by being created through learning.

[0072] Here, "created through learning" means that a basic artificial intelligence model is learned using a learning algorithm using a plurality of learning data, thereby creating a predefined set of operating rules or an artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0073] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations by calculating the results of previous layers and the multiple weights. The multiple weights of the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated during the learning process to reduce or minimize the loss or cost values ​​obtained by the artificial intelligence model.

[0074] Artificial neural networks may include deep neural networks (DNNs), such as, but not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), or deep Q-networks.

[0075] FIG. 2 is a block diagram illustrating a detailed configuration of an electronic device (100) according to an embodiment of the present disclosure. The electronic device (100) may include a memory (110) and a processor (120). The electronic device (100) may further include a display (130), a communication interface (140), a user interface (150), a microphone (160), a speaker (170), and a camera (180). Among the components illustrated in FIG. 2, a detailed description of parts that overlap with the components illustrated in FIG. 1 will be omitted.

[0076] The display (130) is a component that displays content and can be implemented as a variety of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, a PDP (Plasma Display Panel), etc. The display (130) may also include a driving circuit, a backlight unit, etc. that can be implemented as a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. Meanwhile, the display (130) may be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, etc.

[0077] The communication interface (140) is a configuration that performs communication with various types of external devices according to various types of communication methods. For example, the electronic device (100) can perform communication with an external server through the communication interface (140).

[0078] The communication interface (140) may include a Wi-Fi module, a Bluetooth module, an infrared communication module, a wireless communication module, etc. Here, each communication module may be implemented in the form of at least one hardware chip.

[0079] Wi-Fi and Bluetooth modules communicate via Wi-Fi and Bluetooth, respectively. When using a Wi-Fi or Bluetooth module, connection information, such as the SSID and session key, is first transmitted and received. This information is then used to establish a communication connection before various other information can be transmitted and received. Infrared communication modules use infrared data association (IrDA) technology, which wirelessly transmits data over short distances using infrared light, which lies between visible light and millimeter waves.

[0080] In addition to the above-described communication method, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards such as zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.

[0081] Alternatively, the communication interface (140) may include a wired communication interface such as HDMI, DP, Thunderbolt, USB, RGB, D-SUB, DVI, etc.

[0082] In addition, the communication interface (140) may include at least one of a LAN (Local Area Network) module, an Ethernet module, or a wired communication module that performs communication using a pair cable, a coaxial cable, or an optical fiber cable.

[0083] The user interface (150) may be implemented with buttons, a touch pad, a mouse, a keyboard, etc., or may be implemented with a touch screen capable of performing both display and operation input functions. Here, the buttons may be various types of buttons, such as mechanical buttons, touch pads, wheels, etc., formed on any area of ​​the front, side, or back of the main body of the electronic device (100).

[0084] The microphone (160) is configured to receive sound and convert it into an audio signal. The microphone (160) is electrically connected to the processor (120) and can receive sound under the control of the processor (120).

[0085] For example, the microphone (160) may be formed as an integrated unit integrated into the upper side, front side, side direction, etc. of the electronic device (100). Alternatively, the microphone (160) may be provided in a remote control, etc., separate from the electronic device (100). In this case, the remote control may receive sound through the microphone (160) and provide the received sound to the electronic device (100).

[0086] The microphone (160) may include various configurations such as a microphone that collects analog sound, an amplifier circuit that amplifies the collected sound, an A / D conversion circuit that samples the amplified sound and converts it into a digital signal, and a filter circuit that removes noise components from the converted digital signal.

[0087] Meanwhile, the microphone (160) may be implemented in the form of a sound sensor, and any method may be used as long as it has a configuration capable of collecting sound.

[0088] The processor (120) can receive a user's spoken voice through a microphone (160).

[0089] The speaker (170) is a component that outputs various audio data processed by the processor (120) as well as various notification sounds and voice messages.

[0090] The camera (180) is configured to capture still images or moving images. The camera (180) can capture still images at a specific point in time, but can also capture still images continuously. The camera (180) can capture images in at least one direction of the electronic device (100).

[0091] The camera (180) includes a lens, a shutter, an aperture, a solid-state image sensor, an AFE (Analog Front End), and a TG (Timing Generator). The shutter controls the time at which light reflected from a subject enters the camera (180), and the aperture mechanically increases or decreases the size of the opening through which light enters to control the amount of light incident on the lens. When the solid-state image sensor accumulates light reflected from a subject as a photocharge, the image generated by the photocharge is output as an electrical signal. The TG outputs a timing signal for reading out pixel data of the solid-state image sensor, and the AFE samples and digitizes the electrical signal output from the solid-state image sensor.

[0092] As described above, the electronic device (100) tunes a prompt from the user's spoken voice and provides it to the large language model, so that more improved response information can be obtained than in the case where there is no tuning.

[0093] Hereinafter, the operation of the electronic device (100) will be described in more detail with reference to FIGS. 3 to 11. For convenience of explanation, individual embodiments are described in FIGS. 3 to 11. However, the individual embodiments of FIGS. 3 to 11 may be implemented in any combination or with the order changed. In particular, for convenience of explanation, FIGS. 3 to 11 describe prompts being updated step by step, but the steps may be changed or operated individually.

[0094] FIG. 3 is a diagram for explaining a prompt processing process according to one embodiment of the present disclosure.

[0095] The electronic device (100) may be referred to as a prompt tuner, as illustrated in FIG. 3, and the processor (120) may obtain a prompt from a user's spoken voice.

[0096] An external server (200) stores a large language model, and can obtain response information by inputting a prompt provided from an electronic device (100) into the large language model.

[0097] Since the prompt is refined by the electronic device (100), clearer response information can be generated than when it is not refined. For example, a user's speech containing ambiguous expressions may be interpreted by the large language model as meaning not intended by the user. However, a prompt with ambiguous expressions supplemented can be interpreted by the large language model as meaning intended by the user, thereby outputting more appropriate response information.

[0098] For convenience of explanation, in FIG. 3, the electronic device (100) and the external server (200) are described as separate entities, but this is not limiting. For example, the electronic device (100) stores a large language model, and after obtaining a prompt from a user's spoken voice, it may input the prompt into the large language model to obtain response information.

[0099] FIG. 4 and FIG. 5 are drawings for explaining a first word according to one embodiment of the present disclosure.

[0100] The processor (120) can obtain a first word indicating the position of at least one word that is the target of a user command in the user's spoken voice.

[0101] For example, the processor (120) can identify whether a target text exists through a classifier from a user speech, and can identify whether the target is explicitly expressed through a sequence labeler. For example, as illustrated in FIG. 4, the processor (120) can identify "I rode a boat yesterday... (410)" from a user speech such as "A modified summary of only the content about the boat. I rode a boat yesterday..." as at least one word that is the target of the user command "summary."

[0102] The processor (120) can acquire the first word based on the relative positions of the user command and the target of the user command in the user's spoken voice. For example, as illustrated in FIG. 5, the processor (120) can acquire "In the text below (510)" as the first word, since "I rode a boat yesterday..." is located after the expression "summary."

[0103] The processor (120) can obtain a prompt for input into the large language model based on the user's spoken speech and a first word. Here, the processor (120) can add the first word to the user's spoken speech based on the location of the target of the user command. For example, the processor (120) can add "in the text below (510)" before the expression "I rode a boat yesterday..."

[0104] The processor (120) can modify the user's spoken voice into a natural expression through a style transfer. For example, as illustrated in FIG. 5, the processor (120) can update the prompt by adding "to (52-0)" and "to (530)" to "summarize in a modified form", such as "summarize in a modified form."

[0105] Models such as the classifiers in FIGS. 4 and 5 may be implemented as a rule base or as a neural network model.

[0106] FIGS. 6 to 8 are drawings for explaining a method for processing words with ambiguous meanings according to one embodiment of the present disclosure.

[0107] When a preset word is identified in the user's spoken voice, the processor (120) can obtain a second word that describes the preset word.

[0108] For example, the processor (120) can identify a word with ambiguous meaning in a user's spoken voice as a preset word, and obtain a second word that the preset word in the user's spoken voice means among the ambiguous meanings. For example, as illustrated in FIG. 6, the processor (120) can identify lexical morphemes such as nouns, verbs, adjectives, and adverbs from the user's spoken voice through a POS (part of speech) tagger, and identify "pear" as a word with ambiguous meaning.

[0109] The processor (120) may display a UI for selecting whether the meaning of “pear” is “human abdomen (710)” or “fruit (720)” as illustrated in FIG. 7, and if “fruit (720)” is selected, the processor may update the prompt by adding “fruit (810)” as a second word in front of “pear” as illustrated in FIG. 8.

[0110] A model such as the POS tagger in FIGS. 6 to 8 may be implemented as a rule base or as a neural network model.

[0111] FIGS. 9 to 11 are drawings for explaining a method for processing words below a preset score according to one embodiment of the present disclosure.

[0112] The processor (120) may identify words in the user's spoken voice below a preset score as preset words, and may also obtain a second word based on the words below the preset score. For example, the memory (110) may further store a score indicating the user's satisfaction with each of a plurality of words and a second word indicating the meaning of the words below the preset score, and the processor (120) may identify words in the user's spoken voice below a preset score as preset words based on the information stored in the memory (110), and may also obtain a second word indicating the meaning of the words below the preset score. For example, as illustrated in FIG. 9, the processor (120) may identify "modified (910)" as a word below a preset score.

[0113] Here, the score may be obtained based on first sample response information obtained by inputting each of a plurality of sample user speech voices into a large language model, and second sample response information obtained by inputting each of a plurality of sample prompts corresponding to each of the plurality of sample user speech voices into the large language model. For example, as illustrated in FIG. 10, the score may be obtained based on users' evaluations of sample response information obtained by inputting each of a plurality of sample prompts corresponding to each of the plurality of sample user speech voices into the large language model. If the user is dissatisfied, the score of each lexical morpheme included in the sample prompt may be lowered by a preset value. If this operation is repeated several times, the problematic lexical morpheme may have a significantly lower score than other lexical morphemes, and information indicating the meaning of words below the preset score may be stored in the memory (110).

[0114] As shown in FIG. 11, the processor (120) can add "{modified} means {a method of listing important points or words in short bursts when writing} (1110)" to the user's spoken voice as a second word, indicating the meaning of "modified (910)", a word with a score lower than a preset value, through a word dictionary or a generation model.

[0115] Here, the processor (120) can identify a position at which a second word is to be added in the user's spoken speech based on the type of the preset word and the first word. That is, the processor (120) can identify a position at which a second word representing the meaning of a word below a preset score is to be added based on the position of at least one word that is the target of the user command and the words below a preset score. In the example described above, based on the first word, such as "in the text below," the second word can be added before the position of at least one word that is the target of the user command.

[0116] The processor (120) can input a prompt into a large language model to obtain response information corresponding to the prompt, receive the user's satisfaction with the response information, and update a score based on the user's satisfaction. For example, if the user is dissatisfied, the processor (120) can lower the score of each lexical morpheme included in the prompt by a preset value. If this operation is repeated and the score of a specific word falls below the preset score, the processor (120) can additionally store the meaning of the specific word in the memory (110).

[0117] Models such as the word dictionary or generation model in FIGS. 9 to 11 may be implemented as a rule base or as a neural network model.

[0118] FIG. 12 is a flowchart for explaining a control method of an electronic device according to an embodiment of the present disclosure.

[0119] First, a user command is identified from the user's spoken voice (S1210). Then, a first word indicating the position of at least one word that is the target of the user command is obtained from the user's spoken voice (S1220). Then, when a preset word is identified from the user's spoken voice, a second word describing the preset word is obtained (S1230). Then, a prompt for inputting into a large language model (LLM) is obtained based on the user's spoken voice, the first word, and the second word (S1240).

[0120] Additionally, the method may further include a step of inputting a prompt into a large language model to obtain response information corresponding to the prompt.

[0121] And, the step of obtaining a second word (S1230) can identify a word with an ambiguous meaning in the user's spoken voice as a preset word, and obtain a second word that is meant by the preset word in the user's spoken voice among the ambiguous meanings.

[0122] In addition, the step of obtaining the second word (S1230) displays a UI for selecting one of the ambiguous meanings, and when one of the ambiguous meanings is selected, the second word can be obtained based on the selection.

[0123] And, the electronic device stores a score representing the user's satisfaction with each of a plurality of words and a second word representing the meaning of a word below a preset score, and the step of obtaining the second word (S1230) can identify a word below a preset score in the user's spoken voice as a preset word, and obtain the second word based on the word below the preset score.

[0124] Additionally, the method may further include a step of inputting a prompt into a large language model to obtain response information corresponding to the prompt, a step of receiving a user's satisfaction with the response information, and a step of updating a score based on the user's satisfaction.

[0125] In addition, the score can be obtained based on first sample response information obtained by inputting each of a plurality of sample user speech voices into a large language model and second sample response information obtained by inputting each of a plurality of sample prompts corresponding to each of a plurality of sample user speech voices into a large language model.

[0126] Additionally, the step of obtaining a prompt (S1240) can obtain a prompt by changing a word corresponding to a user command into a command type.

[0127] And, the step of obtaining a prompt (S1240) can identify the position where the second word is to be added in the user's spoken voice based on the type of the preset word and the first word.

[0128] Additionally, the method may further include a step of transmitting a prompt to an external server and a step of receiving response information corresponding to the prompt from the external server.

[0129] According to various embodiments of the present disclosure as described above, the electronic device tunes a prompt from a user's spoken voice and provides it to a large language model, thereby obtaining more improved response information than in the case where there is no tuning.

[0130] Meanwhile, according to a temporary example of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include an electronic device (e.g., electronic device (A)) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.

[0131] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0132] Furthermore, according to one embodiment of the present disclosure, the various embodiments described above may be implemented in a computer-readable recording medium or a similar device using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented by the processor itself. In a software implementation, embodiments such as the procedures and functions described herein may be implemented as separate software. Each software may perform one or more functions and operations described herein.

[0133] Meanwhile, computer instructions for performing processing operations of a device according to the various embodiments described above may be stored in a non-transitory computer-readable medium. The computer instructions stored in such a non-transitory computer-readable medium, when executed by a processor of a specific device, cause the specific device to perform processing operations in the device according to the various embodiments described above. A non-transitory computer-readable medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media may include a CD, DVD, hard disk, Blu-ray disk, USB, memory card, or ROM.

[0134] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0135] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, memory; and comprising one or more processors connected to said memory and controlling said electronic device; The above processor, Identifying user commands from user speech voices stored in the above memory, Obtaining a first word indicating the position of at least one word that is the target of the user command from the user's spoken voice; When a preset word is identified in the user's speech, a second word describing the preset word is obtained, An electronic device that obtains a prompt for input into a large language model (LLM) based on the user's spoken voice, the first word and the second word.

2. In paragraph 1, The above memory is, Further save the above huge language model, The above processor, An electronic device that inputs the above prompt into the large language model to obtain response information corresponding to the prompt.

3. In paragraph 1, The above processor, Identify words with ambiguous meanings in the user's speech as the preset words, An electronic device that obtains the second word meaning the preset word in the user's speech voice among the above ambiguous meanings.

4. In paragraph 3, including display; The above processor, Control the display to display a UI for selecting one of the above ambiguous meanings, An electronic device, wherein when one of the above ambiguous meanings is selected, the second word is obtained based on the selection.

5. In paragraph 1, The above memory is, It further stores a score representing the user's satisfaction with each of the multiple words and a second word representing the meaning of words below a preset score. The above processor, Identifying words below the preset score in the user's speech as the preset words, An electronic device that obtains the second word based on words below the preset score.

6. In paragraph 5, The above memory is, Further save the above huge language model, The above processor, By inputting the above prompt into the above giant language model, response information corresponding to the above prompt is obtained, Receive user satisfaction with the above response information, An electronic device that updates the score based on the user's satisfaction.

7. In paragraph 5, The above memory is, Further save the above huge language model, The above score is, An electronic device, wherein the first sample response information is obtained by inputting each of a plurality of sample user speech voices into the large language model, and the second sample response information is obtained by inputting each of the plurality of sample prompts corresponding to each of the plurality of sample user speech voices into the large language model.

8. In paragraph 1, The above processor, An electronic device that obtains the prompt by changing a word corresponding to the user command into a type of command.

9. In paragraph 1, The above processor, An electronic device that identifies a position where the second word is to be added in the user's spoken voice based on the type of the preset word and the first word.

10. In paragraph 1, further comprising a communication interface; The above processor, Control the communication interface to transmit the above prompt to an external server, An electronic device that receives response information corresponding to the prompt from the external server through the communication interface.

11. In a method for controlling an electronic device, A step of identifying a user command from user speech; A step of obtaining a first word indicating the position of at least one word that is the target of the user command in the user's spoken voice; When a preset word is identified in the user's speech voice, a step of obtaining a second word describing the preset word; and A control method, comprising: a step of obtaining a prompt for inputting into a large language model (LLM) based on the user's speech voice, the first word, and the second word.

12. In paragraph 11, A control method further comprising the step of inputting the above prompt into a large language model to obtain response information corresponding to the prompt.

13. In paragraph 11, The step of obtaining the second word is: Identify words with ambiguous meanings in the user's speech as the preset words, A control method for obtaining the second word meaning the preset word in the user's speech voice among the above ambiguous meanings.

14. In paragraph 13, The step of obtaining the second word is: Display a UI for selecting one of the above ambiguous meanings, A control method for obtaining the second word based on the selection when one of the above ambiguous meanings is selected.

15. In paragraph 11, The above electronic device, Stores a score representing the user's satisfaction with each of the multiple words and a second word representing the meaning of words below a preset score. The step of obtaining the second word is: Identifying words below the preset score in the user's speech as the preset words, A control method for obtaining the second word based on words below the preset score.

Citation Information

Patent Citations

  • Standards planning intention recognition method for pre-training large language model tuning

    CN117076661A

  • Infolding Type Hinge Device

    KR1020230006319A

  • Question-answering system based on reconfiguration of dialogue

    KR102280792B1

  • Parameter Efficient Prompt Tuning for Efficient Models at Scale

    US20230325725A1

  • Systems and methods for shared latent space prompt tuning

    US20230419027A1