Method, device, and recording medium for generating response to user input on basis of context

By generating responses based on contextual information and user preferences, the method and device improve user intent recognition, reducing the need for multiple inputs and enhancing satisfaction.

WO2025193060A1PCT designated stage Publication Date: 2025-09-18VTOUCH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099716
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2025-03-12
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing interactive devices often fail to accurately identify user intent, requiring multiple inputs or providing inappropriate responses, leading to user inconvenience and dissatisfaction.

Method used

A method and device that generate responses based on contextual information and user preferences, incorporating a user input processing unit, context information management unit, user preference management unit, and response generation unit to optimize responses.

Benefits of technology

Provides optimized responses that match user intentions by considering contextual information and preferences, enhancing user satisfaction and trust in the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099716_18092025_PF_FP_ABST
    Figure KR2025099716_18092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, and recording medium for generating a response to a user input on the basis of context. The method for generating a response to a user input on the basis of context, according to an embodiment of the present disclosure, comprises the steps of: receiving a user input; searching for and analyzing context information; extracting a user preference by referring to the user input and the context information; and generating a response to the user input on the basis of the context information and the user preference.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device and recording medium for generating a response to user input based on context

[0001] The present disclosure relates to a method, device and recording medium for generating a response to user input based on context.

[0002] Typically, electronic devices generate and output appropriate responses based on various forms of user input, such as peripheral device manipulation, touch, and voice. These devices typically generate and output preset responses to user input. However, even for identical user input, the required response may vary depending on the context or environment in which the input occurs.

[0003] Recently, electronic devices are being used that not only perform a set function but also interact with the user, such as through conversation. These interactive devices incorporate artificial intelligence technology to provide appropriate responses to user input. However, even existing interactive devices often fail to accurately identify the intent of user input, frequently requiring additional user input multiple times or providing inappropriate responses. This causes user inconvenience and lowers satisfaction and trust in the device.

[0004] The present disclosure is intended to solve the problems of the above-described prior art.

[0005] The present disclosure aims to provide a technique for generating an optimized response to user input.

[0006] Additionally, the present disclosure aims to provide a response that matches the user's intention by generating a response to user input based on user preferences along with contextual information.

[0007] A method for generating a response to a user input based on a context according to one embodiment of the present disclosure includes the steps of receiving a user input, searching and analyzing contextual information, extracting a user preference by referring to the user input and the contextual information, and generating a response to the user input based on the contextual information and the user preference.

[0008] According to one embodiment of the present disclosure, context information may include at least one of linkage information including at least one of the time and location of a user input, the user's schedule, screen information regarding items displayed on the screen of the user terminal, connection control information with the user terminal, information regarding the user's surrounding environment, recent conversation information, recent execution information, and companion information.

[0009] According to one embodiment of the present disclosure, in the step of searching and analyzing contextual information, contextual information may be searched with reference to user input at the time when the user input is received.

[0010] According to one embodiment of the present disclosure, user preferences can be preset.

[0011] According to one embodiment of the present disclosure, user preferences can be generated with reference to user input and contextual information.

[0012] According to one embodiment of the present disclosure, the user input may include at least one of a voice input, a button input, a touch input, a gesture input, and a visual input via a user device.

[0013] A method for generating a response to a user input based on a context according to one embodiment of the present disclosure may further include the steps of receiving additional user input for the response and correcting the response based on the additional user input.

[0014] A method for generating a response to a user input based on a context according to one embodiment of the present disclosure may further include, after the step of correcting the response, a step of correcting the user preference or generating a new user preference by referring to the corrected response.

[0015] A device for generating a response to a user input based on a context according to one embodiment of the present disclosure includes a user input processing unit for receiving a user input, a context information management unit for searching and analyzing context information, a user preference management unit for extracting a user preference by referring to the user input and the context information, and a response generation unit for generating a response to the user input based on the context information and the user preference.

[0016] In addition, other methods for implementing the present disclosure, other devices, and recording media for recording a computer program for executing the method are further provided.

[0017] According to one embodiment of the present disclosure, an optimized response to user input can be generated based on context.

[0018] Additionally, according to one embodiment of the present disclosure, a response that matches the user's intention can be provided by generating a response to the user input based on the user's preference along with contextual information.

[0019] FIG. 1 is a schematic diagram illustrating an overall system environment for generating a response to user input based on context according to one embodiment of the present disclosure.

[0020] FIG. 2 is a functional block diagram schematically illustrating the functional configuration of a response generation device according to one embodiment of the present disclosure.

[0021] FIG. 3 is a drawing exemplarily showing a user input made through a user input device according to one embodiment of the present disclosure.

[0022] FIG. 4 is a drawing showing a user input device according to another embodiment of the present disclosure.

[0023] FIG. 5 is a drawing showing a user input device according to another embodiment of the present disclosure.

[0024] FIG. 6 is a diagram exemplarily showing a structure for generating a response to user input according to one embodiment of the present disclosure.

[0025] FIG. 7 is a flowchart illustrating an exemplary method for generating a response to user input according to one embodiment of the present disclosure.

[0026] Figures 8a through 8e illustrate various examples of generating responses by referencing contextual information and user preferences.

[0027] Figures 9a to 9d illustrate various examples of generating or correcting user preferences.

[0028] [Explanation of symbols]

[0029] 100: Full response generation system

[0030] 110: User input device

[0031] 120: Communications network

[0032] 130: User terminal

[0033] 140: Response generating device

[0034] 150: External device

[0035] 201: User Input Processing Unit

[0036] 203: Contextual Information Management Department

[0037] 205: User Preference Management

[0038] 207: Response generation unit

[0039] 209: Communications Department

[0040] 211: Database

[0041] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. Hereinafter, specific descriptions of previously known functions and configurations will be omitted if deemed likely to unnecessarily obscure the gist of the present disclosure. Furthermore, it should be noted that the following description relates only to one embodiment of the present disclosure and that the present disclosure is not limited thereto.

[0042] The terminology used in this disclosure is only used to describe specific embodiments and is not intended to limit the present disclosure. For example, a component expressed in the singular should be understood to include plural components unless the context clearly indicates only the singular. Terms such as “comprise,” “include,” and “have” used in this disclosure are intended to specify the presence of a feature, number, step, motion, component, or combination thereof described in this disclosure, and the use of such terms does not exclude the presence or addition of one or more other features, numbers, steps, motions, components, or combinations thereof.

[0043] In the embodiments of the present disclosure, a "module" or "part" refers to a functional part that performs at least one function or motion, and may be implemented by hardware or software, or a combination of hardware and software. Furthermore, a plurality of "modules" or "parts" may be integrated into at least one software module and implemented by at least one processor, excluding "modules" or "parts" that need to be implemented by specific hardware.

[0044] Additionally, unless otherwise defined, all terms used in this disclosure, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which this disclosure pertains. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with the contextual meaning of the relevant technology, and should not be interpreted in an unduly limiting or expansive manner unless explicitly defined otherwise in this disclosure.

[0045] FIG. 1 is a schematic diagram illustrating an overall system environment for generating a response to user input based on context according to one embodiment of the present disclosure.

[0046] Referring to FIG. 1, the entire system environment (100) for generating a response to a user input may include a user input device (110), a communication network (120), a user terminal (130), a response generating device (140), and an external device (150).

[0047] A user input device (110) according to one embodiment of the present disclosure may be a portable device that can be carried by a user or a wearable device that can be worn by a user, and may be capable of receiving voice and / or visual input. In one embodiment, the user input device (110) may be a ring-shaped smart ring. Although such a user input device (110) is illustrated as being provided separately from the user terminal (130), the user input device (110) does not necessarily need to be a separate device from the user terminal (130), and the user input device (110) may be included in the user terminal (130), or the user terminal (130) may be configured to function as the user input device (110).

[0048] According to one embodiment of the present disclosure, a user input device (110) can communicate with a user terminal (130) via a communication network (120). In the illustrated embodiment, the user input device (110) can receive user inputs such as voice inputs and button inputs, and as described below, it is also possible to receive user inputs such as touch inputs, visual inputs, and gesture inputs via other types of user input devices. A signal regarding a user input received by the user input device (110) can be transmitted to the user terminal (130) via the communication network (120).

[0049] A communication network (120) according to one embodiment of the present disclosure may include any wired or wireless communication network, for example, a TCP / IP communication network. According to one embodiment of the present disclosure, the communication network (120) may be configured as a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), an Internet network, etc., but the present disclosure is not limited thereto. For example, the communication network (120) may be a wireless data communication network that implements, at least in part, a conventional communication method such as WiFi communication, WiFi-Direct communication, Long Term Evolution (LTE) communication, 5G communication, Bluetooth communication (including Bluetooth Low Energy (BLE) communication), infrared communication, ultrasonic communication, etc. As another example, the communication network (120) may be an optical communication network that implements a conventional communication method such as LiFi (Light Fidelity) at least in part.

[0050] A user terminal (130) according to one embodiment of the present disclosure can receive user input and perform processing for response output. The user terminal (130) is a digital device equipped with memory means and a microprocessor to provide computing capabilities, and may be, for example, a smartphone. However, the user terminal (130) of the present disclosure is not limited to the illustrated device, and may be a smart pad, a notebook computer, a personal digital assistant (PDA), a desktop computer, or any other device capable of achieving the purpose of the present disclosure may be used as the user terminal (130) of the present disclosure.

[0051] According to one embodiment of the present disclosure, the user terminal (130) can receive a signal regarding a user input from the user input device (110) via the communication network (120). For example, when a user input is detected by the user input device (110), the corresponding input signal can be transmitted to the user terminal (130) via the communication network (120). In another embodiment, the user terminal (130) can receive a signal regarding a user input via a microphone, camera, or other sensor included within the user terminal.

[0052] According to one embodiment of the present disclosure, a user terminal (130) can communicate with a response generation device (140) via a communication network (120). In one embodiment, the user terminal (130) can transmit a received user input signal to the response generation device (140) via the communication network (120).

[0053] A response generation device (140) according to one embodiment of the present disclosure may perform a function of generating a response to a user input. To this end, the response generation device (140) may communicate with a user terminal (130) via a communication network (120). In one embodiment, the response generation device (140) may be a server system. The response generation device (140) may be an independent physical server, or may be a virtual server such as a cloud server.

[0054] According to one embodiment of the present disclosure, the response generation device (140) is an artificial intelligence server that receives a user input signal from a user terminal (130), recognizes the signal, and then generates a response signal for the user input.

[0055] Although the illustrated embodiment depicts that operations and processing for response generation are performed in a response generation device (140) separate from the user terminal (130), the user terminal (130) may be configured to perform operations such as receiving user input and generating a response, and thus the operations may be processed in the user terminal (130) instead of being performed in an external device. In other words, the user terminal (130) may function as a response generation device (140), directly recognizing a user input signal and performing response generation without communicating with an external device.

[0056] According to one embodiment of the present disclosure, the user terminal (130) may be configured to receive a response signal generated by the response generation device (140) and output a response, or transmit the response to an external device (150). For example, the output response may include transmission of a message via the user terminal (130) or the external device (150), creation and progress of a conversation, control of the user terminal (130), control of the external device (150), etc.

[0057] In one embodiment, the external device (150) to which the response signal is transmitted from the user terminal (130) may be an earphone, a headset, or a speaker. For example, the user terminal (130) may transmit the response signal to an earphone worn by the user, and output a voice message or conduct a conversation through the earphone. In another embodiment, the external device (150) to which the response signal is transmitted may be a home appliance such as a smart TV, an air conditioner, a refrigerator, or a smart car. For example, the user terminal (130) may generate a control signal based on a response signal generated in response to a user input and transmit the control signal to the smart TV, and the smart TV may be controlled (power on / off, channel or volume change, etc.) based on the control signal.

[0058] The external device (150) to which the response signal and / or control signal is transmitted is not limited to the examples above, and may include any electronic device capable of providing an auditory or visual response to the user. It should be understood that any electronic device equipped with wireless communication capabilities and capable of being controlled according to the user's intention is included within the external device (150).

[0059] Control of an application or web installed or running on a user terminal (130) or another user device may be achieved through a response signal generated by a response generation device (140). Hereinafter, an external device that can be connected to and controlled by a user terminal (130) is referred to as a connection control device, and control of a connection control device, an application, a web, etc. is collectively referred to as connection control.

[0060] Meanwhile, although the illustrated embodiment describes that the user input device (110) communicates with the user terminal (130), and the user terminal (130) communicates with the response generation device (140), it may be implemented differently. For example, the user input device (110) may communicate directly with the response generation device (140). In this case, the user input device (110) may transmit a user input signal to the response generation device (140) via the communication network (120), and the response generation device (140) may perform a calculation from the user input signal and then generate and transmit a response signal and / or a control signal. As another example, the user terminal (130) may function as an on-device AI, and may perform a calculation directly from the user input signal without communicating with the response generation device (140) to generate a response signal and / or a control signal. As another example, the user input device (110) may communicate directly with another IoT device, such as an AI speaker, rather than a user terminal (130) or a response generation device (140).

[0061] In addition, the user input device (110) is configured to perform operations such as voice recognition, generation of response signals and / or control signals, etc., so that the operations can be processed in the user input device (110) instead of being performed in an external device such as a user terminal (130) or a response generation device (140).

[0062] FIG. 2 is a functional block diagram schematically illustrating the functional configuration of a response generation device according to one embodiment of the present disclosure.

[0063] Referring to FIG. 2, a response generation device (140) according to one embodiment of the present disclosure may include a user input processing unit (201), a context information management unit (203), a user preference management unit (205), a response generation unit (207), a communication unit (209), and a database (211). The components illustrated in FIG. 2 do not reflect all of the components of the response generation device (140), nor are they essential, and thus, the response generation device (140) may include more or fewer components than the illustrated components.

[0064] According to one embodiment of the present disclosure, at least some of the user input processing unit (201), the context information management unit (203), the user preference management unit (205), the response generation unit (207), and the database (211) may be program modules that communicate with an external system (not shown) via a communication unit (209). These program modules may be included in the response generation device (140) in the form of an operating system, an application program module, or other program modules, and may be physically stored in various known memory devices. In addition, these program modules may be stored in a remote memory device that can communicate with the response generation device (140). Meanwhile, these program modules include, but are not limited to, routines, subroutines, programs, objects, components, data structures, etc. that perform specific tasks or execute specific abstract data types, which will be described later according to the present disclosure.

[0065] A user input processing unit (201) of a response generating device (140) according to one embodiment of the present disclosure may perform a function of receiving a user input signal. In one embodiment, the user input processing unit (201) may receive a user input from a user input device (110) or a user terminal (130) via a communication network (120). Here, the user input may include at least one of a voice input, a button input, a touch input, a gesture input, and a visual input via the user input device (110) or the user terminal (130).

[0066] FIG. 3 is a diagram exemplarily illustrating a user input performed through a user input device according to one embodiment of the present disclosure. Referring to FIG. 3 , the user input device (110) may be configured as a voice input device for voice recognition. In the illustrated embodiment, a smart ring equipped with a microphone is used as the voice input device, and proximity voice generated by the user's close-range speech can be acquired as user input. For example, a motion of moving the voice input device close to the mouth (or lips) can be used as a trigger, and the subsequently input proximity voice can be acquired as user input and provided to the user input processing unit (201) via the user terminal (130).

[0067] Meanwhile, the smart ring is equipped with a button (not shown), allowing user input via the button. For example, button input can be processed as user input for calling interactive AI, user input for sending and receiving emergency data, or user input for controlling other user terminals.

[0068] FIG. 4 is a drawing showing a user input device according to another embodiment of the present disclosure. Referring to FIG. 4, the user input device may be formed in various forms, such as a smart ring type, a touch device (touch panel, touch pad, etc.), a remote control (400a) having a button or microphone, a stylus pen (400b), a smart watch (400c), a smartphone (400d), etc. As illustrated, the user input devices (400a, 400b, 400c, 400d) are configured to have a microphone (401a, 401b, 401c, 401d) and a touch device or button to enable voice input (proximity voice input) and / or touch or button input as user input.

[0069] FIG. 5 is a diagram illustrating a user input device according to another embodiment of the present disclosure. Referring to FIG. 5, the user input device (110) may be configured as a visual input device. In the illustrated embodiment, the visual input device may be configured to acquire visual information in the forward or visual direction of the user. For example, the visual input device may be configured as a camera (501) that can be worn on clothing or a camera (502) that can be worn on the user's head (e.g., earpiece, glasses, hairband). Through this, visual information in the forward or visual direction of the user can be acquired and utilized as user input.

[0070] For example, in the case of a voice input such as “Tell me about the historic site I am looking at” or “Can you tell me the rating of the wine in the upper right corner?”, the user input processing unit (201) can obtain and utilize visual information obtained together with the voice input (i.e., visual information about the historic site, visual information about the wine displayed on the shelf) as user input. As another example, a situation recognized through a visual input device (e.g., movement of a location, change of environment, etc.) can be utilized as user input. In addition, user input such as the user’s gaze and gestures (e.g., a gesture pointing to a specific location with a finger, a gesture moving a finger in one direction, etc.) can also be obtained through the visual input device.

[0071] Meanwhile, the visual input device can also be configured to activate and acquire visual information when triggered. For example, if a voice input device or user terminal (130) linked to the visual input device recognizes the start of a conversation or a voice containing a specific directive (e.g., "here," "there," "this," "that"), the visual input device can be activated to acquire visual information. Accordingly, the visual input device can be operated efficiently by using it only when necessary.

[0072] In addition, the user input device (110) according to the present disclosure may be implemented in another form by including at least one of a microphone, a camera, and other sensors capable of obtaining user input.

[0073] According to one embodiment of the present disclosure, the user input processing unit (201) may perform a function of operationally processing an input signal. For example, if the input signal is received in the form of a voice signal, the user input processing unit (201) may extract characters related to language and emotional information from the voice signal through a voice recognition algorithm. As another example, if the input signal is received in the form of visual information (e.g., a video signal), the user input processing unit (201) may obtain objects, scenes, moods, location information, etc. included in the visual information through an image recognition algorithm and extract characters related to the corresponding information.

[0074] Meanwhile, as described above, the user input device (110) can receive user input when there is a predetermined trigger. Together with this, or separately from this, the user input processing unit (201) can determine whether the acquired signal is noise by determining whether the input is intended by the user. For example, when a predetermined trigger signal for user input (e.g., trigger voice, trigger gesture, etc.) is input first, when a voice larger than a preset volume is input, when a voice of a specific speaker is input, etc., the subsequent user input can be determined as an intended user input, and conversely, when the user input does not meet a predetermined condition, the user input can be determined as noise even if it is received.

[0075] The context information management unit (203) of the response generation device (140) according to one embodiment of the present disclosure may perform a function of searching and analyzing context information. In one embodiment, the context information management unit (203) may search and analyze context information at the time of user input. In addition, the context information management unit (203) may search for context information by referring to the user input. Here, context information refers to information that is not directly included in the user input but is related thereto, and corresponds to information that can be referenced to determine the user's intent from the user input and generate a response.

[0076] For example, if a user input requests that an air conditioner be turned on, the context information management unit (203) may search for and utilize environmental information, including the time and temperature at the time of the user input, as context information to determine an appropriate air conditioner temperature setting. Furthermore, the context information management unit (203) may also search for and utilize recent conversation information as context information. For example, if the user's recent conversation indicates poor health, this may be referenced when determining the air conditioner temperature setting.

[0077] According to one embodiment of the present disclosure, context information may include at least one of linkage information including at least one of the time and location of user input, the user's schedule, screen information regarding items displayed on the screen of the user terminal (e.g., buttons, input windows, etc.), connection control information regarding a connection control device, application, web, etc. connected to the user terminal, information regarding the user's surrounding environment, recent conversation information (e.g., conversation partner, conversation content, etc.), execution information regarding a recent connection control device, application, web, etc. executed, and companion information.

[0078] In one embodiment, contextual information may be obtained from a user terminal (130) or a user input device (110). For example, if the user terminal (130) or the user input device (110) is a smart phone or wearable smart device including functions such as a watch, GPS, or schedule recorder, the contextual information management unit (203) of the response generation device (140) may be configured to obtain such information from the user terminal (130) via the communication network (120). Similarly, connection control information, recent conversation information, recent execution information, and the like may be obtained from the user terminal (130).

[0079] In one embodiment, the screen information among the context information can be obtained from a captured image of the screen of the user terminal or information of a running application. The context information management unit (203) can obtain a captured image of the screen of the user terminal (130) from the user terminal (130) through the communication network (120), and can obtain objects, scenes, moods, location information, etc. included therein and extract characters related to the information. Alternatively, the context information management unit (203) can obtain screen information of the user terminal (130) by receiving information on an application running in the front view from the user terminal (130) (e.g., identifier information of the application, identifier information for an object within the application, etc.).

[0080] In one embodiment, contextual information, such as visual information and companion information, can be obtained from a visual input device. For example, if the contextual information management unit (203) obtains visual information including a book through a camera that can be worn on the user's head, the contextual information management unit (203) can determine contextual information indicating that the user is currently reading. In another example, the contextual information management unit (203) can obtain companion information from visual information about people located around the user obtained through the visual input device. In this way, visual information obtained from a visual input device can be used not only for user input but also to obtain contextual information such as the user's surroundings or companions.

[0081] In one embodiment, contextual information may be obtained from a device connected to the user terminal (130). The user terminal (130) may be connected to a connection control device and / or another user device via a communication network (120), and data related to contextual information obtained from the connection control device and / or the other user device may be provided to the contextual information management unit (203). For example, data regarding the current set temperature and the current indoor temperature obtained from an air conditioner, which is a connection control device, may be provided to the contextual information management unit (203). In addition, when the user terminal (130) is connected to another user device, information regarding the owner of the other user device may be obtained and provided to the contextual information management unit (203). This may be utilized as contextual information, such as information about the user's companion.

[0082] Although the contextual information searched and analyzed by the contextual information management unit (203) has been described as an example above, contextual information is not limited to the examples described above. It should be understood that any information that can be obtained from a user terminal (130), a user input device (110), etc., and can be referenced for responding to user input can be utilized as contextual information without limitation.

[0083] The user preference management unit (205) of the response generation device (140) according to one embodiment of the present disclosure may perform a function of extracting and managing user preferences by referring to user input and context information. User preferences refer to information on a user's preferred tendency for each item related to user input or context. For example, a user's preference for music may include genre information such as "ballad," "dance," "hip-hop," and "pop," and a user's preference for indoor temperature may include numerical information such as "22°C." According to one embodiment of the present disclosure, user preferences may include information that the user dislikes in addition to information that the user prefers.

[0084] In one embodiment, user preferences may be preset. For example, user preferences may be directly entered through a user terminal (130) and recorded and stored in a database (211).

[0085] In one embodiment, user preferences can be generated by referencing user input and contextual information. For example, if a user input such as "It's cold in the room right now, so raise the temperature," is received, the user preference manager (205) can obtain the current room temperature from a connected control device, such as an air conditioner, and generate a temperature 1 to 2 degrees Celsius higher than the current room temperature as the user preference.

[0086] According to one embodiment of the present disclosure, user preferences can be set differently for the same item depending on contextual information. For example, different user preferences can be set depending on the user's location, such as home, work, park, or restaurant, and different user preferences can be set depending on companions, such as family, friends, or coworkers. For example, a user's music preference could be set to "Ballad" if there is no companion, "Dance" if the companion is a friend, and "Children's Song" if the companion is a family member.

[0087] According to one embodiment of the present disclosure, user preferences can be dynamically changed by the user preference management unit (205). In one embodiment, the user preference management unit (205) can correct the user preference by referring to user input, i.e., feedback, in response to a response from the response generation unit (207) described below. For example, when the indoor temperature of the air conditioner is set by referring to the user preference for the user input, and additional user input (feedback) such as “It’s cold in the room now, so turn up the temperature” is received, the user preference management unit (205) can set (correct) a temperature that is 1 to 2 degrees Celsius higher than the indoor temperature set as the existing user preference as the new user preference.

[0088] In one embodiment, user input for generating or correcting user preferences can be obtained in the form of direct input or indirect extraction. For example, user input containing negative feedback, such as "No," "Not that," or "Something else," or positive feedback, such as "Changing the lighting tone definitely makes it more comfortable, thank you," can be utilized as a basis for generating or correcting user preferences as direct input. Furthermore, information regarding positive, negative, and other emotions indirectly extracted from conversational content can also be referenced as a basis for generating or correcting user preferences.

[0089] The response generation unit (207) of the response generation device (140) according to one embodiment of the present disclosure can perform a function of generating a response to a user input based on contextual information and user preferences.

[0090] According to one embodiment of the present disclosure, the response generation unit (207) may include a Large Language Model (LLM) or a Large Multimodal Model (LMM) to generate a response to a user input. The large language model of the response generation unit (207) may be configured to generate and output characters based on a character input, so that when a character extracted from a user input is input in the user input processing unit (201), the corresponding character may be output with reference to contextual information. In addition, the large multimodal model of the response generation unit (207) may be configured to generate and output characters, images, videos, voices, sounds, etc. based on an input of characters, images, videos, voices, sounds, etc., so that the user input processing unit (201) may be configured to output corresponding characters with reference to contextual information corresponding to the user input.

[0091] According to one embodiment of the present disclosure, a response to a user input may be output via a user terminal (130) or a connected control device. The output of the response to the user input may take the form of message generation, conversation transmission or creation, control execution, control suggestion, etc.

[0092] In one embodiment, the transmission of the conversation may be an answer to a user's inquiry, provision of information, a request for additional input, etc., and the creation of the conversation may be in the form of first starting the conversation based on linked information including the user's schedule, information about the surrounding environment, etc. In addition, the execution of the control means generating a control signal for the user terminal (130), a connected control device, an application, the web, etc. For example, a control signal may be generated for controlling the set temperature of an air conditioner or playing music through a speaker according to a user input, a control signal may be generated for editing data such as a schedule stored in the user terminal (130), or a control signal may be generated for controlling an application installed in the user terminal (130) (e.g., an application for ordering food delivery, reserving transportation, etc.).

[0093] The communication unit (209) of the response generation device (140) according to one embodiment of the present disclosure may function to enable the response generation device (140) to communicate with an external communication network. For example, the communication unit (209) may perform a function of receiving a signal according to a user input or transmitting a response signal and a control signal to a user terminal (130).

[0094] The database (211) of the response generation device (140) according to one embodiment of the present disclosure performs a function of storing data that can be used to generate a response. In one embodiment, the database (211) can store contextual information and information related to user preferences for each user.

[0095] Meanwhile, according to one embodiment of the present disclosure, personal information included in contextual information and user preference information may be stored separately in the user terminal (130) rather than in the database (211) of the response generation device (140). Accordingly, personal information can be used and modified only when the user terminal (130) is present, thereby preventing problems related to unauthorized use of personal information.

[0096] FIG. 6 is a diagram exemplarily showing a structure for generating a response to a user input according to one embodiment of the present disclosure, and FIG. 7 is an operational flowchart exemplarily showing a method for generating a response to a user input according to one embodiment of the present disclosure. Each step of the response generation method in the present embodiment does not necessarily have to be performed in the order illustrated, and each step is not necessarily essential. That is, it should be understood that each step of the response generation method according to the present embodiment may be performed in a different order than illustrated, and some steps may be omitted or other steps may be added.

[0097] First, in step (S701), a user input is received. Here, the user input may include at least one of a voice input, a button input, a touch input, a gesture input, and a visual input via a user terminal or a user input device, and may be received via a response generation device (140).

[0098] According to one embodiment of the present disclosure, a user input signal received through a response generation device (140) can be processed in an operational manner. For example, if the input signal is received in the form of a voice signal, the response generation device (140) can extract characters from the voice signal through a voice recognition algorithm, and if the input signal is received in the form of a video signal, the response generation device (140) can obtain information on objects, scenes, moods, locations, etc. included in the video signal through an image recognition algorithm and extract characters from the information. Alternatively, user inputs such as characters, images, videos, voices, and sounds can be used as they are without separate operational processing.

[0099] According to one embodiment of the present disclosure, it is possible to determine whether a signal regarding user input received in step (S701) is noise. For example, if a predetermined trigger signal for user input is first input, if a voice louder than a preset volume is input, or if the voice of a specific speaker is input, the signal may be determined to be intended user input. If the signal does not meet these conditions, the signal may be determined to be noise.

[0100] In step (S703), contextual information is searched and analyzed. In one embodiment, the search and analysis of contextual information may be performed through the response generation device (140). Here, the contextual information may include at least one of linkage information including at least one of the time and location of user input, the user's schedule, screen information regarding items displayed on the screen of the user terminal, connection control information of the user terminal, information regarding the user's surrounding environment, recent conversation information, execution information such as recent connection control devices, and companion information. Meanwhile, the response generation device (140) may search and acquire contextual information from the user terminal (130), the user input device (110), or the connection control device connected to the user terminal (130). In addition, the response generation device (140) may search and analyze contextual information with reference to the user input.

[0101] In step (S705), user preferences are extracted by referring to user input and contextual information. In one embodiment, the response generation device (140) can extract user preferences by referring to user input and contextual information, and record and store the extracted user preferences. The user preferences may be preset or may be generated by referring to user input and contextual information.

[0102] In step (S707), a response to user input is generated based on contextual information and user preferences. In one embodiment, a response to user input may be generated using a large language model (LLM) or a large multimodal model (LMM) of the response generation device (140).

[0103] According to one embodiment of the present disclosure, a response generation device (140) can generate a response to a user input based on contextual information and user preferences, which can be output through a user terminal (130) or a connection control device. The output of the response to the user input can be in the form of message generation, conversation transmission or generation, control execution, control proposal, etc. That is, in response to the user input, voice or text output, screen control, application control, control proposal, etc. can be performed through the user terminal (130) and / or the connection control device.

[0104] Meanwhile, according to one embodiment of the present disclosure, after generating a response, the response can be corrected by referring to additional input or contextual information of the user, and processing can be performed to generate or correct user preferences.

[0105] In step (S709), it is possible to check whether additional user input for the response has been received. In one embodiment, the response generation device (140) terminates the process if there is no additional input from the user, and if additional input is received from the user, it may additionally perform steps for correcting the response or generating or correcting user preferences.

[0106] Specifically, when additional user input is received in step (S711), the response is corrected accordingly. In one embodiment, the response generation device (140) can correct the response based on contextual information and user preferences through the same process as steps (S703) and (S707) described above, but by reflecting additional user input (e.g., negative feedback, etc.).

[0107] Next, in step (S713), if the response is corrected based on additional user input, the user preference can be corrected or a new user preference can be generated. That is, if a response was generated with reference to a preset user preference, but the user determines that the response requires correction, the user preference can be corrected based on the additional user input. If no user preference exists based on the user input and context, a new user preference can be generated.

[0108] Hereinafter, a process of generating a response to a user input, correcting the response, or generating or correcting a user preference according to one embodiment of the present disclosure is described with specific examples.

[0109] Figures 8a through 8e illustrate various examples of generating responses to user input with reference to contextual information and user preferences.

[0110] Fig. 8a illustrates a case where a user input of "party mode" is received. In this case, at the time when the user input is received, contextual information about the current location (living room), contextual information about the schedule (dinner party), contextual information about the environment (party), contextual information about the connected control devices (lights, speakers, home entertainment system), contextual information about recent conversations (invite friends), contextual information about companion information (friends), etc. are acquired, and dance music is extracted as the user's preference by referring to the user input and contextual information. Then, referring to this contextual information and user preference, a control signal for activating various lighting effects in the lighting connected control device, a control signal for playing dance music in the speaker connected control device, and a control signal for preparing operation in the home entertainment system connected control device can be generated and output in response to the user input of "party mode."

[0111] Figure 8b illustrates a case where the user input "Where would be good to go?" is received. In this case, contextual information such as location (study), schedule (dinner with Sam), environment (reading), connected control devices (speaker and Uber), recent conversation partner (Sam) and conversation content (vegan), and companion information (Jamie) are acquired, and the user's preference of providing a 30-minute advance notice of the schedule is extracted. Next, by referencing this contextual information and user preferences, the user's intention to receive recommendations for a place to have dinner with Sam can be identified, and a response signal containing the message "How about a Veggie Weekend for Sam? Confirm and I'll make a reservation and call an Uber." can be generated and output in response to the user input. Afterwards, if there is additional positive user input (feedback), a control signal can be generated to launch a restaurant reservation application and an Uber call application.

[0112] Figure 8c illustrates a case where the user input "Wow, it's hard" is received. In this case, contextual information at the time the user input is received, such as location (near home), schedule (exercise), environment (jogging), connected control devices (lighting, music), recent conversations (5km course, evening), and companion information (none), is acquired, and the user preference (post-exercise bath, lounge music) is extracted with reference to this contextual information (i.e., the user has finished jogging and arrived near home). Then, by referring to the contextual information and user preference, a control signal for controlling the lighting and music can be generated along with a message such as "You will feel more comfortable if you soak in hot water. I have dimmed the lights and played lounge music" in response to the user input.

[0113] FIG. 8d illustrates a case where a user input such as "How should I spend my evening?" is received. In this case, when the user input is received, the contextual information such as location (home, living room), environment (TV viewing), connected control device (TV, DoorDash), recent conversation (XX Chicken, Wonka), and companion information (12-year-old son, June) are acquired, and the user preference related to this (enjoying ordered food while watching TV) is extracted by referring to the contextual information (i.e., watching TV with my son in the living room and talking about chicken and Wonka). Next, by referring to the contextual information and user preference, the message "Shall I order XX Chicken? How about 'Wonka' with June?" can be generated and output as a response to the user input, and if there is positive feedback, the food delivery application can be executed and a control signal for the TV can be generated at the same time.

[0114] Figure 8e illustrates a case where the user input "I'm getting ready for work" is received. In this case, contextual information such as the schedule (9:00 AM meeting), connected control devices (Uber, coffee machine, thermostat), and recent conversation information (car under repair) are acquired at the time the user input is received. Based on this contextual information, the user's preference for using a car during the meeting is extracted. Subsequently, despite the user's preference, the contextual information indicating that the car is under repair can be referenced to generate a response signal containing the message "I turned on the coffee machine. The car is under repair, so I called Uber. It will arrive in 20 minutes."

[0115] Thus, according to the present disclosure, rather than simply performing a pre-entered response to user input, a response tailored to the user's intent can be generated by referencing various contextual information and user preferences. Furthermore, connection control information at the time of user input can be verified to provide appropriate control of the connection control device. Even when a response based on user preferences is not possible, contextual information can be used to generate an optimal response.

[0116] Figures 9a through 9d illustrate various examples of correcting user preferences or generating new user preferences. Specifically, Figures 9a through 9d illustrate various examples of generating or correcting responses by referencing contextual information and user preferences regarding an item when user input is present, and simultaneously reflecting the generated or corrected response in the user preferences (i.e., generating or correcting user preferences).

[0117] Referring to Fig. 9a, if there is user input regarding setting the indoor temperature, a response for setting the indoor temperature can be generated by referring to contextual information (recent conversations, time, weather / temperature, humidity, companions, etc.) and user preference (indoor temperature 22℃), and the response can be corrected if there is additional feedback from the user. For example, if the user preference is 22℃, a response for setting the indoor temperature higher than 22℃ can be generated by referring to recent conversation information containing 'cold' or companion information containing 'young child', and the indoor temperature can be ultimately corrected to 25℃ or 24℃ based on additional user input, i.e., user feedback (not shown). In this way, if a response is generated differently from a previously set user preference by referring to contextual information or a response is corrected based on additional user input (feedback), the user preference can be newly generated or corrected by referring to the corresponding response. For example, if the user has a cold, the user's indoor temperature preference can be set to 25°C, and if the companion includes a young child, the user's indoor temperature preference can be set to 24°C. Similarly, new user preferences can be created based on specific environmental factors, such as time of day, weather, and humidity, or existing user preferences can be modified (edited).

[0118] Referring to Figure 9b, if there is user input regarding bedtime, a response regarding bedtime (generating a conversation or controlling a connection) can be generated by referencing contextual information (schedule, recent conversations, companions, etc.) and the user's preference (weekday bedtime of 11 PM), and the response can be revised based on additional user feedback. For example, if the user's preference is 11 PM, but there is a meeting the next morning or a young child is with them, a response setting the bedtime to an earlier time can be generated or revised. In addition, if there is a project due the next day, a response setting the bedtime to a later time than the user's preference can be generated or revised. Furthermore, a response generated or revised that differs from the user's preference can be referenced to create or correct a new user preference.

[0119] Referring to Fig. 9c, if there is user input regarding music playback, a response can be generated to play music or provide a list of music by referencing contextual information (schedule, companion, recent conversations, etc.) and user preferences (ballad, pop), and the response can be revised based on additional feedback from the user. For example, if the user's music preference is ballad and pop, but hip-hop and dance music are played by referencing a schedule that includes exercise as contextual information, the generated response can be used to generate a user preference for hip-hop and dance music for the music in the exercise situation. In addition, if the companion information includes a young child as contextual information and nursery rhymes are played, the generated response can be used to generate a user preference for nursery rhymes for the music in the situation where a young child is present. Similarly, if a response that is different from the user's preference is generated by referring to contextual information such as companion information (wife, friends), recent conversations (party, stress), schedule (reading), etc., or if the response is corrected by additional user input (feedback), a new user preference can be generated based on the contextual information or an existing user preference can be corrected (changed) by referring to such responses.

[0120] Referring to FIG. 9d, if there is user input regarding video playback, a response may be generated to play a video or provide a list of videos by referring to contextual information (time, location, companion, schedule, etc.) and user preferences (e.g., new technology, science fields). The response may be modified based on additional user feedback. For example, if the user's preference is new technology or science fields, but the response is modified to play videos about travel or scenery based on additional user input (feedback), the user's preference regarding the video may be modified to travel or scenery. Alternatively, if the user plays the video in the bedroom in the morning, the user's preference may be generated as travel or scenery. Similarly, if the response is modified to play videos about family or adventure based on additional user input (feedback), the user's preference regarding the video may be modified to family or adventure. Alternatively, if the user is accompanied by family, the user's preference may be generated as family or adventure.

[0121] In this way, when a response to user input is generated or corrected, it is possible to generate a more optimized response to subsequent user input by referencing this to newly generate or correct user preferences, thereby providing a smart user experience.

[0122] Meanwhile, the embodiments according to the present disclosure described above may be implemented in the form of program commands that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be specially designed and configured for the present disclosure or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. Hardware devices may be changed into one or more software modules to perform processing according to the present disclosure, and vice versa.

[0123] In this way, the idea of ​​the present disclosure should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the claims described below as well as the claims are considered to fall within the scope of the idea of ​​the present disclosure.

Claims

1. A method for generating a response to user input based on context, Step of receiving user input, Steps to explore and analyze contextual information, A step of extracting user preferences by referring to the above user input and the above context information, and A step of generating a response to the user input based on the above context information and the user preference. How to include.

2. In paragraph 1, A method in which the contextual information includes at least one of linkage information including at least one of the time and location of the user input, the user's schedule, screen information regarding matters displayed on the screen of the user terminal, connection control information with the user terminal, information regarding the user's surrounding environment, recent conversation information, recent execution information, and companion information.

3. In paragraph 1, A method in which, in the step of exploring and analyzing contextual information, contextual information is explored with reference to the user input at the time the user input is received.

4. In paragraph 1, The above user preferences are preset, method.

5. In paragraph 1, A method wherein the above user preference is generated by referring to the above user input and the above contextual information.

6. In paragraph 1, A method wherein the user input includes at least one of voice input, button input, touch input, gesture input, and visual input via a user input device or user terminal.

7. In paragraph 1, A step of receiving additional user input for the above response, and Step of correcting the response according to the above user additional input How to include more.

8. In paragraph 7, A method further comprising, after the step of correcting the response, a step of correcting the user preference or generating a new user preference by referring to the corrected response.

9. A recording medium recording a computer program for executing the method according to paragraph 1.

10. A device that generates a response to user input based on context, A user input processing unit that receives user input, Contextual information management department that explores and analyzes contextual information; A user preference management unit that extracts user preferences by referring to the above user input and the above context information, and A response generation unit that generates a response to the user input based on the above context information and the user preference. A device comprising:

Citation Information

Patent Citations

  • Motorcycle battery controller with protective electrical connections

    KR1020240009183A

  • A film for display substrate having improved oxygen gas barrier property and moisture barrier property and manufacturing method thereof

    KR1020240063223A

  • Pressing type nebulizer

    KR102744113B1

  • Methods and systems for improved document processing and information retrieval

    WO2024015321A1

  • KR20190133100A