Electronic device and control method therefor
The electronic device leverages a large language model to interpret user voice commands and generate control commands for connected devices, addressing the complexity of conventional home automation by automating device control without manual settings or additional sensors.
Patent Information
- Application Number
- PCT/KR2024/019775
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-03
AI Technical Summary
Conventional home automation systems require users to manually set operation conditions and device-specific commands for each device, which is cumbersome and necessitates additional sensor installations, especially when controlling multiple devices.
An electronic device equipped with a communication interface, microphone, memory, and processor uses a large language model to interpret user voice commands, retrieves necessary information, and generates control commands for connected devices without manual setting, utilizing a home IoT server to manage device operations.
Enables seamless home automation by interpreting user voice commands, eliminating the need for manual condition and command settings and sensor installations, and simplifying the control of multiple devices through integrated information retrieval and command generation.
Smart Images

Figure KR2024019775_03072025_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device and a method for controlling the same for controlling an external device connected to the electronic device using a large language model.
[0002] Smart home IoT services provide home automation features that can automatically activate devices connected to a smart home IoT server. Existing home automation features automatically activate user-defined device activation conditions when they are met. Therefore, to activate existing home automation, users must first define and configure specific activation conditions and device activation details for each device. In this way, smart home IoT services can operate home automation by activating devices according to the defined device activation details when the conditions are met, based on the user-defined activation conditions and device activation details.
[0003] As described above, conventional home automation technology requires users to manually set operating conditions and configure device operation details based on those conditions, which can be cumbersome. For example, to configure an air conditioner to turn on during hot weather, the user must access a smart home IoT application and input specific values for the temperature at which the air conditioner should turn on and off. Furthermore, building home automation requires the installation of additional sensors to measure external conditions, posing a challenge. Furthermore, when multiple devices need to be activated, the user must specify operating conditions and device operation details for each device, which can be cumbersome.
[0004] According to one embodiment of the present disclosure, an electronic device includes: a communication interface; a microphone; a memory storing at least one instruction; and a processor connected to the communication interface, the microphone, and the memory for controlling the electronic device, wherein the processor executes the at least one instruction so that, when a user voice for controlling an external device is input through the microphone, information on at least one external device related to the user voice is obtained, and when it is determined that additional information necessary for performing an operation related to the user voice exists by inputting information on the user voice into a large language model, additional information related to the user voice is obtained, and a prompt for controlling the at least one external device based on the user voice and the additional information is obtained, and the prompt is input into the large language model to obtain a control command for controlling the at least one external device, and the control command is transmitted to the at least one external device through the communication interface.
[0005] The processor, when the user's voice is input, can obtain list information including a plurality of external devices connected to the home IoT server through the home IoT server connected to the electronic device, and obtain current status information of the plurality of external devices included in the list information.
[0006] The processor may obtain a prompt for controlling the at least one external device based on current status information of the at least one external device among the plurality of external devices, the user voice, and the additional information.
[0007] The processor may obtain a prompt template for generating the prompt, and may obtain the prompt by inserting information about the user voice and the additional information into the prompt template.
[0008] The processor may input information about the user's voice into a large language model, and if it is determined that there is additional information necessary to perform an action related to the user's voice, obtain additional information through at least one of the large language model, the Internet, and a user database.
[0009] The above additional information may include at least one of common sense information obtained through the large language model, external environment information related to the user voice searched through the Internet, and user history information related to the user voice searched through a user database.
[0010] The above processor can map and call a control command obtained through the large language model to an API (Application Programming Interface) related to a home IoT service.
[0011] The above processor can output a message to guide the user with information related to the above control command.
[0012] According to one embodiment of the present disclosure, a method for controlling an electronic device includes: when a user voice for controlling an external device is input, obtaining information on at least one external device related to the user voice; when it is determined that additional information necessary for performing an operation related to the user voice exists by inputting the information on the user voice into a large language model, obtaining additional information related to the user voice; obtaining a prompt for controlling the at least one external device based on the user voice and the additional information; obtaining a control command for controlling the at least one external device by inputting the prompt into the large language model; and transmitting the control command to the at least one external device.
[0013] The step of obtaining information about the at least one external device may include obtaining list information including a plurality of external devices connected to the home IoT server through the home IoT server connected to the electronic device when the user's voice is input, and obtaining current status information of the plurality of external devices included in the list information.
[0014] The step of obtaining the above prompt may obtain a prompt for controlling the at least one external device based on current status information of the at least one external device among the plurality of external devices, the user voice, and the additional information.
[0015] The step of obtaining the above prompt may include obtaining a prompt template for generating the prompt, and inserting information about the user's voice and the additional information into the prompt template to obtain the prompt.
[0016] The step of obtaining the above additional information may include inputting information about the user voice into a large language model, and if it is determined that there is additional information necessary to perform an action related to the user voice, obtaining the additional information through at least one of the large language model, the Internet, and a user database.
[0017] The above additional information may include at least one of common sense information obtained through the large language model, external environment information related to the user voice searched through the Internet, and user history information related to the user voice searched through a user database.
[0018] It may include a step of mapping the control command obtained through the above-mentioned large language model to an API (Application Programming Interface) related to the home IoT service and calling it.
[0019] It may include a step of outputting a message to guide the user with information related to the above control command.
[0020] FIG. 1 is a diagram illustrating a home IoT system according to one embodiment of the present disclosure;
[0021] FIG. 2 is a block diagram showing the configuration of an electronic device according to one embodiment of the present disclosure;
[0022] FIG. 3 is a block diagram showing a configuration for performing a home automation function included in an electronic device according to one embodiment of the present disclosure;
[0023] FIG. 4 is a diagram illustrating a method for performing a weather-related home automation function according to one embodiment of the present disclosure, and
[0024] FIG. 5 is a flowchart for explaining a method for controlling an electronic device according to one embodiment of the present disclosure.
[0025] The present embodiments may be modified and have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0026] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0027] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0028] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0029] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0030] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0031] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0032] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0033] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0034] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0035] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0036] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.
[0037] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0038] According to the present disclosure, a "neural network model" is a model implemented based on the neural network of the human brain, and may refer to an overall model in which artificial neurons that form a network by combining synapses change the strength of the synapses through learning to have problem-solving capabilities. At this time, the neural network model includes multiple layers (or strata), and for example, may include an input layer, an output layer, and a number of hidden layers between them. At this time, the neural network model may include an artificial neural network (ANN) model, a deep neural network (DNN) model, etc.
[0039] According to the present disclosure, a "home IoT (Internet of Things) service" may refer to an intelligent service that connects devices located within a home based on information and communication technology, enabling information exchange and communication between people and devices, and between devices. In this case, the home IoT service may provide a home automation function that automatically configures multiple devices located within the home without requiring the user to manually input settings.
[0040] According to the present disclosure, a "Large Language Model (LLM) (hereinafter referred to as "LLM") is a language model comprised of an artificial neural network with numerous parameters. The LLM can be trained on a significant amount of unlabeled corpus text using self-supervised learning or semi-self-supervised learning. In this case, the LLM can not only generate answers to user inquiries, but also include inference capabilities and the ability to autonomously plan and execute such plans. Meanwhile, the LLM may be referred to by various terms, such as "large language model" or "AI chatbot model."
[0041] According to the present disclosure, a "prompt" may refer to an input for initiating an interaction with a large language model. The prompt may be a text input or a voice input including one or more texts and / or one or more sentences. In one embodiment, the prompt may include natural language text. The natural language text may include various information that the large language model can utilize to generate a response to a user inquiry or to control a home IoT system, such as context, intent, task, and constraints. Meanwhile, the prompt may be referred to by various expressions representing the same / similar concept. For example, the prompt may be replaced with expressions such as "input," "user input," "input phrase," "user command," "directive," "starting sentence," "task query," "trigger sentence," and "message," but is not limited to the examples described above. Meanwhile, the user voice input into the initial LLM (200) may also be a prompt, but for the convenience of explanation, the user voice input into the initial LLM (200) and the prompt generated by additional information are distinguished and used.
[0042]
[0043] Hereinafter, the present disclosure will be described in more detail with reference to the drawings.
[0044] FIG. 1 is a diagram illustrating a home IoT system according to one embodiment of the present disclosure. As illustrated in FIG. 1, the home IoT system (1) may include an electronic device (100), an LLM (200), a home IoT server (10), and a plurality of external devices (20-1, 20-2, 20-3, ...).
[0045] The electronic device (100) may be a smart phone as shown in FIG. 1, but this is only one example, and may include, but is not limited to, a personal computer, a terminal, a portable telephone, a handheld device, a wearable device, etc.
[0046] The home IoT server (10) may be a server for managing a plurality of external devices (20-1, 20-2, 20-3,...) within the home IoT system (1). The home IoT server (10) may include a communication module capable of communicating with another server, a plurality of external devices (20-1, 20-2, 20-3,...) or an electronic device (100), at least one processor capable of processing data received from another server, a plurality of external devices (20-1, 20-2, 20-3,...) or an electronic device (100), and at least one memory capable of storing a program for processing data or processed data. The home IoT server (10) may be implemented as various computing devices such as a workstation, a cloud, a data drive, a data station, etc. The home IoT server (10) can be implemented as one or more servers that are physically or logically separated based on function, detailed configuration of function, or data, etc., and can transmit and receive data and process the transmitted and received data through communication between each server.
[0047] In particular, the home IoT server (10) (or may be referred to as a server, a home server, etc.) may perform functions such as managing user accounts, registering a plurality of external devices (20-1, 20-2, 20-3,...) by linking them to user accounts, and managing or controlling the plurality of registered external devices (20-1, 20-2, 20-3,...). For example, a user may access the home IoT server (10) through an electronic device (100) and create a user account. The user account may be identified by an ID and password set by the user. The home IoT server (10) may register a plurality of external devices (20-1, 20-2, 20-3,...) to the user account according to a set procedure. For example, the home IoT server (10) can register, manage, and control multiple external devices (20-1, 20-2, 20-3,...) by linking their identification information (e.g., serial number or MAC address) to a user account.
[0048] The plurality of external devices (20-1, 20-2, 20-3, ...) may be at least one of various types of home appliances located within a home. For example, the plurality of external devices (20-1, 20-2, 20-3, ...) may include an air conditioner (20-1), a refrigerator (20-2), a washing machine (20-3), etc. as illustrated in FIG. 1, but this is merely an example, and may include at least one of a dishwasher, an electric range, an electric oven, an air conditioner, a clothes manager, a dryer, a microwave oven, but is not limited thereto, and may include various types of home appliances such as a cleaning robot, a vacuum cleaner, a television, etc., which are not illustrated in the drawing. In addition, the aforementioned home appliances are merely examples, and in addition to the aforementioned home appliances, the plurality of external devices may include various products (e.g., manual blinds, automatic windows, automatic doors, etc.) that may be located within a home.
[0049] The LLM (200) may be a language model trained to generate a response to a user's voice based on text obtained through the user's voice or to generate a control command for controlling at least one external device located in the home.
[0050] In particular, the LLM (200) according to one embodiment of the present disclosure can perform a function of identifying whether an additional command is required based on a user's voice, and can perform a function of generating a control command for controlling at least one external device based on the generated prompt.
[0051] Specifically, the electronic device (100) can receive a user voice input for controlling at least one external device among a plurality of external devices (20-1, 20-2, 20-3, ...). At this time, the user voice may include a specific command for controlling at least one external device, such as "make the temperature of the house 20 degrees," but this is only an example, and may include an abstract command, such as "make the home environment comfortable."
[0052] The electronic device (100) can obtain text for the user's voice and input (or transmit) the obtained text to the LLM (200). The LLM (200) can identify whether additional information is required to control an external device based on the obtained text.
[0053] If additional information is identified as being required, the LLM (200) may request the electronic device (100) for the additional information. The electronic device (100) may obtain the additional information through at least one of a large language model, the Internet, and a user database. Furthermore, the electronic device (100) may generate a prompt based on the obtained additional information. The electronic device (100) may generate the prompt using a prompt template.
[0054] The electronic device (100) can input (or transmit) the generated prompt to the LLM (200). The LLM (200) can obtain a control command for controlling at least one external device based on the prompt.
[0055] The electronic device (100) can transmit a control command acquired by the LLM (200) to a home IoT server (10) or at least one external device to perform a service corresponding to the user's voice.
[0056] By the above-described embodiment, the user does not need to set the operating conditions and operation contents one by one to perform the home automation function, and by obtaining additional information through various sources, not only does it eliminate the need for additional sensors, but it also reduces the hassle of specifying the operating conditions and operation contents for each of the various external devices.
[0057] FIG. 2 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure. As illustrated in FIG. 2, the electronic device (100) may include a communication interface (110), a microphone (120), a display (130), memory (140), and a processor (150). Meanwhile, the configuration illustrated in FIG. 2 is merely an example, and other configurations may be added or some configurations may be deleted depending on the type of the electronic device (100).
[0058] The communication interface (110) includes at least one circuit and can communicate with various types of external devices. The communication interface (110) can include at least one of a BLE (Bluetooth Low Energy) module, a Wi-Fi communication module, a cellular communication module, a 3G (third generation) mobile communication module, a 4G (fourth generation) mobile communication module, a 4th generation LTE (Long Term Evolution) communication module, and a 5G (fifth generation) mobile communication module.
[0059] In particular, the communication interface (110) can communicate with an external device (20) included in the home IoT system (1) and a home IoT server (10). Specifically, the communication interface (110) can receive list information including a plurality of external devices (20-1, 20-2, 20-3, ...) connected to the home IoT server (10) from the home IoT server (10). In addition, the communication interface (110) can receive current status information of the plurality of external devices included in the list information from the home IoT server (10). At this time, the communication interface (110) can directly receive information about the external device (20) and information about the current status of the external device (20) from the external device (20).
[0060] The communication interface (110) can communicate with an external LLM (200). For example, the communication interface (110) can transmit a user voice (or text information corresponding to the user voice) to the LLM (200) and receive a signal for requesting additional information from the LLM (200). The communication interface (110) can transmit a generated prompt to the LLM (200) and receive at least one of a control command and a response corresponding to the prompt from the LLM (200).
[0061] The microphone (120) is a configuration for receiving audio generated from an object (e.g., human voice, audio generated from an object, etc.) and converting it into audio data. The microphone (102) can receive audio when activated. The microphone (130) can include various configurations such as a microphone for collecting audio in analog form, an amplifier circuit for amplifying the collected audio, an A / D conversion circuit for sampling the amplified audio and converting it into a digital signal, and a filter circuit for removing noise components from the converted digital signal. In particular, the microphone (120) can receive a user's voice for controlling an external device (20).
[0062] The display (130) is a configuration that displays image data, and may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, the display (130) may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc. The display (130) may also include a driving circuit, a backlight unit, etc., which may be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc. In addition, the display (130) may be implemented as a touch screen with a touch sensor. In particular, the display (130) may output a message for guiding a user with information related to a control command acquired based on a prompt.
[0063] The memory (140) may store an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and instructions or data related to the components of the electronic device (100). In particular, the memory (140) may include a user voice acquisition module (310), a text information acquisition module (310), a device search module (325), an additional information acquisition module (330), a prompt generation module (340), and a control command transmission module (350), as illustrated in FIG. 3, to perform a home automation function using the LLM (200). In particular, when the home automation function is executed using the LLM (200), the electronic device (100) may load data for various modules for performing the home automation function using the LLM (200), which is stored in the non-volatile memory, into the volatile memory. Here, loading means an operation of loading and storing data stored in non-volatile memory into volatile memory so that the processor (150) can access it.
[0064] The memory (140) may store a user database containing user history information. The user history information may include at least one of information regarding the external device's configuration patterns, information regarding the external device's preferred settings, and information regarding the external device's recent settings. While the user database may be stored in the memory (140), this is merely an example, and may be stored externally (e.g., in a home IoT server (10)).
[0065] Meanwhile, LLM (200) may be stored in an external server, but this is only one embodiment, and memory (140) may store LLM (200).
[0066] Meanwhile, the memory (110) may be implemented as a non-volatile memory (e.g., hard disk, SSD (Solid state drive), flash memory), volatile memory (may also include memory within the processor (150)), etc.
[0067] The processor (150) can control the electronic device (100) according to at least one instruction stored in the memory (140).
[0068] In particular, the processor (150) may include a plurality of processors. Specifically, the plurality of processors may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors may control one or any combination of other components of the electronic device and perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in a memory. For example, the plurality of processors may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in a memory.
[0069] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor or by a plurality of processors. That is, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-dedicated processor). For example, an operation of obtaining additional information or an operation of generating a prompt may be performed by the first processor (e.g., a CPU), and an operation of identifying whether additional information is needed through the LLM (200) or an operation of generating a control command may be performed by the second processor (e.g., a GPU or NPU).
[0070] The processor (150) may be implemented as one or more multi-core processors including multiple cores (e.g., homogeneous multi-cores or heterogeneous multi-cores). When the processor (150) is implemented as a multi-core processor, each of the multiple cores included in the multi-core processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multi-core processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multi-core processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0071] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among a plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0072] In embodiments of the present disclosure, a processor may mean a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto.
[0073] Meanwhile, when a user voice for controlling an external device is input through a microphone (120), the processor (150) obtains information on at least one external device related to the user voice. At this time, the information on the external device may include at least one of identification information, location information, performance information, and current status information of the external devices (20-1, 20-2, 20-3, etc.) located in the home. When it is determined that additional information necessary for performing an operation related to the user voice exists by inputting information on the user voice to the LLM (200), the processor (150) obtains additional information related to the user voice. At this time, the additional information may be information necessary for the external device (20) to generate a prompt for performing an operation related to the user voice. The processor (150) obtains a prompt for controlling at least one external device based on the user voice and the additional information. The processor (150) inputs the prompt to the LLM (200) to obtain a control command for controlling the at least one external device. The processor (150) transmits a control command to at least one external device via a communication interface (110).
[0074] In one embodiment, when a user voice is input, the processor (150) may obtain list information including a plurality of external devices connected to the home IoT server (10) through the home IoT server (10) connected to the electronic device (100). The processor (150) may obtain current status information of the plurality of external devices included in the list information. At this time, the current status information of the external device may include at least one of power status information, setting status information, and location status information of the external device.
[0075] In one embodiment, the processor (150) may obtain a prompt to control at least one external device based on current state information, user voice, and additional information about at least one external device among a plurality of external devices.
[0076] In one embodiment, the processor (150) may obtain a prompt template for generating a prompt. The prompt template, which is stored in the memory (140) or the home IoT device (10) for generating the prompt, may include general commands that, when combined with additional information, may execute various user commands. For example, the prompt template may include the general command, "Set (device operation) to execute a user command in (additional information)." The processor (150) may insert information about the user's voice and additional information into the prompt template to obtain the prompt.
[0077] In one embodiment, when it is determined that additional information necessary for performing an operation related to the user voice is present by inputting information about the user voice into the LLM (200), the processor (150) may acquire additional information through at least one of a large language model, the Internet, and a user database. In this case, the additional information may include at least one of common sense information acquired through the LLM (200), external environment information related to the user voice searched through the Internet, and user history information related to the user voice searched through the user database.
[0078] In one embodiment, the processor (150) can map a control command obtained through the LLM (200) to an API (Application Programming Interface) related to a home IoT service and call it.
[0079] In one embodiment, the processor (150) may output a message to guide the user with information related to the control command.
[0080] FIG. 3 is a block diagram illustrating a configuration for performing a home automation function included in an electronic device according to an embodiment of the present disclosure. As illustrated in FIG. 3 , the electronic device (100) may include a user voice acquisition module (310), a text information acquisition module (310), a device search module (325), an additional information acquisition module (330), a prompt generation module (340), and a control command transmission module (350). It should be understood that, depending on the embodiment, some configurations may be deleted or some configurations may be added.
[0081] The user voice acquisition module (310) is a module that acquires the user's voice via a microphone (120). In this case, the user's voice may be a user's voice for controlling an external device within the home IoT system (1). Specifically, the user's voice may not include specific commands that specifically specify the settings of the external device, but may include abstract commands such as "Create a comfortable environment in the house" or "Create a good environment for watching movies."
[0082] In one embodiment, the user voice acquisition module (310) may receive a user voice acquired from a microphone of the external device connected to the electronic device (100) from the external device. While the above-described embodiment describes receiving a user voice, this is merely an example, and it is of course possible to receive user input in the form of text from the user. In this case, if user input in the form of text is received, the text information acquisition module (320) may be omitted.
[0083] The text information acquisition module (320) converts the user's voice into text information. Specifically, the text information acquisition module (320) can input the user's voice into an STT (Speech-to-Text) module to acquire text information corresponding to the user's voice. In addition, the text information acquisition module (320) can transmit the user's voice to a server capable of converting the user's voice into text, and acquire text information corresponding to the user's voice from the server. The text information acquisition module (320) can transmit (or input) the text information corresponding to the user's voice to the LLM (200).
[0084] Meanwhile, in the above-described embodiment, it is described that the electronic device (100) directly obtains text information corresponding to the user's voice, but this is only one embodiment, and the electronic device (100) can directly transmit (or input) the user's voice to the LLM (200), and of course, the user's voice can be converted into text within the LLM (200).
[0085] The device search module (325) can search for at least one external device related to the user's voice and obtain information about the searched at least one external device. Specifically, the device search module (325) can obtain list information about a plurality of external devices (20-1, 20-2, 20-3, etc.) in the home IoT system (1) from the home IoT device (10). At this time, the list information may include information about external devices located in each space in the house. In addition, the device search module (325) can obtain current status information about the plurality of external devices (20-1, 20-2, 20-3, etc.) as well as the list information. The device search module (325) can transmit (input) information about the external devices to the LLM (200).
[0086] The LLM (200) can determine whether additional information is required to perform an action related to the user's voice based on text information corresponding to the user's voice and information about an external device. In this case, the LLM (200) can be trained to determine whether additional information is required based on text information corresponding to the user's voice and information about the external device.
[0087] In one embodiment, when a user voice (or text information) containing an abstract command is input, the LLM (200) may determine that additional information necessary to perform an operation related to the user voice exists. When a user voice containing a specific command is input, the LLM (200) may determine that no additional information necessary to perform an operation related to the user voice is required, and may thus obtain a control command to perform the operation related to the user voice.
[0088] If it is determined that additional information is needed, the LLM (200) may request a newly generated prompt with the additional information from the electronic device (100).
[0089] The additional information acquisition module (330) can acquire additional information for generating a prompt. At this time, the additional information acquisition module (330) can acquire the additional information through at least one of LLM, the Internet, and a user database.
[0090] In one embodiment, the additional information acquisition module (330) can acquire common sense information through the LLM (200). For example, if a user voice command such as "Create a pleasant environment" is acquired, the additional information acquisition module (300) can transmit a query for a "pleasant environment" to the LLM (200) and acquire common sense information about a "pleasant environment" through the LLM (200). That is, the additional information acquisition module (300) can acquire information about "temperature 23 to 25 degrees, humidity 40 to 60%", etc., as common sense information about a "pleasant environment."
[0091] In one embodiment, the additional information acquisition module (330) can acquire external environmental information related to a user's voice searched for via the Internet. For example, if the user's voice "Create a pleasant environment" is acquired, the additional information acquisition module (300) can acquire current weather information for the area where the home IoT system (1) is installed via the Internet.
[0092] In one embodiment, the additional information acquisition module (330) may acquire user history information related to a user voice searched through a user database. For example, if the user voice "Create a pleasant environment" is acquired, the additional information acquisition module (300) may acquire information about the user's preferred indoor environment or information about an external device recently set by the user through the user database.
[0093] The additional information acquisition module (330) may repeatedly perform operations to acquire additional information in order to generate a prompt. That is, until a prompt is generated, i.e., until it is determined that no additional information is needed, the additional information acquisition module (330) may access at least one of the LLM (200), the Internet, and a user database to acquire additional information.
[0094] The prompt generation module (340) can generate a prompt based on at least one of current status information for at least one external device related to the user voice, the user voice (text information), and acquired additional information.
[0095] In one embodiment, the prompt generation module (340) can generate a prompt based on a prompt template. That is, when multiple types of prompt templates are stored, the prompt generation module (340) can search for a prompt template for controlling at least one external device associated with the user's voice. For example, if it is determined that at least one external device associated with the user's voice is an air conditioner, the prompt generation module (340) can obtain a prompt template associated with the air conditioner from among the multiple types of prompt templates stored in advance. In addition, the prompt generation module (340) can obtain a prompt by applying at least one of current status information of the external device, text information corresponding to the user's voice, and additional information to the obtained prompt template associated with the air conditioner.
[0096] In one embodiment, the prompt generation module (340) may obtain a prompt by inputting at least one of current status information about an external device, user voice (text information), and acquired additional information into a prompt generation model trained to generate a prompt. In this case, the prompt generation model may be a neural network model trained to generate a prompt for controlling an external device by inputting at least one of current status information about the external device, user voice (text information), and acquired additional information.
[0097] At this time, the prompt generation module (340) can obtain at least one prompt corresponding to each of at least one external devices. For example, the prompt generation module (340) can generate a first prompt corresponding to an air conditioner and a second prompt corresponding to a humidifier. However, this is merely an example, and it is of course possible to generate one prompt for controlling at least one external device.
[0098] The prompt generation module (340) can input the generated prompt into the LLM (200) to obtain a control command corresponding to the prompt. That is, the LLM (200) can obtain a control command corresponding to the prompt.
[0099] The control command transmission module (350) can transmit the control command acquired by the LLM (200) to at least one external device. At this time, the control command transmission module (350) can map the control command acquired by the LLM (200) to an API (Application Programming Interface) related to the home IoT service and call it.
[0100] Additionally, the control command transmission module (350) can output a message to guide the user with information related to the control command. That is, the control command transmission module (350) can output a message including information related to the control command through the display (130) and can output a message including information related to the control command through the speaker.
[0101] FIG. 4 is a diagram illustrating a method for performing a weather-related home automation function according to one embodiment of the present disclosure.
[0102] As shown in Figure 4, the user can utter the user voice (410) “I will be home in 10 minutes. Make the home environment comfortable.”
[0103] The electronic device (100) can convert user speech (410) into text information. In addition, the electronic device (100) can retrieve information about external devices. For example, the electronic device (100) can retrieve information about external devices (e.g., air conditioners, windows, doors, humidifiers, blinds, lights, etc.) related to temperature, humidity, lighting, etc.
[0104] The electronic device (100) can transmit text information corresponding to the user's voice (410) to the large language model (200). At this time, since the large language model (200) cannot generate a control command with the user's voice (410), it can request additional information from the electronic device (100).
[0105] The electronic device (100) can obtain additional information (420) via the Internet in response to a request for additional information. At this time, the additional information (420) is weather information for the area where the home IoT system (1) is located. For example, the electronic device (100) can obtain additional information (420) such as "Phoenix, 13:00, Aug. 25, 105 degree Fahrenheit, Sunny skies, High 108F. Winds SSW at 5 to 10 mph" as illustrated in FIG. 4. In another embodiment, the electronic device (100) can obtain common sense information about a comfortable environment from the large language model (200) and obtain user history information of external devices related to temperature, humidity, lighting, etc. through the user database.
[0106] The electronic device (100) can generate a prompt using text information and additional information (420) corresponding to the user's voice (410). In particular, the electronic device (100) can generate a prompt by applying the text information and additional information (420) corresponding to the user's voice (410) to a prompt template.
[0107] And, the electronic device (100) can input the generated prompt into the large language model (200), and the large language model (200) can obtain a control command (430) as illustrated in FIG. 4. At this time, the control command (430) can include control commands (on / off information, operation setting information, etc.) for a plurality of external devices located in each space within the home, as illustrated in FIG. 4.
[0108] The electronic device (100) transmits the acquired control command to a corresponding external device so that it can perform an operation related to the user's voice.
[0109] FIG. 5 is a flowchart for explaining a method for controlling an electronic device according to one embodiment of the present disclosure.
[0110] The electronic device (100) acquires the user's voice (S510). Specifically, the electronic device (100) can acquire the user's voice through the microphone (120) of the electronic device (100) or the microphone of an external device.
[0111] The electronic device (100) can input the user's voice into the LLM (200). At this time, the electronic device (100) can directly input the user's voice into the LLM (200), but this is only one embodiment, and the electronic device (100) can obtain text information corresponding to the user's voice and input the obtained text information into the LLM (200).
[0112] Whether additional information is required can be identified (S530). Specifically, the LLM (200) can identify whether control commands can be generated via the user's voice, and if the control commands cannot be generated, it can identify that additional information is required.
[0113] If additional information is identified as unnecessary (S530-N), the electronic device (100) obtains a control command via the LLM (200) (S570).
[0114] If additional information is identified as being required (S530-Y), the electronic device (100) acquires additional information (S540). The electronic device (100) may acquire additional information through at least one of the LLM (200), the Internet, and a user database. Specifically, the additional information may include at least one of common sense information acquired through the LLM (200), external environmental information related to the user's voice searched through the Internet, and user history information related to the user's voice searched through the user database.
[0115] The electronic device (100) generates a prompt (S550). Specifically, the electronic device (100) may generate a prompt for controlling an external device associated with the user's voice, including information about the user's voice and additional information. In one embodiment, the electronic device (100) may obtain a prompt template for generating the prompt, and may insert information about the user's voice and additional information into the prompt template to obtain the prompt.
[0116] The electronic device (100) inputs the generated prompt into the LLM (200) (S560), and the electronic device (100) obtains a control command through the LLM (200) (S570). Then, the electronic device (100) can transmit the obtained control command to at least one external device for performing an operation related to the user's voice.
[0117] Meanwhile, in one embodiment, when a user's voice is input, the electronic device (100) can obtain list information including a plurality of external devices connected to the home IoT server (10) through the home IoT server (10) connected to the electronic device (100). In addition, the electronic device (100) can obtain current status information of the plurality of external devices included in the list information.
[0118] In one embodiment, the electronic device (100) may obtain a prompt for controlling at least one external device based on current state information, user voice, and additional information about at least one external device among a plurality of external devices.
[0119] In one embodiment, the electronic device (100) can map a control command obtained through the LLM (200) to an API (Application Programming Interface) related to a home IoT service and call it.
[0120] In one embodiment, the electronic device (100) may output a message to guide the user with information related to a control command.
[0121] By the embodiment described above, the user does not need to individually set the operating conditions and operation contents of an external device to perform a home automation function, and additional information is obtained through various sources, eliminating the need for additional sensors.
[0122] The artificial intelligence-related functions according to the present disclosure are operated through the processor and memory of the electronic device (100).
[0123] The processor may be composed of one or more processors. In this case, the one or more processors may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an NPU (Neural Processing Unit), but is not limited to the examples of the processors described above.
[0124] CPUs are general-purpose processors capable of performing not only general calculations but also artificial intelligence calculations. Their multi-layered cache structure allows for the efficient execution of complex programs. CPUs are advantageous for serial processing, enabling organic linking of previous and subsequent calculation results through sequential calculations. General-purpose processors are not limited to the examples described above, except where specifically identified as CPUs.
[0125] A GPU is a processor designed for large-scale computations, such as floating-point operations used in graphics processing. It integrates a large number of cores to perform large-scale computations in parallel. In particular, GPUs may be advantageous over CPUs in parallel processing methods, such as convolution operations. Furthermore, GPUs can be used as coprocessors to supplement the functions of CPUs. Processors for large-scale computations are not limited to the examples described above, except in cases where they are specifically referred to as GPUs.
[0126] An NPU is a processor specialized in artificial intelligence computation using artificial neural networks, and each layer of the artificial neural network can be implemented in hardware (e.g., silicon). Since NPUs are designed specifically according to the company's specifications, they have less freedom than CPUs or GPUs, but can efficiently process the AI computations requested by the company. Meanwhile, as a processor specialized in AI computation, an NPU can be implemented in various forms, such as a Tensor Processing Unit (TPU), an Intelligence Processing Unit (IPU), or a Vision Processing Unit (VPU). Except as specifically designated as an NPU, an AI processor is not limited to the examples described above.
[0127] Additionally, one or more processors may be implemented as a System on Chip (SoC). In this case, in addition to one or more processors, the SoC may further include memory and a network interface, such as a bus, for data communication between the processor and the memory.
[0128] When a System on Chip (SoC) included in an electronic device includes multiple processors, the electronic device may use some of the multiple processors to perform operations related to artificial intelligence (e.g., operations related to learning or inference of an artificial intelligence model). For example, the electronic device may use at least one of a GPU, NPU, VPU, TPU, or hardware accelerator specialized in artificial intelligence operations, such as convolution operations or matrix multiplication operations, among the multiple processors to perform operations related to artificial intelligence. However, this is merely an example, and it is of course possible to process operations related to artificial intelligence using a CPU or other general-purpose processor.
[0129] Additionally, electronic devices can perform computations related to artificial intelligence (AI) functions by utilizing multiple cores (e.g., dual cores, quad cores, etc.) contained within a single processor. In particular, electronic devices can perform AI operations, such as convolution operations and matrix multiplication operations, in parallel by utilizing the multiple cores contained within the processor.
[0130] One or more processors are controlled to process input data according to predefined operating rules or artificial intelligence models stored in memory. The predefined operating rules or artificial intelligence models are characterized by being created through learning.
[0131] Here, "created through learning" means that a predefined set of behavioral rules or an AI model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the AI according to the present disclosure is implemented, or through a separate server / system.
[0132] An artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0133] A learning algorithm is a method for training a target device (e.g., a robot) using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.
[0134] Meanwhile, the methods according to various embodiments of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0135] The methods according to various embodiments of the present disclosure may be implemented as software including commands stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call commands stored in the storage medium and operate according to the called commands, and may include an electronic device (e.g., a TV) according to the disclosed embodiments.
[0136] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0137] When the above instruction is executed by the processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter.
[0138] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In electronic devices, communication interface; mike; memory for storing at least one instruction; and A processor connected to the communication interface, the microphone and the memory for controlling the electronic device; The above processor, by executing at least one instruction, When a user's voice is input through the microphone to control an external device, information about at least one external device related to the user's voice is acquired, By inputting information about the user's voice into the Large Language Model, if it is determined that there is additional information necessary to perform an action related to the user's voice, additional information related to the user's voice is acquired, Obtaining a prompt for controlling the at least one external device based on the user voice and the additional information; By inputting the above prompt into the above giant language model, a control command for controlling the at least one external device is obtained, An electronic device that transmits the control command to at least one external device via the communication interface.
2. In paragraph 1, The above processor, When the user voice is input, list information including multiple external devices connected to the home IoT server is obtained through the home IoT server connected to the electronic device, An electronic device that obtains current status information of a plurality of external devices included in the above list information.
3. In paragraph 2, The above processor, An electronic device that obtains a prompt for controlling at least one external device based on current status information of at least one external device among the plurality of external devices, the user voice, and the additional information.
4. In paragraph 1, The above processor, Obtain a prompt template for generating the above prompt, An electronic device for obtaining said prompt by inserting information about said user voice and said additional information into said prompt template.
5. In paragraph 1, The above processor, An electronic device that inputs information about the user's voice into a large language model and determines that there is additional information necessary to perform an action related to the user's voice, and acquires additional information through at least one of the large language model, the Internet, and a user database.
6. In paragraph 4, The above additional information is: An electronic device including at least one of common sense information obtained through the large language model, external environment information related to the user voice searched through the Internet, and user history information related to the user voice searched through a user database.
7. In paragraph 1, The above processor, An electronic device that calls a control command obtained through the above-mentioned large language model by mapping it to an API (Application Programming Interface) related to a home IoT service.
8. In paragraph 1, The above processor, An electronic device that outputs a message to guide a user to information related to the above control command.
9. In a method for controlling an electronic device, When a user voice is input to control an external device, a step of obtaining information about at least one external device related to the user voice; A step of obtaining additional information related to the user voice when it is determined that additional information necessary to perform an action related to the user voice exists by inputting information about the user voice into a large language model; A step of obtaining a prompt for controlling the at least one external device based on the user voice and the additional information; A step of inputting the above prompt into the large language model to obtain a control command for controlling the at least one external device; and A control method comprising the step of transmitting the control command to at least one external device.
10. In paragraph 9, The step of obtaining information about at least one external device comprises: When the user voice is input, list information including multiple external devices connected to the home IoT server is obtained through the home IoT server connected to the electronic device, A control method for obtaining current status information of a plurality of external devices included in the above list information.
11. In paragraph 10, The steps to obtain the above prompt are: A control method for obtaining a prompt for controlling at least one external device based on current status information of at least one external device among the plurality of external devices, the user voice, and the additional information.
12. In paragraph 9, The steps to obtain the above prompt are: Obtain a prompt template for generating the above prompt, A control method for obtaining the prompt by inserting information about the user's voice and the additional information into the prompt template.
13. In paragraph 9, The steps for obtaining the above additional information are: A control method for acquiring additional information through at least one of the giant language model, the Internet, and a user database, when it is determined that additional information necessary to perform an action related to the user voice exists by inputting information about the user voice into the giant language model.
14. In paragraph 13, The above additional information is: A control method including at least one of common sense information obtained through the large language model, external environment information related to the user voice searched through the Internet, and user history information related to the user voice searched through a user database.
15. In paragraph 9, A control method comprising: a step of mapping a control command obtained through the above-mentioned large language model to an API (Application Programming Interface) related to a home IoT service and calling the same;
Citation Information
Patent Citations
History-based key phrase suggestions for voice control of a home automation system
KR1020180064328A
Method of tool break detection in machine tools
KR1020210044372A
Catalyst for hydrogen evolution reaction and preparing method thereof
KR1020240078581A
Underground facility survey system for identifying position change of underground facility
KR102442031B1
Using large language model(s) in generating automated assistant response(s)
WO2023038654A1