Artificial intelligence device and operation method thereof
The AI device addresses the issue of outdated data in generative AI servers by using a document database with up-to-date information to provide accurate responses to user queries, enhancing the accuracy and relevance of information provided to users.
Patent Information
- Application Number
- PCT/KR2023/020803
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-19
AI Technical Summary
Conventional generative AI servers lack up-to-date data, leading to a high probability of providing error information instead of accurate answers, and traditional TVs struggle to find information related to entity names without metadata.
An artificial intelligence device that communicates with electronic devices and AI generation servers, utilizing a document database with up-to-date information to generate prompts for AI servers, ensuring accurate and relevant responses to user queries.
The solution enables the provision of diverse and accurate information desired by users by leveraging document and category databases with up-to-date information, overcoming the limitations of outdated data in conventional systems.
Smart Images

Figure KR2023020803_19062025_PF_FP_ABST
Abstract
Description
Artificial intelligence device and its operating method
[0001] The present disclosure relates to an artificial intelligence device, and more particularly, to an artificial intelligence device capable of providing a response result through a generative artificial intelligence model.
[0002] Digital TV services utilizing wired or wireless networks are becoming more widespread. Digital TV services can offer a variety of services not available with existing analog broadcasting services.
[0003] For example, IPTV (Internet Protocol Television), a type of digital TV service, and smart TV services offer interactivity, allowing users to actively choose the type of program they want to watch and when. Building on this interactivity, IPTV and smart TV services can also offer a variety of additional services, such as internet search, home shopping, and online games.
[0004] Recently, TVs are providing voice recognition services by analyzing voice commands given by users.
[0005] The generative AI models of conventional generative AI servers do not have the latest data, so there is a high probability that they will provide error information rather than the answers that the user wants, using past data or a large amount of web documents.
[0006] Furthermore, conventional TVs retrieved information about entity names one-dimensionally, using metadata centered on TV content. Therefore, in the absence of metadata, it was difficult to find information related to entity names.
[0007] The purpose of this disclosure is to provide up-to-date / linked content by utilizing data from a knowledge document database containing up-to-date information to overcome the limitation of lacking up-to-date information when processing a large language model.
[0008] The purpose of this disclosure is to provide various forms of information linked to object commands when processing a large language model.
[0009] An artificial intelligence device according to an embodiment of the present disclosure may include a communication unit that communicates with an electronic device and one or more generation artificial intelligence (AI) servers, and a processor that obtains an analysis result including an entity name based on text data corresponding to a voice command uttered by a user through the electronic device, obtains latest information on the entity name from a document database, generates a prompt based on the analysis result and the latest information, transmits the generated prompt to the generation AI server, receives a response result corresponding to the prompt from the generation AI server, and transmits the response result to the electronic device.
[0010] An operating method of an artificial intelligence device according to an embodiment of the present disclosure may include a step of obtaining an analysis result including an entity name based on text data corresponding to a voice command uttered by a user through an electronic device, a step of obtaining latest information on the entity name from a document database, a step of generating a prompt based on the analysis result and the latest information, a step of transmitting the generated prompt to a generation AI server, a step of receiving a response result corresponding to the prompt from the generation AI server, and a step of transmitting the response result to the electronic device.
[0011] According to an embodiment of the present disclosure, there is an effect of being able to provide diverse and accurate information desired by a user by using a document DB and a category DB that have the latest information.
[0012] Figure 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.
[0013] FIG. 2 illustrates an artificial intelligence (AI) server according to one embodiment of the present disclosure.
[0014] FIG. 3 is a diagram illustrating the configuration of an artificial intelligence system according to one embodiment of the present disclosure.
[0015] FIG. 4 is a ladder diagram for explaining an operation method of an artificial intelligence system according to an embodiment of the present disclosure.
[0016] FIGS. 5A and 5B are diagrams illustrating an example of providing linked content by utilizing up-to-date information for a voice command uttered by a user according to one embodiment of the present disclosure.
[0017] FIGS. 6A and 6B are diagrams illustrating an example of providing linked content by utilizing up-to-date information for a voice command uttered by a user according to another embodiment of the present disclosure.
[0018] FIGS. 7A and 7B are diagrams illustrating an example of providing linked content by utilizing up-to-date information for a voice command uttered by a user according to another embodiment of the present disclosure.
[0019] Hereinafter, embodiments related to the present invention will be described in more detail with reference to the drawings. The suffixes "module" and "part" used in the following description for components are assigned or used interchangeably solely for the convenience of writing the specification, and do not in themselves have distinct meanings or roles.
[0020] A display device according to an embodiment of the present invention is, for example, an intelligent display device that adds computer-assisted functionality to its broadcast reception function. While faithfully performing the broadcast reception function, it can also be equipped with Internet functionality and other features, providing a more user-friendly interface, such as a manual input device, touch screen, or space remote control. Furthermore, with support for wired or wireless Internet functionality, it can connect to the Internet and a computer, enabling functions such as email, web browsing, banking, or gaming. A standardized, general-purpose operating system can be used for these various functions.
[0021] Accordingly, the display device described in the present invention can perform various user-friendly functions, for example, by allowing various applications to be freely added or deleted on a general-purpose operating system kernel. More specifically, the display device can be a network TV, HBB TV, smart TV, LED TV, OLED TV, etc., and in some cases, it can also be applied to smartphones.
[0022] FIG. 1 is a block diagram illustrating the configuration of a display device according to one embodiment of the present invention.
[0023] Referring to FIG. 1, the display device (100) may include a broadcast receiving unit (130), an external device interface (135), a memory (140), a user input interface (150), a controller (170), a wireless communication interface (173), a display (180), a speaker (185), and a power supply circuit (190).
[0024] The broadcast receiving unit (130) may include a tuner (131), a demodulator (132), and a network interface (133).
[0025] The tuner (131) can select a specific broadcast channel according to a channel selection command. The tuner (131) can receive a broadcast signal for the selected specific broadcast channel.
[0026] The demodulator (132) can separate the received broadcast signal into a video signal, an audio signal, and a data signal related to the broadcast program, and can restore the separated video signal, audio signal, and data signal into a form that can be output.
[0027] The external device interface (135) can receive an application or a list of applications within an adjacent external device and transmit it to the controller (170) or memory (140).
[0028] The external device interface (135) can provide a connection path between the display device (100) and the external device. The external device interface (135) can receive one or more of images and audio output from an external device connected wirelessly or wiredly to the display device (100) and transmit them to the controller (170). The external device interface (135) can include a plurality of external input terminals. The plurality of external input terminals can include an RGB terminal, one or more HDMI (High Definition Multimedia Interface) terminals, and a component terminal.
[0029] A video signal of an external device input through an external device interface (135) can be output through a display (180). A voice signal of an external device input through an external device interface (135) can be output through a speaker (185).
[0030] An external device that can be connected to the external device interface (135) may be any one of a set-top box, a Blu-ray player, a DVD player, a game console, a sound bar, a smartphone, a PC, a USB memory, and a home theater, but this is only an example.
[0031] The network interface (133) may provide an interface for connecting the display device (100) to a wired / wireless network including the Internet. The network interface (133) may transmit or receive data to or from other users or other electronic devices via the connected network or another network linked to the connected network.
[0032] Additionally, some content data stored in the display device (100) can be transmitted to a selected user or electronic device among other users or other electronic devices pre-registered in the display device (100).
[0033] The network interface (133) can access a predetermined web page through a connected network or another network linked to the connected network. That is, by accessing a predetermined web page through a network, data can be transmitted or received with the corresponding server.
[0034] In addition, the network interface (133) can receive content or data provided by a content provider or network operator. That is, the network interface (133) can receive content such as movies, advertisements, games, VOD, broadcast signals, etc. and information related thereto provided from a content provider or network provider via a network.
[0035] Additionally, the network interface (133) can receive firmware update information and update files provided by the network operator, and can transmit data to the Internet or content provider or network operator.
[0036] The network interface (133) can select and receive a desired application from among applications open to the public via a network.
[0037] The memory (140) stores a program for each signal processing and control within the controller (170), and can store signal-processed image, voice, or data signals.
[0038] In addition, the memory (140) may perform a function for temporary storage of video, audio, or data signals input from an external device interface (135) or a network interface (133), and may store information about a specific image through a channel memory function.
[0039] The memory (140) can store an application or a list of applications input from an external device interface (135) or a network interface (133).
[0040] The display device (100) can play content files (video files, still image files, music files, document files, application files, etc.) stored in the memory (140) and provide them to the user.
[0041] The user input interface (150) can transmit a signal input by the user to the controller (170) or transmit a signal from the controller (170) to the user. For example, the user input interface (150) can receive and process control signals such as power on / off, channel selection, and screen setting from the remote control device (200) according to various communication methods such as Bluetooth, Ultra Wideband (WB), ZigBee, Radio Frequency (RF) communication, or infrared (IR) communication, or process control signals from the controller (170) to be transmitted to the remote control device (200).
[0042] In addition, the user input interface (150) can transmit control signals input from local keys (not shown) such as a power key, channel key, volume key, and setting value to the controller (170).
[0043] An image signal processed by the controller (170) can be input to the display (180) and displayed as an image corresponding to the image signal. In addition, an image signal processed by the controller (170) can be input to an external output device through an external device interface (135).
[0044] The voice signal processed by the controller (170) can be output as audio to the speaker (185). In addition, the voice signal processed by the controller (170) can be input to an external output device through the external device interface (135).
[0045] In addition, the controller (170) can control the overall operation within the display device (100).
[0046] In addition, the controller (170) can control the display device (100) by a user command or internal program input through the user input interface (150), and can connect to a network to enable the user to download a desired application or application list into the display device (100).
[0047] The controller (170) enables the user-selected channel information, etc. to be output through a display (180) or speaker (185) together with processed video or audio signals.
[0048] In addition, the controller (170) allows a video signal or audio signal from an external device, for example, a camera or camcorder, input through the external device interface (135) to be output through the display (180) or speaker (185) in accordance with an external device video playback command received through the user input interface (150).
[0049] Meanwhile, the controller (170) can control the display (180) to display an image, for example, a broadcast image input through a tuner (131), an external input image input through an external device interface (135), an image input through a network interface, or an image stored in a memory (140) can be controlled to be displayed on the display (180). In this case, the image displayed on the display (180) can be a still image or a moving image, and can be a 2D image or a 3D image.
[0050] In addition, the controller (170) can control the playback of content stored in the display device (100), received broadcast content, or external input content input from outside, and the content can be in various forms such as broadcast video, external input video, audio file, still image, connected web screen, and document file.
[0051] The wireless communication interface (173) can communicate with an external device through wired or wireless communication. The wireless communication interface (173) can perform short-range communication with the external device. To this end, the wireless communication interface (173) can support short-range communication using at least one of Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies. This wireless communication interface (173) can support wireless communication between the display device (100) and a wireless communication system, between the display device (100) and another display device (100), or between the display device (100) and a network in which the display device (100, or an external server) is located via a short-range wireless communication network (Wireless Area Network). The short-range wireless communication network can be a short-range wireless personal area network (Wireless Personal Area Network).
[0052] Here, the other display device (100) may be a wearable device (e.g., a smartwatch, smart glasses, a head-mounted display (HMD)) or a mobile terminal such as a smart phone that can exchange data with (or be linked to) the display device (100) according to the present invention. The wireless communication interface (173) may detect (or recognize) a wearable device capable of communication around the display device (100).
[0053] Furthermore, if the detected wearable device is a device certified to communicate with the display device (100) according to the present invention, the controller (170) can transmit at least a portion of the data processed in the display device (100) to the wearable device via the wireless communication interface (173). Accordingly, a user of the wearable device can utilize the data processed in the display device (100) via the wearable device.
[0054] The display (180) can generate a driving signal by converting a video signal, data signal, OSD signal processed by the controller (170) or a video signal, data signal, etc. received from an external device interface (135) into R, G, and B signals, respectively.
[0055] Meanwhile, since the display device (100) illustrated in FIG. 1 is merely an embodiment of the present invention, some of the illustrated components may be integrated, added, or omitted depending on the specifications of the display device (100) actually implemented.
[0056] That is, two or more components may be combined into a single component, or a single component may be subdivided into two or more components, as needed. Furthermore, the functions performed by each block are intended to illustrate embodiments of the present invention, and their specific operations or devices do not limit the scope of the present invention.
[0057] According to another embodiment of the present invention, the display device (100) may receive and play back an image through a network interface (133) or an external device interface (135) without having a tuner (131) and a demodulator (132), unlike that shown in FIG. 1.
[0058] For example, the display device (100) may be implemented separately as an image processing device, such as a set-top box, for receiving broadcast signals or contents according to various network services, and a content playback device for playing contents input from the image processing device.
[0059] In this case, the operating method of the display device according to the embodiment of the present invention described below may be performed by any one of the display device (100) described with reference to FIG. 1, as well as an image processing device such as the separated set-top box, or a content playback device having a display (180) and an audio output unit (185).
[0060] FIG. 2 illustrates an artificial intelligence (AI) server according to one embodiment of the present disclosure.
[0061] Referring to FIG. 2, the AI server (200) may refer to a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network.
[0062] The AI server (200) may be composed of multiple servers to perform distributed processing, and may be defined as a 5G network.
[0063] The AI server (200) may be included as part of the configuration of the AI device (100) and may perform at least part of the AI processing together.
[0064] The AI server (200) may include a communication unit (210), a memory (230), a learning processor (240), and a processor (260).
[0065] The communication unit (210) can transmit and receive data with an external device such as a display device (100).
[0066] The memory (230) may include a model storage unit (231).
[0067] The model storage unit (231) can store a model (or artificial neural network, 231a) being learned or learned through the learning processor (240).
[0068] A learning processor (240) can train an artificial neural network (231a) using learning data. The learning model can be used while mounted on the AI server (200) of the artificial neural network, or can be used while mounted on an external device such as a display device (100).
[0069] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (230).
[0070] The processor (260) can use a learning model to infer a result value for new input data and generate a response or control command based on the inferred result value.
[0071] FIG. 3 is a diagram illustrating the configuration of an artificial intelligence system according to one embodiment of the present disclosure.
[0072] The artificial intelligence system (30) may include a display device (100), a natural language processing (NLP) server (200-1), an AI server (200-2), and a generation AI server (300).
[0073] Each of the NLP server (200-1) and the AI server (200-2) may include all of the components of the AI server (200) illustrated in FIG. 2.
[0074] The AI server (200-2) may be referred to as an AI device.
[0075] The display device (100) can obtain a voice command spoken by a user.
[0076] The display device (100) can transmit voice data corresponding to a voice command to the NLP server (200-1).
[0077] The NLP server (200-1) can obtain analysis results of voice commands based on received voice data.
[0078] The NLP server (200-1) can determine whether a response to the analysis result is possible based on the analysis result of the voice command.
[0079] If the NLP server (200-1) determines that a response to the analysis result is impossible, it can transmit text data corresponding to the voice command to the AI server (200-2).
[0080] The AI server (200-2) can obtain the user's intention type and domain based on the received text data.
[0081] The AI server (200-2) can determine whether there is up-to-date information about the acquired entity name.
[0082] If the AI server (200-2) determines that up-to-date information about the entity name exists, it can obtain multiple category items.
[0083] The AI server (200-2) can generate a prompt based on the latest information about the entity name, the user's intention, the entity name, the domain, and multiple category items, and can transmit the generated prompt to the generation AI server (300).
[0084] The NLP server (200-1) and the AI server (200-2) may be configured as a single server. The AI server (200-2) may perform all the functions of the NLP server (200-1).
[0085] The generation AI server (300) can generate a first response result based on the received prompt and transmit the generated first response result to the AI server (200-2).
[0086] The AI server (200-2) can generate a second response result based on the first response result received from the generation AI server (300) and transmit the generated second response result to the NLP server (200-1).
[0087] The NLP server (200-1) can transmit the received second response result to the display device (100).
[0088] The display device (100) can output the received second response result.
[0089] The artificial intelligence system (30) may further include an LLM (Large Language Model) prompt DB (310), a document DB (320), and a category DB (330).
[0090] The LLM prompt DB (310) may include a plurality of prompt commands. Each of the plurality of prompt commands may be a command matching one or more of an entity name, a user's intent type, and a domain.
[0091] The document DB (320) may be a database containing up-to-date information on people or content. The information stored in the document DB (320) may be periodically updated.
[0092] The document DB (320) may be a website server that provides knowledge documents.
[0093] The category DB (330) may store in advance multiple category items that match the user's intention type and domain.
[0094] FIG. 4 is a ladder diagram for explaining an operation method of an artificial intelligence system according to an embodiment of the present disclosure.
[0095] Each of the NLP server (200-1) and the AI server (200-2) may include all of the components of the AI server (200) illustrated in FIG. 2.
[0096] The controller (170) of the display device (100) can obtain a voice command (S401).
[0097] In one embodiment, the controller (170) can receive a voice command spoken by a user through a microphone (not shown).
[0098] The controller (170) can receive a voice command from a remote control device (200). The remote control device (200) can receive a voice command spoken by a user and transmit it to the display device (100).
[0099] The controller (170) of the display device (100) can transmit voice data corresponding to a voice command to the NLP server (200-1) via the network interface (133) (S403).
[0100] The processor (260) of the NLP server (200-1) can obtain the analysis result of the voice command based on the received voice data (S605).
[0101] In one embodiment, the processor (260) can convert voice data into text data using a speech to text (STT) engine stored in the memory (230).
[0102] In another embodiment, the processor (260) can transmit voice data to an STT server and receive converted text data from the STT server.
[0103] The processor (260) can obtain analysis results of the converted text data using a natural language processing engine (NLP).
[0104] The analysis results may include an entity name. In one embodiment, the entity name may be any one of a person's name, an object's name, or a content's name.
[0105] The processor (260) of the NLP server (200-1) can determine whether a response to the analysis result is possible based on the analysis result of the voice command (S407).
[0106] In one embodiment, the processor (260) may determine that a response is possible if a search result for the analysis result of the voice command is available, and may determine that a response is impossible if a search result for the analysis result is not available.
[0107] If the NLP server (200-1) determines that a response to the analysis result is impossible, it can transmit text data corresponding to the voice command to the AI server (200-2) (S409).
[0108] The NLP server (200-1) can transmit text data and analysis results together to the AI server (200-2). The analysis results may include entity names included in the text data.
[0109] The processor (260) of the AI server (200-2) can obtain the user's intention type and domain based on the received text data (S411).
[0110] In one embodiment, the user's intent may be either content search or query response. That is, the user's intent type may be either a search type for content search or a query response type for receiving a response to a query.
[0111] In one embodiment, a domain may be a genre, field, or content related to the entity name. For example, if the entity name is the name of a sports player, the domain may be "sports." Another example is if the entity name is the title of a movie, the domain may be "movies."
[0112] That is, a domain can represent either a genre or a type of content.
[0113] The processor (260) of the AI server (200-2) can determine whether there is up-to-date information about the acquired entity name (S413).
[0114] The processor (260) can determine whether the latest information on the entity name exists from the document DB (320). The processor (260) can inquire whether the latest information on the entity name exists in the document DB (320), and if the latest information exists, can receive the latest information from the document DB (320).
[0115] The document DB (320) may be a server of a website that stores the latest knowledge.
[0116] If the processor (260) of the AI server (200-2) determines that the latest information on the entity name exists, it can obtain multiple category items (S415).
[0117] The processor (260) can obtain multiple category items for the domain of the entity name.
[0118] For example, if the entity name is a person's name and the domain is sports, multiple category entries may include the entity name's current team, past teams, popular games, and latest games.
[0119] As another example, if the entity name is the singer's name and the domain is music, multiple category entries may include the entity name's latest songs, popular songs, and album names.
[0120] The processor (260) can obtain a plurality of category items from the category DB (330). The processor (260) can transmit an entity name and domain to the category DB (330) through the communication unit (210), and can receive a plurality of category items from the category DB (330).
[0121] The processor (260) of the AI server (200-2) can generate a prompt based on the latest information on the entity name, the user's intention, the entity name, the domain, and multiple category items (S417), and can transmit the generated prompt to the generation AI server (300) (S419).
[0122] The processor (260) can generate a prompt for a search command by combining the latest information about the entity name, the user's intent type, the entity name, the domain, and multiple category items.
[0123] For example, a prompt might be a command to provide a response tailored to the user's intent using the entity name, domain, multiple category items, and up-to-date information about the entity name.
[0124] A prompt can be a text command that is input to a generative AI model.
[0125] The processor (260) can generate a prompt based on a prompt command extracted from the LLM prompt DB (310), the latest information on entity names extracted from the document DB (320), and a plurality of category items extracted from the category DB (330).
[0126] The generation AI server (300) can generate a first response result based on the received prompt and transmit the generated first response result to the AI server (200-2) (S421).
[0127] The generation AI server (300) may be equipped with a generation AI model.
[0128] Generative AI models can be models that use deep learning or machine learning to generate responses as output to prompts.
[0129] A generative AI model can be a natural language generation AI model that outputs natural language.
[0130] A natural language generation AI model can generate a first response result in response to a received prompt.
[0131] Natural language generation AI models can be trained using deep learning architectures such as recurrent neural networks (RNNs) and transformation models using attention mechanisms through large-scale text data.
[0132] Data for training natural language generation AI models can be collected from the web or various text sources.
[0133] A natural language generation AI model can take a received prompt as input and generate one or more sentences containing the content.
[0134] The generation AI server (300) can obtain one or more generated sentences as a first response result and transmit the obtained first response result to the AI server (200-2).
[0135] The processor (260) of the AI server (200-2) can generate a second response result based on the first response result received from the generation AI server (300) (S423) and transmit the generated second response result to the NLP server (200-1) (S425).
[0136] The processor (260) can obtain a second response result including content information related to the first response result based on the first response result.
[0137] The processor (260) may request content information corresponding to a plurality of category items included in the first response result from an external server and receive the content information from the external server. The external server may be either a web server or a streaming content provider server, but this is merely an example.
[0138] The content information may include one or more of the name, access address, thumbnail image, or video of the content corresponding to each of the plurality of category items.
[0139] In one embodiment, the processor (260) of the AI server (200-2) can directly transmit the second response result to the display device (100) via the communication unit (210).
[0140] The NLP server (200-1) can transmit the received second response result to the display device (100) (S427).
[0141] The display device (100) can output the received second response result (S429).
[0142] The controller (170) of the display device (100) can display the second response result on the display (180). The display device (100) can display content corresponding to each category item in the form of a user interface (UI).
[0143] Meanwhile, if the NLP server (200-1) determines that a response is possible based on the analysis results of voice data, it can obtain a third response result and transmit the obtained third response result to the display device (100) (S431).
[0144] The third response result may be a result that the NLP server (200-1) can provide based on the analysis results of the voice command uttered by the user. The third response result may include information regarding the function control of the display device (100).
[0145] The display device (100) can output the received third response result (S433).
[0146] FIGS. 5A and 5B are diagrams illustrating an example of providing linked content by utilizing up-to-date information for a voice command uttered by a user according to one embodiment of the present disclosure.
[0147] In particular, FIGS. 5a and 5b are diagrams illustrating examples of providing response results for inquiring about the team of a sports player named James.
[0148] In Figures 5a and 5b, the user<What team does James belong to> Speak the voice command (510).
[0149] The display device (100) can transmit voice data corresponding to a voice command (510) uttered by the user to the NLP server (200-1).
[0150] The NLP server (200-1) can convert voice data into text data using a voice-to-text conversion engine, and can obtain analysis results (520) for the converted text data using a natural language processing engine.
[0151] In one embodiment, the analysis results (520) may include a user query, an action (unknown), and an entity name (James).
[0152] The NLP server (200-1) may generate an analysis result (520) including an action (unknown) indicating that it cannot respond because there is no answer to the question about James's team based on the analysis of the text data.
[0153] However, the NLP server (200-1) can transmit text data and entity name (James) corresponding to the voice command to the AI server (200-2).
[0154] The AI server (200-2) can obtain additional analysis results (530) including the user's intention type and the domain of the entity name based on text data and entity name (James).
[0155] The AI server (200-2) can determine whether the user's intention type is a search type (or content search type) or a question response type based on intent analysis of text data.
[0156] The AI server (200-2) can analyze a user's intention type through speech act analysis. Speech act analysis may be a process of analyzing the intent of text data by considering the context and meaning of the text data.
[0157] In Fig. 5a, the user's intent type may be a question-response type, and the entity name domain may be sports.
[0158] The AI server (200-2) can obtain the latest information about the entity name (James) through the document DB (320). The AI server (200-2) can request the latest information about the entity name (James) from the document DB (320) and receive the latest information about the entity name (James) from the document DB (320).
[0159] The AI server (200-2) can receive a prompt command from the LLM prompt DB (310).
[0160] If the latest information on the entity name (James) exists, the AI server (200-2) can request a prompt command suitable for the intent type and domain from the LLM prompt DB (310) and receive the prompt command from the LLM prompt DB (310).
[0161] The prompt command may be a command extracted from the LLM prompt DB (310). Here, the prompt command is<I'll also give you the latest information. Please answer based on this information> It could be.
[0162] <document content> may contain the latest information about the entity name (James) extracted from the document DB (320).
[0163] The AI server (200-2) can generate a prompt (540) that combines text data corresponding to the voice command spoken by the user, prompt commands, and the latest information about the entity name (James).
[0164] Prompt (540) is text data corresponding to a voice command spoken by the user.<user query> , the prompt command <command word> and latest information on the entity name (James)<document content> may include.
[0165] The AI server (200-2) can transmit a prompt (540) to the generation AI server (300) and receive a response result (550) corresponding to the prompt (540) from the generation AI server (300).
[0166] The response result (550) may include up-to-date information about the team to which a player named James belongs.
[0167] The AI server (200-2) can transmit the response result (550) to the NLP server (200-1) or directly to the display device (100).
[0168] Referring to FIG. 5b, the display device (100) can display a response result (550) on the display (180). The response result (550) can include information about the team that player James belongs to, which is a response to the voice command (510) uttered by the user.
[0169] The generation AI model of the conventional generation AI server (300) does not have the latest data, so there is a high probability that it will provide error information rather than the answer the user wants, using past data or a large amount of web documents.
[0170] According to an embodiment of the present disclosure, there is an effect that can provide the user with the information he or she wants accurately by using a document DB (320) that has the latest information, thereby overcoming these limitations.
[0171] FIGS. 6A and 6B are diagrams illustrating an example of providing linked content by utilizing up-to-date information for a voice command uttered by a user according to another embodiment of the present disclosure.
[0172] In particular, FIGS. 6a and 6b show that the user <james>This diagram illustrates an example of a response result provided to a sports player who uttered the word "."
[0173] In Figures 6a and 6b, the user <james>Speak the voice command (610).
[0174] The display device (100) can transmit voice data corresponding to a voice command (610) uttered by the user to the NLP server (200-1).
[0175] The NLP server (200-1) can convert voice data into text data using a voice-to-text conversion engine, and can obtain analysis results (520) for the converted text data using a natural language processing engine.
[0176] In one embodiment, the analysis results (620) may include a user query, an action (unknown), and an entity name (James).
[0177] The NLP server (200-1) may generate an analysis result (620) that includes an action (unknown) indicating that it cannot respond because there is no action definition for a sports player named James based on the analysis of the text data.
[0178] However, the NLP server (200-1) can transmit the entity name (James) to the AI server (200-2).
[0179] The AI server (200-2) can obtain additional analysis results (630) including the user's intention type and the domain of the entity name based on the entity name (James).
[0180] The AI server (200-2) can determine whether the user's intention type is a search type or a question response type based on intent analysis of text data.
[0181] The AI server (200-2) can analyze a user's intention type through speech act analysis. Speech act analysis may be a process of analyzing the intent of text data by considering the context and meaning of the text data.
[0182] The AI server (200-2) can extract a domain for an entity name from an entity name DB (not shown) that includes entity names and domains matching the entity names.
[0183] In Figure 6a, the user's intent type may be a search type, and the entity name domain may be sports.
[0184] The AI server (200-2) can obtain the latest information about the entity name (James) through the document DB (320). The AI server (200-2) can request the latest information about the entity name (James) from the document DB (320) and receive the latest information about the entity name (James) from the document DB (320).
[0185] The AI server (200-2) can transmit information about the intent type and domain to the category DB (330) and receive a plurality of category items matching the intent type and domain from the category DB (330).
[0186] The category DB (330) may pre-store multiple category items matching the intent type and domain. The multiple category items may include current team items, past team items, and popular game items.
[0187] The AI server (200-2) can receive a prompt command from the LLM prompt DB (310).
[0188] If the latest information on the entity name (James) exists, the AI server (200-2) can request a prompt command suitable for the intent type and domain from the LLM prompt DB (310) and receive the prompt command from the LLM prompt DB (310).
[0189] The prompt command may be a command extracted from the LLM prompt DB (310). Here, the prompt command is <I'll also give you + {개체 명} + sports play's information. Please answer {현재 소속팀, 과거 소속팀, 인기 경기} using this information + {document content}> It could be.
[0190] <document content> can display the latest information about the entity name (James) extracted from the document DB (320).
[0191] The AI server (200-2) can generate a prompt (640) that combines a prompt command, an entity name (James), multiple category items, and the latest information.
[0192] The AI server (200-2) can transmit a prompt (640) to the generation AI server (300) and receive a response result (650) corresponding to the prompt (640) from the generation AI server (300).
[0193] The generation AI server (300) can obtain a response result (650) corresponding to the prompt (640) using the generation AI model, and transmit the obtained response result (650) to the AI server (200-2).
[0194] The response result (650) may contain up-to-date information about a player named James.
[0195] The AI server (200-2) can transmit the response result (650) to the NLP server (200-1) or directly to the display device (100).
[0196] Referring to FIG. 6b, the display device (100) can display a response result (650) on the display (180). The response result (650) can include information about the team that player James belongs to, which is a response to the voice command (710) uttered by the user.
[0197] The response result (650) may include the name of the current team (651), the name of the past team (653), and popular game information (655) of the player named James.
[0198] Popular match information (655) may include one or more of videos and access addresses for popular matches played by James for his current or past teams.
[0199] Conventional display devices (100) used metadata centered on TV content to one-dimensionally retrieve information about object names. Therefore, in the absence of metadata, there was a problem in finding information related to object names.
[0200] According to an embodiment of the present disclosure, there is an effect of being able to provide various and accurate information desired by a user by using a document DB (320) and a category DB (330) that have the latest information.
[0201] FIGS. 7A and 7B are diagrams illustrating an example of providing linked content by utilizing up-to-date information for a voice command uttered by a user according to another embodiment of the present disclosure.
[0202] In particular, FIGS. 7a and 7b show that the user <newj>This diagram illustrates an example of a response result provided to a sports player who uttered the word "."
[0203] In Figures 7a and 7b, the user <newj>Speak the voice command (710).
[0204] The display device (100) can transmit voice data corresponding to a voice command (710) uttered by the user to the NLP server (200-1).
[0205] The NLP server (200-1) can convert voice data into text data using a voice-to-text conversion engine, and can obtain analysis results (720) for the converted text data using a natural language processing engine.
[0206] In one embodiment, the analysis results (720) may include a user query, an action (unknown), and an entity name (NEWJ).
[0207] The NLP server (200-1) may generate an analysis result (720) including an action (unknown) indicating that it cannot respond because there is no action definition for the singer named NEWJ based on the analysis of the text data.
[0208] However, the NLP server (200-1) can transmit the entity name (NEWJ) to the AI server (200-2).
[0209] The AI server (200-2) can obtain additional analysis results (730) including the user's intention type and the domain of the entity name based on the entity name (NEWJ).
[0210] The AI server (200-2) can determine whether the user's intention type is a search type or a question response type based on intent analysis of text data.
[0211] The AI server (200-2) can analyze a user's intention type through speech act analysis. Speech act analysis may be a process of analyzing the intent of text data by considering the context and meaning of the text data.
[0212] The AI server (200-2) can extract a domain for an entity name from an entity name DB (not shown) that includes entity names and domains matching the entity names.
[0213] In Fig. 7a, the user's intent type may be a search type, and the domain of the entity name may be music.
[0214] The AI server (200-2) can obtain the latest information on the entity name (NEWJ) through the document DB (320). The AI server (200-2) can request the latest information on the entity name (NEWJ) from the document DB (320) and receive the latest information on the entity name (NEWJ) from the document DB (320).
[0215] The AI server (200-2) can transmit information about the intent type and domain to the category DB (330), and receive a plurality of category items matching the intent type and domain from the category DB (330). The plurality of category items can include a latest song item, a popular song item, and an album name item.
[0216] The category DB (330) may store multiple category items matching the intent type and domain in advance.
[0217] The AI server (200-2) can receive a prompt command from the LLM prompt DB (310).
[0218] If the latest information on the entity name (NEWJ) exists, the AI server (200-2) can request a prompt command suitable for the intent type and domain from the LLM prompt DB (310) and receive the prompt command from the LLM prompt DB (310).
[0219] The prompt command may be a command extracted from the LLM prompt DB (310). Here, the prompt command is <I'll also give you + {개체 명} + musical artist's information. Please answer {최신 곡, 인기 곡, 앨범 명칭} using this information + {document content}> It could be.
[0220] <document content> It can display the latest information on the entity name (NEWJ) extracted from the document DB (320).
[0221] The AI server (200-2) can generate a prompt (740) that combines a prompt command, an entity name (NEWJ), multiple category items, and the latest information.
[0222] The AI server (200-2) can transmit a prompt (740) to the generation AI server (300) and receive a response result (750) corresponding to the prompt (740) from the generation AI server (300).
[0223] The generation AI server (300) can obtain a response result (750) corresponding to the prompt (740) using the generation AI model, and transmit the obtained response result (750) to the AI server (200-2).
[0224] The response result (750) may include the latest information about a singer named NEWJ.
[0225] The AI server (200-2) can transmit the response result (750) to the NLP server (200-1) or directly to the display device (100).
[0226] Referring to FIG. 7b, the display device (100) can display a response result (750) on the display (180). The response result (750) can include information about the team of the NEWJ singer, which is a response to the voice command (710) uttered by the user.
[0227] The response result (750) may include a latest song tab (751), a popular song tab (753), and an album title tab (755) of a singer named NEWJ.
[0228] The latest song tab (751) may be a tab that provides one or more of a video or access address for playing the latest songs of a NEWJ singer.
[0229] When the latest song tab (751) is selected, the display device (100) can display latest song information (757) including one or more of a video and an access address for playing the latest songs of NEWJ singers.
[0230] The popular song tab (753) may be a tab for providing one or more of a video or access address for playing popular songs of NEWJ singers.
[0231] The album name tab (755) may be a tab for providing one or more of a video or access address for playing songs included in the corresponding album of the NEWJ singer.
[0232] Conventional display devices (100) used metadata centered on TV content to one-dimensionally retrieve information about object names. Therefore, in the absence of metadata, there was a problem in finding information related to object names.
[0233] According to an embodiment of the present disclosure, there is an effect of being able to provide various and accurate information desired by a user by using a document DB (320) and a category DB (330) that have the latest information.
[0234] According to one embodiment of the present disclosure, the above-described method can be implemented as processor-readable code on a medium in which a program is recorded. Examples of processor-readable media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices.
[0235] The display device described above is not limited to the configuration and method of the embodiments described above, and the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.< / newj> < / newj> < / james> < / james>
Claims
1. In artificial intelligence devices, A communication unit for communicating with an electronic device and one or more artificial intelligence (AI) servers; and A processor comprising: a processor for obtaining an analysis result including an entity name based on text data corresponding to a voice command uttered by a user through the electronic device; obtaining the latest information on the entity name from a document database; generating a prompt based on the analysis result and the latest information; transmitting the generated prompt to the generated AI server; receiving a response result corresponding to the prompt from the generated AI server; and transmitting the response result to the electronic device. Artificial intelligence device.
2. In paragraph 1, The results of the above analysis are Include the user's intent type and domain, The above intent type is either a question-response type or a search type, The above domain is Indicates any of the genres or fields related to the above entity name. Artificial intelligence device.
3. In paragraph 2, The above processor Transmitting the intent type and the domain to the prompt database, and receiving a prompt command from the prompt database; Generating the prompt based on the above text data, the above prompt command and the above latest information. Artificial intelligence device.
4. In paragraph 2, The above processor Transmitting the intent type and the domain to the prompt database, and receiving a prompt command from the prompt database; Transmitting the intent type and the domain to a category database, and including a plurality of category items from the category database, Generating the prompt based on the above entity name, the above prompt command, the above multiple category items, and the above latest information. Artificial intelligence device.
5. In paragraph 4, The above response results are Access address corresponding to each of the above multiple category items, including at least one video Artificial intelligence device.
6. In paragraph 5, If the above entity name is the name of a sports player and the above domain is sports, the above multiple category items include the current team items, past team items, and popular game items of the sports player. Artificial intelligence device.
7. In paragraph 5, If the above entity name is the singer's name and the above domain is music, the above multiple category items include the singer's latest song items, popular song items, and album name items. Artificial intelligence device.
8. In the method of operating an artificial intelligence device, A step of obtaining an analysis result including an entity name based on text data corresponding to a voice command uttered by a user through an electronic device; A step of obtaining the latest information on the above entity name from a document database; A step of generating a prompt based on the analysis results and the latest information; Step of transmitting the generated AI to the generated server; A step of receiving a response result corresponding to the prompt from the above-mentioned generating AI server; and comprising a step of transmitting the response result to the electronic device; Method of operation of an artificial intelligence device.
9. In paragraph 1, The results of the above analysis are Include the user's intent type and domain, The above intent type is either a question-response type or a search type, The above domain is Indicates any of the genres or fields related to the above entity name. Method of operation of an artificial intelligence device.
10. In paragraph 9, The above processor Transmitting the intent type and the domain to the prompt database, and receiving a prompt command from the prompt database; Generating the prompt based on the above text data, the above prompt command and the above latest information. Method of operation of an artificial intelligence device.
11. In paragraph 9, The above processor Transmitting the intent type and the domain to the prompt database, and receiving a prompt command from the prompt database; Transmitting the intent type and the domain to a category database, and including a plurality of category items from the category database, Generating the prompt based on the above entity name, the above prompt command, the above multiple category items, and the above latest information. Method of operation of an artificial intelligence device.
12. In paragraph 11, The above response results are Access address corresponding to each of the above multiple category items, including at least one video Method of operation of an artificial intelligence device.
13. In paragraph 12, If the above entity name is the name of a sports player and the above domain is sports, the above multiple category items include the current team items, past team items, and popular game items of the sports player. Method of operation of an artificial intelligence device.
14. In paragraph 12, If the above entity name is the singer's name and the above domain is music, the above multiple category items include the singer's latest song items, popular song items, and album name items. Method of operation of an artificial intelligence device.
Citation Information
Patent Citations
Information processing device, information processing method, and information processing program
JP2022039757A
Polyester hollow fiber with excellent sound absorption
KR1020210138411A
Electrohydraulic pressure-regulating valve
KR1020220115883A
Method for generating and utilizing a training dataset for deep learning based generative ai system using super-large ai
KR102570178B1
KR20210097347A