Artificial intelligence device and operation method thereof

By combining the responses of display device manufacturers and external generative AI servers, the problem of mismatch between speech recognition results and user intent was solved, resulting in a better user experience and diversified services.

CN121816612APending Publication Date: 2026-04-07LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, speech recognition services only display speech recognition results from the most suitable AI server, which may not match the user's speech intent, resulting in a poor user experience.

Method used

By combining responses from servers operated by the display device manufacturer and external generative AI servers, information matching the user's verbal intent is generated, and various services are provided by leveraging the display device manufacturer's NLP services and external generative AI server services.

Benefits of technology

It improved the user experience by providing responses that match the user's verbal intent, thereby enhancing the diversity and accuracy of the service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121816612A_ABST
    Figure CN121816612A_ABST
Patent Text Reader

Abstract

An artificial intelligence device according to an embodiment of the present disclosure may comprise: a communication unit for communicating with an electronic device and one or more generative artificial intelligence (AI) servers; and a processor to: receive a voice command issued by a user from an electronic device; obtaining an analysis result of the received voice command; when it is determined that a response from one or more generative AI servers is required based on the obtained analysis result, generating a prompt based on the analysis result; sending the generated hint to one or more generative AI servers; receiving a response result corresponding to the prompt from the one or more generative AI servers; and transmitting the analysis result and the response result to the electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to artificial intelligence devices, and more specifically, to artificial intelligence devices capable of generating images that match a user's speech. Background Technology

[0002] Digital TV services using wired or wireless communication networks are becoming increasingly common. Digital TV services can offer a variety of services that are not available with existing analog broadcasting services.

[0003] For example, IPTV (Internet Protocol Television) and Smart TV services (which are types of digital TV services) offer interactivity that allows users to actively choose the type of programs they watch, the viewing time, and so on. IPTV and Smart TV services can also leverage this interactivity to offer various additional services, such as internet search, home shopping, and online gaming.

[0004] Recently, TVs have been providing voice recognition services by analyzing the voice commands issued by users.

[0005] Typically, during speech recognition, only the speech recognition results provided from the most suitable AI server are displayed.

[0006] However, there is a problem that the speech recognition results may not match the user's intended speech. Summary of the Invention

[0007] Technical issues

[0008] The purpose of this disclosure is to provide both the response results of a server operated by the manufacturer of the display device and the response results of an external AI server when providing information that matches the user's verbal intent.

[0009] The purpose of this disclosure is to provide various services to users by combining response results from an external generative AI server to voice commands issued by the user.

[0010] Technical solution

[0011] An artificial intelligence device according to embodiments of the present disclosure includes: a communication unit that communicates with an electronic device and one or more generative artificial intelligence (AI) servers; and a processor that receives voice commands issued by a user from the electronic device; obtains analysis results of the received voice commands; generates prompts based on the analysis results when it is determined that a response from one or more generative AI servers is required; sends the generated prompts to one or more generative AI servers; receives response results corresponding to the prompts from one or more generative AI servers; and sends the analysis results and response results to the electronic device.

[0012] A method for operating an artificial intelligence device according to embodiments of the present disclosure includes: receiving a voice command issued by a user from an electronic device; obtaining an analysis result of the received voice command; generating a prompt based on the analysis result when it is determined that a response from one or more generative AI servers is required based on the obtained analysis result; sending the generated prompt to one or more generative AI servers; receiving a response result corresponding to the prompt from one or more generative AI servers; and sending the analysis result and the response result to the electronic device.

[0013] Technical effect

[0014] According to embodiments of this disclosure, the meaning of the user's words can be determined. Figure 1 This involves utilizing the NLP services of display device manufacturers and external server services that provide information received from external generative AI. Therefore, by providing various response results that match the user's verbal intent, the user experience can be improved. Attached Figure Description

[0015] Figure 1 This is a block diagram illustrating the configuration of a display device according to an embodiment of the present invention.

[0016] Figure 2 This is a block diagram of a remote control device according to an embodiment of the present invention.

[0017] Figure 3 An example of a practical configuration of a remote control device according to an embodiment of the present invention is shown.

[0018] Figure 4 An example of using a remote control device according to an embodiment of the present invention is shown.

[0019] Figure 5 An artificial intelligence (AI) server according to an embodiment of the present disclosure is shown.

[0020] Figure 6 It is a ladder diagram used to illustrate the operation method of an AI system according to an embodiment of the present disclosure.

[0021] Figure 7a This diagram illustrates the process by which an AI server provides analysis results based on user voice commands, according to existing technologies.

[0022] Figures 7b to 7d This is a diagram illustrating the process of providing analysis results of an AI server and a first response result of a first generative AI server based on user voice commands according to an embodiment of the present disclosure.

[0023] Figure 8This is a ladder diagram used to illustrate the operation method of an AI system according to another embodiment of the present disclosure.

[0024] Figure 9 This is a diagram illustrating the process of providing analysis results from an AI server based on a user's voice command, and providing first and second response results from a first generative AI server and a second generative AI server, respectively, according to embodiments of the present disclosure.

[0025] Figure 10 This is a ladder diagram used to illustrate the operation method of an AI system according to another embodiment of the present disclosure.

[0026] Figures 11a to 11c This is a diagram illustrating the process of providing the response result of another generative AI when the response result of another generative AI is required, according to an embodiment of the present disclosure. Detailed Implementation

[0027] In the following, embodiments relating to this disclosure will be described in detail with reference to the accompanying drawings. The suffixes “module” and “unit” used for components in the following description are assigned or combined for ease of writing and do not have a unique meaning or function in themselves.

[0028] The display device according to embodiments of this disclosure is an intelligent display device that adds computer-aided functions, such as broadcast reception, and while maintaining fidelity to the broadcast reception function, adds internet functionality, etc., and may have a more convenient interface, such as a manual input device, a touch screen, or a spatial remote control. Furthermore, when supporting wired or wireless internet functionality, it can connect to the internet and a computer and perform functions such as email, web browsing, banking, or gaming. A standardized, general-purpose operating system can be used for these various functions.

[0029] Therefore, the display device described in this disclosure can perform a variety of user-friendly functions, for example, because various applications can be freely added or removed on a general-purpose OS kernel. More specifically, the display device can be, for example, a network TV, HBBTV, smart TV, LED TV, OLED TV, etc., and in some cases, it can also be applied to smartphones.

[0030] Figure 1 This is a block diagram illustrating the configuration of a display device according to an embodiment of the present disclosure.

[0031] Reference Figure 1 The display device 100 may include a broadcast receiver 130, an external device interface 135, a memory 140, a user input interface 150, a controller 170, a wireless communication interface 173, a display 180, a speaker 185, and a power supply circuit 190.

[0032] The broadcast receiver section 130 may include a tuner 131, a demodulator 132, and a network interface 133.

[0033] Tuner 131 can select a specific broadcast channel according to a channel selection command. Tuner 131 can receive broadcast signals for the selected specific broadcast channel.

[0034] Demodulator 132 can divide the received broadcast signal into image signal, audio signal and broadcast program related data signal, and can restore the divided image signal, audio signal and data signal into an output usable form.

[0035] The external device interface 135 can receive applications or application lists from adjacent external devices and deliver the applications or application lists to the controller 170 or the memory 140.

[0036] External device interface 135 provides a connection path between display device 100 and external devices. External device interface 135 can receive at least one of image or audio output from an external device wirelessly or wiredly connected to display device 100, and deliver the received image or audio to a controller. External device interface 135 may include multiple external input terminals. The multiple external input terminals may include RGB terminals, at least one High Definition Multimedia Interface (HDMI) terminal, and component terminals.

[0037] Image signals from external devices input through external device interface 135 can be output through display 180. Audio signals from external devices input through external device interface 135 can be output through speaker 185.

[0038] The external device that can be connected to the external device interface 135 can be one of a set-top box, Blu-ray player, DVD player, game console, soundbar, smartphone, PC, USB storage device, and home theater system, but these are merely examples.

[0039] Network interface 133 can provide an interface for connecting the display device 100 to a wired / wireless network, including the Internet. Network interface 133 can send data to or receive data from another user or another electronic device by accessing the network or by linking to another network accessing the network.

[0040] Additionally, some content data stored in the display device 100 can be sent to a user or electronic device selected from other users or other electronic devices pre-registered in the display device 100.

[0041] Network interface 133 can access a predetermined webpage by connecting to a network or linking to another network connected to the network. In other words, network interface 133 can send data to or receive data from a corresponding server by accessing the predetermined webpage via the network.

[0042] Network interface 133 can receive content or data provided by content providers or network operators. In other words, network interface 133 can receive content (such as movies, advertisements, games, VOD and broadcast signals) and related information provided by content providers or network operators via the network.

[0043] In addition, network interface 133 can receive firmware update information and update files provided by network operators, and can send data to the Internet or content providers or network operators.

[0044] Network interface 133 can select and receive desired applications from applications open to the air via the network.

[0045] The memory 140 can store processed image, voice, or data signals stored by a program for use in each signal processing and control in the controller 170.

[0046] Additionally, the memory 140 can perform functions for temporarily storing image, voice, or data signals output from the external device interface 135 or the network interface 133, and can store information about a predetermined image through the channel memory function.

[0047] The memory 140 can store applications or application lists input from the external device interface 135 or the network interface 133.

[0048] The display device 100 can play content files (e.g., video files, still image files, music files, document files, application files, etc.) stored in the memory 140, and can provide content files to the user.

[0049] User input interface 150 can send signals input by the user to controller 170, or send signals from controller 170 to the user. For example, user input interface 150 can receive or process control signals such as power on / off, channel selection, and screen settings from remote control device 200, or send control signals from controller 170 to remote control device 200, according to various communication methods such as Bluetooth, ultra-wideband (WB), ZigBee, radio frequency (RF), and IR communication methods.

[0050] Additionally, the user input interface 150 can send control signals to the controller 170 from local keys (not shown), such as the power button, channel key, volume key, and settings key.

[0051] The image signal processed by the controller 170 can be input to the display 180 and displayed as an image corresponding to the image signal. Furthermore, the image signal processed by the controller 170 can be input to an external output device via the external device interface 135.

[0052] The voice signal processed by the controller 170 can be output to the speaker 185. In addition, the voice signal processed by the controller 170 can be input to an external output device through the external device interface 135.

[0053] In addition, the controller 170 can control the overall operation of the display device 100.

[0054] Additionally, the controller 170 can control the display device 100 via user commands or internal programs input through the user input interface 150, and can access the network to download desired applications or application lists to the display device 100.

[0055] The controller 170 can output user-selected channel information and processed image or voice signals via a display 180 or a speaker 185.

[0056] In addition, the controller 170 can output image or audio signals from external devices (such as cameras or video cameras) input through the external device interface 135, the display 180, or the speaker 185, based on the external device image playback command received through the user input interface 150.

[0057] Furthermore, the controller 170 can control the display 180 to display images, including broadcast images input via the tuner 131, externally input images input via the external device interface 135, images input via the network interface, or images stored in the memory 140. In this case, the image displayed on the display 180 can be a still image or a video, and can also be a 2D image or a 3D image.

[0058] Additionally, the controller 170 can play content stored in the display device 100, received broadcast content, and externally input content, and the content can be in various formats, such as broadcast images, externally input images, audio files, still images, accessed web screens, and document files.

[0059] Furthermore, the wireless communication interface 173 can perform wired or wireless communication with external devices. The wireless communication interface 173 can perform short-range communication with external devices. For this purpose, the wireless communication interface 173 can utilize Bluetooth. TMThe wireless communication interface 173 can support short-range communication via at least one of the following technologies: Bluetooth Low Energy (BLE), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wideband (UWB), ZigBee, Near Field Communication (NFC), Wi-Fi, Wi-Fi Direct, and Wireless Universal Serial Bus (USB). The wireless communication interface 173 can support wireless communication between the display device 100 and a wireless communication system, between the display device 100 and another display device 100, or between a network including the display device 100 and another display device 100 (or an external server) via a wireless local area network (WLAN). The WLAN can be a wireless personal area network (WPAN).

[0060] In this document, another display device 100 may be a mobile terminal capable of exchanging (or communicating) data with the display device 100, such as a wearable device (e.g., a smartwatch, smart glasses, and head-mounted display (HMD)) or a smartphone. The wireless communication interface 173 may detect (or identify) wearable devices capable of communication around the display device 100.

[0061] Furthermore, if the detected wearable device is certified to communicate with the display device 100, the controller 170 can transmit at least a portion of the data processed in the display device 100 to the wearable device via the wireless communication interface 173. Therefore, the user of the wearable device can use the data processed by the display device 100 through the wearable device.

[0062] The display 180 can convert image signals, data signals, or on-screen display (OSD) signals processed in the controller 170, or image signals or data signals received in the external device interface 135, into R, G, and B signals to generate drive signals.

[0063] also, Figure 1 The display device 100 shown is only one embodiment of this disclosure. Therefore, according to the specifications of the actual implemented display device 100, some of the components shown may be integrated, added, or omitted.

[0064] In other words, if needed, two or more components can be integrated into one component, or a component can be divided into two or more components. Furthermore, the functions performed by each block are used to describe embodiments of this disclosure, and their specific operations or apparatus do not limit the scope of this disclosure.

[0065] According to another embodiment of this disclosure, with Figure 1 Unlike other devices, the display device 100 can receive images via network interface 133 or external device interface 135 and play them without the tuner 131 and demodulator 132.

[0066] For example, the display device 100 can be divided into an image processing device (such as a set-top box) for receiving broadcast signals or content according to various network services and a content playback device for playing content input from the image processing device.

[0067] In this case, the operation method of the display device according to the embodiments of the present disclosure described below can be obtained by referring to... Figure 1 The described display device, image processing device such as a split set-top box, and content playback device including a display 180 and a speaker 185 are used to perform this function.

[0068] Reference Figure 2 and Figure 3 A remote control device according to an embodiment of the present disclosure is described.

[0069] Figure 2 This is a block diagram illustrating a remote control device according to an embodiment of the present disclosure, and Figure 3 This is an illustration showing the actual configuration of a remote control device according to an embodiment of the present disclosure.

[0070] First, refer to Figure 2 The remote control device 200 may include a fingerprint recognition device 210, a wireless communication interface 220, a user input interface 230, a sensor 240, an output interface 250, a power supply circuit 260, a memory 270, a controller 280, and a microphone 290.

[0071] Reference Figure 2 The wireless communication interface 220 sends / receives signals to / from any of the display devices according to the above embodiments of the present disclosure.

[0072] The remote control device 200 may include a radio frequency (RF) circuit 221 capable of transmitting or receiving signals to or from the display device 100 according to an RF communication standard, and an IR circuit 223 capable of transmitting or receiving signals to or from the display device 100 according to an IR communication standard. Furthermore, the remote control device 200 may include a Bluetooth circuit 225 capable of transmitting or receiving signals to or from the display device 100 according to a Bluetooth communication standard. Additionally, the remote control device 200 may include an NFC circuit 227 capable of transmitting or receiving signals to or from the display device 100 according to an NFC communication standard, and a wireless LAN (WLAN) circuit 229 capable of transmitting or receiving signals to or from the display device 100 according to a WLAN communication standard.

[0073] In addition, the remote control device 200 can send a signal containing information about the movement of the remote control device 200 to the display device 100 via the wireless communication interface 220.

[0074] In addition, the remote control device 200 can receive signals sent from the display device 100 via the RF circuit 221, and if necessary, can send commands for power-on / power-off, channel change and volume change to the display device 100 via the IR circuit 223.

[0075] User input interface 230 may be configured with a keyboard, buttons, a touchpad, or a touchscreen. A user can operate user input interface 230 to input commands related to display device 100 to remote control device 200. If user input interface 230 includes hard buttons, the user can input commands related to display device 100 to remote control device 200 by pressing the hard buttons. This will refer to... Figure 3 Describe it.

[0076] Reference Figure 3 The remote control device 200 may include multiple buttons. These buttons may include a fingerprint recognition button 212, a power button 231, a home button 232, a live button 233, an external input button 234, a volume control button 235, a voice recognition button 236, a channel change button 237, an OK button 238, and a back button 239.

[0077] The fingerprint recognition button 212 can be a button for recognizing a user's fingerprint. According to embodiments of this disclosure, the fingerprint recognition button 212 can perform a pressing operation and receive a pressing operation and a fingerprint recognition operation.

[0078] The power button 231 can be a button used to turn the power of the display device 100 on / off.

[0079] The home button 232 can be a button used to move to the home screen of the display device 100.

[0080] The Live button 233 can be a button used to display live broadcast programs.

[0081] External input button 234 can be a button for receiving external input connected to display device 100.

[0082] The volume control button 235 can be a button used to control the volume output from the display device 100.

[0083] The voice recognition button 236 can be a button used to receive the user's voice and recognize the received voice.

[0084] The channel change button 237 can be a button used to receive broadcast signals from a specific broadcast channel.

[0085] OK button 238 can be a button for selecting a specific function, and back button 239 can be a button for returning to the previous screen.

[0086] Describe again Figure 2 .

[0087] If the user input interface 230 includes a touchscreen, the user can touch the soft keys on the touchscreen to input commands related to the display device 100 to the remote control device 200. Additionally, the user input interface 230 may include various user-operable input interfaces, such as scroll keys and microswitches, and this implementation does not limit the scope of this disclosure.

[0088] Sensor 240 may include gyroscope sensor 241 or accelerometer sensor 243. Gyroscope sensor 241 can sense information about the movement of remote control device 200.

[0089] For example, the gyroscope sensor 241 can sense information about the operation of the remote control device 200 based on the x-axis, y-axis, and z-axis, and the accelerometer sensor 243 can sense information about the moving speed of the remote control device 200. Furthermore, the remote control device 200 may also include a distance measurement sensor that senses the distance relative to the display 180 of the display device 100.

[0090] The output interface 250 can output image or voice signals in response to the operation of the user input interface 230, or it can output image or voice signals corresponding to signals sent from the display device 100.

[0091] Users can identify whether the user input interface 230 is being operated or whether the display device 100 is being controlled through the output interface 250.

[0092] For example, if the user input interface 230 is manipulated or sends / receives signals to / from the display device 100 via the wireless communication interface 220, the output interface 250 may include an LED 251 for blinking, a vibrator 253 for generating vibrations, a speaker 255 for outputting sound, or a display 257 for outputting images.

[0093] In addition, the power supply circuit 260 supplies power to the remote control device 200, and stops supplying power if the remote control device 200 does not move within a predetermined time, thereby reducing power waste.

[0094] If a predetermined key provided on the remote control 200 is pressed, the power circuit 260 can restore power supply.

[0095] The memory 270 can store various programs and application data required to control or operate the remote control device 200.

[0096] If the remote control device 200 wirelessly transmits / receives signals through the display device 100 and the RF circuit 221, then the remote control device 200 and the display device 100 transmit / receive signals through a predetermined frequency band.

[0097] The controller 280 of the remote control device 200 can store and refer to information in the memory 270 about the frequency band used to send / receive signals to / from the display device 100 paired with the remote control device 200.

[0098] The controller 280 controls general matters related to the control of the remote control device 200. The controller 280 can transmit signals corresponding to predetermined key operations of the user input interface 230 or signals corresponding to movement of the remote control device 200 sensed by the sensor 240 to the display device 100 via the wireless communication interface 220.

[0099] In addition, the microphone 290 of the remote control device 200 can acquire voice.

[0100] Multiple microphones (290) can be provided.

[0101] Next, the description Figure 4 .

[0102] Figure 4 This is an illustration showing an example of a remote control device utilizing an embodiment of the present disclosure.

[0103] Figure 4 (a) shows that a pointer 205 corresponding to the remote control 200 is displayed on the display 180.

[0104] The user can move or rotate the remote control device 200 vertically or horizontally. The pointer 205 displayed on the display 180 of the display device 100 corresponds to the movement of the remote control device 200. Since the pointer 205 is moved and displayed according to the movement in 3D space as shown in the figure, the remote control device 200 can be referred to as a spatial remote control device.

[0105] Figure 4 (b) shows that if the user moves the remote control 200, the pointer 205 displayed on the display 180 of the display device 100 moves to the left in accordance with the movement of the remote control 200.

[0106] Information regarding the movement of the remote control device 200, detected by the sensors of the remote control device 200, is sent to the display device 100. The display device 100 can calculate the coordinates of the pointer 205 based on the information regarding the movement of the remote control device 200. The display device 100 can display the pointer 205 to match the calculated coordinates.

[0107] Figure 4(c) shows that when a specific button in the remote control 200 is pressed, the user removes the remote control 200 from the display 180. As a result, the selected area in the display 180 corresponding to the pointer 205 can be magnified and displayed at an enlarged size.

[0108] On the other hand, if the user moves the remote control 200 closer to the display 180, the selection area corresponding to the pointer 205 on the display 180 can be shrunk and displayed at a reduced size.

[0109] On the other hand, if the remote control 200 is moved away from the display 180, the selection area can be reduced, while if the remote control 200 is moved closer to the display 180, the selection area can be enlarged.

[0110] Additionally, pressing a specific button on the remote control 200 can exclude the recognition of vertical or horizontal movement. In other words, if the remote control 200 moves away from or closer to the display 180, up, down, left, or right movement can be ignored, and only forward and backward movement can be recognized. When the specific button on the remote control 200 is not pressed, only the pointer 205 moves according to the up, down, left, or right movement of the remote control 200.

[0111] Furthermore, the moving speed or direction of the pointer 205 can correspond to the moving speed or direction of the remote control device 200.

[0112] Furthermore, in this specification, "pointer" refers to an object displayed on the display 180 in response to the operation of the remote control device 200. Therefore, various forms of objects are possible besides the arrow form shown as pointer 205 in the diagram. For example, the aforementioned concepts include points, cursors, prompts, and rough outlines. Pointer 205 can then be displayed corresponding to a point on the horizontal and vertical axes of the display 180, and can also be displayed corresponding to multiple points such as lines and surfaces.

[0113] Figure 5 An artificial intelligence (AI) server according to an embodiment of the present disclosure is shown.

[0114] Reference Figure 5 AI Server 500 can refer to a device that uses machine learning algorithms to train artificial neural networks or uses trained artificial neural networks.

[0115] The AI ​​Server 500 can consist of multiple servers to perform distributed processing and can be defined as a 5G network.

[0116] AI server 500 may be included as part of the configuration of AI device 100 to perform at least a portion of AI processing together.

[0117] AI server 500 may include communication unit 510, memory 530, learning processor 540 and processor 560.

[0118] The communication unit 510 can send data to and receive data from an external device such as the display device 100.

[0119] The memory 530 may include a model storage unit 531.

[0120] The model storage unit 531 can store a model being trained by the learning processor 540 or a trained model (or an artificial neural network 531a).

[0121] The learning processor 540 can use training data to train the artificial neural network 531a. The learning model can be used when installed on the AI ​​server 500 of the artificial neural network, or it can be installed and used on an external device such as the display device 100.

[0122] The learning model can be implemented in hardware, software, or a combination of hardware and software. When the learning model is partially or wholly implemented in software, one or more instructions constituting the learning model can be stored in memory 530.

[0123] The processor 560 can use a learning model to infer the result value for new input data and generate a response or control command based on the inferred result value.

[0124] Figure 6 It is a ladder diagram used to illustrate the operation method of an AI system according to an embodiment of the present disclosure.

[0125] Reference Figure 6 The AI ​​system may include a display device 100, an AI server 500, and a first generative AI server 610.

[0126] AI Server 500 can be a Natural Language Processing (NLP) server.

[0127] The controller 170 of the display device 100 can receive voice commands (S601).

[0128] In one implementation, the controller 170 may receive voice commands issued by a user via a microphone (not shown).

[0129] The controller 170 can receive voice commands from the remote control 200. The remote control 200 can receive voice commands issued by the user and send them to the display device 100.

[0130] The controller 170 of the display device 100 can send voice data corresponding to the voice command to the AI ​​server 500 via the network interface 133 (S603).

[0131] The processor 560 of the AI ​​server 500 can obtain the analysis results of the voice command based on the received voice data (S605).

[0132] In one implementation, processor 560 can use a speech-to-text (STT) engine stored in memory 530 to convert speech data into text data.

[0133] In another embodiment, processor 560 can send voice data to an STT server and receive converted text data from the STT server.

[0134] The processor 560 can use a natural language processing (NLP) engine to obtain analysis results from the converted text data.

[0135] The analysis results can include intent analysis results regarding the intent of the voice command and search results based on the intent analysis results.

[0136] For example, if the voice command is "Tell me the plot of episode A", the intent analysis result is a request to tell you the plot of episode A, and the search results can be search content about the plot of episode A.

[0137] The processor 560 of the AI ​​server 500 can determine whether a response from generative AI is needed based on the analysis results of the voice command (S607).

[0138] In one implementation, the processor 560 can determine whether a response from generative AI is needed based on the intent analysis results of the voice command.

[0139] When the intent analysis result of a voice command is an intent to control the function of the display device 100, the processor 560 can determine that a response from generative AI is not required.

[0140] When the intent analysis result of the voice command is not an intent to control the function of the display device 100, the processor 560 can determine that a response from generative AI is needed.

[0141] When the intent analysis result of a voice command is not an intent for controlling the function of the display device 100, it can be an intent for content search.

[0142] For example, when the intent of a voice command is to control the function of display device 100 (such as screen brightness or volume control), processor 560 can determine that a response from generative AI is not needed. This is because display device 100 can handle it on its own without the help of generative AI.

[0143] When the intent of a voice command is a content search on display device 100 or a travel itinerary search for a specific area, processor 560 can determine that a response from generative AI is needed. This is because display device 100 has difficulty resolving this on its own without the assistance of generative AI.

[0144] When the processor 560 of the AI ​​server 500 determines that a generative AI response is not required, it can send the analysis results to the display device 100 via the communication unit 510 (S609).

[0145] The processor 560 can send the intent analysis results and the search results corresponding to the intent analysis results to the display device 100 via the communication unit 510.

[0146] The processor 560 can send only the search results corresponding to the intent analysis results to the display device 100 via the communication unit 510.

[0147] The controller 170 of the display device 100 can output the received analysis results (S611).

[0148] The controller 170 can display the search results included in the analysis results on the display 180.

[0149] When the processor 560 of the AI ​​server 500 determines that a response from the generative AI is needed, it can generate a prompt for the response from the generative AI (S613).

[0150] The processor 560 can generate prompts for obtaining response results from generative AI models provided by the generative AI server.

[0151] Generative AI models can be models that generate content using deep learning or machine learning in response to commands.

[0152] Prompts can be text commands entered into the generative AI model.

[0153] Hints can be commands used to enable AI models for natural language generation to produce natural language.

[0154] The processor 560 can generate prompts based on the intent analysis results of voice commands.

[0155] The processor 560 can generate prompts for requested content by combining keywords included in the intent analysis results of voice commands.

[0156] Furthermore, when the processor 560 determines that a response from generative AI is needed, it can determine the display format for displaying the analysis results and the first response results.

[0157] The display format can be one or more of the format that indicates the location, size, and layout of the area used to display the analysis and response results of the AI ​​server 500.

[0158] The processor 560 can determine any one of a plurality of display formats based on the type of analysis results and the type of response results, and can send information about the determined display format to the display device 100.

[0159] Here, "type" refers to the type of content, which can be any of text, images, video, and audio.

[0160] The processor 560 of the AI ​​server 500 can send a prompt to the first generative AI server 610 via the communication unit 510 (S615), and can receive a first response result from the first generative AI server 610 (S617).

[0161] In one implementation, the first generative AI server 610 may be a server that provides search results for content.

[0162] In another embodiment, the first generative AI server 610 may be a server equipped with a natural language generation AI model that outputs natural language.

[0163] Natural language generation AI models can generate an initial response in response to received prompts.

[0164] Natural language generation AI models can be models trained using deep learning architectures such as recurrent neural networks (RNNs) or transformation models that use attention mechanisms on large-scale text data.

[0165] Data can be collected from the web or various text sources to train natural language generation AI models.

[0166] Natural language generation AI models can generate one or more sentences containing content corresponding to intent analysis results by taking received prompts as input.

[0167] The first generative AI server 610 can obtain one or more generated sentences as the first response result and send the obtained first response result to the AI ​​server 500.

[0168] The processor 560 of the AI ​​server 500 can send the analysis results and first response results of the voice command to the display device 100 via the communication unit 510 (S619).

[0169] In one embodiment, the processor 560 can sequentially send the analysis results and the first response results to the display device 100 via the communication unit 510.

[0170] In other words, processor 560 can first send the analysis results to display device 100, and then send the first response result to display device 100. This is because the response time of the first generative AI server 610 may be required.

[0171] In another embodiment, the processor 560 can simultaneously send the analysis results and the first response results to the display device 100 via the communication unit 510.

[0172] Furthermore, the processor 560 can filter some information from the analysis results and the first response results, and send the filtered information to the display device 100. The filtered information may include information of a sexual nature or information such as profanity.

[0173] The controller 170 of the display device 100 can output the received analysis results and the first response results (S621).

[0174] The controller 170 can display the analysis results and the first response results on the display 180.

[0175] The controller 170 can display the analysis results and the first response results on the display 180 according to the display format received from the AI ​​server 500.

[0176] Figure 7a This diagram illustrates the process by which an AI server provides analysis results based on user voice commands, according to existing technologies. Figures 7b to 7d This is a diagram illustrating the process of providing analysis results of an AI server and a first response result of a first generative AI server based on user voice commands according to an embodiment of the present disclosure.

[0177] Reference Figure 7a The display device 100 is playing video content 700.

[0178] Users can press the microphone button on the remote control 200 and say the voice command "Tell me the plot of A".

[0179] The remote control device 200 can send received voice commands to the display device 100.

[0180] The display device 100 can send voice data corresponding to the voice command to the AI ​​server 500, and receive a first plot search result 710 from the AI ​​server 500 for the plot of content A included in the analysis result of the voice command.

[0181] The display device 100 can display the first plot search result 710 on the display 180.

[0182] The display device 100 can display the first plot search result 710 by overlaying the first plot search result 710 onto the content video 700.

[0183] Reference Figure 7b The display device 100 is playing video content 700.

[0184] Users can press the microphone button on the remote control 200 and say the voice command "Tell me the plot of A".

[0185] The remote control device 200 can send received voice commands to the display device 100.

[0186] The display device 100 can send voice data corresponding to the voice command to the AI ​​server 500, and receive a first plot search result 710 from the AI ​​server 500 for the plot of content A included in the analysis result of the voice command.

[0187] When AI server 500 determines that the intent analysis results of a voice command are relevant to a content search, it can generate a prompt requesting a plot search for content with name A.

[0188] AI server 500 can send a request to the first generative AI server 610 for hints on a plot search for content with name A.

[0189] The first generative AI server 610 can input the received prompts into a natural language generation AI model to generate a second plot search result that includes the plot of content A.

[0190] The first generative AI server 610 can send the second plot search result to the AI ​​server 500, and the AI ​​server 500 can send the second plot search result 730 to the display device 100.

[0191] The display device 100 can display the first plot search result 710 and the second plot search result 730 on the display 180.

[0192] The display device 100 can display the plot search results 710 by overlaying the plot search results 710 onto the content video 700.

[0193] The display device 100 can display the first plot search result 710 and the second plot search result 730 on the display 180 according to the display format received from the AI ​​server 500. Here, the display format may be that the first plot search result 710 is displayed at the bottom of the display screen and the second plot search result 730 is displayed at the top of the display screen.

[0194] In this way, according to the embodiments of this disclosure, the NLP services of the display device 100 manufacturer and the external server services that provide information received from external generative AI can be utilized based on the user's verbal intent.

[0195] Providing a variety of response results that match the user's verbal intent can improve the user experience.

[0196] Figure 7c This is an example of a user saying the voice command "<Recommend a comedy movie>" on the home screen or content provider screen.

[0197] The display device 100 can display on the screen 750 of the display 180 a first search result 751 for a comedy movie received from the AI ​​server 500 and a second search result 753 for a comedy movie received from the first generative AI server 610.

[0198] The AI ​​server 500 can generate prompts by reflecting the intent analysis results of voice commands (comedy movie recommendations) and user preference information.

[0199] User preference information may include one or more of their preferred genres, preferred actors, and preferred directors. This user preference information can be collected in advance by sending it to the AI ​​server 500 via the display device 100.

[0200] For example, suggestions could include commands to search for comedy movies with a sports theme.

[0201] AI server 500 can send the generated prompts to first generative AI server 610, and first generative AI server 610 can receive second search results 753 for comedy movies with sports themes.

[0202] Each of the first search result 751 and the second search result 753 may include thumbnails of one or more comedy movies.

[0203] The screen 750 may also include a feedback phrase 755 for the second search result 753. The AI ​​server 500 may send the second search result 753 and the feedback phrase 755 explaining the search intent of the second search result 753 together to the display device 100.

[0204] In addition, the display device 100 can display the first search result 751 and the second search result 753 in a half view according to the display format received from the AI ​​server 500, and display a feedback phrase 755 at the bottom of the screen 750.

[0205] Figure 7d This is an example of a user issuing a voice command, "<Recommend a three-day tour of Rome>", on the home screen or content provider screen.

[0206] The display device 100 can display on the screen 750 of the display 180 a third search result 771 for the tourist destination of Rome received from the AI ​​server 500 and a fourth search result 773 for the tourist destination of Rome received from the first generative AI server 610.

[0207] The third search result 771 can include multiple videos (which contain information about Rome as a tourist destination) or the URL of each of the multiple videos.

[0208] AI server 500 can generate prompts based on the intent analysis results of the voice command (Rome tourist destination recommendation) and the third search result 771. Here, prompts may include commands that recommend travel itineraries using multiple videos about Rome as a tourist destination.

[0209] AI server 500 can send the generated prompts to first generative AI server 610, and can receive from first generative AI server 610 a fourth search result 773 indicating a three-day Rome tour itinerary.

[0210] The fourth search result, 773, can include 3-day travel destination itinerary information.

[0211] Screen 770 may also include a feedback phrase 775 for the fourth search result 773.

[0212] In addition, the display device 100 can display the third search result 771 and the fourth search result 773 in a horizontal full view according to the display format received from the AI ​​server 500, and display a feedback phrase 775 at the top of the screen 770.

[0213] Figure 8 This is a ladder diagram used to illustrate the operation method of an AI system according to another embodiment of the present disclosure.

[0214] Reference Figure 8 The AI ​​system may include a display device 100, an AI server 500, a first generative AI server 610, and a second generative AI server 630.

[0215] AI Server 500 can be a Natural Language Processing (NLP) server.

[0216] The controller 170 of the display device 100 can receive voice commands (S801).

[0217] In one implementation, the controller 170 may receive voice commands issued by a user via a microphone (not shown).

[0218] The controller 170 can receive voice commands from the remote control 200. The remote control 200 can receive voice commands issued by the user and send them to the display device 100.

[0219] The controller 170 of the display device 100 can send voice data corresponding to the voice command to the AI ​​server 500 via the network interface 133 (S803).

[0220] The processor 560 of the AI ​​server 500 can obtain the analysis results of the voice command based on the received voice data (S805).

[0221] In one implementation, processor 560 can use a speech-to-text (STT) engine stored in memory 530 to convert speech data into text data.

[0222] In another embodiment, processor 560 can send voice data to an STT server and receive converted text data from the STT server.

[0223] The processor 560 can use a natural language processing (NLP) engine to obtain analysis results from the converted text data.

[0224] The analysis results can include intent analysis results regarding the intent of the voice command and search results based on the intent analysis results.

[0225] The processor 560 of the AI ​​server 500 can determine whether a response from generative AI is needed based on the analysis results of the voice command (S807).

[0226] In one implementation, the processor 560 can determine whether a response from generative AI is needed based on the intent analysis results of the voice command.

[0227] When the intent analysis result of a voice command is an intent to control the function of the display device 100, the processor 560 can determine that a response from generative AI is not required.

[0228] When the intent analysis result of the voice command is not an intent to control the function of the display device 100, the processor 560 can determine that a response from generative AI is needed.

[0229] When the intent analysis result of a voice command is not an intent for controlling the function of the display device 100, it can be an intent for content search.

[0230] For example, when the intent of a voice command is to control the function of display device 100 (such as screen brightness or volume control), processor 560 can determine that a response from generative AI is not needed. This is because display device 100 can handle it on its own without the help of generative AI.

[0231] When the intent of a voice command is a content search on display device 100 or a travel itinerary search for a specific area, processor 560 can determine that a response from generative AI is needed. This is because display device 100 has difficulty resolving this on its own without the assistance of generative AI.

[0232] When the processor 560 of the AI ​​server 500 determines that a generative AI response is not required, it can send the analysis results to the display device 100 via the communication unit 510 (S809).

[0233] The processor 560 can send the intent analysis results and the search results corresponding to the intent analysis results to the display device 100 via the communication unit 510.

[0234] The processor 560 can send only the search results corresponding to the intent analysis results to the display device 100 via the communication unit 510.

[0235] The controller 170 of the display device 100 can output the received analysis results (S811).

[0236] The controller 170 can display the search results included in the analysis results on the display 180.

[0237] When the processor 560 of the AI ​​server 500 determines that a response from the generative AI is needed, it can generate a prompt for the response from the first generative AI server 610 (S813).

[0238] The processor 560 can generate prompts for the response results obtained from the generative AI model set in the first generative AI server 610.

[0239] Generative AI models can be models that generate content using deep learning or machine learning in response to commands. These commands can include one or more text or images.

[0240] The processor 560 can generate prompts based on the intent analysis results of voice commands.

[0241] The processor 560 can generate prompts for requested content by combining keywords included in the intent analysis results of voice commands.

[0242] The processor 560 of the AI ​​server 500 can send a prompt to the first generative AI server 610 via the communication unit 510 (S815), and can receive a first response result from the first generative AI server 610 (S817).

[0243] In one implementation, the first generative AI server 610 may be a server that provides search results for content.

[0244] In another embodiment, the first generative AI server 610 may be a server equipped with a natural language generation AI model that outputs natural language.

[0245] Natural language generation AI models can generate an initial response in response to received prompts.

[0246] Natural language generation AI models can be models trained using deep learning architectures such as recurrent neural networks (RNNs) or transformation models that use attention mechanisms on large-scale text data.

[0247] Data can be collected from the web or various text sources to train natural language generation AI models.

[0248] Natural language generation AI models can generate one or more sentences containing content corresponding to intent analysis results by taking received prompts as input.

[0249] The first generative AI server 610 can obtain one or more generated sentences as the first response result and send the obtained first response result to the AI ​​server 500.

[0250] The processor 560 of the AI ​​server 500 can send a prompt to the second generative AI server 630 via the communication unit 510 (S819), and can receive a second response result from the second generative AI server 630 (S821).

[0251] In other words, processor 560 can send the same prompts to the first generative AI server 610 and the second generative AI server 630.

[0252] The second generative AI server 630 can be a server that provides search results for content.

[0253] In another embodiment, the second generative AI server 630 may be a server equipped with a natural language generation AI model that outputs natural language.

[0254] The second generative AI server 630 may be a server equipped with generative AI models of the same type as the first generative AI server 610.

[0255] The second generative AI server 630 can obtain one or more sentences as a second response result to the prompt, and send the obtained second response result to the AI ​​server 500.

[0256] The processor 560 of the AI ​​server 500 can generate a third response result based on the first response result and the second response result (S823).

[0257] In one implementation, the processor 560 can generate a third response result by combining the first response result and the second response result.

[0258] In another embodiment, the processor 560 may obtain the remaining content of the first response result and the second response result after excluding duplicate parts as the third response result.

[0259] The processor 560 of the AI ​​server 500 can send the third response result to the display device 100 via the communication unit 510 (S825).

[0260] In one implementation, the processor 560 can sequentially send the analysis result and the third response result to the display device 100 via the communication unit 510. That is, the processor 560 can first send the analysis result to the display device 100, and then send the third response result to the display device 100. This is because it may require response time from the first generative AI server 610 and the second generative AI server 630.

[0261] In another embodiment, the processor 560 can simultaneously send the analysis results and the third response results to the display device 100 via the communication unit 510.

[0262] Furthermore, when the processor 560 determines that a response from generative AI is needed, it can determine the display format for displaying the analysis results and the third response results.

[0263] The display format can be a format that indicates the location, size, layout, etc. of the area used to display the analysis results and third response results of AI server 500.

[0264] The processor 560 can determine any one of a plurality of display formats based on the type of the analysis result and the type of the third response result, and can send information about the determined display format to the display device 100.

[0265] Here, "type" refers to the type of content included in the analysis results, which can be any of text, images, video, and audio.

[0266] The controller 170 of the display device 100 can output the received third response result (S827).

[0267] The controller 170 can display the analysis results and the third response results on the display 180.

[0268] The controller 170 can display the analysis results and the third response results on the display 180 according to the display format received from the AI ​​server 500.

[0269] Figure 9 This is a diagram illustrating the process of providing analysis results from an AI server based on a user's voice command, and providing first and second response results from a first generative AI server and a second generative AI server, respectively, according to embodiments of the present disclosure.

[0270] Reference Figure 9 The display device 100 is playing video content 700.

[0271] Users can press the microphone button set on the remote control device 200 and say the voice command "Tell me the plot of A".

[0272] The remote control device 200 can send received voice commands to the display device 100.

[0273] The display device 100 can send voice data corresponding to the voice command to the AI ​​server 500, and receive a first plot search result 710 from the AI ​​server 500 for the plot of content A included in the analysis result of the voice command.

[0274] When AI server 500 determines that the intent analysis results of a voice command are relevant to content search, it can generate a prompt to request a plot search for content with name A.

[0275] AI server 500 can send requests to each of the first generative AI server 610 and the second generative AI server 630 for plot search hints for content with name A.

[0276] Each of the first generative AI server 610 and the second generative AI server 630 can be a server operated by a different operating entity.

[0277] The first generative AI server 610 can input the received prompts into a natural language generation AI model to generate a second plot search result 910 that includes the plot of content A.

[0278] The first generative AI server 610 can send the second plot search result 910 to the AI ​​server 500.

[0279] The second generative AI server 630 can input the received prompts into a natural language generation AI model to generate a third plot search result 930 that includes the plot of content A.

[0280] The second generative AI server 630 can send the third plot search result 930 to the AI ​​server 500.

[0281] AI server 500 can send the second plot search result 910 and the third plot search result 930 to display device 100.

[0282] The display device 100 can display the first plot search result 710, the second plot search result 910, and the third plot search result 930 on the display 180.

[0283] Generative AI search results 900, including second plot search results 910 and third plot search results 930, can be displayed above the first plot search results 710 according to the display format.

[0284] Therefore, according to embodiments of this disclosure, the NLP services of the display device 100 manufacturer and the external server services that provide information received from multiple external generative AI servers can also be utilized based on the user's verbal intent.

[0285] Figure 10 This is a ladder diagram used to illustrate the operation method of an AI system according to another embodiment of the present disclosure.

[0286] Reference Figure 10 The AI ​​system may include a display device 100, an AI server 500, a first generative AI server 610, and a second generative AI server 630.

[0287] AI Server 500 can be a Natural Language Processing (NLP) server.

[0288] The controller 170 of the display device 100 can receive voice commands (S1001).

[0289] In one implementation, the controller 170 may receive voice commands issued by a user via a microphone (not shown).

[0290] The controller 170 can receive voice commands from the remote control 200. The remote control 200 can receive voice commands issued by the user and send them to the display device 100.

[0291] The controller 170 of the display device 100 can send voice data corresponding to the voice command to the AI ​​server 500 via the network interface 133 (S1003).

[0292] The processor 560 of the AI ​​server 500 can obtain the analysis results of the voice command based on the received voice data (S1005).

[0293] In one implementation, processor 560 can use a speech-to-text (STT) engine stored in memory 530 to convert speech data into text data.

[0294] In another embodiment, processor 560 can send voice data to an STT server and receive converted text data from the STT server.

[0295] The processor 560 can use a natural language processing (NLP) engine to obtain analysis results from the converted text data.

[0296] The analysis results can include intent analysis results regarding the intent of the voice command and search results based on the intent analysis results.

[0297] The processor 560 of the AI ​​server 500 can determine whether a response from the generative AI is needed based on the analysis results of the voice command (S1007).

[0298] In one implementation, the processor 560 can determine whether a response from generative AI is needed based on the intent analysis results of the voice command.

[0299] When the intent analysis result of a voice command is an intent to control the function of the display device 100, the processor 560 can determine that a response from generative AI is not required.

[0300] When the intent analysis result of the voice command is not an intent to control the function of the display device 100, the processor 560 can determine that a response from generative AI is needed.

[0301] When the intent analysis result of a voice command is not an intent for controlling the function of the display device 100, it can be an intent for content search.

[0302] For example, when the intent of a voice command is to control the function of display device 100 (such as screen brightness or volume control), processor 560 can determine that a response from generative AI is not needed. This is because display device 100 can handle it on its own without the help of generative AI.

[0303] When the intent of a voice command is a content search on display device 100 or a travel itinerary search for a specific area, processor 560 can determine that a response from generative AI is needed. This is because display device 100 has difficulty resolving this on its own without the assistance of generative AI.

[0304] When the processor 560 of the AI ​​server 500 determines that a generative AI response is not required, it can send the analysis results to the display device 100 via the communication unit 510 (S1009).

[0305] The processor 560 can send the intent analysis results and the search results corresponding to the intent analysis results to the display device 100 via the communication unit 510.

[0306] The processor 560 can send only the search results corresponding to the intent analysis results to the display device 100 via the communication unit 510.

[0307] The controller 170 of the display device 100 can output the received analysis results (S1011).

[0308] The controller 170 can display the search results included in the analysis results on the display 180.

[0309] When the processor 560 of the AI ​​server 500 determines that a response from the generative AI is needed, it can generate a first prompt (S1013) for the response from the first generative AI server 610.

[0310] The processor 560 can generate a first hint for the response results obtained from the generative AI model set in the first generative AI server 610.

[0311] Generative AI models can be models that generate content using deep learning or machine learning in response to commands. These commands can include one or more text or images.

[0312] The processor 560 can generate the first prompt based on the intent analysis results of the voice command.

[0313] The processor 560 can generate a first prompt for the requested content by combining keywords included in the intent analysis results of the voice command.

[0314] The processor 560 of the AI ​​server 500 can send prompts to the first generative AI server 610 via the communication unit 510 (S1015), and can receive a first response result from the first generative AI server 610 (S1017).

[0315] In one implementation, the first generative AI server 610 may be a server that provides search results for content.

[0316] In another embodiment, the first generative AI server 610 may be a server equipped with a natural language generation AI model that outputs natural language.

[0317] Natural language generation AI models can generate an initial response in response to received prompts.

[0318] Natural language generation AI models can be models trained using deep learning architectures such as recurrent neural networks (RNNs) or transformation models that use attention mechanisms on large-scale text data.

[0319] Data can be collected from the web or various text sources to train natural language generation AI models.

[0320] Natural language generation AI models can generate one or more sentences containing content corresponding to intent analysis results by taking received prompts as input.

[0321] The first generative AI server 610 can obtain one or more generated sentences as the first response result and send the obtained first response result to the AI ​​server 500.

[0322] The processor 560 of the AI ​​server 500 can determine whether a response from another generative AI is needed based on the received first response result (S1019).

[0323] In one implementation, when the first response result includes only a specific type of information, the processor 560 can determine that a response from another generative AI is needed. For example, when the first response result includes only text information, the processor 560 can determine that a response from another generative AI is needed to obtain image information.

[0324] In another embodiment, when an additional voice command from a user is received after providing a first response result to the display device 100, the processor 560 may determine that a response from another generative AI is required.

[0325] When the first generative AI server 610 is unable to provide a response to the analysis results of an additional voice command, the processor 560 can determine that a response from another generative AI is needed.

[0326] When the processor 560 of the AI ​​server 500 determines that no response is needed from another generative AI, it can send the analysis results of the voice command and the first response results to the display device 100 through the communication unit 510 (S1021).

[0327] The controller 170 of the display device 100 can output the received analysis results and the first response results (S1023).

[0328] Furthermore, when the processor 560 of the AI ​​server 500 determines that a response from another generative AI is needed, it can generate a second prompt based on the result of the first response (S1025).

[0329] The processor 560 can generate a second prompt by combining the analysis results of the first response and the user's additional voice commands.

[0330] For example, the second prompt may include the analysis results of the image (which is an image) that requests the result of the first response, as well as the command of the analysis results with an additional voice command.

[0331] The processor 560 of the AI ​​server 500 can send a second prompt to the second generative AI server 630 via the communication unit 510 (S1027), and can receive a second response result from the second generative AI server 630 (S1029).

[0332] The second generative AI server 630 can be a server equipped with a generative AI model that provides image analysis results.

[0333] The second generative AI server 630 can obtain one or more sentences as a second response result to the prompt, and send the obtained second response result to the AI ​​server 500.

[0334] The processor 560 of the AI ​​server 500 can send the second response result to the display device 100 via the communication unit 510 (S1031).

[0335] The controller 170 of the display device 100 can output the received second response result (S827).

[0336] The controller 170 can display the second response result on the display 180.

[0337] Figures 11a to 11c This is a diagram illustrating the process of providing the response result of another generative AI when the response result of another generative AI is required, according to an embodiment of the present disclosure.

[0338] Figures 11a to 11c It shows Figure 10 Examples of implementation methods.

[0339] Reference Figure 11a The user issues the first voice command to the remote control device 200: "Picture of a cat drawn on a piano".

[0340] The remote control device 200 sends a first voice command to the display device 100, and the display device 100 can send the first voice data corresponding to the first voice command to the AI ​​server 500.

[0341] AI server 500 can convert the first voice data into the first text data and obtain the analysis results of the converted first text data.

[0342] AI server 500 can send multiple images to display device 100 indicating the search results for the first voice command.

[0343] The display device 100 can display multiple search images 1110 and text phrases indicating that it is waiting for the first response result from the first generative AI server 610 on the screen 1100.

[0344] AI server 500 can generate a first prompt based on the analysis results of the first voice command, and send the generated first prompt to first generative AI server 610. The first prompt may include a command for drawing a picture of a cat on a piano.

[0345] The first generative AI server 610 can generate multiple response images representing "cat on a piano" based on the first prompt, and send the generated multiple response images (first response results) to the AI ​​server 500.

[0346] The display device 100 can display multiple search images 1110 and multiple response images 1130 on the screen 1100.

[0347] Afterward, the user can issue a second voice command to the remote control 200: "Set a second image on the TV".

[0348] The remote control device 200 can send a second voice command to the display device 100.

[0349] The display device 100 can send second voice data corresponding to the second voice command to the AI ​​server 500.

[0350] AI server 500 can determine the required response from another generative AI based on the analysis results of the second speech data. The analysis results of the second speech data may include the intent used to generate the title of the image.

[0351] AI server 500 can generate a second prompt based on the analysis results of the second voice data (setting the title of the second image) and the first response results (multiple response images). The second prompt may include commands for setting a second image included in the multiple response images and the title of that image.

[0352] AI server 500 can send the second prompt to second generative AI server 630, and second generative AI server 630 can generate a title for the image based on the second prompt.

[0353] The second generative AI server 630 can send the generated image title to the AI ​​server 500, and the AI ​​server 500 can send the image title to the display device 100.

[0354] The display device 100 can display a second image 1151 and text 1153 including the image title (Black Cat Playing Piano) on the display 180.

[0355] In this way, according to the embodiments of this disclosure, various services can be provided to users by combining the response results from an external generative AI server to voice commands issued by the user.

[0356] According to embodiments of the present invention, the above method can be implemented as code readable by a processor on a medium on which a program is recorded. Examples of processor-readable media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0357] The above-described display device can be configured by selectively combining all or part of each embodiment, thereby allowing for various modifications rather than being limited to the configuration and methods of the described embodiments.

Claims

1. An artificial intelligence device, the artificial intelligence device comprising: A communication unit configured to communicate with an electronic device and one or more generative artificial intelligence (AI) servers; as well as A processor configured to: receive voice commands issued by a user from the electronic device; and obtain analysis results of the received voice commands; When it is determined, based on the obtained analysis results, that a response from one or more generative AI servers is required, a prompt is generated based on the analysis results; the generated prompt is sent to one or more generative AI servers; and a response corresponding to the prompt is received from one or more generative AI servers. And send the analysis results and the response results to the electronic device.

2. The artificial intelligence device according to claim 1, wherein, The processor is configured to first send the analysis results to the electronic device, and then send the response results to the electronic device.

3. The artificial intelligence device according to claim 1, wherein, The processor is configured to: determine that no response from the one or more generative AI servers is needed when the intent analysis result of the voice command is an intent to control the function of the display device, and determine that a response from the one or more generative AI servers is needed when the intent analysis result of the voice command is not an intent to control the function of the display device.

4. The artificial intelligence device according to claim 1, wherein, The processor is configured to generate the prompt based on the analysis results and the user's preference information.

5. The artificial intelligence device according to claim 1, wherein, The analysis results include intent analysis results and search results corresponding to the intent analysis results, and the processor is configured to generate the prompt based on the intent analysis results and the search results.

6. The artificial intelligence device according to claim 1, wherein, The processor is configured to send the prompt to each of the first generative AI server and the second generative AI server, and to receive a first response result from the first generative AI server and a second response result from the second generative AI server.

7. The artificial intelligence device according to claim 6, wherein, The processor is configured to generate a third response result based on the analysis result, the first response result, and the second response result, and to send the generated third response result to the electronic device.

8. The artificial intelligence device according to claim 6, wherein, The processor is configured to send the analysis results, the first response result, and the second response result to the electronic device.

9. The artificial intelligence device according to claim 1, wherein, The processor is configured to: when it is determined, based on the obtained analysis results, that a response from a first generative AI server is required, generate a first prompt based on the analysis results, send the generated first prompt to the first generative AI server, and receive a first response result from the first generative AI server; and when it is determined, based on the first response result, generate a second prompt, send the second prompt to the second generative AI server, and receive a second response result from the second generative AI server.

10. The artificial intelligence device according to claim 9, wherein, The processor is configured to receive additional voice commands from the electronic device and determine, based on the analysis results of the additional voice commands, whether a response from the second generative AI server is required.

11. The artificial intelligence device according to claim 10, wherein, The processor is configured to: based on the analysis results of the additional voice command, determine that a response from the second generative AI server is needed when the first generative AI server cannot provide a response to the analysis results of the additional voice command.

12. The artificial intelligence device according to claim 11, wherein, The processor is configured to generate the second prompt based on the analysis results of the first response result and the additional voice command.

13. The artificial intelligence device according to claim 1, wherein, The processor is configured to send information to the electronic device regarding the display format for the analysis results and the response results.

14. The artificial intelligence device according to claim 13, wherein, The display format is a format that includes one or more of the position, size, or arrangement of the areas used to display the analysis results and the response results.

15. A method of operating an artificial intelligence device, the method comprising the following steps: Receive voice commands from the user via electronic devices; Obtain the analysis results of the received voice commands; When it is determined, based on the obtained analysis results, that a response from one or more generative AI servers is required, a prompt is generated based on the analysis results; The generated prompts are sent to one or more generative AI servers; Receive response results corresponding to the prompts from one or more generative AI servers; as well as The analysis results and the response results are sent to the electronic device.