Image display device and system thereof
By receiving voice input through an image display device and system, generating and processing text, and using a server for natural language processing and artificial intelligence analysis, the problem of difficulty in constructing and editing long texts in existing technologies has been solved. This enables the function of recommending keywords based on text context, thereby improving the user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2026-03-27
AI Technical Summary
Existing image display devices struggle to effectively utilize user voice input to construct long and edit text, and lack the ability to recommend keywords based on text context.
The image display device and system receive voice input through a remote control device, generate text and process the text context, and use a server for natural language processing and artificial intelligence analysis to recommend keywords.
It enables users to construct and edit long texts using voice input, and recommends keywords based on text context, thus improving the user interaction experience.
Smart Images

Figure CN121753346A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image display device and a system for the image display device. Background Technology
[0002] An image display device is a device capable of displaying images that can be viewed by a user. For example, an image display device may include a television, monitor, or notebook computer equipped with a liquid crystal display (LCD) that uses liquid crystal or an OLED display that uses organic light-emitting diodes (OLED).
[0003] With the development of technology, various services applying speech recognition technology are being developed and provided in various fields. Speech recognition technology is also being widely developed for image display devices such as TVs. In existing technologies, speech recognition is typically performed on words, phrases, or short sentences spoken by the user, and the image display device then performs operations based on the results of the speech recognition.
[0004] On the other hand, with the rapid development of artificial intelligence technology recently, the trend is to develop services that summarize the core content or provide responses by analyzing long texts input by users. Therefore, there is a need to develop a user interface that utilizes the user's spoken voice to construct long texts. Summary of the Invention
[0005] The problem that the invention aims to solve
[0006] One purpose of this disclosure is to address the aforementioned problems and others.
[0007] Another objective is to provide an image display device and a system for such image display device that can utilize a user's voice input to construct long text.
[0008] Another objective is to provide an image display device and a system for editing text using a user's voice input.
[0009] Another objective is to provide an image display device and a system for recommending keywords for an article, taking into account the context of the text.
[0010] Technical solutions to the problem
[0011] An image display apparatus according to an embodiment of the present disclosure for achieving the above-described objectives may include: a display; an external device interface for communicating with a remote control device; a network interface for communicating with at least one server; and a control unit; wherein the control unit, based on an action of receiving a first instruction to activate a function that receives voice input from the remote control device, outputs a predetermined screen through the display including a first object corresponding to the end of the function; the control unit, based on an action of receiving voice input from the remote control device, outputs text corresponding to the voice input received from the remote control device through the predetermined screen; and the control unit, based on an action of receiving a second instruction to select the first object, processes the entire text output through the predetermined screen.
[0012] A system according to an embodiment of this disclosure for achieving the above objectives may include an image display device, a remote control device, and at least one server. The remote control device transmits a first instruction to the image display device to activate a function for receiving voice input. When the function is activated, the remote control device transmits voice input received via a microphone to the image display device. Based on an action received from the remote control device with the first instruction, the image display device outputs a predetermined screen including a first object corresponding to the end of the function. The image display device transmits the voice input received from the remote control device to the server. The image display device outputs text corresponding to the voice input received from the server through the predetermined screen. Based on an action received with a second instruction to select the first object, the entire text output through the predetermined screen is processed.
[0013] Invention Effects
[0014] The effects of the image display device and the system of the image display device according to this disclosure are as follows.
[0015] According to at least one embodiment of this disclosure, long text can be constructed using a user's voice input.
[0016] According to at least one embodiment of this disclosure, text can be edited using a user's voice input.
[0017] According to at least one embodiment of this disclosure, it is possible to recommend text for use in composing an article, taking into account the context of the text.
[0018] The further scope of this disclosure will become apparent from the detailed description below. However, it should be understood that while the detailed description and specific examples indicate preferred embodiments of this disclosure, they are given by way of illustration only, as various changes and modifications within the spirit and scope of this disclosure will be clearly seen by those skilled in the art from this detailed description. Attached Figure Description
[0019] Figure 1 This is a diagram illustrating an image display system according to an embodiment of the present disclosure.
[0020] Figure 2 yes Figure 1 Internal block diagram of the image display device.
[0021] Figure 3 yes Figure 2 Internal block diagram of the control unit.
[0022] Figure 4a It is shown Figure 2 A diagram illustrating the control method of a remote control device. Figure 4b yes Figure 2 An example of an internal block diagram of a remote control device.
[0023] Figure 5 It is used for explanation Figure 1 A diagram of the server.
[0024] Figure 6 This is a flowchart of an operation method of an image display device according to an embodiment of the present disclosure.
[0025] Figure 7 This is a flowchart of an operation method of a system according to an embodiment of the present disclosure.
[0026] Figures 8 to 33 This is a diagram illustrating the operation of an image display device that uses voice input to construct text according to various embodiments of the present disclosure. Detailed Implementation
[0027] The present disclosure will now be described in detail with reference to the accompanying drawings. In the drawings, for clarity and conciseness in explaining the present disclosure, illustrations of parts unrelated to the description have been omitted, and throughout the specification, identical or very similar parts are indicated by the same reference numerals.
[0028] The component suffixes “module” and “unit” used in the following description are given for ease of writing instructions only and do not inherently carry any particular meaning or function. Therefore, “module” and “unit” can be used interchangeably.
[0029] In this application, terms such as “comprising” or “having” are intended to indicate the presence of the stated features, numbers, steps, operations, components, parts or combinations thereof, but should be understood to not exclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.
[0030] Furthermore, in this specification, terms such as "first" and "second" may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another.
[0031] Figure 1 This is a diagram illustrating an image display system according to various embodiments of the present invention.
[0032] Reference Figure 1 The image display system 10 may include an image display device 100 and / or a remote control device 200.
[0033] The image display device 100 can be a device that processes and outputs images. The image display device 100 is not particularly limited as long as it can output a screen corresponding to the image signal, such as a television (TV), a notebook computer, or a monitor.
[0034] The image display device 100 can receive broadcast signals, process the signals, and output a processed broadcast image. When the image display device 100 receives a broadcast signal, it can correspond to a broadcast receiving device.
[0035] The image display device 100 can receive broadcast signals wirelessly via an antenna or wiredly via a cable. For example, the image display device 100 can receive terrestrial broadcast signals, satellite broadcast signals, cable broadcast signals, Internet Protocol Television (IPTV) broadcast signals, etc.
[0036] The remote control device 200 can be connected to the image display device 100 via wired and / or wireless means, and can provide various control signals to the image display device 100. In this case, the remote control device 200 can establish a wired or wireless network with the image display device 100, and send various control signals to the image display device 100 through the established network, or receive signals related to various operations processed in the image display device 100 from the image display device 100.
[0037] For example, various input devices such as mice, keyboards, space remote controls, trackballs, joysticks, etc., can be used as remote control devices 200. Remote control devices 200 can be referred to as external devices. Hereinafter, it is noted that external devices and remote control devices can be used interchangeably as needed.
[0038] The image display device 100 can be connected to a single remote control device 200, or to two or more remote control devices 200 at the same time, and can change the objects displayed on the screen or adjust the state of the screen based on control signals provided from each remote control device 200.
[0039] On the other hand, the image display system 10 may further include at least one server 400. The image display device 100 can transmit and receive data with the server 400. For example, the image display device 100 can transmit and receive data with the server 400 via a network such as the Internet.
[0040] The image display device 100 can transmit data related to operations performed based on user input to the server 400, and the server 400 can store the data received from the image display device 100.
[0041] Figure 2 yes Figure 1 Internal block diagram of the image display device.
[0042] Reference Figure 2 The image display device 100 may include a broadcast receiver 105, an external device interface 130, a network interface 135, a storage unit 140, a user input interface 150, an input unit 160, a control unit 170, a display 180, an audio output unit 185, and / or a power supply unit 190.
[0043] The broadcast receiver 105 may include a tuner unit 110 and a demodulation unit 120.
[0044] Furthermore, unlike the accompanying drawings, the image display device 100 may include only the broadcast receiving unit 105 and the external device interface unit 130 among the broadcast receiving unit 105, the external device interface unit 130, and the network interface unit 135. That is, the image display device 100 may not include the network interface unit 135.
[0045] The tuner unit 110 can select a broadcast signal corresponding to a user-selected channel or all preset channels from broadcast signals received via an antenna (not shown) or a cable (not shown). The tuner unit 110 can convert the selected broadcast signal into an intermediate frequency signal or a baseband image or voice signal.
[0046] For example, if the selected broadcast signal is a digital broadcast signal, the tuner unit 110 can convert it into a digital intermediate frequency (DIF) signal; if it is an analog broadcast signal, it can convert it into an analog baseband image or voice signal (CVBS / SIF). In other words, the tuner unit 110 can process both digital and analog broadcast signals. The analog baseband image or voice signal (CVBS / SIF) output from the tuner unit 110 can be directly input to the control unit 170.
[0047] Furthermore, the tuner unit 110 can sequentially select broadcast signals from all broadcast channels stored through the channel storage function from the received broadcast signals and convert them into intermediate frequency signals or baseband image or voice signals.
[0048] Furthermore, the tuner unit 110 may be equipped with a plurality of tuners to receive broadcast signals from a plurality of channels. Alternatively, it may be a single tuner that simultaneously receives broadcast signals from a plurality of channels.
[0049] The demodulation unit 120 can receive the digital intermediate frequency (DIF) signal converted in the tuner unit 110 and perform demodulation operation.
[0050] The demodulation unit 120 can perform demodulation and channel decoding, and then output a streaming signal (TS). At this time, the streaming signal can be a multiplexed signal of image signal, voice signal or data signal.
[0051] The streaming signal output from the demodulation unit 120 can be input to the control unit 170. The control unit 170 can perform demultiplexing, image / speech signal processing, etc., and then output the image through the display 180 and output the speech through the audio output unit 185.
[0052] The external device interface unit 130 can send or receive data with connected external devices. For this purpose, the external device interface unit 130 may include an A / V input / output unit (not shown).
[0053] The external device interface unit 130 can be connected to external devices (such as DVD (Digital Versatile Disk), Blu-ray, game console, camera, camcorder, computer (laptop), set-top box, etc.) via wired / wireless connection, and can perform input / output operations with external devices.
[0054] Furthermore, the external device interface section 130 can be connected to, for example, Figure 1 The various remote control devices 200 shown establish a communication network, receive control signals related to the operation of the image display device 100 from the remote control device 200, or transmit data related to the operation of the image display device 100 to the remote control device 200.
[0055] The A / V input / output unit can receive image and audio signals from external devices. For example, the A / V input / output unit may include Ethernet terminals, USB terminals, Composite Video Banking Sync (CVBS) terminals, component terminals, S-video terminals (analog), Digital Visual Interface (DVI) terminals, High Definition Multimedia Interface (HDMI) terminals, Mobile High-Definition Link (MHL) terminals, RGB terminals, D-SUB terminals, IEEE 1394 terminals, SPDIF terminals, Liquid HD terminals, etc. Digital signals input through these terminals can be transmitted to the control unit 170. At this time, analog signals input through the CVBS and S-video terminals can be converted into digital signals by an analog-to-digital converter (not shown) and transmitted to the control unit 170.
[0056] The external device interface unit 130 may include a wireless communication unit (not shown) for short-range wireless communication with other electronic devices. Through this wireless communication unit, the external device interface unit 130 can exchange data with a neighboring mobile terminal. For example, in mirror mode, the external device interface unit 130 can receive device information, running application information, application images, etc., from the mobile terminal.
[0057] The external device interface unit 130 can use Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, etc. for short-range wireless communication.
[0058] The network interface unit 135 can provide an interface for connecting the image display device 100 to wired / wireless networks, including the Internet.
[0059] The network interface unit 135 may include a communication module (not shown) for connecting to wired / wireless networks. For example, the network interface unit 135 may include a communication module for WLAN (Wireless Local Area Network) (Wi-Fi), Wibro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), etc.
[0060] The network interface unit 135 can send or receive data with other users or other electronic devices through a connected network or other networks linked to the connected network.
[0061] The network interface unit 135 can receive network content or data provided by content providers or network operators. In other words, the network interface unit 135 can receive content such as movies, advertisements, games, VOD, and broadcasts, as well as related information, from content providers or network providers via the network.
[0062] The network interface unit 135 can receive firmware update information and update files provided by the network operator, and can send data to the Internet, content providers, or network operators.
[0063] The network interface unit 135 can select and receive the required applications from applications that are open to the public via the network.
[0064] The storage unit 140 can store programs for each signal processing and control within the control unit 170, and can also store processed image, voice, or data signals. For example, the storage unit 140 can store application programs designed to perform various tasks that the control unit 170 can handle, and can selectively provide a portion of the stored application programs according to the request of the control unit 170.
[0065] Programs stored in the storage unit 140 are not subject to special restrictions as long as they can be executed by the control unit 170.
[0066] The storage unit 140 can also perform the function of temporarily storing image, voice or data signals received from external devices through the external device interface unit 130.
[0067] The storage unit 140 can store information about a predetermined broadcast channel through a channel storage function such as a channel chart.
[0068] although Figure 2An embodiment is shown in which the storage unit 140 and the control unit 170 are arranged separately, but the scope of the present invention is not limited thereto, and the storage unit 140 may be included in the control unit 170.
[0069] The storage unit 140 may include at least one of volatile memory (e.g., DRAM, SRAM, SDRAM, etc.) and non-volatile memory (e.g., flash memory, hard disk drive (HDD), solid-state drive (SSD), etc.). In various embodiments of the present invention, the storage unit 140 and the memory can be used interchangeably.
[0070] The user input interface unit 150 can transmit user input signals to the control unit 170, or transmit signals from the control unit 170 to the user.
[0071] For example, user input signals (such as power on / off, channel selection, screen settings, etc.) can be sent / received from the remote control device 200, or user input signals input from local buttons (not shown) (such as power button, channel button, volume button, setting value, etc.) can be transmitted to the control unit 170, or user input signals input from a sensor unit (not shown) that senses user gestures can be transmitted to the control unit 170, or signals from the control unit 170 can be transmitted to the sensor unit.
[0072] The input unit 160 may be located on one side of the main body of the image display device 100. For example, the input unit 160 may include a touchpad, physical buttons, etc.
[0073] The input unit 160 can receive various user commands related to the operation of the image display device 100, and can transmit control signals corresponding to the input commands to the control unit 170.
[0074] The input unit 160 may include at least one microphone (not shown) and may receive the user's voice through the microphone.
[0075] The control unit 170 may include at least one processor and use the included processor to control the overall operation of the image display device 100. Here, the processor may be a general-purpose processor, such as a CPU (central processing unit). Of course, the processor may also be a dedicated device, such as an ASIC, or other hardware-based processor.
[0076] The control unit 170 can demultiplex the streams input through the tuner unit 110, demodulation unit 120, external device interface unit 130, or network interface unit 135, or process the demultiplexed signals to generate and output signals for image or voice output.
[0077] The display 180 can convert the image signals, data signals, OSD signals, and control signals processed in the control unit 170 or the image signals, data signals, and control signals received from the external device interface unit 130 to generate drive signals.
[0078] The display 180 may include a display panel (not shown) having a plurality of pixels.
[0079] The plurality of pixels provided on the display panel may have RGB subpixels. Alternatively, the plurality of pixels provided on the display panel may have RGBW subpixels. The display 180 may convert the image signals, data signals, OSD signals, control signals, etc. processed in the control unit 170 to generate drive signals for the plurality of pixels.
[0080] The display 180 can be a PDP (Plasma Display Panel), LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diode), flexible display, or a 3D display. 3D displays 180 can be either glasses-free or glasses-based.
[0081] Furthermore, the display 180 can be configured as a touch screen, serving as both an output device and an input device.
[0082] The audio output unit 185 receives the voice-processed signal from the control unit 170 and outputs it as sound.
[0083] The image signal processed in the control unit 170 can be input to the display 180 and displayed as an image corresponding to the image signal. Furthermore, the image signal processed in the control unit 170 can be input to an external output device via the external device interface unit 130.
[0084] The voice signal processed in the control unit 170 can be output as sound through the audio output unit 185. Furthermore, the voice signal processed in the control unit 170 can be input to an external output device through the external device interface unit 130.
[0085] although Figure 2 As not shown in the diagram, the control unit 170 may include a demultiplexing unit, an image processing unit, etc. This will be referred to later. Figure 3 Describe it.
[0086] Furthermore, the control unit 170 can control the overall operation within the image display device 100. For example, the control unit 170 can control the tuner unit 110 to select (tune) a broadcast corresponding to a channel selected by the user or a preset channel.
[0087] In addition, the control unit 170 can control the image display device 100 through user commands input through the user input interface unit 150 or through an internal program.
[0088] Furthermore, the control unit 170 can control the display 180 to display images. At this time, the image displayed on the display 180 can be a still image or a moving image, and can be a 2D image or a 3D image.
[0089] Furthermore, the control unit 170 can display a predetermined 2D object within the image displayed on the monitor 180. For example, the object can be at least one of the following: a connected webpage (newspaper, magazine, etc.), an EPG (Electronic Program Guide), various menus, widgets, icons, still images, moving images, and text.
[0090] Furthermore, the image display device 100 may further include a camera unit (not shown). The camera unit can capture images of the user. The camera unit can be implemented with a single camera, but is not limited to this; it can also be implemented with multiple cameras. The camera unit can be embedded above the display 180 of the image display device 100, or it can be arranged separately. Image information captured by the camera unit can be input to the control unit 170.
[0091] The control unit 170 can identify the user's position based on the image captured by the imaging unit. For example, the control unit 170 can determine the distance (z-axis coordinate) between the user and the image display device 100. Furthermore, the control unit 170 can determine the x-axis and y-axis coordinates corresponding to the user's position on the display 180.
[0092] The control unit 170 can detect the user's gestures based on images captured by the shooting unit, signals sensed by the sensor unit, or a combination thereof.
[0093] The power supply unit 190 can provide power to the entire image display device 100. In particular, it can power the control unit 170 (which can be implemented as a system on chip (SOC)), the display 180 for image display, and the audio output unit 185 for audio output.
[0094] Specifically, the power supply unit 190 may include a converter (not shown) that converts alternating current to direct current and a DC / DC converter (not shown) that converts the level of direct current.
[0095] The remote control device 200 can transmit user input to the user input interface 150. For this purpose, the remote control device 200 can use Bluetooth, RF (Radio Frequency) communication, infrared (IR) communication, UWB (Ultra-wideband), ZigBee, etc. Furthermore, the remote control device 200 can receive image, voice, or data signals output from the user input interface 150 and display them on the remote control device 200 or output them as voice.
[0096] Furthermore, the aforementioned image display device 100 can be a fixed or mobile digital broadcast receiver capable of receiving digital broadcasts.
[0097] and, Figure 2 The block diagram of the image display device 100 shown is only a block diagram of one embodiment of the present invention. The various structures in the block diagram can be integrated, added or omitted according to the specifications of the actual image display device 100.
[0098] In other words, if necessary, two or more components can be combined into one component, or one component can be subdivided into two or more components. Furthermore, the function performed by each block is for illustrating embodiments of the invention, and its specific operation or device does not limit the scope of the invention.
[0099] Figure 3 yes Figure 2 Internal block diagram of the control unit.
[0100] Reference Figure 3 According to an embodiment of the present invention, the control unit 170 may include a demultiplexing unit 310, an image processing unit 320, a processor 330, an OSD generator 340, a mixer 345, a frame rate conversion unit 350, and / or a formatter 360. Furthermore, it may further include an audio processing unit (not shown) and a data processing unit (not shown).
[0101] The demultiplexing unit 310 can demultiplex the input stream. For example, if an MPEG-2 TS is input, it can be demultiplexed and separated into image, voice, and data signals respectively. Here, the stream signal input to the demultiplexing unit 310 can be a stream signal output from the tuner unit 110, the demodulation unit 120, or the external device interface unit 130.
[0102] The image processing unit 320 can perform image processing on the demultiplexed image signal. For this purpose, the image processing unit 320 may include an image decoder 325 and a scaler 335.
[0103] The image decoder 325 can decode the demultiplexed image signal, and the scaler 335 can scale the image signal so that the resolution of the decoded image signal can be output on the display 180.
[0104] The image decoder 325 may include decoders for various standards. For example, it may include MPEG-2 and H.264 decoders, 3D image decoders for color images and depth images, decoders for multi-view images, etc.
[0105] The processor 330 can control the overall operation within the image display device 100 or the control unit 170. For example, the processor 330 can control the tuner 110 to select (tune) a broadcast corresponding to a user-selected channel or a pre-stored channel.
[0106] Furthermore, the processor 330 can control the image display device 100 through user commands or internal programs input via the user input interface 150.
[0107] In addition, the processor 330 can control data transmission with the network interface unit 135 or the external device interface unit 130.
[0108] Furthermore, the processor 330 can control the operation of the demultiplexing unit 310, image processing unit 320, OSD generator 340, etc. within the control unit 170.
[0109] OSD generator 340 can generate OSD signals based on user input or automatically. For example, based on user input signals input through input unit 160, signals can be generated for displaying various information in graphic or text form on the screen of display 180.
[0110] The generated OSD signal can include various data, such as the user interface screen of the image display device 100, various menu screens, widgets, icons, etc. Furthermore, the generated OSD signal can include 2D objects or 3D objects.
[0111] Furthermore, the OSD generator 340 can generate a pointer that can be displayed on the display 180 based on a pointing signal input from the remote control device 200.
[0112] OSD generator 340 may include a pointer signal processing unit (not shown) for generating pointers. The pointer signal processing unit (not shown) may also be set separately and not within OSD generator 340.
[0113] The mixer 345 can mix the OSD signal generated by the OSD generator 340 with the decoded image signal processed by the image processing unit 320. The mixed image signal can be provided to the frame rate conversion unit 350.
[0114] The frame rate converter (FRC) 350 can convert the frame rate of the input image. Furthermore, the frame rate converter 350 can also output directly without performing additional frame rate conversion.
[0115] The Formatter 360 can arrange the left-eye and right-eye image frames of the frame rate-converted 3D image. It can also output a synchronization signal (Vsync) for turning on the left and right glasses of a 3D viewing device (not shown).
[0116] Furthermore, the formatter 360 can change the format of the input image signal to an image signal for display on the monitor 180 and output it.
[0117] In addition, Formatter 360 can change the format of 3D image signals. For example, it can be changed to any 3D format, such as side-by-side, top / down, frame-sequential, interlaced, checker box, etc.
[0118] Furthermore, the formatter 360 can also convert 2D image signals into 3D image signals. For example, based on a 3D image generation algorithm, edges or optional objects within a 2D image signal can be detected, and objects or optional objects based on the detected edges can be separated and generated as 3D image signals. In this case, the generated 3D image signals can be separated and arranged as a left-eye image signal L and a right-eye image signal R as described above.
[0119] Furthermore, although not shown in the figure, a 3D processor (not shown) for 3D effect signal processing can be further provided after the formatter 360. Such a 3D processor can process the brightness, tint, and color adjustments of the image signal to enhance the 3D effect. For example, signal processing that makes nearby objects sharp and distant objects blurry can be performed. Moreover, the functionality of this 3D processor can be incorporated into the formatter 360 or into the image processing unit 320.
[0120] Furthermore, the audio processing unit (not shown) within the control unit 170 can perform speech processing on the demultiplexed speech signal. For this purpose, the audio processing unit (not shown) may include various decoders.
[0121] In addition, the audio processing unit (not shown) within the control unit 170 can handle bass (Base), treble (Treble), volume adjustment, etc.
[0122] The data processing unit (not shown) within the control unit 170 can perform data processing on the demultiplexed data signal. For example, if the demultiplexed data signal is an encoded data signal, it can be decoded. The encoded data signal may be Electronic Program Guide information that includes broadcast information (e.g., the start and end times of broadcast programs on each channel).
[0123] and, Figure 3 The block diagram of the control unit 170 shown is only a block diagram of one embodiment of the present invention. The various structures in the block diagram can be integrated, added or omitted according to the specifications of the actual implemented control unit 170.
[0124] In particular, the frame rate conversion unit 350 and the formatter 360 may not be located within the control unit 170, but may be set up separately or as a separate module.
[0125] Figure 4a It is shown Figure 2 A diagram illustrating the control method of a remote control device. Figure 4b yes Figure 2 An example of an internal block diagram of a remote control device.
[0126] Referring to FIG4, it can be confirmed that the pointer 205 corresponding to the remote control device 200 is displayed on the display 180 of the image display device 100.
[0127] Figure 4a It is shown Figure 2 A diagram illustrating the control method of a medium-to-long-range control device.
[0128] Reference Figure 4a (a) The user can move the remote control device 200 up, down, left, right, forward, and backward, or rotate it. At this time, the pointer 205 displayed on the monitor 180 of the image display device 100 can be displayed corresponding to the movement of the remote control device 200. As shown, since the pointer 205 moves and is displayed according to the movement of the remote control device 200 in 3D space, it can be called a spatial remote control or a 3D pointing device.
[0129] Reference Figure 4a(b) When the user moves the remote control device 200 to the left, it can be confirmed that the pointer 205 displayed on the display 180 of the image display device 100 also moves to the left in response to the movement of the remote control device 200.
[0130] Information about the movement of the remote control device 200, sensed by the sensors of the remote control device 200, can be transmitted to the image display device 100. The image display device 100 can calculate the coordinates of the pointer 205 from the information about the movement of the remote control device 200. The image display device 100 can display the pointer 205 to correspond to the calculated coordinates.
[0131] Reference Figure 4a (c) While holding down a specific button on the remote control device 200, the user can move the remote control device 200 away from the display 180. As a result, the selected area on the display 180 corresponding to the pointer 205 can be zoomed in and enlarged. Conversely, when the user moves the remote control device 200 closer to the display 180 while holding down the specific button, the selected area on the display 180 corresponding to the pointer 205 can be zoomed out and reduced in size.
[0132] Alternatively, the selected area can be shrunk when the remote control device 200 is away from the display 180, and enlarged when the remote control device 200 is close to the display 180.
[0133] Furthermore, when the user presses a specific button on the remote control device 200, the recognition of up, down, left, and right movements can be excluded. That is, when the remote control device 200 is moved away from or closer to the display 180, up, down, left, and right movements may not be recognized; only forward and backward movements will be recognized. When the user does not press a specific button on the remote control device 200, only up, down, left, and right movements of the remote control device 200 will be recognized, and correspondingly, only the pointer 205 will move.
[0134] Furthermore, the moving speed or direction of the pointer 205 can correspond to the moving speed or direction of the remote control device 200.
[0135] Reference Figure 4b The remote control device 200 may include a wireless communication unit 220, a user input unit 230, a sensor unit 240, an output unit 250, a power supply unit 260, a storage unit 270 and / or a control unit 280.
[0136] The wireless communication unit 220 can send and receive signals with the image display device 100.
[0137] In this embodiment, the remote control device 200 may include an RF (Radio Frequency) module 221, which can transmit and receive signals with the image display device 100 according to the RF communication standard. Furthermore, the remote control device 200 may include an IR (Infrared Radiation) module 223, which can transmit and receive signals with the image display device 100 according to the IR communication standard.
[0138] The remote control device 200 can transmit a signal containing its movement information to the image display device 100 via the RF module 221. The remote control device 200 can also receive signals transmitted by the image display device 100 via the RF module 221.
[0139] The remote control device 200 can transmit commands regarding power on / off, channel change, volume change, etc., to the image display device 100 via the IR module 223.
[0140] The user input unit 230 can be configured as a keyboard, buttons, touchpad, touchscreen, etc. Users can operate the user input unit 230 to input commands related to the image display device 100 to the remote control device 200.
[0141] If the user input unit 230 has a hard key button, the user can input commands related to the image display device 100 to the remote control device 200 by pressing the hard key button.
[0142] If the user input unit 230 has a touch screen, the user can touch the soft keys on the touch screen to input commands related to the image display device 100 to the remote control device 200.
[0143] Furthermore, the user input unit 230 may also include various types of input methods that can be operated by the user, such as scroll keys, toggle keys, etc. This embodiment does not limit the scope of the invention.
[0144] The user input unit 230 may include a microphone. The user can speak into the microphone provided in the user input unit 230. At this time, the microphone provided in the user input unit 230 can receive the user's spoken voice.
[0145] The sensor unit 240 may include a gyroscope sensor 241 or an accelerometer sensor 243. The gyroscope sensor 241 can sense the movement of the remote control device 200.
[0146] The gyroscope sensor 241 can sense information about the operation of the remote control device 200 based on the x, y, and z axes. The accelerometer sensor 243 can sense information such as the moving speed of the remote control device 200. Furthermore, the sensor unit 240 may also include a distance measurement sensor capable of sensing the distance from the display 180.
[0147] The output unit 250 can output an image or sound corresponding to the operation of the user input unit 230 or the signal transmitted from the image display device 100. Through the output unit 250, the user can identify whether the user input unit 230 has been operated or whether the image display device 100 has been controlled.
[0148] The output unit 250 may include an LED module 251 having at least one light-emitting element (e.g., LED), a vibration module 253 that generates vibration, a sound output module 255 that outputs sound, and / or a display module 257 that outputs images.
[0149] The power supply unit 260 can supply power to each component of the remote control device 200. The power supply unit 260 may include at least one battery (not shown).
[0150] If the power supply unit 260 does not detect movement of the remote control device 200 during a specified period of time by the sensor unit 240, it can prevent unnecessary power consumption by interrupting the power supply to each component of the remote control device 200.
[0151] The power supply unit 260 can restore power to each component of the remote control device 200 when a predetermined event occurs. For example, the power supply unit 260 can restore power to each component when a designated key on the remote control device 200 is operated. For example, the power supply unit 260 can restore power to each component of the remote control device 200 when the sensor unit 240 senses movement of the remote control device 200.
[0152] The storage unit 270 can store various types of programs, application data, etc. required for the control or operation of the remote control device 200.
[0153] When the remote control device 200 wirelessly transmits and receives signals with the image display device 100 via the RF module 221, the remote control device 200 and the image display device 100 can transmit and receive signals via a predetermined frequency band. The control unit 280 of the remote control device 200 can store and reference information such as the frequency band for wirelessly transmitting and receiving signals with the image display device 100 paired with the remote control device 200 in the storage unit 270.
[0154] The control unit 280 may include at least one processor and use the processor therein to control the overall operation of the remote control device 200.
[0155] The control unit 280 can transmit control signals corresponding to predetermined key operations of the user input unit 230, or control signals corresponding to the movement of the remote control device 200 sensed by the sensor unit 240, to the image display device 100 via the wireless communication unit 220.
[0156] The user input interface 150 of the image display device 100 may include a wireless communication unit 151 capable of wirelessly transmitting and receiving signals with the remote control device 200, and a coordinate calculation unit 155 capable of calculating the coordinate values of a pointer corresponding to the operation of the remote control device 200.
[0157] The user input interface unit 150 can wirelessly transmit and receive signals with the remote control device 200 via the RF module 152. Furthermore, it can receive signals transmitted by the remote control device 200 according to the IR communication standard via the IR module 153.
[0158] The coordinate calculation unit 155 can perform hand tremor or error correction on the signal received by the wireless communication unit 151 corresponding to the operation of the remote control device 200, and calculate the coordinate values (x, y) of the pointer 205 to be displayed on the display 180.
[0159] Signals sent by the remote control device 200 and input to the image display device 100 via the user input interface 150 can be transmitted to the control unit 170 of the image display device 100. The control unit 170 of the image display device 100 can confirm information about the operation and key operations of the remote control device 200 based on the signals sent by the remote control device 200, and control the image display device 100 accordingly.
[0160] As another example, the remote control device 200 can calculate the pointer coordinate values corresponding to its operation and output them to the user input interface 150 of the image display device 100. In this case, the user input interface 150 of the image display device 100 can transmit the received information about the pointer coordinate values to the control unit 170 without a separate hand tremor or error correction process.
[0161] Furthermore, as another example, the coordinate value calculation unit 115 may be located inside the control unit 170, rather than in the user input interface unit 150 as shown in the figure.
[0162] Figure 5 It is used for explanation Figure 1 A diagram of the server.
[0163] Reference Figure 5 Server 400 may include relay server 410, STT (Speech To Text) server 420, NLP (Natural Language Processing) server 430, AI (Artificial Intelligence) server 440, and / or database 450. This disclosure describes, but is not limited to, the relay server 410, STT server 420, NLP server 430, and AI server 440 as distinct from each other. For example, two or more of the relay server 410, STT server 420, NLP server 430, and AI server 440 may be comprised of a single server.
[0164] The relay server 410 can communicate with the image display device 100. The relay server 410 can transfer data between the STT server 420, the NLP server 430, and the image display device 100. The relay server 410 can store at least a portion of the data transferred between the STT server 420, the NLP server 430, and the image display device 100.
[0165] STT server 420 can receive voice data. STT server 420 can convert voice data into text data. STT server 420 can transmit text data to image display device 100 via relay server 410. STT server 420 can also be named ASR (Automatic Speech Recognition) server.
[0166] The STT server 420 can utilize language models to improve the accuracy of speech-to-text conversion. A language model can refer to a model that calculates the probability of a text or the probability of the next word appearing given previous words. For example, language models can include probabilistic language models such as unigram models, bigram models, and N-gram models. That is, the STT server 420 can use language models to determine whether the text data converted from speech data has been appropriately converted, thereby improving the accuracy of the text-to-speech conversion.
[0167] NLP server 430 can receive text data. NLP server 430 can perform intent analysis on the received text data. NLP server 430 can transmit intent analysis information representing the execution result of intent analysis to image display device 100 via relay server 410.
[0168] According to one embodiment, the NLP server 430 can generate intent analysis information by sequentially performing morphological analysis, syntactic analysis, speech act analysis, and dialogue processing steps on text data. The morphological analysis step classifies the text data corresponding to the user's spoken words into the smallest meaningful unit, the morphological unit, and determines the part of speech of each classified morphological element. The syntactic analysis step uses the results of the morphological analysis step to distinguish the text data into noun phrases, verb phrases, adjective phrases, etc., and determines the relationships between the distinguished phrases. Through the syntactic analysis step, the subject, object, and modifiers of the user's spoken words can be determined. The speech act analysis step uses the results of the syntactic analysis step to analyze the intent of the user's spoken words. Specifically, the speech act analysis step determines whether the user's intent is a question, a request, or simply an emotional expression. The dialogue processing step uses the results of the speech act analysis step to determine whether to respond to the user's speech, provide a reply, or ask for additional information.
[0169] AI server 440 can transmit responses to requests from NLP server 430 to relay server 410 or NLP server 430. For example, when AI server 440 receives a content retrieval request from NLP server 430, it can transmit at least one piece of content corresponding to the received content retrieval request to NLP server 430. For example, when AI server 440 receives a recommendation request from NLP server 430 regarding keywords (hereinafter, associated keywords) associated with specified keywords, it can transmit at least one associated keyword corresponding to the specified keywords to NLP server 430.
[0170] According to one embodiment, the AI server 440 may include a plurality of AI agent servers. Here, the plurality of AI agent servers may each correspond to a type of request from the NLP server 430. For example, when the AI server 440 receives a text association request from the NLP server 430, it may utilize a first AI agent server; when it receives a video association request, it may utilize a second AI agent server.
[0171] Database 450 can store data generated for the response of AI server 440. Database 450 may include a plurality of sub-databases 451 to 453. For example, database 450 may include a first database 451 storing text such as syllables and words, a second database 452 storing speech, a third database 453 storing various learning models, and so on.
[0172] Figure 6This is a flowchart of an operation method of an image display device according to an embodiment of the present disclosure.
[0173] Reference Figure 6 The image display device 100 can activate the function of receiving voice input (hereinafter, voice recognition function) during operation S601. The image display device 100 can activate the voice recognition function based on a command to activate the voice recognition function (hereinafter, start command) received from the remote control device 200. For example, when the user presses a specific button included in the remote control device 200, the remote control device 200 can transmit the start command to the image display device 100. On the other hand, when the voice recognition function of the image display device 100 is activated, the remote control device 200 can activate the microphone included in the user input section 230.
[0174] When the voice recognition function is activated, the image display device 100 can output a specified screen corresponding to the voice recognition function via the display 180. In this case, the specified screen may include at least one object corresponding to the voice recognition function. For example, the specified screen may include an object corresponding to the end of the voice recognition function (hereinafter, the end object).
[0175] The image display device 100 can determine whether voice input has been received during operation S602. For example, if a user's voice is input to the remote control device 200 via a microphone included in the remote control device 200, the remote control device 200 can transmit the data about the input user's voice to the image display device 100.
[0176] When the image display device 100 receives voice input during operation S603, it can transmit the voice input to the server 400. Additionally, the image display device 100 can receive text corresponding to the voice input from the server 400. For example, when the image display device 100 transmits voice data to the server 400, the STT server 420 can convert the voice data into text data and transmit it to the image display device 100.
[0177] According to one embodiment, the image display device 100 can transmit voice data in preset units such as syllables and words to the server 400. That is, when a user speaks an article, the image display device 100 can transmit voice data in preset units to the server 400 and receive text data in preset units from the server 400 while receiving voice input corresponding to the article from the remote control device 200.
[0178] During operation S604, the image display device 100 can output text corresponding to the voice input via the display 180. The image display device 100 can display the text corresponding to the voice input via a predetermined screen corresponding to the voice recognition function output via the display 180.
[0179] In operation S605, the image display device 100 can determine whether the voice input received from the remote control device 200 corresponds to an article. For example, if a preset minimum time has elapsed after the voice input is received from the remote control device 200, the image display device 100 can determine that the voice input received from the remote control device 200 constitutes an article. For example, if the text corresponding to the voice input includes preset words, the image display device 100 can determine that the voice input received from the remote control device 200 constitutes an article. Here, the preset words can include words, sentences, words or sentences following the received article, conjunctions (e.g., and, therefore, but, then), and meaningless words (e.g., function words) located in the middle or at the end of the article.
[0180] In operation S606, the image display device 100 can perform intent analysis on the voice input when it corresponds to a text. For example, the image display device 100 can transmit the text corresponding to the voice input to the NLP server 420 via the relay server 410. For example, if the relay server 410 stores the text corresponding to the voice input converted in the STT server 420, the image display device 100 can request intent analysis of the text corresponding to the voice input stored in the relay server 410 from the server 400.
[0181] This disclosure describes how the image display device 100 determines whether the voice input corresponds to an article, but is not limited thereto. For example, the relay server 410 can determine whether the voice input received from the image display device 100 corresponds to an article. In addition, the relay server 410 can transmit the text corresponding to the voice input stored in the relay server 410 to the NLP server 420.
[0182] In operation S607, the image display device 100 can determine whether the voice input is an input related to editing the current text based on the intent analysis information received from the NLP server 430. For example, the image display device 100 can determine whether the intent analysis information received from the NLP server 430 includes an editing type corresponding to the voice input. Here, the editing type may include text insertion, deletion, modification, replacement, etc.
[0183] In operation S608, when the voice input is for text editing, the image display device 100 can delete the text corresponding to the voice input for text editing on a designated screen. Furthermore, the image display device 100 can edit the currently displayed text on the designated screen based on the voice input for text editing. For example, the intent analysis information received from the NLP server 430 may include keywords included in the voice input, text components corresponding to the keywords, and the editing type corresponding to the voice input. In this case, the image display device 100 can edit the text displayed on the designated screen based on the keywords, text components, editing type, etc., included in the intent analysis information.
[0184] During operation S609, the image display device 100 can determine whether it has received a command to end the voice recognition function (hereinafter, an end command). For example, the image display device 100 can determine that it has received an end command if it receives an input from the remote control device 200 to select the end object to be displayed on a specified screen.
[0185] According to one embodiment, if the image display device 100 receives a start command again before receiving an end command, it can initialize the entire text output through the specified screen until it receives a start command again.
[0186] On the other hand, during operation S610, the image display device 100 can determine whether a predetermined time has elapsed since the reception of voice input if no voice input has been received. Here, the predetermined time may exceed a preset minimum time.
[0187] In operation S611, if a predetermined time has elapsed after receiving voice input, the image display device 100 can output at least one keyword (hereinafter, recommended keyword) associated with the voice input through a predetermined screen. The image display device 100 can then output an object (hereinafter, recommended object) corresponding to the recommended keyword through a predetermined screen.
[0188] For example, if a predetermined time has elapsed after receiving the first voice input, the image display device 100 may request recommended keywords related to the first voice input from the server 400. At this time, the image display device 100 may transmit the text corresponding to the first voice input to the server 400. The server 400 may transmit related keywords about the first voice input, included in the text received from the image display device 100, as recommended keywords to the image display device 100. Alternatively, if a predetermined time has elapsed after receiving the first voice input, the server 400 may also transmit recommended keywords related to the first voice input to the image display device 100.
[0189] During operation S612, the image display device 100 can determine whether it has received input from the remote control device 200 to select a recommended keyword to be displayed on a specified screen. For example, when the coordinates of the pointer 205 correspond to the recommended object, the image display device 100 can determine whether it has received input from the remote control device 200 to select an object.
[0190] In operation S613, the image display device 100 can output text corresponding to the selected recommended keyword when a recommended keyword is selected. On the other hand, when a first recommended keyword is selected, the image display device 100 outputs at least one second recommended keyword associated with the selected first recommended keyword.
[0191] On the other hand, during operation S613, the image display device 100 can process the entire text output through the designated screen upon receiving an end command. In this disclosure, the processing of the entire text output through the designated screen as input for a service utilizing conversational artificial intelligence is used as an example for explanation.
[0192] Figure 7 This is a flowchart illustrating an operation method of a system according to an embodiment of this disclosure. Regarding... Figure 6 The content described herein is repeated, and detailed explanations are omitted.
[0193] Reference Figure 7 During operation S701, the image display device 100 can activate the voice recognition function. For example, the image display device 100 can activate the voice recognition function based on a start command received from the remote control device 200.
[0194] The image display device 100 can receive voice input during S702 operation. For example, the image display device 100 can receive data about the user's voice input via a microphone included in the remote control device 200 from the remote control device 200.
[0195] During S703 operation, the image display device 100 can transmit received voice input to the relay server 410. For example, while receiving voice input, the image display device 100 can transmit voice data in preset units such as syllables and words to the relay server 410.
[0196] In S704 operation, relay server 410 can transmit voice input received from image display device 100 to STT server 420.
[0197] In S705 operation, the STT server 430 can convert voice input received from the relay server 410 into text.
[0198] In S706 operation, STT server 430 can transmit text corresponding to voice input to relay server 410.
[0199] During operation S707, relay server 410 can transmit text received from STT server 420 to image display device 100. According to one embodiment, relay server 410 can store text received from STT server 420.
[0200] During S708 operation, the image display device 100 can output text corresponding to the voice input received from the relay server 410 through a specified screen.
[0201] In operation S709, when the voice input corresponds to a text, the image display device 100 can request intent analysis of the voice input from the relay server 410. For example, if the relay server 410 stores text corresponding to the voice input converted in the STT server 420, the image display device 100 can request intent analysis of the text corresponding to the voice input stored in the relay server 410. For example, if the relay server 410 does not store text corresponding to the voice input converted in the STT server 420, the image display device 100 can transmit the text corresponding to the voice input to the relay server 410.
[0202] During S710 operation, relay server 410 can transmit text corresponding to voice input to NLP server 430.
[0203] In operation S711, the NLP server 430 can perform intent analysis on the text corresponding to the voice input. For example, the NLP server 430 can determine the keywords in the voice input, the textual components of the keywords, and the intent of the text through intent analysis of the voice input. At this time, if the intent of the text is to edit the current text, the NLP server 430 can determine the editing type corresponding to the voice input.
[0204] In S712 operation, NLP server 430 can transmit intent analysis results, including intent analysis information, to relay server 410.
[0205] In S713 operation, relay server 410 can transmit the intent analysis results received from NLP server 430 to image display device 100.
[0206] In operation S714, the image display device 100 can perform operations based on the intent analysis results. For example, if the voice input is not related to editing the current text, the image display device 100 can maintain the text displayed on the designated screen. Alternatively, if the voice input is related to editing the current text, the image display device 100 can delete the text corresponding to the voice input about editing the text on the designated screen. Furthermore, the image display device 100 can edit the current text displayed on the designated screen based on the voice input about editing the text.
[0207] On the other hand, during operation S715, the image display device 100 can request recommended keywords from the relay server 410. For example, if a predetermined time has elapsed after receiving the first voice input, the image display device 100 can request recommended keywords related to the first voice input from the relay server 410. At this time, the image display device 100 can transmit the text corresponding to the first voice input to the relay server 410. Alternatively, if the relay server 410 stores the text corresponding to the first voice input, the image display device 100 can omit the transmission of the text corresponding to the first voice input from the relay server 410.
[0208] During operation S716, relay server 410 can request recommended keywords from AI server 440. For example, relay server 410 can transmit text corresponding to the first voice input to AI server 440. On the other hand, relay server 410 can also transmit text corresponding to the first voice input to NLP server 430 and request recommended keywords. In this case, NLP server 430 can transmit the intent analysis results regarding the text corresponding to the first voice input to AI server 440.
[0209] During S717 operation, AI server 440 can determine recommended keywords. For example, AI server 440 can use data included in database 450 to determine related keywords corresponding to the keywords included in the first voice input as recommended keywords.
[0210] According to one embodiment, AI server 440 can determine recommended keywords based on a preset pattern constituting an article. For example, when the received text corresponds to an article, AI server 440 can determine keywords related to the subject of the next article associated with the keywords of the article as recommended keywords. For example, when the received text is an incomplete article including a subject, AI server 440 can determine keywords related to the object associated with the subject included in the text as recommended keywords.
[0211] AI server 440 can transmit recommended keywords to relay server 410 during S718 operation.
[0212] In S719 operation, relay server 410 can transmit recommended keywords received from AI server 440 to image display device 100.
[0213] During S720 operation, the image display device 100 can output recommended keywords. The image display device 100 can output recommended objects corresponding to the recommended keywords through a specified screen.
[0214] On the other hand, during S730 operation, the image display device 100 can process the entire text output through the designated screen upon receiving an end command. For example, the image display device 100 can process the entire text output through the designated screen as input for a service utilizing conversational artificial intelligence.
[0215] Figures 8 to 33 This is a diagram illustrating the operation of an image display device that uses voice input to construct text according to various embodiments of the present disclosure.
[0216] Reference Figure 8 The image display device 100 can output a specified screen 800 corresponding to the voice recognition function based on a start command received from the remote control device 200 to activate the voice recognition function. At this time, the microphone of the remote control device 200 can be activated to receive voice input.
[0217] The specified screen 800 may include a first indicator 810 indicating a standby state regarding voice input, an end object 820, and / or an example of voice input.
[0218] Reference Figure 9 When the user says "Please make a fairy tale" as the first voice input, the remote control device 200 can transmit the first voice input received through the microphone to the image display device 100.
[0219] The image display device 100 can transmit a first voice input received from the remote control device 200 to the server 400. Based on the text corresponding to the first voice input received from the server 400, the image display device 100 can output text 910 corresponding to the first voice input through a designated screen 800. According to one embodiment, the image display device 100 can transmit the first voice input as voice data in preset units such as syllables and words to the server 400. In this case, the image display device 100 can sequentially output text corresponding to the voice data received from the server 400 through the designated screen 800.
[0220] During the receipt of first voice input from remote control device 200, image display device 100 may change a first indicator 810, which is included in a specified screen 800, to a second indicator 815 indicating the reception status of the voice input.
[0221] Reference Figure 10 The image display device 100 can change the second indicator 815 included in the specified screen 800 back to the first indicator 810 based on the completion of receiving the first voice input.
[0222] The image display device 100 can request intent analysis from the server 400 regarding the first voice input, "Please make a fairy tale." At this time, the first voice input, "Please make a fairy tale," is not input regarding editing the current text; therefore, the image display device 100 can maintain the text 910 corresponding to the first voice input output through the designated screen 800. On the other hand, the image display device 100 can receive data from the server 400 regarding keywords included in the first voice input, textual components corresponding to those keywords, etc.
[0223] According to one embodiment, the image display device 100 can output an object (hereinafter, additional object) 830 corresponding to the additional reception of voice input through a predetermined screen 800. If it is determined that the first voice input corresponds to a text, the image display device 100 can output the additional object 830 through the predetermined screen 800. If it is determined that the first voice input is not related to the editing of the current text, the image display device 100 can output the additional object 830 through the predetermined screen 800. On the other hand, after the additional object 830 is displayed on the predetermined screen 800, the microphone of the remote control device 200 can be disabled until the image display device 100 receives input from the remote control device 200 selecting the additional object 830.
[0224] On the other hand, according to one embodiment, if the image display device 100 does not receive input from the remote control device 200 to select the additional object 830 during a specified time period while the additional object 830 is displayed on the specified screen 800, it can be determined that an end command has been received.
[0225] Reference Figure 11When a user selects an additional object 830 using the pointer 205 displayed on the designated screen 800, the microphone of the remote control device 200 can be activated. Conversely, when input indicating selection of an additional object 830 is received from the remote control device 200, the image display device 100 can output a third indicator 835 on the designated screen 800 indicating that the image display device 100 has received voice input. For example, when a user selects an additional object 830 using the pointer 205 displayed on the designated screen 800, the image display device 100 can change the additional object 830 included on the designated screen 800 to the third indicator 835.
[0226] Reference Figure 12 and Figure 13 When a user speaks the story of "the evil witch bullying the kind pig" as a second voice input, the remote control device 200 can transmit the second voice input received through the microphone to the image display device 100. During the receipt of the second voice input from the remote control device 200, the image display device 100 can display a second indicator 815 on a designated screen 800.
[0227] The image display device 100 can transmit the second voice input as pre-set units such as syllables and words to the server 400. At this time, the image display device 100 can sequentially output text corresponding to the voice data received from the server 400 through a designated screen 800. Therefore, during the reception of the second voice input, text 915 corresponding to a portion of the first and second voice inputs can be output on the designated screen 800.
[0228] Upon receiving the second voice input, the image display device 100 can output text 920 corresponding to the first and second voice inputs through a designated screen 800.
[0229] The image display device 100 can request intent analysis from the server 400 regarding the "story of the evil witch bullying the kind pig" as the second voice input. At this time, the "story of the evil witch bullying the kind pig" as the second voice input is not input related to editing the current text; therefore, the image display device 100 can maintain the text 920 corresponding to the first and second voice inputs output through the designated screen 800. On the other hand, the image display device 100 can receive data from the server 400 including keywords from the second voice input, text components corresponding to the keywords, etc.
[0230] On the other hand, the image display device 100 can display the first indicator 810 on the designated screen 800 upon receiving the second voice input. Furthermore, since the second voice input corresponds to a text and is not related to editing the current text, the appended object 830 can be displayed on the designated screen 800.
[0231] On the other hand, according to one embodiment, the image display device 100 may not use a user interface regarding the added object 830. In this case, the image display device 100 may activate the microphone of the remote control device 200 after receiving a start command until an end command is received. Hereinafter, an embodiment without a user interface regarding the added object 830 will be described.
[0232] Reference Figure 14 When the pointer 205 points to text displayed on the designated screen 800, the image display device 100 can define the portion corresponding to the pointer 205 as the editing area 1400. At this time, the portion corresponding to the pointer 205 can be determined based on parts of speech such as nouns, adjectives, and verbs, and / or text components such as subjects, objects, and modifiers included in the text displayed on the designated screen 800. For example, if the coordinates of the pointer 205 correspond to the word "witch" in the second voice input displayed on the designated screen 800, the portion corresponding to "witch" can be set as the editing area 1400.
[0233] On the other hand, when the editing area 1400 is set, the image display device 100 can display objects associated with editing of the editing area 1400 on the designated screen 800. Objects associated with editing of the editing area 1400 may include objects corresponding to text modification (hereinafter, modification objects) 841, objects corresponding to text deletion (hereinafter, deletion objects) 842, etc.
[0234] Reference Figure 15 The image display device 100 can adjust the text portion corresponding to the editing area 1400. For example, the image display device 100 can adjust the text portion corresponding to the editing area 1400 based on input from the directional keys received from the remote control device 200.
[0235] On the other hand, according to one embodiment, the image display device 100 can adjust the text portion corresponding to the editing area 1400 based on the change in the position of the pointer 205. For example, when the image display device 100 receives input that the editing area 1400 has been selected via the remote control device 200, it can enlarge or shrink the text portion corresponding to the editing area 1400 to the left, corresponding to the leftward movement of the pointer 205.
[0236] Reference Figure 16When the text corresponding to "Bad Witch" is set as the editing area 1400, the user can use the pointer 205 displayed on the designated screen 800 to select the deletion object 842. At this time, the image display device 100 can display the deleted text 925 of "Bad Witch" on the designated screen 800.
[0237] On the other hand, refer to Figure 17 When the text corresponding to "Wicked Witch" is set as the editing area 1400, the user can select the object 841 to modify using the pointer 205 displayed on the designated screen 800. At this time, the image display device 100 can output a third indicator 835 indicating that the image display device 100 has received voice input through the designated screen 800. For example, when the user selects an additional object 830 using the pointer 205 displayed on the designated screen 800, the image display device 100 can change the object associated with editing included on the designated screen 800 to the third indicator 835.
[0238] Reference Figure 18 When a user speaks "a hunter with a terrifying gun" as a third voice input, the remote control device 200 can transmit the third voice input received via the microphone to the image display device 100. The image display device 100 can then transmit the third voice input received from the remote control device 200 to the server 400. The image display device 100 can then output text corresponding to the voice data received from the server 400 through a designated screen 800. At this time, the designated screen 800 can display text 930 where "evil witch" has changed to "hunter with a terrifying gun".
[0239] On the other hand, refer to Figure 19 and Figure 20 When the designated screen 800 displays text 920 corresponding to the first and second voice inputs, and the user speaks the fourth voice input, "Change the kind pig to a cute rabbit," the remote control device 200 can transmit the fourth voice input received through the microphone to the image display device 100. The image display device 100 can transmit the fourth voice input to the STT server 420 and receive the text corresponding to the fourth voice input. Based on the text corresponding to the fourth voice input, the image display device 100 can output text 940 corresponding to the first, second, and fourth voice inputs through the designated screen 800.
[0240] The image display device 100 can request intent analysis of the fourth voice input from the server 400. The NLP server 430 can perform intent analysis on the fourth voice input. At this time, the NLP server 430 can determine the intent of the text corresponding to the fourth voice input as editing the current text based on the phrase "want to modify" included in the fourth voice input. In addition, the NLP server can determine the editing type as text replacement, identify the keyword before replacement as "kind pig", identify the keyword after replacement as "cute rabbit", and determine that both keywords are noun phrases.
[0241] Reference Figure 21 The image display device 100 can determine that the fourth voice input is an input about text editing based on the intent analysis information received from the server 400. At this time, the image display device 100 can delete the text corresponding to "change the kind pig to a cute rabbit" as the fourth voice input from the designated screen 800.
[0242] The image display device 100 can display text 945, which changes "kind pig" to "cute rabbit" in text 920 corresponding to the first voice input and the second voice input, on a designated screen 800 based on keywords included in intent analysis information received from the server 400.
[0243] On the other hand, refer to Figure 22 and Figure 23 When the designated screen 800 displays text 920 corresponding to the first and second voice inputs, and the user speaks "I want to change it to a cute rabbit" as the fifth voice input, the remote control device 200 can transmit the fifth voice input received through the microphone to the image display device 100. The image display device 100 can transmit the fifth voice input to the STT server 420 and receive the text corresponding to the fifth voice input. Based on the text corresponding to the fifth voice input, the image display device 100 can output text 950 corresponding to the first, second, and fifth voice inputs through the designated screen 800.
[0244] The image display device 100 can request intent analysis of the fifth voice input from the server 400. The NLP server 430 can perform intent analysis on the fifth voice input. At this time, the NLP server 430 can determine the intent of the text corresponding to the fifth voice input as editing the current text based on the phrase "want to modify" included in the fifth voice input. In addition, the NLP server can identify the edited keyword as "cute rabbit", identify the keyword as a noun phrase composed of adjectives and nouns, and identify the editing type as text replacement.
[0245] Reference Figure 24The image display device 100 can determine, based on the intent analysis information received from the server 400, that the fifth voice input is an input related to text editing. At this time, the image display device 100 can delete the text corresponding to "want to change to a cute rabbit" that is the fifth voice input from the designated screen 800.
[0246] On the other hand, the image display device 100 can identify the noun phrase "cute rabbit" as a post-edit keyword based on keywords included in the intent analysis information received from the server 400. At this time, the image display device 100 can identify "bad witch" and "kind pig" (composed of adjectives and nouns) in the text 920 corresponding to the first and second voice inputs as pre-edit keywords, based on the fact that "cute rabbit" is a noun phrase composed of adjectives and nouns.
[0247] The image display device 100 can display fourth indicators 851 and 852 representing the pre-editing keywords on a designated screen, based on the fact that there are multiple pre-editing keywords. The fourth indicators 851 and 852 can be displayed adjacent to the text portion determined by the pre-editing keywords. Alternatively, the image display device 100 can display the portion determined by the pre-editing keywords as an editing area. When "bad witch" and "good pig" are determined as pre-editing keywords, the user can use the pointer 205 displayed on the designated screen 800 to select one of the fourth indicators 851 and 852 representing the pre-editing keywords. Alternatively, the user can use the pointer 205 displayed on the designated screen 800 to select one of the portions determined by the pre-editing keywords displayed as an editing area.
[0248] Reference Figure 25 When the user selects "bad witch" from "bad witch" and "kind pig" determined by the keywords before editing, the image display device 100 can display the text 955 of "bad witch" changed to "cute rabbit" in the text 920 corresponding to the first voice input and the second voice input on the designated screen 800.
[0249] On the other hand, according to one embodiment, a user can use voice input to edit a region of text output through a designated screen 800. For example, if a user says "I want to delete from the Bad Witch to the story," the intent of the text can be determined as text editing based on the "want to delete" included in the voice input in the NLP server 430. Furthermore, the NLP server can determine the editing type as text deletion and identify "Bad Witch" and "story" as keywords indicating the deletion range. Additionally, the image display device 100 can display the text from "Bad Witch" to "story" in the text 920 corresponding to the first and second voice inputs on the designated screen 800 based on the keywords included in the intent analysis information received from the server 400.
[0250] On the other hand, refer to Figure 26 and Figure 27 When a user utters "Please make a fairy tale" as the first voice input, the image display device 100 can display the text 810 corresponding to the first voice input on a designated screen 800. If it is determined that the first voice input corresponds to an article and is not related to editing the current text, the image display device 100 can display a fifth indicator 861 representing the article on the designated screen 800. The fifth indicator 861 corresponding to the first voice input can be arranged adjacent to the text portion corresponding to the first voice input.
[0251] Furthermore, when the user speaks the story of "the evil witch bullying the kind pig" as the second voice input, the image display device 100 can display the text 820 corresponding to the first and second voice inputs on a designated screen 800. At this time, if it is determined that the second voice input corresponds to the text and is not an input about editing the current text, the image display device 100 can display a fifth indicator 862 representing the text on the designated screen 800. The fifth indicator 862 corresponding to the second voice input can be arranged adjacent to the text portion corresponding to the second voice input.
[0252] Reference Figure 28 When the fifth indicator 861 and 862 are displayed on the designated screen 800, and the user says "I want to modify it twice" as the sixth voice input, the remote control device 200 can transmit the sixth voice input received through the microphone to the image display device 100. The image display device 100 can transmit the sixth voice input to the STT server 420 and receive the text corresponding to the sixth voice input. Based on the text corresponding to the sixth voice input, the image display device 100 can output the text 960 corresponding to the first voice input, the second voice input, and the sixth voice input through the designated screen 800.
[0253] The image display device 100 can request intent analysis of the sixth voice input from the server 400. The NLP server 430 can perform intent analysis on the sixth voice input. At this time, the NLP server 430 can determine that the intent of the article corresponding to the fourth voice input is to edit the current text, based on the phrase "want to modify" included in the sixth voice input. In addition, the NLP server can determine the editing type as text modification and identify the keyword that is the object of modification as "twice".
[0254] Reference Figure 29 Based on the intent analysis information received from the server 400, the image display device 100 can determine that the sixth voice input is an input about text editing. At this time, the image display device 100 can delete the text corresponding to "want to modify twice" as the sixth voice input in the designated screen 800.
[0255] The image display device 100 can determine, based on keywords included in intent analysis information received from the server 400, that the fifth indicator 862 corresponding to the second voice input has been selected among the fifth indicators 861 and 862.
[0256] The image display device 100 can display sixth indicators 871 to 876, representing text components that constitute the text corresponding to the second voice input, on a designated screen 800, based on the selection of a fifth indicator 862 corresponding to the second voice input. The sixth indicators 871 to 876 can be respectively arranged adjacent to the text components that constitute the text corresponding to the second voice input.
[0257] When a user selects one of the sixth indicators 871 to 876 using the pointer 205 displayed on the designated screen 800 or by using voice input, the image display device 100 can perform editing on the text component corresponding to the selected sixth indicator. For example, the image display device 100 can output the modified object 841, deleted object 842, etc., to a position adjacent to the text component corresponding to the selected sixth indicator.
[0258] On the other hand, according to one embodiment, the image display device 100 can also perform editing on the article corresponding to the selected fifth indicator 861 or 862 based on the selection of one of the fifth indicators 861 or 862. For example, the image display device 100 can output the modified object 841, deleted object 842, etc., to a position adjacent to the article corresponding to the selected fifth indicator.
[0259] On the other hand, refer to Figure 30The image display device 100 can display text 910 corresponding to the first voice input ("Please make a fairy tale") on a designated screen 800 based on the reception of this first voice input. The image display device 100 can determine whether a predetermined time has elapsed since receiving the first voice input. If a predetermined time has elapsed since receiving the first voice input, the image display device 100 can request recommended keywords related to the first voice input from the server 400.
[0260] NLP server 430 can perform intent analysis on the text corresponding to the first voice input. At this time, NLP server 430 can determine the intent of the article corresponding to the first voice input as the user request category of writing a fairy tale, based on the words "fairy tale" and "make" included in the first voice input.
[0261] AI server 440 can determine recommended keywords based on intent analysis results received from NLP server 430, specifically those related to "fairy tale" included in the first speech input. For example, AI server 440 can use a database corresponding to "fairy tale" among multiple databases 450. In this case, AI server 440 can utilize pre-defined patterns that constitute the text to determine recommended keywords. For example, AI server 440 can determine recommended keywords based on the intent of the text corresponding to the first speech input being writing, specifically those related to the background of "fairy tale," such as "in the jungle," "in the sea," and "in the castle."
[0262] Reference Figure 31 The image display device 100 can output recommended keywords related to the first voice input received from the server 400 through a designated screen 800. At this time, the image display device 100 can output recommended objects 3110, 3120, and 3130 corresponding to the recommended keywords received from the server 400 through a designated area 870 of the designated screen 800. The user can use the pointer 205 displayed on the designated screen 800 to select one of the recommended objects 3110, 3120, and 3130.
[0263] Reference Figure 32 When the user selects the object 3130 corresponding to "in the castle" among the recommended objects 3110, 3120, and 3130, the image display device 100 can display the text 971 corresponding to "in the castle" as the first voice input and recommended keyword on the designated screen 800.
[0264] On the other hand, the image display device 100 can request recommended keywords from the server 400 again regarding the first voice input and "in the castle". The NLP server 430 can perform intent analysis on the text 971 corresponding to the first voice input and "in the castle". At this time, the NLP server 430 can determine that the intent of the article corresponding to the first voice input is that the user requests the category of writing fairy tales, and "in the castle" is an adverbial phrase corresponding to the background of the article.
[0265] Based on the intent analysis results received from the NLP server 430, the AI server 440 can identify related keywords corresponding to "fairy tale" (a keyword included in the first speech input) and "castle" (the background of the article) as recommended keywords. At this time, the AI server 440 can utilize pre-defined patterns that constitute the article to determine the recommended keywords. For example, based on the article category being "fairy tale" and the background being "castle," the AI server 440 can identify related keywords such as "prince," "princess," and "magician" (the subjects associated with the "castle" as the background of the "fairy tale") as recommended keywords.
[0266] The image display device 100 can output recommended keywords related to the first voice input received from the server 400 and the phrase "in the castle" through a designated screen 800. At this time, the image display device 100 can output recommended objects 3210, 3220, and 3230 corresponding to the recommended keywords received from the server 400 through a designated area 870 of the designated screen 800. The user can select one of the recommended objects 3210, 3220, and 3230 using the pointer 205 displayed on the designated screen 800.
[0267] Reference Figure 33 When the user selects the object 3220 corresponding to "Princess" from the recommended objects 3210, 3220, and 3230, the image display device 100 can display the text 972 of "Princess" on the specified screen 800.
[0268] On the other hand, the image display device 100 can request the server 400 again for recommended keywords regarding the first voice input and "in the castle" and "princess". The NLP server 430 can perform intent analysis on the text 972 with the added "princess". At this time, the NLP server 430 can determine that the intent of the article corresponding to the first voice input is that the user requested the category of writing fairy tales, "in the castle" is an adverbial phrase corresponding to the background of the article, and "princess" is the subject of the article.
[0269] Based on the intent analysis results received from the NLP server 430, the AI server 440 can identify related keywords as recommended keywords corresponding to the keywords "fairy tale" included in the first speech input, "castle" as the background of the article, and "princess" as the subject of the article. At this time, the AI server 440 can utilize pre-defined patterns that constitute the article to determine the recommended keywords. For example, based on the article category being "fairy tale," the background being "castle," and the subject being "princess," the AI server 440 can identify related keywords such as "ball," "prince," and "food" as recommended keywords, considering the background of "fairy tale" as "castle" and the object associated with the subject "princess."
[0270] The image display device 100 can output recommended keywords related to the first voice input received from the server 400 and the phrases "in the castle" and "princess" through a designated screen 800. At this time, the image display device 100 can output recommended objects 3310, 3320, and 3330 corresponding to the recommended keywords received from the server 400 through a designated area 870 of the designated screen 800. The user can select one of the recommended objects 3310, 3320, and 3330 using the pointer 205 displayed on the designated screen 800.
[0271] As described above, according to at least one embodiment of the present disclosure, long text can be constructed using a user's voice input.
[0272] Furthermore, according to at least one embodiment of this disclosure, text can be edited using the user's voice input.
[0273] Furthermore, according to at least one embodiment of this disclosure, it is possible to recommend text for use in composing an article, taking into account the context of the text.
[0274] Reference Figures 1 to 33 An image display device 100 according to one aspect of this disclosure includes: a display 180; an external device interface 130 for communicating with a remote control device 200; a network interface 135 for communicating with at least one server 400; and a control unit 170; the control unit 170, based on receiving a first instruction to activate a function for receiving voice input from the remote control device 200, outputs a predetermined screen 800 through the display 180 including a first object 820 corresponding to the end of the function; outputs text corresponding to the voice input received from the remote control device 200 through the predetermined screen 800 based on the voice input received from the remote control device 200; and processes the entire text output through the predetermined screen 800 based on receiving a second instruction to select the first object 820.
[0275] Furthermore, according to one aspect of this disclosure, the control unit 170 can output a second object 830 corresponding to the additional reception of the voice input through the designated screen 800 when the first voice input received from the remote control device 200 corresponds to an article; and when a third instruction to select the second object 830 is received from the remote control device 200, the control unit 170 can additionally receive a second voice input from the remote control device 200.
[0276] Furthermore, according to one aspect of this disclosure, if the control unit 170 receives the first instruction again before receiving the second instruction, it can initialize the entire text output through the specified screen 800.
[0277] Furthermore, according to one aspect of this disclosure, the control unit 170 can, upon receiving a first voice input corresponding to text editing from the remote control device 200, edit the text corresponding to a second voice input that was output through the designated screen 800 before the first voice input was received, using at least one keyword included in the first voice input.
[0278] In the system 10 of this disclosure, the control unit 170 can output text corresponding to the first voice input through the designated screen 800 based on receiving the first voice input from the remote control device 200; if it is determined that the first voice input corresponds to the editing of the text, the text corresponding to the first voice input is deleted in the designated screen 800.
[0279] Furthermore, according to one aspect of this disclosure, the control unit 170 can edit the text corresponding to the second voice input based on the keyword when the keyword is included in the text corresponding to the second voice input; and edit the text corresponding to the second voice input based on the article components corresponding to the keyword when the keyword is not included in the text corresponding to the second voice input.
[0280] Furthermore, according to one aspect of this disclosure, if a predetermined time has elapsed after receiving a predetermined voice input, the control unit 170 may request the server 400 to transmit at least one first keyword associated with the predetermined voice input; output the at least one first keyword received from the server 400 through the predetermined screen 800; and, upon receiving a third instruction from the remote control device 200 to select one of the at least one first keyword, request the server 400 to transmit at least one second keyword associated with the predetermined voice input and the selected first keyword.
[0281] Furthermore, according to one aspect of this disclosure, the control unit 170 can, when the text corresponding to the voice input constitutes an article, output a first indicator representing the article to a position adjacent to the constituted article via the specified screen 800; and edit the specified article corresponding to the selected first indicator based on the input received from the remote control device 200 to select the first indicator.
[0282] Additionally, according to one aspect of this disclosure, the control unit 170 may, based on input received from the remote control device 200 to select the first indicator, output a second indicator representing the text components constituting the prescribed text to a position adjacent to the text components.
[0283] The system 10 of this disclosure includes an image display device 100, a remote control device 200, and at least one server 400. The remote control device 200 can transmit a first instruction to the image display device 100 to activate a function for receiving voice input. When the function is activated, the remote control device 100 transmits voice input received through a microphone to the image display device 100. Based on receiving the first instruction from the remote control device 200, the image display device 100 outputs a predetermined screen 800 including a first object 820 corresponding to the end of the function. The remote control device 200 transmits the voice input received from the remote control device 200 to the server 400. The predetermined screen 800 outputs text corresponding to the voice input received from the server 400. Based on receiving a second instruction to select the first object 820, the system processes the entire text output through the predetermined screen 800.
[0284] Additionally, according to one aspect of this disclosure, the server 400 may include: a first server 420 for converting speech to text; a second server 430 for processing text; and a third server 410 for transmitting data between the first server 420, the second server 430, and the image display device 100; wherein, if one of the first server 420 and the third server 410 transmits text corresponding to the speech input to the second server 430 when the text corresponding to the speech input received from the image display device 100 corresponds to an article; and the second server 430 performs intent analysis on the text corresponding to the speech input.
[0285] Furthermore, according to one aspect of this disclosure, when the text corresponding to the voice input does not correspond to text editing, the second server 430 transmits keywords included in the voice input and text components corresponding to the keywords to the image display device 100; when the text corresponding to the voice input corresponds to text editing, it transmits the editing type corresponding to the voice input and editing-related keywords to the image display device 100; while receiving the first voice input from the remote control device 200, the image display device 100 outputs text corresponding to the first voice input through the designated screen 800; when the first voice input corresponds to text editing, the text corresponding to the first voice input is deleted in the designated screen 800; based on the editing type corresponding to the first voice input, and using the editing-related keywords, the text corresponding to the second voice input that was output through the designated screen 800 before receiving the first voice input is edited.
[0286] Furthermore, according to one aspect of this disclosure, if a predetermined time has elapsed after receiving a predetermined voice input, the image display device 100 requests the server 400 to transmit at least one first keyword associated with the predetermined voice input; the predetermined screen 800 outputs at least one first keyword received from the server 400; and upon receiving a third instruction from the remote control device 200 to select one of the at least one first keyword, the device requests the server 400 to transmit at least one second keyword associated with the predetermined voice input and the selected first keyword.
[0287] Additionally, according to one aspect of this disclosure, the server 400 can determine a plurality of databases corresponding to the keywords included in the first voice input; determine at least one first keyword based on data about a first article component in the databases; and determine at least one second keyword based on data about a second article component different from the first article component in the databases.
[0288] Furthermore, according to one aspect of this disclosure, the image display device 100 can, when the text corresponding to the voice input constitutes an article, output a first indicator representing the article to a position adjacent to the constituted article via the specified screen 800; and edit the specified article corresponding to the selected first indicator based on the input received to select the first indicator.
[0289] The accompanying drawings are only for the purpose of understanding the embodiments disclosed in this specification and should not be construed as limiting the technical ideas disclosed in this specification to the drawings. They should be understood to include all variations, equivalents and substitutions within the spirit and scope of this disclosure.
[0290] Furthermore, the operating methods of this disclosure can be implemented as processor-readable code on a processor-readable recording medium. Processor-readable recording media include all types of recording devices that store data readable by a processor. Examples of processor-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc., and also include those implemented in the form of carrier waveforms, such as those transmitted via the Internet. In addition, processor-readable recording media can be distributed across computer systems connected via a network, and processor-readable code can be stored and executed in a distributed manner.
[0291] Furthermore, although preferred embodiments of the present disclosure have been shown and described above, the present disclosure is not limited to the specific embodiments described above, and it is natural that those skilled in the art can make various modified embodiments without departing from the spirit of the present disclosure as claimed in the claims, and these modified embodiments should not be understood solely from the technical ideas or prospects of the present disclosure.
Claims
1. An image display device, characterized in that, include: monitor; External device interface section, for communication with remote control device; The network interface unit communicates with at least one server. as well as Control Department; Based on the action of receiving a first instruction to activate the function of receiving voice input from the remote control device, the control unit outputs a predetermined screen through the display, including a first object corresponding to the end of the function. Based on the action of receiving voice input from the remote control device, the control unit outputs text corresponding to the voice input received from the remote control device through the specified screen; The control unit processes the entire text output through the specified screen based on the action of receiving a second instruction to select the first object.
2. The image display device according to claim 1, characterized in that, If the first voice input received by the control unit from the remote control device corresponds to an article, the control unit outputs a second object corresponding to the additional reception of the voice input through the specified screen. If the control unit receives a third instruction from the remote control device to select the second object, it shall additionally receive a second voice input from the remote control device.
3. The image display device according to claim 1, characterized in that, If the control unit receives the first instruction again before receiving the second instruction, it initializes the entire text output through the specified screen.
4. The image display device according to claim 1, characterized in that, If the control unit receives a first voice input corresponding to text editing from the remote control device, it uses at least one keyword included in the first voice input to edit the text corresponding to the second voice input that was output through the predetermined screen before the first voice input was received.
5. Based on the action of receiving the first voice input from the remote control device, the control unit outputs text corresponding to the first voice input through the specified screen; If the control unit determines that the first voice input corresponds to the editing of the text, it deletes the text corresponding to the first voice input in the specified screen.
6. The image display device according to claim 4, characterized in that, If the keyword is included in the text corresponding to the second voice input, the control unit edits the text corresponding to the second voice input based on the keyword; If the keyword is not included in the text corresponding to the second voice input, the control unit edits the text corresponding to the second voice input based on the article components corresponding to the keyword.
7. The image display device according to claim 1, characterized in that, If a specified time has elapsed after receiving the specified voice input, the control unit requests the server to transmit at least one first keyword associated with the specified voice input; The control unit outputs at least one first keyword received from the server through the specified screen; If the control unit receives a third instruction from the remote control device to select at least one of the first keywords, it requests the server to transmit at least one second keyword associated with the specified voice input and the selected first keyword.
8. The image display device according to claim 1, characterized in that, If the control unit forms an article with the text corresponding to the voice input, it outputs a first indicator representing the article to a position adjacent to the article through the specified screen. The control unit edits a specified text corresponding to the selected first indicator based on the action of receiving input from the remote control device to select the first indicator.
9. The image display device according to claim 8, characterized in that, Based on the action of receiving input from the remote control device to select the first indicator, the control unit outputs a second indicator representing the text components that constitute the specified article to a position adjacent to the text components.
10. A system comprising an image display device, a remote control device, and at least one server, characterized in that, The remote control device transmits a first instruction to the image display device to activate the function of receiving voice input. When the function is activated, the remote control device will transmit the voice input received through the microphone to the image display device; Based on the action of receiving the first instruction from the remote control device, the image display device outputs a specified screen including a first object corresponding to the end of the function; The image display device transmits the voice input received from the remote control device to the server; The image display device outputs text corresponding to the voice input received from the server through the specified screen; Based on the action of receiving the second instruction to select the first object, the entire text output through the specified screen is processed.
11. The system according to claim 10, characterized in that, The server includes: The first server converts speech to text; The second server processes text; and A third server transmits data between the first server, the second server, and the image display device; If the text corresponding to the voice input received from the image display device corresponds to an article, then one of the first server and the third server will transmit the text corresponding to the voice input to the second server. The second server performs intent analysis on the text corresponding to the voice input.
12. The system according to claim 11, characterized in that, If the text corresponding to the voice input does not correspond to the text editing, the second server will transmit the keywords included in the voice input and the text components corresponding to the keywords to the image display device. If the text corresponding to the voice input corresponds to the text editing, then the second server transmits the editing type and editing-related keywords corresponding to the voice input to the image display device; During the period when the image display device receives the first voice input from the remote control device, the image display device outputs text corresponding to the first voice input through the specified screen. If the first voice input corresponds to the editing of the text, the image display device deletes the text corresponding to the first voice input from the specified screen; The image display device, based on the editing type corresponding to the first voice input, uses the editing-related keywords to edit the text corresponding to the second voice input that was output through the specified screen before receiving the first voice input.
13. The system according to claim 10, characterized in that, If a specified time has elapsed after receiving the specified voice input, the image display device requests the server to transmit at least one first keyword associated with the specified voice input; The image display device outputs at least one of the first keywords received from the server through the specified screen; If the image display device receives a third instruction from the remote control device to select at least one of the first keywords, it requests the server to transmit at least one second keyword associated with the specified voice input and the selected first keyword.
14. The system according to claim 13, characterized in that, The server determines a plurality of databases corresponding to the keywords included in the first voice input; The server determines at least one of the first keywords based on data about the first article components in the database; The server determines at least one second keyword based on data in the database regarding second article components that are different from the first article component.
15. The system according to claim 10, characterized in that, If the text corresponding to the voice input constitutes an article, the image display device outputs a first indicator representing the article to a position adjacent to the constituted article through the specified screen; The image display device edits a prescribed text corresponding to the selected first indicator based on the action of receiving input that selects the first indicator.