Display apparatus and control method thereof

By recognizing the user's voice and determining the preset language in the display device, text objects in different languages ​​are displayed, thus solving the voice control limitations when the system language and the display device language are different, and realizing the convenience of cross-language voice operation.

CN116072115BActive Publication Date: 2026-02-10SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310065267.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-07-19
Filing Date
2018-05-11
Publication Date
2026-02-10
Estimated Expiration
2038-05-11

AI Technical Summary

Technical Problem

When the system language and the language on the display device are different, voice control is limited and cannot effectively recognize and operate hyperlink text.

Method used

A display device and its control method are provided, which recognizes user speech through a processor, determines a preset language, displays text objects in different languages, and performs corresponding operations based on the speech recognition results. It supports speech recognition in multiple languages, including communication with external devices and collaboration with servers.

Benefits of technology

It enables cross-language voice control, accurately recognizes and manipulates text objects on the display device, and improves user experience and ease of operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116072115B_ABST
    Figure CN116072115B_ABST
Patent Text Reader

Abstract

A display device is provided. The display device according to an embodiment includes a display; and a processor configured to control the display to display a UI screen including a plurality of text objects, control the display to display a text object in a language different from a preset language among the plurality of text objects and a preset number, and in response to a recognition result of a displayed number included in a voice spoken by a user, perform an operation related to a text object corresponding to the displayed number.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a continuation of the patent application with the application date of May 11, 2018, and the application number of 201810467513.1.

[0002] Cross Reference to Related Applications

[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 505,363, filed May 12, 2017, in the U.S. Patent and Trademark Office, and Korean Patent Application No. 10-2017-0091494, filed July 19, 2017, in the Korean Intellectual Property Office, the disclosures of which are incorporated herein in their entireties by reference. TECHNICAL FIELD

[0004] Apparatuses and methods consistent with embodiments of the present application relate to a display apparatus and a control method thereof, and more particularly, to a display apparatus supporting voice recognition of content in various languages and a control method thereof. BACKGROUND

[0005] With the development of electronic technology, various types of display apparatuses have been developed. In particular, various electronic devices such as televisions, mobile phones, personal computers, notebook computers, laptop computers, tablet computers, smart phones, and personal digital assistants are widely used.

[0006] Recently, voice recognition technology has been developed to more conveniently and intuitively control a display apparatus.

[0007] Conventionally, a display apparatus controlled by a user's voice performs voice recognition by using a voice recognition engine. However, the voice recognition engine varies according to a language used, and thus a voice recognition engine for use can be determined in advance. Generally, a system language of the display apparatus is determined as a language to be used for voice recognition.

[0008] However, assuming that English is used in hyperlink text displayed on the display apparatus and Korean is used as the system language of the display apparatus, even if the user utters a voice corresponding to the hyperlink text, the voice is converted into Korean text by a Korean voice recognition engine. Thus, there is a problem in that the hyperlink text cannot be selected.

[0009] Therefore, when the system language is different from the language on the display apparatus, there is a limitation in controlling the display apparatus by voice. SUMMARY

[0010] Aspects of the example embodiments relate to a display apparatus providing voice recognition control for content in various languages and a control method thereof.

[0011] According to an aspect of an exemplary embodiment, there is provided a display device including a display; and a processor configured to control the display to display a user interface including a plurality of text objects, control the display to display a text object among the plurality of text objects in a language different from a preset language and a preset symbol, and in response to a recognition result of a symbol included in a voice spoken by a user, perform an operation related to a text object corresponding to the symbol.

[0012] The processor is further configured to set a language set in a setting menu of the display device as the preset language, or set a language most frequently used with respect to the plurality of text objects as the preset language.

[0013] The user interface can be a web page, and the processor can be further configured to set a language corresponding to language information of the web page as the preset language.

[0014] The processor can be further configured to determine a text object having at least two languages among the plurality of text objects as a text object in a language different from the preset language based on a ratio of the at least two languages.

[0015] The processor can be further configured to control the display to display the symbol adjacent to a text object corresponding to the symbol.

[0016] The display device can further include a communicator, and the processor can be further configured to control the display to display the symbol while the communicator receives a signal corresponding to selection of a specific button of an external device.

[0017] The external device can include a microphone, the communicator can be configured to receive a voice signal corresponding to a voice input through the microphone of the external device, and the processor can be further configured to, in response to a recognition result of the received voice signal including the symbol, perform an operation related to a text object corresponding to the symbol.

[0018] The processor can be further configured to, in response to a recognition result of the received voice signal including a text corresponding to one of the plurality of text objects, perform an operation related to the text object.

[0019] The operation related to the text object can include an operation of displaying a web page having a URL address corresponding to the text object or executing an application corresponding to the text object.

[0020] The plurality of text objects can be included in an execution screen of a first application, and the processor can be further configured to, while displaying the execution screen of the first application, in response to determining that an object corresponding to a recognition result of a voice spoken by the user is not included in the execution screen of the first application, execute a second application different from the first application and perform an operation corresponding to the voice recognition result.

[0021] The second application can provide a search result of a search word, and the processor can be further configured to, while displaying the execution screen of the first application, in response to determining that the object corresponding to the recognition result of the voice spoken by the user is not included in the execution screen of the first application, execute the second application and provide a search result using a text corresponding to the voice recognition result as a search word.

[0022] The display device can further include a communicator configured to perform communication with a server that performs voice recognition of a plurality of different languages, and the processor can be further configured to control the communicator to provide the server with a voice signal corresponding to a voice spoken by the user and information about the preset language, and in response to a voice recognition result including a displayed number received from the server, perform an operation related to a text object corresponding to the symbol.

[0023] The processor can be further configured to, in response to the voice recognition result including a text corresponding to one of the plurality of text objects received from the server, perform an operation related to the text object.

[0024] According to an aspect of an exemplary embodiment, there is provided a control method for a display device, the method including displaying a user interface including a plurality of text objects, displaying a text object in a language different from a preset language and a symbol, and in response to a recognition result including the symbol of a voice spoken by a user, performing an operation related to a text object corresponding to the symbol.

[0025] The method can further include setting a language set in a setting menu of the display device as the preset language, or setting a language most frequently used with respect to the plurality of text objects as the preset language.

[0026] The plurality of text objects are included in a web page and for the control method of the display device can further include setting a language corresponding to language information of the web page as the preset language.

[0027] The method can further include determining a text object among the plurality of text objects in which at least two languages are expressed as a text object in which a language different from the preset language is expressed based on a ratio of the at least two languages.

[0028] The displaying of the text object and the displayed number can include displaying the symbol adjacent to the text object corresponding to the symbol.

[0029] The displaying of the text object and the displayed number can include controlling the display to display the symbol at the same time as receiving a signal corresponding to selection of a specific button of the external device from the external device.

[0030] Performing the operation related to the text object can include displaying a web page having a URL address corresponding to the text object and executing an application corresponding to the text object.

[0031] The plurality of text objects can be included in an execution screen of a first application, and the method can further include, while displaying the execution screen of the first application, in response to determining that an object corresponding to a recognition result of a voice spoken by the user is not included in the execution screen of the first application, executing a second application different from the first application and performing an operation corresponding to the voice recognition result.

[0032] The method can further include providing information on a voice signal corresponding to the voice spoken by the user and a preset language to a server configured to perform voice recognition in a plurality of different languages, and performing an operation related to the text object can include, in response to the voice recognition result including the displayed number received through the server, performing an operation related to a text object corresponding to the displayed number.

[0033] According to an aspect of an exemplary embodiment, there is provided a non-transitory computer readable medium having embodied thereon a program for executing a method of controlling a display device, the method can include: controlling the display device to display a user interface including a plurality of text objects; having displayed a text object in which a language different from a preset language is expressed and a preset number; and in response to a recognition result including a symbol of a voice spoken by a user, performing an operation related to a text object corresponding to the symbol. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 and Figure 2 are views illustrating a method for inputting a voice command to a display device according to an exemplary embodiment of the present disclosure;

[0035] Figure 3is a view illustrating a voice recognition system according to an exemplary embodiment of the present disclosure;

[0036] Figure 4 is a block diagram illustrating a configuration of a display apparatus according to an exemplary embodiment of the present disclosure;

[0037] Figure 5 , Figure 6 and Figure 7 is a view illustrating displaying a number for selecting an object according to an exemplary embodiment of the present disclosure;

[0038] Figure 8 and Figure 9 is a view illustrating a voice search method according to an exemplary embodiment of the present disclosure;

[0039] Figure 10 is a block diagram of a display apparatus according to an exemplary embodiment of the present disclosure; and

[0040] Figure 11 is a flowchart of a method of controlling a display apparatus according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] Before describing the present disclosure in detail, a method of depicting the specification and drawings will be described.

[0042] All terms used in the present specification, including technical or scientific terms, have the same meanings that are commonly understood by those skilled in the art. However, the terms can be interpreted based on the intent of the skilled in the art, the legal or technical interpretation, and the new technology development, and additionally, some terms can be arbitrarily selected by the present applicant. These terms can be explained based on the context of the present specification and the knowledge of those skilled in the art, and unless otherwise denoted, can be interpreted based on the entire content of the present specification and the knowledge of those skilled in the art.

[0043] Terms such as "first", "second", etc. can be used to describe various elements, but the elements should not be limited by the terms. The terms are used only to distinguish one element from the other elements. The use of such ordinal numbers should not be construed as limiting the meaning of the terms. For example, components associated with such ordinal numbers should not be limited by the order of use, the order of placement, etc. Each ordinal number can be used interchangeably if necessary.

[0044] The terms used in the present application are used only to describe particular exemplary embodiments, and are not intended to be limiting. The singular forms are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms such as "including" or "having," etc., are intended to indicate the existence of the disclosed features, numbers, operations, actions, components, parts, or combinations thereof listed in the specification, and are not intended to preclude the possibility that one or more other features, numbers, operations, actions, components, parts, or combinations thereof can exist or can be added.

[0045] In exemplary embodiments, a "module", "unit", or "component" configured to perform at least one function or operation can be implemented as hardware such as a processor or an integrated circuit, software stored in a memory, downloaded from a memory and executed by a processor reading from the memory, or a combination thereof. In addition, a plurality of "modules", "units", or "components" can be integrated into at least one module or chip, and can be implemented as at least one processor other than the "module", "unit", or "component" that should be implemented in a specific hardware.

[0046] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0047] Figure 1 is a view showing a display device controlled by voice recognition according to an exemplary embodiment of the present disclosure.

[0048] Referring to Figure 1 , the display device 100 can be a television (TV) as shown in Figure 1 , but is not limited thereto. The display device 100 can be embodied as any kind of apparatus capable of displaying information and images, such as a smart phone, a desktop PC, a notebook computer or a tablet computer, a smart watch or other user peripheral device, a navigation device, a refrigerator or a home appliance, etc.

[0049] The display device 100 can perform an operation or execute a command based on a recognition result of a voice spoken by a user. For example, when a user says "switch to channel 7", the display device 100 can tune to channel 7 and display a program on channel 7, and when a user says "turn off power", power of the display device 100 can be turned off.

[0050] Accordingly, the user can perceive that the display apparatus 100 can operate like the display apparatus communicates with the user. For example, when the user asks "What is the name of the broadcast program?", the display apparatus can output a response message "The name of the broadcast program is xxx" through voice or text. When the user asks "How is the weather today?" through voice, the display apparatus can output a message "Please tell me the location where you want to know the temperature" through voice or text, and in response thereto, when the user answers "Seoul", the display apparatus 100 can output a message "The temperature of Seoul is xxx" through voice or text.

[0051] As shown in Figure 1 , the display apparatus 100 can receive a user voice through a microphone connected to or attached to the display apparatus 100. The display apparatus 100 can receive a voice signal corresponding to a voice received through a microphone of an external apparatus such as a PC or a smart phone from the external apparatus. A detailed description of this point will be made with reference to Figure 2 .

[0052] Figure 2 is a view showing a display system according to an exemplary embodiment of the present disclosure.

[0053] Referring to Figure 2 , the display system can include the display apparatus 100 and the external apparatus 200.

[0054] As described above, Figure 1 , the display apparatus 100 can operate according to a voice recognition result.

[0055] Figure 2 An instance in which the external apparatus 200 is embodied as a remote controller is shown, although the external apparatus 200 can be embodied as an electronic apparatus such as a smart phone, a tablet PC, a smart watch, etc.

[0056] The external apparatus 200 can include a microphone and transmit a signal corresponding to a voice input through the microphone to the display apparatus 100. The signal can correspond to a voice of the user or text corresponding to the voice of the user converted into text by the external apparatus 200. For example, the external apparatus 200 can transmit a voice signal to the display apparatus 100 using a wireless communication method such as infrared (IR), RF, Bluetooth, WiFi, etc.

[0057] The external apparatus 200 can be enabled when a predetermined event occurs, thereby saving power. For example, when the microphone button 210 of the external apparatus 200 is pressed, the microphone can be enabled, and when the microphone button 210 is released, the microphone can be disabled. In other words, only when the microphone button 210 is pressed, the microphone can receive a voice.

[0058] The external server can perform recognition of a voice received through a microphone of the display device 100 or a microphone of the external device 200.

[0059] Figure 3 is a view illustrating a voice recognition system according to an exemplary embodiment of the disclosure.

[0060] Referring to Figure 3 , the voice recognition system 200 can include the display device 100 and the server 300. As described with respect to Figure 2 , the system can further include the external device 200.

[0061] The display device 100 can operate according to a voice recognition result as Figure 1 described. The display device 100 and / or the external device 200 can transmit a voice signal corresponding to a voice input through a microphone of the display device 100 or a microphone of the external device 200 to the server 300.

[0062] The display device 100 can transmit information indicating which language to recognize the voice signal based on (hereinafter referred to as "language information") to the server 300 along with the voice signal. Although the same voice signal is input, a voice recognition result can vary according to which language voice recognition engine is used.

[0063] The server 300 can perform voice recognition in a plurality of different languages. The server 300 can include various voice recognition engines corresponding to respective languages. For example, the server 300 can include a Korean voice recognition engine, an English voice recognition engine, a Japanese voice recognition engine, etc. The server 300 can perform voice recognition by using a voice recognition engine corresponding to a voice signal and language information in response to the voice signal and the language information received from the display device 100.

[0064] The server 300 can transmit a voice recognition result to the display device 100, and the display device 100 can perform an operation corresponding to the voice recognition result received from the server 300.

[0065] For example, when text included in the voice recognition result received from the server 300 corresponds to a text object included in the display device 100, the display device 100 can perform an operation related to the text object. For example, when the text included in the voice recognition result corresponds to a text object in a web page, the display device 100 can display a web page having a URL address corresponding to the text object. However, the disclosure is not limited thereto, but can select a user interface (UI) object provided by various applications of the display device 100 through voice recognition, and can perform a corresponding operation.

[0066] The server 300 can be embodied as one server, but the server 300 can be embodied as a plurality of servers corresponding to a plurality of languages, respectively. For example, a Korean voice recognition server and an English voice recognition server can be provided separately.

[0067] In the described example, the voice recognition can be performed by the server 300 which is separate from the display device 100, but according to another embodiment, the display device 100 or the external device 200 can function as the server 300. In other words, the display device 100 or the external device 200 can be implemented integrally with the server 300.

[0068] Figure 4 is a block diagram illustrating a display device according to an exemplary embodiment of the present disclosure.

[0069] The display device 100 can include a display 110 and a processor 120.

[0070] The display 110 can be implemented as a liquid crystal display (LCD), such as a cathode ray tube (CRT), a plasma display panel (PDP), an organic light emitting diode (OLED), a transparent OLED (TOLED), etc. In addition, the display 110 can be implemented as a touch screen capable of sensing a touch operation of a user.

[0071] The processor 120 can control the overall operation of the display device 100.

[0072] For example, the processor 120 can be a central processing unit (CPU) or a microprocessor which communicates with a RAM, a ROM, and a system bus. The ROM can store a command set for system booting. The CPU can copy an operating system stored in a storage of the display device 100 to the RAM according to a command stored in the ROM, execute the operating system and perform system booting. When the booting is completed, the CPU can copy various applications stored in the storage to the RAM, execute the applications and perform various operations. Although the processor 120 is described as including only one CPU in the foregoing description, the processor 120 can be embodied as a plurality of CPUs (or DSPs, SoCs, etc.) or processor cores.

[0073] In response to receiving a user command for selecting an object displayed on the display 110, the processor 120 can perform an operation related to the object selected by the user command. The object can be any one of selectable objects, such as a hyperlink or an icon. The operation related to the selected object can be, for example, an operation of displaying a page, a document, an image, etc. connected to the hyperlink, or an operation of executing a program corresponding to the icon.

[0074] User commands for selecting objects can be commands entered via various input devices (e.g., mouse, keyboard, touchpad, etc.) connected to the display device 100, or voice commands corresponding to voice spoken by the user.

[0075] Although Figure 4 Although not shown, display device 100 may also include a voice receiver for receiving user voice. The voice receiver can directly receive user voice and generate a voice signal via a microphone, or receive an electronic voice signal from external device 200. When the voice receiver receives an electronic voice signal from external device 200, it can be embodied as a communicator for performing wired / wireless communication with external device 200. The voice receiver may not be included in display device 100. For example, the voice signal corresponding to voice input via the microphone of external device 200 can be transmitted to server 300 via another device other than display device 100, or it can be transmitted directly from external device 200 to server 300. In this case, display device 100 may only receive the voice recognition result from server 300.

[0076] The processor 120 can control the display 110 to display text objects and numbers in a language different from the preset language in the text objects displayed on the display 110.

[0077] The preset language can refer to the basic language used for speech recognition (the language of the speech recognition engine to be used for speech recognition). The preset language can be set manually by the user or automatically. When the preset language is set manually by the user, for example, the language (or system language) that is set to be used in the settings menu of the display device 100 can be set to the basic language used for speech recognition.

[0078] When a preset language is automatically set, the processor 120 can identify the language primarily used for text objects displayed on the display 110 and set the language as the base language for speech recognition.

[0079] Specifically, the processor 120 can analyze the type of characters (e.g., Korean or letters) contained in each of a plurality of text objects displayed on the display 110, and set the language of the characters primarily used for the plurality of text objects as the base language for speech recognition.

[0080] According to another embodiment, when the text object displayed on display 110 is included in a webpage, processor 120 can set the language corresponding to the language information of the webpage as the base language for speech recognition. The language information of the webpage can be confirmed through the lang attribute of HTML (e.g., ...).

[0081] When a base language is set for speech recognition, processor 120 can control display 110 to display text objects and preset numbers in a language different from the base language. The user can select a text object by speaking the preset number displayed on display 110. Additionally, since images may not be selected by voice, processor 120 can control display 110 to display image objects and preset numbers.

[0082] The processor 120 can identify text objects presented in a language other than the primary language used for speech recognition as text objects presented in a language different from the primary language used for speech recognition. If the ratio of preset languages ​​is less than a predetermined ratio, the processor 120 can identify text objects presented in at least two languages ​​as text objects presented in a language different from the primary language used for speech recognition.

[0083] Figure 5 It is a view showing the screen displayed on the display device.

[0084] See Figure 5 A UI screen including multiple text objects 51 to 59 can be displayed on the display 110. When the primary language used for speech recognition is English, the processor 120 can control the display to show text objects 51 to 56 and preset numbers ① to ⑥ in a language other than English. Preset numbers ① to ⑥ can be displayed adjacent to their corresponding text objects 51 to 56. Text objects 51 and 58 in English can be displayed with specific icons 57a and 58a to inform the user that text objects 51 and 58 can be selected by speaking the text included in them. Figure 5 As shown, icons 57a and 58a can be represented by "T", but are not limited to this, and can be represented in various forms such as "text".

[0085] Regarding text objects 59 presented in at least two languages, processor 120 can determine whether the English ratio is greater than a predetermined ratio (e.g., 50%), and if the ratio is less than the predetermined ratio, control the display to show text objects 59 presented in at least two languages ​​as well as numbers. Figure 5 The text object 59 can be in Korean and English, but because the proportion of English is greater than a predetermined proportion (e.g., 50%), it may not be displayed together with the numbers. Instead, by saying the text included in the text object, an icon 59a indicating that the text object is selectable can be displayed adjacent to the text object 59.

[0086] See Figure 5The numbers are shown in a form such as “①”, but the form of the numbers is not limited. For example, a square or circle may surround the number “1”, or the number may simply be represented by “1”. According to another embodiment of this disclosure, the numbers may be expressed by words of the base language used for speech recognition. If the base language used for speech recognition is English, the number may be represented by “one”, or if the language is Spanish, the number may be represented by “uno”.

[0087] Despite Figure 5 Not shown, but may be further displayed on display 100 along with phrases that encourage the user to say the number, such as “You can select an object corresponding to the number”.

[0088] According to another exemplary embodiment, if the first word of a text object presented in at least two languages ​​is different from the language used for speech recognition, the processor 120 can determine that the text object is different from the text object presented in the base language used for speech recognition.

[0089] Figure 6 It is a view showing the screen displayed on the monitor.

[0090] See Figure 6 The processor 120 can display a UI screen including multiple text objects 61 to 63 on the display 110. When the language to be used for speech recognition is Korean, the processor 120 can determine that the text object 61, which is presented in at least two languages, is presented in a language different from the base language used for speech recognition, because the first word "AAA" of the text object 61 is English, not Korean, which is the base language used for speech recognition. Therefore, the processor 120 can control the display 110 to display the text object 61 and the number ①.

[0091] According to the reference Figure 6 In one exemplary implementation, even if the ratio of the base language used for speech recognition in a text object presented in at least two languages ​​is greater than a predetermined ratio, a number may be displayed if the first word of the text object is not in the base language used for speech recognition. Conversely, even if the ratio of the base language used for speech recognition in a text object presented in at least two languages ​​is less than a predetermined ratio, a number may not be displayed if the first word of the text object is in the base language used for speech recognition. This is because a user might speak the first word of the text object to select it.

[0092] According to another exemplary embodiment, image objects can be selected without voice. Therefore, numbers can be displayed alongside image objects.

[0093] Figure 7 It is a view showing the screen displayed on the monitor.

[0094] See Figure 7 The processor 120 can display a first image object 71, a second image object 72, a third image object 74, a first text object 73, and a second text object 75 on the display 110. The processor 120 can control the display 110 to display image object 71 and the number ①.

[0095] According to another exemplary embodiment, when multiple objects displayed on display 110 each have a URL link, processor 120 can compare the URL links of the multiple objects. If an object with the same URL link cannot be selected by voice recognition, processor 120 can control display 110 to display a number and one of the multiple objects, and if any of the multiple objects can be selected by voice recognition, processor 120 can control display 110 not to display a number.

[0096] Specifically, when multiple objects with the same URL link that cannot be selected via voice recognition are displayed on monitor 110 (i.e., text objects or image objects presented in a language different from the underlying language used for voice recognition), a number can be displayed near one of the multiple objects. See also Figure 7 The second image object 72 cannot be selected by voice, and the first text object 73 can be a language different from Korean, which is the basic language used for speech recognition. Therefore, since neither the second image object 72 nor the first text object 73 can be selected by voice, but both connect to the same URL link when selected, the number ② can be displayed near either the second image object 72 or the first text object 73. This is to reduce the number of numbers displayed on the display 110.

[0097] To reduce the number of numbers displayed on display 110, according to another exemplary embodiment, multiple objects with the same URL address can be displayed on display 110, and if any of the multiple objects is a text object in the basic language, the numbers may not be displayed. See also Figure 7 The processor 120 can compare the URL address of the third image object 74 with the URL address of the second text object 75, and if it is determined that the URL address of the third image object 74 is the same as the URL address of the second text object 75, and the second text object 75 is a text object presented as Korean, the basic language for speech recognition, then the processor 120 can control the display 110 not to display numbers near the third image object 74.

[0098] If the recognition result of the user's spoken speech includes specific text displayed on display 110, then processor 120 can perform operations related to the text object corresponding to said text. See also Figure 5If the user says “voice recognition”, the processor 120 can control the display 110 to display a page with a URL address corresponding to the text object 59.

[0099] According to an exemplary embodiment, when the recognition result of the user's spoken voice includes text included in at least two of a plurality of text objects typically displayed on display 110, processor 120 may display numbers near each text object, and when the user speaks the displayed numbers, perform operations related to the text objects corresponding to the numbers.

[0100] See Figure 5 When the speech recognition result of the user's spoken words includes the text "speech recognition," the processor 120 can search for text objects containing the phrase "speech recognition" from the displayed text objects. When multiple text objects 57 and 58 are found, the processor 120 can control the display 110 to display preset numbers near each of the text objects 57 and 58. For example, when the number 7 is displayed near text object 57 and the number 8 is displayed near text object 58, the user can select text object 57 by speaking the number "7." When the speech recognition result includes numbers displayed on the display 110, the processor 120 can perform operations related to the text object or image object corresponding to the number.

[0101] See Figure 6 If the user says "one", the processor 120 can control the display 110 to display a page with a URL address corresponding to the text object 61.

[0102] The user's spoken voice can be input via the microphone of display device 100 or the microphone of external device 200. When the user's voice is input via the microphone of external device 200, display device 100 may include a communicator for communicating with external device 200, which includes the microphone, and the communicator can receive a voice signal corresponding to the voice input via the microphone of external device 200. If the recognition result of the voice signal received from external device 200 via the communicator includes a number displayed on display 110, processor 120 can perform operations related to a text object corresponding to said number. See also Figure 6 When a user speaks "a" through the microphone of external device 200, external device 200 can transmit the voice signal to display device 100, and processor 120 can control display 110 to display a page with a URL address corresponding to text object 61 based on the voice recognition result of the received voice signal.

[0103] Numbers corresponding to text or image objects can be displayed during a predetermined time period. According to an exemplary embodiment, when a signal corresponding to the selection of a specific button is received from the external device 200, the processor 120 can control the display 110 to display numbers. In other words, numbers can be displayed only when the user presses a specific button on the external device 200. The specific button may be, for example… Figure 2 The external device 200 described herein includes a microphone button 210.

[0104] According to another exemplary embodiment, if the voice input through the microphone of display device 100 includes a predetermined keyword (e.g., “Hi TV”), processor 120 can control display 110 to display numbers, and if a predetermined time period elapses in response to no further voice input through the microphone of display device 100, the displayed numbers are removed.

[0105] The above implementation describes displaying numbers, but the indicators do not have to be numbers; they can be anything the user can see and read (meaningful or meaningless words). For example, a, b, and c... could be displayed instead of 1, 2, and 3. Alternatively, any other symbols can be used.

[0106] According to another exemplary embodiment, when a webpage displayed on monitor 110 includes a search window, the user can easily perform a search by speaking the word to be searched or specific keywords used to perform the search function. For example, when a webpage displayed on monitor 110 includes a search window, the user can speak "xxx search", "search xxx", etc., to display the search results for "xxx" on monitor 110.

[0107] To this end, processor 120 can detect a search term input window from the webpage displayed on monitor 110. Specifically, processor 120 can search for input objects from the objects on the webpage displayed on monitor 110. Input tags on HTML can be input objects. Input tags can have various attributes, but the type attribute can explicitly define the input characteristics. When the type is "search", the object can correspond to the search term input window.

[0108] However, when the object's type is "text," it cannot be immediately determined whether the object is a search term input window. Since typical input objects are of text type, it is difficult to determine whether the object is a search term input window or a typical input window. Therefore, further processing is required to determine whether the object is a search term input window.

[0109] When the object's type is "text," information about the object's additional attributes can be referenced to determine if the object is a search term input window. The object can be identified as a search term input window when its title or area label includes the keyword "search."

[0110] Processor 120 can determine whether the recognition result of the user's spoken speech includes a specific keyword. The specific keyword could be "search," "retrieve," etc. In response to determining that the specific keyword is included, processor 120 can confirm the position of the specific keyword to more clearly determine the user's intent. If at least one word appears before or after the specific keyword, the user is likely to search for that at least one word. If the speech recognition result only includes specific words such as "search" or "retrieve," the user is less likely to search for said word.

[0111] The process of determining the user's intent can be performed by the display device 100 or by the server 300, and the result can be provided to the display device 100.

[0112] If the user's search intent is determined, the processor 120 can set words (other than specific keywords) as search terms, input the set search terms into the search term input window detected by the above processing, and perform a search. For example, as Figure 8 As shown, if a webpage including the search term input window 810 is displayed on the monitor 110, the processor 120 can detect the search term input window 810, and if the user says "search for puppy" by voice, the processor 120 can set "puppy" as the search term in the speech recognition results of the spoken speech, input the search term into the search term input window 810 and perform the search.

[0113] The search term input window from the webpage displayed on the monitor 110 can be detected before or after determining that the speech recognition result includes specific keywords.

[0114] Figure 9 This is a view illustrating methods for entering search terms. For example, the methods may include a method for searching multiple search term input windows on a webpage.

[0115] See Figure 9A webpage can have two search term input windows. The first search term input window 910 can be used for news searches, and the second search term input window 920 can be used for stock information searches. The processor 120 can perform a search using the search term input window that is displayed when the user speaks a phrase including the search term, based on information about the object's location and the screen layout. For example, when the first search term input window 910 is displayed on the monitor 110 and the user speaks a phrase including the search term and specific keywords, the processor 120 can input the search term into the first search term input window 910, and after the screen scrolls, when the second search term input window 920 is displayed on the monitor 110 and the user speaks a phrase including the search term and specific keywords, the processor 120 can input the search term into the second search term input window 920. In other words, when multiple search term input windows exist on a webpage, the currently displayed search term input window can be used to perform a search.

[0116] Voice control can be performed based on the screen of display 110. Essentially, applications on the display 110 screen can be used to perform functions based on voice commands. However, when the input voice command does not match the object included on the display screen or is unrelated to the function of the application displayed on the screen, another application can be executed, and the function based on the voice command can be performed.

[0117] For example, when the currently running application is a web browsing application and the user's spoken voice does not match an object on the webpage displayed by the web browsing application, the processor 120 can execute another pre-defined application and perform a search function corresponding to the user's spoken voice. The pre-defined application can be an application that provides search functionality, such as an application that provides search results for text corresponding to the spoken voice using a search engine, or an application that provides search results for video-on-demand (VOD) content based on the text corresponding to the spoken voice. Before executing the pre-defined application, the processor 120 can display a UI message for receiving user consent: "No results corresponding to xxx exist on the screen. Would you like to search for xxx on the internet?", or, after the user enters their consent on the UI, provide search results by executing an internet search application.

[0118] Display device 100 may include a voice processor for processing voice recognition results received from server 300 and an application unit for executing applications provided in display device 100. The voice processor may provide the voice recognition results received from server 300 to the application unit. When the recognition results are provided while executing a first application of the application unit and displaying the screen of the first application on display 110, the first application may perform the aforementioned operations based on the voice recognition results received from the voice processor. For example, when "search" is included in the voice recognition results, a search may be performed on text or image objects corresponding to numbers included in the voice recognition results, on text objects corresponding to words included in the voice recognition results, or on a search performed after keywords are entered in a search window.

[0119] If there is no operation to be performed by the first application using the speech recognition result received from the speech processor—that is, no text object or image object corresponding to the speech recognition result, or no search window—the first application can output an instruction to the speech processor indicating such a result. The speech processor can then control the application unit to execute a second application, which performs operations related to the speech recognition result. For example, the second application could be an application that provides search results for a specific search term. The application unit can execute the second application and provide search results for the text included in the speech recognition result used as the search term.

[0120] Figure 10 This is a block diagram illustrating the configuration of the display device. (In the description...) Figure 10 When, will be omitted Figure 4 Redundant description.

[0121] See Figure 10 Examples of display device 100 may include, but are not limited to, analog TVs, digital TVs, 3D-TVs, smart TVs, LED TVs, OLED TVs, plasma TVs, monitors, screen TVs with fixed curvature screens, flexible TVs with fixed curvature screens, curved TVs with fixed curvature screens, and / or TVs with variable curvature whose screen curvature changes according to received user input. As discussed above, display device 100 may be any type of display device, including PCs, smartphones, etc.

[0122] Display device 100 may include display 110, processor 120, tuner 130, communicator 140, microphone 150, input / output unit 160, audio output unit 170, and storage device 180.

[0123] Tuner 130 can select a channel by tuning to the frequency of a channel received by display device 100 from multiple radio wave components that are amplified, mixed, and resonated from a broadcast signal received in a wired / wireless manner. The broadcast signal may include video, audio, or additional data (e.g., electronic program guide (EPG)).

[0124] Tuner 130 can receive video, audio, and data in the frequency band corresponding to the channel number input by the user.

[0125] Tuner 130 can receive broadcast signals from various sources such as terrestrial broadcasting, cable broadcasting, or satellite broadcasting. Tuner 130 can also receive broadcast signals from various sources such as analog broadcasting or digital broadcasting.

[0126] The tuner 130 may be integrated with the display device 100 as an integral unit in a general shape or implemented as an add-on device (e.g., a set-top box or tuner connected to the input / output unit 160), the add-on device including a tuner unit electrically connected to the display device 100.

[0127] The communicator 140 can perform communication with various types of external devices according to various communication methods. The communicator 140 can connect to external devices via a local area network (LAN) or the Internet, and can also connect via wireless communication (e.g., Z-wave, 4LoWPAN, RFID, LTE D2D, BLE, GPRS, zero gravity, Edge Zigbee, ANT+, NFC, IrDA, DECT, WLAN, Bluetooth, WiFi, Wi-Fi Direct, GSM, UMTS, LTE, WiBRO, etc.). The communicator 140 may include various communication chips, such as Wi-Fi chip 141, Bluetooth chip 142, NFC chip 143, and wireless communication chip 144. The Wi-Fi chip 141, Bluetooth chip 142, and NFC chip 143 can communicate with each other using WiFi, Bluetooth, or NFC, respectively. The wireless communication chip 174 can be a chip that performs communication according to various communication standards such as IEEE, ZigBee, 3G, 3GPP, and LTE. The communicator 140 may also include an optical receiver 145 capable of receiving control signals (e.g., IR pulses) from an external device 200.

[0128] The processor 120 can transmit voice signals and language information (information about the basic language used for speech recognition) to the server 300 via the communicator 140, and the processor 120 can receive the speech recognition results performed on the voice signals by the server 300 using a speech recognition engine for the language corresponding to the language information.

[0129] The microphone 150 can receive spoken voice from a user and generate a voice signal corresponding to the received voice. The microphone 150 can be integrated with or separate from the display device 100. A separate microphone 150 can be electrically connected to the display device 100.

[0130] When the microphone is not included in the display device 100, the display device 100 can receive a voice signal corresponding to the voice input through the microphone of the external device 200 from the external device 200 via the communicator 140. The communicator 140 can receive the voice signal from the external device 200 using WiFi, Bluetooth, or the like.

[0131] The input / output unit 160 can be connected to a device. The input / output unit 160 may include at least one of a High Definition Multimedia Interface (HDMI) port 161, a component input jack 162, and a USB port 163. Additionally, the input / output unit 160 may include at least one of ports such as RGB, DVI, HDMI, DP, and PWM.

[0132] The audio output unit 170 can output audio, such as audio included in a broadcast signal received via the tuner 130, audio input via the communicator 140, the input / output unit 160, etc., or audio included in an audio file stored in the storage device 180. The audio output unit 170 may include a speaker 171 and a headphone output terminal 172.

[0133] The storage device 180 may include various application programs, data, and software modules for driving and controlling the display device 100 under the control of the processor 120. For example, the storage device 180 may include a web page parsing module, a JavaScript module, a graphics processing module, a speech recognition result processing module, an input processing module, etc., for parsing web page content data received via the Internet.

[0134] When the display device 100 itself, rather than the external server 300, performs speech recognition, the storage device 180 can store speech recognition modules including various speech recognition engines for various languages.

[0135] Storage device 180 can store data used to form various UI screens provided by display 110. Storage device 180 can also store data used to generate control signals corresponding to various user interactions.

[0136] Storage device 180 can be implemented as non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD). Storage device 180 can be implemented not only as a storage medium in display device 100, but also as an external storage medium such as a micro SD card, USB memory, or network server via a network.

[0137] The processor 120 can control the overall operation of the display device 100, control the signal flow between internal components in the display device 100, and process data.

[0138] Processor 120 may include RAM 121, ROM 122, CPU 123, and bus 124. RAM 121, ROM 122, and CPU 123 can be interconnected via bus 124. Processor 120 may be implemented as a system-on-a-chip (SoC).

[0139] CPU 123 can access storage device 180 and use the operating system stored in storage device 180 to perform booting. In addition, CPU 123 can perform various operations by using various programs, contents and data stored in storage device 180.

[0140] ROM 122 can store a set of commands for system booting. If a power-on command is input and power is supplied, CPU 123 can copy the operating system stored in storage device 180 to RAM 121 according to the commands stored in ROM 122, execute the operating system, and boot the system. When booting is complete, CPU 123 can copy various programs stored in storage device 180 to RAM 121, execute the applications copied to RAM 121, and perform various operations.

[0141] The processor 120 can perform various operations using modules stored in the storage device 180. For example, the processor 120 can parse and process web page content data received via the Internet and display the overall layout of the content and objects on the display 110.

[0142] When the voice recognition function is enabled, the processor 120 can analyze the objects in the web page content, search for objects that can be controlled by voice, perform preprocessing on information about the object's location, object-related operations, and text in the object, and store the preprocessing results in the storage device 180.

[0143] Processor 120 can control display 110 to display selectable objects (which can be voice-controlled) to be identified based on preprocessed object information. For example, processor 120 can control display 110 to display a color that distinguishes the voice-controlled object from other objects.

[0144] Processor 120 can recognize speech input through microphone 150 as text using a speech recognition engine. Processor 120 can use a speech recognition engine with a preset language (the basic language used for speech recognition). Processor 120 can transmit information about the speech signal and the basic language used for speech recognition to server 300, and receive text from server 300 as the speech recognition result.

[0145] The processor 120 can search for objects in the preprocessed objects that correspond to the speech recognition results and instruct the selection of objects at the locations of the searched objects. For example, the processor 120 can control the display to highlight the selected object by voice. The processor 120 can perform operations related to the objects corresponding to the speech recognition results based on the preprocessed object information and output the results through the display 110 or the audio output unit 170.

[0146] Figure 11 This is a flowchart illustrating a method for controlling a display device according to an exemplary embodiment of the present disclosure.

[0147] Figure 11 The flowchart shown illustrates the operations processed by the display device 100 described herein. Therefore, although repeated descriptions are omitted below, the description of the display device 100 can be applied to... Figure 11 The flowchart.

[0148] See Figure 11 In step S1110, the display device 100 can display a UI screen that includes multiple text objects.

[0149] In step S1120, the display device 100 may display a text object in a language different from a preset language among a plurality of text objects displayed on the display device, as well as a preset number. The preset language may refer to a pre-determined basic language for speech recognition. The basic language may be a default language, or it may be manually set by the user, or it may be automatically set based on the language used for the objects displayed on the display 110. When the basic language is automatically set, optical character recognition (OCR) may be applied to the objects displayed on the display device 100 to confirm the language used for the objects.

[0150] In step S1130, when the recognition result of the user's spoken voice includes the displayed number, an operation related to the text object corresponding to the displayed number can be performed.

[0151] The recognition result of the user's spoken speech can be obtained from the speech recognition of the display device itself, or by sending a speech recognition request to an external server that performs speech recognition for multiple different languages. By sending a speech recognition request, the display device 100 can provide the external server with information about the speech signal corresponding to the user's spoken speech and the basic language used for speech recognition, and when it receives the speech recognition result, including the displayed digits, from the external server, it performs operations related to the text object corresponding to the displayed digits.

[0152] For example, when the text object is hyperlink text in a webpage, an operation can be performed to display a webpage with a URL address corresponding to the text object, and if the text object is an icon used to execute an application, the application can be executed.

[0153] A UI screen including multiple text objects can be the execution screen of a first application. The execution screen of the first application can be any screen provided by the first application. While displaying the execution screen of the first application, if it is determined that the object corresponding to the recognition result of the user's spoken speech does not exist on the execution screen of the first application, the display device can execute a second application different from the first application and perform the operation corresponding to the speech recognition result. The first application can be a web browsing application, and the second application can be an application for performing searches in various sources, such as the Internet, data stored on the display device, VOD content, and channel information (e.g., EPG). For example, when the displayed webpage does not contain an object corresponding to the speech recognition, the display device can execute another application and provide search results corresponding to the speech recognition (e.g., search engine results, VOD search results, channel search results, etc.).

[0154] According to the above exemplary implementation, objects in various languages ​​can be controlled by voice, and voice searches can be easily performed.

[0155] The exemplary embodiments described above can be implemented using software, hardware, or a combination thereof on a recording medium readable by a computer or similar device. Depending on the hardware implementation, the exemplary embodiments described herein can be implemented using at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, and an electrical unit for performing other functions. In some cases, the exemplary embodiments described herein can be implemented by the processor 120 itself. Depending on the software implementation, exemplary embodiments such as the processes and functions described herein can be implemented in separate software modules. Each software module can perform one or more functions and operations described herein.

[0156] Computer instructions for performing the processing operations in the display device 100 according to the exemplary embodiments of the present disclosure described above may be stored on a non-transitory computer-readable medium. The computer instructions stored on the non-volatile computer-readable medium, when executed by a processor of a particular device, cause the processor and other components of the particular device to perform processing operations in the display device 100 according to the various embodiments described above.

[0157] Non-volatile computer-readable media refers to media that store data semi-permanently and can be read by a device, rather than media that store data for short periods, such as registers, caches, and memory. Specific examples of non-volatile computer-readable media include CDs, DVDs, hard drives, Blu-ray discs, USB drives, memory cards, and ROMs.

[0158] Although exemplary embodiments have been shown and described, those skilled in the art will understand that changes can be made to these exemplary embodiments without departing from the principles and spirit of this disclosure. However, the scope of the invention is not limited to the detailed description in the specification, but is defined by the scope of the claims, and those skilled in the art will understand that various changes in form and detail can be made without departing from the spirit and scope of the invention as set forth in the following specification.

Claims

1. A display device, comprising: monitor; A communication unit for communicating with an external control device, the external control device including a microphone and a microphone button for activating the microphone; as well as The processor is configured as follows: Based on user input received through the external control device, the web browsing application on the display device is run. Identify multiple hyperlink objects included in the first webpage displayed by the application through this webpage browsing. Extract the text keywords from the multiple hyperlink objects. When the microphone button of the external control device is pressed, an icon including a number is displayed in the link of at least one of the identified hyperlink objects, wherein the icon including the number is displayed near hyperlink objects containing text that cannot be recognized by speech. User voice input is received via a microphone included in the external control device. Process the user's voice input to obtain text information corresponding to the user's voice input. The text information input by the user's voice is compared with the numbers included in the icon and the text keywords extracted from the plurality of hyperlink objects. Determine whether one of the multiple hyperlink objects matches the text information input by the user's voice. Based on the determination that the hyperlink object among the plurality of hyperlink objects matches the text information input by the user's voice, the web browsing application is controlled to provide a second webpage corresponding to the hyperlink object. Based on the determination that the hyperlink object does not match the text information input by the user's voice, information indicating that the text information does not match is received, and Based on the information indicating that the text information does not match, while the web browsing application is running, a search application different from the web browsing application is run. This search application is used to perform search operations related to the text information input by the user's voice through an external server.

2. The display device as claimed in claim 1, wherein, The search application includes a video search application for performing searches of video content associated with the text information via the external server.

3. The display device as claimed in claim 1, wherein, The processor is configured as follows: While the web browsing application displays the first webpage in the foreground, the microphone button related to voice recognition on the external device is pressed, controlling the web browsing application to analyze multiple hyperlink objects included in the first webpage, and displaying symbols near the multiple hyperlink objects to guide the user's speech.

4. The display device as claimed in claim 3, wherein, The microphone button associated with voice recognition on this external device is the button to activate the microphone of this external device.

5. The display device as claimed in claim 3, wherein, The symbols used to guide the user's speech include the number, and The processor is configured to, when the text information includes a second number, identify, based on the second number included in the text information and the symbol including the number, a hyperlink object among a plurality of hyperlink objects included in the first webpage that corresponds to the user's speech.

6. The display device as claimed in claim 3, wherein, The symbols used to guide the user's speech include this icon, and The processor is configured to, when the text information includes text, identify, based on the text included in the text information and the symbol including the icon, a hyperlink object among multiple hyperlink objects included in the first webpage that corresponds to the user's words.

7. The display device as claimed in claim 6, wherein, Display a symbol including the icon near content that includes text capable of speech recognition.

8. The display device as claimed in claim 1, wherein, The processor is configured as follows: If no hyperlink object corresponding to the text information is recognized in the first webpage, the display is controlled to show a user interface (UI) asking whether permission to use the search application is granted. The search application runs based on user voice input to the UI.

9. The display device as claimed in claim 1, wherein, The processor is further configured to control the retrieval application to provide search results using the text information.

10. A display method performed by a display device, the display method comprising: The web browsing application on the display device is run based on user input received through an external control device including a microphone and a microphone button for activating the microphone; Identify multiple hyperlink objects included in the first webpage displayed by the application through this webpage; Extract the text keywords from the multiple hyperlink objects; When the microphone button of the external control device is pressed, an icon including a number is displayed in the link of at least one of the identified multiple hyperlink objects, wherein the icon including the number is displayed near the hyperlink object containing text that cannot be recognized by speech. The external control device receives user voice input via a microphone. Process the user's voice input to obtain text information corresponding to the user's voice input; The text information input by the user's voice is compared with the numbers included in the icon and the text keywords extracted from the plurality of hyperlink objects; Determine whether the hyperlink object among the plurality of hyperlink objects matches the text information input by the user's voice. Based on the determination that the hyperlink object among the plurality of hyperlink objects matches the text information input by the user's voice, the web browsing application is controlled to provide a second webpage corresponding to the hyperlink object; Based on the determination that the hyperlink object does not match the text information input by the user's voice, information indicating that the text information does not match is received; and Based on the information indicating that the text information does not match, while the web browsing application is running, a search application different from the web browsing application is run. This search application is used to perform search operations related to the text information input by the user's voice through an external server.

11. The display method as described in claim 10, wherein, The search application includes a video search application for performing searches of video content associated with the text information via the external server.

12. The display method of claim 10, further comprising: While the web browsing application displays the first webpage in the foreground, the microphone button related to voice recognition on the external device is pressed, controlling the web browsing application to analyze multiple hyperlink objects included in the first webpage, and displaying symbols near the multiple hyperlink objects to guide the user's speech.

13. The display method as described in claim 12, wherein, The microphone button associated with voice recognition on this external device is the button to activate the microphone of this external device.

14. The display method as described in claim 12, wherein, The symbols used to guide the user's speech include the number, and The display method further includes, when the text information includes a second number, identifying, based on the second number included in the text information and the symbol including the number, a hyperlink object among a plurality of hyperlink objects included in the first webpage that corresponds to the user's speech.

15. The display method as described in claim 12, wherein, The symbols used to guide the user's speech include this icon, and The display method further includes, when the text information includes text, identifying the hyperlink object that corresponds to the user's words among a plurality of hyperlink objects included in the first webpage, based on the text included in the text information and the symbol including the icon.

16. The display method as described in claim 15, wherein, Display a symbol including the icon near content that includes text capable of speech recognition.

17. The display method as described in claim 10, wherein, Running this search application includes: If no hyperlink corresponding to the text information is recognized on the first webpage, a user interface (UI) prompt will appear asking whether permission to use the search application is granted. The search application runs based on user voice input to the UI.

18. The display method as described in claim 10, wherein, Running this search application includes: Controls the search application to provide search results using this text information.

19. A computer-readable medium storing a computer-readable program, wherein, When this computer-readable program is run on a display device, the display device: The web browsing application on the display device is run based on user input received through an external control device including a microphone and a microphone button for activating the microphone; Identify multiple hyperlink objects included in the first webpage displayed by the application through this webpage; Extract the text keywords from the multiple hyperlink objects; When the microphone button of the external control device is pressed, an icon including a number is displayed in the link of at least one of the identified multiple hyperlink objects, wherein the icon including the number is displayed near the hyperlink object containing text that cannot be recognized by speech. The external control device receives user voice input via a microphone. Process the user's voice input to obtain text information corresponding to the user's voice input; The text information input by the user's voice is compared with the numbers included in the icon and the text keywords extracted from the plurality of hyperlink objects; Determine whether the hyperlink object among the plurality of hyperlink objects matches the text information input by the user's voice. Based on the determination that the hyperlink object among the plurality of hyperlink objects matches the text information input by the user's voice, the web browsing application is controlled to provide a second webpage corresponding to the hyperlink object; Based on the determination that the hyperlink object does not match the text information input by the user's voice, information indicating that the text information does not match is received; and Based on the information indicating that the text information does not match, while the web browsing application is running, a search application different from the web browsing application is run. This search application is used to perform search operations related to the text information input by the user's voice through an external server.

20. The computer-readable medium of claim 19, wherein, The search application includes a video search application for performing searches of video content associated with the text information via the external server.

21. The computer-readable medium of claim 19, wherein, When the computer-readable program is run on a display device, it further enables the display device to: While the web browsing application displays the first webpage in the foreground, the microphone button related to voice recognition on the external device is pressed, controlling the web browsing application to analyze multiple hyperlink objects included in the first webpage, and displaying symbols near the multiple hyperlink objects to guide the user's speech.

22. The computer-readable medium of claim 21, wherein, The microphone button associated with voice recognition on this external device is the button to activate the microphone of this external device.

23. The computer-readable medium of claim 21, wherein, The symbols used to guide the user's speech include the number, and Wherein, when the computer-readable program is run on the display device, the display device further causes the display device to: When the text information includes a second number, based on the second number included in the text information and the symbol including the number, identify the hyperlink object that corresponds to the user's speech among the multiple hyperlink objects included in the first webpage.

24. The computer-readable medium of claim 21, wherein, The symbols used to guide the user's speech include this icon, and Wherein, when the computer-readable program is run on the display device, the display device further causes the display device to: When the text information includes text, based on the text included in the text information and the symbol including the icon, identify the hyperlink object that corresponds to the user's words among the multiple hyperlink objects included in the first webpage.

25. The computer-readable medium of claim 24, wherein, Display a symbol including the icon near content that includes text capable of speech recognition.

26. The computer-readable medium of claim 19, wherein, When the computer-readable program is run on a display device, it further enables the display device to: If no hyperlink corresponding to the text information is recognized on the first webpage, a user interface (UI) prompt will appear asking whether permission to use the search application is granted. The search application runs based on user voice input to the UI.

27. The computer-readable medium of claim 19, wherein, When the computer-readable program is run on a display device, it further enables the display device to: Controls the search application to provide search results using this text information.

Citation Information

Patent Citations

  • Apparatus for managing for taking medicine

    KR1020170091494A

  • Control method and control device

    CN103888799A