Voice input device and voice input method
The voice input device and method automatically recognize keywords and select options within a pop-up menu or list box, eliminating the need for mode switching and reducing errors in voice input operations.
Patent Information
- Application Number
- PCT/JP2023/046999
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Existing voice input systems require users to manually switch modes for voice input, which is cumbersome and prone to unintended conversations or malfunctions due to ambient environmental sounds.
A voice input device and method that utilizes a keyword recognition unit to identify keywords from input voice and extract options from a UI element's option list, allowing voice input operations without the need for mode switching by automatically recognizing and selecting options within a pop-up menu or list box.
Enables seamless voice input operations by automatically selecting options within a pop-up menu or list box, reducing user effort and minimizing errors by ensuring voice input is only accepted for recognized keywords.
Smart Images

Figure JP2023046999_03072025_PF_FP_ABST
Abstract
Description
Voice input device and voice input method
[0001] The present invention relates to a voice input device and a voice input method.
[0002] When performing voice input into an information processing device such as a personal computer (PC) or a mobile terminal device, it is necessary to switch the mode for voice input beforehand. For example, to input voice into a text input field, it is necessary to click a microphone icon or the like located in the text field. Also, to input voice into a voice assistant, it is necessary to speak a specific phrase in advance. Furthermore, Patent Document 1 describes a device that recognizes voice commands and controls a mobile terminal device based on the recognized commands.
[0003] JP 2009-252238 A
[0004] As described above, when performing voice input, it is necessary to switch the mode for voice input in advance. This is to prevent malfunctions due to conversations that are not intended as voice input or surrounding environmental sounds. However, switching the mode every time voice input is performed is a cumbersome task for the user.
[0005] In view of the above-mentioned problems, an object of the present invention is to provide a voice input device and a voice input method that allow a voice input operation to be easily performed without setting a voice input mode.
[0006] A voice input device according to one aspect of the present invention includes a keyword recognition unit that recognizes keywords from input voice, and an extraction unit that extracts options from an option list belonging to a UI element to be input based on the extracted keywords.
[0007] A speech input method according to one aspect of the present invention includes a step of recognizing a keyword from input speech, and a step of extracting an option from an option list belonging to a UI element to be input based on the extracted keyword.
[0008] According to the present invention, the user can perform a voice input operation by operating a specific target UI element without setting the mode to voice input mode.
[0009] 1 is a block diagram showing the configuration of an information processing apparatus to which a voice input device according to an embodiment of the present invention can be applied. FIG. 1 is a diagram illustrating a voice input operation according to a first embodiment of the present invention. FIG. 1 is an explanatory diagram of a voice input operation according to the first embodiment of the present invention. FIG. 1 is an explanatory diagram of a voice input operation according to the first embodiment of the present invention. FIG. 2 is an explanatory diagram of a voice input operation according to the first embodiment of the present invention. FIG. 3 is an explanatory diagram of a voice input operation according to the first embodiment of the present invention. FIG. 4 is a flowchart showing processing of a voice input operation according to the first embodiment of the present invention. FIG. 5 is a flowchart showing another example of processing of a voice input operation according to the first embodiment of the present invention. FIG. 6 is an explanatory diagram of a voice input operation according to a second embodiment of the present invention. FIG. 7 is an explanatory diagram of a voice input operation according to the second embodiment of the present invention. FIG. 8 is an explanatory diagram of a voice input operation according to the second embodiment of the present invention. FIG. 9 is an explanatory diagram of a voice input operation according to the second embodiment of the present invention. It is an explanatory diagram of a voice input operation according to a fourth embodiment of the present invention.It is an explanatory diagram of a voice input operation according to the fourth embodiment of the present invention.It is an explanatory diagram of a modification of the present invention.It is a block diagram based on the function of a voice input device according to another aspect of the present invention.
[0010]
[0023] Hereinafter, embodiments of the present invention will be described with reference to the drawings. <First Embodiment> Fig. 1 is a block diagram showing the configuration of an information processing device 1 to which a voice input device according to an embodiment of the present invention can be applied. As shown in Fig. 1, the information processing device 1 includes a control unit 11, a storage unit 12, a display unit 13, an input unit 14, a voice input unit 15, and an image input unit 16.
[0011] The control unit 11 includes a central processing unit (CPU) and executes various processes according to programs stored in the storage unit 12. In this embodiment, the control unit 11 implements a voice input processing unit 20 that executes processes based on voice input from the voice input unit 15. More specifically, when the control unit 11 executes a process to execute a specific UI (user interface) element, a voice input operation is activated. The control unit 11 then recognizes a keyword from the input voice and extracts an option from a choice list belonging to the UI element to be input based on the extracted keyword. This allows a user to select a desired option from the choice list through voice input operation without performing a mode setting operation. The control unit 11 also functions as a keyword recognition unit that recognizes a keyword from the input voice. The UI element is an element that constitutes a user interface and may be any of a button, a scroll bar, a menu, a check box, etc. The UI element may also be a context menu. In this embodiment, a case where the UI element is a pop-up window, a list box, a pull-down menu, etc. will be described as an example.
[0012] The storage unit 12 is configured by a storage medium, such as a hard disk drive (HDD), a flash memory, an electrically erasable programmable read-only memory (EEPROM), a random access read / write memory (RAM), a read-only memory (ROM), or any combination of these storage media. The storage unit 12 stores, for example, an option list.
[0013] The display unit 13 is made up of an LCD (Liquid Crystal Display), an organic EL (Electroluminescence) display, or the like, and displays various types of information on the screen.
[0014] The input unit 14 corresponds to a keyboard, a mouse, a touch panel, etc., and accepts various input operations.
[0015] The audio input unit 15 includes a microphone and receives external sounds as audio signals, while the image input unit 16 includes a camera and receives images of the surroundings.
[0016] As described above, in the information processing device 1 according to the first embodiment of the present invention, when a process for executing a specific UI element is performed, the voice input operation is turned on, and a choice from the choice list to which the UI element to be input belongs can be selected by the voice input operation. This will be described below.
[0017] In the following description, "pop-up" refers to a display element that appears to jump out to the foreground. Examples of pop-up displays include pop-up windows and pop-up menus. A pop-up window refers to a window that pops up in response to a specific operation, such as a mouse click. A pop-up menu refers to a UI element that displays a list of options in a pop-up window and accepts input for the options in the pop-up menu. Here, a pop-up menu is considered to be one form of a pop-up window.
[0018] The focus refers to an element among the UI elements that is ready to receive input from the user.
[0019] A control is a part that makes up a UI. In this example, the pop-up menu and list box are examples of controls.
[0020] The cursor represents an object to be operated on to select an option within a control.
[0021] The mouse pointer refers to a mark that moves in response to the movement of a pointing device such as a mouse.
[0022] In the voice input operation according to the first embodiment of the present invention, when a pop-up menu is displayed as a specific UI element, voice input is turned on.
[0023] 2 to 5 are explanatory diagrams of a voice input operation according to the first embodiment of the present invention. In FIG. 2, a pop-up menu 101 is displayed as a UI element for the user to select a hometown. In the state shown in FIG. 2, the pop-up menu 101 is closed. In this state, the voice input operation is turned off, and there is no reaction to the voice input operation from the voice input unit 15. In other words, even if the user utters a voice in the state shown in FIG. 2, there is no reaction to the voice input. Also, in the state shown in FIG. 2, a check mark 103 is added to "Aomori," and "Aomori" is selected as the hometown.
[0024] As shown in Fig. 3, when the user clicks on the pop-up menu 101 with the mouse, a pop-up menu 101a is displayed. Options such as "Hokkaido," "Aomori," "Akita," ... "Okinawa" are displayed in the pop-up menu 101a. When the pop-up menu 101a is displayed in this manner, a voice input operation is enabled. When the user performs a voice input operation by uttering a voice included in an option, an option corresponding to the voice input is selected from the options included in the pop-up menu 101a.
[0025] For example, suppose the user utters "Akita." In this case, as shown in Figure 4, options belonging to the pop-up menu 101 are extracted using "Akita" as a keyword, and "Akita" is selected as the place of origin. The cursor 102 then moves to the selected option. Note that even if the user utters an option not in the pop-up menu 101a, such as "Nihon," there will be no response to the voice input.
[0026] As described above, in this embodiment, the user's voice input operation is accepted only for the options in the pop-up menu 101a while the pop-up menu 101a is popped up. According to this embodiment, the voice input operation can be performed by displaying the pop-up menu 101a in a pop-up state, eliminating the need for the user to switch modes. That is, the control unit 11 accepts voice input only for keywords included in the options in the pop-up menu. This prevents speech from being accepted for keywords other than those included in the options in the pop-up menu. Even if speech is spoken in the vicinity, input through the voice input unit 15, and recognized, keywords other than those included in the options are not accepted.
[0027] The following two methods are conceivable as operations when voice input is made while the pop-up menu 101a is displayed as a pop-up.
[0028] (A1) Change the selected item. (A2) Without changing the selected item, change the selected item by placing the cursor on the option and pressing the Enter key.
[0029] In method (A1), for example, if the user utters "Akita," the position of the check mark 103 is changed from "Aomori" to "Akita" as shown in Fig. 5, and the option is confirmed as "Akita." In method (A2), for example, if the user utters "Akita," the cursor 102 is placed on the option "Akita" as shown in Fig. 4, and the check mark 103 is not changed. In this method, the option is confirmed as "Akita" when the user finally performs an operation such as pressing the Enter key.
[0030] Method (A1) allows for easy voice input operation, but makes it difficult to correct an incorrect voice input operation. Method (A2) allows for easy redo of a voice input operation when an incorrect voice input occurs, but requires an operation to confirm the selection. Either method may be used as the voice input method while the pop-up menu 101a is displayed as a pop-up. Also, method (A1) and method (A2) may be optionally selectable.
[0031] FIG. 6 is a flowchart showing a process of a voice input operation according to the first embodiment of the present invention.
[0032] (Step S101) The control unit 11 initially sets the state of acceptance of voice input to OFF.
[0033] (Step S102) The control unit 11 determines whether or not an input operation has been performed from the input unit 14. If there has been no input operation (step S102: No), the control unit 11 returns the process to step S102, and if there has been an input operation (step S102: Yes), the control unit 11 proceeds to step S103.
[0034] (Step S103) The control unit 11 determines whether the executed input operation is an operation to click on a pop-up menu. If the operation to click on a pop-up menu has not been performed (Step S103: No), the control unit 11 returns the process to Step S102. If the operation to click on a pop-up menu has been performed (Step S103: Yes), the control unit 11 proceeds to Step S104.
[0035] (Step S104) When the pop-up menu is clicked in step S102, the control unit 11 displays the pop-up menu as a pop-up, and the process proceeds to step S105.
[0036] (Step S105) The control unit 11 sets the state of acceptance of voice input to ON, and proceeds to step S106.
[0037] (Step S106) The control unit 11 determines whether or not a voice input operation has been performed from the voice input unit 15. If a voice input operation has been performed (Step S106: Yes), the control unit 11 proceeds to step S107, and if a voice input operation has not been performed (Step S106: No), the control unit 11 proceeds to step S110.
[0038] (Step S107) If a voice input operation is performed from the voice input unit 15 (step S106: Yes), the control unit 11 recognizes the input voice, converts it into a character string, and proceeds to step S108.
[0039] (Step S108) The control unit 11 compares the character string of the options serving as keywords in the pop-up menu with the character string input by voice, and determines whether or not there is a character string corresponding to the character string input by voice among the character strings of the options serving as keywords in the pop-up menu. If there is a character string corresponding to the character string input by voice among the character strings of the options serving as keywords in the pop-up menu (Step S108: Yes), the control unit 11 proceeds to step S109, and if there is no corresponding character string (Step S108: No), the control unit 11 returns to step S106.
[0040] (Step S109) The control unit 11 selects, as an option, a character string in the pop-up menu that corresponds to the character string input by voice, and proceeds to step S111.
[0041] (Step S110) If there is no voice input operation in step S106 (step S106: No), the control unit 11 determines whether or not there has been an input operation from the input unit 14. If there has been no input operation (step S110: No), the control unit 11 returns the process to step S106, and if there has been an input operation (step S110: Yes), the control unit 11 proceeds to step S111.
[0042] (Step S111) The control unit 11 determines whether or not an operation to close the pop-up menu has been performed. If an operation to close the pop-up menu has not been performed (Step S111: No), the control unit 11 returns the process to Step S106. If an operation to close the pop-up menu has been performed (Step S111: Yes), the control unit 11 proceeds to Step S112.
[0043] (Step S112) The control unit 11 ends the display of the pop-up menu, and proceeds to step S113.
[0044] (Step S113) The control unit 11 sets the state of acceptance of voice input to OFF, and returns the process to step S102.
[0045] In this embodiment, voice input is turned off when the pop-up menu is not displayed. Therefore, even if the user speaks when the pop-up menu is closed, voice input is not performed. Therefore, in order to inform the user that voice input is not possible when the pop-up menu 101 is closed, a warning indicating "keyword-based input is not accepted" may be displayed on the display unit 13 when voice input is attempted when the pop-up menu is closed.
[0046] Furthermore, even if the user inputs a voice after the pop-up menu is displayed, there are cases where the voice is not included in the keywords to be selected or the voice cannot be recognized. Therefore, even if the user inputs a voice while the pop-up menu is displayed, if the voice is not included in the keywords to be selected or the voice cannot be recognized, a warning indicating that "keyword-based input cannot be accepted" may be displayed on the display unit 13.
[0047] 7 is another example of a flowchart showing a voice input operation according to the first embodiment of the present invention. In FIG. 7, the same processes as those in FIG. 6 are denoted by the same numbers, and the description thereof will be omitted.
[0048] In this example, in step S150, the control unit 11 determines whether or not a voice input operation has been performed while the pop-up menu is not displayed. If a voice input operation has been performed (step S150: Yes), the control unit 11 displays a warning indicating that voice input is not being accepted (step S151). This notifies the user that voice input is not being accepted while the pop-up menu is not displayed.
[0049] In this example, if the control unit 11 determines in step S108 that the voice input is not included in the options (step S108: No), it displays a warning indicating that the voice cannot be recognized. This allows the user to be informed that the voice cannot be recognized if the voice input is not included in the options or the voice cannot be recognized, even if there is a voice input while the pop-up menu is displayed.
[0050] Second Embodiment Next, a description will be given of a voice input according to a second embodiment of the present invention. In a voice input operation according to the second embodiment of the present invention, if a list box is focused as a specific UI element, the voice input is turned on.
[0051] 8 and 9 are explanatory diagrams of a voice input operation according to the second embodiment of the present invention.
[0052] 8, a list box 201 is displayed as a UI element for the user to select a hometown. The list box 201 displays options such as "Hokkaido," "Aomori," "Akita," etc. In the state shown in FIG. 8, the focus is not on the pop-up menu 201. In this state, voice input operation is turned off, and there is no response to voices from the voice input unit 15.
[0053] When the user clicks on the list box 201 with the mouse, the focus (indicated by a thick line) is placed on the list box 201, as shown in Fig. 9. In the list box 201a that has the focus, the voice input operation is turned on. In other words, when the user performs a voice input operation by uttering a voice included in an option in the list box 201a that has the focus, the option corresponding to the voice input is selected from the options included in the list box 201a.
[0054] In this way, in this embodiment, only the options in the focused list box 201a are configured to accept voice input operations from the user. According to this embodiment, by clicking on the list box 201 and focusing on the list box 201, voice input operations can be performed, eliminating the need for the user to switch modes.
[0055] The following methods can be used to display the options selected in the list box.
[0056] (B1) Move the cursor to an option that satisfies the condition and display it. (B2) Color the option that satisfies the condition and display it.
[0057] For example, when searching for a document file in a list box that displays a list of documents or files, there may be cases where the user wants to extract the document file using criteria such as "year" or "month." A list box 301 shown in FIG. 10 displays document files as options. In this case, when selecting a document file from the options in the list box 301 using "2021" as a criterion, in this embodiment, the user focuses on the list box 301 and utters "2021" as a voice input.
[0058] 11 and 12 show the selection results. In FIG. 11, as shown in method (B1), the cursor 302 is moved to the first item in the document file for the date that matches the voice input. The document file positioned after the cursor 302 becomes the document file for "2021" to be selected.
[0059] 12, as shown in the method (B2), document files that match the search target are displayed in color 303. That is, in FIG. 12, the document file colored 303 is the document file for "2021" to be selected.
[0060] Furthermore, in order to prevent erroneous operation when using voice input in a list box, it may be possible to set additional conditions other than the focus being on the item.
[0061] In other words, once the focus is achieved, it often remains focused for a long time. A UI element that has the focus but remains focused for a long time without any user operation is likely not the target of user operation. It is also possible that the user is operating a different UI element rather than the focused UI element. For example, in today's PC operating systems, scrolling with the mouse wheel is often possible when the mouse is positioned over an unfocused UI element. When a mouse wheel scroll operation is performed over an unfocused UI element, it is likely that the user is operating the UI element where the mouse pointer is positioned, rather than the focused UI element. Therefore, the following condition can be added:
[0062] (C1) The condition that the mouse pointer is inside or near the UI element is added as a condition for enabling voice input, or this takes priority over the presence or absence of the cursor. (C2) The condition that the user's line of sight is inside or near the UI element is added as a condition for enabling voice input, or this takes priority over the presence or absence of the cursor. The user's line of sight can be obtained by photographing the user's face using the image input unit 16. (C3) If no voice operation is performed on the UI element where the cursor is located for a certain period of time or more, voice input is ignored until the next operation.
[0063] Furthermore, when selecting using a list box, there are options that are not displayed in the list box. For example, in Fig. 13, the options displayed in list box 401 range from "Hokkaido" to "Miyagi," but the options in list box 401 should actually range from "Hokkaido" to "Okinawa." Options other than those displayed in list box 401 are located in areas 402 and 403 outside the display area, as shown in Fig. 14.
[0064] As the number of options in the list box increases, the number of options located in the areas 402 and 403 outside the list box 401 also increases. When performing voice input, including options located in the areas 402 and 403 outside the list box 401 increases the likelihood of malfunction due to surrounding voices. It is also difficult for the user to select an option that is not displayed in the list box 401. Therefore, when using voice input for a list box, it is possible to have the voice input respond in the following cases:
[0065] (D1) Voice input responds only to the options currently displayed on the screen as a list box. (D2) Voice input responds to all options contained in the list box. (D3) Each UI element has an attribute that determines whether it responds to all options or only to the options displayed.
[0066] FIG. 15 is a flowchart showing a process of a voice input operation according to the second embodiment of the present invention.
[0067] (Step S201) The control unit 11 initially sets the state of acceptance of voice input to OFF.
[0068] (Step S202) The control unit 11 determines whether or not an input operation has been performed from the input unit 14. If there has been no input operation (step S202: No), the control unit 11 returns the process to step S201, and if there has been an input operation (step S202: Yes), the control unit 11 proceeds to step S203.
[0069] (Step S203) The control unit 11 determines whether the executed input operation is an operation of clicking a list box. If an operation of clicking a list box has not been performed (Step S203: No), the control unit 11 returns the process to Step S202. If an operation of clicking a list box has been performed (Step S203: Yes), the control unit 11 proceeds to Step S204.
[0070] (Step S204) When the list box is clicked in step S203 (step S204: Yes), the control unit 11 focuses on the list box, displays it, and proceeds to step S205.
[0071] (Step S205) The control unit 11 sets the state of acceptance of voice input to ON, and proceeds to step S206.
[0072] (Step S206) The control unit 11 determines whether or not a voice input operation has been performed from the voice input unit 15. If a voice input operation has been performed (step S206: Yes), the control unit 11 proceeds to step S207, and if a voice input operation has not been performed (step S206: No), the control unit 11 proceeds to step S210.
[0073] (Step S207) If a voice input operation is performed from the voice input unit 15 (step S206: Yes), the control unit 11 recognizes the input voice, converts it into a character string, and proceeds to step S208.
[0074] (Step S208) The control unit 11 compares the character string of the options serving as keywords in the list box with the voice-input character string and determines whether or not there is a character string corresponding to the voice-input character string among the character strings of the options serving as keywords in the list box. If there is a character string corresponding to the voice-input character string among the character strings of the options serving as keywords in the list box (Step S208: Yes), the control unit 11 proceeds to step S209. If there is no corresponding character string (Step S208: No), the control unit 11 returns to step S206. Here, the control unit 11 may accept voice input only for keywords recognized from the input voice that correspond to the options in the list box. As a result, it is possible to not accept voice input for character strings obtained by voice recognition other than the options included in the list box, and to accept only the options. This makes it possible to not accept voice input even if something other than the options is spoken. Here, the control unit 11 may accept voice input only for keywords included in the list popped up by the pop-up menu, or may accept voice input only for the options displayed on the screen in the pop-up list.
[0075] (Step S209) The control unit 11 selects, as an option, a character string in the list box that corresponds to the character string input by voice, and returns the process to step S206.
[0076] (Step S210) If there is no voice input operation in step S207 (step S206: No), the control unit 11 determines whether or not there has been an input operation from the input unit 14. If there has been no input operation (step S210: No), the control unit 11 returns the process to step S206, and if there has been an input operation (step S210: Yes), the control unit 11 proceeds to step S211.
[0077] (Step S211) The control unit 11 determines whether or not an operation to remove the focus from the list box has been performed. If an operation to remove the focus from the list box has not been performed (Step S211: No), the control unit 11 returns the process to Step S206. If an operation to remove the focus from the list box has been performed (Step S211: Yes), the control unit 11 proceeds to Step S212.
[0078] (Step S212) The control unit 11 removes the focus from the list box, and proceeds to step S213.
[0079] (Step S213) The control unit 11 sets the state of acceptance of voice input to OFF, and returns the process to step S202.
[0080] In this embodiment, voice input is turned off when the list menu is not focused. Therefore, even if the user speaks when the list menu is not focused, voice input is not performed. Therefore, in order to inform the user that voice input is not possible when the list menu is not focused, a warning indicating "keyword-based input is not accepted" may be displayed on the display unit 13 when voice input is performed when the list menu is not focused.
[0081] Furthermore, even if the user inputs a voice while the list menu is focused, there are cases where the input voice is not included in the keywords to be selected, or the voice cannot be recognized. Therefore, even if the input voice is included in the keywords while the list menu is focused, if the input voice is not included in the keywords or cannot be recognized, a warning indicating that "keyword-based input cannot be accepted" may be displayed on the display unit 13.
[0082] 16 is another example of a flowchart showing a voice input operation according to the first embodiment of the present invention. In FIG. 16, the same processes as those in FIG. 15 are denoted by the same numbers, and the description thereof will be omitted.
[0083] In this example, in step S250, the control unit 11 determines whether a voice input operation has been performed when the list menu is not focused. If a voice input operation has been performed (step S250: Yes), the control unit 11 displays a warning indicating that voice input is not being accepted (step S251). This notifies the user that voice input is not being accepted when the list menu is not focused.
[0084] In this example, if the control unit 11 determines in step S208 that the voice input is not included in the options (step S208: No), it displays a warning indicating that the voice cannot be identified (step S252). As a result, even if a voice input is made with the focus on the list menu, if the input voice is not included in the options that are keywords or if voice recognition is not possible, the user can be informed that the voice cannot be identified.
[0085] <Third Embodiment> In the above-described embodiments, an option is selected by comparing a voice-input character string with character strings included as options. However, it is conceivable to expand the character strings to be compared and selected to include attribute information added to the options.
[0086] 17 to 21 are explanatory diagrams of a voice input operation according to the third embodiment of the present invention. Fig. 17 shows an example of the relationship between character strings that are options and attribute information added to the options.
[0087] In FIG. 17 , restaurant names are options, and menu items are added as attribute information. In this example, "hamburger steak" and "omelette rice" are added as attribute information to the option "Family Restaurant A." The option list is data in which options and attribute information are associated with each other, and may be stored in the storage unit 12. In this case, if the character strings to be subjected to speech recognition are not only the options themselves but also the attribute information added to the options, the system will respond not only to the character string "Family Restaurant A" included as an option in the list, but also to voice input of menu items such as "hamburger steak" and "omelette rice," which are attribute information.
[0088] For example, suppose a list box 501 with restaurant names as options is generated as shown in Fig. 18 from information on restaurant names and menu items as shown in Fig. 17. This list box 501 responds to voice input while the focus is on it. That is, in the list box 501 shown in Fig. 18, if a voice input such as "Family Restaurant A" is made while the focus is on it, the option "Family Restaurant A" is selected and a cursor 502 is displayed on the option "Family Restaurant A," as shown in Fig. 19. The control unit 11 then reads attribute information corresponding to "Family Restaurant A" from the storage unit 12 and displays it in the list, so that the menu items available at "Family Restaurant A" are displayed to the right of the restaurant name.
[0089] The list box 501 also responds to attribute information added to the options. For example, if attribute information such as "ramen" is input by voice while the list box 501 is focused, the control unit 11 responds to the voice information and displays attribute information containing "ramen" as a keyword and the names of restaurants corresponding to the attribute information in the list box 511, as shown in Fig. 20. This allows the control unit 11 to display the names of restaurants that offer "ramen" as a menu item and a list box 511 for selecting the ramen served at the restaurant.
[0090] When such hierarchical information is to be input by voice, the available options may be displayed to the user. For example, a message such as "Please input the name of the restaurant or the menu item you would like to eat" may be displayed on the display screen.
[0091] Furthermore, the hierarchy may be further multi-layered. For example, a hierarchy representing the flavor of ramen may be set below the "ramen" hierarchy, and data representing the flavor may be used as attribute information, and a hierarchy representing the flavor may be displayed, such as "miso," "soy sauce," and "salt." In this case, the storage unit 12 may store data in which options are associated with attribute information divided into multiple hierarchies.
[0092] Here, since the possibility of operating erroneously increases when the number of options displayed increases, it is conceivable to limit the response to the selected item or items below for hierarchical information.
[0093] For example, suppose there is a list box 601 for selecting store information, as shown in Fig. 21. This list box 601 contains information on a hierarchy 602 of store categories such as "clothing," "restaurants," and "miscellaneous goods," and below that, a hierarchy 603 contains information on the names of stores belonging to each category.
[0094] In this embodiment, before an option in the category hierarchy 602 is selected, the voice input responds to the names of stores in the hierarchy below it, 603. In this case, the number of options becomes very large because information on all store names is included.
[0095] When a category level 602 is selected, the voice input only responds to store names in the level 603 below the selected category. For example, when "Restaurant" is selected from the category level 602, the voice input only responds to store names "Family Restaurant A," "Family Restaurant B," and "Dining C" below the level of "Restaurant." In this way, when the control unit 11 extracts a keyword recognized from the input voice that corresponds to an option, the next voice input may be limited to keywords corresponding to character strings included in attribute information belonging to a level below the extracted keyword. The control unit 11 then extracts the keywords received through the voice recognition process as options. By storing hierarchical attribute information associated with the options, it is possible to extract keywords obtained by voice recognition from among keywords included in attribute information belonging to a level below the keyword selected as an option. The attribute information belonging to a level below the keyword may be extracted from keywords included in attribute information one level below the keyword, or keywords included in any level of attribute information belonging to the keyword may be extracted.
[0096] In the above-described embodiments, an option is selected by comparing the voice-input character string with the character strings included as options, but the options included in a pop-up menu or list box are not limited to character strings. Options may also be displayed as icons or illustrations.
[0097] 22 and 23 are explanatory diagrams of a voice input operation according to the fourth embodiment of the present invention. For example, as shown in FIG. 22, shapes represented by circles or triangles may be included in a pop-up menu or list box. In this case, as shown in FIG. 23, a character string representing the pronunciation of each shape may be added as attribute information and stored in the storage unit 12 as reference data. In this case, for example, a circle shape may respond to a voice input such as "circle." The control unit 11 performs voice recognition on the input voice and, from the character strings obtained as a result of the voice recognition, extracts keywords that match the character string representing the attribute information based on the reference data stored in the storage unit 12. The control unit 11 may then accept voice input of only the extracted keywords and extract options (e.g., shapes) corresponding to the extracted keywords by referring to the reference data stored in the storage unit 12.
[0098] The selectable graphic may be a meaningful graphic such as an icon or a photograph. For example, on a restaurant menu, a ramen icon may be associated with the character string "ramen," and when the user says "ramen," the ramen icon may be displayed and the ramen selected.
[0099] For example, in a menu for selecting clothing, clothing may be contrasted with a photo, and when the user utters the sound "checkered sweater," a photo of the checked sweater may be displayed and the clothing may be selected. Furthermore, in addition to icons and photos, a logo may be associated with a character string.
[0100] <Modifications and Applications> Although the embodiments of the present invention have been described in detail with reference to the drawings, the specific configurations are not limited to these embodiments. Furthermore, these embodiments may be combined.
[0101] For example, in the first embodiment described above, in the case of a pop-up menu, voice input is accepted only for options in the pop-up menu while the menu is popped up. Furthermore, in the second embodiment, in the case of a list box, voice input is accepted only for options in the list box while the menu is focused. These may be combined to accept voice input only for options in the pop-up menu while the menu is focused. In this way, the control unit 11 may accept voice input only for keywords included in the list box.
[0102] 24, when the focus is on a pop-up menu 701, the user may be allowed to input voice commands to select an option from the pop-up menu 701. In this way, the control unit 11 may be allowed to input voice commands only for keywords included in the pop-up menu.
[0103] FIG. 25 is a functional block diagram of a voice input device according to another aspect of the present invention.
[0104] As shown in FIG. 25 , a voice input device according to another aspect of the present invention includes a keyword recognition unit 1001 that recognizes keywords from input voice, and an extraction unit 1002 that extracts options from an option list belonging to a UI element to be input based on the extracted keyword.
[0105] All or part of the voice input device in the above-described embodiments may be implemented by a computer. In this case, a program for implementing this function may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into a computer system. Furthermore, "computer-readable recording medium" may also include media that dynamically store programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or telephone lines, or media that store programs for a fixed period of time, such as volatile memory within a computer system that serves as a server or client. The program may also be designed to implement part of the above-described functions, or may be capable of implementing the above-described functions in combination with a program already stored in the computer system, or may be implemented using a programmable logic device such as an FPGA.
[0106] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.
[0107] REFERENCE SIGNS LIST 11 control unit 12 storage unit 13 display unit 15 voice input unit 16 image input unit 20 voice input processing unit 1001 keyword recognition unit 1002 extraction unit
Claims
1. A voice input device having a keyword recognition unit that recognizes a keyword from an input voice, and an extraction unit that extracts an option from an option list belonging to a UI element to be input, based on the extracted keyword.
2. The voice input device according to claim 1, wherein the UI element is a pop-up window, and the extraction unit extracts an option from an option list belonging to the popped-up pop-up window.
3. The voice input device according to claim 1, wherein the UI element is a list box, and the extraction unit extracts an option from an option list belonging to the list box.
4. The voice input device according to claim 1, wherein the UI element to be input is a UI element with focus, and the extraction unit extracts an option from an option list belonging to the UI element with focus.
5. The voice input device according to claim 1, further comprising a UI element detection unit that detects the UI element to be input based on an input operation among the UI elements within the range displayed on the display screen, and the extraction unit extracts an option from an option list belonging to the detected UI element.
6. The voice input device according to claim 5, wherein the UI element detection unit detects a UI element near the position of the mouse pointer as the UI element to be input.
7. The voice input device according to claim 5, further comprising a gaze detection unit that detects the user's gaze, and the UI element detection unit detects a UI element corresponding to the tip of the detected gaze as the UI element to be input.
8. The voice input device according to claim 5, wherein the UI element detection unit detects the UI element to be input from among the UI elements within the display range of the display screen after the display screen is scrolled.
9. The option list includes an option and attribute information associated with the option, and the extraction unit extracts an option or attribute information based on the keyword from among the options or attribute information included in the option list. The voice input device according to claim 1.
10. The voice input device according to claim 9, further comprising a first display control unit that, when the attribute information is extracted, displays the extracted attribute information and the option corresponding to the extracted attribute information on the display screen.
11. The option list has a subordinate option list which is an option list of another hierarchy associated with the options included in the option list, and the extraction unit extracts an option based on the keyword from among the subordinate option lists belonging to the option extracted from the option list. The voice input device according to claim 1.
12. It has a reference data storage unit that stores reference data in which a figure that is the option and the figure and a character string are associated, and the extraction unit refers to the reference data storage unit and extracts, as an option, a figure corresponding to the character string according to the keyword. The voice input device according to claim 1.
13. The voice input device according to claim 1, further comprising a second display control unit that displays the extracted option on a display screen in a display mode different from that of the unextracted option.
14. The keyword recognition unit accepts voice input only for keywords recognized from the input voice and corresponding to the options included in the pop-up menu of the pop-up window, and the extraction unit extracts an option according to the keyword accepted by the keyword recognition unit. The voice input device according to claim 2.
15. The keyword recognition unit accepts voice input only for keywords included in the pop-up window when the pop-up menu of the pop-up window is popped up, and the extraction unit extracts an option according to the keyword accepted by the keyword recognition unit. The voice input device according to claim 2.
16. The keyword recognition unit accepts voice input only for keywords included in the list box of the pop-up window, and the extraction unit extracts an option according to the keyword accepted by the keyword recognition unit. The voice input device according to claim 2.
17. The keyword recognition unit accepts voice input only for keywords recognized from the input voice and corresponding to the character strings stored in the reference data storage unit, and the extraction unit extracts, as an option, a figure corresponding to the character string according to the keyword accepted by the keyword recognition unit. The voice input device according to claim 12.
18. A voice input device according to claim 1, comprising a storage unit that stores data in which options are associated with attribute information divided into a plurality of hierarchies, wherein when the keyword recognition unit extracts a keyword recognized from the input voice and corresponding to the option, the keyword recognition unit accepts only the next voice input for a character string included in the attribute information belonging to the lower layer of the extracted keyword.
19. A voice input device according to any one of claims 1 to 13, further comprising a message display unit that displays a message indicating that an input based on the keyword cannot be accepted when a keyword is recognized in a state where there is no inputtable UI element among the UI elements.
20. A voice input device according to any one of claims 1 to 13, further comprising a message display unit that displays a message indicating that an input based on the keyword cannot be accepted when there is no option corresponding to the recognized keyword.
21. A voice input method including a step of recognizing a keyword from an input voice, and a step of extracting an option from an option list belonging to a UI element to be input based on the extracted keyword.
Citation Information
Patent Citations
Mobile terminal and its menu control method
JP2009252238A
Electronic device and control method thereof
JP5535298B2
Long-distance extension of digital assistant services
JP6606301B1
Information processing device, computer control method and control program
JP6828231B1
System and method for speech-based navigation and interaction with a device's visible screen elements using a corresponding view hierarchy
US9600227B2