Speech processing system, speech processing method, and speech processing program

The voice processing system addresses the challenge of user convenience by displaying operation screens with assistance information, receiving voice inputs, and executing commands, thereby improving the usability of voice-controlled systems.

JP7813511B2Active Publication Date: 2026-02-13SHARP KK
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2020150854
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-09-08
Publication Date
2026-02-13
Estimated Expiration
2040-09-08

AI Technical Summary

Technical Problem

Conventional voice processing systems lack user convenience due to difficulty in grasping which voice commands are recognizable and which parts of the operation screen can be operated by voice commands.

Method used

A voice processing system that displays an operation screen, presents assistance information, receives user voice input, identifies and executes commands, and includes a support information unit to provide operation assistance.

Benefits of technology

Improves user convenience by allowing users to quickly determine which operations can be performed via voice commands and what commands to use, enhancing the overall usability of voice-controlled systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813511000001
    Figure 0007813511000001
  • Figure 0007813511000002
    Figure 0007813511000002
  • Figure 0007813511000003
    Figure 0007813511000003
Patent Text Reader

Abstract

To provide a voice processing system, a voice processing method, and a voice processing program capable of improving the convenience of operation by a voice command.SOLUTION: The voice processing system has a display processing unit that displays an operation screen of an operation target application that is an operation target of a user, a support information presentation unit that presents operation support information for the operation target application in association with the operation screen, a voice receiving unit that receives voice of the user, a command specifying unit that specifies a first command for the operation target application based on the voice received from the voice receiving unit, and a command execution unit that executes the first command specified by the command specifying unit for the operation target application.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice processing system, a voice processing method, and a voice processing program. [Background technology]

[0002] In recent years, a voice processing system has become known that can recognize a user's voice and execute a predetermined command corresponding to the voice. For example, when a document is displayed on a display device by a predetermined application and the user utters a voice command to turn (advance) a page of the document, the voice processing system executes a command to turn the page of the document in response to the voice.

[0003] In the voice processing system, a technique has been proposed in the past to display a list of voice commands that can be voice recognized when voice recognition fails (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 5234160 Summary of the Invention [Problem to be solved by the invention]

[0005] However, with conventional technologies, it is difficult for a user to grasp voice commands that can be recognized by voice prior to the voice recognition process. Furthermore, it is difficult for a user to grasp which parts of an operation screen displayed on a display device can be operated by the voice commands. Thus, conventional voice processing systems have the problem of poor convenience in operation by voice commands.

[0006] An object of the present invention is to provide a voice processing system, a voice processing method, and a voice processing program that can improve the convenience of operations using voice commands. [Means for solving the problem]

[0007] A voice processing system according to one aspect of the present invention is a voice processing system that executes a predetermined command based on a user's voice, and includes: a display processing unit that displays an operation screen of an application to be operated that is the target of operation for the user; an assistance information presentation unit that presents operation assistance information for the application to be operated in association with the operation screen; a voice receiving unit that receives the user's voice; a command identification unit that identifies a first command for the application to be operated based on the voice received by the voice receiving unit; and a command execution unit that executes the first command identified by the command identification unit for the application to be operated.

[0008] Another aspect of the present invention is a voice processing method for executing a predetermined command based on a user's voice, the method being performed by one or more processors, which includes a display step for displaying an operation screen of an application to be operated by the user, an assistance information presentation step for presenting operation assistance information for the application to be operated in association with the operation screen, a voice receiving step for receiving the user's voice, a command identification step for identifying a first command for the application to be operated based on the voice received in the voice receiving step, and a command execution step for executing the first command identified in the command identification step for the application to be operated.

[0009] A voice processing program according to another aspect of the present invention is a voice processing program that executes a predetermined command based on a user's voice, and is a program for causing one or more processors to execute the following steps: a display step that displays an operation screen of an application to be operated by the user; an assistance information presentation step that presents operation assistance information for the application to be operated in association with the operation screen; a voice receiving step that receives the user's voice; a command identification step that identifies a first command for the application to be operated based on the voice received in the voice receiving step; and a command execution step that executes the first command identified in the command identification step for the application to be operated. [Effects of the Invention]

[0010] According to the present invention, there are provided a voice processing system, a voice processing method, and a voice processing program that can improve the convenience of operations using voice commands. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a functional block diagram showing the configuration of a speech processing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing an example of command information used in the voice processing system according to the embodiment of the present invention. [Figure 3] FIG. 3 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 4] FIG. 4 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 5] FIG. 5 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 6] FIG. 6 is a flowchart illustrating an example of a procedure for audio processing in the audio processing system according to the embodiment of the present invention. [Figure 7] FIG. 7 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 8] FIG. 8 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 9] FIG. 9 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 10] FIG. 10 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 11] FIG. 11 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. [Figure 12] FIG. 12 is a diagram showing an example of a display screen displayed on a display device in the speech processing system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings. Note that the following embodiment is an example of the present invention, and does not limit the technical scope of the present invention.

[0013] [Speech processing system 100] FIG. 1 is a diagram showing a schematic configuration of a voice processing system according to an embodiment of the present invention. The voice processing system 100 includes a voice processing device 1, a cloud server 2, and a display device 3. The voice processing device 1 is a microphone / speaker device equipped with a speaker 13 and a microphone 14, and is, for example, an AI speaker or a smart speaker. The voice processing device 1, the cloud server 2, and the display device 3 are connected to each other via a network N1. The network N1 is a communication network such as the Internet, a LAN, a WAN, or a public telephone line. The cloud server 2 is constructed, for example, with one or more data servers (virtual servers). Note that the cloud server 2 may be replaced with a single physical server. The voice processing system 100 is capable of executing predetermined commands based on a user's voice.

[0014] [Speech processing device 1] 1, the voice processing device 1 includes a control unit 11, a storage unit 12, a speaker 13, a microphone 14, and a communication interface 15. The voice processing device 1 is placed on a table, for example, and acquires a user's voice via the microphone 14 and outputs voice to the user from the speaker 13.

[0015] The communication interface 15 is a communication interface for connecting the voice processing device 1 to the network N1 by wire or wirelessly and for executing data communication in accordance with a predetermined communication protocol with other devices (e.g., the cloud server 2 and the display device 3) via the network N1. The communication interface 15 may also be a communication interface capable of realizing a video conference system (described later).

[0016] The storage unit 12 is a non-volatile storage unit such as a flash memory that stores various types of information. The storage unit 12 stores control programs such as a voice processing program for causing the control unit 11 to execute voice processing (see FIG. 6 ), which will be described later. For example, the voice processing program is distributed from the cloud server 2 and stored therein. Alternatively, the voice processing program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, read by a reading device (not shown) such as a CD drive or DVD drive provided in the voice processing device 1, and stored in the storage unit 12.

[0017] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various types of arithmetic processing. The ROM stores in advance control programs such as a BIOS and an OS that cause the CPU to execute various types of processing. The RAM stores various types of information and is used as a temporary storage memory (work area) for the various types of processing executed by the CPU. The control unit 11 controls the audio processing device 1 by having the CPU execute various control programs that are pre-stored in the ROM or the storage unit 12.

[0018] Specifically, the control unit 11 includes various processing units such as an audio receiving unit 111, an audio determining unit 112, and an audio transmitting unit 113. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units included in the control unit 11 may be configured with electronic circuits. The audio processing program may be a program for causing multiple processors to function as the various processing units.

[0019] The voice receiving unit 111 receives voice spoken by a user who uses the voice processing device 1. The voice receiving unit 111 is an example of a voice receiving unit of the present invention. The user speaks, for example, a specific word (also called an activation word or a wake-up word) for the voice processing device 1 to start accepting voice commands, various voice commands (command voices) to instruct the voice processing device 1, and the like. The voice receiving unit 111 receives various voices spoken by the user.

[0020] The voice determination unit 112 determines whether or not the voice includes the specific word based on the voice received by the voice receiving unit 111. For example, the voice determination unit 112 performs voice recognition on the voice received by the voice receiving unit 111 and converts it into text data. Then, the voice determination unit 112 determines whether or not the specific word is included at the beginning of the text data.

[0021] The voice transmitting unit 113 executes a transmission process for the voice received by the voice receiving unit 111, based on the determination result by the voice determining unit 112. Specifically, when the voice determining unit 112 determines that the voice received by the voice receiving unit 111 contains the specific word, the voice transmitting unit 113 transmits text data of a keyword (a command keyword) that is included in the voice and follows the specific word to the cloud server 2. On the other hand, when the voice determining unit 112 determines that the voice received by the voice receiving unit 111 does not contain the specific word, the voice transmitting unit 113 does not transmit the voice to the cloud server 2. As a result, when the specific word is uttered, the command keyword is transmitted to the cloud server 2, thereby preventing voice of normal conversation that does not contain the specific word from being transmitted to the cloud server 2 by mistake.

[0022] [Cloud Server 2] As shown in FIG. 1, the cloud server 2 includes a control unit 21, a storage unit 22, a communication interface 23, and the like.

[0023] The communication interface 23 is a communication interface for connecting the cloud server 2 to the network N1 via a wired or wireless connection and for performing data communication with other devices (e.g., an audio processing device 1, a display device 3) via the network N1 in accordance with a predetermined communication protocol.

[0024] The storage unit 22 is a non-volatile storage unit such as a flash memory that stores various types of information. The storage unit 22 stores control programs such as a voice processing program for causing the control unit 21 to execute voice processing (see FIG. 6), which will be described later. For example, the voice processing program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, and may be read by a reading device (not shown) such as a CD drive or DVD drive provided in the cloud server 2 and stored in the storage unit 22. The storage unit 22 also stores text data of the command keywords received from the voice processing device 1, etc.

[0025] Furthermore, command information D1 is stored in the storage unit 22. FIG. 2 shows an example of the command information D1. Information such as an application to be operated, a voice command, and an effect is registered in the command information D1 in association with one another. The application to be operated is an application that the user operates on the display device 3. The application to be operated may run on the cloud server 2 and receive operations for the display device 3, or may be installed on the display device 3 and run. In this embodiment, the applications to be operated include a "voice application" that starts and ends voice processing that executes voice commands in response to the user's voice, "Power Point" (registered trademark) that can display and edit various documents in a slide format, and "Pensoft" that can be written on a touch panel with a touch pen or the like.

[0026] The voice command is a command that can be executed in the voice processing system 100, and is registered for each application to be operated. The voice command corresponds to the command keyword. The effect is information indicating the operation content to be executed by the voice command. For example, when the first page of a document is displayed on the display device 3 using "Power Point," if the user utters the voice command (command keyword) "Move to next page," the voice processing system 100 executes the voice command, and the second page of the document is displayed on the display device 3.

[0027] In another embodiment, part or all of the information in the command information D1 may be stored in either the voice processing device 1 or the display device 3, or may be stored in a distributed manner in these multiple devices. In another embodiment, the information may be stored in a server accessible from the voice processing system 100. In this case, the voice processing system 100 may acquire the information from the server and perform various processes such as the voice processing described below (see FIG. 6).

[0028] The control unit 21 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various types of arithmetic processing. The ROM stores in advance control programs such as a BIOS and an OS that cause the CPU to execute various types of processing. The RAM stores various types of information and is used as a temporary storage memory (work area) for the various types of processing executed by the CPU. The control unit 21 controls the cloud server 2 by having the CPU execute various control programs that are pre-stored in the ROM or the storage unit 22.

[0029] 1, the control unit 21 includes various processing units such as a voice receiving unit 211, a command identification unit 212, and a command processing unit 213. The control unit 21 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units included in the control unit 21 may be configured with electronic circuits. The control program may be a program for causing a plurality of processors to function as the various processing units.

[0030] The voice receiving unit 211 receives the command keyword corresponding to the voice command transmitted from the voice processing device 1. The command keyword is a word (text data) following a specific word included at the beginning of text data of the voice received by the voice processing device 1. Specifically, when the voice processing device 1 detects the specific word and transmits the command keyword to the cloud server 2, the cloud server 2 receives the command keyword.

[0031] The command identification unit 212 identifies a voice command based on the command keyword received by the voice receiving unit 211. The command identification unit 212 is an example of the command identification unit 212 of the present invention. For example, the command identification unit 212 identifies a voice command corresponding to the command keyword with reference to command information D1 (see FIG. 2). When the user utters the command keyword corresponding to a predetermined voice command for the application to be operated, the command identification unit 212 identifies a voice command (corresponding to a first command of the present invention) for the application to be operated based on the command keyword. The command identification unit 212 is an example of the command identification unit of the present invention.

[0032] In this embodiment, a plurality of voice commands are registered in the command information D1 in advance, and the voice command that matches the command keyword is identified from the command information D1, but the method of identifying the voice command is not limited to this. For example, the command identification unit 212 may identify the voice command by interpreting the meaning of the user's instruction based on predetermined terms included in the command keyword, the phrases and syntax of the entire command keyword, etc. For example, the command identification unit 212 may identify the voice command from the command keyword using known methods such as morphological analysis, syntactic analysis, semantic analysis, and machine learning.

[0033] The command processing unit 213 stores information about the voice command identified by the command identification unit 212 in a command storage area (queue) corresponding to the display device 3. For example, the storage unit 22 includes one or more command storage areas corresponding to the display device 3. Here, the storage unit 22 includes a queue K1 corresponding to the display device 3. Note that if the voice processing system 100 includes multiple display devices 3, the storage unit 22 may store a queue for each display device 3.

[0034] For example, the command processing unit 213 stores information on the voice command “Move to next page” identified by the command identification unit 212 in the queue K1 corresponding to the display device 3.

[0035] The data (voice commands) stored in the queue K1 are retrieved by the display device 3 corresponding to the queue K1, and the display device 3 executes the voice commands.

[0036] [Display device 3] As shown in FIG. 2, the display device 3 includes a control unit 31, a storage unit 32, an operation unit 33, a display unit 34, a communication interface 35, and the like.

[0037] The operation unit 33 is a mouse, keyboard, touch panel, or the like that accepts operations by the user of the display device 3. The display unit 34 is a display panel such as a liquid crystal display or an organic EL display that displays various information. The operation unit 33 and the display unit 34 may be an integrated user interface.

[0038] The communication interface 35 is a communication interface for connecting the display device 3 to the network N1 via a wired or wireless connection and for performing data communication with other devices (e.g., an audio processing device 1, a cloud server 2) via the network N1 in accordance with a predetermined communication protocol.

[0039] The storage unit 32 is a non-volatile storage unit such as a flash memory that stores various types of information. The storage unit 32 stores control programs such as an audio processing program for causing the control unit 31 to execute audio processing (see FIG. 6 ), which will be described later. For example, the audio processing program may be non-temporarily recorded on a computer-readable recording medium such as a CD or DVD, read by a reading device (not shown) such as a CD drive or DVD drive provided in the display device 3, and stored in the storage unit 32.

[0040] The control unit 31 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various types of arithmetic processing. The ROM stores in advance control programs such as a BIOS and an OS that cause the CPU to execute various types of processing. The RAM stores various types of information and is used as a temporary storage memory (work area) for the various types of processing executed by the CPU. The control unit 31 controls the display device 3 by having the CPU execute various control programs that are pre-stored in the ROM or the storage unit 32.

[0041] Specifically, the control unit 31 includes various processing units such as an operation reception unit 311, a display processing unit 312, a command acquisition unit 313, a command execution unit 314, and a support information presentation unit 315. The control unit 31 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units included in the control unit 31 may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the various processing units.

[0042] The operation accepting unit 311 accepts various operations from the user. Specifically, the operation accepting unit 311 accepts operations on the operation unit 33 by the user. For example, the operation accepting unit 311 accepts an operation to start a predetermined application (such as an application to be operated), an operation on an operation screen operated by the application to be operated, an operation to open a predetermined file, etc. The operation accepting unit 311 also accepts an operation from the user to request the presentation of operation support information, which will be described later.

[0043] The display processing unit 312 displays various types of information on the display unit 34. For example, the display processing unit 312 displays an operation screen of the operation target application that is the user's operation target on the display unit 34. FIGS. 3 and 4 show examples of the operation screen displayed on the display unit 34. In the example shown in FIG. 3, an operation screen of the operation target application AP1 of "voice application" and an operation screen of the operation target application AP2 of "Power Point" are displayed. In the example shown in FIG. 4, an operation screen of the operation target application AP1, an operation screen of the operation target application AP2, and an operation screen of the operation target application AP3 of "Pensoft" are displayed.

[0044] A list of a plurality of files F1 that can be displayed is displayed on the operation screen of the operation target application AP1. The user can specify a desired file from the list by voice or the like. An operation button B1 for requesting the presentation of the operation support information is also displayed on the operation screen of the operation target application AP1. When requesting the presentation of the operation support information, the user selects (presses) the operation button B1 with a finger, a touch pen, a mouse, or the like.

[0045] The command acquisition unit 313 acquires a voice command stored in a command storage area (queue K1) of the cloud server 2. Specifically, the command acquisition unit 313 monitors the queue K1 corresponding to the display device 3, and acquires a voice command when the voice command is stored in the queue K1. For example, when the operation button B1 is pressed, the command acquisition unit 313 periodically (for example, every 5 seconds) queries the queue K1 to acquire the voice command. Note that the command processing unit 213 of the cloud server 2 may transmit data related to the voice command to the display device 3, and the command acquisition unit 313 may acquire the voice command.

[0046] The command execution unit 314 executes the voice command identified by the command identification unit 212 of the cloud server 2 for the operation target application. The command execution unit 314 is an example of a command execution unit of the present invention. Specifically, the command execution unit 314 executes the voice command acquired by the command acquisition unit 313. For example, the command execution unit 314 executes the voice command acquired by the command acquisition unit 313 from the queue K1.

[0047] For example, when the first page of a document is displayed on the display unit 34 of the display device 3 using "Power Point," if the user utters the voice command (command keyword) "Move to next page," the command execution unit 314 executes the voice command acquired from the queue K1 by the command acquisition unit 313. As a result, the second page of the document is displayed on the display unit 34 of the display device 3.

[0048] Here, in each operation screen shown in Figures 3 and 4, it is difficult for the user to grasp at a glance which operation screen of the application to be operated can be operated by voice commands, and what voice commands can be used to operate the operation screen.

[0049] Therefore, the support information presenting unit 315 presents information (operation support information) for supporting user operation to the user operating the operation screen. Specifically, the support information presenting unit 315 presents the operation support information for the operation target application in association with the operation screen. Furthermore, the support information presenting unit 315 may present the operation support information when the operation receiving unit 311 receives an operation from the user requesting presentation of the operation support information. For example, the support information presenting unit 315 may present the operation support information when the user presses the operation button B1 on the operation screen shown in FIG. 4. Furthermore, for example, the support information presenting unit 315 may present the operation support information when the user utters a voice to start voice processing and the voice receiving unit 211 of the cloud server 2 receives the voice. The support information presenting unit 315 is an example of a support information presenting unit of the present invention.

[0050] FIG. 5 shows an example of the operation screen including the operation support information. FIG. 5 shows the operation support information corresponding to the operation screen of FIG. 4. The support information presenting unit 315 presents the operation support information corresponding to one or more commands for the target application in association with the operation screen. For example, as shown in FIG. 5, the support information presenting unit 315 presents operation support information H1 corresponding to a voice command for a target application AP1 of "voice application" in association with the operation screen of the target application AP1. Furthermore, the support information presenting unit 315 presents operation support information H2 corresponding to a voice command for a target application AP2 of "Power Point" in association with the operation screen of the target application AP2. Furthermore, the support information presenting unit 315 presents operation support information H3 corresponding to a voice command for a target application AP3 of "Pensoft" in association with the operation screen of the target application AP3. Each of the operation support information H1, H2, and H3 is composed of a speech bubble object image and text information of the voice command. The support information presenting unit 315 displays each piece of operation support information H1 so that at least a portion thereof overlaps with the operation screen of the operation target application AP1, displays each piece of operation support information H2 so that at least a portion thereof overlaps with the operation screen of the operation target application AP2, and displays each piece of operation support information H3 so that at least a portion thereof overlaps with the operation screen of the operation target application AP3. Furthermore, when there are multiple pieces of operation support information for the operation screen, the support information presenting unit 315 displays the multiple pieces of operation support information side by side.

[0051] When the user presses the operation button B1 again, the support information presenting unit 315 may erase (hide) all of the operation support information.

[0052] According to this configuration, for example, the user can see at a glance that they can operate each operation screen of the target applications AP1, AP2, and AP3, and can also see at a glance the types (contents) of voice commands that can be executed on each operation screen.

[0053] [Audio Processing] Hereinafter, an example of the procedure of the audio processing executed by the control unit 11 of the audio processing device 1, the control unit 21 of the cloud server 2, and the control unit 31 of the display device 3 will be described with reference to FIG.

[0054] The present invention can be understood as an invention of a voice processing method that executes one or more steps included in the voice processing. Furthermore, one or more steps included in the voice processing described herein may be omitted as appropriate. Furthermore, the steps in the voice processing may be executed in a different order as long as the same operational effect is achieved. Furthermore, while the description here takes as an example a case where the steps in the voice processing are executed by control units 11, 21, and 31, in other embodiments, the steps in the voice processing may be executed in a distributed manner by one or more processors.

[0055] Here, for example, it is assumed that the operation screens shown in FIG. 4 are displayed on the display unit 34 of the display device 3, and the user is able to operate the operation screens of the applications to be operated by voice.

[0056] In step S11, the control unit 31 determines whether or not the operation target application that can be operated by the user exists on the display device 3. If the operation target application exists (S11: Yes), the process proceeds to step S12. On the other hand, if the operation target application does not exist (S11: No), the process proceeds to step S14. For example, as shown in FIG. 4, when an operation screen of at least one operation target application is displayed on the display device 3, the control unit 31 determines that the operation target application exists.

[0057] In step S12, the control unit 31 of the display device 3 determines whether or not an operation requesting the presentation of the operational assistance information has been received from the user. If an operation requesting the presentation of the operational assistance information has been received from the user (S12: Yes), the process proceeds to step S13. On the other hand, if an operation requesting the presentation of the operational assistance information has not been received from the user (S12: No), the process proceeds to step S14. For example, when the user presses the operation button B1 on the operation screen shown in FIG. 4, the control unit 31 determines that an operation requesting the presentation of the operational assistance information has been received from the user. Note that the operation button B1 may be displayed within the operation screen of any of the target applications, or may be displayed outside the operation screen of the target application.

[0058] In step S13, the control unit 31 presents information (operation support information) for supporting the user operation to the user operating the operation screen. Specifically, the control unit 31 presents the operation support information for the application to be operated in association with the operation screen.

[0059] 5, the control unit 31 presents operation support information H1 corresponding to a voice command for a target application AP1 of "voice application" in association with the operation screen of the target application AP1, presents operation support information H2 corresponding to a voice command for a target application AP2 of "Power Point" in association with the operation screen of the target application AP2, and presents operation support information H3 corresponding to a voice command for a target application AP3 of "Pensoft" in association with the operation screen of the target application AP3. Step S13 is an example of a support information presenting step of the present invention.

[0060] In step S14, the control unit 11 of the voice processing device 1 determines whether or not the user's voice has been received. If the control unit 11 has received the user's voice (S14: Yes), the process proceeds to step S15. On the other hand, if the control unit 11 has not received the user's voice (S14: No), the process returns to step S11. Step S14 is an example of a voice receiving step of the present invention.

[0061] In step S15, the control unit 11 determines whether the specific word is included in the received voice based on the received voice. For example, the control unit 11 performs voice recognition on the received voice and converts it into text data, and determines whether the specific word is included at the beginning of the text data. If the specific word is included in the voice (S15: Yes), the process proceeds to step S16. If the specific word is not included in the voice (S15: No), the process returns to step S11.

[0062] In step S16, the control unit 11 transmits to the cloud server 2 text data of a keyword (command keyword) that is included in the voice and follows the specific word.

[0063] Next, in step S17, the control unit 21 of the cloud server 2 receives the command keyword transmitted from the voice processing device 1 and identifies a voice command based on the command keyword. For example, the control unit 21 identifies a voice command corresponding to the command keyword with reference to command information D1 shown in Fig. 2. Step S17 is an example of a command identification step of the present invention.

[0064] Next, in step S18, the control unit 11 stores information on the identified voice command in the queue K1 corresponding to the display device 3.

[0065] Next, in step S19, the control unit 31 of the display device 3 executes the voice command identified for the application to be operated. Specifically, the control unit 31 acquires the voice command from the queue K1 corresponding to the display device 3 and executes the voice command. Step S19 is an example of a command execution step of the present invention. In this way, the voice processing system 100 executes the voice processing.

[0066] As described above, the voice processing system 100 according to this embodiment displays an operation screen of a target application that is the user's operation target, and presents operation assistance information for the target application in association with the operation screen. The voice processing system 100 also receives the user's voice, identifies a first command for the target application based on the voice, and executes the first command for the target application. This allows the user to quickly determine which operation screens can be operated by voice commands, what voice commands can be used to operate the operation screens, and so on. This improves the convenience of operations using voice commands.

[0067] The present invention is not limited to the above-described embodiment, and other embodiments of the present invention will be described below.

[0068] Here, when a plurality of operation screens corresponding to the same application to be operated are displayed on the display device 3, it becomes difficult for the user to know at a glance which operation screen can be operated by a voice command, what voice command can be used to operate the operation screen, etc. For example, as shown in Fig. 7, when two operation screens of the application AP2 to be operated, "Power Point," are displayed on the display device 3, it becomes difficult for the user to know at a glance which operation screen can be operated by a voice command, what voice command can be used to operate the operation screen, etc.

[0069] Therefore, in the voice processing system 100 according to another embodiment, when a plurality of operation screens corresponding to the same application to be operated are displayed on the display device 3, the control unit 31 (support information presenting unit 315) of the display device 3 presents screen identification information capable of identifying the plurality of operation screens, in association with each of the operation screens. The screen identification information is an example of the operation support information of the present invention. For example, as shown in FIG. 8, the control unit 31 displays screen identification information H21 with a red frame (shown by a "bold line" in FIG. 8 for convenience) on one operation screen, and displays screen identification information H31 with a blue frame (shown by a "dotted line" in FIG. 8 for convenience) on the other operation screen. This allows, for example, a user to identify an operation screen on which a voice command is to be executed from the two operation screens by the screen identification information, and to specify the operation screen by the screen identification information. For example, when a user utters the voice command (command keyword) "Move to next page by red," the operation screen at the top of the figure is designated, and the voice command for that operation screen is identified, causing the page of the document displayed on that operation screen to be turned to the next page.

[0070] When the user presses the operation button B1, for example, the control unit 31 displays the screen identification information H21 and H31.

[0071] Furthermore, when the user presses, for example, operation button B1, the control unit 31 may display, in addition to screen identification information H21 and H31, operation support information H1, H2, and H3, which are composed of a speech bubble object image and text information of the voice command, as shown in FIG. 9.

[0072] Furthermore, the screen identification information is not limited to identification information corresponding to a color, and may be identification information corresponding to a number, as shown in Figures 10 and 11. In this case, for example, when a user utters a voice command (command keyword) of "Move to next page by two," the operation screen at the bottom in the figure is designated, and the voice command for that operation screen is identified. Furthermore, the screen identification information may be identification information corresponding to the position of the operation screen (upper side, lower side, left side, right side, etc.), the line type of the outer frame, and the line width.

[0073] In another embodiment, the control unit 31 (support information presenting unit 315) of the display device 3 may present, on the operation screen, text information (operation support information) corresponding to one or more voice commands that the command executing unit 314 can currently execute, in an identifiable manner. For example, in the example shown in Fig. 12, when the last page of a document is displayed on the operation screen of the target application AP2 of "Power Point," there is no next page, and therefore the command executing unit 314 cannot execute the voice command "Move to next page." Furthermore, the support information presenting unit 315 deletes (hides) the operation support information H2 corresponding to the voice command "Move to next page," and presents only the operation support information H2 corresponding to the voice command that can currently be executed.

[0074] Also, in FIG. 12, if there are no executable voice commands on the operation screen of the target application AP3, "Excel," the support information presenting unit 315 may present operation support information H33 indicating that voice commands for the operation screen of the target application AP3 are not accepted.

[0075] Furthermore, as another embodiment, the control unit 31 (support information presenting unit 315) of the display device 3 may present, on the operation screen, only operation support information corresponding to a voice command that is used more frequently than a predetermined frequency out of one or more voice commands in a identifiable manner. Furthermore, the support information presenting unit 315 may present, on the operation screen, only operation support information corresponding to a predetermined number (for example, five) of voice commands that are most frequently used out of one or more voice commands in a identifiable manner.

[0076] In another embodiment, the control unit 31 (support information presenting unit 315) of the display device 3 may present, in the plurality of pieces of operation support information shown in Fig. 5, operation support information corresponding to the next voice command that the user can operate, the next voice command that the user cannot operate, and the voice command that the user may operate, in a identifiable manner on the operation screen. For example, on the operation screen of the operation target application AP2 of "Power Point", the support information presenting unit 315 may blink the operation support information H2 corresponding to the next voice command "Move to next page" that can be operated, and may gray out the operation support information H2 corresponding to the next voice command "Move to previous page" that cannot be operated. In this way, candidates for the next operation content may be suggested to the user.

[0077] In another embodiment, the control unit 31 (support information presenting unit 315) of the display device 3 may display the operation support information in association with the position of the operation target. For example, when an operation button (object image) for advancing a page is displayed on the operation screen of the operation target application AP2, the support information presenting unit 315 displays a part (speech bubble portion) of the balloon object image of the operation support information so as to overlap the operation button. This allows the user to easily understand the command keyword (command voice) corresponding to the content of the operation he or she wishes to perform.

[0078] The voice processing system of the present invention can be applied to a video conference system. For example, the voice processing system 100 includes a first voice processing device 1 and a first display device 3 arranged in a first conference room, and a second voice processing device 1 and a second display device 3 arranged in a second conference room. The first voice processing device 1 and the first display device 3, the second voice processing device 1 and the second display device 3, and the cloud server 2 are connected to each other via a network N1, thereby realizing a video conference in the first conference room and the second conference room. In the video conference, for example, the display processing unit 312 of the first display device 3 displays two operation screens of the target application AP2 of "Power Point" (see FIG. 8, etc.). The display processing unit 312 of the second display device 3 also displays two operation screens similar to those of the first display device 3, i.e., two operation screens of the target application AP2 of "Power Point". In this case, the support information presenting unit 315 of the first display device 3 displays screen identification information H21 and H31, which can identify the two operation screens, in association with the respective operation screens on the first display device 3. Similarly, the support information presenting unit 315 of the second display device 3 displays screen identification information H21, H31 that can identify the two operation screens in association with the respective operation screens on the second display device 3. In this way, each of the multiple display devices 3 constituting the video conference system executes the above-mentioned processes by the control unit 31. This makes it possible to improve the convenience of operations by voice commands for each user participating in the video conference.

[0079] Furthermore, the voice processing system of the present invention can be constructed by freely combining the above-described embodiments within the scope of the invention described in each claim, or by appropriately modifying or omitting parts of each embodiment. [Explanation of symbols]

[0080] 1: Audio processing device 2: Cloud server 3:Display device 100: Audio processing system 111: Audio receiving unit 112: Audio determination unit 113: Audio transmission unit 211: Audio receiving unit 212: Command specification section 213: Command processing section 311: Operation reception section 312: Display processing unit 313: Command acquisition unit 314: Command execution section 315: Support information presentation section AP1: Application to be operated AP2: Application to be operated AP3: Application to be operated B1: Operation buttons H1: Operation support information H2: Operation support information H3: Operation support information

Claims

1. A voice processing system that executes a predetermined command based on a user's voice, a display processing unit that displays an operation screen of an operation target application that is an operation target of the user; a support information presenting unit that presents a plurality of pieces of operation support information for the operation target application in association with the operation screen; a voice receiving unit for receiving the user's voice; a command specifying unit that specifies a command for the application to be operated based on the voice received by the voice receiving unit; a command execution unit that executes the command identified by the command identification unit for the operation target application; Equipped with The support information presentation unit displays, in the application to be operated, the plurality of pieces of operation support information corresponding to each of the plurality of commands that can be executed based on the user's voice, in association with the operation screen, and when a first screen is displayed, displays, in the same display mode, first operation support information corresponding to a first command that the user can operate next on the first screen and second operation support information corresponding to a second command that the user can operate next on the first screen, and when transitioning from the first screen to a second screen, displays, in different display modes, the first operation support information corresponding to the first command that the user can operate next on the second screen and the second operation support information corresponding to the second command that the user cannot operate next on the second screen.

2. the support information presenting unit causes the plurality of pieces of operation support information corresponding to the commands for the operating target application to be displayed on the operation screen in association with each other, identifies the first command and the second command based on an operation state of the operating target application, and displays the first operation support information corresponding to the first command and the second operation support information corresponding to the second command in a display mode according to the operation state. The audio processing system of claim 1 .

3. the support information presenting unit presents text information of a plurality of specific words corresponding to the plurality of commands, in association with the operation screen; 3. The speech processing system according to claim 1 or 2.

4. the support information presenting unit presents the text information corresponding to a command currently executable by the command executing unit among the plurality of commands in an identifiable manner associated with the operation screen. The audio processing system of claim 3 .

5. When the display processing unit displays a plurality of the operation screens corresponding to the same operation target application, the support information presenting unit presents screen identification information capable of identifying the plurality of operation screens in association with each of the operation screens; The speech processing system according to any one of claims 1 to 4.

6. the display processing unit displays the plurality of operation screens corresponding to the same operation target application on a first display device and a second display device that are communicably connected to each other via a network, the support information presenter presents screen identification information capable of identifying the plurality of operation screens in association with each of the operation screens on the first display device and the second display device, respectively. The speech processing system according to any one of claims 1 to 4.

7. further comprising an operation receiving unit that receives a predetermined operation by the user; the support information presenting unit presents the operation support information when the operation accepting unit accepts an operation requesting presentation of the operation support information from the user.

7. A speech processing system according to claim 1.

8. the assistance information presenting unit presents the operation assistance information when the voice of the user is received from the voice receiving unit.

7. A speech processing system according to claim 1.

9. 1. A voice processing method for executing a predetermined command based on a user's voice, comprising: a display step of displaying an operation screen of an operation target application that is an operation target of the user; a support information presenting step of presenting a plurality of pieces of operation support information for the operation target application in association with the operation screen; a voice receiving step of receiving the voice of the user; a command specifying step of specifying a command for the application to be operated based on the voice received in the voice receiving step; a command execution step of executing the command identified in the command identification step for the operation target application; executed by one or more processors, a voice processing method in which, in the support information presentation step, in the application to be operated, the plurality of pieces of operation support information corresponding to each of the plurality of commands that can be executed based on the user's voice are displayed in association with the operation screen, and when a first screen is displayed, first operation support information corresponding to a first command that the user can operate next on the first screen and second operation support information corresponding to a second command that the user can operate next on the first screen are displayed in the same display mode, and when a transition is made from the first screen to a second screen, the first operation support information corresponding to the first command that the user can operate next on the second screen and the second operation support information corresponding to the second command that the user cannot operate next on the second screen are displayed in different display modes.

10. A voice processing program that executes a predetermined command based on a user's voice, a display step of displaying an operation screen of an operation target application that is an operation target of the user; a support information presenting step of presenting a plurality of pieces of operation support information for the operation target application in association with the operation screen; a voice receiving step of receiving the voice of the user; a command specifying step of specifying a command for the application to be operated based on the voice received in the voice receiving step; a command execution step of executing the command identified in the command identification step for the operation target application; A speech processing program for causing one or more processors to execute the above. a voice processing program in which, in the support information presentation step, in the application to be operated, the plurality of pieces of operation support information corresponding to each of the plurality of commands that can be executed based on the user's voice are displayed in association with the operation screen, and when a first screen is displayed, first operation support information corresponding to a first command that the user can operate next on the first screen and second operation support information corresponding to a second command that the user can operate next on the first screen are displayed in the same display mode, and when transitioning from the first screen to a second screen, the first operation support information corresponding to the first command that the user can operate next on the second screen and the second operation support information corresponding to the second command that the user cannot operate next on the second screen are displayed in different display modes.

Citation Information

Patent Citations

  • Breather valve

    JP1977034160A

  • Window management method

    JP1994149528A

  • Multiwindow display controller

    JP1995200235A

  • Method and device for performing user function by using voice recognition

    JP2013143151A

  • Electronic device and display program

    JP2014128931A