Information processing device, information processing method and program

JP2025020383A5Pending Publication Date: 2025-07-23CASIO COMPUTER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024198001
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

In the prior art, users need to continuously issue wake word to operate voice recognition devices, resulting in inconvenience and trouble, especially when they are busy.

Method used

After detecting the wake word, the corresponding control process is performed and the user is allowed to no longer issue the wake word for a certain period of time, so as to continue inputting the voice command.

Benefits of technology

Users can no longer issue wake word for a specific time after executing voice commands, which improves operation convenience and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To allow a user to omit uttering wake words when giving instructions to a device.SOLUTION: The information processing device 100 is provided with a voice acquisition unit for acquiring a voice signal and a control unit 110. The control unit 110, when it is determined that first recognition data derived from the voice signal includes a wake word, executes first control processing corresponding to the control information, if it is determined that the control information that is information related to the second recognition data derived from the voice signal after the wake word is included, and executes second control processing corresponding to the control information included in the third recognition data, if it is determined that the control information is included in the third recognition data derived from the voice signal acquired by the voice acquisition unit, during the first period after the predetermined conditions are met or while the first control processing is being executed.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] In a voice recognition device such as a smart speaker or a smartphone, when a user utters a so-called wake word, the device can respond to the user's subsequent voice instructions. For example, the device can reply to the user's voice or start various application programs according to the user's instructions. Patent Document 1 also discloses a technology that allows multiple cloud services to be used separately using multiple wake words. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2019-86535 A Summary of the Invention [Problem to be solved by the invention]

[0004] The technology disclosed in Patent Document 1 enables multiple cloud services to be used selectively by transmitting voice data after a first wake word to a first platform and voice data after a second wake word to a second platform. However, with conventional technologies including the technology disclosed in Patent Document 1, the user is always required to utter the wake word in order to give voice instructions to these devices.

[0005] The present invention has been made in consideration of the above-mentioned situation, and aims to provide an information processing device, an information processing method, and a program that enable a user to omit speaking a wake word when issuing instructions to a device. [Means for solving the problem]

[0006] In order to achieve the above object, one aspect of the information processing device according to the present invention is A voice acquisition unit that acquires a voice signal; A control unit, The control unit is If it is determined that the first recognition data derived from the voice signal includes a wake word, if it is determined that second recognition data derived from the voice signal after the wake word includes control information, the control information being information related to a control process, execute a first control process according to the control information; When it is determined that the control information is included in third recognition data derived from a voice signal acquired by the voice acquisition unit during a first period after a predetermined condition is satisfied or during execution of the first control process, execute a second control process according to control information included in the third recognition data; It is characterized by: Effect of the Invention

[0007] According to the present invention, it is possible for a user to omit uttering a wake word when issuing an instruction to a device. [Brief description of the drawings]

[0008] [Figure 1] 1 is a block diagram showing an example of a functional configuration of an information processing device according to a first embodiment. [Diagram 2] 4 is a diagram showing an example of an operation performed when a user gives a voice instruction to the information processing device according to the first embodiment. FIG. [Diagram 3] FIG. 11 is a diagram showing another example of the operation when a user gives a voice instruction to the information processing device in accordance with the first embodiment. [Figure 4] 4 is a diagram showing an example of how the amount of elapsed time is displayed in the information processing device according to the first embodiment. FIG. [Diagram 5] 4 is an example of a flowchart of a voice command recognition process according to the first embodiment. [Figure 6] FIG. 11 is a diagram showing an example of an operation performed when a user gives a voice instruction to the information processing device according to the second embodiment. [Figure 7] FIG. 11 is a diagram showing an example of an extraction parameter table according to the second embodiment. [Figure 8] 13 is an example of a flowchart of a voice command recognition process according to the second embodiment. [Figure 9] FIG. 11 is a diagram showing an example of an operation performed when a user gives a voice instruction to the information processing device according to the third embodiment. [Figure 10] FIG. 11 is a diagram showing an example of a behavior table according to the third embodiment. [Figure 11] 13 is an example of a first portion of a flowchart of a voice command recognition process according to the third embodiment. [Figure 12] 13 is an example of a second part of the flowchart of the voice command recognition process according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] An information processing device and the like according to an embodiment will be described with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals.

[0010] (Embodiment 1) The information processing device according to the first embodiment is an electronic device, such as a smartphone, that allows a user to give various instructions (such as starting various application programs) by voice.

[0011] As shown in FIG. 1, the information processing device 100 includes a control unit 110, a storage unit 120, an input unit 130, an output unit 140, a communication unit 150, and a sensor unit 160.

[0012] The control unit 110 is configured with a processor such as a CPU (Central Processing Unit). The control unit 110 executes processes for implementing various functions of the smartphone and voice command recognition processes (described later) using programs stored in the storage unit 120. The control unit 110 also supports multi-threading and can execute multiple processes in parallel.

[0013] The storage unit 120 stores programs executed by the control unit 110 and necessary data. The storage unit 120 may include, but is not limited to, a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, etc. Note that the storage unit 120 may be provided inside the control unit 110.

[0014] The input unit 130 is a user interface such as a microphone, a push button switch, a touch panel, etc., and accepts operational input from a user. When the input unit 130 includes a touch panel, the touch panel may be integrated with the display of the output unit 140. The microphone functions as a voice acquisition unit that acquires a voice signal.

[0015] The output unit 140 includes a display such as a liquid crystal display or an organic EL (Electro-Luminescence) display, and displays a display screen, an operation screen, and the like that provide the functions of the information processing device 100. The output unit 140 also includes a sound output means such as a speaker, and can read out e-mails, for example. The output unit 140 may also include a vibrator that generates vibrations.

[0016] The communication unit 150 is a network interface compatible with, for example, a wireless LAN (Local Area Network), LTE (Long Term Evolution), etc. The information processing device 100 can communicate with the Internet and other information processing devices via the communication unit 150.

[0017] The sensor unit 160 includes devices for detecting various values ​​related to the user's movements and the surrounding environment, such as a heart rate sensor, a temperature sensor, an air pressure sensor, an acceleration sensor, a gyro sensor, a GPS (Global Positioning System) device, etc. The control unit 110 can acquire values ​​detected by each device included in the sensor unit 160 as detection values ​​at any timing. However, the sensor unit 160 does not have to include all of the above sensors, and may not include, for example, a temperature sensor or an air pressure sensor.

[0018] The heart rate sensor detects a pulse wave, for example, by a PPG (Photoplethysmography) sensor equipped with an LED (Light Emitting Diode) and a PD (Photodiode). The control unit 110 can acquire the heart rate by measuring the pulse rate (heart rate) per unit time (for example, one minute) based on the pulse wave detected by the heart rate sensor. The temperature sensor includes, for example, a thermistor and can measure body temperature. The air pressure sensor includes, for example, a piezo-resistance type IC (Integrated Circuit) and can measure the surrounding air pressure.

[0019] The acceleration sensor detects acceleration in each direction of three orthogonal axes (X-axis, Y-axis, Z-axis) of the information processing device 100. The gyro sensor detects angular velocity of rotation about each of the three orthogonal axes (X-axis, Y-axis, Z-axis) of the information processing device 100. The GPS device acquires the current position of the information processing device 100 (e.g., three-dimensional data of latitude, longitude, and altitude).

[0020] When issuing a voice instruction to the information processing device 100, the user basically utters a key phrase called a wake word (such as "OK Google" or "Hey Siri") and then utters the content of the instruction. By having the user utter the wake word, the information processing device 100 is prevented from erroneously recognizing voices that are not instructions for the information processing device 100 (for example, conversations between family members, sounds from a television, etc.).

[0021] However, in situations where it is clear that instructions will be given to the information processing device 100 by voice (for example, situations where instructions are expected to be given continuously), uttering the wake word is an unnecessary hassle for the user. Also, there are information processing devices that accept voice instructions after pressing a button instead of uttering the wake word, but this is inconvenient when your hands are dirty and you do not want to touch the screen or buttons, such as when cooking. Therefore, in situations where it is expected that instructions will be received by voice, the information processing device 100 accepts voice instructions without the need for a wake word.

[0022] For example, in the example shown in FIG. 2, the user first utters a wake word (in this example, "Hey, smartphone") and then utters "Tell me when 5 minutes are up," causing the information processing device 100 to start a timer. Note that the instruction uttered by the user after the wake word is also called a voice command. "Tell me when 5 minutes are up" is an example of a voice command. In addition, since the information processing device 100 executes some control process (e.g., an application program) in response to the voice command, the voice command is also called control information related to the control process. In the example shown in FIG. 2, the information processing device 100 recognizes the voice command, starts a timer as an application program, sets the timer time to 5 minutes, and starts the set timer.

[0023] Returning to Fig. 2, after five minutes have elapsed, the information processing device 100 emits a beep to notify the user that five minutes have elapsed, and accepts the next instruction (voice command) without the need for a wake word during a predetermined period (e.g., one minute) after the timer has finished executing. In this example, the user issues an instruction to "read out the next step" without a wake word, and the information processing device 100 accepts the instruction and reads out the text sentence set as the next step.

[0024] In this example, since it is considered likely that the user will again issue some kind of instruction to the information processing device 100 after the timer has finished executing, the information processing device 100 accepts instructions (voice commands) without the need for a wake word for a predetermined period of time after the timer has finished executing.

[0025] In the example shown in FIG. 3, the user first utters the wake word, and then utters "Play the first step of cooking", causing the information processing device 100 to start playing a video (image) corresponding to the first step in a cooking commentary video. The information processing device 100 is assumed to be configured to pause playback when it detects a chapter attached to the end of the video corresponding to the first step in the commentary video. Note that a chapter represents a division attached to a change in scene of a video (image). For example, when a video is made up of multiple configurations, chapters are attached to the video content at predetermined points on the time axis, such as the start of the table of contents, the start of the first step, the end of the first step (same as the start of the second step), and the end of the second step. The commentary video may include a procedure for simmering for five minutes, for example. If a user tries to work according to the procedure of this commentary video, it is assumed that the user will give an instruction to the information processing device 100 such as "five-minute timer" at the start of simmering, so in the example shown in FIG. 3, the user gives an instruction to "five-minute timer".

[0026] After five minutes have passed, the information processing device 100 emits a beep to notify the user that five minutes have passed, and accepts instructions again without the need for a wake word for a predetermined period of time (e.g., one minute) after the timer has finished executing. In this example, the user issues an instruction without a wake word to "play the next cooking step," and the information processing device 100 accepts the instruction and starts playing a video corresponding to the next step in the cooking instruction video.

[0027] In this example, since it is considered likely that the user will again issue some instruction to the information processing device 100 after playing the video corresponding to each step of the instruction video or after the timer has finished executing (time is up) in accordance with the cooking steps, the information processing device 100 accepts the next instruction without the need for a wake word during a specified period after the command given by voice (voice command for timer, video playback, etc.) has finished executing.

[0028] As shown in FIG. 2 and FIG. 3, when a user instructs the information processing device 100 with a predetermined voice command, there is a high possibility that the user will instruct the information processing device 100 again with a voice command after the completion of the execution of the application program executed by the voice command. Therefore, when the information processing device 100 receives a voice command from the user, the information processing device 100 requires a wake word at first, but receives the voice command without the wake word in a predetermined period after the execution of the predetermined voice command after the wake word. Here, the predetermined voice command is, for example, a timer, video playback, video playback pause, video playback end, etc. In other words, the information processing device 100 acquires a wake word and a voice command the first time, and executes a process (timer, video playback, etc.) corresponding to the first voice command. In addition, the information processing device 100 receives subsequent voice commands without the wake word in a predetermined period after the execution of the process corresponding to the first voice command is completed (stopped).

[0029] Also, the predetermined period during which a voice command is accepted without a wake word may be a fixed length (for example, 1 minute), but may be changed according to the contents of the voice command (the contents and type of the application program to be started by the voice command). For example, if the thing to be started by the voice command is a timer of 1 minute or less, it is considered that the user is likely to wait for the timer to time out and immediately issue the next voice command, so the predetermined period may be set to a short period (for example, 30 seconds). Conversely, if the thing to be started by the voice command is a timer of 5 minutes or more, it is considered that the user may be engaged in another task and may not notice the timer time out, so the predetermined period may be set to a long period (for example, 3 minutes). Also, the predetermined period may be a time proportional to the time until the execution of the application program to be started by the voice command is completed. For example, if the predetermined period is set to half the time until the execution is completed, the predetermined period after the 5-minute timer is 2.5 minutes, and the predetermined period after playing a 2-minute video is 1 minute.

[0030] The predetermined period may be set depending on the type of application executed by the voice command. For example, in an instructional video teaching how to cook, when the playback is paused at a certain step (a chapter assigned at the end of the first step), the user may not have completed the work according to the instruction content, so the predetermined period may be set to 3 minutes, which is longer than the default period (for example, 1 minute). Furthermore, the information processing device 100 may set the length of the predetermined period depending on the type of such instructional video (how to cook, how to draw, training methods and technique introductions for sports such as soccer). In this case, the information processing device 100 may set the length of the predetermined period depending on the type of instructional video by acquiring the title and tag information (hashtags, etc.) set in the instructional video, for example, 3 minutes if the instructional video is about how to cook, and 2 minutes if the instructional video is about soccer training methods.

[0031] Furthermore, the predetermined period may be changed depending on the type of application program estimated to be started next. When performing such processing, the control unit 110 stores a history of application programs started by voice commands in the storage unit 120. Then, based on this history, the control unit 110 can estimate that the application program that has been started most frequently among the application programs started after the application program currently being started by a voice command and is being executed is the application program to be started next.

[0032] In addition, the date and time (timestamp) when the application program was launched may also be recorded in this history, and a specified period may be determined based on the difference in the launch dates and times of each application program (for example, for each application, the average time from the completion of execution of the previous application to its launch by a voice command may be calculated, and the specified period may be set as twice the average time).

[0033] After a predetermined period of time has elapsed, a wake word is required when issuing a voice command to the information processing device 100. For this reason, the information processing device 100 may output to the user how much of the predetermined period has elapsed and the amount of time that has elapsed (for example, by displaying the remaining time on a display, informing the user of the remaining time by voice, or by vibrating the device).

[0034] For example, as shown in Fig. 4, the information processing device 100 may output the amount of elapsed time by changing the color of the icon on the display according to the amount of time that has elapsed in a predetermined period, such as displaying a blue icon 211 if not much time has elapsed (e.g., more than 2 / 3 of the time remaining), a yellow icon 212 if about half the time has elapsed (e.g., more than 1 / 3 but less than 2 / 3 of the time remaining), and a red icon 213 if a considerable amount of time has elapsed (e.g., less than 1 / 3 of the time remaining). Also, as shown in Fig. 4, the information processing device 100 may output the amount of elapsed time by displaying a time bar 221 on the display that shortens in length according to the amount of time that has elapsed in a predetermined period.

[0035] Such a process (voice command recognition process) that enables a voice command to be accepted without the need for a wake word will be described with reference to Fig. 5. This process is started when the information processing device 100 is started and is ready to accept a voice command, and is executed in parallel with other processes.

[0036] First, the control unit 110 acquires a voice signal from the microphone of the input unit 130, analyzes it (voice recognition), and derives first recognition data (step S101). Then, the control unit 110 determines whether the first recognition data includes a wake word (step S102). If the wake word is not included (step S102; No), the process returns to step S101.

[0037] If the wake word is included (step S102; Yes), the control unit 110 acquires a voice signal uttered by the user after the wake word from the microphone of the input unit 130, analyzes it (voice recognition), and derives second recognition data (step S103). Then, the control unit 110 determines whether the second recognition data includes a voice command (control information that is information related to the application program (control process)) (step S104). If the voice command is not included (step S104; No), the process returns to step S101.

[0038] If a voice command is included (step S104; Yes), the control unit 110 executes an application program corresponding to the voice command (initially the first control process, but if it returns from step S109, the second control process) in parallel with the voice command recognition process by multi-thread processing, and waits until the execution is completed (step S105). Note that the completion of execution means that the timer has timed out, in the case of video playback, and that playback has progressed up to the specified point (for example, the boundary between the next process (next video or image)). In other words, the completion of execution means that the instructions given by the voice command have been completed.

[0039] Then, the control unit 110 sets the timer to a first period until the timer runs out (step S106). The first period is the above-mentioned predetermined period, for example, one minute. Next, the control unit 110 outputs the remaining time on the timer through the output unit 140 (step S107). In this step, for example, a display such as icons 211, 212, 213 or a time bar 221 shown in FIG. 4 may be performed.

[0040] Then, the control unit 110 acquires a voice signal from the microphone of the input unit 130, analyzes it (voice recognition), and derives third recognition data (step S108). Then, the control unit 110 determines whether or not the third recognition data includes a voice command (step S109). If a voice command is included (step S109; Yes), the process returns to step S105. As described above, in step S105, the control unit 110 executes an application program (second control process) corresponding to the voice command. Therefore, if the control unit 110 determines that the third recognition data includes a voice command (control information), the control unit 110 executes the second control process regardless of the presence or absence of a wake word in the third recognition data.

[0041] If no voice command is included (step S109; No), the control unit 110 judges whether the time measured by the timer has passed the first period (step S110). If the first period has not passed (step S110; No), the process returns to step S107. If the first period has passed (step S110; Yes), the process returns to step S101.

[0042] In the above process, for all voice commands, the voice command is accepted without the wake word during the first period after the execution of the application program corresponding to the voice command (activated by the voice command) is completed. However, during the first period after the execution of the application program is completed, only a predetermined voice command may be accepted without the wake word. If this is desired, the control unit 110 may determine in step S109 whether the voice command indicated by the third recognition data is the predetermined voice command.

[0043] In the above process, the wake word can be omitted for a predetermined period after the execution of the application program executed by the voice command is completed. However, the condition for making the wake word omissible is not limited to a predetermined period. For example, the wake word can be omitted if the attitude (movement, position) of the information processing device 100 has not changed for a certain period (which may be different from the above-mentioned predetermined period). This is because, for example, when the user installs the information processing device 100 at an angle that is easy to see in the kitchen, if the attitude of the information processing device 100 remains the same, it is considered that the user is continuing cooking. Furthermore, by detecting the movement of the user's arms, etc., it can be determined whether the user is continuing a related task, so that if the user is continuing a related task, the wake word can be omitted, and if the user has completed the related task and is likely to be doing something else, the wake word cannot be omitted.

[0044] In the above process, in step S105, the control unit 110 waits until the execution of the application program is completed, but during this waiting (during the execution of the application program), the control unit 110 may perform the same process as in step S108 (acquire a voice signal from the microphone of the input unit 130, analyze (voice recognition), and acquire the third recognition data). In this case, the control unit 110 may perform a process (misrecognition prevention process) to prevent the voice output from the application program being executed (for example, the voice output during video playback) from being erroneously recognized as a voice command. Possible methods for the misrecognition prevention process include a method of adding voice data in the opposite phase to the voice data output from the application program to the voice signal from the microphone (thereby canceling the voice output from the application program), a method of registering the voice of a user (not limited to one person) in advance, and not accepting any voice other than the registered voice as a voice command, and the like.

[0045] As described above, in the voice command recognition process of this embodiment, the information processing device 100 analyzes the voice signal acquired by the microphone of the input unit 130, and if the analyzed voice signal includes a wake word and a voice command, executes an application program according to the voice signal. Then, the information processing device 100 becomes able to accept a voice command without a wake word for a predetermined period after detecting the end of the executed application program (for example, when a timer expires or playback stops due to chapter detection in video playback). Therefore, the user can omit speaking the wake word when issuing an instruction to the information processing device 100.

[0046] (Embodiment 2) In the first embodiment, after the operation of the application program executed by the voice command is completed, the user can omit speaking the wake word. Here, a second embodiment will be described in which the user can omit speaking content by using data (voice signal, text data, etc.) output by the application program.

[0047] For example, in the example shown in FIG. 6, the user first utters the wake word and then utters "Play the first step of cooking," causing the information processing device 101 according to the second embodiment to start playing a video corresponding to the first step of a cooking instruction video. Then, the control unit 110 of the information processing device 101 recognizes the voice output from the video. In this example, the video includes a step of simmering for 10 minutes over medium heat, and a voice saying "Simmer for 10 minutes over medium heat" is present. Then, the control unit 110 extracts "10 minutes," which is a parameter indicating "time," from the voice data acquired by voice recognition from the voice signal output during video playback, and stores the parameter in the storage unit 120.

[0048] Then, when the control unit 110 detects the chapter assigned to the end of the video corresponding to the first step, it pauses the playback of the video. Then, the control unit 110 accepts the next instruction without the wake word during a predetermined period (predetermined period 1) after the pause. Suppose the user is trying to follow the steps of this instruction video and issues a voice instruction saying "timer" to the information processing device 101 within the predetermined period 1 (for example, when simmering begins). Then, the control unit 110 applies the parameter "10 minutes" voice-recognized from the video to a timer application program started by the voice command, and a 10-minute timer is set.

[0049] Then, after 10 minutes have passed, control unit 110 emits a beep to notify the user that 10 minutes have passed, and during a predetermined period of time after the timer has finished executing (predetermined period 2), the next instruction is accepted without the need for a wake word. In this example, the user issues an instruction to "play the next cooking step" without a wake word, and control unit 110 accepts the instruction and starts playing a video corresponding to the next step (the second step) in the cooking instruction video.

[0050] In this instructional video, there is a procedure for cutting carrots using a method called "Twisted plum," and a voice is heard saying, "Carrots cut using a twisted plum..." Then, the control unit 110 extracts "Twisted plum," which is a parameter indicating the "name of the method of cutting vegetables," from the voice data acquired by voice recognition from the voice signal output during playback of the video, and stores the parameter in the storage unit 120.

[0051] Then, when the control unit 110 detects the chapter assigned to the end of the video corresponding to the second step, it pauses the playback of the video. Then, the control unit 110 accepts the next instruction without the wake word during a predetermined period after the pause (predetermined period 3). Suppose the user is trying to follow the steps of this instruction video and wants to know how to cut a twisted plum. Then, when the user issues an instruction of "how to cut" to the information processing device 101 within the predetermined period 3, the control unit 110 applies the parameter "twisted plum" voice-recognized from the video to a video search application program started in response to the voice command, and a video of "how to cut a twisted plum" is searched for. Then, the control unit 110 accepts the next instruction without the wake word during a predetermined period after the search (predetermined period 4).

[0052] In this way, in the information processing device 101 of the second embodiment, not only can the wake word be omitted, but also parameters (control parameters) to be applied to an application program (control process) in response to a voice command can be automatically acquired.

[0053] The functional configuration of the information processing device 101 according to the second embodiment is the same as that of the information processing device 100 according to the first embodiment, as shown in FIG. 1. However, the storage unit 120 of the information processing device 101 is provided with an extracted parameter table 121 and a parameter buffer which is a buffer (storage area) for temporarily storing parameters to be applied to an application program (control process). The extracted parameter table 121 stores parameters extracted as parameters from data (voice signal, text data, etc.) output by an application program started by a voice command. The parameter buffer stores parameters (time, etc.) extracted from the data (voice data, text data, etc.) in a voice command recognition process to be described later.

[0054] As shown in FIG. 7, the extracted parameter table 121 defines "extracted parameters" (parameters extracted from data (voice data, text data, etc.) output by an application program started by a voice command), "user voice" (voice command uttered by the user after execution of an application program started by a voice command), and "started application" (application program started by applying the "extracted parameters" when the "user voice" is uttered).

[0055] For example, in FIG. 7, "time" of the "extracted parameter", "timer" of the "user voice", and "timer at the time" of the "launched application" are defined in association with each other. Here, when the user utters "timer", the control unit 110 judges whether or not "time" (10 minutes in FIG. 6), which is one type of parameter, is stored in the parameter buffer. Then, when the control unit 110 judges that a parameter corresponding to "time" is stored in the parameter buffer, it reads out the parameter, and based on the extracted parameter table 121, starts a timer as an application, sets the parameter (10 minutes), and starts the timer.

[0056] In addition, in the next line in FIG. 7, the "name of the method of cutting vegetables" in the "extracted parameter", the "cutting method" in the "user voice", and the "video search of the method of cutting the vegetables" in the "started application" are associated with each other. Here, when the user utters "cutting method", the control unit 110 judges whether or not the "name of the method of cutting vegetables" (twisted plum in FIG. 6), which is a type of parameter, is stored in the parameter buffer. Then, when the parameter (twisted plum, for example) corresponding to the "name of the method of cutting vegetables" is stored in the parameter buffer as a result of the judgment, the control unit 110 reads the parameter, and starts a video search as an application based on the extracted parameter table 121, sets the parameter (twisted plum) as a search keyword, and starts a video search. In this example, the "extracted parameter" is defined as the "name of the method of cutting vegetables", but it is not necessary to be limited to such a definition. For example, since there are only a limited number of basic ways to cut vegetables (thin slices, round slices, half-moon slices, etc.) and decorative cuts (twisted plums, etc.), the names of specific cutting methods can be individually defined as “extraction parameters” to construct the extraction parameter table 121.

[0057] The same is true for the other examples shown in FIG. 7, but these are merely examples of the extracted parameter table 121, and the extracted parameter table 121 may be expanded or modified as desired.

[0058] As described above, when the data (audio signal, text data, etc.) output from the application executed by the voice command includes a parameter related to an item defined as an extracted parameter in the extracted parameter table 121, the control unit 110 stores the parameter in the parameter buffer. Then, the control unit 110 judges whether or not a parameter corresponding to the voice command (information on the application program) uttered by the user is stored in the parameter buffer. Then, when the parameter corresponding to the voice command is stored in the parameter buffer, the control unit 110 reads the parameter, and applies the parameter to the application program corresponding to the voice command based on the extracted parameter table 121 to start the application program (start a timer at a set time, search for a video with a specific keyword, etc.). As a result, the information processing device 101 does not require a wake word for the user's speech to the information processing device 101 during a predetermined period after the application program is executed by the voice command, and makes it possible to omit the contents of parameters (time, name, etc.) that should originally be included in the voice command.

[0059] The voice command recognition process according to the second embodiment will be described with reference to Fig. 8. This process is started when the information processing device 101 is booted and becomes ready to accept a voice command, and is executed in parallel with other processes.

[0060] First, the processes from step S201 to step S204 are similar to the processes from step S101 to step S104 in the voice command recognition process (FIG. 5) according to the first embodiment, and therefore the description thereof will be omitted.

[0061] In step S205, the control unit 110 starts an application program corresponding to the voice command and executes the application program in parallel with the voice command recognition process by multi-thread processing. Then, the control unit 110 analyzes (recognizes) data output by the execution of the application program (output data such as voice signals and text data) as first output information (step S206).

[0062] Then, the control unit 110 judges whether or not there is a correlation between the words included in the first output information and the contents of the items defined as extraction parameters in the extraction parameter table 121 (step S207). The judgment of whether or not there is a correlation is made for all of the defined extraction parameters. That is, in the extraction parameter table 121 of FIG. 7, whether or not there is a correlation is judged for 11 types of items such as time and names of ways to cut vegetables. If there is no correlation as a result of the judgment (step S207; No), the process proceeds to step S209.

[0063] Here, being related means that the words included in the first output information are related to the items defined as extraction parameters in the extraction parameter table 121. In other words, not only when the words included in the first output information completely match the items defined as extraction parameters, but also when it is determined that the words match with a certain degree of range (latitude), such as synonyms and dialects, are considered to be related here. As an example of a match with a certain degree of range, the vegetable daikon and the Okinawa dialect for daikon, daekney, are considered to match. Similarly, other names that change between the present and the past, such as "ruler" and "ruler," names with relatively high recognition of nicknames, such as "product or service names" and "nicknames of product or service names," and names with relatively high recognition of abbreviated names, such as "product or service names" and "names with abbreviated product or service names," are also considered to match.

[0064] If there is an association between the word included in the first output information and the item defined as the extraction parameter (step S207; Yes), the control unit 110 stores the word included in the first output information as the extraction parameter in the parameter buffer of the memory unit 120 (step S208), and proceeds to step S209.

[0065] In step S209, control unit 110 determines whether or not the execution of the application program that started in step S205 has been completed. If the execution has not been completed (step S209; No), the process returns to step S206.

[0066] When the execution of the application program is completed (step S209; Yes), the process proceeds to step S210. The process from step S210 to step S212 is similar to the process from step S106 to step S108 in the voice command recognition process (FIG. 5) according to the first embodiment, and therefore the description thereof will be omitted.

[0067] In step S213, the control unit 110 determines whether or not the parameters corresponding to the third recognition data acquired in step S212 are stored in the parameter buffer. If the control unit 110 determines that the parameters corresponding to the third recognition data are not stored in the parameter buffer (step S213; No), the control unit 110 cannot execute the application defined in the extracted parameter table 121 using the third recognition data and the parameters stored in the parameter buffer, and proceeds to step S215.

[0068] If the extracted parameters and the third recognition data exist in the extracted parameter table 121 (step S213; Yes), the control unit 110 determines that the extracted parameters (control parameters) are applicable to the application program (second control process) defined as the "start application" in the extracted parameter table 121, and applies the extracted parameters to execute the application program in parallel with the voice command recognition process by multi-thread processing (step S214). Then, the process returns to step S206.

[0069] The processes in steps S215 and S216 are similar to those in steps S109 and S110 in the voice command recognition process (FIG. 5) according to the first embodiment, and therefore will not be described.

[0070] By the above voice command recognition process, if there is a parameter related to an item defined as an extracted parameter in the extracted parameter table in the data (voice signal, text data, etc.) output from the application program executed by the voice command, the control unit 110 stores the parameter in the parameter buffer. Then, the control unit 110 judges whether or not a parameter corresponding to the voice command (information on the application program) uttered by the user is stored in the parameter buffer. Then, if a parameter corresponding to the voice command is stored in the parameter buffer, the control unit 110 reads out the parameter, and applies the parameter to the application program corresponding to the voice command based on the extracted parameter table 121 to execute the application program. As a result, during a predetermined period after the execution of the application program by the voice command, the information processing device 101 can start an appropriate application program by a voice command that does not require a wake word and omits the specification of a parameter in response to the user's utterance to the information processing device 101.

[0071] In the above-mentioned second embodiment, the parameters (extracted parameters) stored in the parameter buffer are extracted by analyzing data output from an application program, but the present invention is not limited to this. For example, when an application program for playing a video is executed by a voice command, text data (hashtags, etc.) attached to the video, text data obtained by character recognition from an image, etc. may be used as the extracted parameters instead of or in addition to the extracted parameters obtained by analyzing the audio output from the video.

[0072] In the above-mentioned second embodiment, the control unit 110 executes an application program after recognizing that the user has uttered a voice command. However, the control unit 110 may estimate the next application program to be started according to the parameters stored in the parameter buffer, and start the estimated application program in advance in the background. This allows the application program to respond instantly after the user has uttered a voice command. In this case, if no voice command is uttered even after a predetermined period of time has passed, the control unit 110 automatically terminates the application program started in the background.

[0073] (Embodiment 3) In the second embodiment, parameters extracted from data output from an application program are also used to reduce the effort required for the user to speak, but we will now explain the third embodiment, which can reduce the amount of speech required from the user based on the user's behavior.

[0074] For example, in the example shown in FIG. 9, the user first utters the wake word and then utters "read out the email", which causes the information processing device 102 according to the third embodiment to read out the contents of the received email. Then, the control unit 110 of the information processing device 102 analyzes the text data of the email. In this example, it is assumed that the address of a meeting place is described in the email. Then, the control unit 110 extracts "△△1-2-3, XX-ku, Tokyo", which is a parameter of "address, place name, facility name, etc." from the email, and stores it in the storage unit 120. Note that, specifically, the control unit 110 stores the extracted parameters in the parameter buffer based on the data output from the application program and the extracted parameter table 121, as in the second embodiment.

[0075] Then, the control unit 110 completes reading out the email, and within a predetermined period thereafter (predetermined period 1), it accepts the next instruction without the wake word. Suppose the user wants to go to the address provided in the email, and issues a voice command saying "map" to start an application program for map display and navigation. The control unit 110 then applies the address of the meeting place, "1-2-3, △△, XX-ku, Tokyo," which is a parameter extracted from the email, to the map display application program that is started in response to the voice command, and a map of the area around this address is displayed.

[0076] Then, for a predetermined period of time (predetermined period 2) after the map is displayed, control unit 110 monitors the user's behavior. In this example, the user starts moving to the meeting place. Then, for a predetermined period of time (predetermined period 3) after the user's movement is detected, control unit 110 accepts the next instruction without the need for a wake word. In this example, the user issues the instruction "Balance" without the wake word, and control unit 110 accepts the instruction, launches the application program of the transportation IC card, and outputs "It's 2500 yen."

[0077] After this output, control unit 110 monitors the user's actions for a predetermined period of time (predetermined period 4). In this example, the user uses the IC card to exit the ticket gate. Then, control unit 110 accepts the next instruction without the wake word again for a predetermined period of time (predetermined period 5) after the user uses the IC card. In this example, the user issues the instruction "send email" without the wake word, and control unit 110 accepts this instruction and uses an email application program to send an email notifying the user that they have left the station.

[0078] In this way, in the third embodiment, not only can the wake word be omitted, but also the user's behavior can be monitored and an application program inferred from the behavior can be launched.

[0079] 1, the functional configuration of the information processing device 102 according to the third embodiment is similar to that of the information processing device 101 according to the second embodiment. However, in the storage unit 120 of the information processing device 102, in addition to the storage areas of the extracted parameter table 121 and parameter buffer provided in the information processing device 101, storage areas of the behavior table 122 and behavior buffer are also prepared. The behavior table 122 stores user behaviors related to applications that have completed execution, etc. Also, the behavior buffer is a buffer that stores detected user behaviors.

[0080] As shown in FIG. 10, the behavior table 122 defines "executed application" (an application program that has been launched by a voice command and has completed execution), "user action" (a user action that is presumed to be performed after the execution of the "executed application" has been completed), "user voice" (a voice command that is presumed to be uttered by the user after the "user action"), and "launched application" (an application program that is launched based on the "user action" or "user voice" when the "user voice" is uttered).

[0081] For example, in FIG. 10, the "completed application" is defined as "map," the "user action" is defined as "move or start navigation," the "user voice" is defined as "balance," and the "launch application" is defined as "(transportation IC card) balance output." Normally, the voice command "balance" would be considered to output the balance by some IC card application, but IC card applications include shopping applications handled by chains such as convenience stores, and transportation applications handled by chains of transportation facilities. In this example, since the "completed application" is defined as "map" and the "user action" is defined as "move or start navigation," the target IC card is presumed to be a transportation IC card, and the control unit 110 will output the balance of the transportation IC card.

[0082] 10, the next line defines "executed application" as "output balance," "user action" as "use IC card," "user voice" as "send email," and "launched application" as "send email (that the user has left the station)." Normally, there are various possibilities for what kind of email should be sent in response to the voice command "send email," but in this example, since the "executed application" is "output balance" and the "user action" is "use IC card," it is presumed that the user has used the IC card to exit the station ticket gate, and control unit 110 will send an email notifying the user that they have left the station.

[0083] Also, information not defined in the action table 122 (the destination of the e-mail in this example) may be set based on the history of application programs that have been started up to that point. For example, in the example shown in Fig. 9 above, since the process starts with receiving an e-mail in the information processing device 102, the control unit 110 may set the destination of the last e-mail to be sent as a reply to the first e-mail received (or a reply to everyone including CC (Carbon Copy)).

[0084] In this way, by using the action table 122, the control unit 110 can determine the next application program to be started based on information on what application program has been executed and what the user's action or voice was in response to it, thereby further reducing the user's effort. Note that the action table 122 shown in Fig. 10 is merely an example, and may be expanded or modified as desired.

[0085] In the third embodiment, not only can the user omit speaking the wake word, but also, by using the behavior table 122, the application program can be started with the contents taking into account the user's behavior.

[0086] The voice command recognition process according to the third embodiment will be described with reference to Fig. 11 and Fig. 12. This process is started when the information processing device 102 is booted and ready to accept a voice command, and is executed in parallel with other processes.

[0087] First, among the processes from step S301 to step S316 (processing shown in FIG. 11), except for step S315, the processes are similar to the processes from step S201 to step S216 (excluding step S215) in the voice command recognition process (FIG. 8) according to embodiment 2, and therefore their explanation will be omitted.

[0088] In step S315, the control unit 110 determines whether or not the third recognition data includes a voice command. If the voice command is not included (step S315; No), the process proceeds to step S316, as in the second embodiment. If the voice command is included (step S315; Yes), the process proceeds to Fig. 12, where the control unit 110 executes an application program (second control process) that is started in response to the voice command in parallel with the voice command recognition process by multithread processing (step S318).

[0089] Then, the control unit 110 determines whether the execution of the application program started in step S318 or step S331 is completed (step S319). If the execution is not completed (step S319; No), the control unit 110 returns to step S319 and waits until the execution is completed.

[0090] When the execution of the application program is completed (step S319; Yes), the control unit 110 sets a timer whose time is up to the second period (step S320). The second period is a predetermined period for monitoring the user's behavior as described above, for example, 10 minutes. Then, the control unit 110 outputs the remaining time of the timer on the output unit 140 (step S321). In this step, for example, a display such as icons 211, 212, 213 or a time bar 221 shown in FIG. 4 may be performed. Also, the output method (output mode (display, audio output, vibration, etc.), font when displayed, color and size of the icon, time bar, etc.) may be different from the output method in steps S311 and S327 so as to be distinguished from the timer of the first period (a predetermined period in which the wake word can be omitted).

[0091] Next, the control unit 110 refers to the action table 122 to monitor a user action corresponding to the application program whose execution has been completed in step S319 (step S322), and determines whether or not the user action has been detected (step S323).

[0092] If the user action is not detected (step S323; No), the control unit 110 judges whether the time measured by the timer has passed the second period (step S324). If the second period has not passed (step S324; No), the control unit 110 returns to step S321. If the second period has passed (step S324; Yes), the control unit 110 returns to step S301.

[0093] On the other hand, if the user action is detected (step S323; Yes), the control unit 110 stores the detected action in the action buffer of the storage unit 120 (step S325).

[0094] Then, control unit 110 sets a timer whose time is up to a first period (step S326). As described above, the first period is a predetermined period during which the user can omit the wake word, for example, 10 minutes. Next, control unit 110 outputs the remaining time on the timer via output unit 140 (step S327). In this step, for example, a display such as icons 211, 212, 213 or a time bar 221 shown in FIG. 4 may be performed.

[0095] Then, the control unit 110 acquires a voice signal from the microphone of the input unit 130, analyzes it (voice recognition), and acquires third recognition data (step S328). Next, the control unit 110 acquires the user action stored in the action buffer (step S329). Then, the control unit 110 judges whether or not the third recognition data includes a voice command, and whether or not the user action acquired in step S329 and the voice command included in the third recognition data exist in the action table 122 as "user action" and "user voice", respectively (step S330). Note that, in this judgment, similarly to step S207 of the voice command recognition process (FIG. 8) of the above-mentioned second embodiment, if the user action in the action buffer and the voice command in the third recognition data are related to the "user action" and "user voice" in the action table 122, respectively, the user action and the voice command may be judged to exist in the action table 122, respectively.

[0096] If the user action and voice command are present in the action table 122 (step S330; Yes), the control unit 110 executes an application program (second control process) defined as a "launch application" corresponding to the "user action" and "user voice" in accordance with the action table 122 in parallel with the voice command recognition process by multi-thread processing (step S331), and proceeds to step S319.

[0097] In addition, the "start application" in the behavior table 122 includes not only the application program to be started, but also information on what parameters to apply at the time of start based on the corresponding "user behavior" and "user voice." Therefore, in step S331, the control unit 110 can apply appropriate parameters based on the information on the "start application" defined in the behavior table 122 to execute the application program.

[0098] If the third recognition data does not include a voice command, or if the user action acquired in step S329 and the voice command included in the third recognition data do not exist in the action table 122 (step S330; No), the control unit 110 determines whether the third recognition data includes a voice command (step S332). If the voice command is included (step S332; Yes), the process returns to step S318.

[0099] If a voice command is not included (step S332; No), the control unit 110 judges whether the time measured by the timer has passed the first period (step S333). If the first period has not passed (step S333; No), the process returns to step S327. If the first period has passed (step S333; Yes), the process returns to step S301.

[0100] With the above voice command recognition process, when a user performs an action in accordance with the "user action" defined in the action table 122, not only can the user omit the wake word, but also an appropriate application program matching the user's action can be launched.

[0101] (Other variations) The information processing device 100 is not limited to a smartphone, and may be realized by a computer such as a smart watch equipped with the sensor unit 160, a portable tablet, or a PC (Personal Computer). Specifically, in the above embodiment, the program such as the voice command recognition process executed by the control unit 110 is described as being stored in advance in the storage unit 120. However, the program may be stored in a non-transitory computer-readable recording medium such as a flexible disk, a CD-ROM (Compact Disc Read Only Memory), a DVD (Digital Versatile Disc), an MO (Magneto-Optical disc), a memory card, or a USB memory and distributed, and the program may be read and installed in a computer to configure a computer capable of executing each of the above-mentioned processes.

[0102] Furthermore, the program may be superimposed on a carrier wave and applied via a communication medium such as the Internet. For example, the program may be distributed by posting it on a bulletin board system (BBS) on a communication network. The program may then be started and executed under the control of an OS (Operating System) in the same manner as other application programs, thereby enabling the above-mentioned processes to be performed.

[0103] In addition, the control unit 110 may be configured by any processor alone, such as a single processor, a multiprocessor, or a multi-core processor, or may be configured by combining any of these processors with a processing circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).

[0104] Although the preferred embodiment of the present invention has been described above, the present invention is not limited to the specific embodiment, and the present invention includes the inventions described in the claims and their equivalents. The inventions described in the original claims of this application are listed below.

[0105] (Appendix 1) A voice acquisition unit that acquires a voice signal; A control unit, The control unit is If it is determined that the first recognition data derived from the voice signal includes a wake word, if it is determined that second recognition data derived from the voice signal after the wake word includes control information, the control information being information related to a control process, execute a first control process according to the control information; When it is determined that the control information is included in third recognition data derived from the voice signal acquired by the voice acquisition unit during a first period after a predetermined condition is satisfied or during the execution of the first control process, a second control process is executed according to the control information included in the third recognition data. 23. An information processing apparatus comprising:

[0106] (Appendix 2) When the control unit determines that the control information is included in the third recognition data derived from the voice signal, the control unit executes the second control process even if the third recognition data does not include the wake word. 2. The information processing device according to claim 1 .

[0107] (Appendix 3) The predetermined condition is satisfied when the execution of the first control process is completed. 2. The information processing device according to claim 1 .

[0108] (Appendix 4) The control unit is When it is determined that the first output information outputted by executing the first control process includes a control parameter related to the control process, applying the control parameter when the second control process is executed; 3. The information processing device according to claim 2.

[0109] (Appendix 5) The control unit is determining whether there is a correlation between the information included in the first output information and the information included in the parameter table; When it is determined that there is the association, it is determined that the first output information includes the control parameter. 5. The information processing device according to claim 4.

[0110] (Appendix 6) The control unit is Predicting a subsequent behavior of the user based on the second control process; the predetermined condition is satisfied when the execution of the first control process is completed or the estimated behavior is detected; The control unit further includes: When it is detected that the user has performed the estimated behavior during a second period after the execution of the second control process is completed, if it is determined that the control information is included in the third recognition data derived from the voice signal acquired by the voice acquisition unit during the first period after the detection, a new second control process is executed according to the control information included in the third recognition data. 2. The information processing device according to claim 1 .

[0111] (Appendix 7) The control unit is Predicting the user's behavior based on the second control process and a behavior table; when it is detected that the user has performed the estimated behavior during a second period after the execution of the second control process is completed, determining whether or not there is a correlation between control information included in the third recognition data derived from a voice signal acquired by the voice acquisition unit during the first period after the detection and information included in the behavior table; When it is determined that there is the association, a new second control process is executed in accordance with control information included in the third recognition data. 7. The information processing device according to claim 6,

[0112] (Appendix 8) The control unit is outputting the amount of time that has elapsed since the start of the first period until the end of the first period; 8. An information processing device according to any one of claims 1 to 7.

[0113] (Appendix 9) A control unit of an information processing device including a voice acquisition unit that acquires a voice signal and a control unit, If it is determined that the first recognition data derived from the voice signal includes a wake word, if it is determined that second recognition data derived from the voice signal after the wake word includes control information, the control information being information related to a control process, execute a first control process according to the control information; When it is determined that the control information is included in third recognition data derived from a voice signal acquired by the voice acquisition unit during a first period after a predetermined condition is satisfied or during execution of the first control process, execute a second control process according to control information included in the third recognition data; 23. An information processing method comprising:

[0114] (Appendix 10) A control unit of an information processing device including a voice acquisition unit that acquires a voice signal and a control unit, If it is determined that the first recognition data derived from the voice signal includes a wake word, if it is determined that second recognition data derived from the voice signal after the wake word includes control information, the control information being information related to a control process, execute a first control process according to the control information; When it is determined that the control information is included in third recognition data derived from a voice signal acquired by the voice acquisition unit during a first period after a predetermined condition is satisfied or during execution of the first control process, execute a second control process according to control information included in the third recognition data; A program characterized by executing a process. [Explanation of symbols]

[0115] 100: information processing device, 110: control unit, 120: storage unit, 121: extracted parameter table, 122: behavior table, 130: input unit, 140: output unit, 150: communication unit, 160: sensor unit, 211, 212, 213: icons, 221: time bar

Claims

An information processing apparatus including a control unit, which, within a predetermined period that is either during the execution of control related to the content of processing based on a wake word and control information included in recognition data derived from an acquired audio signal or within a predetermined period after the completion of the control, if a user's action estimated based on the control is detected and second control information is included in the recognition data of the audio signal acquired within the predetermined period, starts new control using the second control information as a parameter without recognizing the wake word.

2. The control unit when determining that control parameters related to control processing are included in the output information output by executing the processing, applies the control parameters during the execution of the control. The information processing apparatus according to claim 1, characterized by the above.

3. The control unit determines whether there is a correlation between the information included in the output information and the information included in a parameter table, and when determining that there is a correlation, determines that the control parameters are included in the output information. The information processing apparatus according to claim 2, characterized by the above.

4. When a control unit of an information processing apparatus within a predetermined period that is either during the execution of control related to the content of processing based on a wake word and control information included in recognition data derived from an acquired audio signal or within a predetermined period after the completion of the control, if a user's action estimated based on the control is detected and second control information is included in the recognition data of the audio signal acquired within the predetermined period, starts new control using the second control information as a parameter without recognizing the wake word. An information processing method characterized by the above.

5. A program for causing a control unit of an information processing apparatus within a predetermined period that is either during the execution of control related to the content of processing based on a wake word and control information included in recognition data derived from an acquired audio signal or within a predetermined period after the completion of the control, if a user's action estimated based on the control is detected and second control information is included in the recognition data of the audio signal acquired within the predetermined period, to start new control using the second control information as a parameter without recognizing the wake word.