Information processing device, information processing method, and program
The voice recognition system addresses fast response and erroneous recognition by using both short and long voice patterns within specific conditions, enhancing accuracy and speed in command recognition.
Patent Information
- Application Number
- PCT/JP2025/025721
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-03
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Existing voice recognition systems face issues with fast response times and erroneous recognition when only the first syllable of a word is recognized as a voice command.
A voice recognition system that identifies commands based on a comparison between voice patterns detected within a predetermined condition, using both shorter and longer candidate voice patterns to enhance accuracy and speed.
Reduces erroneous recognition while providing a fast response by recognizing voice commands efficiently, allowing for clear user intent indication.
Smart Images

Figure JP2025025721_29012026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present invention relates to an information processing device, an information processing method, and a program.
[0002] Devices that can be operated by voice recognition are being developed, and it is desirable for such devices to have a short response time to speech utterances.
[0003] Patent Document 1 describes a technique for performing speech recognition by targeting only the first syllable of a word input as a voice command.
[0004] Japanese Patent Application Laid-Open No. 2000-112490
[0005] However, if only the first syllable of a word is unconditionally recognized, there is a possibility that the user's utterance may be recognized as a voice command contrary to the user's intention.
[0006] One example of a problem to be solved by the present invention is to provide a fast response according to the selection result while suppressing erroneous recognition in speech recognition.
[0007] The invention described in claim 1 is an information processing device comprising: a voice recognition unit that identifies a selected command from a plurality of commands based on a comparison result between a voice pattern detected within a period that satisfies a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands; the voice recognition unit identifies the selected command based on a comparison result between a voice pattern detected within a period that does not satisfy the first condition and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands; and for each of the plurality of commands, the second candidate voice pattern is longer than the first candidate voice pattern for the same command.
[0008] The invention described in claim 11 is an information processing method executed by one or more computers, including a voice recognition step of identifying a selected command from a plurality of commands based on a comparison result between a voice pattern detected within a period satisfying a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands; in the voice recognition step, the selected command is identified based on a comparison result between a voice pattern detected within a period not satisfying the first condition and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands; and for each of the plurality of commands, the second candidate voice pattern is longer than the first candidate voice pattern for the same command.
[0009] The invention described in claim 12 causes a computer to function as a voice recognition means that identifies a selected command from a plurality of commands based on a comparison result between a voice pattern detected within a period that satisfies a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands; the voice recognition means identifies the selected command based on a comparison result between a voice pattern detected within a period that does not satisfy the first condition and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands; and for each of the plurality of commands, the second candidate voice pattern is a program that is longer than the first candidate voice pattern for the same command.
[0010] FIG. 1 is a diagram illustrating a functional configuration of an information processing device according to a first embodiment. FIG. 2 is a diagram illustrating an information processing method according to the first embodiment. FIG. 3 is a diagram illustrating a configuration of a system according to the first embodiment. FIG. 4 is a diagram illustrating a computer for realizing an information processing device. FIG. 5 is a diagram illustrating an image displayed on a display unit. FIG. 6 is a diagram illustrating an image displayed on a display unit. FIG. 7 is a diagram illustrating an image displayed on a display unit. FIG. 8 is a diagram illustrating an image displayed on a display unit. FIG. 9 is a diagram illustrating a configuration of reference data held in a storage unit. FIG. 10 is a flowchart illustrating a processing flow executed by an information processing device according to a first embodiment. FIG. 11 is a flowchart illustrating processing contents of a determination unit and a voice recognition unit in step S103. FIG. 12 is a diagram illustrating a functional configuration of an information processing device according to an embodiment b. FIG. 13 is a diagram illustrating an information processing method according to a second embodiment. FIG. 14 is a flowchart illustrating a processing flow executed by an information processing device according to a second embodiment.
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0012] First Embodiment FIG. 1 is a diagram illustrating an example of the functional configuration of an information processing device 10 according to the first embodiment. The information processing device 10 according to the present embodiment includes a voice recognition unit 130. The voice recognition unit 130 identifies a selected command from among a plurality of commands based on a comparison result between a first voice pattern detected within a period that satisfies a predetermined first condition and one or more of a plurality of first candidate voice patterns. Each of the plurality of first candidate voice patterns is associated with one of the plurality of commands.
[0013] 2 is a diagram illustrating an information processing method according to this embodiment. The information processing method according to this embodiment is executed by one or more computers. The information processing method according to this embodiment includes a voice recognition step S10. In the voice recognition step S10, the one or more computers identify a selected command from the plurality of commands based on a comparison result between a first voice pattern detected within a period satisfying a predetermined first condition and one or more of a plurality of first candidate voice patterns. Each of the plurality of first candidate voice patterns is associated with one of the plurality of commands.
[0014] The information processing method according to this embodiment can be executed by the information processing device 10 according to this embodiment. The information processing device 10 and the information processing device according to this embodiment will be described in detail below.
[0015] 3 is a diagram illustrating an example of the configuration of a system 50 according to this embodiment. The system 50 according to this embodiment includes an information processing device 10, an operation unit 20, a sound collection unit 22, a display unit 30, and a sound output unit 32. In the example of FIG. 3, the information processing device 10 further includes a determination unit 110 and an operation processing unit 150.
[0016] The hardware configuration of the information processing device 10 will be described below. Each functional component of the information processing device 10 (the determination unit 110, the voice recognition unit 130, and the operation processing unit 150) may be realized by hardware (e.g., a hardwired electronic circuit) that realizes the functional component, or by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). A case where each functional component of the information processing device 10 is realized by a combination of hardware and software will be further described below.
[0017] FIG. 4 is a diagram illustrating an example of a computer 1000 for implementing the information processing device 10. The computer 1000 may be any computer. For example, the computer 1000 may be a system-on-chip (SoC), a personal computer (PC), a server machine, a tablet terminal, a smartphone, or the like. The computer 1000 may be a dedicated computer designed for implementing the information processing device 10, or a general-purpose computer. The information processing device 10 may be implemented by a single computer 1000 or by a combination of multiple computers 1000. When the information processing device 10 is implemented by a combination of multiple computers 1000, some of the multiple computers 1000 may be provided in a mobile object, and the other units may be provided as a server machine separate from the mobile object. In other words, some of the multiple functional components of the information processing device 10 (the determination unit 110, the voice recognition unit 130, and the operation processing unit 150) may be implemented by a computer 1000 provided in a mobile object, and the other units may be implemented as a server machine.
[0018] The computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path through which the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 transmit and receive data to and from each other. However, the method of interconnecting the processor 1040 and other components is not limited to bus connection. The processor 1040 may be any of various processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device implemented using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device implemented using a hard disk, a solid state drive (SSD), a memory card, a read-only memory (ROM), or the like.
[0019] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, an input device such as a microphone and an output device such as a display are connected to the input / output interface 1100. The input / output interface 1100 may be connected to the input device or output device wirelessly or by wire. In the example of Fig. 3, an operation unit 20, a sound collection unit 22, a display unit 30, and a sound output unit 32 are connected to the input / output interface 1100 of the computer 1000 that realizes the information processing device 10.
[0020] The network interface 1120 is an interface for connecting the computer 1000 to a network. Examples of this communication network include a LAN (Local Area Network) and a WAN (Wide Area Network). The network interface 1120 may connect to the network wirelessly or by wire.
[0021] When the information processing device 10 is realized by combining a plurality of computers 1000 , the plurality of information processing devices 10 are connected to each other via a network interface 1120 or an input / output interface 1100 .
[0022] The storage device 1080 stores program modules that realize the various functional components of the information processing device 10. The processor 1040 reads these program modules into the memory 1060 and executes them to realize the functions corresponding to the respective program modules.
[0023] The system 50 can function as a navigation system for a mobile object. A user of the system 50 and the information processing device 10 can be a driver of the mobile object or a passenger other than the driver of the mobile object. The mobile object is not particularly limited, but can be, for example, a vehicle. Examples of vehicles include bicycles, motorcycles, and automobiles with three or more wheels.
[0024] The operation unit 20 is, for example, a physical button. Alternatively, the operation unit 20 and the display unit 30 may be integrated and realized by a touch panel. The information processing device 10 acquires operation information indicating operations performed on the operation unit 20 from the operation unit 20. The operation unit 20 is provided in a position where it can be operated by at least the driver of the mobile body. The operation unit 20 may also be provided in a position where it can be operated by a person other than the driver of the mobile body. The system 50 may include multiple operation units 20. In this case, the system 50 may include multiple types of operation units 20, such as physical buttons and touch panels.
[0025] The operation unit 20 may be configured integrally with the information processing device 10. For example, a computer 1000 that configures at least some of the functions of the information processing device 10 may be housed in a housing in which the operation unit 20 is provided. However, the operation unit 20 may be provided separately from the information processing device 10. In this case, the operation unit 20 independent of the information processing device 10 may be attached in a position that is easy for the driver to operate.
[0026] When the operation unit 20 is a button, the button is preferably provided on the steering wheel of the vehicle. This allows the driver to easily operate the button while driving. As described above, a device containing a computer 1000 that realizes at least some of the functions of the information processing device 10 in a housing in which the operation unit 20 is provided may be attached to the steering wheel, or the operation unit 20 independent of the information processing device 10 may be attached to the steering wheel.
[0027] As another example, the operation unit 20 may be a sensor for detecting the user's actions. Examples of the sensor include an acceleration sensor and a camera. The sensor is worn on the driver's head, for example. The sensor may be provided in a headset worn by the driver. If the sensor is an acceleration sensor, the driver's head movements can be detected as gestures. If the sensor is a camera, a change in the camera orientation caused by the driver's head movements can be detected by analyzing an image captured by the camera, and the driver's head movements can be detected as gestures. Note that a camera may be attached to a moving object at a position where it can capture an image of the driver's head, and head movements captured by the camera may be detected.
[0028] The sound collection unit 22 is, for example, a microphone. The sound collection unit 22 acquires sound and converts it into sound data. The sound data indicates the waveform of the sound acquired by the sound collection unit 22. The sound data is input to the information processing device 10 from the sound collection unit 22. Note that A / D conversion of the sound signal may be performed by the sound collection unit 22 or by the information processing device 10. The sound collection unit 22 is configured to be able to acquire at least sound uttered by a driver riding in the mobile object. The sound collection unit 22 may also be configured to be able to further acquire sound uttered by a person other than the driver riding in the mobile object. The system 50 may include multiple sound collection units 22.
[0029] The display unit 30 is, for example, a display. The display unit 30 is configured to display various images under the control of the information processing device 10. The information processing device 10 causes the display unit 30 to display, for example, a map for navigation, information for operation, etc. The display unit 30 is provided in a position visible to the driver. Specifically, the display unit 30 is preferably provided near the steering wheel of the mobile object or near the meter of the mobile object so that it is easily visible to the driver.
[0030] As described above, when the operation unit 20 and the display unit 30 are integrated and realized by a touch panel, virtual buttons for operation are displayed on the display unit 30. Then, the user can perform button operations by touching an area on the touch panel that corresponds to the button.
[0031] Examples of the sound output unit 32 include a speaker, an earphone, a headphone, etc. When the sound output unit 32 is an earphone or a headphone, the sound output unit 32 is worn by the driver, and the driver can hear the sound from the sound output unit 32. The sound output unit 32 is configured to be controlled by the information processing device 10 and to output various sounds. The information processing device 10 causes the sound output unit 32 to output, for example, voice guidance for navigation.
[0032] Each functional component of the information processing device 10 will be described in detail below.
[0033] The determination unit 110 determines whether the first condition is satisfied. The determination unit 110 will be described in detail later.
[0034] The voice recognition unit 130 processes the received voice data to identify selection results for multiple commands, and the operation processing unit 150 then executes processing based on the commands identified by the voice recognition unit.
[0035] For example, multiple states are predefined. Then, multiple commands are associated with each of the multiple states. In each state, the operation processing unit 150 executes processing based on the selection results for the multiple commands associated with that state. The processing executed by the operation processing unit 150 includes state transitions and the execution of designated processing other than state transitions. For example, the multiple states are organized into a hierarchy. There are two types of commands: commands that cause a state transition (called "transition commands") and commands that cause a processing other than a state transition (called "execution commands"). A transition command can be considered to be a command at the top or middle layer of the hierarchy. An execution command can be considered to be a command at the bottom layer of the hierarchy.
[0036] When a transition command is selected, the operation processing unit 150 transitions the state according to the selection result of the command, and displays an image associated with the state after the transition on the display unit 30. For example, at least some of the commands associated with the state are presented by the image associated with the state. This allows the user to check the commands presented in the image and further select a desired command from among those commands.
[0037] When an execution command is selected, the operation processing unit 150 executes a process associated with the execution command. Examples of processes associated with execution commands include a process for displaying a specific spot on a map, a process related to a route (resetting a route, changing a destination, etc.), a process related to guidance (navigation) (changing the scale of a map, changing the volume of a guidance voice, etc.), etc. However, the processes associated with execution commands are not limited to these examples.
[0038] 5 to 8 are diagrams illustrating examples of images displayed on the display unit 30. The flow of processing executed by the information processing device 10 will be further explained using these diagrams. The following example shows the flow up to displaying additional spots on a map, but the processing of the information processing device 10 is not limited to this example.
[0039] 5, a mark 91 indicating the current location is displayed superimposed on a map 90. The map 90 and the mark 91 change in accordance with the movement of the mobile object. For example, when the user performs a predetermined operation (for example, pressing a button or uttering a wake word), the image displayed on the display unit 30 changes as shown in FIG. 6, and the voice recognition unit 130 enters a state in which it can accept a command input by voice.
[0040] In FIG. 6, multiple selectable commands ("convenience store," "restroom," and "gas station") are displayed. The user speaks one of the multiple displayed commands. The voice recognition unit 130 recognizes the user's voice and identifies the spoken command. If the operation unit 20 and the display unit 30 are implemented as a touch panel, the user may select one of the commands by touching it on the screen instead of speaking.
[0041] For example, suppose "Convenience Store" is selected from the commands shown in Fig. 6. Then, the image displayed on the operation unit 20 changes to that shown in Fig. 7. In Fig. 7, multiple selectable commands ("Aiueo Store," "Kakikuke Shop," and "Sa-Shi-Suse Mart") are also displayed. These commands are used to further narrow down the "Convenience Store" selected in the state of Fig. 6, and each indicates the name of a convenience store chain.
[0042] 6, the user speaks one of the multiple commands displayed. The voice recognition unit 130 recognizes the user's voice and identifies the spoken command. If the operation unit 20 and the display unit 30 are implemented as touch panels, the user may select one of the commands by touching it on the screen instead of speaking.
[0043] When any command is selected in the state of Fig. 7, the operation processing unit 150 searches for stores in the chain corresponding to the command and displays them on the map as shown in Fig. 8. The flow of the process executed by the information processing device 10 has been described above.
[0044] When specifying a command by voice, shortening the time from the user's utterance to the operation processing unit 150's response can reduce user stress and improve convenience. To this end, the voice recognition unit 130 identifies a selected command from among multiple commands based on the results of comparing the first voice pattern detected within a period in which the first condition is satisfied with one or more of multiple first candidate voice patterns. That is, the voice recognition unit 130 identifies the selected command by recognizing voice corresponding to a portion of the beginning of the command within the period in which the first condition is satisfied. In other words, the voice recognition unit 130 does not need to recognize voice corresponding to the entire command within the period in which the first condition is satisfied. By recognizing voice corresponding to a portion of the beginning of the command, the voice recognition unit 130 can identify the selected command before the user finishes uttering the entire command. Furthermore, performing such voice recognition only when the first condition is satisfied reduces erroneous recognition, such as recognizing unrelated utterances as commands. That is, by realizing a situation in which the first condition is satisfied, the user can clearly indicate their intention to input voice.
[0045] The determination unit 110 determines whether or not the first condition is satisfied. The speech recognition unit 130 acquires information indicating whether or not the first condition is satisfied from the determination unit 110. An example of the first condition will be described below.
[0046] In one example, the first condition is that a predetermined button is pressed. Note that a state in which a button is pressed includes a state in which the button is pressed within a predetermined time (for example, within 10 seconds or within 30 seconds). In this example, the determination unit 110 acquires, as operation information, information indicating whether the button is pressed from the operation unit 20, and makes a determination based on that information. As described above, the operation unit 20 may be a physical button or a button that is temporarily displayed on a touch panel. In this way, the user can clearly indicate their intention to issue a command by continuing to press the button.
[0047] In another example, the first condition is determined by the surrounding environment, for example, the first condition is at least one of the surrounding environment being quiet and the surrounding environment being noisy.
[0048] In this example, the determination unit 110 obtains a detection result of the level of noise around the user using, for example, a sound sensor such as a microphone installed in the vehicle. Then, for example, the determination unit 110 calculates an average value of the detection results within a predetermined time period (for example, within 10 seconds or within 30 seconds) and determines whether the first condition is satisfied based on whether the average value is equal to or greater than a threshold value. Alternatively, the determination unit 110 may determine whether the first condition is satisfied based on whether other audio output (for example, music) is being output. For example, when the operation processing unit 150 causes the sound output unit 32 to output content such as music in response to a user's voice input or button operation, whether other audio output is being output can be determined by obtaining a control signal from the operation processing unit 150 to the sound output unit 32. This control signal is generated, for example, based on the user's input.
[0049] As shown in Figures 6 and 7, multiple selectable commands may be displayed simultaneously. In this example, the first condition may be determined by the number of times the multiple selectable commands have been displayed to the user within a predetermined period. For example, the first condition may be that a screen including multiple selectable commands has been displayed a predetermined number of times or more within a predetermined period (e.g., within the past month or the past year). Information that can identify this number of displays, such as the display history of each screen (e.g., information indicating the date and time when each screen was displayed for each screen), is stored, for example, in storage unit 101.
[0050] In this example, the determination unit 110 obtains the number of times a screen containing multiple selectable commands is displayed to the user over a predetermined period of time, and determines whether the first condition is met based on whether the number of times the screen is displayed is greater than or equal to a threshold value.
[0051] In another example, the first condition is determined by the speed of the vehicle, for example, the first condition is at least one of the vehicle being stationary and the vehicle being traveling at a low speed.
[0052] In this example, the determination unit 110 obtains the detection result of the vehicle speed from a sensor, and determines that the first condition is satisfied if the detection result is equal to or less than a threshold value. In this case, the threshold value is, for example, 20 km / h. The determination unit 110 obtains the detection result of the vehicle speed via a Controller Area Network (CAN), which is a vehicle control and management system, from a speed sensor connected to the CAN.
[0053] The first condition may be that the vehicle is traveling at a high speed. In this example, the determination unit 110 obtains the detection result of the vehicle speed using the above-mentioned method, and determines that the first condition is met if the detection result is equal to or greater than a threshold value. In this case, the threshold value is, for example, 80 km / h.
[0054] In another example, the first condition is determined by the signal strength of the human voice, for example, the first condition is at least one of the signal strength of the acquired human voice being equal to or greater than a predetermined value or the signal strength of the acquired human voice being equal to or less than a predetermined value.
[0055] In this example, the determination unit 110 first determines whether the sound acquired by the sound sensor is a human voice. If the sound is determined to be a human voice, the determination unit 110 then obtains the detection result of the signal strength of the human voice and determines whether the first condition is satisfied depending on whether the detection result is equal to or greater than a threshold.
[0056] The determination of whether the first condition is satisfied may be repeatedly performed by the determination unit 110, or may be performed by the determination unit 110 when a human voice is detected.
[0057] In another example, the first condition is that a predetermined time T has elapsed since a predetermined gesture was made. The gesture may be, for example, a motion of nodding the head multiple times. Note that a head shake may be a driving motion, and therefore is preferably not treated as a gesture. The time T is, for example, between 3 and 10 seconds.
[0058] In this example, the determination unit 110 acquires, as operation information, information indicating the detection result by the sensor from the operation unit 20. By analyzing the detection result by the sensor, the determination unit 110 can detect a gesture by the movement of the driver's head, as described above. The determination unit 110 records the time when the gesture is detected, and determines that the first condition is satisfied until a predetermined time T has elapsed from that time.
[0059] The voice recognition performed by the voice recognition unit 130 will be described below. When the voice recognition unit 130 acquires voice data from the sound collection unit 22, it detects a voice pattern indicated in the voice data. A voice pattern indicates a waveform of a part of the voice waveform indicated in the voice data whose amplitude is equal to or exceeds a predetermined threshold, or a feature quantity of the waveform of that part. The voice recognition unit 130 performs voice recognition by comparing the detected voice pattern with candidate voice patterns prepared in advance. In other words, among multiple candidate voice patterns, the voice recognition unit 130 identifies a candidate voice pattern that is highly similar to the detected voice pattern. Then, the voice recognition unit 130 identifies a command associated with the candidate voice pattern as the selected command.
[0060] The voice recognition unit 130 may determine the similarity for all or only some of the selectable commands in that state. When determining the similarity for all selectable commands, the voice recognition unit 130 determines the command with the highest similarity as the selected command. When determining the similarity for only some of the selectable commands, the voice recognition unit 130 determines the similarity for each command in turn, and when a similarity higher than a predetermined standard is determined, the voice recognition unit 130 determines the similarity for that command as the selected command. Then, it is not necessary to determine the similarity for the remaining commands.
[0061] In the example of Fig. 3, the information processing device 10 includes a storage unit 101. The storage unit 101 is realized, for example, by using a storage device 1080 of a computer 1000 that realizes the information processing device 10. The voice recognition unit 130 is able to access the storage unit 101. Note that Fig. 3 shows an example in which the storage unit 101 is provided in the information processing device 10, but the storage unit 101 may be provided outside the information processing device 10. Furthermore, the storage unit 101 may be provided as a part of the system 50, or may be provided outside the system 50.
[0062] 9 is a diagram illustrating an example of the configuration of reference data stored in the storage unit 101. In the reference data, a plurality of commands are associated with each of a plurality of states. Each command is also associated with a first candidate voice pattern and a second candidate voice pattern. The table illustrated in FIG. 9 shows the data names of the candidate voice patterns.
[0063] As described above, the speech recognition unit 130 identifies a selected command from among the multiple commands based on the comparison result between the first speech pattern detected within the period satisfying the first condition and one or more of the multiple first candidate speech patterns. Each of the multiple first candidate speech patterns corresponds to, for example, a speech with the minimum number of initial syllables that can identify the command associated with that first candidate speech pattern. In other words, if the initial syllables of the multiple selectable commands are all different, the first candidate speech patterns for those multiple commands may all be speech patterns corresponding to one syllable. If the multiple selectable commands include two or more commands that have the same initial first syllable but different second syllables, the first candidate speech patterns for those two or more commands are speech patterns corresponding to two syllables. This enables speech recognition using the shortest speech patterns.
[0064] For each of the multiple commands, the second candidate voice pattern is a voice pattern longer than the first candidate voice pattern. Specifically, the second candidate voice pattern may be a voice pattern corresponding to the entire command. In other words, for each command, the first candidate voice pattern can be considered a portion of the second candidate voice pattern. The voice recognition unit 130 identifies the selected command based on the results of comparing the voice pattern detected during the period in which the first condition is not satisfied with one or more of the multiple second candidate voice patterns, each associated with one of the multiple commands. This enables highly accurate voice recognition when the first condition is not satisfied.
[0065] The speech recognition unit 130 may be configured to identify a speech period in the speech data where the amplitude of the speech waveform is equal to or greater than a predetermined threshold, and generate detected speech data indicating the speech waveform of that speech period. In this case, the speech recognition unit 130 may compare a portion of the second candidate speech pattern, including at least the beginning thereof, with a portion of the speech waveform including the beginning of that speech period before completing the generation of the detected speech data (e.g., before identifying the end point of that speech period). If the comparison result satisfies a predetermined criterion, the speech recognition unit 130 may identify the command corresponding to the second candidate speech pattern as the selected command.
[0066] The operation processing unit 150 acquires information indicating the selected command identified by the voice recognition unit 130. The operation processing unit 150 can also access the storage unit 101. As shown in Fig. 9, the reference data associates each of a plurality of commands with a process to be executed by the operation processing unit 150. The operation processing unit 150 uses the reference data to identify and execute the process corresponding to the selected command.
[0067] The use of the first and second candidate voice patterns will be further explained using specific examples. Specific examples of the first condition are as described above. The following explanation will be given based on these specific examples.
[0068] For example, the first condition is determined by the surrounding environment, for example, the first condition is at least one of the surrounding environment being quiet and the surrounding environment being noisy.
[0069] In this example, the determination unit 110 obtains the detection result of the noise level around the user using, for example, a sound sensor such as a microphone installed in the vehicle. The determination unit 110 then calculates the average value of the detection results within a predetermined time period (e.g., within 10 seconds or within 30 seconds) and determines whether the first condition is met based on whether the average value is equal to or greater than a threshold. The operation processing unit 150 then identifies the command selected using the first candidate voice pattern if the average value is equal to or less than the threshold, and identifies the command selected using the second candidate voice pattern if the average value is equal to or greater than the threshold. Alternatively, the operation processing unit 150 identifies the command selected using the first candidate voice pattern if no other audio output (e.g., music) is being output, and identifies the command selected using the second candidate voice pattern if other audio output is being output.
[0070] As shown in Figures 6 and 7, multiple selectable commands may be displayed simultaneously. In this example, the first condition may be determined based on the number of times the multiple selectable commands have been displayed to the user within a predetermined period. For example, the first condition may be that a screen including multiple selectable commands has been displayed a predetermined number of times or more within a predetermined period (e.g., within the past month or the past year).
[0071] In this example, the operation processing unit 150 obtains the number of times a screen containing multiple selectable commands is displayed to the user over a predetermined period of time, and if the number of times a screen is displayed is equal to or greater than a threshold, identifies the command selected using the first candidate voice pattern, and if the detection result is equal to or less than the threshold, identifies the command selected using the second candidate voice pattern.
[0072] In another example, the first condition is determined by the speed of the vehicle, for example, the first condition is at least one of the vehicle being stationary and the vehicle being traveling at a low speed.
[0073] In this example, the determination unit 110 obtains the detection result of the vehicle speed using a sensor, and determines that the first condition is met if the detection result is equal to or less than a threshold value. In this case, the threshold value is, for example, 20 km / h. The operation processing unit 150 then identifies the selected command using the first candidate voice pattern if the detection result is equal to or less than the threshold value, and identifies the selected command using the second candidate voice pattern if the detection result is equal to or greater than the threshold value.
[0074] The first condition may be that the vehicle is traveling at a high speed. In this example, if the detection result is equal to or greater than a threshold, it is determined that the first condition is met. In this case, the threshold is, for example, 80 km / h. Then, if the detection result is equal to or greater than the threshold, the operation processing unit 150 identifies the selected command using the first candidate voice pattern, and if the detection result is equal to or less than the threshold, it identifies the selected command using the second candidate voice pattern.
[0075] In another example, the first condition is determined by the signal strength of the human voice, for example, the first condition is at least one of the signal strength of the acquired human voice being equal to or greater than a predetermined value or the signal strength of the acquired human voice being equal to or less than a predetermined value.
[0076] In this example, the determination unit 110 first determines whether the sound acquired by the sound sensor is a human voice. If the determination unit 110 determines that the sound is a human voice, it acquires the detection result of the signal strength of the human voice and determines whether it is equal to or greater than a threshold. If the detection result is equal to or greater than the threshold, the operation processing unit 150 identifies the command selected using the first candidate voice pattern, and if the detection result is equal to or less than the threshold, it identifies the command selected using the second candidate voice pattern.
[0077] The operation processing unit 150 may further identify a command by using a second condition. For example, the operation processing unit 150 may identify a command selected by using a first candidate voice pattern when the second condition is satisfied, and may identify a command selected by using a second candidate voice pattern when the second condition is not satisfied.
[0078] The second condition is, for example, whether or not the selected command can be identified using the first candidate voice pattern the first time. Specifically, when the above-mentioned first condition is satisfied and the selected command cannot be identified using the first candidate voice pattern the first time (for example, when voice recognition cannot be performed within a predetermined time), the selected command may be identified using the second candidate voice pattern.
[0079] The second condition is not limited to this example, and the above-described first condition may be used as the second condition.
[0080] The determination of whether the above-mentioned first and second conditions are satisfied may be made periodically by the determination unit 110, or may be made by the determination unit 110 when a human voice is detected.
[0081] FIG. 10 is a flowchart illustrating the flow of processing executed by the information processing device 10 according to this embodiment. In step S101, the voice recognition unit 130 determines whether an operation has been performed to enable command input reception. Examples of operations to enable command input reception include pressing a predetermined button and detecting a wake word. The voice recognition unit 130 can determine whether a predetermined button has been pressed based on operation information from the operation unit 20. The button may be a physical button or a button temporarily displayed on a touch panel. The voice recognition unit 130 can also detect a predetermined wake word by analyzing audio data from the sound collection unit 22.
[0082] The voice recognition unit 130 repeats step S101 until an operation for enabling command input is performed. When the voice recognition unit 130 detects that an operation for enabling command input has been performed (Yes in step S101), the operation processing unit 150 causes the display unit 30 to display a plurality of selectable commands (step S102). Note that the plurality of commands that can be initially selected after the operation for enabling command input is performed may be determined in advance.
[0083] In step S103, the voice recognition unit 130 accepts a voice input of a command.
[0084] 11 is a flowchart illustrating the process in step S103 by the determination unit 110 and the speech recognition unit 130. First, the determination unit 110 determines whether or not a first condition is satisfied (step S201).
[0085] If the first condition is satisfied (Yes in step S201), the voice recognition unit 130 reads out first candidate voice patterns associated with each of the selectable commands from the storage unit 101. The voice recognition unit 130 also compares the voice data obtained by the sound collection unit 22 with the read out first candidate voice patterns to identify the selected command from the selectable commands, as described above (step S202).
[0086] On the other hand, if the first condition is not satisfied (No in step S201), the voice recognition unit 130 reads out second candidate voice patterns associated with each of the selectable commands from the storage unit 101. The voice recognition unit 130 also compares the voice patterns in the voice data obtained by the sound collection unit 22 with the read out second candidate voice patterns to identify a selected command from the selectable commands, as described above (step S203). Once the selected command is identified, step S103 ends.
[0087] Returning to FIG. 10 , step S104 is executed following step S103. In step S104, the operation processing unit 150 uses the reference data to identify a process associated with the selected command and executes the identified process. That is, if the identified process is a state transition (Yes in step S104), the operation processing unit 150 executes the state transition process associated with the selected command and causes the display unit 30 to display multiple commands associated with the state after the transition (S102). On the other hand, if the identified process is not a state transition (No in step S104), the operation processing unit 150 executes the process associated with the selected command (step S105). Then, the voice recognition unit 130 stops accepting command inputs (step S106). Note that steps S101 to S106 are repeated unless the operation of the information processing device 10 is turned off.
[0088] As described above, the information processing device 10 and information processing method according to this embodiment identify a selected command from among a plurality of commands based on the comparison result between the first voice pattern detected within a period that satisfies a first condition and one or more of a plurality of first candidate voice patterns. This makes it possible to reduce recognition errors while providing a fast response according to the selection result.
[0089] Second Embodiment FIG. 12 is a diagram illustrating a functional configuration of an information processing device 10 according to a second embodiment. The information processing device 10 according to this embodiment includes a voice recognition unit 130 and an operation processing unit 150. The operation processing unit 150 executes processing based on selection results for multiple commands associated with each of multiple states. The voice recognition unit 130 processes received voice data to identify the selection results for the multiple commands. The operation processing unit 150 transitions the state according to the selection results. The operation processing unit 150 also displays an image associated with the state on the display unit. After identifying the selection result in the first state during a period in which a predetermined first condition is satisfied, the voice recognition unit 130 can accept voice data for identifying the selection results for the multiple commands associated with the second state, regardless of whether an image associated with the second state after the transition is displayed on the display unit.
[0090] FIG. 13 is a diagram illustrating an information processing method according to this embodiment. The information processing method according to this embodiment is executed by one or more computers. The information processing method according to this embodiment includes an operation processing step S21 and a voice recognition step S20. In the operation processing step S21, the one or more computers execute processing based on selection results for multiple commands associated with each of multiple states. In the voice recognition step S20, the one or more computers identify the selection results for the multiple commands by processing the received voice data. In the operation processing step S21, the one or more computers transition the state in accordance with the selection results and display an image associated with the state on a display unit. In the voice recognition step S20, the one or more computers can, within a period in which a predetermined first condition is satisfied, receive voice data for identifying the selection results for the multiple commands associated with the second state, regardless of whether an image associated with a second state after the transition is displayed on the display unit after identifying the selection result in the first state.
[0091] The information processing method according to this embodiment can be executed by the information processing device 10 according to this embodiment. The information processing device 10 and the system 50 according to this embodiment are the same as the information processing device 10 and the system 50 according to the first embodiment, respectively, except for the processing performed by the voice recognition unit 130. The information processing device 10 and the information processing device according to this embodiment will be described in detail below.
[0092] Similar to the first embodiment, an example of the configuration of the system 50 according to this embodiment is shown in Fig. 3. The system 50 according to this embodiment includes an information processing device 10, an operation unit 20, a sound collection unit 22, a display unit 30, and a sound output unit 32. The information processing device 10 according to this embodiment may further include a determination unit 110 and a storage unit 101.
[0093] The hardware configuration of a computer that realizes the information processing device 10 according to this embodiment is shown in Fig. 4, for example, similarly to the information processing device 10. However, a storage device 1080 of the computer 1000 that realizes the information processing device 10 according to this embodiment stores program modules that realize the functions of each functional component of the information processing device 10 according to this embodiment.
[0094] The determination unit 110, operation processing unit 150, operation unit 20, sound collection unit 22, display unit 30, and sound output unit 32 are as described in the first embodiment.
[0095] The flow of processing executed by the information processing device 10 according to this embodiment is similar to that described with reference to FIGS. 5 to 8 in the first embodiment.
[0096] In this embodiment, at least some of the multiple commands associated with a state are presented by an image associated with that state. Here, if multiple states (hierarchies) must be passed through before a desired process can be executed, the user may feel stressed if, after specifying one command, the next command cannot be specified until the image on the display unit 30 changes. In particular, a user familiar with the information processing device 10 may be able to determine the next command to be selected even before the next selectable command is displayed on the display unit 30. Therefore, it is preferable to be able to specify the next command without waiting for the next selectable command to be displayed on the display unit 30.
[0097] The voice recognition unit 130 according to the present embodiment, after identifying a selection result in the first state during a period in which the first condition is satisfied, can accept voice data for identifying a selection result for multiple commands associated with the second state, regardless of whether an image associated with the second state after the transition is displayed on the display unit. That is, the voice recognition unit 130 according to the present embodiment, after identifying a selected command during a period in which the first condition is satisfied, can continuously accept voice data for identifying a next selected command. Therefore, it is possible to suppress erroneous voice recognition while reducing the hassle of waiting for a screen transition.
[0098] 14 is a flowchart illustrating the flow of processing executed by the information processing device 10 according to this embodiment. Step S301 is the same as step S101 in FIG.
[0099] When the voice recognition unit 130 detects that an operation has been performed to make the device ready to accept command input (Yes in step S301), the judgment unit 110 judges whether the first condition is satisfied (step S302).
[0100] If the first condition is not satisfied (No in step S302), similar to step S102 in FIG. 10 , the operation processing unit 150 displays a plurality of selectable commands on the display unit 30 (step S303). Then, in step S304, the voice recognition unit 130 accepts a voice input of a command. That is, the voice recognition unit 130 acquires voice data of a voice uttered while the plurality of selectable commands are displayed on the display unit 30, and performs voice recognition. In step S304, the voice recognition unit 130 performs voice recognition using the second candidate voice pattern. Once the selected command is identified, the process proceeds to step S306.
[0101] If the first condition is satisfied (Yes in step S302), the voice recognition unit 130 accepts a voice input of a command in step S305. At this time, the operation processing unit 150 simultaneously performs processing for displaying a plurality of selectable commands on the display unit 30. That is, the voice recognition unit 130 can acquire voice data of a voice uttered when the plurality of selectable commands are not displayed on the display unit 30, and perform voice recognition. The relationship between the timing at which the plurality of commands are displayed on the display unit 30 and the timing at which the command is uttered is not important, and the voice recognition unit 130 can accept both a command uttered before the plurality of commands are displayed on the display unit 30 and a command uttered while the plurality of commands are displayed. Once the selected command is identified, the process proceeds to step S306.
[0102] In step S306, the operation processing unit 150 uses the reference data to identify a process associated with the selected command. If the identified process is a state transition (Yes in step S306), the process returns to step S302. On the other hand, if the identified process is not a state transition (No in step S306), the operation processing unit 150 executes the process associated with the selected command (step S307). Then, the voice recognition unit 130 stops accepting command input (step S308). Note that steps S301 to S308 are repeated unless the operation of the information processing device 10 is turned off.
[0103] The speech recognition unit 130 according to this embodiment may not use the first candidate speech pattern for speech recognition when the first condition is satisfied, or may use the first candidate speech pattern as in the first embodiment.
[0104] If the first condition is satisfied and the first candidate speech pattern is not used for speech recognition, the speech recognition unit 130 uses the second candidate speech pattern for speech recognition in step S305. In this case, the reference data does not need to include the first candidate speech pattern.
[0105] When the first condition is satisfied, if the first candidate voice pattern is to be used for voice recognition as in the first embodiment, the voice recognition unit 130 uses the first candidate voice pattern in the voice recognition in step S305. That is, the voice recognition unit 130 identifies a selected command from among the multiple commands based on a comparison result between the first detected voice pattern of the voice data received within the period in which the first condition is satisfied and one or more of the multiple first candidate voice patterns, each of which is associated with one of the multiple commands.
[0106] When using the information processing device 10 according to this embodiment, the user can also utter a command for a first state and a command for a second state in succession. In this case, there is almost no silent interval between the two commands in the voice data. In response to this, the voice recognition unit 130 can identify the start point of the second command using a second candidate voice pattern, which is a voice pattern corresponding to the entire command. That is, when the voice recognition unit 130 receives voice data in the first state and identifies the selected command, it identifies a portion of the voice pattern indicated by the voice data that corresponds to the second candidate voice pattern for that command. The portion following the identified portion is then treated as the voice pattern of the voice data received for the next, second state.
[0107] As described above, according to the information processing device 10 and the information processing method of the present embodiment, after the voice recognition unit 130 has identified a selection result in a first state during a period in which a predetermined first condition is satisfied, the voice recognition unit 130 can accept voice data for identifying a selection result for multiple commands associated with the second state, regardless of whether an image associated with the second state after the transition is displayed on the display unit. This reduces the hassle of the user having to wait for a screen transition.
[0108] Although the embodiments and examples have been described above with reference to the drawings, these are merely examples of the present invention, and various configurations other than those described above can also be adopted.
[0109] In addition, although the flowcharts used in the above description show multiple steps (processes) in a sequential order, the order of steps executed in each embodiment is not limited to the order shown. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of content. Furthermore, the above-described embodiments can be combined as long as the content is not contradictory.
[0110] Examples of reference forms are given below. 1-1. An information processing device including a voice recognition unit that identifies a selected command from a plurality of commands based on a comparison result between a first voice pattern detected within a period satisfying a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands. 1-2. The information processing device described in 1-1., wherein each of the plurality of first candidate voice patterns corresponds to a voice with a minimum number of initial syllables that can identify the command associated with that first candidate voice pattern. 1-3. The information processing device described in 1-1. or 1-2., wherein the first condition is that a predetermined button is pressed. 1-4. The information processing device described in 1-3., wherein the button is provided on a handle of a mobile object. 1-5. The information processing device described in 1-4., wherein the mobile object is a motorcycle. 1-6. 1-1. or 1-2. 1-7. The information processing device described in any one of 1-1. to 1-6., wherein the first condition is within a predetermined time period after a predetermined gesture is made. 1-7. The information processing device described in any one of 1-1. to 1-6., further comprising a determination unit that determines whether the first condition is satisfied. 1-8. The information processing device described in any one of 1-1. to 1-7., wherein the voice recognition unit identifies the selected command based on a comparison result between a voice pattern detected within a period in which the first condition is not satisfied and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands, and the second candidate voice pattern is longer than the first candidate voice pattern for each of the plurality of commands. 1-9. The information processing device described in any one of 1-1. to 1-8., further comprising an operation processing unit that executes processing based on the command identified by the voice recognition unit.1-10. An information processing device according to any one of 1-1. to 1-9., wherein, within a period during which the first condition is satisfied, the voice recognition unit is capable of continuously accepting voice data for identifying a next selected command after identifying the selected command. 1-11. An information processing method executed by one or more computers, comprising a voice recognition step of identifying a selected command from among the plurality of commands based on a comparison result between a first voice pattern detected within a period during which a predetermined first condition is satisfied and one or more of a plurality of first candidate voice patterns each associated with one of the plurality of commands. 1-12. A program that causes a computer to function as voice recognition means that identifies a selected command from among the plurality of commands based on a comparison result between a first voice pattern detected within a period during which a predetermined first condition is satisfied and one or more of a plurality of first candidate voice patterns each associated with one of the plurality of commands. 2-1. 2-2. An information processing device comprising: an operation processing unit that executes processing based on selection results for a plurality of commands associated with each of a plurality of states; and a voice recognition unit that processes received voice data to identify the selection results for the plurality of commands, wherein the operation processing unit transitions the state according to the selection results and displays an image associated with the state on a display unit, and the voice recognition unit is capable of accepting voice data for identifying the selection results for the plurality of commands associated with the second state, after identifying the selection result in the first state, within a period in which a predetermined first condition is satisfied, regardless of whether an image associated with the second state after the transition is displayed on the display unit. 2-2. An information processing device according to 2-1, wherein at least some of the plurality of commands associated with the state are presented by the image associated with that state. 2-3. An information processing device according to 2-1 or 2-2, wherein the first condition is that a predetermined button is pressed.2-4. An information processing device according to 2-3., wherein the button is provided on a handle of a mobile object. 2-5. An information processing device according to 2-4., wherein the mobile object is a motorcycle. 2-6. An information processing device according to 2-1. or 2-2., wherein the first condition is that a predetermined time period has passed since a predetermined gesture was made. 2-7. An information processing device according to any one of 2-1. to 2-6., wherein the voice recognition unit identifies a selected command from the plurality of commands based on a comparison result between a first detected voice pattern of the voice data received within a period satisfying the first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands. 2-8. A system comprising: an information processing device according to any one of 2-1. to 2-7.; and the display unit. 2-9. An information processing method executed by one or more computers, comprising: an operation processing step of executing processing based on selection results for a plurality of commands associated with each of a plurality of states; and a voice recognition step of identifying the selection results for the plurality of commands by processing received voice data, wherein in the operation processing step, a state is transitioned in accordance with the selection results, and an image associated with the state is displayed on a display unit, and in the voice recognition step, after identifying the selection result in a first state within a period in which a predetermined first condition is satisfied, voice data can be received to identify the selection result for the plurality of commands associated with the second state, regardless of whether an image associated with the second state after the transition is displayed on the display unit.2-10. A program that causes a computer to function as: operation processing means that executes processing based on selection results for a plurality of commands associated with each of a plurality of states; and voice recognition means that processes accepted voice data to identify the selection results for the plurality of commands, wherein the operation processing means transitions the state according to the selection results and causes an image associated with the state to be displayed on a display unit, and the voice recognition means is capable of accepting voice data for identifying the selection results for the plurality of commands associated with the second state, after identifying the selection result in the first state, within a period that satisfies a predetermined first condition, regardless of whether an image associated with the second state after the transition is displayed on the display unit.
[0111] This application claims priority based on Japanese Patent Application No. 2024-116979 filed on July 22, 2024, and PCT / JP2025 / 007494 filed on March 3, 2025, the disclosures of which are incorporated herein in their entirety.
[0112] REFERENCE SIGNS LIST 10 Information processing device 20 Operation unit 22 Sound collection unit 30 Display unit 32 Sound output unit 50 System 90 Map 91 Mark 101 Storage unit 110 Determination unit 130 Voice recognition unit 150 Operation processing unit 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface
Claims
1. An information processing device comprising: a voice recognition unit that identifies a selected command from among a plurality of commands based on a comparison result between a voice pattern detected within a period that satisfies a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands; wherein the voice recognition unit identifies the selected command based on a comparison result between a voice pattern detected within a period that does not satisfy the first condition and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands; and wherein, for each of the plurality of commands, the second candidate voice pattern is longer than the first candidate voice pattern for the same command.
2. An information processing device according to claim 1, wherein each of the plurality of first candidate voice patterns corresponds to a voice with a minimum number of initial syllables that can identify a command associated with that first candidate voice pattern.
3. An information processing device according to claim 1 or 2, wherein the first condition is that a predetermined button is pressed.
4. An information processing device according to claim 3, wherein the button is provided on a handle of a mobile object.
5. An information processing device according to claim 4, wherein the moving body is a motorcycle.
6. An information processing device according to claim 1 or 2, wherein the first condition is that a predetermined time has elapsed since a predetermined gesture was made.
7. An information processing device according to claim 1 or 2, wherein the first condition is that the volume of noise around the user is equal to or less than a predetermined value.
8. An information processing device according to any one of claims 1 to 7, further comprising a determination unit that determines whether or not the first condition is satisfied.
9. An information processing device according to any one of claims 1 to 8, further comprising an operation processing unit that executes processing based on the command identified by the voice recognition unit.
10. An information processing device according to any one of claims 1 to 9, wherein, within a period in which the first condition is satisfied, the voice recognition unit is capable of continuously accepting voice data for identifying a next selected command after identifying the selected command.
11. An information processing method executed by one or more computers, comprising: a voice recognition step of identifying a selected command from among a plurality of commands based on a comparison result between a voice pattern detected within a period satisfying a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands; in the voice recognition step, the selected command is identified based on a comparison result between a voice pattern detected within a period not satisfying the first condition and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands; and, for each of the plurality of commands, the second candidate voice pattern is longer than the first candidate voice pattern for the same command.
12. A program that causes a computer to function as a voice recognition means that identifies a selected command from a plurality of commands based on a comparison result between a voice pattern detected within a period that satisfies a predetermined first condition and one or more of a plurality of first candidate voice patterns, each of which is associated with one of the plurality of commands; the voice recognition means identifies the selected command based on a comparison result between a voice pattern detected within a period that does not satisfy the first condition and one or more of a plurality of second candidate voice patterns, each of which is associated with one of the plurality of commands; and for each of the plurality of commands, the second candidate voice pattern is longer than the first candidate voice pattern for the same command.
Citation Information
Patent Citations
Speech recognition device
JP2004301875A
Information processing device, information processing method, program, and moving body
WO2019069731A1