Input / Output Devices

The input/output device adjusts its response output based on surrounding conditions to prevent detection, addressing the inconvenience of fixed responses in conventional speech recognition devices by reducing audio and visual visibility when necessary.

JP2026048974APending Publication Date: 2026-03-17PIONEER IP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Conventional speech recognition devices respond in a fixed manner without considering the speaker's surroundings, leading to inconvenience when users want to avoid others hearing or seeing the response.

Method used

An input/output device that adjusts its response output based on surrounding conditions, using audio and display units to prevent the response from being heard or seen by others when the input voice level is low, and includes mechanisms to calculate the signal-to-noise ratio for ambient noise.

Benefits of technology

Enables the device to change its output to be less noticeable in private settings, maintaining recognition accuracy while ensuring the response is not heard or seen by others, thus enhancing user convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048974000001_ABST
    Figure 2026048974000001_ABST
Patent Text Reader

Abstract

The present invention provides an input / output device that can change its response to an input according to the surrounding conditions and output a modified version of that response. [Solution] In the voice recognition device 1, the level check unit 31 detects the level of the audio signal output from the microphone 2, and the use case determination unit 33 determines whether the detected audio signal level is lower than a predetermined audio signal level. If the detected audio signal level is lower than a predetermined audio signal level, the volume of sound output from the speaker is reduced and the brightness of the display device is lowered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an input / output device used in a voice recognition device that recognizes spoken voice and the like.

Background Art

[0002] In recent years, in in-vehicle devices, portable devices, and the like, in order to enable easy operation only by voice without the need for operations such as buttons, many devices incorporate a voice recognition device (voice recognition function).

[0003] In this type of voice recognition device, for an input voice, a processing result corresponding to the input voice is output as a response in display information such as voice information or an image, or a response such as an acceptance of an input or a recognition result is output in voice information or display information. Such a response method has been performed in a fixed manner, such as a certain voice level or a certain brightness, without considering the situation around the speaker.

[0004] As a method of operating a voice recognition device in consideration of the situation around the speaker, the method described in Patent Document 1 can be cited as an example. The voice recognition device described in Patent Document 1 amplifies the input voice level to an appropriate level according to the usage state of the portable information terminal device, making it possible to prevent a decrease in the recognition rate.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] <了 Conventional speech recognition devices always respond in a predetermined format without considering the speaker's surroundings. Therefore, if the speaker does not want others to hear or see the content of their response, the device cannot be used, and the only option is to refrain from using the speech recognition device. Consequently, in such situations, users have to operate the device using buttons or other means, which is inconvenient.

[0007] The speech recognition device described in Patent Document 1 controls the input speech level solely to prevent a decrease in recognition rate, and does not take into account the response from the speech recognition device as described above.

[0008] Therefore, in view of the above-mentioned problems, the object of the present invention is to provide an input / output device that can change the response to an input according to the surrounding conditions and output a corresponding response. [Means for solving the problem]

[0009] To solve the above problems, the invention described in claim 1 is characterized by comprising: an audio output unit that outputs sound; a display unit that displays an image; an audio signal output unit that outputs the user's spoken voice as an audio signal to an audio recognition means; a response information acquisition unit that acquires response information according to the recognition result of the audio recognition means; and a control unit that, when the audio level of the audio signal is lower than a predetermined audio level, prevents the audio output unit from outputting sound based on the response information and displays an image based on the response information using the display unit.

[0010] The invention described in claim 4 is an input / output method performed by an input / output device comprising an audio output unit and a display unit, characterized in that it includes: an audio signal output step of outputting the user's spoken voice as an audio signal to an audio recognition means; a response information acquisition step of acquiring response information corresponding to the recognition result of the audio recognition means; and a control step of not outputting audio based on the response information by the audio output unit and displaying an image based on the response information by the display unit when the audio level of the audio signal is lower than a predetermined audio level.

[0011] The invention described in claim 5 is an input / output program executed by a computer of an input / output device comprising an audio output unit and a display unit, characterized in that the computer functions as: an audio signal output unit that outputs the user's spoken voice as an audio signal to an audio recognition means; a response information acquisition unit that acquires response information corresponding to the recognition result of the audio recognition means; and a control unit that, when the audio level of the audio signal is lower than a predetermined audio level, prevents the audio output unit from outputting audio based on the response information and displays an image based on the response information using the display unit.

[0012] The invention described in claim 6 is characterized by storing the input / output program described in claim 5. [Brief explanation of the drawing]

[0013] [Figure 1] This is a configuration diagram of an input / output device according to a first embodiment of the present invention. [Figure 2] Figure 1 is a flowchart illustrating the operation of the input / output device. [Figure 3] This is a configuration diagram of an input / output device according to a second embodiment of the present invention. [Figure 4] Figure 2 is a flowchart illustrating the operation of the input / output device. [Figure 5] This is a configuration diagram of an input / output device according to another embodiment of the present invention. [Figure 6] This is a configuration diagram of an input / output device according to another embodiment of the present invention. [Modes for carrying out the invention]

[0014] The following describes an input / output device according to one embodiment of the present invention. The input / output device according to one embodiment of the present invention includes a first sound collection means for collecting spoken input sound, a first output means for outputting the input sound collected by the first sound collection means to a speech recognition means, a response acquisition means for acquiring a response from the speech recognition means, and a second output means for outputting the response acquired by the response acquisition means. Furthermore, it includes a speech level comparison means for detecting the input sound level, which is the sound level of the input sound collected by the first sound collection means, and comparing the input sound level with a predetermined sound level, and a control means for changing the output of the second output means so that the response is less likely to be recognized by the surroundings when the input sound level compared by the speech level comparison means is lower than the predetermined sound level. In this way, when the sound level of the input sound is low, it is possible to change the output of the second output means so that it is less likely to be recognized by the surroundings, based on the judgment that the speech recognition response should not be heard or seen by the surroundings.Therefore, the response to the input can be changed and output according to the surrounding circumstances.

[0015] Furthermore, the second output means includes an audio output means that outputs the response as sound, and the control means may reduce the volume of the sound output from the audio output means if the result of the comparison by the audio level comparison means is lower than a predetermined audio level. In this way, the volume of the sound output from the audio output means, such as a speaker, can be reduced when the voice recognition response should not be heard by those around.

[0016] Furthermore, the second output means may further include a display means for displaying the response as an image, and the control means may stop the display on the display means and reduce the volume output from the audio output means if the result of the comparison by the audio level comparison means is lower than a predetermined audio level. In this way, when both the audio output means and the display means are present, the display on the display means can be stopped and the volume output from the audio output means such as a speaker can be reduced.

[0017] Further, the second output means further has an output interface for outputting the response as sound from an external audio output means, and the control means may be configured to output the response only to the output interface when the result compared by the audio level comparison means is smaller than a predetermined audio level. By doing so, when it is not desired for the response of the voice recognition to be heard by the surroundings, sound can be output only from an external audio output means such as an earphone.

[0018] Further, the second output means has display means for displaying the response as an image, and the control means may be configured to change the display of the display means so that the image is less recognizable from the surroundings when the result compared by the audio level comparison means is smaller than a predetermined audio level. By doing so, when it is not desired for the response of the voice recognition to be seen by the surroundings, for example, the brightness and viewing angle of display means such as a liquid crystal display can be changed.

[0019] Further, the second output means further has audio output means for outputting the response as sound, and the control means may stop the output of the audio output means and change the display of the display means so that the image is less recognizable from the surroundings when the result compared by the audio level comparison means is smaller than a predetermined audio level. By doing so, when both the audio output means and the display means are provided, it is possible to stop the output of the sound from the audio output means and make the display of the display device less recognizable.

[0020] Also, an input / output device according to an embodiment of the present invention includes: a first sound collection means for collecting the input voice that has been spoken; a second sound collection means for collecting ambient sound other than the input voice; a first output means for outputting the input voice collected by the first sound collection means to a voice recognition means; an ambient sound level detection means for detecting an ambient sound level which is the voice level of the ambient sound collected by the second sound collection means; a response acquisition means for acquiring a response from the voice recognition means; and a second output means for outputting the response acquired by the response acquisition means. And it further includes: a ratio calculation means for detecting an input voice level which is the voice level of the input voice collected by the first sound collection means, and calculating a ratio between the input voice level and the ambient sound level detected by the ambient sound level detection means; and a control means for changing the output of the second output means so that it becomes difficult to recognize a response from the surroundings when the ratio calculated by the ratio calculation means is smaller than a predetermined value. By doing so, the situation around the speaker can be judged from the ratio between the input voice and the ambient sound. That is, when the ratio (S / N ratio) of the input voice level of the spoken voice to the ambient sound level is small, it can be judged that there are many people around and the speaker is speaking in a low voice, so the output of the output means can be changed because it is not desired for the response of the voice recognition to be heard or seen by the surroundings.

[0021] Also, an input / output method according to an embodiment of the present invention is an input / output method in an input / output device that outputs a response from a voice recognition means to a spoken input voice, and includes: a voice level comparison step of detecting an input voice level which is the voice level of the voice collected by a first sound collection means that collects the input voice, and comparing the input voice level with a predetermined voice level; and a control step of changing the output of the response so that it becomes difficult to recognize the response of the voice recognition means from the surroundings when the input voice level compared in the voice level comparison step is smaller than the predetermined voice level. By doing so, when the voice level of the input voice is small, it can be judged that it is not desired for the response of the voice recognition to be heard or seen by the surroundings, and the output of the response can be changed. Therefore, the response to the input can be changed and output according to the situation around.

[0022] Alternatively, the above-described input / output method may be configured as an input / output program that is executed by a computer. In this way, the computer can be used to determine that if the audio level of the input voice is low, the voice recognition response should not be heard or seen by those around, and the output of the response can be changed accordingly. Therefore, the response to the input can be changed and output according to the surrounding circumstances.

[0023] Furthermore, the aforementioned speech recognition program may be stored on a computer-readable recording medium. This allows the program to be distributed independently, in addition to being incorporated into a device, and makes version upgrades and other modifications easier.

[0024] Furthermore, an input / output method according to one embodiment of the present invention is an input / output method in an input / output device that outputs a response from a speech recognition means to spoken input sound, and includes a ratio calculation step of calculating the ratio between the input sound level and the ambient sound level, which is the sound level of the sound collected by a first sound collection means that collects the input sound, and the ambient sound level, which is the sound level of the ambient sound collected by a second sound collection means that collects ambient sounds other than the input sound, and when the ratio compared in the ratio calculation step is smaller than a predetermined value, a control step of changing the output of the response so that the response of the speech recognition means is less likely to be recognized from the surroundings. In this way, the situation around the speaker can be determined from the ratio of the input sound and the ambient sound. In other words, if the ratio of the spoken input sound level to the ambient sound level (S / N ratio) is small, it can be determined that there are many people around and the speaker is speaking in a quiet voice, so the output of the output means can be changed if the speaker does not want the response of the speech recognition to be heard or seen by those around them.

[0025] Alternatively, the above-described input / output method may be configured as an input / output program executed by a computer. In this way, the computer can be used to determine if the speech recognition response should not be heard or seen by others when the signal-to-noise ratio is low, and change the output of the response accordingly. Therefore, the response to the input can be changed and output according to the surrounding circumstances.

[0026] Furthermore, the aforementioned speech recognition program may be stored on a computer-readable recording medium. This allows the program to be distributed independently, in addition to being incorporated into a device, and makes version upgrades and other modifications easier. [Examples]

[0027] A speech recognition device having an input / output device according to a first embodiment of the present invention will be described with reference to Figures 1 and 2. As shown in Figure 1, the speech recognition device 1 includes a microphone 2, a control device 3, and an external output device 4.

[0028] The microphone 2, acting as the first sound collection means, collects the voice (input voice) spoken by the user, converts it into an electrical signal, and outputs it as an audio signal to the control device 3.

[0029] The control device 3 includes a level check unit 31, a speech recognition engine unit 32, and a use case determination unit 33. The control device 3 is composed of, for example, a microcomputer (MCU), a digital signal processor (DSP), or an ASIC (Application Specific Integrated Circuit).

[0030] The first output means, the level check unit 31 as an audio level comparison means, outputs the audio signal input from the microphone 2 to the speech recognition engine unit 32. That is, it outputs the input audio collected by the first sound collection means to the speech recognition means. The level check unit 31 detects the level of the audio signal input from the microphone 2 and outputs it as the input audio level to the use case determination unit 33. That is, it detects the input audio level, which is the audio level of the input audio collected by the first sound collection means. In this specification, the level of the audio signal indicates the loudness of the sound in question, for example, the maximum or average value of the amplitude of the audio signal.

[0031] The speech recognition engine unit 32 converts the audio signal input from the level check unit 31 into a digital signal and performs speech recognition processing (the level check unit 31 may also perform the conversion into a digital signal). The speech recognition processing is not particularly limited and can use any known method, such as statistical methods, dynamic time stretching, or hidden Markov models. The speech recognition engine unit 32 outputs a response regarding the result of the speech recognition processing to the external output device 4. The response regarding the result of the speech recognition processing is not limited to audio information or display information relating to the response to the spoken audio content, but also includes audio information or display information indicating that the audio was recognized, or audio information or display information indicating that the audio was not recognized, or audio information or display information prompting input such as the next command.

[0032] Furthermore, if the speech recognition engine unit 32 determines that the speech recognition result is a command for another processing unit (not shown), it outputs the command to that other processing unit. Note that this other processing unit is not limited to being integrated with the speech recognition device 1, but may be detachable or communicate wirelessly or wired via a network or the like. In the configuration shown in Figure 1, the control device 3 includes the speech recognition engine unit 32, so the speech recognition engine unit 32 also serves as the speech recognition means and the response acquisition means for acquiring responses from the speech recognition means.

[0033] The use case determination unit 33, acting as both an audio level comparison means and a control means, determines that if the input audio level detected by the level check unit 31 is lower than a predetermined audio signal level (a predetermined audio level), it is in private mode, which indicates a situation where the user does not want their voice recognition response to be heard or seen by others. The unit then outputs a control signal to the external output device 4 to change its output to one corresponding to this private mode. In other words, it compares the input audio level with a predetermined audio level. When the input audio level compared by the audio level comparison means is lower than the predetermined audio level, it changes the output of the second output means so that the response is less likely to be recognized by others.

[0034] Furthermore, since a low input audio level may reduce the recognition rate in the speech recognition engine unit 32, it is desirable to set the predetermined audio signal level within a range that does not reduce the recognition rate in the speech recognition engine unit 32. Alternatively, speech recognition processing may be performed after applying processing to reduce the influence of ambient noise, such as the processing described in Patent Document 1.

[0035] In Figure 1, the control device 3 is shown as an integrated unit comprising a level check unit 31, a voice recognition engine unit 32, and a use case determination unit 33. However, it is not limited to this configuration. For example, each of these components may be composed of separate parts (microcontroller, DSP, ASIC, etc.).

[0036] The external output device 4, as a second output means, includes an audio output unit 41 as an audio output means and a display unit 42 as a display means. The audio output unit 41 includes a speaker that outputs responses as audio information from the responses related to the results of the speech recognition processing output from the speech recognition engine unit 32, and an amplifier that controls the volume output to the speaker. The display unit 42 includes a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display that displays responses as display information from the responses related to the results of the speech recognition processing output from the speech recognition engine unit 32 as images (including information containing only text), and a driver circuit that controls the display of the display device. In other words, the external output device 4 outputs the responses acquired by the response acquisition means.

[0037] Then, when the use case determination unit 33 determines that it is in private mode and receives a control signal that changes the output, the audio output unit 41 changes the amplification ratio of the amplifier, etc., so that the sound output from the speaker becomes quieter. In other words, if the result of the comparison by the audio level comparison means is lower than a predetermined audio level, the sound output from the audio output means is reduced. In addition, the driver circuit of the display unit 42 controls the brightness of the display device to decrease. In other words, if the result of the comparison by the audio level comparison means is lower than a predetermined audio level, the display of the display means is changed so that the image is difficult to recognize from the surroundings.

[0038] As is clear from the above description, the microphone 2, level check unit 31, use case determination unit 33, and external output device 4 constitute the input / output device 10 according to the first embodiment of the present invention.

[0039] Next, the operation of the input / output device 10 with the above configuration will be explained with reference to the flowchart in Figure 2. The flowchart shown in Figure 2 is executed by the control device 3.

[0040] First, in step S11, the audio signal of the input audio is input from microphone 2 to level check unit 31, and the process proceeds to step S12.

[0041] Next, in step S12, the level check unit 31 detects the input audio level of the audio signal input from microphone 2 and outputs it to the use case determination unit 33, and the process proceeds to step S13.

[0042] Next, in step S13, the use case determination unit 33 compares the input audio level detected by the level check unit 31 with a predetermined audio signal level. If the level is lower than the predetermined audio signal level (YES), the process proceeds to step S14; if the level is equal to or greater than the predetermined audio signal level (NO), the process proceeds to step S15. In other words, steps S12 and S13 function as audio level comparison steps.

[0043] Next, in step S14, since it was determined in step S13 that the audio signal level was lower than a predetermined level, the use case determination unit 33 changes the output of the external output device 4 to make it difficult to perceive from the surroundings in private mode (output control). Specifically, as described above, the audio output unit 41 changes the amplification factor of the amplifier, etc., so that the sound output from the speaker is lower than the default volume, and the display unit 42 causes the driver circuit to control the brightness of the display device to be lower than the default brightness. In other words, this step functions as a control step. Here, the default volume and brightness are the volume and brightness in the initial state of the voice recognition device 1.

[0044] On the other hand, in step S15, since it was determined in step S13 that the volume and brightness were above a predetermined level, the use case determination unit 33 sets the volume and brightness to the default values ​​for normal mode. In other words, if the volume and brightness were at the default values ​​before this step was executed, they remain unchanged. If the volume and brightness were lower than the default values ​​before this step was executed, they are returned to the default values.

[0045] In this embodiment, the voice recognition device 1 detects the input audio level output from the microphone 2 using a level check unit 31, and the use case determination unit 33 determines whether the detected input audio level is lower than a predetermined audio signal level. If the input audio level is lower than a predetermined audio signal level, the sound output from the speaker is reduced, and the brightness of the display device is lowered. In this way, when the input audio level is low, it can be determined that the voice recognition response should not be heard or seen by others, and the sound can be reduced or the brightness lowered accordingly. Therefore, the response to the input can be changed and output according to the surrounding circumstances. [Examples]

[0046] Next, a speech recognition device 1 according to a second embodiment of the present invention will be described with reference to Figures 3 and 4. Note that parts identical to those in the first embodiment described above are denoted by the same reference numerals and their descriptions are omitted.

[0047] The input / output device 10 in this embodiment has a microphone 5 added to the speech recognition device 1 shown in Figure 1. The microphone 5, as a second sound collection means, does not collect the voice spoken by the user, but rather collects ambient sounds (sounds around the speech recognition device 1). In other words, it collects ambient sounds other than the spoken input voice.

[0048] The ambient sound collected by microphone 5 is detected at a level check unit 31, and the level of that audio signal (ambient sound level) is output to the use case determination unit 33. In other words, the level check unit 31 functions as an ambient sound level detection means that detects the ambient sound level, which is the audio level of the ambient sound collected by the second sound collection means.

[0049] The use case determination unit 33 calculates the ratio (S / N ratio) between the input audio level collected by the microphone 2 detected by the level check unit 31 and the ambient sound level. In this embodiment, the S / N ratio is the value obtained by dividing the input audio level by the ambient sound level (input audio level / ambient sound level). If the calculated S / N ratio is smaller than a predetermined value, it is determined to be in private mode, and a control signal is output to the external output device 4 to change the output to one corresponding to private mode. In other words, the use case determination unit 33 functions as a ratio calculation means.

[0050] In other words, a low S / N ratio means that ambient noise is relatively loud compared to the user's speech, so it can be inferred that the user is speaking softly in a situation with many people around. Therefore, if the S / N ratio is low, it is determined that the user does not want the response of the speech recognition engine unit 32 to be heard or seen by those around them, and the system is activated in private mode. The operation of the external output device 4 in private mode is the same as in the first embodiment. That is, the sound output from the speaker is reduced, and the brightness of the image displayed on the display device is lowered so that it is difficult for those around to recognize it.

[0051] Next, the operation of the speech recognition device 1 in this embodiment will be explained with reference to the flowchart in Figure 4. The flowchart shown in Figure 4 is executed by the control device 3.

[0052] First, in step S21, audio signals are input from microphones 2 and 5 to the level check unit 31, and the process proceeds to step S12.

[0053] Next, in step S22, the level check unit 31 detects the input audio level of the audio signal input from microphone 2 and the ambient sound level of the audio signal input from microphone 5, outputs them to the use case determination unit 33, and proceeds to step S23.

[0054] Next, in step S23, the use case determination unit 33 calculates the ratio (S / N ratio) between the input audio level and the ambient sound level detected by the level check unit 31. If the S / N ratio is less than a predetermined value (YES), the process proceeds to step S24; if it is greater than or equal to the predetermined value (NO), the process proceeds to step S25. In other words, steps S22 and S23 function as a ratio calculation process.

[0055] Steps S24 and S25 are the same as steps S14 and S15 in Figure 2.

[0056] In this embodiment, the voice recognition device 1 detects the input voice level and the ambient sound level output from the microphone 5 (ambient sound level) in the level check unit 31, and the use case determination unit 33 determines whether the ratio of the input voice level to the ambient sound level (S / N ratio) is smaller than a predetermined value. If the S / N ratio is smaller than a predetermined value, the device reduces the volume of the sound output from the speaker and lowers the brightness of the display device, for example. In this way, if the S / N ratio is small, the device can determine that it does not want the voice recognition response to be heard or seen by others, and therefore reduce the volume or brightness. Thus, the response to the input can be changed according to the surrounding conditions.

[0057] In the two embodiments described above, the brightness of the display device of the display unit 42 was reduced to make the displayed image difficult to recognize from the surroundings. However, the invention is not limited to this, and for example, the viewing angle of the display device may be narrowed. In this case, for example, a filter that changes the polarization direction by changing the orientation state of the liquid crystal by applying a voltage to the liquid crystal element may be provided on the surface of the display device.

[0058] Furthermore, in the two embodiments described above, the control of both the audio output unit 41 and the display unit 42 was changed, but it is also possible to change only one of them.

[0059] Furthermore, as in the second embodiment described above, when both a speaker (audio output unit 41) and a display device (display unit 42) are present, if private mode is detected, the display device may be stopped (the screen may be turned off), and the sound output by the speaker may be reduced. Alternatively, the sound output from the speaker may be stopped, and the brightness of the display device may be reduced or the viewing angle narrowed. In other words, when both an audio output means and a display means are present, stopping the operation of one of them is also included in changing the output so that the response is less likely to be recognized from the surroundings.

[0060] Furthermore, the speech recognition engine unit 32 is not limited to the configuration included in the control device 3 as shown in Figures 1 and 3, but may also be provided in an external server that communicates wirelessly or via wired connection over a network, for example. An example of this is shown in Figure 5. In Figure 5, a communication unit 34 is provided in the control device 3. The communication unit 34 outputs the voice signal input from the level check unit 31 to the speech recognition engine unit 21 provided in the server 20 connected to the Internet 30. The communication unit 34 then outputs the response input from the speech recognition engine unit 21 to an external output device 4 or other processing device. In the case shown in Figure 5, the communication unit 34 functions as a first output means and a response acquisition means.

[0061] Furthermore, as shown in Figure 6, the device may have terminals for connecting external audio output means 6 such as earphones or headphones, and output interfaces 43 such as circuits and antennas for wireless communication with the external audio output means 6 using Bluetooth® or similar technologies.

[0062] The output interface 43 shown in Figure 6 can be switched between the audio output unit 41 and the selector switch 44. In other words, when earphones or headphones are connected, the selector switch 44 is switched to the output interface 43 side so that no sound is output from the speaker of the audio output unit.

[0063] In the case where the output interface 43 shown in Figure 6 is present, if private mode is detected, the display on the display device may be stopped, and the sound (voice signal) of the response from the voice recognition engine unit 32 may be output only from the output interface. In this way, if the voice recognition response is not to be heard by others, only the sound can be output from an external audio output means such as earphones or headphones.

[0064] Furthermore, by configuring the level check unit 31 and the use case determination unit 33 with a computer such as a microcontroller, and using the flowcharts shown in Figures 2 and 4 as computer programs, the system can be configured as an input / output program.

[0065] Furthermore, the present invention is not limited to the above embodiments. That is, those skilled in the art can implement the invention in various ways, without departing from the core principles, in accordance with conventionally known knowledge. As long as such modifications still possess the configuration of the input / output device of the present invention, they are of course included within the scope of the present invention. [Explanation of Symbols]

[0066] 2. Microphone (first sound collection means) 31 Level check unit (first output means, audio level comparison means, ambient sound level detection means) 32. Speech recognition engine unit (response acquisition means) 33. Use case determination unit (voice level comparison means, control means, ratio calculation means) 4. External output device (second output means) 41 Audio output unit (second output means, audio output means) 42 Display section (second output means, display means) 5. Microphone (second sound collection means) 6. External audio output means 10 Input / Output Devices S12 Level check (audio level comparison process) S13 Lower than the predetermined audio signal level (audio level comparison step) S14 Private mode (control process) S22 Level check (ratio calculation process) S23 Less than the specified value (ratio calculation process) S24 Private Mode (Control Process)

Claims

[Claim 1] A first output means that outputs the spoken input audio to a speech recognition means, A second output means that outputs a response from the aforementioned speech recognition means, A sound level comparison means for comparing the input sound level of the aforementioned input sound with a predetermined sound level, If the input audio level compared by the audio level comparison means is smaller than the predetermined audio level, the control means controls the second output means to suppress the response. An input / output device characterized by comprising the following features.

Citation Information

Patent Citations

  • Speech recognition device, method, and portable information terminal device using speech recognition method

    JP4299768B2