Voice response apparatus, voice response method, and storage medium
By detecting the volume of user and ambient sounds and generating an appropriate response voice volume, the problem of recognition accuracy of existing voice dialogue devices under different volume conditions is solved, and a highly efficient voice response effect is achieved.
Patent Information
- Application Number
- CN202111169732.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-04
- Filing Date
- 2021-10-08
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-10-08
AI Technical Summary
Existing voice interaction devices struggle to achieve accurate voice recognition when the input voice volume is too high or too low, and the response voice volume cannot be flexibly adjusted, affecting the user's communication experience.
The system detects the volume of the user's voice and ambient sounds using a microphone, generates response content using a processor, determines the output volume of the response voice based on the volume, and outputs the adjusted response voice through a speaker.
It enables precise control of the volume of the responding voice under different ambient volume conditions, improving the accuracy of voice recognition and the effectiveness of user communication.
Smart Images

Figure CN114724537B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to a voice response device, a voice response method, and a voice response program. Background Technology
[0002] AI speakers (smart speakers) and other voice-activated devices (voice response devices) take the user's voice as input and perform speech recognition on the content of the input voice. The voice-activated device then outputs a response, generated based on the speech recognition result relative to the input voice. Typically, voice-activated devices struggle to obtain accurate speech recognition results when the input voice volume is too high or too low. It is argued that voice-activated devices may be able to control the volume of the speaker's (user's) voice by controlling the volume of the output response voice. This is because the speaker can adjust the volume of their voice based on the volume of the person they are speaking to.
[0003] However, conventional voice interaction devices cannot flexibly change the volume of the response voice because the response volume is preset or user-specified. Furthermore, voice interaction devices use microphones to collect not only the speaker's voice but also other sounds. Therefore, voice interaction devices suffer from the following problem: even if the response voice volume can be easily set to correspond to the input voice volume, it is not easy to improve the accuracy of speech recognition. Summary of the Invention
[0004] To address the aforementioned issues, a speech response device, speech response method, and speech response program capable of achieving high-precision speech response are provided.
[0005] According to one embodiment, the voice response device includes a microphone, a processor, and a speaker. Sound is input to the microphone. The processor generates response content in the form of speech corresponding to a user's voice detected from the sound input by the microphone, and determines a volume for outputting the response content as response speech based on an input volume (the volume of the user's voice) and the volume of ambient sounds other than the user's voice. The speaker outputs the response speech at the volume determined by the processor.
[0006] According to an embodiment, a voice response method for a voice response device includes the following steps: acquiring sound input to a microphone, detecting a user's voice from the sound input to the microphone, generating response content corresponding to the user's voice detected from the sound input to the microphone, determining a volume for outputting the response content as speech based on the volume of the user's voice and the volume of ambient sounds other than the user's voice, and outputting the response speech of the response content from a speaker at the determined volume.
[0007] According to an embodiment, a storage medium is provided that stores a voice response program, the voice response program causing a computer to perform the following processes: acquiring sound input to a microphone, detecting a user's voice from the sound input to the microphone, generating response content corresponding to the user's voice detected from the sound input to the microphone, determining a volume for outputting the response content as speech based on the volume of the user's voice and the volume of ambient sounds other than the user's voice, and outputting the response speech of the response content from a speaker at the determined volume. Attached Figure Description
[0008] Figure 1 This is a diagram that schematically illustrates a structural example of the voice response device according to the embodiment.
[0009] Figure 2 This is a block diagram illustrating a structural example of the control system of the voice response device according to the embodiment.
[0010] Figure 3 This is a diagram illustrating an example of a function used in the implementation of a voice response device to determine the response volume based on the input volume when the ambient volume is below a threshold.
[0011] Figure 4 This is a diagram illustrating an example of a function that determines the response volume based on the input volume when the ambient volume is above a threshold, according to the voice response device involved in the implementation.
[0012] Figure 5 This is a diagram illustrating an example of a table used by the voice response device according to the embodiment to select a function corresponding to the ambient volume and the input volume.
[0013] Figure 6 This is a flowchart used to explain an example of the operation of the voice response device according to the embodiment.
[0014] Figure 7 This is a flowchart illustrating the calculation and processing of the response volume of the voice response device according to the embodiment.
[0015] Figure 8 This is a flowchart illustrating the calculation and processing of the response volume of the voice response device according to the embodiment.
[0016] Explanation of reference numerals in the attached figures
[0017] 1. Voice response device (voice dialogue device); 2. Microphone; 3. Speaker; 11. Processor; 12. Main storage device; 13. Auxiliary storage device; 14. Voice processing circuit. Detailed Implementation
[0018] The embodiments will now be described with reference to the accompanying drawings.
[0019] Figure 1 This is a diagram used to provide a general description of the voice response device 1 according to the embodiment.
[0020] like Figure 1 As shown, the voice response device 1 according to the embodiment includes a microphone 2 and a speaker 3. The voice response device 1 is a device that outputs a response voice from the speaker 3 corresponding to the voice of the speaker input to the microphone 2.
[0021] The voice response device 1 is, for example, a voice dialogue device referred to as an AI speaker. Alternatively, the voice response device 1 can also be an information processing device such as a smartphone, tablet, or personal computer. Furthermore, the voice response device 1 can also be a structure in which either or both of the microphone 2 and the speaker 3 are connected to the information processing device.
[0022] The voice response device 1 collects sound, including the speaker's voice and ambient sound, via the microphone 2. The voice response device 1 detects the speaker's voice (input speech) based on the sound collected by the microphone 2. The voice response device 1 recognizes the content of the input speech (the speaker's speech content) by performing speech recognition on the detected input speech. The voice response device 1 generates a response speech based on the recognized content of the input speech.
[0023] Furthermore, the voice response device 1 according to this embodiment measures (calculates) the volume of the speaker's voice (input voice) and the volume of other sounds (ambient sounds). The voice response device 1 maintains multiple functions (or tables) for determining the volume of the response voice. The multiple functions for determining the volume of the response voice are set according to a combination of the volume of the ambient sounds and the volume of the input voice. The voice response device 1 selects a function (or table) based on the volume of the input voice measured according to the sound collected by the microphone and the volume of the ambient sounds. The voice response device 1 determines the volume of the response voice corresponding to the volume of the input voice according to the selected function. The voice response device outputs the response content generated corresponding to the content of the input voice as the response voice with the volume determined according to the volume of the input voice and the volume of the ambient sounds from the speaker 3.
[0024] Next, the structure of the voice response device 1 according to the embodiment will be described.
[0025] Figure 2 This is a block diagram illustrating a structural example of the voice response device 1 according to the embodiment.
[0026] like Figure 2 As shown, the voice response device 1 includes a processor 11, a main storage device 12, an auxiliary storage device 13, a voice processing circuit 14, a microphone 2, and a speaker 3.
[0027] Processor 11 is responsible for the overall control of voice response device 1. Processor 11 is, for example, a CPU. Processor 11 performs various processes described later by executing programs. For example, processor 11 performs various processes such as motion control of voice response device 1, voice detection, voice recognition, generation of response sentences, measurement of the volume of input voice, measurement of the volume of ambient sound, calculation of the volume of response voice, and generation of response waveforms.
[0028] Main storage device 12 is the main memory for storing data. Main storage device 12 is composed of, for example, RAM (Random Memory). Main storage device 12 temporarily stores data being processed by processor 11. In addition, main storage device 12 can also store data required for program execution and program execution results. Furthermore, main storage device 12 also functions as a buffer memory for temporarily storing data.
[0029] For example, the main storage device 12 functions as a memory that stores information representing the volume of ambient sounds calculated based on the sounds collected by the microphone. For instance, the main storage device 12 stores speech data obtained by processing the sounds collected by the microphone 2 using the speech processing circuit 14. Furthermore, the main storage device 12 may also store the calculation result of the volume of the speaker's voice (input speech) contained in the sounds collected by the microphone. Additionally, the main storage device 12 may also store information representing the volume of the response speech determined based on the volume of the input speech and the volume of the ambient sounds.
[0030] Auxiliary storage device 13 is a storage device for storing data. Auxiliary storage device 13 includes non-volatile memory such as ROM (read-only memory) that cannot be rewritten, as well as rewritable non-volatile memory. As rewritable non-volatile memory, it is composed of, for example, HDD (hard disk drive), SSD (solid-state drive), EEPROM (trademarked), or flash ROM.
[0031] The auxiliary storage device 13 stores programs and control data executed by the processor 11. For example, the auxiliary storage device 13 stores a voice response program for outputting response voice corresponding to the input voice. The voice response program includes programs for various processes such as voice detection, voice recognition, intent parsing, generation of response statements, calculation of input volume, calculation of ambient volume, calculation of response volume, and generation of response waveforms, as described later. Furthermore, some or all of the processing performed by the processor 11 executing the program described later can also be executed by hardware such as processing circuitry.
[0032] In addition, Figure 2 In the example shown, auxiliary storage device 13 stores function table 13a for selecting functions that determine the volume of the response speech corresponding to the volume of the input speech, taking into account the volume of ambient sound (ambient volume). Function table 13a will be explained in detail later.
[0033] Microphone 2 collects (acquires) sound. For example, microphone 2 inputs the collected sound as an analog signal (analog waveform) and outputs the analog signal of the input sound to the speech processing circuit 14.
[0034] The voice processing circuit 14 takes the analog sound signal collected by the microphone 2 as input and outputs the analog sound signal as digital data. The voice processing circuit 14 includes an AD converter, etc., that digitizes analog waveforms.
[0035] Alternatively, microphone 2 can be an external device connected to voice response device 1. When microphone 2 is used as an external device, voice processing circuit 14 only needs to have an interface for voice input to connect microphone 2.
[0036] Speaker 3 outputs speech. Speaker 3 generates a response speech based on a response waveform supplied from processor 11. Speaker 3 controls the volume via processor 11. For example, speaker 3 emits a response speech based on a response waveform whose amplitude has been adjusted by processor 11 according to the volume of the response speech.
[0037] Furthermore, the speaker 3 can also be an external device connected to the voice response device 1. When the speaker 3 is used as an external device, the voice response device 1 only needs to have an interface that outputs a signal representing the waveform of the sound that should be output to the speaker 3.
[0038] Next, the function for determining the volume (response volume) of the response voice in the voice response device 1 according to the embodiment will be described.
[0039] The voice response device 1 recognizes the sounds emitted by a speaker and outputs a voice response to the speaker's speech (input statement). The voice response device 1 generates a response content to the speaker's sounds and determines the response volume using a function selected based on the volume of the input speech (input volume) and the volume of ambient sounds (ambient volume). That is, the voice response device 1 maintains multiple functions corresponding to the volume of ambient sounds as functions for determining the response volume based on the input volume. The voice response device 1 selects a function suitable for the volume of ambient sounds from among the multiple functions and determines the response volume based on the input volume.
[0040] Figure 3 as well as Figure 4 This is a diagram illustrating an example of a function (filter) used to determine the volume of the response speech (response volume) corresponding to the volume of the input speech (input volume) V.
[0041] Figure 3 An example of a function (the first function) for determining the response volume based on the input volume when the ambient sound volume (ambient volume) S is less than the threshold Ts (s < Ts) is shown. Additionally, Figure 4 An example of a function (the second function) for determining the response volume based on the input volume when the ambient volume S is above the threshold Ts (S≥Ts) is shown.
[0042] exist Figure 3In the example shown, function FA is used to determine the response volume based on the input volume when the ambient volume S is less than the threshold Ts (S < Ts). Function FA's characteristics change with respect to the thresholds Tva, Tvb, Tvc, and Tvd of the input volume V. Function FA consists of functions FAa, FAb, FAc, FAd, and FAe, which are divided into five intervals by the four thresholds Tva, Tvb, Tvc, and Tvd of the input volume V.
[0043] Function FAa is used to determine the response volume based on the input volume when the ambient volume S is less than the threshold Ts (S < Ts) and the input volume V is less than the threshold Tva (V < Tva). Function FAb is used to determine the response volume based on the input volume when the ambient volume S is less than the threshold Ts (S < Ts) and the input volume V is above the threshold Tva but less than the threshold Tvb (Tva ≤ V < Tvb).
[0044] The function FAc determines the response volume based on the input volume when the ambient volume S is less than the threshold Ts (S < Ts) and the input volume V is greater than or equal to the threshold Tvb but less than the threshold Tvc (Tvb ≤ V < Tvc). The function FAd determines the response volume based on the input volume when the ambient volume S is less than the threshold Ts (S < Ts) and the input volume V is greater than or equal to the threshold Tvc but less than the threshold Tvd (Tvc ≤ V < Tvd). The function FAe determines the response volume based on the input volume when the ambient volume S is less than the threshold Ts (S < Ts) and the input volume V is greater than or equal to the threshold Tvd (Tvd ≤ V).
[0045] exist Figure 4 In the example shown, the function FB is used to determine the response volume based on the input volume when the ambient volume S is above the threshold Ts (Ts≤S). The function FB's characteristics change with respect to three thresholds Tvi, Tvj, and Tvk of the input volume V. The function FB consists of functions FBa, FBb, FBc, and FBd, which are divided into four intervals by the three thresholds Tvi, Tvj, and Tvk of the input volume V.
[0046] The function FBa is used to determine the response volume based on the input volume when the ambient volume S is above the threshold Ts (Ts≤S) and the input volume V is below the threshold Tvi (V<Tvi). The function FBb is used to determine the response volume based on the input volume when the ambient volume S is above the threshold Ts (Ts≤S) and the input volume V is above the threshold Tvi but below the threshold Tvj (Tvi≤V<Tvj).
[0047] The function FBc is used to determine the response volume based on the input volume when the ambient volume S is above the threshold Ts (Ts≤S) and the input volume V is above the threshold Tvj but below the threshold Tvk (Tvj≤V<Tvk). The function FBd is used to determine the response volume based on the input volume when the ambient volume S is above the threshold Ts (Ts≤S) and the input volume V is above the threshold Tvk (Tvk≤V).
[0048] Figure 5 This is a diagram illustrating a structural example of a function table 13a used in the implementation of the voice response device 1 to select a size suitable for ambient volume and input volume.
[0049] Figure 5 The function table 13a shown represents the function based on the ambient volume and the input volume. Figure 3 as well as Figure 4 The functions selected from the functions shown. For example, Figure 2 As shown, Figure 5 The function table 13a shown is stored in the auxiliary storage device 13 of the voice response device 1. The voice response device 1 selects a function corresponding to the ambient volume S and the input volume V by referring to the function table 13a. The voice response device 1 uses the function selected based on the ambient volume S and the input volume V and determines the response volume based on the input volume.
[0050] For example, when S < Ts and V < Tva, voice response device 1 uses function FAa to determine the response volume based on the input volume. When S < Ts and Tva ≤ V < Tvb, voice response device 1 uses function FAb and determines the response volume based on the input volume. When S < Ts and Tvb ≤ V < Tvc, voice response device 1 uses function FAc and determines the response volume based on the input volume. When S < Ts and Tvc ≤ V < Tvd, voice response device 1 uses function FAd and determines the response volume based on the input volume. When S < Ts and Tvd ≤ V, voice response device 1 uses function FAe and determines the response volume based on the input volume.
[0051] Furthermore, when Ts ≤ S and V < Tvi, voice response device 1 uses function FBa and determines the response volume based on the input volume. When Ts ≤ S and Tvi ≤ V < Tvj, voice response device 1 uses function FBb and determines the response volume based on the input volume. When Ts ≤ S and Tvj ≤ V < Tvk, voice response device 1 uses function FBc and determines the response volume based on the input volume. When Ts ≤ S and Tvk ≤ V, voice response device 1 uses function FBd and determines the response volume based on the input volume.
[0052] Next, the operation of the voice response device 1 according to the embodiment will be explained.
[0053] Figure 6 This is a flowchart illustrating an example of the operation of the voice response device 1 according to the embodiment in processing the voice output response speech in response to a speaker (user).
[0054] The processor 11 of the voice response device 1 inputs the sound collected by the microphone 2 as input sound data (ACT11). The microphone 2 supplies a signal representing the analog waveform of the collected sound to the voice processing circuit 14. The voice processing circuit 14 digitizes the signal representing the analog waveform input from the microphone 2. The voice processing circuit 14 supplies the digitized digital signal as sound data to the processor 11. The processor 11 acquires the input sound data obtained by digitizing the sound collected by the microphone 2 through the voice processing circuit 14.
[0055] If the input sound data is acquired, the processor 11 uses speech detection processing to detect whether the input sound data includes the speaker's voice (ACT12). The processor 11 performs speech detection processing to detect whether the input sound includes the speaker's voice by executing a speech detection program.
[0056] If no speaker's voice is detected from the input sound (ACT12, No), the processor 11 calculates (measures) the volume of the ambient sound (ambient volume) based on the sound data of the input sound (ACT13). If no speaker's voice is detected from the input sound, the input sound is an ambient sound that does not include the speaker's voice (sounds other than the speaker's voice). If the input sound is an ambient sound, the processor 11 calculates the volume based on the sound data of the input sound. If the input sound is an ambient sound, the processor 11 stores the calculated volume of the input sound as the ambient volume S in the main storage device 12 or the auxiliary storage device 13 (ACT14).
[0057] In this embodiment, the processor 11 stores the ambient volume S, calculated from the input sound (ambient sound) excluding the speaker's voice during a period to infer the ambient volume when the speaker is speaking. Therefore, the processor 11 can also store the calculated ambient volume S over an already stored ambient volume (past ambient volume). Additionally, the processor 11 can store the ambient volume S for a predetermined period from the current point in time. Furthermore, the processor 11 can also store the average ambient volume calculated over the predetermined period from the current point in time as the ambient volume S.
[0058] If the speaker's voice is detected in the input sound (ACT12, yes), the processor 11 performs the processing of generating response content (response statement) (ACT15-17) and the processing of calculating the response volume (ACT18-19).
[0059] Processor 11 performs speech recognition processing, content parsing processing, and response sentence generation as processing to generate response content. That is, processor 11 performs speech recognition (ACT15) to recognize the speaker's voice (input speech) contained in the input sound. Processor 11 extracts the speaker's voice from the input sound and recognizes the language (input sentence) spoken by the speaker from the extracted speaker's voice. For example, processor 11 recognizes the language spoken by the speaker by referring to the pronunciation of pre-set language (words).
[0060] If the processor 11 receives an input statement as a speech recognition result of the speaker's voice, it performs intent parsing processing (ACT16) to parse the meaning of the input statement obtained as a speech recognition result. As intent parsing processing, the processor 11 parses the meaning of the input statement (the user's intent contained in the input statement) based on the recognition results of the words contained in the input statement.
[0061] For example, processor 11 determines whether the input statement is a question, a request or request, or a greeting. If the processor 11 determines that the input statement is a question, it determines the content of the question contained in the input statement. If the processor 11 determines that the input statement is a request, it determines the content of the request contained in the input statement. If the processor 11 determines that the input statement is a greeting, it determines the content of the greeting contained in the input statement.
[0062] If the processor 11 parses the meaning of the speaker's voice (input statement), it generates a response content (response statement) relative to the input statement (ACT17). For example, if the processor 11 determines that the input statement contains a question, it generates a response statement corresponding to the question. Furthermore, if the processor 11 determines that the input statement contains a request from the speaker, it generates a response statement based on the speaker's request. Additionally, if the processor 11 determines that the input statement contains a greeting (understood as a greeting from the speaker), it generates a response statement that corresponds to the greeting from the speaker.
[0063] On the other hand, the processor 11 performs calculation processing of input volume V and calculation processing of response volume as a process for calculating response volume. The processor 11 calculates the volume V of the speaker's voice (input speech) detected in the input sound (ACT18). For example, the processor 11 extracts the components of the speaker's voice (input speech) from the sound data of the input sound and calculates the volume (input volume) V of the extracted input speech.
[0064] If the input volume V is calculated, the processor 11 performs a process (ACT19) to calculate the response volume based on the calculated input volume V and the ambient volume S. The processor 11 calculates the response volume relative to the input volume based on a function selected according to the input volume V and the ambient volume S. The process for calculating the response volume (response volume calculation process) will be explained in detail later.
[0065] The processor 11 generates a response waveform (ACT20) to be emitted by the speaker 3, based on the response statement generated in ACT17 and the response volume calculated in ACT19. For example, the processor 11 generates a response waveform for emitting the response statement generated in ACT17 as the response voice. The processor 11 adjusts the amplitude of the response waveform used to emit the generated response voice according to the response volume calculated in ACT19. If a response waveform is generated, the processor 11 outputs the generated response waveform from the speaker 3 (ACT21).
[0066] Next, the calculation and processing of the response volume of the voice response device 1 according to the embodiment will be described in detail.
[0067] Figure 7 as well as Figure 8 This is a flowchart illustrating the calculation and processing of the response volume of the voice response device 1 according to the embodiment.
[0068] In the response volume calculation process, the processor 11 obtains the current input volume V calculated in ACT18 above (ACT31). In addition, the processor 11 obtains the ambient volume S stored in the main storage device 12 or the auxiliary storage device 13 (ACT32).
[0069] If the input volume V and ambient volume S are obtained, the processor 11 refers to... Figure 5 From the function table shown, select the function corresponding to the input volume V and ambient volume S. Figure 7 as well as Figure 8 In the processing example shown, processor 11 according to Figure 5 Use the function table 13a shown to select a function.
[0070] Furthermore, the function used to determine the response volume based on the input volume while taking ambient volume into account is not limited to Figure 3 as well as Figure 4 The structure shown can also be appropriately set according to the application. Furthermore, the thresholds for ambient volume and input volume are not limited to... Figure 3 , Figure 4 as well as Figure 5 The content shown can also be appropriately set according to the function.
[0071] exist Figure 7 as well as Figure 8 In the processing example shown, processor 11 refers to Figure 5 The table shown indicates whether the ambient volume S is below the threshold Ts (ACT33).
[0072] When the ambient volume S is less than the threshold Ts (S < Ts) (ACT33, yes), processor 11 applies the function FA for the case where the ambient volume S is small. Figure 3 In the example shown, function FA is composed of five functions FAa, FAb, FAc, FAd, and FAe, which are partitioned by thresholds Tva, Tvb, Tvc, and Tvd. Processor 11 is based on Figure 5 The table shown compares the input volume V with thresholds Tva, Tvb, Tvc, and Tvd, and selects a function from FAa, FAb, FAc, FAd, and FAe.
[0073] That is, if S < Ts (ACT33, Yes), processor 11 determines whether the input volume V is insufficient to the threshold Tva (ACT41). If it is determined that the input volume V is insufficient to the threshold Tva (ACT41, Yes), processor 11 determines that the ambient volume S < threshold Ts and the input volume V < threshold Tva. If S < Ts and V < Tva, processor 11 selects function FAa (ACT42).
[0074] If the input volume V is determined to be less than the threshold Tva (ACT41, No), the processor 11 determines whether the input volume V is less than the threshold Tvb (ACT43). If the input volume V is determined to be less than the threshold Tvb (ACT43, Yes), the processor 11 determines that the ambient volume S < threshold Ts and the threshold Tva ≤ input volume V < threshold Tvb. If S < Ts and Tva ≤ V < Tvb, the processor 11 selects the function FAb (ACT44).
[0075] If the input volume V is determined to be less than the threshold Tvb (ACT43, No), the processor 11 determines whether the input volume V is less than the threshold Tvc (ACT45). If the input volume V is determined to be less than the threshold Tvc (ACT45, Yes), the processor 11 determines that the ambient volume S < threshold Ts and the threshold Tvb ≤ input volume V < threshold Tvc. If S < Ts and Tvb ≤ V < Tvc, the processor 11 selects the function FAc (ACT44).
[0076] If the input volume V is determined to be less than the threshold Tvc (ACT45, No), the processor 11 determines whether the input volume V is less than the threshold Tvd (ACT47). If the input volume V is determined to be less than the threshold Tvd (ACT47, Yes), the processor 11 determines that the ambient volume S < threshold Ts and the threshold Tvc ≤ input volume V < threshold Tvd. If S < Ts and Tvc ≤ V < Tvd, the processor 11 selects the function FAd (ACT48).
[0077] If the input volume V is determined to be below the threshold Tvd (ACT47, No), the processor 11 determines that the ambient volume S < threshold Ts and the threshold Tvd ≤ input volume V because the input volume V is above the threshold Tvd. If S < Ts and Tvd ≤ V, the processor 11 selects the function FAe (ACT49).
[0078] On the other hand, if the ambient volume S is not below the threshold Ts, in other words, if the ambient volume S is above the threshold Ts (ACT33, No), the processor 11 applies the function FB for the case where the ambient volume S is large. Figure 4 In the example shown, the function FB consists of four functions FBa, FBb, FBc, and FBd, partitioned by thresholds Tvi, Tvj, and Tvk of the input volume V. Processor 11 is based on... Figure 5 The function table 13a shown compares the input volume V with the thresholds Tvi, Tvj, and Tvk, and selects a function from FBa, FBb, FBc, and FBd.
[0079] That is, if S < Ts (ACT33, No), processor 11 determines whether the input volume V is insufficient to the threshold Tvi (ACT51). If the input volume V is insufficient to the threshold Tvi (ACT51, Yes), processor 11 determines that the ambient volume S ≥ threshold Ts and the input volume V < threshold Tvi. If S ≥ Ts and V < Tvi, processor 11 selects function FBa (ACT52).
[0080] If the input volume V is determined to be less than the threshold Tvi (ACT51, No), the processor 11 determines whether the input volume V is less than the threshold Tvj (ACT53). If the input volume V is determined to be less than the threshold Tvj (ACT53, Yes), the processor 11 determines that the ambient volume S ≥ the threshold Ts and the threshold Tvi ≤ the input volume V < the threshold Tvj. If S ≥ Ts and Tvi ≤ V < Tvj, the processor 11 selects the function FBb (ACT54).
[0081] If the input volume V is determined to be less than the threshold Tvj (ACT53, No), the processor 11 determines whether the input volume V is less than the threshold Tvk (ACT55). If the input volume V is determined to be less than the threshold Tvk (ACT55, Yes), the processor 11 determines that the ambient volume S ≥ the threshold Ts and the threshold Tvj ≤ the input volume V < the threshold Tvk. If S < Ts and Tvj ≤ V < Tvk, the processor 11 selects the function FBc (ACT56).
[0082] If the input volume V is determined to be below the threshold Tvk (ACT55, No), the processor 11 determines that the ambient volume S ≥ the threshold Ts and the threshold Tvk ≤ the input volume V because the input volume V is above the threshold Tvk. If S ≥ Ts and Tvk ≤ V, the processor 11 selects the function FBd (ACT57).
[0083] If a function corresponding to the ambient volume S and the input volume V is selected, the processor 11 determines the response speech based on the selected function (ACT60). That is, the processor 11 uses the selected function to calculate the response volume corresponding to the input volume V. Thus, the processor 11 is able to calculate the response volume corresponding to the input volume while taking the ambient volume into account.
[0084] As described above, the voice response device according to the embodiment detects the user's voice in the sound input to the microphone. The voice response device generates response content (response statement) as a response voice to the user's voice. Furthermore, the voice response device calculates a response volume based on the input volume, which is the volume of the user's voice, and the volume of ambient sounds other than the user's voice. The voice response device outputs the response voice from a speaker at the calculated response volume.
[0085] That is, the voice response device according to the embodiment can take into account the volume of ambient sound and output a response voice with a response volume corresponding to the input volume. Therefore, it is expected that the volume of the speaker's (user's) voice can be controlled in accordance with the volume of the response voice output by the voice response device. The voice response device can guide the volume of the user's voice to a volume suitable for voice recognition, enabling high-accuracy voice recognition.
[0086] Furthermore, the voice response device according to the embodiment maintains multiple functions selected based on the ambient volume. When the ambient volume is below a threshold, the voice response device determines the volume of the response voice based on a first function and the input volume. When the ambient volume is above the threshold, the voice response device determines the volume of the response voice based on a second function different from the first function and the input volume. Thus, the voice response device according to the embodiment can set a response volume corresponding to the ambient sound level. As a result, even in environments where the ambient volume cannot be predicted in advance, the voice response device can guide the volume of the user's voice to a level suitable for speech recognition.
[0087] Furthermore, the voice response device according to the embodiment stores multiple functions selected based on the ambient volume and the input volume in a storage device. The voice response device determines the volume of the responding voice based on one of the multiple functions selected according to the ambient volume and the input volume, and according to the input volume. Therefore, the voice response device can select a function based on the ambient volume and the input volume, and can easily guide the volume of the user's voice to a level suitable for voice recognition.
[0088] Furthermore, in the above embodiment, the case where a program for processor execution is pre-stored in the device's memory has been described. However, the program executed by the processor can also be downloaded to the device from a network or installed on the device from a storage medium. The storage medium can be a storage medium capable of storing programs and readable by the device, such as a CD-ROM. Additionally, functions obtained through pre-installation or download can be implemented in conjunction with the device's internal OS (operating system).
[0089] While several embodiments have been described, these embodiments are merely illustrative and not intended to limit the scope of the invention. These embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included within the scope and spirit of the invention, and likewise within the scope of the invention as described in the claims and its equivalents.
Claims
1. A voice response device, characterized in that, have: Microphone, for inputting sound; The processor generates response content in speech form corresponding to the user's voice detected from the sound input by the microphone, and determines the volume for outputting the response content as response speech based on the input volume of the user's voice and the volume of ambient sounds other than the user's voice. as well as The speaker outputs the response voice at a volume determined by the processor. The processor determines the volume of the response speech based on a first function and the input volume when the volume of the ambient sound is below a threshold, and determines the volume of the response speech based on a second function different from the first function and the input volume when the volume of the ambient sound is above the threshold.
2. The voice response device according to claim 1, characterized in that, It also includes an auxiliary storage device that stores multiple functions corresponding to the volume of ambient sound and the volume of the input sound. The processor determines the volume of the response speech based on a function selected from a plurality of functions stored in the auxiliary storage device according to the volume of the ambient sound and the input volume, and based on the input volume.
3. The voice response device according to claim 1 or 2, characterized in that, It has a memory that stores the volume of the sound input from the microphone as the volume of the ambient sound when no user's voice is detected. When a user's voice is detected from the sound input from the microphone, the processor calculates an input volume as the volume of the user's voice, and determines the volume of the response voice based on the input volume and the volume of the ambient sound stored in the memory.
4. The voice response device according to claim 3, characterized in that, The memory stores the average volume of the ambient sound over a specified period of time as the volume of the sound input from the microphone when no user's voice is detected.
5. The voice response device according to claim 1, characterized in that, Both the first function and the second function are composed of multiple functions divided by multiple thresholds related to the input volume.
6. A voice response method for a voice response device, characterized in that, The following steps are involved: Get the sound input to the microphone. The system detects the user's voice from the sound input to the microphone. Generate response content corresponding to the user's voice detected from the input to the microphone. If the volume of ambient sounds other than the user's voice is below a threshold, the volume for outputting the response content in speech is determined based on a first function and the input volume, which is the volume of the user's voice. If the volume of ambient sounds is above the threshold, the volume for outputting the response content in speech is determined based on a second function different from the first function and the input volume. The response speech of the response content is output from the speaker at a determined volume.
7. A storage medium storing a voice response program, characterized in that, The voice response program causes the computer to perform the following processes: Get the sound input to the microphone. The system detects the user's voice from the sound input to the microphone. Generate response content corresponding to the user's voice detected from the input to the microphone. If the volume of ambient sounds other than the user's voice is below a threshold, the volume for outputting the response content in speech is determined based on a first function and the input volume, which is the volume of the user's voice. If the volume of ambient sounds is above the threshold, the volume for outputting the response content in speech is determined based on a second function different from the first function and the input volume. The response speech of the response content is output from the speaker at a determined volume.
8. The storage medium according to claim 7, characterized in that, The voice response program causes the computer to perform the following processes: If the volume of the ambient sound is below a threshold, the volume of the response speech is determined based on a first function and the input volume. If the volume of the ambient sound is above the threshold, the volume of the response speech is determined based on a second function different from the first function and the input volume.
Citation Information
Patent Citations
Information processing apparatus, information processing system, and information processing method, and program
US20200388268A1