Speech conversation system
Patent Information
- Application Number
- PCT/JP2026/006518
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-02-24
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026006518_01102026_PF_FP_ABST
Abstract
Description
Voice conversation system
[0001] One aspect of the present disclosure relates to a voice conversation system.
[0002] Patent Document 1 below shows an example configuration of a voice conversation system that enables a plurality of users to hold conversations with each other via voice.
[0003] Japanese Unexamined Patent Publication No. 2016-4066
[0004] In a voice conversation system, it is desired to determine whether the content indicated by the voice uttered by a speaker is conveyed to a listener.
[0005] A voice conversation system according to one aspect of the present disclosure includes: an acoustic signal acquisition unit that acquires an acoustic signal derived from a voice uttered by a speaker; a signal processing unit that derives acoustic-related information related to the acoustic signal by processing the acoustic signal; and a determination unit that determines, based on the acoustic-related information, whether the content indicated by the voice uttered by the speaker is conveyed to the listener.
[0006] According to one aspect of the present disclosure, in a voice conversation system, it is possible to determine whether the content indicated by the voice uttered by a speaker is conveyed to a listener.
[0007] An example configuration of the voice conversation system according to Embodiment 1 is shown. The figure is for explaining the influence of noise added to an acoustic signal derived from a voice uttered by a speaker. An example configuration of a voice conversation system as a modified example of Embodiment 1 is shown. An example configuration of the voice conversation system according to Embodiment 2 is shown. An example configuration of the voice conversation system according to Embodiment 3 is shown. An example configuration of the voice conversation system according to Embodiment 4 is shown.
[0008] [Embodiment 1] Embodiment 1 will be described below. For convenience of explanation, components having the same functions as the components described in Embodiment 1 are denoted by the same reference numerals in the following embodiments, and the description thereof will not be repeated. For the sake of brevity, descriptions of known technical matters are also appropriately omitted.
[0009] In this specification, each component and each numerical value is merely illustrative, unless otherwise consistent with the content. Therefore, unless otherwise consistent with the content, the positional and connection relationships of each component are not limited to the examples shown in the figures.
[0010] (Example of the configuration of the voice conversation system 1) Figure 1 shows an example of the configuration of the voice conversation system 1 in Embodiment 1. The voice conversation system 1 is a voice conversation system that enables multiple users to converse with each other by voice. In Embodiment 1, for the sake of clarity, we will illustrate a case in which two users are using the voice conversation system 1, and one of the two users is the speaker and the other is the listener.
[0011] The voice conversation system 1 in the example shown in Figure 1 comprises a speaker-side terminal device 10 and a listener-side terminal device 90. The speaker-side terminal device 10 is a terminal device used by the speaker in the voice conversation system 1. On the other hand, the listener-side terminal device 90 is a terminal device used by the listener in the voice conversation system 1.
[0012] The speaker-side terminal device 10 and the listener-side terminal device 90 may be portable terminal devices or stationary terminal devices. The speaker-side terminal device 10 and the listener-side terminal device 90 may be connected to each other by any communication interface, as long as it is possible to allow the listener to hear the voice emitted by the speaker when using the voice conversation system 1.
[0013] In the example shown in Figure 1, the voice conversation system 1 comprises an acoustic signal acquisition unit 11, a signal processing unit 12, a determination unit 13, a notification unit 14, and an acoustic signal playback unit 91. In the example shown in Figure 1, the speaker-side terminal device 10 comprises an acoustic signal acquisition unit 11, a signal processing unit 12, a determination unit 13, and a notification unit 14, while the receiver-side terminal device 90 comprises an acoustic signal playback unit 91.
[0014] The acoustic signal acquisition unit 11 acquires an acoustic signal originating from the speech emitted by the speaker. The acoustic signal acquisition unit 11 may be any type of acoustic input device (e.g., a microphone). Figure 1 illustrates a case where the speaker-side terminal device 10 has the acoustic signal acquisition unit 11. However, unlike the example in Figure 1, the acoustic signal acquisition unit 11 may be located outside the speaker-side terminal device 10. In this case, the acoustic signal acquisition unit 11 only needs to be connected to the speaker-side terminal device 10 in a communicative manner.
[0015] In the voice conversation system 1, the acoustic signal reproduction unit 91 acquires an acoustic signal from the acoustic signal acquisition unit 11. The acoustic signal reproduction unit 91 reproduces the acoustic signal. This allows the listener to hear the sound corresponding to the voice emitted by the speaker. The acoustic signal reproduction unit 91 may be any type of acoustic output device (e.g., a loudspeaker).
[0016] Figure 1 illustrates a case where the receiver terminal device 90 has an acoustic signal reproduction unit 91. However, unlike the example in Figure 1, the acoustic signal reproduction unit 91 may be located outside the receiver terminal device 90. In this case, the acoustic signal reproduction unit 91 only needs to be connected to the receiver terminal device 90 in a communicative manner.
[0017] The signal processing unit 12 acquires an acoustic signal from the acoustic signal acquisition unit 11. The signal processing unit 12 derives information related to the acoustic signal by processing the acoustic signal. In this specification, this information is referred to as acoustic-related information. Various specific examples of acoustic-related information will be described in Embodiment 2 and subsequent embodiments.
[0018] The determination unit 13 acquires acoustic information from the signal processing unit 12. Based on the acoustic information, the determination unit 13 determines whether or not the content indicated by the voice emitted by the speaker has been conveyed to the receiver. Specific examples of methods for determining whether or not the content has been conveyed to the receiver will be described in Embodiment 2 and subsequent embodiments.
[0019] The notification unit 14 executes a notification based on the determination result of the determination unit 13. Specific examples of notifications by the notification unit 14 will be described later. Figure 1 illustrates a case where the speaker-side terminal device 10 has a notification unit 14. However, unlike the example in Figure 1, the notification unit 14 may be located outside the speaker-side terminal device 10. In this case, the notification unit 14 only needs to be connected to the speaker-side terminal device 10 in a communicative manner.
[0020] (Problems in conventional voice conversation systems) Figure 2 is a diagram illustrating the effect of noise added to the acoustic signal originating from the speech emitted by the speaker. In the reference embodiment, the problems in conventional voice conversation systems will be described with reference to Figure 2. Unlike the voice conversation system 1 of Embodiment 1, the conventional voice system does not have a signal processing unit 12, a determination unit 13, and a notification unit 14.
[0021] In the example in Figure 2, the code SP represents the speaker, and the code RP represents the listener. In the example in Figure 2, the code AS represents the acoustic signal acquired by the acoustic signal acquisition unit 11, which originates from the speech emitted by the speaker SP. In the example in Figure 2, the acoustic signal AS corresponds to the speech "Can you hear me?" emitted by the speaker SP. The code NZ represents the noise added to the acoustic signal AS.
[0022] As an example, the acoustic signal acquisition unit 11 may be a wearable device (e.g., a shoulder-worn or earphone-type audio input / output device). A typical example of noise NZ in this case is noise generated when a part of the speaker SP's body comes into contact with the housing of the acoustic signal acquisition unit 11. For example, noise NZ as a rubbing sound is generated when the speaker SP's hair comes into contact with the housing. As another example, noise NZ as a rubbing sound is also generated when the clothing worn by the speaker SP comes into contact with the housing.
[0023] In the example in Figure 2, reference numeral 210 indicates the case where the acoustic signal AS is not accompanied by noise NZ (i.e., an ideal example). In the example of reference numeral 210, the series of sounds in Japanese, "Kikoemasu ka" (Can you hear me?), represented by the acoustic signal AS, can be heard without any problems by the receiver RP. Note that "Kikoemasu ka" in Japanese means "Can you hear me?". In the following explanation, unless otherwise specified, it is assumed that the speaker SP is speaking Japanese.
[0024] As described above, in the example of code 210, the series of voices "Can you hear me?" represented by the acoustic signal AS are heard without any problems by the receiver RP. Therefore, in the example of code 210, the receiver RP can understand the content indicated by the voice emitted from the speaker SP. In other words, in the example of code 210, the content indicated by the voice emitted from the speaker SP is transmitted to the receiver RP without any problems.
[0025] In the example in Figure 2, reference numeral 220 indicates the case where a small amount of noise NZ is added to the acoustic signal AS. In the example of reference numeral 220, the part of the series of sounds "Can you hear me?" represented by the acoustic signal AS that corresponds to one syllable "su" is obscured by the noise NZ.
[0026] Therefore, in the example of symbol 220, the remaining part of the sequence of sounds "Kikoemasu ka" (Can you hear me?), excluding the part corresponding to the single syllable "su", is heard by the receiver RP. In other words, in the example of symbol 220, the sound "Kikoemaka" is heard by the receiver RP. In the example of Figure 2, one full-width space in the sound heard by the receiver RP represents one syllable that is not heard by the receiver RP.
[0027] In the example of symbol 220, not all of the series of sounds "Can you hear me?" uttered by the speaker SP are heard by the receiver RP. However, the majority of the series of sounds, "Can you hear me?", are heard by the receiver RP.
[0028] Therefore, it is thought that the receiver RP can infer that the content indicated by the series of sounds uttered by the speaker SP is "Can you hear me?" by supplementing the single syllable "su" that the receiver RP could not hear.
[0029] Therefore, even in the example of code 220, the listener RP may be able to grasp the content indicated by the voice emitted from the speaker SP. Thus, even in the example of code 220, the content indicated by the voice emitted from the speaker SP may be transmitted to the listener RP. In this way, even if noise NZ is added to the acoustic signal AS, if the noise NZ is minor, the content indicated by the voice emitted from the speaker SP may be transmitted to the listener RP.
[0030] In the example in Figure 2, reference numeral 230 indicates a case where a large amount of noise NZ is added to the acoustic signal AS. The example of reference numeral 230 is a counterpart to the example of reference numeral 220. In the example of reference numeral 230, the parts of the series of sounds "Kikoemasu ka" represented by the acoustic signal AS that correspond to the four syllables "ko," "e," "ma," and "su" are obscured by the noise NZ.
[0031] Therefore, in the example of symbol 220, the remaining part of the sequence of sounds "Kikoemasu ka" (Can you hear me?), excluding the parts corresponding to the four syllables "ko," "e," "ma," and "su," will be heard by the receiver RP. In other words, in the example of symbol 230, the sound "kika" will be heard by the receiver RP.
[0032] Thus, in the example of symbol 230, the majority of the series of sounds "Can you hear me?" uttered by the speaker SP is not heard by the receiver RP. Therefore, in the example of symbol 230, it is difficult for the receiver RP to infer that the content indicated by the series of sounds uttered by the speaker SP is "Can you hear me?" by supplementing the four syllables "ko," "e," "ma," and "su" that the receiver RP could not hear.
[0033] From this, in the example of reference numeral 230, it is difficult for the listener RP to grasp the content indicated by the voice emitted from the speaker SP. Therefore, in the example of reference numeral 230, it can be said that the voice emitted from the speaker SP is not transmitted to the listener RP. Thus, when there is a large amount of noise NZ added to the acoustic signal AS, the possibility of the voice emitted from the speaker SP being transmitted to the listener RP is low.
[0034] As can be seen from the example in Figure 2, the likelihood that the content indicated by the voice emitted from the speaker SP will be transmitted to the listener RP depends on the degree of noise NZ added to the acoustic signal AS. Thus, the likelihood that the content indicated by the voice emitted from the speaker SP will be transmitted to the listener RP depends on the surrounding environment when the voice conversation system is in use.
[0035] However, conventional voice conversation systems do not incorporate a mechanism to inform the speaker (SP) whether or not the content indicated by the voice they emit has been conveyed to the receiver (RP).
[0036] Therefore, for example, if the speaker SP is unaware that the aforementioned friction noise NZ is occurring, even if the speaker SP utters the sound appropriately, there is a risk that the content indicated by the sound will not be conveyed to the receiver RP. In addition, since the speaker SP may not be aware that the content indicated by the sound they uttered is not being conveyed to the receiver RP, there is concern that this situation may persist without being resolved.
[0037] (Effects of Embodiment 1) As described above, conventional voice conversation systems have the problem that they cannot determine whether or not the content indicated by the voice emitted by the speaker is being conveyed to the receiver. Based on this problem in conventional voice conversation systems, the inventors of the present invention have created a novel voice conversation system, voice conversation system 1. As described above, unlike conventional voice conversation systems, voice conversation system 1 is equipped with a signal processing unit 12 and a determination unit 13.
[0038] As described above, the determination unit 13 can determine whether or not the content indicated by the voice emitted by the speaker is being conveyed to the receiver, based on the acoustic information acquired from the signal processing unit 12. Thus, Embodiment 1 provides a voice conversation system (e.g., voice conversation system 1) that can determine whether or not the content indicated by the voice emitted by the speaker is being conveyed to the receiver.
[0039] Furthermore, as described above, unlike conventional voice conversation systems, the voice conversation system 1 also includes a notification unit 14 that executes notifications based on the determination result of the determination unit 13. For example, if the determination unit 13 outputs a determination result indicating that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver, the notification unit 14 may notify the speaker that the content has not been conveyed to the receiver.
[0040] The manner in which the notification unit 14 provides notification is not particularly limited. For example, the notification unit 14 may provide notification through an auditory notification (e.g., outputting an alarm sound). As another example, the notification unit 14 may provide notification through a visual notification (e.g., flashing of a red lamp, or a warning message popping up on the display screen of the speaker-side terminal device 10). As yet another example, the notification unit 14 may provide notification through a tactile notification (e.g., vibration of a vibrator). The notification unit 14 may also provide notification in a multimodal manner, combining at least two of the auditory, visual, and tactile notification methods.
[0041] By notifying the speaker from the notification unit 14 that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver, the speaker can be prompted to improve the situation in which the content has not been conveyed to the receiver.
[0042] It should be noted that, as described above, noise as rubbing sound can be a major factor that impedes transmission of the content represented by the speech uttered by a speaker to a listener. Therefore, as an example, when the determination unit 13 outputs a determination result indicating that the content represented by the speech uttered by the speaker is not transmitted to the listener, the notification unit 14 may notify the speaker that the content is not transmitted to the listener, and may also notify the speaker of a message prompting removal of the rubbing sound.
[0043] An example of a message prompting removal of rubbing sound includes the message: "The microphone may be in contact with some object. If it is in contact, please move the microphone away from the contacting object." The message prompting removal of rubbing sound may be a text message or a voice message.
[0044] It should be noted that the notification unit 14 may notify the listener that the speech uttered by the speaker has not been transmitted to the listener. In this case, for example, the listener who has received the notification can use the text chat tool of the voice conversation system 1 to inform the speaker that the content represented by the speech uttered by the speaker is not transmitted to the listener, thereby prompting the speaker to improve the situation where the content is not transmitted to the listener.
[0045] As described above, in the example of Embodiment 1, when the determination unit 13 outputs a determination result indicating that the content represented by the speech uttered by the speaker is not transmitted to the listener, the notification unit 14 only needs to notify at least one of the speaker and the listener that the content is not transmitted to the listener.
[0046] [Modified Example] FIG. 1 illustrates a case where the speaker-side terminal device 10 includes the signal processing unit 12 and the determination unit 13. That is, a case where the determination unit 13 is mounted in the same device as the signal processing unit 12 is illustrated. However, unlike the example in FIG. 1, the determination unit 13 does not necessarily have to be mounted in the same device as the signal processing unit 12.
[0047] FIG. 3 shows a configuration example of a voice conversation system 1V as a modification of the first embodiment. The voice conversation system 1V includes a speaker-side terminal device 10V instead of the speaker-side terminal device 10. The voice conversation system 1V, unlike the voice conversation system 1, includes a server SV.
[0048] The speaker-side terminal device 10V in the example of FIG. 3 includes an acoustic signal acquisition unit 11, a signal processing unit 12, and a notification unit 14. That is, unlike the speaker-side terminal device 10, the speaker-side terminal device 10V does not include a determination unit 13. Instead, in the speaker-side terminal device 10V, the server SV includes the determination unit 13.
[0049] The server SV in the example of FIG. 3 only needs to be communicably connected to the speaker-side terminal device 10V and the listener-side terminal device 90. As described above, the determination unit in the voice conversation system according to one aspect of the present disclosure may be located outside the speaker-side terminal device.
[0050] As described in each embodiment below, the determination unit according to one aspect of the present disclosure may include a deep learning model. Therefore, when the deep learning model is a large-scale model, depending on the computational resources of the speaker-side terminal device, it may not necessarily be easy to provide the determination unit in the speaker-side terminal device. In such a case, the determination unit may be provided in a server having sufficiently larger computational resources than the speaker-side terminal device.
[0051] Thereby, even when the computational resources of the speaker-side terminal device are not necessarily large, it is possible to determine whether the voice uttered by the speaker is transmitted to the listener using the determination unit as a large-scale model. Therefore, even when the computational resources of the speaker-side terminal device are not necessarily large, a highly reliable determination result can be obtained.
[0052] [Second Embodiment] FIG. 4 shows a configuration example of a voice conversation system 2 according to the second embodiment. The voice conversation system 2 includes a speaker-side terminal device 20 instead of the speaker-side terminal device 10. The speaker-side terminal device 20 includes a signal processing unit 22 instead of the signal processing unit 12, and includes a determination unit 23 instead of the determination unit 13.
[0053] In one aspect of this disclosure, the acoustic information may be information relating to the language corresponding to the speech emitted by the speaker. In this specification, information relating to the language corresponding to the speech emitted by the speaker is referred to as language information.
[0054] Embodiment 2 illustrates a case where acoustic information is language-related information. The signal processing unit 22 derives language-related information by processing acoustic signals originating from speech emitted by the speaker.
[0055] Text data obtained by transcribing speech emitted by a speaker is an example of language-related information. Therefore, as an example, in Embodiment 2, the signal processing unit 22 may generate (derive) text data by performing speech recognition processing on acoustic signals derived from speech emitted by a speaker. The signal processing unit 22 may then output this text data as language-related information. For this reason, the signal processing unit 22 in Embodiment 2 may be called a speech recognition unit. The signal processing unit 22 is an example of a speech recognition unit that generates text data.
[0056] The signal processing unit 22 may generate text data using any STT (Speech-To-Text) algorithm. Therefore, as an example, any machine learning model having STT functionality can be used as the signal processing unit 22. Examples of machine learning models having STT functionality include Whisper and Speech Recognition.
[0057] In Embodiment 2, the determination unit 23 acquires language-related information from the signal processing unit 22. Based on the language-related information, the determination unit 23 determines whether the content indicated by the voice emitted by the speaker is conveyed to the receiver. Specifically, the determination unit 23 determines whether the content indicated by the voice is conveyed to the receiver by determining, based on the language-related information, whether the linguistic content indicated by the voice is semantically understandable.
[0058] In this specification, "semantic comprehensibility" means the possibility that the linguistic content represented by the sound corresponding to the acoustic signal can be interpreted as representing some meaning in natural language.
[0059] As an example, consider the case where text data as language-related information is supplied from the signal processing unit 22 to the determination unit 23. In this case, the determination unit 23 may determine whether or not the linguistic content indicated by the voice emitted by the speaker is semantically understandable by interpreting the meaning of the text data.
[0060] Therefore, the determination unit 23 may include any type of deep learning model capable of interpreting the meaning of text data. An example of such a deep learning model is a language model. For example, the determination unit 23 may include an LLM (Large Language Model).
[0061] It is preferable that the LLM has an architecture that is useful for semantic interpretation of text data. Examples of such architectures include Attention and Transformer.
[0062] Let us again refer to the example of reference numeral 210 in Figure 2 described above. In the example of reference numeral 210, the signal processing unit 22 generates the text data "Can you hear me?" by performing speech recognition processing on the acoustic signal AS, which does not have noise NZ added to it.
[0063] The determination unit 23 obtains the text data "Can you hear me?" from the signal processing unit 22 and interprets the meaning of the text data. In the example of reference numeral 210, there are no missing characters in the text data "Can you hear me?". Therefore, the determination unit 23 can interpret the text data "Can you hear me?" as representing some meaning in natural language. Specifically, the determination unit 23 interprets it as meaning an appeal, "Can you hear me?".
[0064] The determination unit 23 determines that the linguistic content expressed by the sound corresponding to the acoustic signal is semantically understandable if it interprets that the linguistic content expresses some meaning in natural language. In Embodiment 2, if the determination unit 23 determines that the linguistic content expressed by the sound corresponding to the acoustic signal is semantically understandable, it determines that the content expressed by the sound has been conveyed to the recipient. Therefore, in the example of reference numeral 210, the determination unit 23 determines that the content has been conveyed to the recipient.
[0065] Next, let us refer again to the example of reference numeral 220 in Figure 2. In the example of reference numeral 220, the part of the series of sounds "Can you hear me?" represented by the acoustic signal AS that corresponds to one syllable "su" is obscured by noise NZ. Therefore, in the example of reference numeral 220, the signal processing unit 22 generates the text data "kikoemaka" by performing speech recognition processing on the acoustic signal AS.
[0066] The determination unit 23 obtains the text data "kikoemaka" from the signal processing unit 22 and interprets the meaning of the text data. In the example of reference numeral 220, the determination unit 23 interprets the text data "kikoemaka" as text data with one character "su" missing from the text data "kikoemasu ka". From this, the determination unit 23 interprets the text data "kikoemaka" as meaning the call "Kikoemasu ka".
[0067] Therefore, in the example of reference numeral 220, the determination unit 23 determines that the linguistic content indicated by the sound corresponding to the acoustic signal is semantically understandable. For this reason, in the example of reference numeral 220, the determination unit 23 determines that the content indicated by the sound is conveyed to the recipient.
[0068] Next, let us refer again to the example of reference numeral 230 in Figure 2. In the example of reference numeral 230, the parts of the series of sounds "Kikoemasu ka" represented by the acoustic signal AS that correspond to the four syllables "ko", "e", "ma", and "su" are obscured by noise NZ. Therefore, in the example of reference numeral 230, the signal processing unit 22 generates the text data "kika" by performing speech recognition processing on the acoustic signal AS.
[0069] The determination unit 23 obtains the text data "kika" from the signal processing unit 22 and interprets the meaning of the text data. In the example of reference numeral 230, the determination unit 23 cannot interpret the text data "kikoemaka" as text data from "kikoemasu ka" with the four characters "ko," "e," "ma," and "su" missing. Therefore, the determination unit 23 cannot interpret the text data "kikoemaka" as representing any meaning in natural language.
[0070] Therefore, in the example of reference numeral 230, unlike the examples of reference numeral 210 and 220, the determination unit 23 determines that the linguistic content indicated by the sound corresponding to the acoustic signal is not semantically understandable. For this reason, in the example of reference numeral 230, the determination unit 23 determines that the content indicated by the sound has not been conveyed to the recipient. Then, in the example of reference numeral 230, the notification unit 14, upon receiving the determination result from the determination unit 23, notifies at least one of the speaker and the recipient that the content has not been conveyed to the recipient.
[0071] [Embodiment 3] Figure 5 shows an example configuration of the voice conversation system 3 in Embodiment 3. The voice conversation system 3 includes a speaker-side terminal device 30 instead of a speaker-side terminal device 10. The speaker-side terminal device 30 includes a signal processing unit 32 instead of a signal processing unit 12, and a determination unit 33 instead of a determination unit 13.
[0072] In this specification, information indicating the number of phonemes included in the voice uttered by a speaker in one utterance of the speaker is referred to as phoneme number information. Phoneme number information is another example of language-related information. As described above, the language-related information is not limited to the text data described in Embodiment 2. Embodiment 3 illustrates a case where the language-related information is phoneme number information.
[0073] In Embodiment 3, the signal processing unit 32 derives the phoneme number information by performing speech recognition processing on an acoustic signal derived from a voice uttered by a speaker. Then, the signal processing unit 32 outputs the phoneme number information as the language-related information. The signal processing unit 32 in Embodiment 3 is another example of a speech recognition unit. The signal processing unit 32 is an example of a speech recognition unit that derives the number of phonemes.
[0074] The number of characters representing one word in a certain language may vary depending on the type of characters representing the word in the language. As an example, the case of Japanese will be described. For example, the number of characters of the Japanese kanji word "犬" (meaning dog) is 1. On the other hand, the number of characters of the Japanese hiragana word "いぬ" (meaning dog) is 2. Therefore, Embodiment 3 focuses on phoneme number information. According to the phoneme number information, it is possible to eliminate the influence of the difference in the type of characters representing one word in a certain language (e.g., Japanese) and allow the determination unit 33 to perform determination processing.
[0075] As an example, the signal processing unit 32 uses any of the aforementioned STT (Speech-To-Text) algorithms to generate text data obtained by transcribing the voice uttered by the speaker in one utterance of the speaker. Next, the signal processing unit 32 performs morphological analysis on the text data. Next, the signal processing unit 32 collates the text data obtained after performing the morphological analysis with a given phoneme dictionary. As a result of the collation with the phoneme dictionary, the signal processing unit 32 derives an estimated value of the number of phonemes of the voice represented by the text data as the phoneme number information.
[0076] As another example, the signal processing unit 32 may include a machine learning model that estimates the number of phonemes contained in the speech represented by the text data. Therefore, the signal processing unit 32 can derive phoneme count information without requiring the text data after morphological analysis to be compared with a phoneme dictionary.
[0077] In Embodiment 3, the determination unit 33 acquires phoneme number information from the signal processing unit 32. Based on the phoneme number information, the determination unit 33 determines whether the linguistic content expressed by the speech emitted by the speaker is semantically understandable. As described above, the phoneme number information eliminates the influence of differences in the types of characters used to represent a word in a given language. Therefore, the determination unit 33 can perform the determination of whether the linguistic content expressed by the speech emitted by the speaker is semantically understandable, with the influence of differences in the types of characters used to represent a word in a given language already eliminated.
[0078] The number of phonemes contained in the speech emitted by a speaker in a single utterance may vary depending on the linguistic content expressed by that speech. For example, the determination unit 33 may derive an evaluation index as a value obtained by dividing the number of phonemes contained in the speech emitted by a speaker in a single utterance by the duration of that utterance. The determination unit 33 may derive this evaluation index from the phoneme count information.
[0079] As described with reference to Figure 2 above, as the amount of noise added to the acoustic signal increases, more syllables of the speech emitted by the speaker are obscured by that noise. Therefore, as the amount of noise increases, the number of missing characters in the text data increases.
[0080] Therefore, when the noise added to the acoustic signal is minor, the evaluation index described above takes a relatively large value. In this case, it can be expected that the linguistic content expressed by the speech emitted by the speaker is semantically understandable. On the other hand, as the noise added to the acoustic signal increases, the evaluation index decreases. Therefore, when the evaluation index is relatively small, it can be considered that the linguistic content expressed by the speech is not semantically understandable.
[0081] As an example, the determination unit 33 may determine whether the linguistic content expressed by the speech emitted by the speaker is semantically understandable by comparing the evaluation index with a predetermined threshold. The threshold in Embodiment 3 may be appropriately set by the speech conversation system 3 based on knowledge in the field of natural language processing.
[0082] Specifically, the determination unit 33 determines that if the evaluation index is above a threshold, the linguistic content expressed by the voice emitted by the speaker is semantically understandable. In other words, if the evaluation index is above a threshold, the determination unit 33 determines that the content expressed by the voice is being conveyed to the receiver.
[0083] On the other hand, if the evaluation index is below the threshold, the determination unit 33 determines that the linguistic content expressed by the voice emitted by the speaker is not semantically understandable. In other words, if the evaluation index is below the threshold, the determination unit 33 determines that the content expressed by the voice is not conveyed to the receiver.
[0084] Let us refer again to the examples in Figure 2 described above. In the example of reference numeral 210 in Figure 2, no noise is added to the acoustic signal. Therefore, the evaluation index in the example of reference numeral 210 takes a large value. From this, the determination unit 33 determines that in the example of reference numeral 210, the evaluation index is above the threshold. Therefore, in the example of reference numeral 210, the determination unit 33 determines that the content indicated by the voice emitted by the speaker is being conveyed to the receiver.
[0085] In the example of reference numeral 220 in Figure 2, noise is added to the acoustic signal, but this noise is minor. Therefore, the evaluation index in the example of reference numeral 220 does not decrease significantly compared to the evaluation index in the example of reference numeral 210. Consequently, in the example of reference numeral 220, the determination unit 33 determines that the evaluation index is above the threshold. Therefore, in the example of reference numeral 220, the determination unit 33 determines that the content indicated by the voice emitted by the speaker is being conveyed to the receiver.
[0086] In the example of reference numeral 230 in Figure 2, a significant amount of noise is added to the acoustic signal. As a result, the evaluation index in the example of reference numeral 230 is significantly lower than the evaluation index in the example of reference numeral 220. Therefore, unlike the examples of reference numeral 210 and 220, in the example of reference numeral 230, the determination unit 33 determines that the evaluation index is below the threshold.
[0087] Therefore, in the example of reference numeral 230, unlike the examples of reference numeral 210 and 220, the determination unit 33 determines that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver. Then, in the example of reference numeral 230, the notification unit 14, upon receiving the determination result of the determination unit 33, notifies at least one of the speaker and the receiver that the content has not been conveyed to the receiver.
[0088] [Embodiment 4] Figure 6 shows an example configuration of the voice conversation system 4 in Embodiment 4. The voice conversation system 4 includes a speaker-side terminal device 40 instead of a speaker-side terminal device 10. The speaker-side terminal device 40 includes a signal processing unit 42 instead of a signal processing unit 12, and a determination unit 43 instead of a determination unit 13.
[0089] Acoustic-related information according to one aspect of this disclosure may be a feature quantity (acoustic feature quantity) that indicates the properties of an acoustic signal derived from an acoustic signal originating from speech emitted by a speaker. In Embodiment 4, the signal processing unit 42 derives an acoustic feature quantity as acoustic-related information by processing an acoustic signal originating from speech emitted by a speaker.
[0090] One example of an acoustic feature is the Mel-Frequency Cepstrum Coefficient (MFCC). Therefore, the signal processing unit 42 may derive the MFCC by processing the acoustic signal. Another example of an acoustic feature is the filter bank feature. Therefore, the signal processing unit 42 may derive the filter bank feature by processing the acoustic signal.
[0091] In Embodiment 4, the determination unit 43 acquires acoustic features as acoustic-related information from the signal processing unit 42. Based on the acoustic features, the determination unit 43 determines whether or not the content indicated by the voice emitted by the speaker is conveyed to the receiver. Therefore, the determination unit 43 may include an acoustic model. In this specification, an acoustic model means a statistical model that has learned the correspondence between acoustic features and phonemes.
[0092] For example, an acoustic model can be a deep learning model. Examples of deep learning models used as acoustic models include DNN (Deep Neural Network), CNN (Convolutional Neural Network), and RNN (Recurrent Neural Network).
[0093] However, it should be noted that the acoustic model relating to one aspect of this disclosure does not necessarily have to be a deep learning model. An example of an acoustic model that is not a deep learning model is the Hidden Markov Model (HMM).
[0094] The determination unit 43 uses an acoustic model to estimate which phoneme corresponds to each time frame of the acoustic signal, based on acoustic features. Phonemes are vowels (e.g., "a" and "o") and consonants (e.g., "k" and "m"). The determination unit 43 then connects the estimated phonemes to generate a string. For example, if the estimated phonemes are in the order of "a", "k", and "a", the determination unit 43 connects these phonemes to generate the string "aka". Next, the determination unit 43 generates the Japanese hiragana string "aka" from the string "aka", for example, according to the rules of Japanese romanization. The Japanese hiragana string "aka" means "red" in Japanese. From this, if the determination unit 43 is able to generate a string of a certain language (e.g., Japanese) from the phonemes estimated by the acoustic model, it may determine that the content indicated by the speech emitted by the speaker has been conveyed to the receiver.
[0095] On the other hand, there are cases where a sequence of phonemes estimated by an acoustic model cannot be converted into a string. For example, if the sequence of phonemes estimated by the acoustic model contains consecutive consonants, it may not be possible to convert the sequence of phonemes into a string. For instance, if the estimated phonemes are in the order of "w", "k", and "m", the determination unit 43 can connect these phonemes to generate the string "wkm", but it cannot convert this string into a Japanese string. Therefore, if the determination unit 43 is unable to generate a string of a certain language (e.g., Japanese) from the estimated sequence of phonemes, it may determine that the content indicated by the speech emitted by the speaker has not been conveyed to the receiver.
[0096] As another example, the determination unit 43 may determine whether a string generated from phonemes estimated by the acoustic model exists as a language by comparing it with a predetermined dictionary. This language may be Japanese or any other language. If the determination unit 43 determines that the string generated from phonemes estimated by the acoustic model exists as a language, it may determine that the content indicated by the voice emitted by the speaker has been conveyed to the receiver. As an example, consider the case where the determination unit 43 generates the string "aka" or "aka" from phonemes estimated by the acoustic model. In this case, the determination unit 43 determines that the string exists as a language by comparing it with a predetermined dictionary. Therefore, the determination unit 43 determines that the content indicated by the voice emitted by the speaker has been conveyed to the receiver.
[0097] On the other hand, if the determination unit 43 determines that the content indicated by the voice emitted by the speaker has not been conveyed to the listener, the determination unit 43 may determine that the string generated from the phonemes estimated by the acoustic model is not a string that exists as a language. As an example, consider the case where the determination unit 43 generates the string "wkm" from the phonemes estimated by the acoustic model. In this case, the determination unit 43 determines that the string is not a string that exists as a language by comparing it with a predetermined dictionary. Therefore, the determination unit 43 determines that the content indicated by the voice emitted by the speaker has not been conveyed to the listener.
[0098] As described above, the determination unit 43 may determine whether the content indicated by the voice emitted by the speaker is being conveyed to the receiver by determining whether the sequence of phonemes estimated by the acoustic model can be converted into a string of characters. Alternatively, the determination unit 43 may determine whether the content indicated by the voice emitted by the speaker is being conveyed to the receiver by determining whether the string of characters generated from the sequence of phonemes estimated by the acoustic model is a string of characters that exists as a language.
[0099] Furthermore, human speech may include acoustic information corresponding to sounds other than words (e.g., coughs, breathing sounds, and hesitations). Typically, this non-verbal acoustic information is difficult to convert into appropriate strings. Therefore, it is not always appropriate to immediately conclude that the content of the speech is not being conveyed to the listener simply because the speech contains phonemes that cannot be converted into strings.
[0100] As another example, the determination unit 43 may calculate the ratio of the amount of phonemes that were converted into a string to the amount of phonemes estimated by the acoustic model. If the determination unit 43 is equal to or greater than a predetermined threshold, it may determine that the content indicated by the voice emitted by the speaker has been conveyed to the receiver. On the other hand, if the determination unit 43 is less than a predetermined threshold, it may determine that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver.
[0101] [Modification] (1) A speech conversation system according to one aspect of the present disclosure may derive intermediate expressions useful for semantic analysis of linguistic content indicated by speech emitted by a speaker as language-related information.
[0102] As an example, the signal processing unit of a speech conversation system according to one aspect of this disclosure may derive a word graph as an intermediate representation, which shows the relationships between multiple words contained in speech uttered by a speaker. For example, the signal processing unit may derive the word graph using an STT algorithm and a pronunciation dictionary.
[0103] In this case, the determination unit (e.g., LLM) of the speech conversation system according to one aspect of the present disclosure can determine whether the linguistic content indicated by the speech emitted by the speaker is semantically understandable, based on a word graph supplied by the signal processing unit. Therefore, the accuracy of the determination process in the determination unit is improved.
[0104] As another example, a signal processing unit of a speech conversation system according to one aspect of the present disclosure may derive a semantic network as an intermediate representation that represents the degree of linguistic semantic relevance between multiple words included in a word graph. For example, the signal processing unit may derive the semantic network using a word graph and a language dictionary.
[0105] In this case, a determination unit (e.g., LLM) of a speech conversation system according to one aspect of the present disclosure can determine whether the linguistic content indicated by the speech emitted by the speaker is semantically understandable, based on a semantic network supplied by the signal processing unit. The semantic network is language-related information that contains more information than a word graph. Therefore, the accuracy of the determination process in the determination unit is further improved.
[0106] (2) A speech conversation system according to one aspect of the present disclosure may include a deep learning model (multimodal deep learning model) that integrates an acoustic model, a pronunciation dictionary, and a language model. In this specification, such deep learning model will be referred to as an integrated deep learning model.
[0107] According to the integrated deep learning model, the linguistic content expressed by the speech emitted by the speaker can be derived end-to-end. Therefore, the integrated deep learning model makes it possible to integrate the signal processing unit and the determination unit of a speech conversation system according to one aspect of this disclosure.
[0108] As can be understood from the description of Embodiment 2, it is preferable that the integrated deep learning model has an architecture that is useful for semantic interpretation of text data. Therefore, the integrated deep learning model may include architectures such as Attention and Transformer.
[0109] The process of deriving the linguistic content expressed by speech from speech emitted by a speaker in an end-to-end manner generally requires significant computing resources. For this reason, in a speech conversation system according to one aspect of this disclosure, it is preferable that the integrated deep learning model be installed on a server with significant computing resources (e.g., Server SV in Figure 3).
[0110] However, if the speaker-side terminal device in a speech conversation system according to one aspect of this disclosure has large computing resources, an integrated deep learning model may be provided on the speaker-side terminal device.
[0111] [Example of implementation using software] The functions of the voice conversation systems 1 to 4 (hereinafter referred to as "devices" for convenience) can be realized by programs that cause a computer to function as the device, and by programs that cause a computer to function as each control block of the device (in particular, the signal processing units 12 to 42 and the determination units 13 to 43).
[0112] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.
[0113] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.
[0114] Some or all of the functions of each of the above control blocks can also be implemented by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in one aspect of this disclosure. In addition, it is also possible to implement the functions of each of the above control blocks by, for example, a quantum computer.
[0115] Each of the processes described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server).
[0116] [Summary] The voice conversation system according to Embodiment 1 of the present disclosure comprises: an acoustic signal acquisition unit that acquires an acoustic signal originating from a voice emitted by a speaker; a signal processing unit that derives acoustic-related information relating to the acoustic signal by processing the acoustic signal; and a determination unit that determines, based on the acoustic-related information, whether or not the content indicated by the voice emitted by the speaker has been conveyed to the receiver.
[0117] In the voice conversation system according to aspect 2 of the present disclosure, in aspect 1, the signal processing unit may derive language-related information relating to the language corresponding to the voice emitted by the speaker as sound-related information, and the determination unit may determine whether the content expressed by the voice is conveyed to the receiver by determining, based on the language-related information, whether the linguistic content expressed by the voice emitted by the speaker is semantically understandable.
[0118] In the voice conversation system according to aspect 3 of the present disclosure, in aspect 2, the signal processing unit may generate text data by performing speech recognition processing on the acoustic signal, output the generated text data as language-related information, and the determination unit may determine whether the linguistic content indicated by the voice emitted by the speaker is semantically understandable by interpreting the meaning of the text data.
[0119] In the voice conversation system according to aspect 4 of this disclosure, in aspect 3, the determination unit may include a Large Language Model (LLM).
[0120] In the voice conversation system according to aspect 5 of the present disclosure, in aspect 2, the signal processing unit may derive phoneme count information indicating the number of phonemes contained in the speech emitted by the speaker in a single utterance by the speaker by performing speech recognition processing on the acoustic signal, and may output the derived phoneme count information as language-related information, and the determination unit may determine, based on the phoneme count information, whether or not the linguistic content indicated by the speech emitted by the speaker is semantically understandable.
[0121] In the voice conversation system according to aspect 6 of the present disclosure, in aspect 5, the determination unit may derive an evaluation index from the phoneme count information, which is a value obtained by dividing the number of phonemes contained in the voice emitted by the speaker in a single utterance by the duration of the utterance, and if the evaluation index is less than a threshold, the determination unit may determine that the linguistic content indicated by the voice emitted by the speaker is not semantically understandable.
[0122] In the voice conversation system according to embodiment 7 of the present disclosure, in embodiment 1, the signal processing unit may derive acoustic feature quantities as acoustic-related information by processing the acoustic signal, and the determination unit may determine, based on the acoustic feature quantities, whether or not the content indicated by the voice emitted by the speaker has been conveyed to the receiver.
[0123] In the voice conversation system according to aspect 8 of the present disclosure, in aspect 7, the acoustic feature quantity may be a Mel-frequency cepstrum coefficient or a filter bank feature quantity.
[0124] The voice conversation system according to aspect 9 of the present disclosure may include a speaker-side terminal device having the signal processing unit and the determination unit in any one of aspects 1 to 8.
[0125] The voice conversation system according to embodiment 10 of the present disclosure may, in any one of embodiments 1 to 8, include a speaker-side terminal device having the signal processing unit and a determination unit located outside the speaker-side terminal device.
[0126] The voice conversation system according to embodiment 11 of the present disclosure may include a notification unit that performs a notification based on the determination result of the determination unit in any one of embodiments 1 to 10.
[0127] In the voice conversation system according to embodiment 12 of the present disclosure, in embodiment 11, if the determination unit outputs a determination result indicating that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver, the notification unit may notify at least one of the speaker and the receiver that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver.
[0128] (Cross-reference of related applications) This application claims priority to Japanese Patent Application No. 2025-050429, filed on 25 March 2025, and by reference thereto, all of its contents are included herein.
[0129] [Additional Notes] One aspect of this disclosure is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included within the technical scope of one aspect of this disclosure. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.
[0130] 1, 1V, 2, 3, 4 Voice conversation system 10, 10V, 20, 30, 40 Speaker-side terminal device 11 Acoustic signal acquisition unit 12, 22, 32, 42 Signal processing unit 13, 23, 33, 43 Judgment unit 14 Notification unit 90 Receiver-side terminal device 91 Acoustic signal playback unit AS Acoustic signal NZ Noise SP Speaker RP Receiver SV Server
Claims
1. A voice conversation system comprising: an acoustic signal acquisition unit that acquires an acoustic signal originating from a voice emitted by a speaker; a signal processing unit that processes the acoustic signal to derive acoustic-related information relating to the acoustic signal; and a determination unit that determines, based on the acoustic-related information, whether or not the content indicated by the voice emitted by the speaker has been conveyed to the receiver.
2. The voice conversation system according to claim 1, wherein the signal processing unit derives language-related information relating to the language corresponding to the voice emitted by the speaker as sound-related information, and the determination unit determines whether the linguistic content indicated by the voice emitted by the speaker is semantically understandable based on the language-related information, thereby determining whether the content indicated by the voice is conveyed to the receiver.
3. The voice conversation system according to claim 2, wherein the signal processing unit generates text data by performing speech recognition processing on the acoustic signal, outputs the generated text data as language-related information, and the determination unit determines whether or not the linguistic content indicated by the voice emitted by the speaker is semantically understandable by interpreting the meaning of the text data.
4. The voice conversation system according to claim 3, wherein the determination unit includes an LLM (Large Language Model).
5. The voice conversation system according to claim 2, wherein the signal processing unit derives phoneme count information indicating the number of phonemes contained in the speech emitted by the speaker in a single utterance by the speaker by performing speech recognition processing on the acoustic signal, outputs the derived phoneme count information as language-related information, and the determination unit determines, based on the phoneme count information, whether or not the linguistic content indicated by the speech emitted by the speaker is semantically understandable.
6. The voice conversation system according to claim 5, wherein the determination unit derives an evaluation index from the phoneme count information, which is a value obtained by dividing the number of phonemes contained in the sound emitted by the speaker in a single utterance by the duration of the utterance, and if the evaluation index is less than a threshold, the determination unit determines that the linguistic content indicated by the sound emitted by the speaker is not semantically understandable.
7. The voice conversation system according to claim 1, wherein the signal processing unit derives acoustic feature quantities as acoustic-related information by processing the acoustic signal, and the determination unit determines, based on the acoustic feature quantities, whether or not the content indicated by the voice emitted by the speaker has been conveyed to the receiver.
8. The speech conversation system according to claim 7, wherein the acoustic feature is a Mel-frequency cepstrum coefficient or a filter bank feature.
9. The voice conversation system according to any one of claims 1 to 8, comprising a speaker-side terminal device having the signal processing unit and the determination unit.
10. A voice conversation system according to any one of claims 1 to 8, comprising: a speaker-side terminal device having the signal processing unit; and a determination unit located outside the speaker-side terminal device.
11. The voice conversation system according to any one of claims 1 to 10, further comprising a notification unit that performs notification based on the determination result of the determination unit.
12. The voice conversation system according to claim 11, wherein if the determination unit outputs a determination result indicating that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver, the notification unit notifies at least one of the speaker and the receiver that the content indicated by the voice emitted by the speaker has not been conveyed to the receiver.