Two-way communication system, method, and program
The bidirectional communication system adjusts speech rates based on emotion, normal rate comparison, or time to synchronize conversation pace, addressing the issue of inappropriate speech speeds in conventional systems, improving conversation flow and efficiency.
Patent Information
- Application Number
- PCT/JP2025/002164
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-31
AI Technical Summary
Conventional bidirectional voice communication systems fail to adjust speech rates appropriately, leading to issues where individuals speaking quickly or slowly may not synchronize effectively, impacting conversation flow and efficiency.
A bidirectional communication system that adjusts speech rates by changing the collected voice's speed based on user emotion, normal speech rate comparison, or remaining time, using automatic response devices or user instructions, to synchronize conversation pace.
The system ensures participants unconsciously adjust their speech rate to match the adjusted pace, enhancing conversation synchronization and efficiency.
Smart Images

Figure JP2025002164_31072025_PF_FP_ABST
Abstract
Description
Two-way communication system, method, and program
[0001] The present disclosure relates to a two-way communication system, method, and program.
[0002] Conventionally, there has been known a technology for two-way audio communication between a plurality of terminals (for example, video conferencing, etc.). For example, Patent Literature 1 describes a neck-worn terminal equipped with a microphone and a speaker wirelessly communicating with another neck-worn terminal.
[0003] Patent No. 6786139
[0004] However, there are cases where the speaking speed of the users of each device is not appropriate. For example, you may want to prompt a person who is speaking too quickly to slow down, or a person who is speaking too slowly to speed up.
[0005] The present disclosure aims to make a speaker speak slower or faster in two-way audio communication between multiple terminals.
[0006] A two-way communication system according to a first aspect of the present disclosure is a two-way communication system for two-way communication of voice between multiple terminals, which changes the speech rate of voice collected by at least one terminal or voice generated by an automatic answering device, and transmits the voice with the changed speech rate to a terminal other than the terminal from which the voice was collected.
[0007] According to a first aspect of the present disclosure, a person who hears speech whose speech rate has been changed is lured into speaking slower without realizing it, or is lured into speaking faster without realizing it, due to the speech rate of the speech.
[0008] A second aspect of the present disclosure is a two-way communication system according to the first aspect, wherein the speech rate of the voice collected by at least one of the terminals or the voice generated by the automatic answering device is changed in response to an instruction from a user of one of the terminals or based on a judgment of the two-way communication system.
[0009] According to the second aspect of the present disclosure, any user can instruct a change in speech rate, or the speech rate can be changed automatically.
[0010] A third aspect of the present disclosure is a two-way communication system according to the second aspect, wherein the speech rate of a voice collected at a terminal other than the terminal of the user or a voice generated by the automatic answering device is changed based on the results of an analysis of the emotions of the user of the terminal at which the voice is collected.
[0011] According to the third aspect of the present disclosure, it is possible to change the speech rate according to the emotion of the terminal user.
[0012] A fourth aspect of the present disclosure is a two-way communication system according to the third aspect, wherein when the emotion of the user of the terminal where the voice is collected is impatience, the speech rate of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is reduced.
[0013] According to the fourth aspect of the present disclosure, when a person is in a hurry and hears a voice with a slowed-down speaking speed, the person is lured by the slowed-down speaking speed of the voice and begins to speak slower without realizing it.
[0014] A fifth aspect of the present disclosure is a two-way communication system according to the second aspect, wherein the speech speed of the voice collected at a terminal other than the user's terminal or the voice generated by the automatic answering device is changed based on the difference between the normal speech speed of the user of the terminal where the voice is collected and the speech speed of the voice collected at the terminal where the voice is collected.
[0015] According to the fifth aspect of the present disclosure, it is possible to change the speech speed in accordance with the difference from the normal speech speed.
[0016] A sixth aspect of the present disclosure is a two-way communication system according to the fifth aspect, wherein, when the speech rate of the voice collected at the terminal collecting the voice is faster than the normal speech rate of the user of the terminal collecting the voice, the speech rate of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is reduced.
[0017] According to the sixth aspect of the present disclosure, when a person who speaks faster than normal hears a voice with a slowed-down speaking rate, that person is lured into speaking slower without realizing it by the slowed-down speaking rate of the voice.
[0018] A seventh aspect of the present disclosure is a two-way communication system according to the fifth aspect, wherein, when the speech rate of the voice collected at the terminal collecting the voice is slower than the normal speech rate of the user of the terminal collecting the voice, the speech rate of the voice collected at a terminal other than the terminal of the user or the voice generated by the automatic answering device is increased.
[0019] According to the seventh aspect of the present disclosure, when a person who speaks slower than normal hears speech with an increased speaking speed, that person is lured by the increased speaking speed of the speech and begins to speak faster without realizing it.
[0020] An eighth aspect of the present disclosure is a two-way communication system according to the second aspect, wherein the speech rate of the voice collected by the at least one terminal or the voice generated by the automatic answering device is changed based on the remaining time of the conference or work performed by the two-way voice communication.
[0021] According to the eighth aspect of the present disclosure, the speech rate can be changed according to the remaining time of a meeting or task.
[0022] A ninth aspect of the present disclosure is a two-way communication system according to the eighth aspect, wherein when the remaining time of the conference or work via the two-way voice communication falls below a threshold, the speech rate of the voice collected by all terminals and the voice generated by the automatic answering device is increased.
[0023] According to a ninth aspect of the present disclosure, by listening to the voice with the increased speaking speed, everyone is lured into speaking faster without realizing it, allowing the meeting or work to be completed within the remaining time.
[0024] A method according to a tenth aspect of the present disclosure is a method executed by a two-way communication system that performs two-way voice communication between multiple terminals, and includes the steps of: changing the speech rate of voice collected by at least one terminal or voice generated by the automatic answering device; and transmitting the voice with the changed speech rate to a terminal other than the terminal from which the voice was collected.
[0025] A program according to an eleventh aspect of the present disclosure causes a computer to execute the following steps: a procedure for changing the speech rate of speech collected by at least one of a plurality of terminals that perform two-way speech communication or speech generated by the automatic answering device; and a procedure for transmitting the speech with the changed speech rate to a terminal other than the terminal that collected the speech.
[0026] FIG. 1 is a diagram illustrating an overall configuration according to an embodiment of the present disclosure. FIG. 2 is an example of a wearable device according to an embodiment of the present disclosure. FIG. 3 is a diagram illustrating a hardware configuration of a wearable device according to an embodiment of the present disclosure. FIG. 4 is a diagram illustrating a hardware configuration of a two-way communication management device according to an embodiment of the present disclosure. FIG. 5 is a functional block diagram of a two-way communication system according to an embodiment of the present disclosure. FIG. 6 is a sequence diagram (in the case of a remote supporter terminal) of speech rate conversion processing based on emotion analysis according to an embodiment of the present disclosure. FIG. 7 is a sequence diagram (in the case of a remote supporter terminal) of speech rate conversion processing based on a comparison with a normal speech rate according to an embodiment of the present disclosure. FIG. 8 is a sequence diagram (in the case of a remote supporter terminal) of speech rate conversion processing based on the remaining time of a meeting or task according to an embodiment of the present disclosure. FIG. 9 is a sequence diagram (in the case of an automatic answering device) of speech rate conversion processing based on a comparison with a normal speech rate according to an embodiment of the present disclosure. FIG. 10 is a sequence diagram (in the case of an automatic answering device) of speech rate conversion processing based on the remaining time of a meeting or task according to an embodiment of the present disclosure. FIG. 11 is a diagram for explaining a speech section and a non-speech section according to an embodiment of the present disclosure. FIG. 12 is a diagram for explaining a method for lowering the speech rate of audio according to an embodiment of the present disclosure. FIG. 10 is a diagram illustrating a method for increasing the speech rate according to an embodiment of the present disclosure.
[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0028] <Explanation of Terms> In this specification, a "terminal" refers to a terminal for two-way communication of at least sound between multiple terminals. For example, a terminal is a wearable terminal equipped with a microphone function and a speaker function. Note that a terminal may be a personal computer, smartphone, tablet, etc. equipped with a microphone function and a speaker function, or a personal computer, smartphone, tablet, etc. connected to a microphone and a speaker. In this specification, "speech rate" refers to the speed of speech, and is, for example, the value obtained by dividing the number of moras in speech by the duration of the speech (for example, the number of moras per second, the number of moras per minute, etc.).
[0029] <Overview> In a two-way communication system according to an embodiment of the present disclosure, the speech rate of audio collected by at least one terminal is changed, and the audio with the changed speech rate is transmitted to a terminal other than the terminal from which the audio was collected. (Note that this includes cases where the device managing the two-way communication (two-way communication management device 20) converts the speech rate of the audio, where the terminal from which the audio was collected converts the speech rate of the audio, and where a terminal other than the terminal from which the audio was collected converts the speech rate of the audio.) When a speaker hears audio with a slowed speech rate, they are lured into speaking more slowly by the slowed speech rate, or when a speaker hears audio with an increased speech rate, they are lured into speaking more quickly by the increased speech rate. Note that speakers are known to synchronize with the speech rate of the other party (see, for example, "Komatsu Takanori and Morikawa Koji, 'Speech Rate Entrainment in Interactive Communication between Humans and Artificial Objects,' Research Report, Information Processing Society of Japan, 2004-ICS-137(10), October 29, 2004"). The speech rate of the voice generated by the automatic answering device may be changed.
[0030] <Overall Configuration> Fig. 1 is a diagram showing the overall configuration according to one embodiment of the present disclosure. The two-way communication system 1 includes wearable terminals 10A, 10B, and 10C (hereinafter collectively referred to as wearable terminals 10), a two-way communication management device 20, and at least one of a remote supporter terminal 30 and an automatic response device 40. Note that Fig. 1 illustrates a case where there are three wearable terminals 10 and one remote supporter terminal 30, but the number of terminals is not limited to this. Each of these will be described below.
[0031] <<Wearable Terminal>> The wearable terminal 10 is a terminal (computer) equipped with a microphone function and a speaker function, and used by a field worker (it is assumed that field worker 11A is wearing the wearable terminal 10A, field worker 11B is wearing the wearable terminal 10B, and field worker 11C is wearing the wearable terminal 10C. Hereinafter, field workers 11A, 11B, and 11C are collectively referred to as field workers 11). For example, the wearable terminal 10 is a neck-worn terminal. The wearable terminal 10 can send and receive data to and from the two-way communication management device 20 via any network.
[0032] <<Two-way Communication Management Device>> The two-way communication management device 20 is a device that manages at least two-way audio communication between multiple terminals (in the example of FIG. 1, the wearable terminals 10A, 10B, and 10C and the remote supporter terminal 30). The two-way communication management device 20 is composed of one or more computers.
[0033] <<Remote Supporter Terminal>> The remote supporter terminal 30 is a terminal used by the remote supporter 31. For example, the remote supporter 31 is a person who supports the work of the on-site worker 11 at a remote location away from the site where the on-site worker 11 is working. For example, the remote supporter terminal 30 is a personal computer. The remote supporter terminal 30 can send and receive data with the two-way communication management device 20 via any network.
[0034] <<Automatic Answering Device>> The automatic answering device 40 uses artificial intelligence (AI) or the like to perform speech recognition of the voice of the field worker 11 wearing the wearable device 10, and generates a voice that indicates a response (e.g., an answer to the question) to the content of the voice (e.g., a question). The automatic answering device 40 is composed of one or more computers. Note that the automatic answering device 40 and the two-way communication management device 20 may be implemented in a single device.
[0035] <Configuration of Wearable Terminal> Fig. 2 shows an example of a wearable terminal 10 according to an embodiment of the present disclosure. For example, the wearable terminal 10 is a terminal shaped like a neck strap as shown in Fig. 2. For example, the wearable terminal 10 includes a sound input unit (microphone) 105, a sound output unit (speaker) 106, and an earphone jack 100 as shown in Fig. 2. Furthermore, the wearable terminal 10 may include an operation unit 104 and an imaging unit (camera) 107 as shown in Fig. 2. Each unit will be described in detail with reference to Fig. 3.
[0036] <Hardware Configuration> An example of the hardware configuration of the wearable terminal 10 will be described below with reference to FIG. 3, and an example of the hardware configuration of the two-way communication management device 20 will be described with reference to FIG.
[0037] 3 is a diagram illustrating the hardware configuration of a wearable terminal 10 according to an embodiment of the present disclosure. The wearable terminal 10 can include a control unit (processor) 101, a storage unit (memory) 102, a communication unit 103, an operation unit 104, a sound input unit (microphone) 105, a sound output unit (speaker) 106, an imaging unit (camera) 107, and various sensors 108. Each of these will be described below.
[0038] The control unit (processor) 101 is a processor that controls the wearable terminal 10. For example, the control unit (processor) 101 is a central processing unit (CPU), a graphics processing unit (GPU), or the like.
[0039] The storage unit (memory) 102 is a memory that stores any data.
[0040] The communication unit 103 connects to any network and communicates with other computers (such as the bidirectional communication management device 20).
[0041] The operation unit 104 is a button or the like that allows the field worker 11 to input instructions to the wearable terminal 10 .
[0042] The sound input unit (microphone) 105 collects sound.
[0043] The sound output unit (speaker) 106 outputs sound (for example, sound collected by another terminal and acquired via the two-way communication management device 20).
[0044] The imaging unit (camera) 107 captures still images and moving images.
[0045] The various sensors 108 are one or more arbitrary sensors such as an acceleration sensor, a GPS (Global Positioning System), and the like.
[0046] 4 is a diagram showing the hardware configuration of the two-way communication management device 20 according to an embodiment of the present disclosure. The same applies to the remote supporter terminal 30 and the automatic response device 40. The two-way communication management device 20 can include a control unit (processor) 201, a storage unit (memory) 202, and a communication unit 203. Each of these will be described below.
[0047] The control unit (processor) 201 is a processor that controls the two-way communication management device 20. For example, the control unit (processor) 201 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0048] The storage unit (memory) 202 is a memory that stores any data.
[0049] The communication unit 203 connects to any network and communicates with other computers (such as the wearable terminal 10 and the remote supporter terminal 30).
[0050] FIG. 5 is a functional block diagram of a two-way communication system 1 according to an embodiment of the present disclosure.
[0051] For example, the wearable terminal 10 includes a sound receiving unit 111, a sound collection unit 112, and a sound transmission unit 113. For example, the wearable terminal 10 functions as the sound receiving unit 111, the sound collection unit 112, and the sound transmission unit 113 by executing a program.
[0052] For example, the two-way communication management device 20 includes a voice receiving unit 211, a determination unit 212, a speech speed control unit 213, and a voice transmitting unit 214. For example, the two-way communication management device 20 functions as the voice receiving unit 211, the determination unit 212, the speech speed control unit 213, and the voice transmitting unit 214 by executing a program.
[0053] For example, the remote supporter terminal 30 includes a voice receiving unit 311, a voice collection unit 312, and a voice transmission unit 313. For example, the remote supporter terminal 30 functions as the voice receiving unit 311, the voice collection unit 312, and the voice transmission unit 313 by executing a program.
[0054] For example, the automatic answering device 40 includes a voice receiving unit 411, a voice recognition unit 412, a voice generation unit 413, and a voice transmission unit 414. For example, the automatic answering device 40 functions as the voice receiving unit 411, the voice recognition unit 412, the voice generation unit 413, and the voice transmission unit 414 by executing a program.
[0055] Note that Figure 5 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0056] [Wearable Terminal] The audio receiving unit 111 receives audio from the two-way communication management device 20 and reproduces it through a speaker or the like.
[0057] The sound collection unit 112 acquires the sound (that is, the sound collected by the microphone) uttered by the person wearing the wearable device 10 (that is, the field worker 11).
[0058] The voice transmitting unit 113 transmits the voice to the two-way communication management device 20 .
[0059] [Two-Way Communication Management Device] The voice receiving unit 211 receives voice from the wearable terminal 10 , the remote supporter terminal 30 , and the automatic response device 40 .
[0060] The determining unit 212 determines whether or not to change the speech rate of the voice.
[0061] The speech speed conversion unit 213 changes the speech speed of the voice (specifically, decreases or increases the speech speed of the voice).
[0062] As described above, the wearable terminal 10 and the remote supporter terminal 30 may convert the speech speed of the voice (that is, the wearable terminal 10 and the remote supporter terminal 30 may be provided with the determination unit 212 and the speech speed conversion unit 213).
[0063] The voice transmitting unit 214 transmits the voice to the wearable terminal 10 , the remote supporter terminal 30 , and the automatic response device 40 .
[0064] [Remote Supporter Terminal] The voice receiving unit 311 receives voice from the two-way communication management device 20 and reproduces it on a speaker or the like.
[0065] The sound collection unit 312 acquires the sound (that is, the sound collected by the microphone) uttered by the user of the remote supporter terminal 30 (that is, the remote supporter 31).
[0066] The voice transmitting unit 313 transmits the voice to the two-way communication management device 20 .
[0067] [Automatic Answering Device] The voice receiving unit 411 receives voice from the two-way communication management device 20 .
[0068] The voice recognition unit 412 performs voice recognition on the voice received by the voice receiving unit 411. For example, the voice recognition unit 412 uses AI or the like to perform voice recognition on the voice received by the voice receiving unit 411 (for example, the voice of the field worker 11 wearing the wearable terminal 10 asking a question) and converts the voice into text.
[0069] The voice generation unit 413 generates voice indicating a response to the result of voice recognition by the voice recognition unit 412. For example, the voice generation unit 413 uses AI (e.g., generation AI that generates an answer to a question) or the like to generate text of a response (e.g., an answer to the question) to text (e.g., a question) that is the result of voice recognition, and generates voice of the text indicating the response (i.e., performs voice synthesis).
[0070] The voice transmitting unit 414 transmits the voice generated by the voice generating unit 413 to the bidirectional communication management device 20 .
[0071] Here, an example of changing the speech rate will be described. For example, the speech rate can be changed in response to an instruction from a user of one of the terminals (e.g., the remote supporter terminal 30) or based on the judgment of the two-way communication management device 20. As examples of changing the speech rate based on the judgment of the two-way communication management device 20, "speech rate conversion based on the user's emotions," "speech rate conversion based on a comparison with the normal speech rate," and "speech rate conversion based on the remaining time of a meeting or task" will be described.
[0072] [Speech Rate Conversion Based on User's Emotions] For example, the speech rate can be changed based on the results of analyzing the emotions of the person wearing the wearable device 10 (i.e., the field worker 11). For example, the determination unit 212 of the two-way communication management device 20 analyzes the emotions of the person (i.e., the field worker 11) who has spoken the voice (i.e., the voice collected by the microphone) of the person wearing the wearable device 10, and when the emotion is a specific emotion (e.g., impatience), determines that the speech rate of the voice of the person in the conversation (i.e., the remote supporter 31) or the speech rate of the voice generated by the automatic response device 40 should be changed (e.g., lower the speech rate).
[0073] [Speech Speed Conversion Based on Comparison with Normal Speech Speed] For example, the speech speed can be changed based on the difference between the speech speed of the voice collected by the wearable device 10 and the normal speech speed of the person wearing the wearable device 10 (i.e., the field worker 11). For example, when the speech speed of the person wearing the wearable device 10 (i.e., the field worker 11) is faster than normal, the determination unit 212 of the two-way communication management device 20 determines that the speech speed of the voice of the other party (i.e., the remote supporter 31) or the speech speed of the voice generated by the automatic answering device 40 should be changed (e.g., lower the speech speed). Furthermore, for example, when the speech speed of the person wearing the wearable device 10 (i.e., the field worker 11) is slower than normal, the determination unit 212 of the two-way communication management device 20 determines that the speech speed of the other party (i.e., the remote supporter 31) or the speech speed of the voice generated by the automatic answering device 40 should be changed (e.g., increase the speech speed).
[0074] [Converting speech speed based on remaining time of a meeting or task] For example, the speech speed can be changed based on the remaining time of a meeting or task via two-way voice communication. For example, when the determination unit 212 of the two-way communication management device 20 detects that the remaining time of the meeting or task is equal to or less than a threshold, it determines that the speech speed of the voices of all terminals and the speech speed of the voice generated by the automatic answering device 40 should be changed (for example, by increasing the speech speed).
[0075] <Processing Method> A speech speed conversion processing method will be described below with reference to Figures 6 to 8. Note that the speech speed of the speaker may be controlled (for example, by gradually slowing down or gradually speeding up) by repeating speech speed conversion.
[0076] 6 is a sequence diagram (for a remote supporter terminal) of a speech rate conversion process based on emotion analysis according to an embodiment of the present disclosure. While FIG. 6 illustrates an example in which speech rate conversion is performed on speech collected by the remote supporter terminal 30 based on the results of analyzing the emotion of the user of the wearable device 10 (i.e., the field worker 11), speech rate conversion may also be performed on speech collected by a terminal other than a given terminal based on the results of analyzing the emotion of the user of the given terminal.
[0077] Note that Figure 6 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0078] In step 101 (S101), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0079] In step 102 (S102), the wearable terminal 10 transmits the voice collected in S101 to the two-way communication management device 20.
[0080] In step 103 (S103), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S102.
[0081] In step 104 (S104), the two-way communication management device 20 transmits the voice received in S103 to the remote supporter terminal 30 (that is, transmits the voice to the remote supporter 31 without speech speed conversion).
[0082] In step 105 (S105), the remote supporter terminal 30 receives the voice transmitted by the two-way communication management device 20 in S104.
[0083] In step 106 (S106), the remote supporter terminal 30 outputs the voice received in S105 (for example, by playing it on a speaker).
[0084] In step 107 (S107), the two-way communication management device 20 analyzes the emotion of the person who made the voice (i.e., the field worker 11) based on the voice received in S103. For example, the two-way communication management device 20 can analyze emotion based on the frequency, volume, quality, and speaking speed of the voice. Note that the two-way communication management device 20 may also analyze the emotion of the person wearing the wearable terminal 10 based on something other than voice (for example, based on biometric information, etc.). The two-way communication management device 20 saves the results of the emotion analysis.
[0085] S107 may be executed simultaneously with S104 to S106 or before S104 to S106.
[0086] In step 108 (S108), the remote supporter terminal 30 collects the voice uttered by the user of the remote supporter terminal 30 (that is, the remote supporter 31).
[0087] In step 109 (S109), the remote supporter terminal 30 transmits the voice collected in S108 to the two-way communication management device 20.
[0088] In step 110 (S110), the two-way communication management device 20 receives the voice transmitted by the remote supporter terminal 30 in S109.
[0089] In step 111 (S111), the two-way communication management device 20 converts the speech rate of the voice received in S110. Specifically, the two-way communication management device 20 reads the result of the emotion analysis in S107 and converts the speech rate of the voice based on the result. For example, if the emotion is a specific emotion, the two-way communication management device 20 changes the speech rate of the voice (for example, if the emotion is impatience (e.g., panic, losing one's usual composure, being flustered, being confused, etc.) then the speech rate is reduced).
[0090] In step 112 (S112), the two-way communication management device 20 transmits the voice whose speech rate has been changed (for example, slowed down) in S111 to the wearable terminal 10.
[0091] In step 113 (S113), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S112.
[0092] In step 114 (S114), the wearable device 10 outputs (e.g., plays on a speaker) the speech received in S113 with the speech rate changed (e.g., slowed down). Therefore, for example, the on-site worker 11 who has been speaking quickly due to impatience may be influenced by the speech of the remote supporter 31 whose speech rate has been slowed down and start speaking more slowly.
[0093] 7 is a sequence diagram (for a remote supporter terminal) of a speech speed conversion process based on a comparison with a normal speech speed according to an embodiment of the present disclosure. In FIG. 7, an example of converting the speech speed of the voice collected by the remote supporter terminal 30 based on the result of a comparison with the normal speech speed of the user of the wearable terminal 10 (i.e., the field worker 11) is described. However, speech speed conversion may be performed on a terminal other than a given terminal based on the result of a comparison with the normal speech speed of the user of the given terminal.
[0094] Note that Figure 7 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0095] In step 200 (S200), the two-way communication management device 20 stores information on the normal speaking speed of the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0096] In step 201 (S201), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0097] In step 202 ( S202 ), the wearable terminal 10 transmits the voice collected in S201 to the two-way communication management device 20 .
[0098] In step 203 (S203), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S202.
[0099] In step 204 (S204), the two-way communication management device 20 transmits the voice received in S203 to the remote supporter terminal 30 (that is, transmits the voice to the remote supporter 31 without speech speed conversion).
[0100] In step 205 (S205), the remote supporter terminal 30 receives the voice transmitted by the two-way communication management device 20 in S204.
[0101] In step 206 (S206), the remote supporter terminal 30 outputs the voice received in S205 (for example, plays it on a speaker).
[0102] In step 207 (S207), the two-way communication management device 20 compares the speech rate of the voice received in S203 with the normal speech rate (the normal speech rate saved in S200) of the person who made the voice (i.e., the field worker 11). The two-way communication management device 20 saves the result of the comparison with the normal speech rate.
[0103] S207 may be executed simultaneously with S204 to S206 or before S204 to S206.
[0104] In step 208 (S208), the remote supporter terminal 30 collects the voice uttered by the user of the remote supporter terminal 30 (that is, the remote supporter 31).
[0105] In step 209 (S209), the remote supporter terminal 30 transmits the voice collected in S208 to the two-way communication management device 20.
[0106] In step 210 (S210), the two-way communication management device 20 receives the voice transmitted by the remote supporter terminal 30 in S209.
[0107] In step 211 (S211), the two-way communication management device 20 converts the speech rate of the voice received in S210. Specifically, the two-way communication management device 20 reads the result of the comparison with the normal speech rate in S207 and converts the speech rate of the voice based on the result. For example, if the speech rate is faster than normal, the two-way communication management device 20 lowers the speech rate of the voice. Also, for example, if the speech rate is slower than normal, the two-way communication management device 20 raises the speech rate of the voice.
[0108] In step 212 (S212), the two-way communication management device 20 transmits the voice whose speech rate has been changed in S211 to the wearable terminal 10.
[0109] In step 213 (S213), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S212.
[0110] In step 214 (S214), the wearable device 10 outputs the speech received in S213 with the speech rate changed (for example, by playing it through a speaker). Therefore, for example, the on-site worker 11 who has been speaking faster than usual will be influenced by the speech of the remote supporter 31 whose speech rate has been slowed down and will also start speaking more slowly. Also, for example, the on-site worker 11 who has been speaking more slowly than usual will be influenced by the speech of the remote supporter 31 whose speech rate has been increased and will also start speaking more quickly.
[0111] FIG. 8 is a sequence diagram (in the case of a remote supporter terminal) of a speech speed conversion process based on the remaining time of a meeting or task according to an embodiment of the present disclosure.
[0112] Note that Figure 8 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0113] In step 301 (S301), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0114] In step 302 (S302), the wearable terminal 10 transmits the voice collected in S301 to the two-way communication management device 20.
[0115] In step 303 (S303), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S302.
[0116] In step 304 (S304), the two-way communication management device 20 detects the remaining time of the conference or work by two-way voice communication. If speech speed conversion is necessary (for example, if the remaining time is equal to or less than the threshold), the process proceeds to step 305. If speech speed conversion is not necessary (for example, if the remaining time is not equal to or less than the threshold), the process proceeds to step 306.
[0117] In step 305 (S305), the two-way communication management device 20 changes the speech rate of the voice received in S303 (for example, increases the speech rate).
[0118] In step 306 (S306), the two-way communication management device 20 transmits to the remote supporter terminal 30 the voice whose speech rate has been changed (for example, increased) in S305 (if speech rate conversion is required), or the voice received in S303 (if speech rate conversion is not required).
[0119] In step 307 (S307), the remote supporter terminal 30 receives the voice transmitted by the two-way communication management device 20 in S306.
[0120] In step 308 (S308), the remote supporter terminal 30 outputs (e.g., plays on a speaker) the speech received in S307 with the speech rate changed (e.g., increased). As a result, the remote supporter 31, influenced by the speech of the on-site worker 11 with the increased speech rate, also begins to speak faster (as a result, the meeting or work can be completed within the remaining time).
[0121] In step 309 (S309), the remote supporter terminal 30 collects the voice uttered by the person using the remote supporter terminal 30 (that is, the field worker 11).
[0122] In step 310 (S310), the remote supporter terminal 30 transmits the voice collected in S309 to the two-way communication management device 20.
[0123] In step 311 (S311), the two-way communication management device 20 receives the voice transmitted by the remote supporter terminal 30 in S310.
[0124] In step 312 (S312), the two-way communication management device 20 detects the remaining time of the conference or work by two-way voice communication. If speech speed conversion is necessary (for example, if the remaining time is equal to or less than the threshold), the process proceeds to step 313. If speech speed conversion is not necessary (for example, if the remaining time is not equal to or less than the threshold), the process proceeds to step 314.
[0125] In step 313 (S313), the two-way communication management device 20 changes the speech rate of the voice received in S311 (for example, increases the speech rate).
[0126] In step 314 (S314), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed (for example, increased) in S313 (if speech rate conversion is required), or the voice received in S311 (if speech rate conversion is not required).
[0127] In step 315 (S315), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S314.
[0128] In step 316 (S316), the wearable device 10 outputs (e.g., plays on a speaker) the speech received in S315 with the speech rate changed (e.g., increased). As a result, the on-site worker 11, influenced by the speech of the remote supporter 31 with the increased speech rate, also begins to speak faster (as a result, the meeting or work can be completed within the remaining time).
[0129] A method for converting the speech speed of speech generated by an automatic answering device will be described below with reference to Figures 9 to 11. Note that the speech speed of the speaker may be controlled (for example, by gradually slowing down or gradually speeding up) by repeating the speech speed conversion.
[0130] For example, speech speed conversion is used when the automatic answering device 40 generates an answer to a question from the field worker 11 and causes the wearable device 10 of the field worker 11 to output a voice indicating the answer. Also, speech speed conversion is used when the automatic answering device 40 analyzes a real-time video of a work operation captured by the wearable device 10 of the field worker 11 and causes the wearable device 10 of the field worker 11 to output work instructions based on the results of the analysis by voice to the wearable device 10 of the field worker 11. Also, speech speed conversion is used when the automatic answering device 40 causes the wearable device 10 of the field worker 11 to output work procedures described in a work procedure manual by voice (to provide voice guidance).
[0131] 9 is a sequence diagram (in the case of an automatic answering device) of a speech rate conversion process based on emotion analysis according to an embodiment of the present disclosure. In FIG. 9, an example of converting the speech rate of a voice generated by the automatic answering device 40 based on the result of analyzing the emotion of the user of the wearable device 10 (i.e., the field worker 11) is described.
[0132] Note that Figure 9 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0133] In step 1001 (S1001), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0134] In step 1002 (S1002), the wearable terminal 10 transmits the voice collected in S1001 to the two-way communication management device 20.
[0135] In step 1003 (S1003), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S1002.
[0136] In step 1004 (S1004), the two-way communication management device 20 transmits the voice received in S1003 to the automatic answering device 40 (that is, transmits the voice to the automatic answering device 40 without speech speed conversion).
[0137] In step 1005 (S1005), the automatic answering device 40 receives the voice transmitted by the two-way communication management device 20 in S1004.
[0138] In step 1006 (S1006), the automatic answering device 40 performs speech recognition on the speech received in S1005. For example, the automatic answering device 40 uses AI or the like to perform speech recognition on the speech received in S1005 (e.g., the speech of the field worker 11 wearing the wearable device 10 asking a question) and converts the speech into text.
[0139] In step 1007 (S1007), the two-way communication management device 20 analyzes the emotion of the person who made the voice (i.e., the field worker 11) based on the voice received in S1003. For example, the two-way communication management device 20 can analyze emotion based on the frequency, volume, quality, and speaking speed of the voice. Note that the two-way communication management device 20 may also analyze the emotion of the person wearing the wearable terminal 10 based on something other than voice (for example, based on biometric information, etc.). The two-way communication management device 20 saves the results of the emotion analysis.
[0140] Note that S1007 may be executed simultaneously with S1004 to S1006 or before S1004 to S1006.
[0141] In step 1008 (S1008), the automatic response device 40 generates a voice indicating a response to the result of the voice recognition in S1006. For example, the automatic response device 40 uses AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to the text (e.g., a question) that is the result of the voice recognition, and generates a voice of the text indicating the response (i.e., performs voice synthesis).
[0142] In step 1009 (S1009), the automatic answering device 40 transmits the voice generated in S1008 to the two-way communication management device 20.
[0143] In step 1010 (S1010), the two-way communication management device 20 receives the voice transmitted by the automatic answering device 40 in S1009.
[0144] In step 1011 (S1011), the two-way communication management device 20 converts the speech rate of the voice received in S1010. Specifically, the two-way communication management device 20 reads the result of the emotion analysis in S1007 and converts the speech rate of the voice based on the result. For example, if the emotion is a specific emotion, the two-way communication management device 20 changes the speech rate of the voice (for example, if the emotion is impatience (e.g., panic, losing one's usual composure, being flustered, being confused, etc.), the speech rate is reduced).
[0145] In step 1012 (S1012), the two-way communication management device 20 transmits the voice whose speech rate has been changed (for example, slowed down) in S1011 to the wearable terminal 10.
[0146] In step 1013 (S1013), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S1012.
[0147] In step 1014 (S1014), the wearable device 10 outputs (for example, plays on a speaker) the speech received in S1013 with the speech rate changed (for example, slowed down). Therefore, for example, the field worker 11 who has been speaking quickly due to impatience may be persuaded by the speech of the automatic answering device 40 with the speech rate slowed down and start speaking more slowly.
[0148] 10 is a sequence diagram (in the case of an automatic answering device) of a speech speed conversion process based on a comparison with a normal speech speed according to an embodiment of the present disclosure. Fig. 10 illustrates an example of converting the speech speed of the voice generated by the automatic answering device 40 based on the result of a comparison with the normal speech speed of the user of the wearable device 10 (i.e., the field worker 11).
[0149] Note that Figure 10 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0150] In step 2000 (S2000), the two-way communication management device 20 stores information on the normal speaking speed of the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0151] In step 2001 (S2001), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0152] In step 2002 (S2002), the wearable terminal 10 transmits the voice collected in S2001 to the two-way communication management device 20.
[0153] In step 2003 (S2003), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S2002.
[0154] In step 2004 (S2004), the two-way communication management device 20 transmits the voice received in S2003 to the automatic answering device 40 (that is, transmits the voice to the automatic answering device 40 without speech speed conversion).
[0155] In step 2005 (S2005), the automatic answering device 40 receives the voice transmitted by the two-way communication management device 20 in S2004.
[0156] In step 2006 (S2006), the automatic answering device 40 performs speech recognition on the speech received in S2005. For example, the automatic answering device 40 uses AI or the like to perform speech recognition on the speech received in S2005 (e.g., the speech of the field worker 11 wearing the wearable device 10 asking a question) and converts the speech into text.
[0157] In step 2007 (S2007), the two-way communication management device 20 compares the speech rate of the voice received in S2003 with the normal speech rate (the normal speech rate saved in S2000) of the person who made the voice (i.e., the field worker 11). The two-way communication management device 20 saves the result of the comparison with the normal speech rate.
[0158] Note that S2007 may be executed simultaneously with S2004 to S2006 or before S2004 to S2006.
[0159] In step 2008 (S2008), the automatic answering device 40 generates a voice that indicates a response to the result of the voice recognition in S2006. For example, the automatic answering device 40 uses AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to the text (e.g., a question) that is the result of the voice recognition, and generates a voice of the text that indicates the response (i.e., performs voice synthesis).
[0160] In step 2009 (S2009), the automatic answering device 40 transmits the voice generated in S2008 to the two-way communication management device 20.
[0161] In step 2010 (S2010), the two-way communication management device 20 receives the voice transmitted by the automatic answering device 40 in S2009.
[0162] In step 2011 (S2011), the two-way communication management device 20 converts the speech rate of the voice received in S2010. Specifically, the two-way communication management device 20 reads the result of the comparison with the normal speech rate in S2007 and converts the speech rate of the voice based on the result. For example, if the speech rate is faster than normal, the two-way communication management device 20 lowers the speech rate of the voice. Also, for example, if the speech rate is slower than normal, the two-way communication management device 20 raises the speech rate of the voice.
[0163] In step 2012 (S2012), the two-way communication management device 20 transmits the voice whose speech rate has been changed in S2011 to the wearable terminal 10.
[0164] In step 2013 (S2013), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S2012.
[0165] In step 2014 (S2014), the wearable device 10 outputs the voice received in S2013 with the speech rate changed (for example, by playing it through a speaker). Therefore, for example, a field worker 11 who has been speaking faster than usual will be persuaded by the voice of the automatic answering device 40 with the speech rate slowed down and will also start speaking more slowly. Also, for example, a field worker 11 who has been speaking more slowly than usual will be persuaded by the voice of the automatic answering device 40 with the speech rate increased and will also start speaking more quickly.
[0166] FIG. 11 is a sequence diagram of a speech speed conversion process based on the remaining time of a meeting or task according to an embodiment of the present disclosure (in the case of an automatic answering device).
[0167] Note that Figure 11 explains the case where the two-way communication management device 20 converts the speech speed of the voice, but the terminal that collects the voice may also convert the speech speed of the voice, or a terminal other than the terminal that collects the voice may convert the speech speed of the voice.
[0168] In step 3001 (S3001), the wearable terminal 10 collects the voice uttered by the person wearing the wearable terminal 10 (i.e., the field worker 11).
[0169] In step 3002 (S3002), the wearable terminal 10 transmits the voice collected in S3001 to the two-way communication management device 20.
[0170] In step 3003 (S3003), the two-way communication management device 20 receives the voice transmitted by the wearable terminal 10 in S3002.
[0171] In step 3004 (S3004), the two-way communication management device 20 transmits the voice received in S3003 to the automatic answering device 40.
[0172] In step 3005 (S3005), the automatic answering device 40 receives the voice transmitted by the two-way communication management device 20 in S3004.
[0173] In step 3006 (S3006), the automatic answering device 40 performs speech recognition on the speech received in S3005. For example, the automatic answering device 40 uses AI or the like to perform speech recognition on the speech received in S3005 (e.g., the speech of the field worker 11 wearing the wearable device 10 asking a question) and converts the speech into text.
[0174] In step 3007 (S3007), the automatic answering device 40 generates a voice indicating a response to the result of the voice recognition in S3006. For example, the automatic answering device 40 uses AI (e.g., a generation AI that generates an answer to a question) or the like to generate a text of a response (e.g., an answer to the question) to the text (e.g., a question) that is the result of the voice recognition, and generates a voice of the text indicating the response (i.e., performs voice synthesis).
[0175] In step 3008 (S3008), the automatic answering device 40 transmits the voice generated in S3007 to the two-way communication management device 20.
[0176] In step 3009 (S3009), the two-way communication management device 20 receives the voice transmitted by the automatic answering device 40 in S3008.
[0177] In step 3010 (S3010), the two-way communication management device 20 detects the remaining time of the conference or work by two-way voice communication. If speech speed conversion is necessary (for example, if the remaining time is equal to or less than the threshold), the process proceeds to step 3011. If speech speed conversion is not necessary (for example, if the remaining time is not equal to or less than the threshold), the process proceeds to step 3012.
[0178] In step 3011 (S3011), the two-way communication management device 20 changes the speech rate of the voice received in S3009 (for example, increases the speech rate).
[0179] In step 3012 (S3012), the two-way communication management device 20 transmits to the wearable terminal 10 the voice whose speech rate has been changed (for example, increased) in S3011 (if speech rate conversion is required), or the voice received in S309 (if speech rate conversion is not required).
[0180] In step 3013 (S3013), the wearable terminal 10 receives the voice transmitted by the two-way communication management device 20 in S3012.
[0181] In step 3014 (S3014), the wearable device 10 outputs (for example, plays on a speaker) the speech received in S3013 with the speech rate changed (for example, increased). As a result, the field worker 11, influenced by the speech rate increased from the automatic answering device 40, also begins to speak faster (as a result, the meeting or work can be completed within the remaining time).
[0182] Hereinafter, the speech section and the non-speech section will be described with reference to FIG. 12, the method for lowering the speech rate will be described with reference to FIG. 13, and the method for increasing the speech rate will be described with reference to FIG.
[0183] 12 to 14 indicate that the user of the user's terminal is speaking (i.e., the sound is being collected by the user's terminal) during the "elapsed time (milliseconds)." The speech includes a section where speech (voice) is detected (a speech section, also called a voice section) and a section where speech (voice) is not detected (a non-speech section, also called a non-voice section).
[0184] The shaded areas of "Communication & Processing" in Figures 12 to 14 indicate that communication and processing are taking place from one's own terminal to a terminal other than one's own terminal (the other terminal) during the "Elapsed Time (milliseconds)" (for example, the two-way communication management device 20 is performing the processing).
[0185] The shaded areas of "Playback by other party" in FIGS. 12 to 14 indicate that the voice of the user of the user's own terminal is being played back at the other party's terminal during the "Elapsed time (milliseconds)".
[0186] Figure 12 shows a case where both the speech section and the non-speech section of "own utterance" are sent to the other party's terminal without extracting the speech section from "own utterance" (in Figure 12, both the speech section "aaa" and the non-speech section "aa" are sent to the other party's terminal and played back on the other party's terminal).
[0187] In one embodiment of the present disclosure, a speech section may be detected using VAD (Voice Activity Detection), and only the speech section may be sent to the other terminal. Hereinafter, with reference to Fig. 13, a case will be described in which the speech rate of only the speech section is lowered and sent to the other terminal (i.e., the non-speech section is not sent to the other terminal), and with reference to Fig. 14, a case will be described in which the speech rate of only the speech section is increased and sent to the other terminal (i.e., the non-speech section is not sent to the other terminal).
[0188] 13 is a diagram illustrating a method for slowing down the speech rate according to an embodiment of the present disclosure. As shown in FIG. 13, when a speech section ends and a non-speech section begins, the speech rate of the speech section is slowed down and sent to the other party's terminal for playback (i.e., the non-speech section is not sent to the other party's terminal). Note that the process of slowing down the speech rate of the speech section may be performed by the two-way communication management device 20 in the "Communication & Processing" section of FIG. 13, by the user's own terminal, or by the other party's terminal.
[0189] 14 is a diagram illustrating a method for increasing the speech rate according to an embodiment of the present disclosure. As shown in FIG. 14, when a speech section ends and a non-speech section begins, the speech rate of the speech section is increased and sent to the other party's terminal for playback (i.e., the non-speech section is not sent to the other party's terminal). The process of increasing the speech rate of the speech section may be performed by the two-way communication management device 20 in "Communication & Processing" in FIG. 14, by the user's own terminal, or by the other party's terminal.
[0190] However, if the speech rate is increased too much, the speech of two speech sections will be played back consecutively without any non-speech section, making it difficult for the listener to hear. Therefore, a buffer (a buffer of a predetermined minimum length) may be provided between the speech of the two speech sections (i.e., a non-speech section may be provided).
[0191] Although the embodiments have been described above, it will be understood that various changes in form and details can be made without departing from the spirit and scope of the claims.
[0192] This international application claims priority to Japanese Patent Application No. 2024-010382, filed on January 26, 2024, the entire contents of which are incorporated herein by reference.
[0193] 1 Two-way communication system 10A Wearable terminal 10B Wearable terminal 10C Wearable terminal 11A Field worker 11B Field worker 11C Field worker 20 Two-way communication management device 30 Remote supporter terminal 31 Remote supporter 40 Automatic response device 101 Control unit (processor) 102 Storage unit (memory) 103 Communication unit 104 Operation unit 105 Sound input unit (microphone) 106 Sound output unit (speaker) 107 Imaging unit (camera) 108 Various sensors 100 Earphone jack 111 Audio receiving unit 112 Audio collection unit 113 Audio transmitting unit 201 Control unit (processor) 202 Storage unit (memory) 203 Communication unit 211 Audio receiving unit 212 Determination unit 213 Speech speed conversion unit 214 Voice transmission unit 311 Voice reception unit 312 Voice collection unit 313 Voice transmission unit 411 Voice reception unit 412 Voice recognition unit 413 Voice generation unit 414 Voice transmission unit
Claims
1. A two-way communication system for two-way communication of voice between a plurality of terminals, which changes the speaking speed of the voice collected by at least one terminal or the voice generated by an automatic response device, and transmits the voice with the changed speaking speed to a terminal other than the terminal where the voice was collected.
2. The two-way communication system according to claim 1, wherein the speaking speed of the voice collected by at least one terminal or the voice generated by the automatic response device is changed in response to an instruction from a user of any terminal or based on a determination of the two-way communication system.
3. The two-way communication system according to claim 2, wherein the speaking speed of the voice collected by a terminal other than the user's terminal or the voice generated by the automatic response device is changed based on the result of analyzing the emotion of the user of the terminal where the voice was collected.
4. The two-way communication system according to claim 3, wherein when the emotion of the user of the terminal where the voice was collected is anxiety, the speaking speed of the voice collected by a terminal other than the user's terminal or the voice generated by the automatic response device is lowered.
5. The two-way communication system according to claim 2, wherein the speaking speed of the voice collected by a terminal other than the user's terminal or the voice generated by the automatic response device is changed based on the difference between the normal speaking speed of the user of the terminal where the voice was collected and the speaking speed of the voice collected by the terminal where the voice was collected.
6. The two-way communication system according to claim 5, wherein when the speaking speed of the voice collected by the terminal where the voice was collected is faster than the normal speaking speed of the user of the terminal where the voice was collected, the speaking speed of the voice collected by a terminal other than the user's terminal or the voice generated by the automatic response device is lowered.
7. The two-way communication system according to claim 5, wherein when the speaking speed of the voice collected by the terminal where the voice was collected is slower than the normal speaking speed of the user of the terminal where the voice was collected, the speaking speed of the voice collected by a terminal other than the user's terminal or the voice generated by the automatic response device is increased.
8. The two-way communication system according to claim 2, wherein the speaking speed of the voice collected by at least one terminal or the voice generated by the automatic response device is changed based on the remaining time of the meeting or work by the two-way communication of the voice.
9. The two-way communication system according to claim 8, wherein when the remaining time of the conference or work by the two-way communication of the voice becomes equal to or less than a threshold value, the speech rate of the voice collected by all terminals and the voice generated by the automatic response device is increased.
10. A method executed by a two-way communication system for two-way communication of voice between a plurality of terminals, the method including: changing a speech rate of voice collected by at least one terminal or voice generated by an automatic response device; and transmitting the voice with the changed speech rate to a terminal other than the terminal where the voice was collected.
11. A program for causing a computer to execute a procedure for changing a speech rate of voice collected by at least one terminal among a plurality of terminals that perform two-way communication of voice or voice generated by an automatic response device, and a procedure for transmitting the voice with the changed speech rate to a terminal other than the terminal where the voice was collected.
Citation Information
Patent Citations
Voice interactive device, input voice optimizing method in the device and input voice optimizing processing program in the device
JP2003150194A
Apparatus and method for speech processing, recording medium, and program
JP2004258290A
Voice communication apparatus and voice communication system
JP2008032933A
Automatic voice recognition / voice conversion system
JP2014095753A
Aviation control voice communication device and voice processing method
JP2014228691A